Computer vision · 6 min read
Transfer learning starting points, why a small nearby dataset beats a huge general one
Four checkpoints for the same hard hat model, and the small one trained on people won. Pick a base model by what it has seen, then feed it the site's frames.
Summary
This post compares four starting checkpoints for one hard hat model on a site camera, from random weights to a huge general set to a small set of people, and explains why the small nearby one reached the best result in the fewest epochs. It draws the rule for choosing a base model for a plant or yard camera by the domain the checkpoint has seen rather than its size, and argues that the site's own reviewed frames outweigh either. It is for engineers picking a starting point for a custom detector.
Esdras Ntuyenabo · Engineer · Oct 4, 2026

Scaffold with two workers, guardrails, ladder and plank decks boxed, from a customer site camera
The site camera over the scaffold yard sees the crew go up the ladders every morning at 7 am, and the hard hat model that watches it had to be trained from somewhere. The engineer had four places to start: random weights, a checkpoint trained on a huge general set of everyday objects, a checkpoint trained on an equally large set of overhead imagery, and a small checkpoint trained on a few thousand frames of people. The obvious bet was the huge general set. The small set of people won, and by a margin that changed how the team picks base models.
Transfer learning reuses filters that already know edges and textures
A detector trained from random weights has to learn everything from the site's frames, including the things every image shares: edges and corners, gradients, and the textures of cloth and metal. A checkpoint trained on any large set already has those in its early layers, and fine-tuning from it means the 7 am ladder frames only have to teach the later layers what a hard hat on a head looks like. That is transfer learning, and it is the reason a few hundred checked frames are enough for a usable model where tens of thousands would be needed from nothing.
The question is which checkpoint. All four carry the early layers. What differs is what the later layers already know.
Four starting points for the same hard hat model
Random weights did worst by a distance on the 7 am ladder frames, as expected, and were still improving slowly when the others had finished. The huge general set did well: a strong start and a good result. The overhead imagery set, as large as the general one, did worse than the general one, because its later layers had learned fields and roofs from above and had to unlearn them before they could learn a person from the side.
The small set of people, a fraction of the size of the general set, reached the best result and reached it first. Its later layers already knew heads, shoulders and the way a person stands on a deck, and a hard hat is a small change to a head the model already finds. The site's frames refined something rather than building it.
The scaffolders' hats are colour-coded by trade on that site, and the model learned the colour before it learned the hat. That is a frame problem rather than a checkpoint problem, and it is fixed by frames.
Fewer epochs from a nearby checkpoint beat more from a distant one
An epoch is one pass through the training set, and the number of epochs is the training budget. The nearby checkpoint reached its best score on the yard's frames in a small number of epochs and then flattened. The distant one kept climbing slowly for much longer and flattened lower. Running the distant checkpoint on through April did not close the gap, because the gap was in what the later layers had to unlearn, and a longer run on the site's few hundred frames cannot supply what the pretraining left out.
The practical reading is that the nearby checkpoint is cheaper as well as better. A shorter run is less compute per version, and a model that reaches its best quickly is a model that can be retrained often.
Domain similarity is what the checkpoint carries and size comes second
The size of a pretraining set is the number people quote because it is the number printed on the model card. Similarity is what matters, and it is rarely printed. The later layers of a checkpoint are specific to what it was trained on. A checkpoint trained on people brings features for people, as it did on the scaffold yard in March; one trained on overhead fields brings features for fields, however many of them it saw.
So the audit before fine-tuning is of the pretraining domain rather than its size. What was in the frames, at what angle, at what scale, in what light. A checkpoint whose frames look like the site camera's is the one to start from, and a checkpoint that is enormous and unlike the site is a slower route to a worse model.
Pick the base for a plant or yard camera by what it has seen
For a yard camera like the one over the scaffold at 7 am, a checkpoint trained on people or on PPE from another site is the starting point, and the general set is the fallback. For a conveyor camera with parts on it, a checkpoint from an industrial set is the first choice if one exists, and the general set otherwise, since it contains machinery and tools. For a thermal camera there is usually nothing nearby, and the honest plan is more of the site's own frames rather than a search for a checkpoint that does not exist.
Whichever base is chosen, the lesson on how much footage you need applies unchanged: the frames that teach the model are the ones that sample distinct conditions on that camera, and a good checkpoint reduces how many are needed without changing which ones.
The site's own reviewed frames matter more than either
The four checkpoints differed at the start. After a few hundred checked frames from the scaffold yard and a month of corrections, the difference had narrowed, and after a season it was hard to tell which base a version had started from. The frames from the camera, labeled to a written rule and checked by a person, are what the model is made of in the end.
LexData takes the hard hat model through its whole life. You type what to look for, Lexi puts a box on every frame, and a person checks each label before anything trains on it. The model then watches the site camera, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime.
The lesson on retraining without starting over makes the point that corrections on frames the model got wrong are worth many times a label on a random frame, and each new version starts from the last one rather than from any of the four checkpoints. My own view is that the choice of base model matters for the first version and for almost nothing after it, and that a team arguing about checkpoints for a week would do better to spend the week reviewing the doubted frames.
See it on your own footage.
Start with your footageMore in Computer vision

Computer vision · 7 min read
Five computer vision applications in production, on the cameras a site already owns
Detect, count, inspect, track and read are the five jobs a fixed camera can be given, and each is a different question with a different label behind it.
Ayman Quadir · Oct 4, 2026

Computer vision · 6 min read
Computer vision projects worth building on the cameras you already have
A plant, a utility, a grower, a store and a warehouse each have a project that starts on an existing camera, produces a decision, and has someone to act on it.
Ayman Quadir · Oct 4, 2026

Computer vision · 6 min read
How to choose an object detection model architecture for a camera on your own site
Start from the decision the plant has to make, then the box beside the recorder the model has to fit, and only then the family. The benchmark comes last.
Andreas Ohrvall · Oct 4, 2026