Computer vision · 6 min read
Transfer learning from a pretrained checkpoint, why a few hundred frames are enough when you start warm
A pretrained backbone already knows edges and shapes. A few hundred reviewed frames from a site camera teach the head what a guardrail is. Retrains begin warm.
Summary
This post explains why a detector for guardrails and hard hats on a scaffold camera can be trained from a few hundred reviewed frames rather than tens of thousands, by starting from a pretrained checkpoint. It covers the split between backbone and head, when to freeze and when to fine-tune, how overfitting shows on held-out frames, and how the retrain keeps the warm start. It is for teams who have been told they need more footage than they have.
Rob Hickey · Chief AI Officer · Oct 3, 2026

Scaffold with two workers, guardrails, ladder and plank decks boxed, from a customer site camera
The site camera on the tower crane's mast looks down on a scaffold on the Marsh Lane site that goes up two lifts a week. The safety manager wants a model that flags a deck with a guardrail missing and a person on it without a hard hat. There are about four hundred labeled frames, all from that camera, checked by the site's own safety officer. Someone has told the team that four hundred frames is nowhere near enough to train a detector.
They are right about training from nothing, and wrong about what happens in practice, because nobody trains from nothing. The model starts from a checkpoint that has already seen an enormous general set of images, and the four hundred frames teach it the part it does not know.
A pretrained backbone already knows edges, textures and shapes
A detector has two parts. The backbone takes the frame and turns it into features: edges at every angle, textures, corners, the shapes that edges make, and at deeper layers the arrangements of those shapes that tend to mean an object. A backbone pretrained on a large general dataset, COCO or something like it, has learned all of that without ever seeing a scaffold, because edges and textures are the same in every picture.
That is what a checkpoint is: the backbone's weights, saved after that general training, available as a starting point. Starting from it means the model already has a way of seeing, and the site's frames only have to teach it what to look for. The how much footage lesson counts in situations rather than frames; with a warm start, the situations the frames need to cover are the scaffold's own, and the general visual world is already covered.
The head is the part that learns the guardrail
The second part of the detector, the head, takes the backbone's features and turns them into boxes and classes. It is small compared with the backbone, and it is the part that is specific to the job: guardrail, missing guardrail, hard hat, person. On a fresh checkpoint the head knows nothing about any of these, and the four hundred frames are for it.
A head can learn a new class from a few hundred examples because it is not learning what a guardrail looks like from pixels. It is learning which combinations of features the backbone already produces tend to sit on a guardrail. That is a much smaller thing to learn, and the labeling with Lexi doc describes the labeling for it: a box on each rail and each hat, checked by a person, the class list short and stable.
The scaffold's safety netting changes colour when the contractor changes, from green to orange the week the second firm arrived. The backbone's features handled that without a single new label, because a rail behind orange netting has the same edges as a rail behind green.
Freezing the backbone keeps what was learned and fine-tuning bends it
With a warm start there is a choice about how much of the model to train. Freezing the backbone leaves its weights as the checkpoint had them and trains only the head. The model learns quickly, cannot forget what the checkpoint knew, and cannot adapt the backbone to anything specific about the scaffold camera's view. Fine-tuning unfreezes some or all of the backbone and lets those weights move too, at a lower learning rate, so the features themselves bend toward the site's frames.
My own view is to freeze first and unfreeze only when the validation curve says so. A frozen backbone with four hundred frames gives a model that works on the Marsh Lane scaffold by the afternoon, and the curve on the held-out frames tells you whether it has stopped improving with the head alone. If it has, unfreeze the last few layers of the backbone and train again. If it has not, there is nothing to gain by giving the model more freedom than the frames can constrain.
A few hundred reviewed frames from the site camera is enough to adapt
The count that matters is instances per class, from the camera the model will watch, which at Marsh Lane is the one on the mast. Four hundred frames from the mast camera with several rails and a few people in each is a few thousand rail instances and several hundred hard hats, at the mast camera's angle and distance, in the site's own light. That is enough for a head to learn, and it is more useful than forty thousand frames of scaffolds from other sites at other angles.
What the four hundred frames do not contain will show up as misses. A deck at the top lift in fog, a worker in a white hard hat against a white sky, a rail hidden behind a hoist. Those are the situations the site adds over the following weeks, and they arrive through review rather than through a second collection effort.
Overfitting shows on the held-out frames before it shows on site
The risk with a small dataset is overfitting: the head memorises the four hundred frames instead of learning what a guardrail looks like. It shows as a training score that keeps rising while the score on held-out frames flattens and then falls. On a warm start the risk is smaller than on a cold one, because the backbone's features are fixed or barely moving and the head has less room to memorise. It is still there, and the held-out set is the only place it is visible.
The held-out frames have to be honest for that to work. Frames from the same minute as the training frames will score well however overfit the head is. Frames from a different day, or better still from the camera's newest position after the crane climbed on the Friday, are the ones that show the fall. Training stops at the epoch where the held-out score peaks, and the version from that epoch is the one that goes on the camera.
The retrain keeps the warm start rather than starting cold
Once the model is on the mast camera at Marsh Lane, frames it is unsure of come back to the safety officer, the corrections retrain it, and the new version replaces the old one with no downtime. The retrain starts from the current version, not from the original checkpoint and not from nothing. The retraining without starting over lesson argues that every correction is a deposit toward the next version, and a warm start is what makes the deposit compound: each version begins where the last one ended, with the site's own frames on top.
The dataset keeps the original four hundred frames alongside the corrections, so that a version trained on the foggy top-lift frames does not forget the clear ones. By the fourth retrain the dataset was around nine hundred frames, still all from the one camera, and the model had not once been trained from cold.
See it on your own footage.
Start with your footageMore in Computer vision

Computer vision · 6 min read
What CLIP is and why typing a description finds the frame
A shared space for pictures and words lets an operator type a sentence and get the matching frames from a recorder. A way to look, and the boxes come after.
Rob Hickey · Oct 3, 2026

Computer vision · 6 min read
An end-to-end object detection workflow starts with decisions, and the first frame is labeled last
Tag, box, mask or keypoints. A class list with occlusion and size rules. A metric and an acceptance sentence. Then, and only then, the first camera's frames.
Ayman Quadir · Oct 3, 2026

Computer vision · 7 min read
Image recognition AI, explained on the cameras a site already owns
A tag, a box, a mask or keypoints, on a substation camera, a line camera and an aisle camera. How a model gets there, and what its score does not promise.
Esdras Ntuyenabo · Oct 3, 2026