Labeling · 7 min read
Representative training data for computer vision on a site camera
A PPE and fall model that scored well at the pilot site missed the crew at the second. The labeled set had one kind of worker in it. The site had every kind.
Summary
This post follows a PPE and fall protection model from a pilot site to a second one, where the crew looked nothing like the workers in the labeled set and the recall fell. It argues for a labeled set that holds workers of every build, skin tone and mobility aid, for per-subgroup recall instead of one average, and for treating a new site's crew as the test the pilot never ran. It is for safety teams rolling a site camera model out beyond the first site.
Ayman Quadir · Head of Product · Oct 3, 2026

Scaffold tower beside site containers with workers, ladders and plank decks boxed, from a customer site camera
The model watching the pilot site's cameras did what it was built for. Every worker on the scaffold got a box, every hard hat and vest was found, and when someone stepped onto a deck without their fall protection clipped on, the frame reached the site manager's phone inside the minute. Six weeks of that, and in May the model was rolled out to the second site, forty miles away, same contractor, same rule.
On the second site it missed people. A worker in a wheelchair at the ground-floor station, checking deliveries, was boxed as a piece of equipment when he was boxed at all. Two of the steel crew, both heavier built than anyone at the pilot, were found on some frames and lost on others. A worker with a darker skin tone under a white hard hat, seen from the pole camera in the afternoon glare, was missed often enough that the vest rule on him never fired.
The model had not changed. The labeled set it was trained on held one kind of worker, the kind the pilot site employed, and the second site employed every kind.
This is not face recognition and the gap is still there
It is worth being clear about what the camera is doing, because the phrase people reach for is the wrong one. A site safety model is not doing face recognition. It does not identify anyone, and the site safety and PPE question is whether the person in the frame is wearing what the zone requires, not who they are. Nobody is in a database. Nothing about a face is matched, on the pilot site or the second one in May.
And the model still has to find the person before it can check the gear, and finding a person is a learned appearance. If every person in the training set walked, stood at a certain height and build, and wore skin of a narrow range of tones under the site's lighting, then that is what "person" means to the model. A person outside that range is a shape it has weak evidence for. The problem that face recognition systems became known for, a training set that stood in for a narrower world than the one deployed into, arrives on a site camera by the same road, with no face involved.
One average hides the subgroup that was missed
The pilot's recall was a single number, and a good one. Recall on a set where nine in ten workers look alike is recall on those nine, and the tenth can be missed on every frame without moving the figure much. The site safety page says compliance is judged on the worst moment, and a model that is right most of the time can still miss every violation. The same is true across people: a model that is right on most workers can miss every violation by the ones it was not trained to see.
So the held-out set is scored per subgroup, and the subgroups are written down before scoring. Build, height, skin tone, mobility aid, the kinds of gear and clothing the sites' crews wear. On the second site's footage from the pole camera, labeled by a person over a Monday, the recall for workers using a mobility aid was the lowest row in the table, and the recall for the heaviest build the second lowest. The average across all rows was close to the pilot's number.
My own view is that a safety model should never be reported as one recall figure, and that the person signing it off should ask for the table. A row that is empty, because the held-out set had nobody in that subgroup, is as much of a finding as a low one.
The labeled set has to hold every worker the sites will employ
The fix started with the frames. The second site's footage was labeled, every worker boxed by a person. The frames with the workers the model had missed were the first batch: the wheelchair user at the delivery station, the steel crew on the scaffold, the afternoon glare frames from the pole camera. A few hundred frames from a crew the pilot had never had, and the class "person" in the training set grew to mean more people.
A labeled set does not become representative by accident and it does not become representative by volume. A thousand more frames from the pilot site would have added a thousand more of the same workers. The footage lesson makes this point in general, that the unit of a dataset is distinct conditions, and on a site camera the crew is a condition, as much as the weather or the mount height. The set has to be checked against the people the sites will actually employ, and the gaps filled on purpose from footage of those people.
There is a version of this that the guideline has to carry too. A worker using crutches on a site is a person, boxed as a person, gear checked as gear, and a labeler who has never been told that will box the crutches and hesitate over the rest. The ruling, with a picture, went into the guideline the same week in June.
A new site's crew is the test the pilot never ran
The drift catalog files a new site as covariate shift arriving all at once, the same task and the same definitions on a different input distribution, with the new site's detection profile differing from the pilot's from day one rather than degrading into it. The usual reading of that is mounting heights and equipment vintages, the things the June rollout checklist already had on it. The crew is in the same list. The second site was a different distribution of people, and the model met it on the first morning.
Which means the way to catch it is the way the catalog gives for every new site: label a window from the new site before it goes live, score it per subgroup, and fold it in. A model that has seen three sites' crews generalises to a fourth far better than one that has seen a single crew very thoroughly. Compare each site against its own history and its own table rather than against a fleet average, because the good sites carry the bad one and the average hides it.
Somebody at the contractor now keeps the subgroup table for every site on one page, and it is the first thing checked before a site's cameras are switched from pilot to live.
The corrections keep the set honest after rollout
After the second site's frames were folded in and the model retrained in July, the delivery station worker was boxed on every frame and the vest rule on the steel crew fired when it should. The frames the model doubts on either site come back to a person, and when they cluster on one kind of worker, that cluster is the next gap. The correction on each is a label that says this person is a person, and the new version learns it with no downtime between the old one and the new.
The rule at the third site is that nothing goes live until the table has a number in every row. It is a slower rollout than the second site had. It is also the only one the contractor's safety lead will sign.
See it on your own footage.
Start with your footageMore in Labeling

Labeling · 7 min read
Aerial dataset augmentation for drone frames where there is no up
A tower seen straight down has no top or bottom, so rotations and both flips are safe. Scale for altitude, brightness for sun, move every box with its pixels.
Esdras Ntuyenabo · Oct 3, 2026

Labeling · 6 min read
A collaborative data annotation workflow run as a pipeline
Batch the bottling line's frames by camera and shift, assign so nothing is boxed twice, attach the guideline, review every label, train on the approved set.
Rajiya Sultana · Oct 3, 2026

Labeling · 7 min read
Dataset health check for computer vision, what to look at before anything trains
A scratch dataset where every scratch sits in the centre of the frame will train a model that looks in the centre. Five counts to read before the first epoch.
Rajiya Sultana · Oct 3, 2026