Labeling · 6 min read
Labeling outdoor surveillance footage, the person forty pixels tall on the perimeter camera
Tight boxes on distant people and vehicles, occlusion labeled instead of skipped, classes the intrusion rule can use, and the rain and night frames kept in.
Summary
This post is about labeling frames from a fixed perimeter camera at a remote oil and gas site, where a person is forty pixels tall and the camera sees rain, dust and night for most of the year. It covers tight boxes on small objects, labeling occlusion rather than skipping it, class names the intrusion rule can use, and why the weather frames belong in the set. It is for security and operations teams building their first perimeter model.
Sheikh Srijon · GTM Lead · Sep 28, 2026

Thermal camera on a fence line at night, two people at the fence and vehicles boxed, from a customer site camera
Camera 3, on the north fence of the well pad, looks along the fence line towards the access track. A person walking the track is about forty pixels tall at the far end, a pickup is perhaps twice that, and for most of the year the frame is dusty, wet or dark. The site wants a model that says when someone is inside the fence who should not be, and the labeled set is where that model's judgement comes from.
Labeling this camera is a different job from labeling a warehouse aisle, and most of the difference is size.
A bounding box on a forty-pixel person has no room for error
On the aisle camera a loose box costs a strip of floor. On camera 3 a box that is five pixels too big on each side has doubled the area, and most of what is inside it is gravel. The model learns gravel as part of what a person looks like, and at the far end of the track it starts finding people in the gravel.
So the rule for this camera is that the box touches the person: head, feet, the wider of the shoulders and the outstretched arm. A vehicle gets the same treatment, bumper to bumper, roof to tyres, with the dust cloud behind it left out. The reviewer zooms to check, because at forty pixels a tight box and a loose one look the same at normal size.
The classes are few. Person, vehicle, and later an animal class after the first model kept finding people at dusk that turned out to be the site's regular fox.
Occlusion is labeled rather than skipped
The fence in front of camera 3 hides things. A person walking behind the chain-link is striped by it, a pickup behind the gate is half a pickup, and a figure between the two tanks is a head and a shoulder. The instinct is to skip those frames because the object is not clearly visible, and the instinct is wrong: those are the frames the intrusion rule exists for. Someone inside the fence is, almost by definition, partly behind something.
The rule is that a person gets a box whenever a person is recognisable, on the visible part, tagged as occluded. The tag lets the reviewer check the boundary cases as a group and lets the model learn that a striped half-person at the fence is still a person. A frame skipped because the person was hard to see teaches the model the opposite.
Class names are chosen for the rule that will use them
The hazard zone intrusion use case is a person box and a zone drawn once in the camera's own coordinates, and the whole decision is whether a point on the box crosses a line. That shapes the classes. The rule needs "person" and it needs "vehicle". It does not need "contractor" and "employee", because the camera cannot tell them apart at forty pixels and a class the labelers cannot apply consistently is a class the model cannot learn.
Where a distinction matters downstream, it goes into a tag rather than a class. A vehicle is a vehicle; whether it is the site's own pickup or an unknown one is a question for the gate log, and a label that tries to answer it from the far end of the track is a guess.
An aside from the well pad: the class list on the first day had "intruder" on it. It came off by the afternoon, because nobody could label an intruder from a frame. A person inside the fence at 2 am is a person, and the rule decides the rest.
Night and rain frames go into the set on purpose
The first sample from camera 3 was drawn from clear days, because those were the frames where a labeler could see what was there. The first model trained on them worked on clear days. Then the dust season started, and detections thinned out for a week before anyone touched anything.
The drift catalog describes exactly this: conditions the model rarely saw in training arrive for a week and detection falls off, then recovers, and the week gets dismissed as noise instead of being collected. On the well pad the fix was to go back to the recorder and pull frames from the rain, the dust, the fog that sits on the pad until mid-morning, and the nights, and to label those with the same care as the clear days. The night frames are the hardest to label and the most important, because the intrusion rule matters most at night.
My own view is that the weather frames should be the first sample, not the second. A model trained on the worst conditions the camera sees tends to handle the good ones; the reverse is never true.
The precision on each label is where the site's number comes from
Every box on the fence camera is checked by a second person before anything trains on it, at zoom, against the rule. The check is slow on this camera and it is where the value is. The 12k+ precision image annotations delivered in our oil and gas work were labeled this way, with the reviewer's time spent on the far end of the track rather than on the easy frames near the gate.
The reviewer's list for this set is short and specific. Gravel inside a person box. A vehicle box that took in its dust cloud. A striped figure at the fence with no box. A fox labeled as a person. A frame from the rain that was skipped because it was hard.
The frames the model doubts come from the fence at night
Once the model is on camera 3, the frames it is unsure of come back to a person. On the north fence they are almost all from after dark and during weather: the headlights on the access track, the figure at the gate in fog, the tarpaulin on the tank that flaps in wind. A person confirms or corrects each one, and the corrections retrain the model. When they cross the project's threshold a new version trains and replaces the old one with no downtime.
The correction rate is the signal. When the reviewer's corrections rise in the first week of the dust season, the model is seeing what the set never had, and the frames coming back are already the ones the next version needs. The clear-day set the project started with is, by the second winter, the smallest part of what the model has seen.
See it on your own footage.
Start with your footageMore in Labeling

Labeling · 6 min read
AI labeled vs human labeled data, how much a model can do before a person has to look
On a shelf dataset the model's proposals are mostly right. On safety cones in a yard they are mostly wrong. Measure the gap, then keep a person on every label.
Finn Ellingwood · Sep 28, 2026

Labeling · 6 min read
Bounding boxes in computer vision, what a box teaches and what it reports
The same rectangle on a forklift is the lesson at training time and the answer at run time. Corner and centre formats, the shifted-box bug, floor in the box.
Esdras Ntuyenabo · Sep 28, 2026

Labeling · 6 min read
Which words find the forklift, measured instead of guessed
Five ways to say forklift, five sets of boxes on aisle 6. Score each phrase on a small labeled set before Lexi labels the whole archive with the winner.
Sheikh Srijon · Sep 28, 2026