Labeling · 6 min read
Natural language image annotation, type what to look for and check what comes back
Hard hat and reflective vest, typed once and run across a site camera's frames. The boxes get checked one by one, and the lanyard clip gets drawn by hand.
Summary
This post follows a site safety camera through natural language annotation, from typing hard hat and reflective vest to reviewing the boxes that come back and drawing by hand the lanyard clip the prompt could not find. It concludes that the prompt sets the scope of the project, the review pass sets its quality, and the reviewed set is what the first model trains on. It is for site and safety teams starting a labeling run.
Esdras Ntuyenabo · Engineer · Sep 28, 2026

Worker on a scaffold deck, hard hat and fall-arrest gear boxed, from a customer site camera
The camera on the tower crane mast looks down on the scaffold at the east elevation, and the safety manager wants to know, every day, whether everyone on the deck has a hard hat and a vest on. Two years of footage sit on the recorder. The old way to start was a week of someone drawing boxes on hats.
The new way is to type "hard hat" and "reflective vest" into the project, and to wait while Lexi runs both across a sample of frames from the mast camera. Boxes come back in minutes. The rest of this post is about what happens to those boxes, because the minutes are the smallest part of the work.
Type hard hat and the mast camera's frames come back boxed
A natural language prompt is matched against the image rather than a trained class. The model has learned to place words and image regions in the same space, so "hard hat" pulls out the regions of a frame that look like the pictures the words were seen beside. On the mast camera, Lexi returns a box on every hat on the deck, a yellow bucket by the ladder, and the top of a hi-vis jacket hung on the guardrail.
The prompt wording decides what comes back. "Hard hat" and "helmet" return nearly the same boxes. "PPE" returns hats, vests, gloves and a fire extinguisher. "Person wearing a hard hat" returns people, which is a different question, and on this camera the safety manager wants the hat and the person separately so the rule can say a person without a hat.
The labeling doc puts it as writing the prompt like a work order. The words the site uses, one object per word.
Every box is reviewed one by one before it means anything
The boxes from the prompt are proposals. A reviewer opens the frames from the mast camera and goes through them: confirm the hat, confirm the vest, delete the bucket. Tighten the box that took in the worker's shoulder. Tag the vest on the far worker as occluded because the scaffold tube crosses it.
Confirming is fast. Most of the reviewer's time goes on the boxes that are nearly right and the ones that are missing. The prompt found the hats on the near deck and missed two on the fourth lift, where the netting blurs them. Those two have to be drawn by hand, because a hat that is never boxed is a hat the model learns to ignore.
The review is not optional and it is not a sample. Every label in the set is checked by a person before anything trains on it. That is how labels come back at up to 99.9% accuracy, and on a safety camera it is the difference between an alert the site trusts and one it turns off.
The lanyard clip is the class the prompt cannot find
The safety manager's third class is the lanyard clip, the carabiner that should be attached to the lifeline whenever a worker is on the outer deck. It is small, it is metal against metal, and the words for it are not ones the model has seen beside many pictures. "Lanyard clip" returns the scaffold couplers when Lexi runs it. "Carabiner" returns nothing useful at this distance.
So the clip is drawn by hand, on every frame in the sample where it is visible, by the reviewer who already knows the deck. It is the slowest class in the set by a distance and the one the site cares about most, and the prompt saved no time on it at all.
That pattern holds on every camera. The prompt covers the common, visually distinct objects and leaves the rare, small, specific ones to a person. An aside from this site: the reviewer's notes on the clip frames included the observation that the clip is easiest to see in the late afternoon, when the low sun catches it, and that note became a rule for which frames to sample next.
Computer vision projects get their scope from what the prompt found
By the end of Monday afternoon the safety manager knows three things. Hats and vests are easy on this camera. The clip is hard. And the far end of the deck under the netting is hard for everything. That is a scope, and most computer vision projects would have taken a month of labeling to learn it.
The scope changes the plan. The hat and vest rule can go live on the first model. The clip needs more frames, sampled in the afternoon light, drawn by hand, and the model for it will follow. The netted end of the deck needs the camera looked at, because no label fixes a view the camera cannot see.
My own view is that the prompt is more valuable as a scoping tool than as a labeling tool. The boxes it draws save an afternoon. The list of what it could not draw saves a quarter.
The reviewed set is what the first model trains on
The prompt is not the model that goes on the camera. It is slow per frame, it needs a large model to run, and its answers move with the wording. The mast camera needs a detector that looks at every frame, every couple of seconds, on hardware at the site, and gives the same answer to the same frame on Tuesday that it gave on Monday. That detector trains on the reviewed set.
LexData takes that model through its whole life. You type what to look for, Lexi puts a box on every frame, and a person checks each label before anything trains on it. The model then watches the cameras you already have, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime. The accuracy you launched with is the accuracy you keep.
The site safety and PPE use case has the same shape on other sites: a person box, a hat box, and a rule that fires when the first exists without the second.
The prompt comes back out when the site changes
Once the hat model is on the mast camera, the frames it is unsure of come back to the reviewer in LexAnnotate. Most of them are the ones the prompt struggled with too: the netted end, the low winter sun, the new subcontractor's white hats that the first set never saw. The reviewer corrects them, the corrections retrain the model, and the next version knows the white hats.
When the site adds a class, a new prompt runs across the recent frames and the same afternoon repeats, shorter each time, because the reviewer now knows which end of the deck to look at first.
See it on your own footage.
Start with your footageMore in Labeling

Labeling · 6 min read
AI labeled vs human labeled data, how much a model can do before a person has to look
On a shelf dataset the model's proposals are mostly right. On safety cones in a yard they are mostly wrong. Measure the gap, then keep a person on every label.
Finn Ellingwood · Sep 28, 2026

Labeling · 6 min read
Bounding boxes in computer vision, what a box teaches and what it reports
The same rectangle on a forklift is the lesson at training time and the answer at run time. Corner and centre formats, the shifted-box bug, floor in the box.
Esdras Ntuyenabo · Sep 28, 2026

Labeling · 6 min read
Which words find the forklift, measured instead of guessed
Five ways to say forklift, five sets of boxes on aisle 6. Score each phrase on a small labeled set before Lexi labels the whole archive with the winner.
Sheikh Srijon · Sep 28, 2026