Skip to content
LexDataLexData
PlatformIndustriesCustomers
DocsThe Field GuideBlogWhy models drift
AboutCareersSecurityContact
Log inStart now
← All posts

Computer vision · 6 min read

An end-to-end object detection workflow starts with decisions, and the first frame is labeled last

Tag, box, mask or keypoints. A class list with occlusion and size rules. A metric and an acceptance sentence. Then, and only then, the first camera's frames.

Summary

This post lays out the decisions a team makes before labeling the first frame of a detection project, using a site camera under a tower crane as the running case: which output the question needs, what the class list says about occlusion and minimum size, which metric and acceptance sentence the model will be judged by, and which camera goes first. It concludes that the class list is the most important document in the project and that labeling begins only when it is written. It is for teams about to start a detection project on cameras they already have.

Ayman Quadir · Head of Product · Oct 3, 2026

Crane lift over a fenced area from a pole camera, crane, load, hard hats and workers boxed, generated scene with detections from our model

The site manager's question is short: is anyone standing under the load when the crane is lifting. The pole camera at the north corner of the Hollins Road site sees the crane's swing radius and most of the ground beneath it, and there is a year of recorder footage nobody has looked at. The instinct is to open the labeling tool and start drawing boxes on people and loads.

The projects that go live in days are the ones where the first frame is labeled last. Everything that follows is a decision the team makes on paper, in order, before a single box is drawn on the crane camera.

01Decide whether the question needs a tag, a box, a mask or pose estimation

The question sets the output. "Is anyone under the load" on the Hollins Road camera needs to know where people are and where the load is, which is two boxes and a comparison between them. It does not need to know how much of the ground is shadowed by the load, which would be a mask. It does not need to know whether a worker's arm is raised to signal the crane operator, which would be pose estimation, points on the body rather than a box around it.

Each of those is a different model, a different label and a different cost, and the what a vision model can and cannot see lesson makes the point that the question has to be written down first because it decides all three. The crane camera's question is a box question. If the site later wants the signaller's arm, that is a second project with its own labels, and deciding that now stops the first project from labeling for both.

02Write the class list with the occlusion and minimum size rules

The class list is the most important document in the project, and my own view is that most detection projects that fail were lost here rather than at training. For the crane camera it is short: person, load, crane hook. Every entry then needs two rules the labelers will meet on the first day and answer differently unless told.

The occlusion rule says what to do when a person is half behind a container. Boxed as a person, with the box on the visible part? Boxed on where the whole person would be? Skipped? Any of these can be right; only one can be the convention.

The minimum size rule says how small a person can be in the frame and still be boxed. A figure at the far edge of the swing radius is a few pixels tall, and a labeler who boxes it teaches the model to find things the camera cannot resolve. The labeling with Lexi doc calls this writing the prompt like a work order, and the work order has to cover the edge cases before the first labeler meets them.

The site calls the load a "pick", and the class list says load, and the labelers are told both, because the site's own people will be the ones reviewing the frames.

03Pick the metric and the acceptance sentence before training

A model is going to come out of training with a number, and the team has to have decided in advance which number matters and what value of it means the model is ready. On the Hollins Road crane camera, the cost of a miss is a person under a load who was not flagged, and the cost of a false alarm is a supervisor's phone buzzing during a clean lift. Those are not equal, and the metric should say so: recall on the person class, at a threshold where the false alarm rate is one the supervisor will tolerate.

The acceptance sentence writes that down in plain words. "The model flags a person under a suspended load in at least this fraction of the held-out frames, with no more than this many false alerts per shift." Written before training, it is a target. Written after, it is a description of whatever the model happened to do.

04Choose the camera whose frames go first

There are six cameras on the site, and the temptation is to label frames from all of them for coverage. The better first step is one camera, the north pole camera that sees the swing radius best, and enough of its frames to cover the situations that camera meets. Morning and afternoon light, the crane at each end of its arc, the ground with and without containers stacked on it, the Thursday it rained.

One camera first means the first model can be on that camera and producing frames for review while the others are still being planned. It also gives the honest hold-out: the second camera, added later, is the test of whether the model learned people under loads or learned the north corner of the site.

05Label the first camera's frames and check every one

Only now does the labeling tool open. You type what to look for, in the words of the class list, Lexi puts a box on every person and load in every frame, and a person checks each label before anything trains on it. The check is against the class list's rules: the half-hidden person boxed the way the occlusion rule says, the far figure skipped the way the size rule says. A box that follows a different convention is corrected, and if the same correction is made three times, the class list gets a line.

This is where the days go, and it is the only place they should. The decisions above took an afternoon on paper. The labeling takes what it takes, and it takes less when nobody is deciding the occlusion rule while drawing.

06Put the first version on the camera and let the review queue set the second

The first model goes onto the north camera at Hollins Road and is judged by the acceptance sentence on held-out frames. If it passes, the rule goes live: a person inside the load's zone while the hook is raised, critical, to the site supervisor in Slack, with a cooldown. From there, frames the model is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime. The quickstart covers the mechanics from workspace to first alert.

The review queue is also what tells the team what to label next. If the doubted frames are mostly from dusk, dusk frames go into the next batch. If they are mostly the second camera, the second camera was different enough to need its own window of labels. On the site under the crane, the first version was on the north camera within the week, and the second camera's frames were the second week's work, chosen by what the first version could not decide.

See it on your own footage.

Start with your footage

More in Computer vision

Computer vision · 6 min read

What CLIP is and why typing a description finds the frame

A shared space for pictures and words lets an operator type a sentence and get the matching frames from a recorder. A way to look, and the boxes come after.

Rob Hickey · Oct 3, 2026

Computer vision · 7 min read

Image recognition AI, explained on the cameras a site already owns

A tag, a box, a mask or keypoints, on a substation camera, a line camera and an aisle camera. How a model gets there, and what its score does not promise.

Esdras Ntuyenabo · Oct 3, 2026

Computer vision · 7 min read

Image segmentation and the questions on a weld bead that only a mask can answer

A box finds the weld. A mask measures it. Semantic, instance and panoptic told apart on a weld cell and a board camera, with what each costs to label.

Esdras Ntuyenabo · Oct 3, 2026

LexData
LexData

Product

  • Platform
  • Industries
  • Use cases

Resources

  • Docs
  • The Field Guide
  • Blog
  • Why models drift

Industries

  • Energy & utilities
  • Oil & gas
  • Agriculture
  • Manufacturing
  • Insurance
  • Retail
  • Robotics

Company

  • About
  • Customers
  • Careers
  • Contact

Trust

  • Security
  • Privacy
  • Terms

Stay updated

What we learn running vision models in production.

See everything.
Miss nothing.

Stay updated

What we learn running vision models in production.

Terms of use & Privacy policy

© 2026 LexData Labs · All rights reserved