Skip to content
LexDataLexData
PlatformIndustriesCustomers
DocsThe Field GuideBlogWhy models drift
AboutCareersSecurityContact
Log inStart now
← All posts

Labeling · 7 min read

How to label image data for computer vision so the model learns what you meant

Tight boxes on the forklift, the visible part of a hidden pallet, a written rule for the frame edge, and why a forgotten box teaches the model it is background.

Summary

This post walks through labeling frames from a warehouse aisle camera, with the forklift, the pallets and the people as the classes. It covers tight boxes, occlusion, objects cut by the frame edge, class names, and the reviewer pass, and it argues that a missing box is the most damaging label in the set because it teaches the model the object is background. It is for anyone labeling their first frames or writing the guide for people who will.

Stephen Biswas · Engineer · Sep 28, 2026

Warehouse aisle from a high camera, a forklift and racked pallets boxed, generated scene with detections from our model

The camera at the end of aisle 7 looks down the racking from about five metres up. In a frame from the Monday morning shift there is a forklift halfway along with a pallet on its forks, and a second pallet on the floor half hidden behind the first. A person is walking towards the camera, and the back end of a second forklift is leaving the frame on the right. Three classes, five objects, and at least four decisions a labeler has to make before the first box is drawn.

A model learns exactly what the boxes teach it, and the boxes teach whatever the labeler decided, including the decisions they did not know they were making.

The box touches the forklift on all four sides

A tight box is the one rule everyone has heard and most sets still break. The box on the forklift in aisle 7 touches the top of the overhead guard, the bottom of the tyres, the tip of the forks and the back of the counterweight. Nothing of the forklift is outside it and as little of the floor as possible is inside it.

The reason is what the model learns from the pixels inside the box. Floor inside the box is presented to the model as forklift. A set of boxes that each carry a strip of floor trains a model that expects floor around a forklift, and returns loose boxes of its own, evenly, on every frame. On a counting camera that is a nuisance. On a camera that decides whether the forklift is in the pedestrian lane it is the difference between an alert and none.

The forks are the common mistake. They are thin, they are the same grey as the floor, and a labeler in a hurry closes the box at the mast. The rule says forks included, and it says so in writing.

An occluded pallet gets a box on the part you can see

The second pallet on the floor is half behind the first. The labeler can guess where its hidden edge is, and the guide has to say whether to box the guess or the visible part. For aisle 7 the answer is the visible part, tagged as occluded, because the model's job on the live feed is to find pallets from what the camera shows, and the camera does not show hidden edges.

The other rule is that the pallet gets a box at all. A labeler who skips it because it is mostly hidden has just told the model that a pallet behind a pallet is floor. The threshold, how much has to be visible before it earns a box, is written in the guide as a fraction, and reviewers check the boundary cases against it rather than against their own instinct.

Objects at the frame edge need a written rule

The second forklift is leaving on the right and only its counterweight and rear wheel are in the frame. Box it or leave it? Both answers are defensible, and the wrong answer is to let each labeler choose. If half the labelers box the edge cases and half do not, the model learns that an object at the edge is sometimes an object and sometimes not, and it hesitates at the frame edge forever.

The guide for aisle 7 says a box is drawn when a recognisable part is visible, the rear wheel counts, and the box ends at the frame edge rather than where the forklift would be. The rule is short, it is arbitrary, and it is consistent, which is the property that matters.

An aside: on this camera the most argued-over object in the whole set turned out to be a hi-vis jacket hung on the end of the racking, which is not a person and looks exactly like one from five metres up. It got its own class.

A missing bounding box teaches the model the object is background

This is the one that does the most damage and gets the least attention. Every pixel outside every box is presented to the model as not the object. A person walking down aisle 7 without a box is, to the model, a piece of floor that happens to be shaped like a person. Train on enough frames like that and the model learns to leave people alone, which is the opposite of what the safety rule needs.

A forgotten bounding box is worse than a loose one. The loose box teaches a slightly wrong shape. The missing box teaches a wrong answer, and the frame looks fine to anyone scrolling past it, because there is nothing there to look wrong.

So the guide has a rule that sounds obvious: every instance of every class in every frame, including the ones at the back, the ones partly hidden and the ones nobody cares about on that particular day. Frames with genuinely nothing in them, the empty aisle at 3 am, are kept and marked as empty on purpose, so the model learns the empty aisle too.

Class names are the words the floor uses

"Forklift", "pallet", "person". Short names, the ones the shift lead would use on the radio, one per thing. The temptation is a taxonomy: counterbalance forklift, reach truck, pallet with load, pallet empty. Each split halves the number of examples per class and doubles the chances that two labelers disagree about which one they are looking at.

Start with the classes the question needs and split later if a split is needed. The labeling doc puts this as writing the prompt like a work order, and the same words become the class list. On aisle 7, "reach truck" was added in the second month, after the first model kept boxing it as a forklift and the floor said that was wrong.

A reviewer checks every label before anything trains

Every label in the aisle 7 set is looked at by a second person before it goes anywhere near training. The reviewer's checklist is the guide, read as questions: is the box tight, is the occluded pallet tagged, does the edge forklift have a box, is every person boxed, is the class the right one. The reviewer's corrections go back to the labeler, and disagreements that keep coming up become new lines in the guide.

That pass is how labels come back at up to 99.9% accuracy, and it is the part of labeling that most teams skip when the deadline is close. My own view is that a set of a thousand reviewed frames beats a set of three thousand unreviewed ones every time, because the errors in the unreviewed set are the ones nobody can find later.

Lexi proposes the boxes, a labeler checks and corrects them, and a reviewer checks the labeler. The footage lesson covers how many frames that pass needs to cover; the guide covers what each of those frames has to say.

The guide keeps changing after the first model

The guide written before the first frame is a draft. The first model shows where the labelers disagreed, the live camera shows the objects nobody thought to include, and the reviewer's notes show which rules were too vague to follow. The hi-vis jacket, the reach truck and the edge rule for aisle 7 were all added after the first version, and the guide in use today is on its fourth revision.

When the model is on the live feed, the frames it is unsure of come back to a person, and each correction is a label the guide has to be able to explain. If the reviewer cannot say why the correction is right in the guide's words, the guide is missing a line, and that line is the next thing to write.

See it on your own footage.

Start with your footage

More in Labeling

Labeling · 6 min read

AI labeled vs human labeled data, how much a model can do before a person has to look

On a shelf dataset the model's proposals are mostly right. On safety cones in a yard they are mostly wrong. Measure the gap, then keep a person on every label.

Finn Ellingwood · Sep 28, 2026

Labeling · 6 min read

Bounding boxes in computer vision, what a box teaches and what it reports

The same rectangle on a forklift is the lesson at training time and the answer at run time. Corner and centre formats, the shifted-box bug, floor in the box.

Esdras Ntuyenabo · Sep 28, 2026

Labeling · 6 min read

Which words find the forklift, measured instead of guessed

Five ways to say forklift, five sets of boxes on aisle 6. Score each phrase on a small labeled set before Lexi labels the whole archive with the winner.

Sheikh Srijon · Sep 28, 2026

LexData
LexData

Product

  • Platform
  • Industries
  • Use cases

Resources

  • Docs
  • The Field Guide
  • Blog
  • Why models drift

Industries

  • Energy & utilities
  • Oil & gas
  • Agriculture
  • Manufacturing
  • Insurance
  • Retail
  • Robotics

Company

  • About
  • Customers
  • Careers
  • Contact

Trust

  • Security
  • Privacy
  • Terms

Stay updated

What we learn running vision models in production.

See everything.
Miss nothing.

Stay updated

What we learn running vision models in production.

Terms of use & Privacy policy

© 2026 LexData Labs · All rights reserved