Labeling · 7 min read
Data annotation for computer vision is the problem statement, written in boxes
On a warehouse frame, the pallet nobody boxed teaches absence and the loose box teaches where a pallet ends. The schema is the spec, and QA is what holds it.
Summary
This post argues that labeling is where a computer vision problem gets defined, using a warehouse aisle frame in which a pallet nobody boxed becomes a lesson in absence and a loose box becomes a lesson in boundaries. It covers the schema as the written specification, the rules for occlusion and ambiguity, the four label types as four different questions, and the QA pass that makes the specification hold across thousands of frames. It is for teams about to label their first dataset.
Sheikh Srijon · GTM Lead · Oct 2, 2026

Warehouse aisle from a high camera, forklift and pallets boxed, generated scene with detections from our model
A frame from the high camera over aisle 7 of a warehouse shows nine pallets on the racking, a forklift with a tenth on its forks, and a stack of loose cartons on the floor where someone broke a pallet down and did not finish. A labeler boxes the nine pallets on the racks and the one on the forks. She skips the cartons on the floor, because they are not a pallet. She also skips the pallet on the top rack at the far end, because it is small and half behind an upright and she is on her four hundredth frame of the afternoon.
Both of those decisions are now part of the model. Neither was written down anywhere.
That is the argument of this post. Labeling is not a chore that happens before the real work. It is the real work, because every box drawn and every box skipped is a sentence in the specification the model will learn, and a specification nobody wrote is still a specification.
The pallet nobody boxed becomes a lesson in absence
The pallet on the top rack at the far end is in the frame. The model, training on that frame, sees a pallet-shaped object with no box on it, and learns from it exactly what it learns from the empty floor: this is background. Do that on a few hundred frames, and the model has been taught, carefully and consistently, that small pallets on high racks at the far end of the aisle are not pallets.
Nobody meant that. The labeler was tired and the pallet was small. But a missing box is a false negative the model is trained on, and the model has no way to tell a deliberate skip from an accidental one. It only sees what was boxed and what was not.
The labeling guide puts this under the verification pass, and it is the reason Lexi proposes a box on every frame first. The person's job becomes confirming and correcting rather than finding, and the far pallet gets a proposed box the labeler has to actively reject.
A loose box teaches the model where a pallet ends
The cartons on the floor are a different lesson. Suppose the labeler had boxed them as a pallet, because the schema said "pallet" and a broken-down pallet is still sort of one. Now the model has a box around a loose heap with no pallet under it, and it has learned that the class extends to heaps. On the next frame with a pile of returns by the dock door, it draws a pallet.
Or suppose she boxed the pallet on the forks in aisle 7 loosely, with the box taking in half the forklift mast. The model learns that a pallet's boundary includes a bit of forklift, and on frames with no forklift the boxes come out short, or wide, or hesitant. Boundaries are learned from boundaries, and a loose box is not a harmless approximation. It is an instruction about where the thing ends.
The warehouse's labeling lead has a rule she repeats until people are sick of it: the box hugs the wood.
The schema is written before the first box is drawn
What the far pallet and the loose cartons have in common is that the answer to both was decided by a labeler under pressure, when it should have been decided by the schema. The schema is the list of classes and the definition of each. What counts as a pallet, whether a broken-down one counts, whether a pallet on forks is the same class as a pallet on a rack, and what to do with the small one at the far end.
Writing that down is writing the problem statement. The lesson on what a vision model can and cannot see makes the point that most failed vision projects failed on a question nobody wrote down. The schema is where that question either gets written or gets left to four hundred individual decisions per labeler per afternoon.
A schema for aisle 7 can be a page. Class: pallet, meaning a wooden or plastic pallet with or without a load, boxed to the edge of the pallet and not the load. Loose cartons: not a pallet, and boxed as their own class if the model needs to know about them. Pallet on forks: a pallet. Far and small: still a pallet, boxed however small.
Occlusion and ambiguity get a rule before labelers invent one
The two cases every schema has to settle are the thing that is half hidden and the thing that might not be the thing at all. On aisle 7, occlusion is the pallet behind the rack upright: box the visible part, box the whole estimated extent, or skip below some fraction visible. Any of the three can be right. All three in the same dataset is wrong, because the model is being taught three boundaries for one situation.
Ambiguity is the shrink-wrapped block on the floor that might be a pallet under the wrap or might be a stack of cartons. The schema either says which, or says to flag the frame as unclear rather than guess. The flag is the better answer more often than people expect, since a frame marked unclear can be looked at by someone who knows the warehouse, and a frame guessed wrong is in the training set forever.
My own view is that occlusion rules are where most schemas are silent and where most label disagreements come from, and that writing the occlusion rule first, before the class list, would save most teams a relabel.
Boxes, polygons, points and tags are four different questions
Boxes are what aisle 7 needs, because the question is "where is each pallet". The other label types answer other questions, and choosing among them is choosing what the model will be asked. A polygon answers "what is the exact outline", which matters for a spill on the floor and not for a pallet. A point answers "where is the centre of this", enough for counting cartons on a belt. A tag answers "what is true of this whole frame", such as whether the aisle is blocked, with no location at all.
A dataset can carry more than one, and on the platform all four are available with instance tracking across frames, so a pallet that moves down the aisle on the forklift keeps its identity from one frame to the next. But each label type is a different specification, and a team that draws polygons because they look thorough, when the rule only ever needed a box, has paid for precision it will not use.
A QA pass on every label is what makes the specification hold
A schema on a page and a labeler on her four hundredth frame are still two different things, and the QA pass is what closes the gap. Every label, whether a person drew it or Lexi proposed it, is checked before anything trains on it: is the class right, does the box hug the wood, was the occluded pallet handled the way the schema says. Labels come back at up to 99.9% accuracy because of that pass, and the number is the schema being enforced rather than the labelers being perfect.
LexData takes the pallet model through its whole life. You type what to look for, Lexi puts a box on every frame, and a person checks each label before anything trains on it. The model then watches the warehouse cameras, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime. The specification written on aisle 7 is the one the model keeps, because every correction is checked against it too.
The far pallet on the top rack got its box on the QA pass. The labeler, shown it, said she had wondered about that one.
See it on your own footage.
Start with your footageMore in Labeling

Labeling · 7 min read
AI data labeling workflows, three ways to label footage and when each one pays
A pipeline right-of-way survey labeled three ways: every box by hand, Lexi proposing and a person checking, and synthetic frames for the leak nobody has filmed.
Rajiya Sultana · Oct 2, 2026

Labeling · 6 min read
Annotation analytics, the numbers a labeling queue produces besides labels
Three labelers on a month of warehouse footage produce boxes, and also a throughput, an acceptance rate and a map of where the rejections cluster.
Rajiya Sultana · Oct 2, 2026

Labeling · 7 min read
Annotation format conversion between COCO, YOLO and CVAT without losing a box
Three years of line inspection labels from two tools arrive in three formats. The boxes that shift are the ones nobody draws on a frame before training.
Esdras Ntuyenabo · Oct 2, 2026