Skip to content
LexDataLexData
PlatformIndustriesCustomers
DocsThe Field GuideBlogWhy models drift
AboutCareersSecurityContact
Log inStart now
← All posts

Labeling · 6 min read

AI labeled vs human labeled data, how much a model can do before a person has to look

On a shelf dataset the model's proposals are mostly right. On safety cones in a yard they are mostly wrong. Measure the gap, then keep a person on every label.

Summary

This post compares two labeling jobs, a retail shelf and safety cones in a yard, and shows how far a model's proposed labels get on each before a person has to correct them. It argues for measuring where the proposals fail on a sample first, and for keeping a person checking every label before anything trains, whatever the split turns out to be. It is for teams deciding how much of a labeling run to automate.

Finn Ellingwood · Engineer · Sep 28, 2026

Bread shelf with an empty slot flagged and the rack sections boxed, from a customer store camera

Two labeling jobs land on the same desk in the same week. The first is a shelf camera over aisle 4 in a supermarket, and the store wants every empty facing boxed across a month of frames. The second is a yard camera at a depot, and the safety lead wants every traffic cone boxed so a rule can say when a cone has been moved off the pedestrian line.

The person at the desk types "empty shelf" and "traffic cone" into two projects, lets a model propose boxes on a sample of each, and gets two very different afternoons.

The shelf proposals are mostly right and the cone proposals are mostly wrong

On the shelf frames from aisle 4, the proposals are good. An empty facing is a large, dark, rectangular gap in a wall of product, the model has seen many pictures of shelves, and the boxes it draws are close to where a person would draw them. The reviewer confirms most of them, tightens a few, and draws by hand the gaps on the bottom row where the trolley bay hides half the shelf.

On the yard frames, the proposals are a mess. A cone is small, it is forty metres from the camera, it is the same orange as the depot's bollards and the forklift's beacon, and half the cones are knocked over so they do not look like cones. The model boxes the bollards, boxes the beacon, misses the fallen cones and boxes one very confident puddle. The reviewer deletes more than she confirms.

That difference is the whole question. Neither job is "AI labeled" or "human labeled". Both are proposals plus a person, and what differs is how much the person has to do.

Measure where the proposals fail on a sample before deciding anything

The mistake is to decide the split in advance. The right order is to run the proposals on a few hundred frames from the actual camera and have a person correct them. Then count: how many boxes were confirmed as drawn, how many needed moving, how many were deleted, and how many objects had to be drawn from nothing.

On aisle 4 the count said most boxes were confirmed and a few per frame were drawn by hand, all on the bottom row. On the yard it said most boxes were deleted and most cones were drawn by hand. Those two counts decide everything downstream: how long the job takes, what it costs, and where the reviewers should look first.

The count also shows the pattern of failure, which matters more than the rate. Proposals that fail on one class, or one region of the frame, or one lighting condition, are fixable by drawing those by hand and letting the model handle the rest. Proposals that fail at random are not worth running.

An aside from the yard: the fallen cones turned out to be the class the safety lead cared about most, because a fallen cone is one that a vehicle has hit. The model was worst on exactly the thing the rule was for.

Where the proposals break is predictable

The shelf in aisle 4 and the yard are instances of a general pattern. Proposals from a model that has seen the world's pictures work on objects that are common, large, distinct and described by an ordinary word. They break on objects that are rare in the world's pictures, small in the frame, similar to other things nearby, or defined by a relationship rather than an appearance.

"Cone lying on its side" is a relationship. "Facing that should have product in it" is a relationship too, and the shelf proposals only worked because an empty facing happens to look distinct on its own. A shelf where the gap is filled by the neighbouring product leaning over would have been the yard all over again.

The labeling doc covers this from the prompt side: write the prompt in the words the site uses, and expect the rare and specific classes to need a person from the start.

A person checks every label before anything trains, whichever job it is

The split between proposal and correction changes per job. What does not change is that a person looks at every label before it trains anything. On aisle 4 that check is fast, because most labels are confirmations. On the yard it is slow, because most labels are drawn. Either way, nothing goes into a training set that a person has not agreed with.

The reason is what the labels do afterwards. The shelf model will run on the store's cameras for a year, and every loose or missing box in its training set becomes a category of gap it reports wrongly for that year. The cone model will decide whether a vehicle has hit something. Neither is a place for a label that a model proposed and nobody looked at. That check on every label is how labels come back at up to 99.9% accuracy, and it is the same pass regardless of how the box was first drawn.

My own view is that the phrase "AI labeled data" should be retired. There are labels a person drew and labels a person confirmed, and a set that contains any other kind has a problem it does not know about yet.

The right split is the one the sample said

For aisle 4, the job ran with proposals on everything, a reviewer confirming and tightening, and the bottom row drawn by hand. For the yard, the job ran with proposals on the standing cones only, everything else drawn by hand, and a second reviewer on the fallen cones because they were the class the rule depended on.

Both sets went through the same verification pass and both trained a first version within the week. On the platform, the two projects look the same from the outside: a prompt, a set of reviewed labels, a model on the camera. The difference was in the hours, and the sample had said in advance where the hours would go.

The live camera changes the split over time

Once the first version is on the camera, the proposals for the next round come from that version rather than from the general model, and it has seen the yard. Its proposals on cones are better than the general model's were, because it trained on the reviewer's hand-drawn cones. The frames it is unsure of come back to her in LexAnnotate, she corrects them, and the corrections retrain it.

So the yard job, which started with a person drawing almost everything, ends up closer to the shelf job by the third version. The split was never fixed. It moved as the model learned the yard, and the person checking every label is what let it move safely.

See it on your own footage.

Start with your footage

More in Labeling

Labeling · 6 min read

Bounding boxes in computer vision, what a box teaches and what it reports

The same rectangle on a forklift is the lesson at training time and the answer at run time. Corner and centre formats, the shifted-box bug, floor in the box.

Esdras Ntuyenabo · Sep 28, 2026

Labeling · 6 min read

Which words find the forklift, measured instead of guessed

Five ways to say forklift, five sets of boxes on aisle 6. Score each phrase on a small labeled set before Lexi labels the whole archive with the winner.

Sheikh Srijon · Sep 28, 2026

Labeling · 6 min read

How much training data a computer vision model needs from your own cameras

A weld defect detector got useful on a few hundred chosen frames. A sixty-class shelf model needed far more. The difference is the task and the variation.

Sheikh Srijon · Sep 28, 2026

LexData
LexData

Product

  • Platform
  • Industries
  • Use cases

Resources

  • Docs
  • The Field Guide
  • Blog
  • Why models drift

Industries

  • Energy & utilities
  • Oil & gas
  • Agriculture
  • Manufacturing
  • Insurance
  • Retail
  • Robotics

Company

  • About
  • Customers
  • Careers
  • Contact

Trust

  • Security
  • Privacy
  • Terms

Stay updated

What we learn running vision models in production.

See everything.
Miss nothing.

Stay updated

What we learn running vision models in production.

Terms of use & Privacy policy

© 2026 LexData Labs · All rights reserved