Skip to content
LexDataLexData
PlatformIndustriesCustomers
DocsThe Field GuideBlogWhy models drift
AboutCareersSecurityContact
Log inStart now
← All posts

Labeling · 6 min read

Missing vs null annotations, an empty frame is a label and a forgotten box is a lie

An empty corridor, a defect-free shift and a stocked shelf are labels on purpose. A frame where someone forgot the box looks identical and teaches the opposite.

Summary

This post separates two frames that look the same in a dataset: one labeled empty on purpose, and one where a labeler never drew the box. It shows why the first trains a model well and the second trains it to ignore the object, how the common export formats hide the difference, and how a review pass in LexAnnotate tells them apart. It is for anyone auditing a dataset before training on it.

Rajiya Sultana · Engineering Manager · Sep 28, 2026

Dairy case with a shopper, rack sections and carts boxed and no gap flagged, from a customer store camera

The dataset for the shelf gap model has eleven thousand frames, and a quarter of them have no boxes at all. The retail analyst who inherited it cannot tell, for any one of those frames, whether the shelf was full and a labeler marked it so, or whether the labeler opened it at the end of a long Friday and moved on. Both frames sit in the folder as an image with no label, and the export format treats them the same.

The model will not treat them the same. One of those frames is a fact. The other is a false negative wearing a fact's clothes.

An empty frame labeled empty is one of the most useful labels in the set

A null label says, deliberately, that the object is not here. The corridor camera at 3 am with nobody in it. The inspection station on a shift where every part passed. The dairy case in aisle 6 with every facing full. Each is a frame the live camera will see every day, and a model that has never trained on the empty case does not know what empty looks like. It finds gaps in a full shelf because full is unfamiliar.

So the empty frames go in on purpose, marked as empty, and they are cheap. Nobody draws anything. A reviewer confirms that the frame was looked at and that the answer is nothing, and that confirmation is the label.

A forgotten box teaches the model the object is background

A missing label is a frame where the object is present and nobody drew the box. To the training run, every pixel outside a box is not the object, so a forgotten gap on the bottom row of the dairy case is presented to the model as a full shelf. Enough of those and the model learns that bottom-row gaps are fine, which is exactly the gap the store loses sales on.

The partial version is worse because it is harder to find. A frame with three gaps boxed and a fourth missed looks labeled. Anyone scrolling past sees boxes and moves on. The fourth gap is a false negative in the training set, and the model inherits it as a blind spot that shows up months later as a category of gap it never reports.

An aside from the same dataset: the frames with missing boxes clustered on two dates. One was the Friday. The other was the day a new labeler started, before anyone had shown her the bottom row rule.

The export formats cannot tell the two apart

In the common formats, an image with a null label and an image with a forgotten label are written the same way. A COCO JSON file has an entry for the image and no annotation records pointing at it. A Pascal VOC XML file has the image's size and no object element. Neither says whether a person decided the frame was empty or never opened it.

That is why a count of empty images in an export is a starting point rather than an answer. For the shelf set the count was a quarter of the frames. The question was how many of those were the analyst's problem, and the format could not say. What can say is a record kept alongside the labels: who reviewed the frame, when, and whether they marked it empty on purpose.

A review pass in LexAnnotate tells the two apart

The way the shelf set was sorted out was a review pass. Every frame with no boxes went to a reviewer with the question stated plainly: is this frame actually empty? A full dairy case got confirmed as empty of gaps and kept. A frame with a gap on the bottom row got its box drawn and moved to the labeled set. A frame that was blurred or pointed at the floor got marked unusable rather than empty.

In LexAnnotate the empty mark is explicit. A frame the reviewer confirmed has a record of that confirmation, and a frame nobody has opened has none, so the two states are different in the dataset rather than only in someone's memory. The labeling doc describes the verification pass that every label goes through before training; the empty frames go through the same pass, and that is what makes them labels.

My own view, from running these queues, is that the review of empty frames should be done by a different person from the one who labeled the set. A labeler who skipped a box on Friday will skip it again on Monday, for the same reason.

Auditing the set before training is cheaper than finding out after

The order that works is to count the empty frames, look at when and by whom they were labeled, and review the ones from any labeler or any day that stands out. On the shelf set that was the Friday, the new labeler's first day, and one person, and reviewing those frames found most of the missing boxes without opening the rest.

Then a sample of the remaining empty frames gets a second look anyway, because the clustering finds the systematic gaps and a sample finds the random ones.

The millions of annotations behind our labeling work were checked by hand, and the empty frames were checked as carefully as the full ones. An unchecked empty frame is the easiest place for a false negative to hide.

An object detection model on the live shelf keeps sending empty frames back

Once the gap model is on the camera over aisle 6, the same distinction keeps mattering. The model returns the frames it is unsure of, and many of them are near-empty: a full shelf with a shadow that looks like a gap, or a gap the model half sees. A person confirms empty or draws the box, and either way the frame is now a deliberate label rather than a guess.

Those corrections retrain the model. The frame where the reviewer confirmed the shelf was full is as much a correction as the one where she drew a gap, and the next version has learned both. The footage lesson makes the same point about coverage: a set that only holds the interesting frames has not covered the day, and most of the day is empty.

See it on your own footage.

Start with your footage

More in Labeling

Labeling · 6 min read

AI labeled vs human labeled data, how much a model can do before a person has to look

On a shelf dataset the model's proposals are mostly right. On safety cones in a yard they are mostly wrong. Measure the gap, then keep a person on every label.

Finn Ellingwood · Sep 28, 2026

Labeling · 6 min read

Bounding boxes in computer vision, what a box teaches and what it reports

The same rectangle on a forklift is the lesson at training time and the answer at run time. Corner and centre formats, the shifted-box bug, floor in the box.

Esdras Ntuyenabo · Sep 28, 2026

Labeling · 6 min read

Which words find the forklift, measured instead of guessed

Five ways to say forklift, five sets of boxes on aisle 6. Score each phrase on a small labeled set before Lexi labels the whole archive with the winner.

Sheikh Srijon · Sep 28, 2026

LexData
LexData

Product

  • Platform
  • Industries
  • Use cases

Resources

  • Docs
  • The Field Guide
  • Blog
  • Why models drift

Industries

  • Energy & utilities
  • Oil & gas
  • Agriculture
  • Manufacturing
  • Insurance
  • Retail
  • Robotics

Company

  • About
  • Customers
  • Careers
  • Contact

Trust

  • Security
  • Privacy
  • Terms

Stay updated

What we learn running vision models in production.

See everything.
Miss nothing.

Stay updated

What we learn running vision models in production.

Terms of use & Privacy policy

© 2026 LexData Labs · All rights reserved