Labeling · 6 min read
How to identify mislabeled images before the next retrain learns them
A shelf dataset had bleach boxed as beverages for a month and validation never noticed. Find the frames where the model disagrees with the label.
Summary
This post takes a supermarket shelf dataset in which cleaning products were labeled as beverages for weeks, explains why the validation split hid the mistake, and shows two ways to find such frames: outliers among the frames that share a label, and the frames where a trained model disagrees with what a person wrote. It concludes that the QA pass on every label is what stops the same mistake being made twice. It is for teams cleaning a dataset before a retrain.
Sheikh Srijon · GTM Lead · Oct 2, 2026

Dairy case with three empty slots flagged across the rack sections, from a customer store camera
The dataset for the aisle cameras in a supermarket had been growing for a month before anyone opened the beverages class and looked. Somewhere in the second week a labeler had started boxing the bleach and the surface sprays on the end of aisle 9 as beverages. The bottles were the same shape as the water, the same height, the same colour cap, and the guideline had a picture of a water bottle beside the word "beverage" and nothing beside "cleaning". A few hundred frames later the mistake was a pattern.
The model trained on it did well on validation. That was the problem.
Validation hides a mistake that was made consistently
A validation split is a random slice of the same labels. If the bleach was labeled as a beverage in the training frames, it was labeled as a beverage in the validation frames too, and the model that learned to call bleach a beverage was rewarded for it on both. The score for the beverages class looked healthy. The store, when the shelf and planogram check went live on aisle 9, got a report that the cleaning end cap was fully stocked with drinks.
Random errors wash out. A labeler who misses one bottle on one frame has cost the model almost nothing. A labeler who makes the same wrong call on every frame of a product for a month has taught the model a rule, and the model applies it with the same confidence it applies the right ones. A mistake that is consistent is a mistake that validation cannot see, because validation is measured against the mistake.
Frames that look wrong for their label stand out from the rest
The first way to find the frames is to look at the beverages class as a set and ask which members look least like the others. Each labeled crop can be turned into a vector that describes its appearance, and crops with the same label should sit near each other. The water bottles from aisle 3 sat in one cluster. The juice cartons in another. And off to one side, a tight little group of crops that were near each other and far from everything else in the class: the bleach.
An outlier inside a class is a question rather than a verdict. Some are genuine variety, a new pack design or a product photographed under different light. Some are mislabels. The point of the map is that it turns a search across a few thousand crops into a look at a few dozen, and a person can settle a few dozen in an hour. The labeling guide puts the same person at the same step: a label is verified when a person has actually seen it, and the outliers are the labels most worth a second look.
Duplicates show up the same way, as a knot of crops that sit on top of each other. Fifty near-identical frames of the same bottle from one quiet minute of the feed are not fifty examples. They are one example that will be counted fifty times in the score.
A trained model disagreeing with a label is usually worth a look
The second way uses the model against its own training set. Train on the labels as they stand, run the model over the same frames, and list every crop where the prediction and the label disagree. Most of the disagreements are the model being wrong, since it is still learning. A stubborn few are the label being wrong, and they are recognisable because they cluster on one product, one aisle or one week.
On the shelf dataset the list for aisle 9 had the bleach near the top. The model, having also seen the cleaning class labeled correctly on other aisles, was unsure whether the end cap bottles were beverages or cleaning, and its uncertainty on exactly those frames was the tell. A model can only surface this where it has seen the right answer somewhere else in the set, which is one reason a dataset with a few clean examples of every class is easier to audit than one that is thin.
A better version of this uses folds: train on most of the set, check the held-out part, rotate. Each frame is then judged by a model that never trained on it, and the label it disagrees with was never in that model's training. It costs a few training runs, and on a dataset that is about to be retrained on for production it is a cheap price.
My own view is that a team should run this disagreement check before every retrain, on the whole set and not on new frames only. A retrain that folds a month of corrections into a model also folds a month of consistent mistakes, and the corrections are the more visible of the two.
Fixing the frames is a review job, done once and recorded
The mislabeled frames go back to a person. The bleach frames were relabeled as cleaning on a Thursday afternoon, the guideline gained a second picture, and the class definitions gained a sentence about bottle shapes on the cleaning aisle. The frames the reviewer could not settle, a hand soap in a bottle the same shape as a sports drink, were flagged rather than guessed, which is what the review queue is for.
The relabeled frames were then checked again with the same outlier map. This is the step teams skip. Fixing the frames that were found and stopping leaves in the ones that were not, and the map after the fix shows whether the cluster has gone or only shrunk.
Write down what was found. The store's dataset now carries a note that says which frames were relabeled, when, and why, so the next person who sees a strange score on beverages knows there was a month when the class was wrong.
The QA pass is what stops it happening a second time
A person checking each label before it trains would have caught the bleach on the first day, because the second labeler and the reviewer would have had to agree with the first. The consistent mistake needs a single unchecked labeler and a guideline with a gap. Take away either and it does not form.
This is why our labeling pass has a reviewer on every label and a small random sample re-checked from scratch each session. The retail work behind our 4M+ annotations and validations is shelf footage of exactly this kind, hundreds of products that look alike at shelf angle, and the way the accuracy holds is that no label trains without a second pair of eyes on it. The frames the model doubts in production come back through the same pass, and the corrections that retrain it were checked the same way, so a mistake has to get past two people twice to become a rule.
The guideline for the store now has a picture of the bleach under "cleaning". It took a month to earn it.
See it on your own footage.
Start with your footageMore in Labeling

Labeling · 7 min read
AI data labeling workflows, three ways to label footage and when each one pays
A pipeline right-of-way survey labeled three ways: every box by hand, Lexi proposing and a person checking, and synthetic frames for the leak nobody has filmed.
Rajiya Sultana · Oct 2, 2026

Labeling · 6 min read
Annotation analytics, the numbers a labeling queue produces besides labels
Three labelers on a month of warehouse footage produce boxes, and also a throughput, an acceptance rate and a map of where the rejections cluster.
Rajiya Sultana · Oct 2, 2026

Labeling · 7 min read
Annotation format conversion between COCO, YOLO and CVAT without losing a box
Three years of line inspection labels from two tools arrive in three formats. The boxes that shift are the ones nobody draws on a frame before training.
Esdras Ntuyenabo · Oct 2, 2026