Skip to content
LexDataLexData
PlatformIndustriesCustomers
DocsThe Field GuideBlogWhy models drift
AboutCareersSecurityContact
Log inStart now
← All posts

Labeling · 7 min read

Dataset health check for computer vision, what to look at before anything trains

A scratch dataset where every scratch sits in the centre of the frame will train a model that looks in the centre. Five counts to read before the first epoch.

Summary

This post runs a health check on a defect dataset from a stamping line, reading the counts per class and split, the missing and empty labels, the near-duplicates from the recorder, the frame sizes and where in the frame the boxes sit. It concludes that each count predicts a specific failure on the live camera and is cheaper to fix before training than after. It is for teams about to train on a dataset someone else assembled.

Rajiya Sultana · Engineering Manager · Oct 3, 2026

Stamping line, steel sheets on the conveyor passing the inspection station, generated scene with detections from our model

The scratch dataset for the stamping line arrived as a folder of six thousand frames and a COCO file, assembled over a year by whoever was on the inspection station that week. Every frame with a scratch has the scratch in the middle. Not roughly in the middle; in the middle, within a few dozen pixels, on all nine hundred of them. The inspector took them with a phone, and a person photographing a defect centres it without thinking.

A model trained on that folder will learn that scratches live in the centre of the frame, and the first scratch near the sheet edge on the live camera will go past it. That is one of several things a health check finds before the first epoch, and every one of them is cheaper to fix now than after the model has been on the line for a month.

Every defect in the dataset sits in the centre of the frame

Where the boxes fall in the frame is the first thing to plot, and the plot for the stamping dataset is a tight cluster in the middle. The live camera over the conveyor sees the whole sheet, and a scratch can be anywhere on it. A model that has only ever seen centred scratches has weak evidence for the corners.

Spatial concentration is acceptable when the live camera has it too. A fixed camera looking at a fill head will always see the cap in the same place, and a dataset with every cap in that place is faithful. On the stamping line it is an artefact of the phone, and the fix is to replace the phone frames with frames from the conveyor camera, in which the scratches sit wherever the press left them.

The inspector on the Thursday shift still keeps the phone photos, because they are the only record of what a scratch looked like on the sheets that were scrapped in February.

Counts per class and per split show what the model will never learn

The second thing to read is a table: each class, each split, the number of frames and the number of boxes. On the stamping dataset from press 2 the table says scratch has nine hundred boxes, dent has sixty, and burr has four. A model can learn scratch from that. It cannot learn burr, and it will report burr as one of the other two or as nothing.

The split columns matter as much as the totals. If all four burrs are in the training split, the test score for burr is undefined and the report will show a blank that nobody reads. If two are in test, the test score for burr swings between zero and a hundred on a single frame. The how much footage lesson counts instances rather than frames for exactly this reason.

Class imbalance is not always a fault. The line produces far more scratches than burrs, and a dataset that reflects that is honest about the line. What the count tells the team is that burr needs its own collection effort, or needs to be folded into a broader class until enough frames exist, and that decision should be made before training rather than discovered in the test report.

Missing and empty labels are two different findings

A frame with no label file is a missing label. A frame whose label file says there is nothing in it is an empty label. The health check counts both, and they mean opposite things. The empty label is often correct and valuable: a clean sheet, deliberately included so the model learns what no defect looks like. The missing label is a frame nobody got to, and if it goes into training it is treated as a clean sheet whether or not it is one.

On the stamping dataset, two hundred frames had no label file. Forty of them had scratches in them, and left as they were, they would have taught the model that forty scratches were fine.

A third case sits between the two: a label file with a box that has zero width, or a class name spelled two ways, or coordinates outside the frame. Those are labels the tooling accepted and the model will read as noise. The labeling with Lexi doc describes the QA pass on every label; a health check is the same pass applied to the files rather than to the boxes.

Near-duplicate frames from the recorder inflate every count

A third of the stamping dataset came from the conveyor recorder on press 2, sampled every two seconds, and a sheet takes twelve seconds to pass the camera. Six frames per sheet, with the same scratch in each, counted as six scratches. The class table overstates scratch by a factor the team cannot see until neighbouring frames are compared and the runs collapsed.

Duplicates also leak across splits. If a random shuffle put three frames of one sheet in training and three in test, the test score for scratch is measuring memory. Grouping frames by sheet, or by the minute they were taken, before splitting is what keeps the score honest, and the health check is where the grouping is checked.

A frame's size and shape should match the camera it came from

The last table is frame dimensions. The stamping dataset had three shapes in it: the inspector's phone frames in portrait from the February scrap, the conveyor camera's wide frames, and a batch of square crops someone made for a slide deck. Training resizes all of them to one input size, and a scratch that was thin in a wide frame becomes thinner still when the frame is squashed to square, and the phone frames are stretched the other way.

This is not fatal, but it should be known. When the live camera produces one shape, the dataset should be mostly that shape, and the others should be checked for what the resize does to the smallest defects. A scratch that survives at the conveyor camera's resolution and disappears at the training size is a scratch the model will never learn to find.

The check runs again before every retrain

My own view is that a health check is worth more than a second model architecture, and that most teams spend their time the other way round. The findings above are all counts, and none took longer than an afternoon. Each one predicted a specific failure on the line: scratches at the edge, burrs reported as dents, forty scratches taught as clean, a test score borrowed from training.

The check is not a one-off. Once the model is watching press 2, frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old with no downtime. The corrected frames land in the dataset, and they change the counts. A month of corrections that are mostly edge scratches shifts the spatial plot; a batch of returned frames from a new camera brings a new frame shape. The same five tables are read before every retrain, so that what was fixed in the first pass does not creep back in through the review queue.

The burr count was still four in the third retrain. The team collected burrs on purpose after that.

See it on your own footage.

Start with your footage

More in Labeling

Labeling · 7 min read

Aerial dataset augmentation for drone frames where there is no up

A tower seen straight down has no top or bottom, so rotations and both flips are safe. Scale for altitude, brightness for sun, move every box with its pixels.

Esdras Ntuyenabo · Oct 3, 2026

Labeling · 6 min read

A collaborative data annotation workflow run as a pipeline

Batch the bottling line's frames by camera and shift, assign so nothing is boxed twice, attach the guideline, review every label, train on the approved set.

Rajiya Sultana · Oct 3, 2026

Labeling · 6 min read

A dataset quality audit for object detection when mAP hides a weak class

The PPE camera at the rig site scored well on average and missed most bare heads. Audit the labels per class, fix the boxes in review, and retrain on the fixes.

Rob Hickey · Oct 3, 2026

LexData
LexData

Product

  • Platform
  • Industries
  • Use cases

Resources

  • Docs
  • The Field Guide
  • Blog
  • Why models drift

Industries

  • Energy & utilities
  • Oil & gas
  • Agriculture
  • Manufacturing
  • Insurance
  • Retail
  • Robotics

Company

  • About
  • Customers
  • Careers
  • Contact

Trust

  • Security
  • Privacy
  • Terms

Stay updated

What we learn running vision models in production.

See everything.
Miss nothing.

Stay updated

What we learn running vision models in production.

Terms of use & Privacy policy

© 2026 LexData Labs · All rights reserved