Skip to content
LexDataLexData
PlatformIndustriesCustomers
DocsThe Field GuideBlogWhy models drift
AboutCareersSecurityContact
Log inStart now
← All posts

Operations · 6 min read

The class that fails behind a good average, read off the confusion matrix

A shelf model with a healthy mAP and a dairy case full of unreported gaps. The matrix names the cell the failing class fell into, and the cell names the fix.

Summary

This post reads a confusion matrix for a shelf detector whose overall score looked fine while its empty-space class was failing in the dairy case. It walks the matrix cell by cell, pulls precision and recall off the rows and columns, and turns each failing cell into a labeling and retraining decision. It is for retail and store technology teams signing off a shelf model.

Rob Hickey · Chief AI Officer · Sep 28, 2026

Dairy case with three empty slots flagged across the rack sections, from a customer store camera

The shelf model went live in the dairy aisle with three classes, product, empty space and price label, and a mean average precision on the sign-off sheet that nobody argued with. Three weeks later the Tuesday restock list was wrong. The dairy case had four empty slots at 7 am and the list showed one. The score on the sheet had not moved, because the sheet was an average, and the average was mostly products.

Products are easy. Every frame of a dairy case holds dozens of them, the model finds nearly all of them, and their row in the arithmetic drowns everything else. The empty-space class appears a few times a frame and was quietly failing behind a number that looked fine.

A good average is where a failing class hides

Mean average precision averages over classes, and averaging is exactly what lets a rare class fail unnoticed. Three classes at a healthy figure, or two classes near perfect and one near useless, can produce the same mean. The store does not experience the mean. It experiences the Tuesday restock list, and that list is produced entirely by the class that was failing.

So the first move, before anyone touches the model, is to stop looking at one number and start looking at a grid.

The matrix puts every mistake in a named cell

A confusion matrix for a detector is a table with the true classes down the side and the predicted classes across the top, plus one extra row and one extra column for "nothing there". Each labeled thing in the held-out frames lands in one cell. A product predicted as a product sits on the diagonal, an empty slot predicted as a product sits off it, and an empty slot the model drew no box on at all sits in the "missed" column.

Every off-diagonal cell is a kind of mistake with a name. The dairy case on a Tuesday morning produced two kinds that mattered. Empty slots predicted as product: the model saw the back panel of the case, with its printed brand strip, and boxed it as stock. Empty slots missed entirely: a gap in the bottom shelf, in shadow under the lip of the shelf above, that the model drew nothing on.

Neither of those is visible in a mean. Both are visible in a cell.

The empty-space row says where the misses went

Read the row for the class that failed. Across the empty-space row, the count on the diagonal is how many gaps were found as gaps. The count under "product" is how many gaps were mistaken for stock. The count under "missed" is how many were never boxed. On the dairy case at 7 am the "product" cell was the larger of the two, which points at a specific cause. The printed back panel of the case looks like a facing to a model that has seen too few empty slots with that panel behind them.

Then read the column. Down the empty-space column, the count in the product row is how many full facings were boxed as gaps, which is the false alarm the assistant manager walks to. That number was small, which is why nobody had complained. The failure was silent, in the misses, and silent failures are the ones a store finds out about from the restock list.

The dairy case door, as it happens, fogs for about a minute after every restock, and those frames were in the held-out set too.

Precision and recall come straight off the rows and columns

Precision for a class is its diagonal cell divided by the total of its column: of everything the model called empty space, how much was. Recall is the diagonal divided by the total of its row: of every real gap, how many the model found. Precision and recall on the dairy case were lopsided, high precision and poor recall, and the lopsidedness is the diagnosis. A model that is rarely wrong when it speaks and often silent is a model whose threshold is too cautious for this class, or whose training set held too few examples of the thing, or both.

The threshold part is quick to test. The number beside each detection is a ranking, as the lesson on what the number beside a detection means lays out. Lowering the line for empty space alone will move some of the missed gaps onto the diagonal at the cost of a few more false alarms. If the recall on the Tuesday frames barely moves when the line drops, the threshold was never the problem and the frames were.

Each failing cell points at a labeling job

This is where the matrix earns its place: each cell names what to label next. Gaps mistaken for product with the brand panel behind them means the training set needs more frames of that case with that panel visible through a gap, labeled as empty space. Gaps missed in the bottom-shelf shadow means frames from the evening, when the aisle lights are on and the case lights make the shadow, with the gap boxed by a person.

Neither of those is a training knob. Nobody changes a learning rate to fix the bottom shelf. The fix is a few hundred frames from the dairy camera, chosen by the cells, with Lexi putting the first box on each and a person checking each label before it counts. The shelf and planogram compliance use case is built around the same reading: the rare class is the one the store is paying for, and the matrix is how you find out it is starving.

My own view is that the matrix should be printed per camera, and that a model-wide matrix is only slightly better than a model-wide mean. The dairy case fails differently from the bread wall, and the bread wall's diagonal will hide the dairy case's misses just as the average did.

The corrected frames retrain the class that failed

Once the new frames are labeled, the model retrains and the matrix is re-read on the same held-out set, and the empty-space row is the only row anyone looks at. LexData takes the shelf model through its whole life. You type what to look for, Lexi puts a box on every frame, and a person checks each label before anything trains on it. The model then watches the aisle cameras the store already has, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime.

After go-live the matrix is rebuilt from the corrections rather than from a fixed set: every gap a person marks that the model missed is a count in the "missed" cell, live, per camera. The retail work behind our 4M+ annotations and validations is largely this, one shelf and one cell at a time, and the restock list on a Tuesday is the score that matters.

See it on your own footage.

Start with your footage

More in Operations

Operations · 7 min read

Build or buy the layer that keeps a vision model accurate

Two engineers and a pilot can build a detector in a month. The review queue, the versioning and the rollout are what they are still building a year on.

Ayman Quadir · Sep 28, 2026

Operations · 7 min read

Benchmark the model on your own cameras before you believe a score

Two candidate models, two published scores, and a packaging line that only cares which one finds the torn label. The held-out week from your cameras decides.

Rob Hickey · Sep 28, 2026

Operations · 6 min read

Computer vision on the industrial HMI, the camera's verdict on the operator screen without the noise

The verdict as a state, the frame one click away, an alarm the operator acknowledges like any other, and everything else kept off the screen.

Andreas Ohrvall · Sep 28, 2026

LexData
LexData

Product

  • Platform
  • Industries
  • Use cases

Resources

  • Docs
  • The Field Guide
  • Blog
  • Why models drift

Industries

  • Energy & utilities
  • Oil & gas
  • Agriculture
  • Manufacturing
  • Insurance
  • Retail
  • Robotics

Company

  • About
  • Customers
  • Careers
  • Contact

Trust

  • Security
  • Privacy
  • Terms

Stay updated

What we learn running vision models in production.

See everything.
Miss nothing.

Stay updated

What we learn running vision models in production.

Terms of use & Privacy policy

© 2026 LexData Labs · All rights reserved