Skip to content
LexDataLexData
PlatformIndustriesCustomers
DocsThe Field GuideBlogWhy models drift
AboutCareersSecurityContact
Log inStart now
← All posts

Operations · 7 min read

Precision, recall and the mistakes each one lets through

A shelf model that cries wolf and an inspection model that waves the crack past are wrong in opposite ways. The threshold is the dial between them.

Summary

This post explains precision, recall and F1 through two models that fail differently, a shelf-gap detector that alerts on full shelves and a line inspector that passes cracked panels. It concludes that the threshold is a dial between those two mistakes, that it has to be set per rule on a held-out set from the site's own cameras, and that the override rate is the number to watch after go-live. It is for anyone signing off a model before it watches a camera.

Rob Hickey · Chief AI Officer · Sep 28, 2026

Milk shelf with three empty slots flagged and the rack sections boxed, from a customer store camera

The shelf model on aisle 3 sent its first empty-facing alert at 8:20 am on a Monday, and the assistant manager walked over and found a full shelf of oat milk. It sent another at 8:40 am. By Thursday the alerts were arriving every twenty minutes and nobody was walking anywhere. On a stamping line in a different building, an inspection model watched a panel with a hairline crack along the fold go past at 3 pm and said nothing, and the panel was found two stations later by a person with a torch.

Both models had a good score on the day they were signed off. Both were wrong in a way the score did not describe, and the two ways are opposites.

The shelf model is wrong in the way that wastes a walk

Every alert the shelf model sends is a claim: there is a gap here. Precision is the share of those claims that were true. When the assistant manager walks to aisle 3 and finds oat milk, that alert was a false positive, and enough of them in a morning and the channel is muted by lunch.

Low precision is the cheap mistake per event and the expensive one over a week, because its real cost is trust. A person who has been sent to three full shelves stops going to the fourth, and the fourth is the one that was empty. On a store camera the usual causes are mundane: a shadow across the shelf edge at the hour the sun reaches the window, a promotional tag hanging in front of a full facing, a trolley parked where the model expects product.

Nobody at the store calls this precision. They call it the model crying wolf, and that phrase is the more accurate description of the damage.

The inspection model is wrong in the way that ships the defect

Recall is the other direction: of all the cracks that went past, how many did the model flag. The panel at 3 pm was a false negative, and a false negative on an inspection line has no alert attached to it, so nobody knows it happened until the part fails somewhere else.

This is the mistake that matters most on any camera watching for a defect or a hazard, and the one a demo never shows, because a demo is built from the frames where the thing was found. The frames where it was missed are, by definition, absent from the highlight reel. The line's own history is the only place they turn up, in the reject bin two stations downstream and in the rework log on Friday.

A hairline crack along a fold is small, low contrast and rare, which is the worst combination for recall. The model has seen a few dozen of them in training against a great many clean panels, and the safest thing for it to learn, in terms of its training score, is to pass everything.

F1 folds both into one number and hides which one you have

F1 is the harmonic mean of precision and recall, one figure that goes up when both do. It is a convenient thing to put on a sign-off sheet and a poor thing to make a decision from, because the shelf model and the inspection model can carry the same F1 while failing in opposite ways.

My own view is that F1 has no place on a go-live checklist. The two numbers it averages answer two different questions the site's operations lead already has opinions about, and averaging them throws the opinions away. Print precision and recall side by side for each class, and let the person who owns the rule say which one they would trade.

There are rules where the trade is obvious. A crack rule needs recall and can afford a person glancing at a clean panel. A shelf-gap rule for a morning list wants precision, because the list is read by someone with forty other things to do before the doors open.

Precision and recall are two readings of one dial

The model puts a number beside every detection, and the rule chooses a line above which a detection becomes an alert. Move the line down and recall rises, because more of the faint cracks clear it, and precision falls, because more shadows do too. Move it up and the reverse. Each is a reading of the same dial at one position, which is why quoting either without the threshold behind it is quoting half a fact.

The number beside a detection is a ranking of that frame against the others, and reading it as a probability is the mistake that leads to a rule like "alert above nine tenths" being copied from aisle 3 to the stamping line. The lesson on what the number beside a detection means goes through why the same score is near-certain for one class and a coin toss for another. The practical consequence is that each rule gets its own line, set by looking at what the two numbers do across the dial on that camera, and written down with the rule.

The held-out frames have to come from the same cameras

All of this is measured against a set of frames the model never trained on, labeled by a person, and drawn from the cameras the model will watch. A held-out set from a public benchmark scores the model on somebody else's shelves. A held-out set from the store's own aisle 3, across the morning sun and the evening restock, scores the model on the thing that will send the alerts.

Lexi puts the first box on each of those frames and a person checks every one before it counts as truth. The checking is where the score is decided, because a held-out set with the faint cracks left unlabeled will report a recall the line never gets. On the manufacturing lines we run, the sign-off set is a week of frames across all three shifts, and the night shift stays in.

After go-live the number to watch is the override rate

The sign-off score is the baseline, and it starts going stale the day the model is switched on. What replaces it is the rate at which people disagree with the model: frames coming back for review, corrections landing, the assistant manager marking the alert as a full shelf. When that rate rises on one camera and nowhere else, something on that camera has changed, and the score from sign-off day cannot tell you.

LexData takes the model through its whole life. You type what to look for, Lexi puts a box on every frame, and a person checks each label before anything trains on it. The model then watches the cameras the site already has, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime. The alert written as a sentence, with its severity and its cooldown, is where the threshold finally lives, and the override rate on that alert is the precision figure the store actually experiences.

The store, as it happens, keeps a laminated card of aisle numbers taped by the back door, because the alerts name the camera and the cameras were numbered by the installer.

See it on your own footage.

Start with your footage

More in Operations

Operations · 7 min read

Build or buy the layer that keeps a vision model accurate

Two engineers and a pilot can build a detector in a month. The review queue, the versioning and the rollout are what they are still building a year on.

Ayman Quadir · Sep 28, 2026

Operations · 7 min read

Benchmark the model on your own cameras before you believe a score

Two candidate models, two published scores, and a packaging line that only cares which one finds the torn label. The held-out week from your cameras decides.

Rob Hickey · Sep 28, 2026

Operations · 6 min read

Computer vision on the industrial HMI, the camera's verdict on the operator screen without the noise

The verdict as a state, the frame one click away, an alarm the operator acknowledges like any other, and everything else kept off the screen.

Andreas Ohrvall · Sep 28, 2026

LexData
LexData

Product

  • Platform
  • Industries
  • Use cases

Resources

  • Docs
  • The Field Guide
  • Blog
  • Why models drift

Industries

  • Energy & utilities
  • Oil & gas
  • Agriculture
  • Manufacturing
  • Insurance
  • Retail
  • Robotics

Company

  • About
  • Customers
  • Careers
  • Contact

Trust

  • Security
  • Privacy
  • Terms

Stay updated

What we learn running vision models in production.

See everything.
Miss nothing.

Stay updated

What we learn running vision models in production.

Terms of use & Privacy policy

© 2026 LexData Labs · All rights reserved