Skip to content
LexDataLexData
PlatformIndustriesCustomers
DocsThe Field GuideBlogWhy models drift
AboutCareersSecurityContact
Log inStart now
← All posts

Operations · 7 min read

Benchmark the model on your own cameras before you believe a score

Two candidate models, two published scores, and a packaging line that only cares which one finds the torn label. The held-out week from your cameras decides.

Summary

This post takes two candidate detectors with published benchmark scores and compares them the only way that matters, on a held-out week of frames from a packaging line's own cameras. It explains IoU, precision and recall in plain words, reads the result per class, and concludes that a published score was never a measurement of your objects. It is for teams choosing between models before one goes live.

Rob Hickey · Chief AI Officer · Sep 28, 2026

Packaging line with cartons and labels boxed as they pass the scanner, generated scene with detections from our model

Two candidate models arrived on the packaging line in the same week, each with a score from a public benchmark, and one of the scores was a little higher. The line lead's question was simpler than either number: which one finds the torn label on a carton at station 2, at line speed, under the sodium lights, before the carton reaches the scanner and fails the read.

Neither benchmark had ever seen a carton from this line. Neither had a class called "torn label". The higher score was true, and it was true about something else.

The published score was measured on somebody else's objects

A benchmark score is a measurement of a model on a fixed set of pictures with a fixed list of classes, and it is a fair way to rank models against each other on that set. It says nothing about a packaging line, because the set holds no packaging lines, and its classes, however many there are, do not include the one the rule is written around.

Two things follow. The first is that a higher published score is weak evidence about the line, worth about as much as the model's family name. The second is that the only score that means anything here has to be produced on the line's own frames, with the line's own classes. The labels come from a person who knows what a torn label looks like at station 2 and what a fold in the film looks like next to it.

Producing that score takes a week of footage and a few days of labeling, which is less than the cost of picking wrong.

IoU decides whether a box counts as found

Before precision or recall can be counted, something has to decide whether a predicted box matches a labeled one. That something is intersection over union: the area the two boxes share, divided by the area they cover between them. Two identical boxes score one. A predicted box that overlaps the label by half of their combined area scores a half, which is the usual line for calling the detection a match.

On a carton label the line matters. A box that covers the carton instead of the label overlaps the label completely and covers far more besides, so its IoU is low and it is scored as a miss even though the model found roughly the right place. That is the right verdict for a rule that will crop the label and read it, and the wrong verdict for a rule that only needs to know a carton went past. Choose the IoU line to match what the rule will do with the box, and use the same line for both models.

Precision and recall are counted per class, in plain words

With matches decided, precision is the share of the model's boxes that matched a label, and recall is the share of the labels that a box matched. A model with high precision and low recall on "torn label" is one that rarely cries wolf and misses most tears. A model the other way round finds most tears and sends the line lead to a lot of intact cartons at station 2.

Read both numbers per class. The line has four classes on its list: carton, label, torn label and open flap, and the first two appear on every frame while the last two appear a few times an hour. An average over the four is mostly a score on cartons, and both candidates will find cartons. The torn-label row is where they differ, and where the published score had nothing to say.

Precision and recall on that one row, at the threshold the rule will run at, is the whole comparison. Everything else is context.

The held-out week comes from the cameras the model will watch

The frames the two models are scored on have to come from the line's own cameras and from a stretch of time neither model trained on. A week is a good unit because it holds the things a day hides: the Monday changeover, the night shift with half the lights, the Thursday when the film supplier's new roll arrived a shade glossier than the old one. The lesson on how much footage you actually need makes the case that distinct conditions are the unit, and a week across three shifts is a cheap way to collect a lot of them.

Lexi puts a first box on every frame of that week and a person checks each one. The person's attention goes to the rare classes, because a torn label the labeler missed becomes a torn label the better model is punished for finding. Then the week is locked, and neither model sees a frame of it until the comparison runs.

I would rather see a model score honestly low on a hard week than high on a tidy one. The tidy week is the one that gets copied into the sign-off deck, and the hard week is the one the line actually runs.

Both models run on the same frames at the same operating point

The comparison is only fair if everything but the model is held still. Same frames, same labels, same IoU line, same way of matching boxes to labels, and, the one most often skipped, the same threshold on the number each model puts beside a detection. Those numbers are rankings, and they are not on a shared scale between two models, so a fixed line of nine tenths means different things to each. The lesson on what the number beside a detection means goes into why.

The fair method is to sweep the threshold for each model and compare them at the operating point the rule will use, which for a torn-label rule is the point where recall is high enough that the line lead trusts the misses are rare.

Write down the whole protocol before running it, in a paragraph, so that when the second candidate's vendor asks why they lost, the answer is a document rather than a memory.

The scanner at station 2, for what it is worth, still beeps on every carton it reads, and the operators stopped hearing it years ago.

A close result means the frames decide from here on

When the two models land within a few points of each other on the torn-label row, the choice stops being about the model. Either one will be retrained on the line's own corrections within a month, and after that the frames it has learned from matter far more than the architecture it started with. Pick the one that fits where it will run, on the server by the line or on a runner beside the recorder, and put the labeling effort into the week after go-live.

LexData takes the chosen model through its whole life. You type what to look for, Lexi puts a box on every frame, and a person checks each label before anything trains on it. The model then watches the cameras the line already has, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime. The held-out week stays locked as the reference, so each new version is scored against the same Thursday with the glossy film, and the number on the torn-label row is comparable from one version to the next.

See it on your own footage.

Start with your footage

More in Operations

Operations · 7 min read

Build or buy the layer that keeps a vision model accurate

Two engineers and a pilot can build a detector in a month. The review queue, the versioning and the rollout are what they are still building a year on.

Ayman Quadir · Sep 28, 2026

Operations · 6 min read

Computer vision on the industrial HMI, the camera's verdict on the operator screen without the noise

The verdict as a state, the frame one click away, an alarm the operator acknowledges like any other, and everything else kept off the screen.

Andreas Ohrvall · Sep 28, 2026

Operations · 7 min read

A model trained on your warehouse beats one trained on the internet

A general detector calls the forklift a truck and the pallet nothing at all. The held-out week from your own cameras is the only score that counts.

Ayman Quadir · Sep 28, 2026

LexData
LexData

Product

  • Platform
  • Industries
  • Use cases

Resources

  • Docs
  • The Field Guide
  • Blog
  • Why models drift

Industries

  • Energy & utilities
  • Oil & gas
  • Agriculture
  • Manufacturing
  • Insurance
  • Retail
  • Robotics

Company

  • About
  • Customers
  • Careers
  • Contact

Trust

  • Security
  • Privacy
  • Terms

Stay updated

What we learn running vision models in production.

See everything.
Miss nothing.

Stay updated

What we learn running vision models in production.

Terms of use & Privacy policy

© 2026 LexData Labs · All rights reserved