Skip to content
LexDataLexData
PlatformIndustriesCustomers
DocsThe Field GuideBlogWhy models drift
AboutCareersSecurityContact
Log inStart now
← All posts

Computer vision · 6 min read

Object detection in production, what a detector returns and what it learns from

For every frame off the yard camera the detector returns a list of boxes with a class and a score, and the score is a ranking rather than a probability.

Summary

This post explains what an object detector returns for a frame from a yard camera, how it learns from the boxes a person drew, and why loose boxes in the training set come back as loose boxes in production. It then follows the model from a notebook to a fleet of cameras and argues that its most useful output there is the list of frames it doubted. It is for engineers putting a first detector on live cameras.

Esdras Ntuyenabo · Engineer · Oct 4, 2026

Camera on a pole over an industrial yard at dusk, camera, fence and vehicle boxed, generated scene with detections from our model

The camera on the pole at the plant gate looks across the yard to the fence line, and at dusk on a Tuesday it sees a delivery van parked against the fence where nothing should be parked. The detector that watches that camera does not see a van. It produces a list, and everything that happens in the yard afterwards depends on what is in the list and how it was taught to make it.

A detector returns a list of boxes with a class and a score

For each frame the detector returns zero or more entries, and each entry is a class name, four numbers describing a rectangle, and a score. The van at the fence is one entry: vehicle, the rectangle around it, a score. The fence itself, if it is a class, is another. An empty yard at 4 am returns an empty list, which is a valid and useful answer.

Before that list is returned the raw output is tidied. The network proposes many overlapping boxes around every real object, most of them near duplicates, and a suppression step keeps the best box per object and discards the rest. What the yard sees is the tidy list, sampled from the camera about every two seconds, and the rule that says a vehicle at the fence after hours is worth a message to Slack reads that list and nothing else.

The score orders the detections and does not measure a chance

The number beside each box is the part most often misread. A box with a high score is more likely to be right than a box with a low one on frames that resemble the training set, and that ordering is what the score is for. It is not a probability. The same score can be near certain for the vehicle class on this camera and a coin flip for the person class at the far fence at 11 pm. The score was shaped by training on each class separately and by the suppression step afterwards.

The practical consequence is that thresholds are set per class, on the camera's own frames, by looking at what the detections above and below a candidate threshold actually contain. The lesson on reading the number beside a detection covers why a business rule written on the raw score is a rule that behaves differently on every camera.

A detector learns from the bounding box a person drew and nothing else

Training a detector is comparing its list with a person's list. For each labeled frame the person has drawn a bounding box around every vehicle and every person, with a class on each, and the training step measures how far the model's boxes are from those and nudges the weights to close the gap. Repeat over a few hundred frames from the camera at gate 2, many times over, and the model's list comes to resemble the person's.

That is the whole of the teaching. The class list is the model's entire vocabulary, so a forklift in the yard is either a class someone named or a thing the model has no word for. And the boxes are the model's entire idea of what each class looks like, which is where the next section comes from.

Loose boxes teach loose boxes

A box drawn a metre wide of the van on every frame teaches the model that a vehicle includes a metre of yard. A box that stops at a person's shoulders because a bollard hid the legs teaches the model that people are short. Whatever is consistently inside the labeled boxes becomes part of the class, and whatever is consistently left out becomes background, so a training set labeled loosely produces a model that returns loose boxes, and a training set labeled inconsistently produces a model that cannot decide.

The fix is a written rule for the box edge, agreed before labeling starts: the box hugs the visible extent of the object, occluded parts are not guessed, a vehicle includes its mirrors and excludes its shadow. Lexi proposes a box on every frame, and a person checks each one against that rule before anything trains on it. The checking is where the tightness is enforced, and labels checked this way come back at up to 99.9% accuracy.

The gate guard keeps a paper ledger of every vehicle through the gate, with the time. On Fridays the ledger is what the model's count is checked against, and it has settled more arguments than any dashboard.

A fleet of cameras breaks in ways a notebook never shows

In a notebook a detector is a folder of images and a score. In production it is a consumer attached to every camera in the yard, each stream watched in parallel, each frame answered before the next one arrives, on a server or on a runner beside the recorder. The things that go wrong are no longer metrics. They are cameras.

A camera is replaced after a storm with a unit that renders colour differently. The sun sets behind the fence at a new angle in October and the glare washes out the far corner. A skip is parked in the gap and a portion of the yard is quietly gone from every frame. None of these appears in a held-out test set, and all of them appear in the doubted frames within a week.

The lesson on what a vision model can and cannot see has one more limit that a notebook hides. The person at the far fence is a few dozen pixels tall, and below a certain size no detector finds anything, whatever the score on the benchmark said.

The most useful output in production is the list of doubted frames

LexData takes the yard model through its whole life. You type what to look for, Lexi puts a box on every frame, and a person checks each label before anything trains on it. The model then watches the cameras the plant already has, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime.

My own view is that a detector's most valuable output in production is the frames it was unsure about, and that the boxes it was sure of are the less interesting half of its work. The doubted frames are where the new camera, the October glare and the skip show up first. When the corrections on them cross the project's threshold a new version is trained, and the override rate on the gate camera, rising or falling week by week, is the number that says whether the model still matches the yard.

See it on your own footage.

Start with your footage

More in Computer vision

Computer vision · 7 min read

Five computer vision applications in production, on the cameras a site already owns

Detect, count, inspect, track and read are the five jobs a fixed camera can be given, and each is a different question with a different label behind it.

Ayman Quadir · Oct 4, 2026

Computer vision · 6 min read

Computer vision projects worth building on the cameras you already have

A plant, a utility, a grower, a store and a warehouse each have a project that starts on an existing camera, produces a decision, and has someone to act on it.

Ayman Quadir · Oct 4, 2026

Computer vision · 6 min read

How to choose an object detection model architecture for a camera on your own site

Start from the decision the plant has to make, then the box beside the recorder the model has to fit, and only then the family. The benchmark comes last.

Andreas Ohrvall · Oct 4, 2026

LexData
LexData

Product

  • Platform
  • Industries
  • Use cases

Resources

  • Docs
  • The Field Guide
  • Blog
  • Why models drift

Industries

  • Energy & utilities
  • Oil & gas
  • Agriculture
  • Manufacturing
  • Insurance
  • Retail
  • Robotics

Company

  • About
  • Customers
  • Careers
  • Contact

Trust

  • Security
  • Privacy
  • Terms

Stay updated

What we learn running vision models in production.

See everything.
Miss nothing.

Stay updated

What we learn running vision models in production.

Terms of use & Privacy policy

© 2026 LexData Labs · All rights reserved