Skip to content
LexDataLexData
PlatformIndustriesCustomers
DocsThe Field GuideBlogWhy models drift
AboutCareersSecurityContact
Log inStart now
← All posts

Computer vision · 6 min read

How to evaluate object detection models fairly, and why two mAP numbers are rarely comparable

A vendor's score and yours are answers to different questions. The test set, the calculation and the inference conditions each move the number on their own.

Summary

This post explains why a detector's published score and the score a utility measures on its own drone flights are rarely the same number, through the three choices hidden in one figure: the test set, the calculation and the inference conditions. It describes how to build an independent test set from your own flights and argues that after go-live the score that holds is the operator correction rate. It is for teams comparing a vendor's number with their own.

Rob Hickey · Chief AI Officer · Oct 4, 2026

Transmission tower over a country road from a drone, six insulator strings and one corrosion mark boxed, from a customer inspection run

The utility's drone programme received two numbers for the same insulator defect model in the same week in March. The first came with the model, a mean average precision measured on the supplier's own test set. The second was measured by the utility's engineer on frames from its own flights along the line, and it was lower by a margin that started an argument. Neither number was wrong. They were answers to different questions, and the difference between the questions is most of what there is to know about evaluating a detector.

An object detection score is three choices wearing one number

A mean average precision figure is produced by a test set, a calculation and a set of inference conditions, and each of the three can be chosen in ways that move the number without changing the model. A score reported without all three is a number without a unit, and the two figures the utility received in March had none of the three in common.

The test set decides which frames the model is judged on. The calculation decides how much a box has to overlap the labeled one to count, and how the classes are averaged. The inference conditions decide the resolution the frames were fed at, the threshold below which detections were discarded, and whether the frames were tiled. Two scores agree only when all three agree, which between a supplier and a customer they almost never do.

The test set decides most of the number

The supplier's test set was flown in clear weather over towers of one design, with the camera at the distance that suits the model, and the labels drawn by the team that trained it. The utility's flights include tower 17 at dusk with the sun behind it, the corrosion mark on hardware of a design the model has never seen, and a pilot who flew wider than usual because of the wind. The model is being asked a harder question on the utility's frames, and a lower score is the honest answer to it.

There is a second way the test set inflates a number. When the test frames come from the same flights as the training frames, adjacent frames are near duplicates, and the model is being tested on scenes it has already seen. Hold out whole flights, by date, and the number drops to something closer to what the next flight will show.

The pilots name the towers, for what it is worth, and the tower with the persistent corrosion mark has had a name since the second season.

The calculation changes the number without touching the model

The overlap threshold is the first lever. A box counts as correct when it overlaps the labeled box by at least a chosen fraction, and a score at half overlap is always higher than one at a stricter fraction, because a loose box still counts. The COCO benchmark reports an average across a range of thresholds, and a supplier quoting the lenient end is not lying, only choosing.

Averaging across classes is the second lever. A model that finds every insulator and misses most of the corrosion has a respectable average and is failing at the class the utility cares about. The per-class figures are the ones to ask for, and the class that matters is usually the rare one.

The inference conditions have to match the runner

A score measured at full resolution on a GPU server describes a model that will run at reduced resolution on a Jetson beside the recorder only if the runner gets the same frames. It rarely does. Shrinking the frame to fit the device shrinks the corrosion mark below the size any detector resolves, and the score on the runner is a different number from the score on the server. The threshold is the same kind of choice: a low threshold raises recall and the count of false boxes together, and the score depends on where it was set. Fix the resolution, the threshold and the tiling to what the field deployment will use, and measure there.

Build the independent test set from your own flights

The number the utility can use is measured on a test set built from its own flights, held out by flight date, labeled to the same written box rule as the training set by people who did not train the model, and never trained on. It is refreshed as the seasons turn, because a test set from summer flights says little about November, and the most recent flights are the ones held out.

Two things make that set worth its cost. The lesson on retraining without starting over puts them plainly: hold out the future rather than a random slice, and know what better means before a version is promoted. On the drone programme better means recall on the corrosion class at the field threshold, on the held-out flights, with the false boxes per tower held below a number the review team agreed to.

After go-live the score that holds is the correction rate

A test set is a snapshot of the recent past.

Once the model is watching every flight, the score that keeps being measured is the operator correction rate: how many frames come back for review, how many corrections land, how often the reviewer overrides the model's box. When that rate rises on one line section or after one storm, the model has met something the test set did not contain, and the corrections are the frames the next version trains on.

LexData takes the insulator model through its whole life. You type what to look for, Lexi puts a box on every frame, and a person checks each label before anything trains on it. The model then watches the footage from every flight, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime.

The threshold that shaped the vendor's number is a choice the review team makes for itself on its own frames, and the lesson on reading the number beside a detection explains why the same threshold behaves differently on every camera and every class.

My own view is that a supplier's benchmark is useful for exactly one decision, whether the model is worth the cost of running your own test, and that no purchase should turn on it beyond that.

See it on your own footage.

Start with your footage

More in Computer vision

Computer vision · 7 min read

Five computer vision applications in production, on the cameras a site already owns

Detect, count, inspect, track and read are the five jobs a fixed camera can be given, and each is a different question with a different label behind it.

Ayman Quadir · Oct 4, 2026

Computer vision · 6 min read

Computer vision projects worth building on the cameras you already have

A plant, a utility, a grower, a store and a warehouse each have a project that starts on an existing camera, produces a decision, and has someone to act on it.

Ayman Quadir · Oct 4, 2026

Computer vision · 6 min read

How to choose an object detection model architecture for a camera on your own site

Start from the decision the plant has to make, then the box beside the recorder the model has to fit, and only then the family. The benchmark comes last.

Andreas Ohrvall · Oct 4, 2026

LexData
LexData

Product

  • Platform
  • Industries
  • Use cases

Resources

  • Docs
  • The Field Guide
  • Blog
  • Why models drift

Industries

  • Energy & utilities
  • Oil & gas
  • Agriculture
  • Manufacturing
  • Insurance
  • Retail
  • Robotics

Company

  • About
  • Customers
  • Careers
  • Contact

Trust

  • Security
  • Privacy
  • Terms

Stay updated

What we learn running vision models in production.

See everything.
Miss nothing.

Stay updated

What we learn running vision models in production.

Terms of use & Privacy policy

© 2026 LexData Labs · All rights reserved