Skip to content
LexDataLexData
PlatformIndustriesCustomers
DocsThe Field GuideBlogWhy models drift
AboutCareersSecurityContact
Log inStart now
← All posts

Computer vision · 6 min read

Real-time object detection models, transformer or convolutional, and which size for the runner

A DETR-style detector drops the anchors and the suppression step. The size to run is decided by the box beside the camera, which the leaderboard cannot see.

Summary

This post compares the two families of fast detectors, convolutional single-pass models and the transformer detectors descended from DETR, by what each changes about finding an object and what the two mAP figures actually measure. It concludes that the size to deploy is set by the hardware beside the camera and by frames from that camera, and that the leaderboard answers neither. It is for engineers choosing a detector for a yard camera and a small box in a cabinet.

Rob Hickey · Chief AI Officer · Sep 28, 2026

Camera on a pole over an industrial yard at dusk, camera, fence and vehicle boxed, generated scene with detections from our model

Camera 2, on the pole over the yard at the depot, has a small computer in the cabinet at its base, and the question is which detector to put on it to box the vehicles crossing the gate. The engineer has two shortlists open. One is the family of single-pass convolutional detectors that has been the default for a decade. The other is the newer family of transformer detectors, and the leaderboard says the transformer is a point or two ahead.

The leaderboard was measured on a benchmark set, on a data-centre GPU, at a resolution the yard camera does not use. The box in the cabinet is in none of those.

A convolutional detector slides local filters over the frame in one pass

The convolutional family works the way it sounds. Small filters slide over the frame and build up, layer by layer, a map of features at every position. A detection head reads that map and, at every position, proposes boxes from a set of preset shapes called anchors, or in the newer versions from the position directly. The result on camera 2 is many overlapping proposals for each vehicle, and a final step, non-maximum suppression, keeps the best one and discards the rest.

It is fast because every operation is local and regular, which is exactly what a small GPU or a CPU does well. It is also why the family scales cleanly from tiny to large: the same design with fewer or more filters.

A DETR-style detector drops the anchors and the suppression step

The transformer family changes the framing of the problem. Instead of proposing boxes at every position, a DETR-style model carries a fixed set of learned queries, each of which attends to the whole frame and settles on one object, or on nothing. Detection becomes set prediction: the model outputs a set of boxes and is trained to match them one to one against the labeled objects. There are no anchors to design and no suppression step, because the model has already been trained to produce one box per object.

Attention across the whole frame is the other difference. A query looking at the vehicle at the gate can use the context of the fence, the gatehouse and the road, which helps when the vehicle is partly hidden. The original version paid for that with slow training and weak results on small objects. The descendants that are fast enough for a live feed fixed both: attention restricted to a few sampled points, better initial queries, and training that converges in a fraction of the time.

On the yard camera the practical difference shows up at the gate, where two vehicles overlap. The convolutional model proposes boxes on both and relies on the suppression step to keep two. The transformer model's queries have to settle on two, and the training taught them to.

The two mAP figures measure different things

Every leaderboard reports mean average precision, and there are two common versions. The first counts a detection as correct when its box overlaps the labeled box by at least half, and averages precision over the recall range for each class. The second averages that figure across a range of stricter overlap thresholds, from half up to nearly exact. The first says whether the model found the vehicle. The second says how precisely it boxed it.

For the rule at gate 2, the first matters more: a vehicle found with a slightly loose box still crosses the line. For a rule about where the vehicle is in the lane, the second matters. Reading the wrong figure is a common way to pick a model for a job it will not do.

The score beside each detection feeds both figures, and the lesson on what that number means is worth a read before anyone sets a threshold from a benchmark. The score ranks detections against each other on frames like the training set, and the yard is not the training set.

The size is decided by the box in the cabinet

Both families come in sizes, from a model small enough for a single-board computer to one that needs a data-centre GPU. The choice is not the largest that fits. It is the smallest that answers the question at the rate the question needs. The yard camera is sampled every couple of seconds, so a model that runs in a fraction of a second on the cabinet's hardware has room to spare, and a larger one that runs in two seconds has none.

The deployment doc makes the point that the cabinet is not the bench. It is sealed, it gets warm in July, and a model that runs at a comfortable rate on the desk throttles inside it. Video decode is often the bottleneck before the model is, and a camera that is swapped for a higher-resolution one changes the image in ways the model notices first. The size to deploy is the one that has been run in the cabinet, on the yard's own frames, for a week.

An aside: the transformer family scales less evenly than the convolutional one. The small sizes give up more than a proportional share of accuracy, because attention over a whole frame has a fixed cost that shrinking the model does not remove. On a CPU-only box the convolutional family usually wins on that alone.

Evaluate on frames from the yard, on the hardware in the cabinet

The number that decides is not on any leaderboard. It is the model's precision and recall on a held-out set of frames from camera 2, across a week, with the night and the rain in it, run on the box in the cabinet at the rate the gate rule needs. Two candidate models, one from each family, trained on the same labeled frames and measured that way, give an answer the benchmark cannot.

My own view is that the family matters less than most engineers expect and the frames matter more. A convolutional model fine-tuned on a thousand reviewed frames from the yard beats a transformer model that has only seen the benchmark, and the reverse is also true. The architecture sets the ceiling; the frames decide where under it the model lands.

The model on the runner keeps changing after the choice is made

Whichever family goes on the cabinet, it does not stay the version that was chosen. The model watches gate 2, the frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one on the runner with no downtime. The choice between families is made once. The choice of which frames the model has seen is made every week, and it is the one that moves the number the depot cares about.

The model is exported for the device it runs on, described as the device rather than the architecture, and tested in the cabinet before it replaces anything. The engineer's two shortlists are still open in a tab somewhere. The set of reviewed frames from the yard is the document that matters now.

See it on your own footage.

Start with your footage

More in Computer vision

Computer vision · 6 min read

When the outline is the answer and a box will not do

A weld defect whose area decides the rework and a part a robot has to grasp both need a mask. Here is what the mask costs to label and when a box is enough.

Esdras Ntuyenabo · Sep 28, 2026

Computer vision · 6 min read

DETR explained, what the detection transformer changed about finding objects

Detection as set prediction, a fixed set of queries instead of anchors, no suppression step, and the small-object and slow-training problems later fixed.

Finn Ellingwood · Sep 28, 2026

Computer vision · 6 min read

Reading a training run before trusting the model

The loss curve says the run finished. The gap to validation, the recall on the crack class and the frames behind the score say whether the line can trust it.

Rob Hickey · Sep 28, 2026

LexData
LexData

Product

  • Platform
  • Industries
  • Use cases

Resources

  • Docs
  • The Field Guide
  • Blog
  • Why models drift

Industries

  • Energy & utilities
  • Oil & gas
  • Agriculture
  • Manufacturing
  • Insurance
  • Retail
  • Robotics

Company

  • About
  • Customers
  • Careers
  • Contact

Trust

  • Security
  • Privacy
  • Terms

Stay updated

What we learn running vision models in production.

See everything.
Miss nothing.

Stay updated

What we learn running vision models in production.

Terms of use & Privacy policy

© 2026 LexData Labs · All rights reserved