Computer vision · 6 min read
DETR explained, what the detection transformer changed about finding objects
Detection as set prediction, a fixed set of queries instead of anchors, no suppression step, and the small-object and slow-training problems later fixed.
Summary
This post explains the detection transformer, DETR, by what it changed: detection as set prediction with a fixed set of learned queries, one-to-one matching against the labels, and no anchors or suppression step. It covers the backbone that still does the looking, where the original struggled on small objects and training time, what the fast descendants fixed, and what any of it means for a model on a runner beside a recorder. It is for engineers who keep meeting the name and want the mechanism.
Finn Ellingwood · Engineer · Sep 28, 2026

Street scene from a vehicle camera, eleven traffic signs, pedestrians and vehicles boxed, from a customer perception run
A frame from a vehicle camera at a junction holds eleven traffic signs, four pedestrians and six cars, and a detector has to return one box for each of them and none for anything else. For years the way to do that was to propose a great many boxes at every position in the frame and then throw most of them away. In 2020 a paper called Detection Transformer, DETR for short, proposed doing something else: ask the model for a set of boxes directly, and train it so the set is right.
That sounds like a small change of phrasing. It removed two of the most fiddly parts of a detector, and it introduced two problems of its own that took the field a few years to fix.
Detection becomes set prediction
A conventional detector produces a great many candidate boxes per frame, from preset anchor shapes at every position, and a final step called non-maximum suppression keeps the best candidate for each object and discards the overlapping rest. Both parts need tuning. The anchors have to suit the objects' sizes and shapes, and the suppression threshold decides whether two pedestrians standing close together are two detections or one.
DETR replaces both with a fixed set of slots. The model carries a hundred or so learned queries, each of which produces one box and one class, where the class can be "nothing". At training time the model's set is matched one to one against the labeled set by a bipartite matching, the Hungarian algorithm, so that each labeled sign is paired with exactly one query and every leftover query is trained to say nothing. There is nothing to suppress afterwards, because the model was trained to produce one box per object in the first place.
A ResNet backbone still does the looking before the transformer
The transformer does not see pixels. A convolutional backbone, in the original paper a ResNet, runs over the frame first and produces a grid of feature vectors, one per patch of the image, each summarising what is in that patch. The transformer works on that grid. A positional encoding is added so the model has a record of where each patch sits, because attention on its own has no notion of position.
That matters for two practical reasons. The backbone is where most of the compute goes, and swapping it for a smaller one is the first lever when the model has to fit on a runner. And the backbone is where the general picture knowledge lives, so a backbone pretrained on the world's images and fine-tuned on frames from the junction camera is what makes the transformer's queries have something to attend to.
Object queries attend to the whole frame and settle on one object each
The encoder takes the grid of patch features and lets every patch attend to every other, so the feature for the patch containing a sign can absorb the context of the pole, the road and the car beside it. The decoder takes the learned queries and lets each attend across the encoded frame, refining, layer by layer, towards one object. By the last layer each query has either found an object and predicts its box and class, or has settled on "nothing".
The queries in DETR are not assigned regions in advance. They learn, over training, a loose specialisation, some tending towards large objects in the centre, some towards small ones at the edges, and the matching at training time is what pushes them apart. It is a strange thing to watch on a junction frame: the same query finds the near-side sign in one frame and a pedestrian in the next.
An aside: the fixed number of queries is a ceiling on how many objects the model can return per frame. A hundred is plenty for a junction and not for a crowd or a dense shelf, and the number is a setting rather than a law.
The original was slow to train and weak on small objects
Two problems came with the elegance. The first was training time. Because the matching in DETR between queries and objects is learned from scratch, and every query starts out attending everywhere, the original took many times longer to converge than a conventional detector on the same set. The second was small objects. Attention across a whole frame at the backbone's resolution loses the eleven signs at the far end of the junction, which are a few patches each, and the original's results on small objects were well behind the conventional models.
Both were addressed by the descendants. Deformable attention lets each query attend to a small set of sampled points around a reference rather than to every patch, which cuts the cost and lets the model use higher-resolution features where the small signs live. Better initial queries, seeded from the encoder's own proposals rather than learned from nothing, shortened training to something comparable with the conventional family. The versions fast enough for a live feed are built on those two changes.
What it means for a model on the runner beside the recorder
For an engineer putting a detector on a box beside a camera, the transformer family means a simpler pipeline: no anchor design, no suppression threshold to tune per site, and a model whose output is already a clean set of boxes.
It also means a heavier fixed cost at small sizes. Attention has a floor that shrinking the backbone does not remove, and a CPU-only runner often does better with a convolutional model.
The deployment doc covers what the runner needs regardless of family: an export for the device, headroom for the enclosure, and a test on the site's own frames. My own view is that the choice of family is a smaller decision than the choice of which frames the model has seen, and that most teams would do better spending a week on the labeled set than a week on the architecture.
The mechanism does not change what the camera can answer
A transformer detector finds what is visible and consistent, like any detector, and DETR is no exception. The lesson on what vision can see applies unchanged. A sign that is legible to the model is one it was trained on in enough conditions, and a sign that is a few pixels at the far end of the junction is a camera question before it is a model question.
What the model does on the live feed is also unchanged. It watches the camera, the frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime. The queries, the matching and the attention are inside the box. What the box learns from is still the labeled frames, checked by a person, from the road it is going to drive.
See it on your own footage.
Start with your footageMore in Computer vision

Computer vision · 6 min read
When the outline is the answer and a box will not do
A weld defect whose area decides the rework and a part a robot has to grasp both need a mask. Here is what the mask costs to label and when a box is enough.
Esdras Ntuyenabo · Sep 28, 2026

Computer vision · 6 min read
Reading a training run before trusting the model
The loss curve says the run finished. The gap to validation, the recall on the crack class and the frames behind the score say whether the line can trust it.
Rob Hickey · Sep 28, 2026

Computer vision · 6 min read
What mAP measures and what it cannot tell you about your cameras
Mean average precision is built from per-class curves at an overlap line. A high score on the held-out set says nothing about the fence camera at 2 am.
Stephen Biswas · Sep 28, 2026