Computer vision · 6 min read
What YOLO does in one pass and what that costs
One look at the frame is why a single-pass detector fits the box beside the recorder, and why it misses the small cluttered things a two-stage model finds.
Summary
This post explains what a single-pass detector does with a frame, how that differs from a two-stage model that proposes regions and looks again, and where each wins: the single pass on a runner beside the recorder, the two-stage model on small cluttered objects. It concludes that the family name matters far less than the frames the model was trained on. It is for engineers choosing a detector for a fixed camera.
Andreas Ohrvall · CTO · Sep 29, 2026

Bottling line with bottles queued under the fill head, one missing its cap, bottle and cap boxed, generated scene with detections from our model
The camera over the fill head on the bottling line produces a frame every couple of seconds, and every so often one of the bottles in it has no cap. The model that catches that runs on a small box in the cabinet beside the recorder, and it has to finish with one frame before the next one arrives. That constraint, more than any benchmark, is why the model on that box is a single-pass detector.
The name YOLO stands for a design decision: look at the frame once, and predict every box in it from that one look. Most of what the family does well, and most of what it misses, follows from that decision.
A single pass divides the frame into a grid and predicts everything at once
A single-pass detector lays a grid over the frame and, for each cell, predicts whether an object is centred there, what class it is, and where its box is. All of those predictions come out of one forward pass through the network, together, for the whole frame. There is no separate step that asks "where might objects be" before deciding what they are.
The output is a long list of candidate boxes, most of them near-duplicates of each other around every real object, each with a score. A tidy-up step keeps the best box per object and drops the rest, and the result is the boxes the line sees at 2 am: bottle, bottle, bottle, bottle with no cap.
Because everything happens in one pass, the cost per frame is roughly fixed. A frame with one bottle and a frame with forty cost about the same, which is a property a fixed camera on a line can plan around.
A two-stage R-CNN proposes regions first and looks again
The other family works in two steps. An R-CNN style model first proposes regions of the frame that might hold something, then runs a second stage on each region to classify it and tighten its box. The second look is the point. Each candidate gets the network's attention at a scale that suits it, which is why two-stage models tend to do better on small objects packed close together, where the single pass's grid cell is trying to describe three things at once.
The price is that the cost per frame varies with the frame. Forty proposals cost forty second-stage passes, and a cluttered frame takes longer than an empty one. On a server scoring a drone survey after landing that is fine. On a box beside a recorder that has to keep up with the camera, it is the reason the box falls behind.
On a runner beside the recorder the single pass wins on cost
The cabinet box has a fixed budget of compute and a fixed stream of frames, sampled about every two seconds. A single-pass detector with a modest backbone fits inside that budget with room to spare, and its fixed cost per frame means it fits on the busy Monday changeover as well as on a quiet Thursday night. That predictability is worth more on a runner than a few points of accuracy on a benchmark, because a model that keeps up on average and falls behind at the busiest hour is a model that misses the busiest hour.
Where the model runs is a deployment decision. The deployment doc describes how a model leaves the platform for the box, exported as PT, ONNX or TorchScript through a preset that describes the device, a Jetson, a Raspberry Pi, a GPU server, rather than the architecture. The preset is the place to say the model is going to a small box, and the export is sized to it.
Small cluttered objects are where the single pass misses
The fill head is a good case for the single pass because bottles are large in the frame and evenly spaced. The same detector pointed at a tray of caps, hundreds of small round things touching each other, does worse, because several caps land in one grid cell and the cell can only describe so many. The misses cluster in the crowded corners of the frame, and they are the kind of miss a two-stage model, with its second look at each proposal, makes less often.
Slicing the frame into tiles is the usual remedy for a single-pass model on small objects, at a compute cost that multiplies with the tiles. Whether that fits on the Jetson in the cabinet is an arithmetic question. The answer decides whether the small-object camera gets the single pass with tiles, a two-stage model on a bigger server, or a camera moved closer so the caps are no longer small.
The line's own name for a bottle with no cap, for what it is worth, is "a bald one", and the class list uses the plain phrase because the operators do.
The family name matters less than the frames it was trained on
The YOLO family has had many versions, each faster or a little more accurate than the last on the public benchmarks, and the differences between recent versions on a fixed camera are small next to the difference the training frames make. A model from any recent version trained on a few hundred checked frames from the fill head will beat the newest version trained on a public set, every time, because the public set has no fill head in it. The lesson on what a vision model can and cannot see puts it plainly: vision is good at what is visible and consistent, and the fill head's frames are what make the missing cap consistent to the model.
My own view is that the version number stopped being interesting somewhere around the point where the family passed its tenth release. Pick a version that exports to the box, train it on the line's frames, and put the remaining attention into the review queue.
The export fits the box and the corrections keep the model current
LexData takes the fill-head model through its whole life. You type what to look for, Lexi puts a box on every frame, and a person checks each label before anything trains on it. The model then watches the cameras the plant already has, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime.
On the box beside the recorder that last step matters most: the new version arrives sized for the same device, and the bald bottle the model doubted at 2 am is a frame in the set it was trained from. The single pass keeps looking once, and what it sees in that one look is decided by the frames a person checked.
See it on your own footage.
Start with your footageMore in Computer vision

Computer vision · 6 min read
What the big vision models are for, and what runs on the camera
The big models find the frames, draw the first boxes and tighten them, and a person checks. The small model trained on those labels runs beside the recorder.
Andreas Ohrvall · Sep 29, 2026

Computer vision · 6 min read
Why a good detector can still be a bad counter
The queue camera finds every shopper on every frame and still gets the count wrong. Swaps, lost tracks and jitter are tracker mistakes, and review fixes one.
Stephen Biswas · Sep 29, 2026

Computer vision · 6 min read
What semantic segmentation labels and what it costs
Drivable path, corrosion area and crop rows are questions about pixels, and the answer is a class for every one. Painted labels cost more than boxes.
Esdras Ntuyenabo · Sep 29, 2026