Operations · 6 min read
YOLO versus vision language models, decided by the failure you can afford
A detector gives the same box for the same frame, every two seconds, forever. A language model answers a question nobody wrote a class for. The dock needs both.
Summary
This post separates the two questions a loading dock asks of its cameras, one a rule that must fire the same way every time and one a search nobody wrote a class for, and matches each to a fixed-class detector or a vision language model by the failure the dock can tolerate. It concludes that the language model finds the frames and the detector watches the camera, and that a language model should not sit on the path to an alert. It is for the teams deciding what runs on the recorder.
Rob Hickey · Chief AI Officer · Oct 1, 2026

A loading dock from a mounted camera, trucks at the bays and a forklift with a pallet boxed, generated scene with detections from our model
The dock manager at the distribution centre left a sticky note with two questions on the vision engineer's monitor. The first was "is there a forklift in the walkway by bay 3", and she wanted it answered every couple of seconds, every day, with an alert. The second was "find me every frame from last month where a pallet was stacked on its side", and she wanted it answered once, by Friday.
They look like the same kind of question. They are answered by different kinds of model, and choosing one for both is how the dock ends up with an alert that fires late and a search that never finishes.
A fixed-class detector gives the same box for the same frame
A detector in the YOLO family is trained on a fixed list of classes: forklift and person, pallet and truck. Show it a frame and it returns boxes with a class and a number for each, in a fixed format, in a fixed time, and the same frame returns the same boxes tomorrow. That predictability is the whole point. A rule can be written against it: a forklift box overlapping the walkway polygon by bay 3, for longer than the window, is an alert.
It knows nothing outside its list. A pallet on its side is a pallet, if it is anything, and the detector has no way to say "on its side" unless someone labeled that as a class. Its strength and its limit are the same fact.
Pose estimation and boxes are outputs a rule can consume
The detector family includes more than boxes. Pose estimation returns the joints of each person as points, so a rule can say the worker at bay 3 is bending into the walkway rather than standing beside it. Segmentation returns the outline, so the pallet's area is a number. Each is a structured output: fields with types, produced the same way every time, that a rule or a PLC can read without interpretation.
That structure is what an alert path needs. The alert is a rule written as a sentence, with a severity and a cooldown, approved before it goes live, and it compiles against the model's classes. It cannot compile against a paragraph.
A language model answers a question nobody wrote a class for
A vision language model takes the frame and a question in words and gives an answer in words. "Is any pallet in this frame stacked on its side?" gets a yes or no with a reason. "What is unusual about bay 3?" gets a sentence. Nobody had to define a class for a sideways pallet, and next week's question about a torn shrink wrap needs no new model.
The cost of that flexibility is everything the detector had. The answer is prose, and prose has to be parsed before a rule can use it. The same frame can get a differently worded answer twice. The model is large, slow, and expensive per frame in a way that rules it out for every camera every two seconds. And it can be confidently wrong in a way that reads as reasonable, which is the worst failure for anything that fires an alert.
The failure you can afford decides which one runs
For the walkway by bay 3, the failure the dock cannot afford is a forklift that was there and was not flagged. It needs a model that runs on every sampled frame from every camera, in a fixed time, with an output a rule can read, and whose misses can be measured and reduced by training on the frames it missed. That is the detector.
For the sideways pallet search, the failure the dock can afford is a few wrong frames in a list a person will look through anyway. It needs a model that can be asked a question nobody planned for, over frames already recorded, once. That is the language model. My own view is that a language model should not sit on the path to an alert at all, whatever its accuracy on a benchmark, because an alert path has to fail in ways a person can see and count, and prose does not.
The field guide's lesson on what a vision model can and cannot see puts the same distinction as three questions: is there a person in the frame, how many crossed the line, and is this weld going to fail. The first two are detector questions with a tracking costume on the second. The third is a question for a person with the frames in front of them, and the language model's proper job is getting those frames in front of them.
The language model finds the frames and the detector watches the camera
The sticky note's two questions turn out to feed each other. Ask the language model about last month's footage and it returns the frames with a sideways pallet, with the frames behind the answer so a person can check. That is what LexInsight is for: a question of your footage, answered with the evidence.
Those frames are then the seed of something else. If the dock decides a sideways pallet is worth an alert, the returned frames go to a person, who draws the boxes, and a class now exists. From there the detector takes over, and retraining on the corrected frames is what improves it, one review at a time.
LexData takes the walkway model through its whole life. You type what to look for, Lexi puts a box on every frame, and a person checks each label before anything trains on it. The model then watches the dock cameras, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime. The search that found last month's pallets ran once, on Thursday. The forklift rule has fired every couple of seconds since, on the same recorder, and the dock manager has stopped leaving sticky notes.
See it on your own footage.
Start with your footageMore in Operations

Operations · 7 min read
AGPL-3.0 licensing risk for computer vision teams serving a model
A camera streaming to a served detector is the network interaction the licence was written for. What a legal review will ask, and why to pick the weights first.
Ayman Quadir · Oct 1, 2026

Operations · 7 min read
Cloud vs owned GPU inference for computer vision, worked out per camera hour
A plant on three shifts and a retailer with cameras spread across stores get different answers from one sum, and footage leaving the building is a cost too.
Ayman Quadir · Oct 1, 2026

Operations · 7 min read
Computer vision heatmaps drawn from the aisle cameras a store already has
Footpoints from every tracked box, aggregated over a day and mapped onto the floor plan, show where footfall goes. A camera nudged in cleaning shifts the map.
Rajiya Sultana · Oct 1, 2026