Operations · 6 min read
When to add an LLM to a vision pipeline and ask the footage a question
A detector boxes the crushed carton on the depot conveyor. Whether the label is still readable, and where the parcel goes now, is a question for the footage.
Summary
This post uses a conveyor camera at a parcel depot to draw the line between what a detector answers, where and how many, and what needs a question put to the footage, such as whether a damaged label is still readable and where the parcel should be routed. It concludes that the detector runs first because it is cheap and the question runs second because it is not, and that the answer must arrive with the frames behind it. It is for operations teams deciding whether a detector is enough.
Ayman Quadir · Head of Product · Oct 1, 2026

Packing station with a crushed carton and its label boxed, generated scene with detections from our model
At the parcel depot the induction conveyor carries a carton past the camera every couple of seconds, and one in a while arrives with a corner stove in and the label creased across the barcode. The detector above belt 2 boxes it as damaged before it reaches the sorter. That is a good detector doing its job, and it is also the point where the depot's real question begins. Can the label still be read, and if it can, where does this parcel go now?
A detector cannot answer that, and it should not be asked to. The question is about reading, judgement and context, and it is a different kind of work.
Object detection answers where and how many
A detector's output is a box, a class and a position. On the induction belt that covers most of what the depot needs: how many cartons passed, which of them were damaged, whether a label is present and inside the window where the scanner expects it. Every one of those is a region and a rule, and the package and label inspection use case is built on boxes for that reason.
The detector is also fast, which matters at line speed, and it is stable, which matters more. A box on a crushed carton looks the same on the night shift as on the day shift, and a rule about where the label sits does not change its mind.
What the detector does not do is read the label, judge whether the crease has taken out the check digit, or know that parcels for the Leeds hub with an unreadable label go to the manual rework bench rather than the reject chute.
A question about the label is a different kind of work
"Is this label still readable, and where should the parcel go" is a question with the answer spread across the picture and the depot's rules. It needs the characters on the label, a view on whether the barcode is decodable, the routing code, and a decision that depends on whether the parcel was bound for Leeds or for the coast. A model that understands language and images together can take the boxed region and the question and give a sentence back.
That is the case for adding one. There are three signs a team is looking at a question rather than a detection. The answer is a sentence, the rule for it lives in a document rather than a threshold, and the depot would happily let a person answer it if there were enough people.
The signs a team does not need one are as clear. Counting, locating and classifying are detections. A tight time budget between cartons is a detection. A model that has to live on the box beside the belt with no network is a detection. A vision language model is slower and more expensive per frame than a detector by a wide margin, and asked to do a detector's job it does it worse.
Ask the footage and get the frames behind the answer
On our platform the question goes to LexInsight. The detector has already boxed the damaged carton, so the question is asked about a small region of a known frame rather than about the whole belt. The answer comes back with the frames behind it: the crop of the label, the characters it read, and the routing it concluded. A supervisor at the rework bench sees the answer and the evidence together, and can disagree with it.
The depot asks the same question in bulk at the end of the day. How many damaged cartons had readable labels, how many went to rework, and which sender's packaging crushes most often. Each answer is a set of frames a person can open.
A question put to the footage changes no model. Corrections a person makes at review are what retrain the detector. The two paths stay separate on purpose: the detector keeps learning from what people correct, and Insight keeps answering from what the detector found.
The detector runs first because the question is expensive
The order is the architecture. The detector runs on every sampled frame, and it is cheap enough to do so. The question runs only on the frames the detector has flagged, which on the induction belt is a small fraction of the day. Reversing that, asking a language model to look at every carton, is the most common way this pattern fails, because the cost scales with the belt rather than with the problem.
My own view is that the second stage should be added the Monday a team finds itself writing rules against text. A rule of one word, damaged, is a detector's rule. A rule that runs to damaged, with a label that reads as a Leeds code, with the barcode intact, is a question, and a team that keeps extending the detector's rules to cover it ends up with a brittle parser of its own.
The depot's rework staff call the crushed cartons pancakes, and the sorter's reject count is known on the floor as the pancake count. The label reader has not changed the name.
The answer must never guess a routing code
A model that answers in sentences can answer confidently and wrongly, and on a parcel that means a package sent to the wrong hub with a plausible reason attached. Three things keep that in check. The answer arrives with the frames, so the supervisor can see the label the model read. The routing rule is looked up rather than recalled, so the model reads the code and the depot's own table decides where the code goes. And a low-confidence read, a smudged digit the model could not resolve, goes to the rework bench as unreadable rather than as a best guess.
LexData takes the depot's detector through its whole life. You type what to look for, Lexi puts a box on every carton and label in every frame, and a person checks each label before anything trains on it. The model then watches the induction cameras the depot already has, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime. The question layer sits on top of that, and the platform page shows where each part lives.
The rework bench has a rubber stamp that says REROUTED in red. It gets used on the parcels the model read correctly too.
See it on your own footage.
Start with your footageMore in Operations

Operations · 7 min read
AGPL-3.0 licensing risk for computer vision teams serving a model
A camera streaming to a served detector is the network interaction the licence was written for. What a legal review will ask, and why to pick the weights first.
Ayman Quadir · Oct 1, 2026

Operations · 7 min read
Cloud vs owned GPU inference for computer vision, worked out per camera hour
A plant on three shifts and a retailer with cameras spread across stores get different answers from one sum, and footage leaving the building is a cost too.
Ayman Quadir · Oct 1, 2026

Operations · 7 min read
Computer vision heatmaps drawn from the aisle cameras a store already has
Footpoints from every tracked box, aggregated over a day and mapped onto the floor plan, show where footfall goes. A camera nudged in cleaning shifts the map.
Rajiya Sultana · Oct 1, 2026