Operations · 6 min read
Ask the footage a question with a VLM, but let the trained model count
What does the gauge on the pump skid read, what does the label at the dock say. A question gets an answer with frames behind it. Counting all shift is a rule.
Summary
This post separates two things a plant wants from its cameras: an answer to a question about one frame, such as what a gauge reads or what a shipping label says, and a count or a rule that runs on every frame all shift. It argues that a vision language model is the right tool for the first and the wrong tool for the second, and that an answer to a question changes nothing about the model watching the feed. It is for operations leads deciding what to ask and what to automate.
Ayman Quadir · Head of Product · Oct 2, 2026

Analog gauge on a plant pipe, generated scene with detections from our model
The pump skid at the back of the plant has a pressure gauge that nobody has fitted a transducer to, and a camera that looks at it. At 3 pm the shift lead wants to know what it read at lunchtime, because the pump tripped at half past one and the log has nothing between noon and two. He does not want a model trained. He wants a number.
At the loading dock, a different question. The dispatcher wants to know what the shipping label on the pallet in bay 2 says, because the driver is waiting and the paperwork disagrees with the pallet. She also does not want a model trained. She wants the words on the label.
Both are questions about one frame, and both are exactly what a vision language model is for. Neither is what the plant should build its monitoring on.
A question about the gauge is a VLM job, once
The shift lead opens LexInsight, picks the skid camera, and asks what the gauge read at half past one. The answer comes back as a reading, and it comes back with the frames behind it: the frame at half past one with the gauge in it, and the frames either side, so he can see the needle and judge the answer himself. A vision language model is good at this because the question is open, the frame is one, and a person is there to read the result.
The dispatcher does the same with the dock camera and the pallet in bay 2. The label text comes back, the frame comes back, and the paperwork is corrected before the driver has finished his coffee.
What makes a VLM fit here is that the question was not known in advance. Nobody trained a gauge reader for the skid because nobody expected to need one until the pump tripped. The question box is where the unexpected question goes.
Counting cartons all shift needs a trained model
Now the other thing the dock needs. Every pallet that leaves bay 2 has to be counted, every label has to be checked for presence, and the count has to be right at the end of every shift, all year. That is a rule, and it runs on every frame the camera takes, sampled about every two seconds, on every bay, without a person asking.
A VLM is the wrong tool for that, for reasons that have nothing to do with how clever it is. Asked to count cartons on a frame, it will give a number, and on the next frame a slightly different number for the same pallet. It will do so at a cost and a latency per frame that no dock budget survives across a shift. It was built to answer, and a count across a shift is a discipline rather than an answer.
The trained model is the discipline. It finds cartons and labels the same way on every frame. The rule is written as a sentence, "a pallet leaves bay 2 with a carton missing its label", severity routine, approved before it goes live, and it turns those boxes into a count and an alert with the frame attached.
The lesson on what a vision model can and cannot see draws the same line from the other side: "is there a label on this pallet" is a detection problem, and "how many pallets left today" is a tracking problem wearing a detection costume. Neither is a question.
The answer comes with the frames behind it
The part of the question box that matters most in a plant is the frames. An answer that says the gauge read a number is an opinion. An answer that says the gauge read a number and here is the frame is evidence, and the shift lead can disagree with it by looking. Every answer in LexInsight arrives that way, because the person asking is the one who has to act on it and the frame is what they act on.
The same is true of the trained model's alerts. The frame with the box drawn on it is what the dispatcher sees on her phone, and she judges it in a glance.
My own view is that most plants use the question box a great deal in the first fortnight and then rarely, while the rules run all year. That is the right pattern. The questions find out what the plant did not know it needed, and the rules are what it builds once it knows.
Answers do not change the model, corrections do
A thing worth being clear about. When the shift lead reads the gauge through LexInsight, the model watching the skid is unchanged. When the dispatcher reads the label, the carton model is unchanged. The question box reads the footage and answers; it does not teach anything.
What teaches the model is the review. The carton model returns the frames it is unsure of, a label half under shrink wrap or a carton at the edge of the bay, and a person confirms or corrects the box. Those corrections are what the next version trains on. If the dispatcher notices the model missing a new label format, the fix is a correction in the review queue, and a question about the label would leave the model exactly as it was.
LexData takes the carton model through its whole life. You type what to look for, Lexi puts a box on every frame, and a person checks each label before anything trains on it. The model then watches the dock cameras, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime. The question box sits beside that loop and reads from the same footage, and the two are not confused for each other on the platform because they answer different needs.
Where the VLM belongs in a plant's day
The gauge on the skid got its own rule a month later, once the shift lead had asked about it enough times to know the reading mattered. Two classes, the pivot and the needle tip, checked by a person, and a rule about pressure above a margin. The question that started as a one-off became a rule because the plant learned it was not one.
That is the path the question box is good for. Ask about the thing you did not plan for, look at the frames, and when the same question comes back every day, write it as a rule and let the trained model carry it. The skid camera answers the 3 pm question in both ways now: the rule has the reading every couple of seconds, and the question box still has the frames for the day the rule looks wrong.
The dispatcher, for what it is worth, still asks the question box about labels on the pallets from one particular supplier, whose printer has been fading for a year. The rule flags them. She wants to read them.
See it on your own footage.
Start with your footageMore in Operations

Operations · 6 min read
Batch video analysis of archived drone survey footage without a notebook open
Three seasons of right-of-way flights in a cloud bucket. Import the originals, sample the frames, run the model, and review only what it doubted.
Stephen Biswas · Oct 2, 2026

Operations · 6 min read
Semantic search across a hundred live camera feeds with one sentence
An operator types what the rare scene looks like and the yard cameras return the frames that match. Search finds candidates; a person confirms them.
Sheikh Srijon · Oct 2, 2026

Operations · 7 min read
Computer vision event logging that keeps the frame with the prediction
A missed bone fragment on the night shift can only be explained if the frame, the prediction, the model version and the lot were logged together.
Rob Hickey · Oct 2, 2026