Operations · 6 min read
Evaluating a computer vision platform once the pilot ends
The feature spreadsheet cannot tell you which platform survives the second year. Walk the lifecycle instead, from import to a retrain that keeps the old frames.
Summary
This post replaces the feature spreadsheet with the model's lifecycle as the way to judge a computer vision platform after a pilot. It walks through import without re-encoding, a reviewer on every label, running where the cameras are, the correction rate as the drift signal, and a retrain that does not start from nothing. It is for the people choosing the platform a plant or estate will still be using in its second year.
Ayman Quadir · Head of Product · Sep 30, 2026

A camera on a pole over an industrial yard at dusk, generated scene with detections from our model
The pilot at the bottling plant ended on a Friday with a model that found missing caps on line 1 and a spreadsheet with forty rows. Each row was a feature, each column a vendor, and every cell was a tick. The row labelled "AI" was ticked all the way across. So was "annotation". So was "deployment". The spreadsheet could not distinguish the platforms, and the team that made it knew that, which is why the meeting was long.
A spreadsheet of features describes what a platform has. What the plant needs to know is what happens to the cap model in the ninth month, when the supplier changes the cap colour and the line manager has stopped looking at the dashboard.
Computer vision projects fail between the stages
The stages of a vision project are well known: collect footage, label it, train, deploy, watch. Each vendor has a page for each. Computer vision projects rarely fail inside a stage. They fail in the gaps, where frames are exported from one tool and imported into another, where a correction made by an operator on a live frame lands in a spreadsheet nobody reads, where a retrain means finding last quarter's dataset on someone's laptop.
So the evaluation is a walk along the lifecycle with one question at each seam: does the thing that came out of the last stage go into this one without a person carrying it. The rest of this post is that walk, for the cap model on line 1.
Footage has to come in as it was recorded
The plant has three years of recordings on the recorder in the line 1 cabinet, and the pilot's labels in COCO JSON from a contractor. The first test is whether both arrive intact. Footage from S3, Google Drive or an upload, and nothing re-encoded on the way, because a re-encode changes the frames the model will learn from into frames the cameras will never produce. Labels in COCO JSON, YOLO TXT or CVAT XML, imported as they are, so the contractor's month of work is not redone.
A platform that asks for the footage to be converted first has already put a person in the first gap.
A person checks every label before the model trains on it
The pilot's labels were good enough for a pilot. In the second year the model is retrained on corrections made at 3 am by a line operator between two other jobs, and the question is whether anything stands between that correction and the next training run.
The answer should be a reviewer. Lexi proposes the box, a person checks it, a QA pass sits on every label before training, and the labels come back at up to 99.9% accuracy. What the evaluation looks for is the seam again: is the review queue the same queue the training set is built from, or is there an export in between. The quickstart shows this from the first labeled frame, and it is worth doing with the plant's own footage rather than the sample set.
My own view is that a platform demo should run on the buyer's own frames or not at all, because the sample set has been chosen to look good and the line 1 cabinet has not.
The model has to run where the cameras are
The recorder in the line 1 cabinet is on a network built for PLCs. The platform either runs the model there, on a runner beside the recorder, or on the plant's own servers, or in its cloud with the stream sent up, and the plant needs the choice to be the plant's. The footage that may not leave the site should not have to.
Ask where the doubted frames go. With the runner, they are the only thing that leaves, and the alert for a missing cap fires locally before anything has crossed the firewall. Ask what leaves the platform too: the dataset as COCO, the model as PT, ONNX or TorchScript with a device preset for the box in the cabinet, so the plant owns the thing it paid to train.
The correction rate is the drift signal the platform has to show
In the ninth month the cap supplier changes from white to pale grey. The model keeps finding caps, mostly, and starts sending more frames from line 1 to the review queue. The operator corrects a few every shift. That rise, on one camera, from one week, is the signal that the model is drifting, and it is the only signal the plant has, because nobody is labeling every frame to compute an accuracy figure.
A platform that shows the correction rate per camera is showing the plant the truth. One that shows a dashboard of the model's own scores is showing the plant what the model reports about itself, and the model has not been told about the grey caps.
A retrain that keeps what the model already knew
The last seam is the one most pilots never reach. The corrections on the grey caps cross the project's threshold and a new version trains. Does it train on the corrections and the three years of frames, or does someone rebuild a dataset. Does the new version keep a record of what it was trained on. Does it replace the old one on line 1 without stopping the line, and can the old one come back if the new one is worse on the night shift.
LexData takes the cap model through its whole life. You type what to look for, Lexi puts a box on every frame, and a person checks each label before anything trains on it. The model then watches the cameras the plant already has, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime. The platform is that loop, and the field guide's lesson on retraining without starting over is the argument for why the old frames stay in.
The spreadsheet at the bottling plant got one more row after the walk. It asked whether a correction made on line 1 at 3 am reaches the next version without a person exporting anything, and it was the only row with a cell that was not a tick.
See it on your own footage.
Start with your footageMore in Operations

Operations · 6 min read
Camera calibration for computer vision, and why the part at the edge of the frame measures wrong
A straight edge bows at the corner of the frame, so a part that passes in the centre fails at the edge. Calibrate once, and again the day the lens changes.
Stephen Biswas · Sep 30, 2026

Operations · 6 min read
Face blurring for privacy in computer vision, done before the frame is stored
A blur applied after the request arrives is a blur applied to a frame that has already been copied. The step belongs at ingestion, beside the recorder.
Esdras Ntuyenabo · Sep 30, 2026

Operations · 6 min read
How to test computer vision model robustness before the weather does it for you
A vest model trained in June daylight meets October rain. Blur, darken, shift and compress the held-out frames first, and read recall per perturbation.
Rob Hickey · Sep 30, 2026