Operations · 6 min read
Comparing two model versions visually on the same shelf frame before rollout
Version 4 of the shelf gap model finds the shadowed gap version 3 missed, and invents one on the dark packs. Overlay both on one frame and look.
Summary
This post puts two versions of a shelf gap model on the same aisle frame and looks at where they disagree, before the new version replaces the old one. It argues that an F1 score averages away the aisle that got worse, that the frames the model doubted are the right set to compare on, and that the disagreements are where a reviewer's time goes. It is for whoever signs off a new version before it goes live.
Rajiya Sultana · Engineering Manager · Oct 2, 2026

Bread shelf with an empty slot flagged and rack sections boxed, from a customer store camera
Version 4 of the shelf gap model has finished training on the corrections from the last six weeks, and its numbers are better than version 3's on the held-out set. The person who has to approve the rollout has a frame open from the bread aisle camera at 7 am, when the bakery delivery is half unpacked, and she is not looking at the numbers. She is looking at the two sets of boxes.
Version 3 missed the empty slot on the bottom shelf, where the shadow from the shelf above swallows the gap. Version 4 finds it. Version 4 also draws a gap on the top shelf, where a run of dark-packaged rye sits perfectly stocked and, to the model, looks like the back of the fixture.
That frame is the whole argument for looking before rolling out. The number said better. The frame says better here, worse there, and the there is a shelf the store cares about.
An F1 score averages away the aisle that got worse
The held-out set has frames from every aisle in the store, and the F1 score on it went up from version 3 to version 4. That is true and it is also the wrong shape for the decision. A score over the whole set is an average, and the average can rise while the dark-pack shelf on aisle 2 gets worse, because aisle 2 is a few frames among many and the shadowed bottom shelves are more of them.
The store does not experience an average. It experiences the top shelf of aisle 2, where a false gap sends someone to restock a full shelf every morning until they stop trusting the alert. The precision and recall behind the score are closer to useful, and per aisle they are closer still, but the thing that finally tells the reviewer what she needs to know is the frame with both versions on it.
She keeps a list, on paper, of the shelves that have fooled a version before. Dark packs, mirrored freezer doors, the promotional end cap that changes every fortnight. New versions get checked on those shelves first.
Overlay both versions on one frame in two colours
The comparison is mechanical once the frames are chosen. Run version 3 on the frame, run version 4 on the same frame, and draw both sets of boxes over it in two colours. Where the colours sit on top of each other, the versions agree and there is nothing to look at. Where one colour stands alone, one version found something the other did not, and that box is either a catch or an invention.
On the bread aisle frame, the bottom shelf has a box in version 4's colour alone: the catch. The top shelf has a box in version 4's colour alone: the invention. Two boxes, two verdicts, one frame, and the reviewer has learned more about version 4 than the score told her.
The same overlay across the whole comparison set produces a short list of frames where the versions disagree at all, and that list is where the hour goes.
The doubted frames are the set to compare on
Which frames to compare on is the decision that makes the comparison honest. A random sample from the store's footage is mostly frames where both versions agree, because most shelves are stocked and most gaps are obvious. The interesting frames are the ones version 3 was unsure of, the ones it sent back for review over the last six weeks, because those are the conditions it struggled with and the conditions version 4 was trained to fix.
Those doubted frames already have a verdict on them, from the person who reviewed them. So the comparison has a third colour available: what a person said was there. Version 4 is better on a frame when its boxes match the verdict and version 3's did not, and worse when the reverse. On the bread aisle, the shadowed gap was a doubted frame six weeks ago, and the verdict said gap. Version 4 agrees with the person. That is what better means.
The lesson on retraining without starting over makes the case that corrections are the most valuable labels a team will get. They are also the most valuable test set, for the same reason.
Where the versions disagree is where the reviewer looks
An hour spent on the disagreement list is worth more than a day spent on random frames. Each frame on it is a place where the model's behaviour changed, and the reviewer's job is to say whether the change was an improvement. Catches go in one pile and inventions in the other, and the dark-pack shelf goes straight back into labeling as a frame version 5 needs to see, with a box that says stocked.
My own view is that no version should go live on the strength of a score alone, however good, and that the disagreement overlay is the least a store owes its night crew. A score is what the model did on frames it will never see again. The overlay is what it will do on aisle 2 tomorrow.
The comparison also catches the failure nobody plans for: a new version that is better on every gap and has quietly stopped boxing the rack sections it used to box, because a class fell out of the training set. That shows up as a whole colour missing from the overlay, and no score would have said so.
Rollout replaces the old version with no downtime, and keeps it
Version 4 went live with a note: watch aisle 2. The rollout replaced version 3 on the store's cameras with no downtime, and version 3 was kept, with what it was trained on, in case the note turned into a problem. It did not. The dark-pack invention became a correction in the review queue within a day, and it is in the set version 5 will train on.
LexData takes the shelf model through its whole life. You type what to look for, Lexi puts a box on every frame, and a person checks each label before anything trains on it. The model then watches the aisle cameras the store already has, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime. Each version keeps what it was trained on, which is what makes the comparison possible at all, and the platform is where the two sets of boxes get drawn on the same frame.
The frame from the bread aisle at 7 am is now the first thing the reviewer opens for every version. It has fooled two of them.
See it on your own footage.
Start with your footageMore in Operations

Operations · 6 min read
Batch video analysis of archived drone survey footage without a notebook open
Three seasons of right-of-way flights in a cloud bucket. Import the originals, sample the frames, run the model, and review only what it doubted.
Stephen Biswas · Oct 2, 2026

Operations · 6 min read
Semantic search across a hundred live camera feeds with one sentence
An operator types what the rare scene looks like and the yard cameras return the frames that match. Search finds candidates; a person confirms them.
Sheikh Srijon · Oct 2, 2026

Operations · 7 min read
Computer vision event logging that keeps the frame with the prediction
A missed bone fragment on the night shift can only be explained if the frame, the prediction, the model version and the lot were logged together.
Rob Hickey · Oct 2, 2026