Operations · 7 min read
Embedding based anomaly detection for the frame the model has never seen
A car park camera meets its first snowplough and a line camera meets the maintenance crew. Distance from the usual frames is one check that finds them.
Summary
This post takes two fixed cameras that meet something new, a car park camera seeing its first snowplough and a line camera seeing the maintenance crew inside the fence, and explains embedding distance as a check a team can log to surface frames unlike anything the model trained on. It concludes that the distance is a way to find frames for review, while the drift signal the platform reads stays the operator correction rate. It is for teams who want to see the unfamiliar frame before it becomes a wrong answer.
Rob Hickey · Chief AI Officer · Oct 1, 2026

Factory floor with the robot cell fence boxed, generated scene with detections from our model, the kind of frame a crew inside the fence turns unfamiliar
At 5 am on the first snowy morning of the year, the camera over the office car park watches a snowplough come through the gate, push a ridge of snow across the disabled bays and leave. The model that counts cars in the bays has never seen a snowplough. It has never seen snow. It counts the plough as two cars, misses the three cars under the ridge, and reports a car park that is half full when it is nearly empty, and it does so with a straight face.
The same week, on a bottling line in a different building, the maintenance crew opens the robot cell fence at the end of the shift and three people in hi-vis stand where the model has only ever seen a robot arm and bottles. The model boxes one of them as a bottle.
Neither model is broken. Each has met a frame unlike anything it was trained on, and each did what a detector does with such a frame, which is to answer anyway.
An embedding is a position, and a familiar frame sits near its neighbours
Somewhere inside the model, before the boxes come out, every frame is turned into a list of numbers that summarises what the model made of it. That list is the frame's embedding, and it can be treated as a position in a space with as many dimensions as the list has numbers. Two frames that look alike to the model sit close together in that space. Two frames that look different sit far apart.
For a fixed camera the interesting property is how tightly the familiar frames cluster. The car park at 5 am on a Tuesday in October and the car park at 5 am on a Wednesday in October produce embeddings a hair apart. The car park with a snowplough in it produces one that sits out on its own, far from every frame the camera has produced since the model was trained. The distance is measurable, and it can be measured on every sampled frame as it arrives.
Distance from the usual frames is one check a team can log
The check is simple to describe. Keep a set of embeddings from frames the model handled well, the reference set, and for each new frame compute how far its embedding is from the nearest members of that set. Log that distance with the frame's timestamp. Most frames land close. A frame that lands far is one the model has little evidence for, whatever boxes it drew, and it is a candidate for a person to look at.
Two choices shape the check. The reference set should come from the camera as mounted, across the hours and the seasons the model has already seen, so that a dusk frame in November is not flagged as strange merely because the reference set was built in July. And the threshold for "far" should be set from the distances the camera's own ordinary frames produce, since a car park and a bottling line have different spreads, and a threshold borrowed from one is wrong on the other.
The learn section explains what drift is and what it is not, and this check is aimed at one of its four kinds: the world producing frames the training set never covered.
The far frames go to a person, and most of them are boring
What comes out of the check is a short list of frames per day, sorted by distance. On the car park camera on the Tuesday of the snowy week the list was the snowplough, the ridge of snow itself, and a helium balloon that drifted through the frame at lunchtime and got stuck in the hedge. On the bottling line it was the maintenance crew, a dropped crate of empties, and a frame where the hall lights had been switched off with the line still running.
A person looks at the list. The snowplough is a new thing that will come back every winter, and the frames go into the next version as labeled examples. The balloon is noise, and its frames are marked as nothing. The crew inside the fence is a rule the site wants, a person inside the cell, and it becomes an alert as well as a label. Each verdict is a correction, and corrections are what retrain the model.
My own view is that the distance check is a flashlight rather than a verdict. It finds frames a person should see. It says nothing about whether the boxes on them were right, and a team that treats a high distance as an error has confused unfamiliar with wrong.
The balloon stayed in the hedge for three days and was flagged on every sampled frame until the wind took it.
The correction rate is the drift signal the platform reads
A team can log the distance and should. It is a good way to find frames for review. The signal the platform itself reads for drift is a different one, and it is the one that carries the person's judgement: the correction rate. Frames the model is unsure of come back to a person, the person corrects some of them, and the fraction that needed correcting is the number that rises when the world has moved. When the corrections cross the project's threshold a new version is trained on them, keeps what it was trained on, and rolls out with no downtime.
The distance check and the correction rate answer different questions. The distance says a frame is unlike the training set. The correction rate says the model's answers have started to need fixing. A snowplough raises both. A gradual shift in the car park's lighting as the trees leaf out raises the second and barely moves the first, because every frame is only slightly unlike the last.
LexData takes the car park model through its whole life. You type what to look for, Lexi puts a box on every car in every frame, and a person checks each label before anything trains on it. The model then watches the camera the office already has, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime. The platform page shows where the review queue sits in that loop, and the distance check is one way a team can add frames to it.
The check misses the wrong answer on a familiar frame
The limit of the method is the frame that looks ordinary and is not. A car parked across two bays at 9 am, at the usual angle, in the usual light, is a familiar frame with an unusual count, and its embedding sits in the middle of the reference set. The bottling line's cap that is on crooked looks, to the embedding, like every other bottle. The distance check will never surface either, because there is nothing unfamiliar about the picture, only about the answer.
Those are found the other way round: by the model's own doubt on the frame, and by a person disagreeing with a label at review. The distance check finds the snowplough. The review queue finds the crooked cap. A site needs both, and they are looking at different things.
See it on your own footage.
Start with your footageMore in Operations

Operations · 7 min read
AGPL-3.0 licensing risk for computer vision teams serving a model
A camera streaming to a served detector is the network interaction the licence was written for. What a legal review will ask, and why to pick the weights first.
Ayman Quadir · Oct 1, 2026

Operations · 7 min read
Cloud vs owned GPU inference for computer vision, worked out per camera hour
A plant on three shifts and a retailer with cameras spread across stores get different answers from one sum, and footage leaving the building is a cost too.
Ayman Quadir · Oct 1, 2026

Operations · 7 min read
Computer vision heatmaps drawn from the aisle cameras a store already has
Footpoints from every tracked box, aggregated over a day and mapped onto the floor plan, show where footfall goes. A camera nudged in cleaning shifts the map.
Rajiya Sultana · Oct 1, 2026