Operations · 6 min read
How to test computer vision model robustness before the weather does it for you
A vest model trained in June daylight meets October rain. Blur, darken, shift and compress the held-out frames first, and read recall per perturbation.
Summary
This post describes testing a site safety vest model against perturbed held-out frames before rollout, with blur, exposure, colour shift and compression applied one at a time and recall read per perturbation. It concludes that the perturbation table finds the weakness the aggregate score hides, that dawn and rain frames cannot be synthesised and have to be collected, and that the second site needs its own window before it goes live. It is for the teams responsible for a model about to leave the site it was trained on.
Rob Hickey · Chief AI Officer · Sep 30, 2026

Two workers under netting on a scaffold, hard hats boxed, from a customer site camera
The vest model was trained on frames from the site camera at the yard gate in June, from 9 am to 4 pm on the days someone remembered to pull footage. It finds a person without a hi-vis vest in the exclusion zone, and on the held-out frames from the same fortnight it did well enough that the safety manager signed off. The second site opens in October, on the north side of a hill, and the camera there faces the sunrise.
Nothing in the June frames looks like October. The test set agreed with the training set because both were drawn from the same sunny fortnight, and the score is a description of that fortnight.
Perturb the held-out frames before the world does
The cheapest test is to take the held-out frames and break them on purpose, one way at a time. Blur them, as if the lens had a film of dust. Darken them, as if it were dawn. Shift the colour towards blue, as an overcast sky does. Re-encode them at a lower bitrate, as a recorder does when the disk is filling. Each of the four is a separate set, and the model is scored on each separately, against the same labels.
The point is the separation. A single score across all four says the model got worse. Four scores say which thing made it worse, and the answer for the vest model was compression: the hi-vis stripe that made a vest easy to find in June is exactly the fine detail a low bitrate throws away first. Nothing in the aggregate metric would have said so.
My own view is that a model that has never been scored on a compressed frame should not be attached to a recorder, because the recorder will compress the frames whether or not anyone tested for it.
Object detection recall is read per perturbation, per class
For object detection the number to watch is recall, per class, per perturbation, because the failure that matters on a site is a person the model did not find. Precision can wait. A vest box drawn on a traffic cone is a nuisance; a worker without a vest in the zone at 6:30 am who was never boxed is the event the model exists for.
A table with the classes down the side and the perturbations across the top is the whole output of this test. For the vest model it showed that "person" held up under everything, "vest" held up under blur and colour shift, and "vest" fell away under darkness and compression. That is a specific instruction: the next training set needs dark frames and compressed frames, and it does not need more blurred ones. The site safety and PPE compliance use case adds the other half of the job, attributing the gear to a body rather than a scene, and the same table is the way to test that too.
Dawn and rain cannot be synthesised, so they have to be collected
Darkening a June frame is a rehearsal for dawn, and a poor one. Real dawn on the north site has a low sun straight into the lens, long shadows across the zone, and a worker whose vest is lit from behind and reads as a dark shape. Real rain puts drops on the housing and a sheen on every surface, and a wet vest reflects the sky. No perturbation on a sunny frame produces either.
The drift catalog files this under rain, fog and dust: conditions the model rarely saw arrive for a week, detections thin out, and then they recover without anyone touching anything, which is exactly why those days get dismissed as noise rather than kept. The instruction is the opposite. Keep the bad-weather footage. Those frames are rare in a training set and worth more than any number of sunny ones, and they are already on the recorder.
So the test plan has a second half that is a collection plan. Before the October rollout, a window of frames from the north site at dawn, in rain, in the low light after the clocks change, labeled and checked and folded in.
The second site is a new site and needs its own window
Even with the perturbation table clean and the weather frames in, the north site is a different place. Different camera height, a different background, a fence where the June site had a wall, workers in a different contractor's vests. The catalog calls this a new site came online, and the signal is that the north site's detections differ from the June site's from the first day rather than degrading into it.
The fix is the same shape as the weather one. Label a window from the new site before it goes live, and compare the north site against its own history rather than against a fleet average, because the June site's good numbers will carry the north site's bad ones and the average will look fine.
The review queue is the test that runs after rollout
The perturbation table and the collected frames are what can be done before October. After October, the test is the review queue.
LexData takes the vest model through its whole life. You type what to look for, Lexi puts a box on every frame, and a person checks each label before anything trains on it. The model then watches the cameras the site already has, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime.
The dawn frames from the north site that come back in the first week are the perturbation nobody could synthesise, and the correction rate on that one camera at that hour is the robustness number that matters once the model is live.
The site radio crackles at 6:20 every morning as the first crew signs in, and the model's hardest hour starts about ten minutes later.
See it on your own footage.
Start with your footageMore in Operations

Operations · 6 min read
Camera calibration for computer vision, and why the part at the edge of the frame measures wrong
A straight edge bows at the corner of the frame, so a part that passes in the centre fails at the edge. Calibrate once, and again the day the lens changes.
Stephen Biswas · Sep 30, 2026

Operations · 6 min read
Evaluating a computer vision platform once the pilot ends
The feature spreadsheet cannot tell you which platform survives the second year. Walk the lifecycle instead, from import to a retrain that keeps the old frames.
Ayman Quadir · Sep 30, 2026

Operations · 6 min read
Face blurring for privacy in computer vision, done before the frame is stored
A blur applied after the request arrives is a blur applied to a frame that has already been copied. The step belongs at ingestion, beside the recorder.
Esdras Ntuyenabo · Sep 30, 2026