Labeling · 7 min read
Splitting camera footage into train, validation and test so the score means something
Frames a second apart from one camera land on both sides of a random split and the test score lies. Split by camera and by day, and hold out the newest camera.
Summary
This post explains why a random split of recorder footage puts near-identical frames in both the training and the test set, and what to do instead. It concludes that the split should follow cameras and days, that the validation set is only there to decide when training stops, and that the site's newest camera is the honest hold-out. It is written for teams building their first dataset from footage they already record.
Sheikh Srijon · GTM Lead · Oct 3, 2026

Pole camera over an industrial yard at dusk, camera, fence and vehicle boxed, generated scene with detections from our model
A distribution yard has five cameras on its fence, three over the dock doors and two on the gate, and the recorder has been keeping their footage since March. The team pulls a frame every two seconds, labels the trucks and the people, shuffles the frames and cuts the pile into three: most for training, a slice for validation, a slice for test. The first model scores beautifully on the test slice. The first week on the live cameras, it misses a truck at the gate in the rain and boxes a shadow on the dock as a person.
The score was never wrong about the test set. The test set was wrong about the yard.
Consecutive frames from one camera leak the answer into the test
Two frames pulled two seconds apart from the same dock camera are close to identical: the same truck in the same bay, the same light, the same puddle. When a random shuffle puts one of them in training and the other in test, the model is being examined on a question it has already seen the answer to. It does not need to have learned what a truck looks like, only what that truck looked like at 14:06 on a Tuesday.
This is why the test score on shuffled recorder footage sits so far above the number the yard sees. Every frame in the test slice has a near twin in the training slice, and the model is rewarded for memory rather than for the thing the yard wanted, which was to recognise a truck it has never met.
The how much footage lesson makes the same point from the other side: the useful unit is the number of distinct situations, and a thousand frames of one afternoon are close to one situation.
Split by camera and by day, never by frame
The fix is to decide the split before anything is shuffled, and to decide it on the boundaries that matter in the yard. A day is one boundary: the footage from Monday goes wholly to one side and the footage from Wednesday wholly to another, so a weather front or a delivery schedule cannot straddle the cut. A camera is another: frames from the second gate camera stay together, so a framing the model has learned from cannot appear in the set that examines it.
In practice this means grouping frames by camera and by day, then assigning whole groups to train, validation and test. The proportions end up uneven, because days are not the same size and the gate cameras see less traffic than the dock, and that is fine. The score from an uneven, honest split is worth more than a tidy one from a leaky split.
The recorder on the dock runs an hour behind the gate recorder because nobody reset it after the clocks changed in March. The split by day has to know that, or a late shift on one camera lands on the wrong day's side of the cut.
The validation set decides when to stop the epochs
Training runs in passes over the training set, and each pass is an epoch. Left to run on the yard's frames from March to July, a model keeps improving on the frames it is shown long after it has stopped improving on frames it has not, and the validation set is how that turn is seen. After each epoch the model is scored on the validation frames it never trains on; the epoch where that score stops rising is the version to keep.
That is the validation set's whole job. It is consulted after every epoch and it guides choices such as the learning rate, the input size and which of two label conventions works better, so by the end of a project it has been looked at hundreds of times. A number that has been used to steer that many decisions is no longer an unbiased estimate of anything, and it should never be quoted as the model's accuracy.
The test set is read once and then retired
The test set exists so that one number in the project has not been steered. It is held back through every epoch and every experiment, scored once when the model is chosen, and the result is written down as what the yard can expect. Scoring it, adjusting the model, and scoring it again turns it into a second validation set, and the project is left without an honest number.
Teams shrink the test set when frames are scarce, and that is understandable and usually a mistake. A test set of forty frames from one Friday afternoon can swing several points on a single missed truck. If the labeled frames are scarce, the test set is the last place to economise, because the model can be improved later and a false score cannot be taken back.
My own view is that a test set should be locked with the same ceremony as a release: named, dated, and stored where nobody edits it by accident. It is the one part of a dataset whose value comes from being left alone.
The site's newest camera is the honest hold-out
The yard added a sixth camera over the new dock door in August. Its frames are the best test set the team will ever have, because they are precisely the case the model faces after launch: a framing it has not seen, on a site it knows. Holding that camera out entirely, with none of its frames in training or validation, gives a score that predicts the seventh camera better than any random slice can.
The same logic applies in time. Holding out the most recent two weeks, rather than a random two weeks, tests the model against the future rather than against a shuffled past. The retraining without starting over lesson argues that this is the only hold-out that tells you whether a new version is better where it matters.
When the newest camera scores badly, that is information. The model has learned the five old framings well and a sixth is not yet covered, and the fix is a short window of labeled frames from the new door, folded in at the next retrain.
A split that stays with the dataset survives the retrain
A model in production keeps generating candidates for the dataset. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old with no downtime. Each of those corrected frames has to land on one side of the split and stay there, or the leak returns through the back door: a frame the model doubted on Tuesday, corrected and trained on, must not be the frame that examines version four.
The way to keep that straight is to make the split a property of the frame rather than of a run. Each frame carries its camera and its day, and the rule that maps those to train, validation or test is written down once and applied to every frame that arrives afterwards. The what the score beside a detection means lesson covers why the number on a single frame is a ranking rather than a promise; the split is what makes the number across all the frames worth reading at all.
On the yard, the sixth camera stayed held out for the whole first quarter. The team knew its score on that door before they trusted the score anywhere else.
See it on your own footage.
Start with your footageMore in Labeling

Labeling · 7 min read
Aerial dataset augmentation for drone frames where there is no up
A tower seen straight down has no top or bottom, so rotations and both flips are safe. Scale for altitude, brightness for sun, move every box with its pixels.
Esdras Ntuyenabo · Oct 3, 2026

Labeling · 6 min read
A collaborative data annotation workflow run as a pipeline
Batch the bottling line's frames by camera and shift, assign so nothing is boxed twice, attach the guideline, review every label, train on the approved set.
Rajiya Sultana · Oct 3, 2026

Labeling · 7 min read
Dataset health check for computer vision, what to look at before anything trains
A scratch dataset where every scratch sits in the centre of the frame will train a model that looks in the centre. Five counts to read before the first epoch.
Rajiya Sultana · Oct 3, 2026