Labeling · 7 min read
How to reduce dataset size without losing accuracy on recorder footage
A dataset built from a recorder is mostly the same frame repeated. Which frames can go, which must stay, and how to know the model did not notice.
Summary
This post takes a dataset built from three months of recorder footage on a food line and asks which frames can be removed. It concludes that near-duplicates, smeared frames from a dirty lens and frames from another site can go, that the rare frames must stay regardless of count, and that the only proof is the same held-out score before and after. It is for teams whose dataset grew from a recorder rather than from a plan.
Sheikh Srijon · GTM Lead · Oct 3, 2026

Fillets on a blue conveyor at a food processing line, conveyor, fillet and rail boxed, generated scene with detections from our model
The recorder above the fillet line has been sampling a frame every two seconds since June, and the dataset folder now holds a little over a hundred thousand frames. Labeling all of them would take the plant's two reviewers into next spring. Training on all of them takes a night per run. And most of them are, to the eye and to the model, the same picture: blue belt, six fillets, the rail on the left, the same lamp.
The instinct is that more frames must be better. On a recorder dataset that is usually wrong, and the question worth asking is which frames are doing any work.
A recorder dataset is mostly the same frame repeated
Footage from a fixed camera over a steady process changes slowly. A fillet enters the frame, travels the belt for eight seconds and leaves, and in that time the recorder has taken four frames of it that differ by its position and almost nothing else. Across a shift, the belt speed, the lighting and the product are constant, so a Monday's footage is a few hundred distinct situations spread across a folder of forty thousand frames.
The how much footage lesson counts in situations for this reason: the model learns from variety, and a hundred thousand frames of one situation teach it about as much as a hundred. Everything that follows is a way of finding the frames that carry a situation the rest do not.
Near-duplicate frames seconds apart cost time and teach nothing
Two frames of the same fillet at the same belt position, two seconds apart, are near-duplicates. Keeping both means labeling both, and the second label is a copy of the first with the boxes moved a few pixels. It also means the model sees that fillet twice per epoch, which is a small vote for that fillet and against the ones it saw once.
Finding them does not need anything clever. Frames are compared to their neighbours in time, and where the difference is below a threshold, the run is collapsed to one frame. On the fillet line this alone removes most of the folder, and the frames that survive are the ones where something changed: a new fillet entered, a hand reached in, the belt stopped.
The one thing to watch is the threshold on a slow scene. Set it too loose and a fillet that shifted slightly as the belt jerked is kept as a new situation. Set it too tight and a defect that appeared on a fillet between one frame and the next is collapsed into the clean frame before it. The reviewer on the Tuesday shift checks a sample of what was collapsed, not only of what was kept.
Smeared frames from a dirty lens teach the model the smear
The lens over the belt gets wiped at the start of each shift, at 6 am and again at 2 pm, with whatever cloth is nearest, and the first ten minutes of every shift are smeared until the condensation clears. Those frames look like a fillet line seen through a wet window. A model trained on them learns that a fillet is sometimes a pale blur, and lowers its own standard for what counts as one.
Blurred and washed-out frames are found by measuring sharpness and exposure, and the trap is the same as with duplicates: the cut has to be checked. A frame that is blurred because the lens was wet is noise. A frame that is blurred because a fillet was being pulled off the belt by hand is a rare event worth keeping, and a sharpness score cannot tell the two apart. A person can, in a second.
Whether to keep any smeared frames at all depends on what the live camera will see. If the lens is wiped and the first ten minutes of the shift are always smeared, the model will be asked about smeared frames every morning, and a few labeled examples of that are worth more than a clean dataset that never met them.
Frames from another site are outliers that pull the model sideways
Somewhere in the hundred thousand frames is a folder from the trial line at plant 2, copied in during the pilot in July because it was labeled and available. The belt there is white. The lamp is warmer. The fillets are cut differently. To the model these frames are a second dataset stitched onto the first, and a class boundary that has to stretch to cover both is a boundary that fits neither.
Outliers like this are found by looking for frames that sit far from the rest in appearance, and by looking at the folder names, which is faster. Whether they should go depends on whether the second plant is a site the model will ever watch. If it is, the frames belong in a dataset for that site. If it is not, they are pulling a model for the blue belt toward a white one it will never see.
The rare frames stay whatever the count says
Pruning removes what is common. It must never remove what is rare, and on a food line the rare frames are the ones the whole project exists for: the fillet with a bone, the piece that fell off the belt, the glove that entered the frame at 3 am. There may be forty such frames in a hundred thousand. A pipeline that trims by similarity will quite happily keep the clean fillets and drop the one with the defect, because the defect frame is unlike anything else and looks, to a distance measure, like an error.
My own view is that pruning should be done before labeling, when a frame is cheap to keep and cheap to drop, and that the rare classes should be counted before and after with the numbers written down. A dataset that went from a hundred thousand frames to four thousand and still holds all forty defect frames has lost nothing that mattered. One that holds thirty-one has lost the project.
The labeling with Lexi doc describes the review pass as a check on every label. The pruning pass is a check on every deletion, and the rare frames are where that check earns its time.
Measure before and after on the same held-out frames
The only proof that the model did not notice is a score. Before pruning, train on the full set and score on a held-out set of frames the pruning never touched, chosen from a camera and a fortnight in August that stay locked. After pruning, train again and score on the same frames. If the score holds, the removed frames were not carrying anything. If it drops on one class, look at what was removed from that class, because something rare went out with the duplicates.
The held-out set is never pruned. It is tempting, because it is also full of near-duplicates, and a smaller test set runs faster. A test set thinned by the same rule as the training set will agree with the training set about what matters, and the score it gives is no longer independent.
The same discipline carries into the retrain. Frames the model is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime. Those returned frames are the opposite of duplicates: each one is a situation the model did not have, and the retraining without starting over lesson treats them as the most valuable labels a project ever gets. Pruning is what keeps the dataset small enough that those frames are not lost in a hundred thousand copies of the belt.
See it on your own footage.
Start with your footageMore in Labeling

Labeling · 7 min read
Aerial dataset augmentation for drone frames where there is no up
A tower seen straight down has no top or bottom, so rotations and both flips are safe. Scale for altitude, brightness for sun, move every box with its pixels.
Esdras Ntuyenabo · Oct 3, 2026

Labeling · 6 min read
A collaborative data annotation workflow run as a pipeline
Batch the bottling line's frames by camera and shift, assign so nothing is boxed twice, attach the guideline, review every label, train on the approved set.
Rajiya Sultana · Oct 3, 2026

Labeling · 7 min read
Dataset health check for computer vision, what to look at before anything trains
A scratch dataset where every scratch sits in the centre of the frame will train a model that looks in the centre. Five counts to read before the first epoch.
Rajiya Sultana · Oct 3, 2026