Labeling · 7 min read
Data augmentation for object detection on a fixed inspection camera
A stamping line camera varies in known ways. Conveyor speed changes scale and blur, a relamp changes the light. Augment for those and move the boxes too.
Summary
This post chooses augmentations for a fixed inspection camera on a stamping line from what the line actually does: conveyor speed changes the scale and the blur, maintenance changes the lighting, and nothing ever flips a steel sheet upside down. It shows how every bounding box has to follow a geometric change, why the flips the line never produces teach a false world, and why augmentation widens a set of real frames without ever replacing them. It is for the engineer setting up training on line footage.
Stephen Biswas · Engineer · Oct 2, 2026

Stamping line, steel sheets passing the inspection station, generated scene with detections from our model
The inspection camera on a stamping line looks straight down at steel sheets coming off press 2, and the model boxes each sheet and any scratch or dent it finds, along with burrs on a sheared edge. The camera does not move. The sheets arrive in the same place every time. Over a year, what does change is that the conveyor runs faster on some products and slower on others, the maintenance crew relamps the bay twice, and the bay door behind the station opens in summer.
That is the entire list of ways the frames vary. It is short, and it is the list the augmentations should be chosen from, because augmentation for object detection is a way of showing the model the variation the camera will meet, and a fixed camera meets very little.
A fixed camera varies in a few known ways
Augmentation on a general dataset reaches for everything: flips and rotations, crops and colour jitter, blur and noise and cutouts. Each is a guess about how the world might differ from the training frames, and on photographs from anywhere the guesses are reasonable. On the camera over press 2 they are mostly wrong, because the world in front of that camera is nearly fixed and the model has to learn a narrow thing precisely.
The lesson on how much footage you actually need puts the unit of value as distinct conditions. A stamping line camera has a handful of them, and each has a cause on the line, whether a speed, a lamp or a door. The augmentation list is written by walking the line and asking what changes the picture, and on press 2 the answer fitted on a sticky note stuck to the HMI.
The note says faster belt, slower belt, new lamps and an open door. Nothing else.
Conveyor speed changes scale and blur, so augment both
When the conveyor speeds up for a thinner gauge, two things happen to the frame. Each sheet spends less time under the camera and is caught at a slightly different position, and the motion blur along the direction of travel grows. When it slows down, the reverse. A model trained only at one speed has sharp sheets in one place and does less well on the blurred ones a few weeks later.
So the augmentations for speed are a small shift along the belt direction, a small scale change to cover the sheet's apparent size at the camera's fixed height, and a directional blur along the belt axis at a range that matches the fastest product. Directional matters: a blur in every direction teaches the model a world with a shaking camera, and the camera on press 2 does not shake.
Each of these is measured from the line rather than guessed. The fastest belt speed and the camera's exposure give the maximum blur length in pixels, and the augmentation goes up to that and no further.
Maintenance changes the light, so shift brightness and colour
The relamp is the other real change. New lamps are brighter and cooler than the ones they replace, and for a few weeks the frames from press 2 have a different white balance and a different contrast, until the crew relamps another bay and the balance shifts again. The summer door adds daylight from one side and a shadow across the far end of the sheet.
Brightness shifts, contrast shifts and a colour temperature range are the photometric augmentations that cover this, and again the range comes from the line. The plant's relamp log and a frame from each side of the last one give the shift in brightness the model needs to be indifferent to. Going far beyond that, into colour jitter that turns steel green, spends capacity on a world that does not exist.
My own view is that the relamp is the single most common cause of a fixed camera's model quietly degrading, and that a brightness augmentation matched to the last relamp is the cheapest insurance a line can buy.
The bounding box has to follow every geometric change
Every geometric augmentation moves pixels, and every bounding box on the frame has to move with them. A shift along the belt shifts every box by the same pixels. A scale change scales every box about the same centre. A crop drops the boxes that fall outside it and clips the ones on its edge. A clipped box on a scratch that is now half out of frame is a question: keep it as a partial scratch, or drop it because a sliver teaches nothing.
The photometric augmentations move no pixels and leave the boxes alone, which is one reason they are the safer place to start. The geometric ones are where a wrong transform produces a training set of boxes sitting beside their scratches, and a model that learns an offset nobody intended.
The check is to draw a sample of augmented frames with their transformed boxes and look at them. On press 2 the first pass had the boxes scaling about the frame's corner instead of its centre, and every scratch box on a scaled frame sat a few pixels off. Nobody would have found that in a loss curve.
Flips the line never produces teach a false world
A horizontal flip is the most common augmentation in any tutorial, and on press 2 it is wrong. The sheets come off the press with a known orientation, the burr on a sheared edge is always on the same side, and a dent from the die is always in the same corner. A flipped frame shows the model a burr on the wrong edge and a dent in the wrong corner, and the model, trained to be indifferent to that, loses the positional knowledge that makes those defects easy to find on the real line.
Vertical flips and large rotations are wrong for the same reason: the camera is mounted, the belt runs one way, and a sheet upside down has never happened. Small rotations of a degree or two are fine, because sheets do sit slightly askew on the belt.
The rule is that an augmentation is allowed if the line could produce that frame. If the line could not, the augmentation is teaching a world the model will never see, at the cost of the one it will.
Augmentation widens the frames, and never replaces them
None of this substitutes for footage from the camera. Augmentation makes copies of the frames the line already gave, varied within a range the line already showed. It cannot produce the scratch that appears when a new die goes in, or the burr pattern from a supplier's new coil, because those are not in any frame yet. The doubted frames from the review queue are where those arrive.
LexData takes the press model through its whole life. You type what to look for, Lexi puts a box on every frame, and a person checks each label before anything trains on it. The model then watches the inspection camera, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime. Each retrain runs the same augmentation over a set that now includes the new die's scratches, and the sticky note on the HMI gets a new line when the line gets a new way to vary.
Our manufacturing work holds 99%+ accuracy in production on cameras like this one, and the augmentation list behind each of them is short, specific, and copied from the line rather than from a tutorial.
See it on your own footage.
Start with your footageMore in Labeling

Labeling · 7 min read
AI data labeling workflows, three ways to label footage and when each one pays
A pipeline right-of-way survey labeled three ways: every box by hand, Lexi proposing and a person checking, and synthetic frames for the leak nobody has filmed.
Rajiya Sultana · Oct 2, 2026

Labeling · 6 min read
Annotation analytics, the numbers a labeling queue produces besides labels
Three labelers on a month of warehouse footage produce boxes, and also a throughput, an acceptance rate and a map of where the rejections cluster.
Rajiya Sultana · Oct 2, 2026

Labeling · 7 min read
Annotation format conversion between COCO, YOLO and CVAT without losing a box
Three years of line inspection labels from two tools arrive in three formats. The boxes that shift are the ones nobody draws on a frame before training.
Esdras Ntuyenabo · Oct 2, 2026