Skip to content
LexDataLexData
PlatformIndustriesCustomers
DocsThe Field GuideBlogWhy models drift
AboutCareersSecurityContact
Log inStart now
← All posts

Labeling · 6 min read

Image preprocessing vs data augmentation, and why only one runs at the camera

Resize, orientation and normalisation have to match between the labeled frames and the runner. Augmentation runs in training only, never at the camera.

Summary

This post separates preprocessing, the resize, orientation and normalisation that must be applied identically to the labeled frames and to the frames the runner feeds the model, from augmentation, which exists only to widen the training set and never runs at the camera. It uses a bottling line whose model quietly degraded because the runner prepared frames differently from the labeled ones, and shows the comparison that found it. It is for anyone putting a trained model on their own hardware.

Stephen Biswas · Engineer · Oct 2, 2026

Bottling line, bottles queued under the fill head, generated scene with detections from our model

The model over bottling line 2 found every missing cap for three months in the cloud. Then the plant moved it to a runner beside the recorder, as planned, and within a week the corrections in the review queue had doubled. The bottles were the same. The camera was the same. The model file was the same file, checked byte for byte.

What had changed was a line of code nobody thought of as part of the model: the runner resized each frame by stretching it to a square, and the frames the model had been trained on had been resized by fitting and padding. Every bottle the model now saw was a little too wide.

That line of code is preprocessing, and the mistake is treating it as an implementation detail rather than as part of the model's definition.

Preprocessing is a contract between labeled frames and the camera

A trained model expects its input in a particular shape: a particular size, a particular orientation, pixel values scaled a particular way, channels in a particular order. The labeled frames it trained on were put into that shape by a preprocessing step. At inference, the frames from the camera have to be put into exactly the same shape by exactly the same step, or the model is being asked a question in a language slightly different from the one it learned.

That makes preprocessing a contract rather than a convenience. Whatever was done to the frames before labeling and training is what has to be done to the frames from the camera on line 2, on every sample, for as long as that version runs. Change one side and the other side is wrong, without an error, and the model degrades in the direction of whatever moved.

The controls engineer on line 2 has since written the resize method on the same masking tape as the model version. It is the kind of thing that only seems obvious after a week of extra corrections.

Resize, orientation and normalisation have to match on the runner

Three things go into that contract on nearly every model, and each has its own way of failing quietly.

The resize. A camera frame at 1920x1080 goes into a square input. Fit and pad keeps every bottle the shape it was and adds black bars. Stretch makes every bottle wider. Crop loses the edges of the line. The model was trained on one of these, and the runner has to do the same one, with the same interpolation.

The orientation. A frame with a rotation tag in its metadata is shown upright by a viewer and read on its side by a runtime that ignores the tag. The labels were drawn upright.

The normalisation. Pixel values in the range of zero to a maximum, or scaled to a unit range, or shifted by a per-channel mean, in red-green-blue order or the reverse. Get any of that wrong and the model receives images it has never seen the like of, and produces confident nothing.

The deployment guide is where those choices are pinned to the device preset, so a model exported for a Jetson, a Raspberry Pi or a GPU server carries its own preparation with it. The failure on line 2 happened because the plant wrote its own runner loop around the file, and the loop resized its own way.

Data augmentation runs in training and never on the runner

Augmentation is a different thing that gets confused with preprocessing because it also changes pixels. In training, each labeled frame is copied with a random change, a flip, a shift in brightness, a small rotation, a blur, and the model sees the copies as well as the original. The point is to widen the set: a few thousand bottling frames become a few thousand frames plus all the versions of them under lighting the camera might meet next winter.

Data augmentation never runs at the camera. The runner feeds the model the frame as the camera took it, preprocessed by the contract and nothing else. Flipping the frame from line 2 at inference would show the model a mirror image of the line for no reason, and brightening it would move it away from the frames it was trained on rather than toward them.

The rule is short: preprocessing is applied to every frame on both sides; augmentation is applied to training copies on one side. Anything applied on the runner that was not applied to the labeled frames is a bug, however sensible it looks.

Augment for the variation the line will meet

Which augmentations to use is a question about the line rather than about the model. A fixed camera over a conveyor has a known set of ways it varies: the light changes when the bay door opens, the belt speed changes the blur, maintenance swaps a bulb and the colour temperature shifts. Those are the augmentations worth applying, because they are the conditions the camera will produce.

A vertical flip is not one of them. Bottles on line 2 are never upside down, and teaching the model that they might be spends capacity on a world that does not exist. The lesson on how much footage you actually need makes the same point from the other direction: the unit that matters is distinct conditions. Augmentation manufactures the conditions the footage is short of. It should never invent ones the line does not have.

My own view is that most teams over-augment and under-collect. A brightness shift is cheap and a week of frames from the night shift is not, and the second is the one that fixes the night shift.

The mismatch shows up as corrections, and a frame comparison finds it

What made the bottling line's problem visible was the review queue. The model returned more frames it was unsure of, the person checking them corrected more boxes, and the override rate rose in a week from a level that had held for three months. Nothing in the world had changed, which pointed at the pipeline.

The test that found it is the same one that finds any preprocessing mismatch. Take one frame from the camera on line 2, run it through the training side's preprocessing and the runner's, and compare the two prepared images before they reach the model. On line 2 they were different sizes in one axis, and the whole story was in that difference.

LexData takes the cap model through its whole life. You type what to look for, Lexi puts a box on every frame, and a person checks each label before anything trains on it. The model then watches the line camera, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime. On line 2 the corrections did their job in a different way: they were the alarm, and the fix was a line of code rather than a retrain.

See it on your own footage.

Start with your footage

More in Labeling

Labeling · 7 min read

AI data labeling workflows, three ways to label footage and when each one pays

A pipeline right-of-way survey labeled three ways: every box by hand, Lexi proposing and a person checking, and synthetic frames for the leak nobody has filmed.

Rajiya Sultana · Oct 2, 2026

Labeling · 6 min read

Annotation analytics, the numbers a labeling queue produces besides labels

Three labelers on a month of warehouse footage produce boxes, and also a throughput, an acceptance rate and a map of where the rejections cluster.

Rajiya Sultana · Oct 2, 2026

Labeling · 7 min read

Annotation format conversion between COCO, YOLO and CVAT without losing a box

Three years of line inspection labels from two tools arrive in three formats. The boxes that shift are the ones nobody draws on a frame before training.

Esdras Ntuyenabo · Oct 2, 2026

LexData
LexData

Product

  • Platform
  • Industries
  • Use cases

Resources

  • Docs
  • The Field Guide
  • Blog
  • Why models drift

Industries

  • Energy & utilities
  • Oil & gas
  • Agriculture
  • Manufacturing
  • Insurance
  • Retail
  • Robotics

Company

  • About
  • Customers
  • Careers
  • Contact

Trust

  • Security
  • Privacy
  • Terms

Stay updated

What we learn running vision models in production.

See everything.
Miss nothing.

Stay updated

What we learn running vision models in production.

Terms of use & Privacy policy

© 2026 LexData Labs · All rights reserved