Labeling · 7 min read
Synthetic image data for the defect that happens once in a million units
A crushed carton corner too rare to photograph. Rendered variants across lighting and severity, mixed with real frames, evaluated only on real ones.
Summary
This post is about using synthetic frames for a defect on a packaging line that is too rare to photograph, a crushed carton corner from a case packer jam. It covers rendering variants across lighting and severity, mixing them with real frames rather than training on synthetic alone, evaluating only on real frames, and the review loop that catches what the synthetic frames taught that was false. It is for teams whose rare class has fewer examples than the model needs.
Sheikh Srijon · GTM Lead · Oct 2, 2026

Packaging line, cartons with labels past the scanner, generated scene with detections from our model
The case packer on packaging line 5 jams a few times a year, and when it does, one carton comes out with a crushed corner that the camera over the outfeed is supposed to catch before it is palletised. A few times a year, across a line running around the clock, is a defect in something like one carton in a million. The plant has photographs of four of them. The model needs hundreds.
Staging the defect is possible and slow: crush a carton, run it past the camera, repeat with the light in a different state. The quality engineer did that for an afternoon and got sixty frames, all under the afternoon light, all with the same carton.
What the model needs is the crushed corner under every condition the line produces, and synthetic frames are the only way to get there before the case packer jams a few hundred more times.
The defect is too rare to photograph and too costly to stage
The reason synthetic frames earn a place here and not on most lines is the arithmetic of the rare class. The lesson on how much footage you actually need says the practical floor for a class is a few hundred labeled instances, and that the rare class sets the schedule. On the outfeed camera of line 5 the rare class produces itself a few times a year. Waiting is a schedule measured in decades.
Staging gets part of the way, and the sixty frames from the afternoon are real and valuable. But they cover one carton, one light, one crush. The line has cartons in several sizes, a bay door that changes the light twice a shift, and crushes that range from a folded corner to a caved-in side. Sixty frames from one afternoon do not teach any of that.
The four photographs from the real jams sit in a folder named after the dates. Two of them were taken on a phone, at an angle the outfeed camera never sees.
Synthetic frames vary the lighting and the severity on purpose
A synthetic frame for line 5 starts with a real frame from the outfeed camera, a good carton passing the scanner, and puts a crushed corner on it. Rendered from a model of the carton, or generated from the staged examples, the crush is placed on the carton where the case packer would put it, at a severity chosen from a range, under the lighting the real frame already has. The result is a frame the camera could have taken, with a defect the camera has almost never seen.
The variation is the point. A few hundred synthetic frames that all show the same crush under the same light teach the model the same thing sixty staged frames did. The useful set walks through severity from a barely folded corner to a side caved in, across the carton sizes the line runs. It covers the morning light, the bay-door light and the night shift's sodium lamps, with the crush on the near corner and on the far one.
Each synthetic frame carries its own box, placed where the crush was put, so labeling costs nothing. That is the other reason synthetic frames are tempting, and it is where the first mistake usually hides.
Mix synthetic with real frames, and never train on synthetic alone
A model trained on synthetic crushed corners alone learns synthetic crushed corners. Whatever the rendering does that a real crush does not, a slightly too-clean edge, a shadow that always falls the same way, becomes the thing the model looks for. On the line, the real crush from the next jam has none of those tells, and the model walks past it.
So the synthetic frames go into the training set alongside every real frame the line has: the sixty staged ones, the four photographs that were taken from the right angle, and the good cartons the camera on line 5 sees all day, every day. The real frames anchor what a carton looks like. The synthetic ones supply the rare class in volume. The proportion is worth tuning, and a set that is mostly synthetic on the rare class is usually fine as long as the good class is entirely real.
My own view is that synthetic frames are a way to get a first model that can start doubting, and that they should be treated as scaffolding to be taken down as real examples arrive.
Evaluate only on real frames, or the number means nothing
The evaluation set never contains a synthetic frame. Recall on the crushed corner class is measured against the staged sixty, held out, and against every real jam line 5 produces from now on. A model that finds its own rendered crushes has proved that the renderer is consistent, which nobody doubted.
This is the rule most often broken, because the real evaluation set is tiny and the number that comes out of it is noisy. Sixty frames is a small test, and a recall figure on it moves a lot when one more frame is caught or missed. That is honest noise, and the alternative, a smooth number from a synthetic test set, is a smooth number about nothing. The surface defect detection use case lives with the same small real sets on every line it covers, and the small real number is the one that gets reported.
The engineer who ran the staging afternoon keeps the sixty frames in a folder marked "do not train". It is the most important folder in the project.
The review loop catches what synthetic taught that was false
The first model trained on the mixed set went on the outfeed camera and started returning the frames it was unsure of. Among them, within a fortnight, was a pattern: cartons with a printed fold line near the corner, perfectly intact, flagged as crushed. The synthetic crushes had all been rendered with a sharp crease at the edge of the crushed region, and the model had learned that a sharp crease near a corner means a crush. The printed fold line was a sharp crease near a corner.
Nobody would have found that by looking at the synthetic frames. It was found because a person on the line looked at the doubted frames, said "that carton is fine", and the correction went back. The renderer was adjusted to soften the crease, the false lesson was diluted by real corrections, and the next version stopped flagging the fold line.
LexData takes the outfeed model through its whole life. You type what to look for, Lexi puts a box on every frame, and a person checks each label before anything trains on it. The model then watches the outfeed camera, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime. For a model raised partly on synthetic frames, that review is where the synthetic and the real get reconciled, one correction at a time.
Synthetic frames are a bridge to the first real hundred
The case packer will keep jamming a few times a year, and each jam now produces a real frame with a real crush, caught or doubted by a model that has already seen roughly what to look for. Each one goes into the training set as a real example, and each one lets a few synthetic frames retire. The crushed-corner class on line 5 is on its way from four photographs to a real hundred, and it will get there in years rather than decades because the synthetic frames let the model start.
The folder of four photographs is still there. The two taken on a phone are still at the wrong angle.
See it on your own footage.
Start with your footageMore in Labeling

Labeling · 7 min read
AI data labeling workflows, three ways to label footage and when each one pays
A pipeline right-of-way survey labeled three ways: every box by hand, Lexi proposing and a person checking, and synthetic frames for the leak nobody has filmed.
Rajiya Sultana · Oct 2, 2026

Labeling · 6 min read
Annotation analytics, the numbers a labeling queue produces besides labels
Three labelers on a month of warehouse footage produce boxes, and also a throughput, an acceptance rate and a map of where the rejections cluster.
Rajiya Sultana · Oct 2, 2026

Labeling · 7 min read
Annotation format conversion between COCO, YOLO and CVAT without losing a box
Three years of line inspection labels from two tools arrive in three formats. The boxes that shift are the ones nobody draws on a frame before training.
Esdras Ntuyenabo · Oct 2, 2026