Labeling · 7 min read
Handling imbalanced classes when the defect is one part in a thousand
A pinched cable once in a thousand housings. Why the loss can afford to ignore it, how to collect the rare frames on purpose, and the one honest metric.
Summary
This post takes an assembly line where the defect that matters, a pinched cable at the housing edge, appears once in a thousand parts, and explains why a model trained on the line's footage learns to ignore it. It covers collecting the rare frames on purpose, oversampling and loss weighting within reason, augmenting the rare class while evaluating only on real frames, and per-class recall as the one number worth reporting. It is for the engineer whose model is accurate and useless.
Rob Hickey · Chief AI Officer · Oct 2, 2026

Assembly station, housing and cable boxed, generated scene with detections from our model
The camera over station 4 on an assembly line sees a housing go past every few seconds with a cable routed along its edge. About once in a thousand housings the cable is pinched under the lid, and a pinched cable is a warranty return six months later. The first model trained on a week of the line's footage reported an accuracy that looked wonderful. It had found roughly one pinched cable. It had looked at a week.
An accuracy figure on this line is almost entirely a measure of how well the model recognises a good housing, which was never in question.
The loss can afford to ignore the class that matters
Training works by lowering a loss, and the loss is added up over every example. On station 4's footage, the examples are overwhelmingly good housings, and a model that says "good" to everything gets nearly all of them right and pays a tiny penalty for the handful of pinched cables it missed. From the loss's point of view, that is an excellent model. The rare class contributes so little to the total that the training can afford to ignore it, and it does.
The result is a model that has learned the good housing in exquisite detail and the pinched cable barely at all. Its confidence on the defect, on the rare occasions it fires, is low, and its boxes on the defect are hesitant, because it has seen a dozen examples against a week's worth of the other thing.
The quality engineer on station 4 keeps the pinched housings from the last year on a shelf in the QA office. There are eleven of them, each with a tag. That shelf turned out to be the most valuable thing in the plant for this project.
Per-class recall is the only honest number on this line
Before fixing the model, fix the number. Overall accuracy on station 4 is the fraction of housings the model got right, and since nearly every housing is good, a model that never fires scores nearly perfectly. The number that describes what the line actually needs is recall on the pinched cable class alone: of the pinched cables that went past, how many did the model catch.
That number was low for the first model, and it was the number nobody had looked at. Precision on the rare class matters too, since a model that flags every tenth housing to catch every pinch is a model the line will switch off, but recall is the one that maps to the warranty return. Report it per class, report it on real frames, and stop reporting the overall figure to anyone who might mistake it for progress.
The surface defect detection use case has the same shape on every line it covers: the defect is rare by design, and the metric has to be the rare class's own.
Collect the rare frames on purpose, from the reject bin
The single most effective thing is also the least clever: go and get more examples of the defect. The eleven housings on the QA shelf go back under the camera at station 4, one at a time, in the lighting the line has, rotated through the positions a housing takes on the belt. That is a few hundred frames of the rare class from an afternoon, where the line would have taken a year to produce them.
The reject bin is the second source. Every pinched cable an inspector has caught in the last quarter went somewhere, and the frames from the moment it passed the camera are on the recorder. The lesson on how much footage you actually need says the timeline is set by the least common class, and this is what that means in practice: the rare class's timeline is set by how fast someone can find or stage examples of it.
None of this is a substitute for the line's own footage. It is the fastest way to get the rare class from a dozen examples to a few hundred, which is where a model starts to have something to learn from.
Oversample the rare class and weight its loss, within reason
With a few hundred pinched-cable frames in hand against a week of good housings from station 4, the ratio is still lopsided, and two training levers help. Oversampling shows the rare frames more often during training, so the model sees a pinched cable as frequently as it needs to rather than as rarely as the line produces one. Loss weighting makes each miss on the rare class cost more, so the training can no longer afford to ignore it.
Both have a limit. Oversample too hard and the model memorises the few hundred rare frames rather than learning what a pinch looks like, and its recall on a new pinch it has never seen is no better than before. Weight the loss too hard and the model starts seeing pinches in shadows. The right setting is found by watching per-class recall and precision on a held-out set of real frames, and moving one lever at a time.
My own view is that oversampling is the safer of the two on a line camera, because it changes what the model sees rather than what it is punished for. A model that has seen the rare class often is easier to reason about than one that has been shouted at about it.
Augment the rare frames, and evaluate only on real ones
The few hundred frames staged at station 4 can be widened by augmentation: brightness shifts for the bay door opening, small rotations for the housing sitting askew on the belt, blur for the belt speeding up. That gives the model the rare class under the conditions the line will present, which the afternoon of staging could not cover on its own.
The rule that keeps this honest is that the augmented frames never appear in evaluation. Per-class recall is measured on real pinched cables from the line and nothing else.
A model that catches its own augmented copies has proved nothing. A model that catches the eleven real ones from the QA shelf, held out from training, has proved something.
The defect rate moves, and the thresholds move with it
Once the model is catching pinched cables, a different problem arrives. A supplier changes the cable jacket, or a fixture wears, and the defect rate on station 4 goes from one in a thousand to one in two hundred for a fortnight. The model is finding them correctly. The alert volume triples, and the line lead thinks the model has broken.
The drift catalog calls this changing defect rates, and the point of it is that nothing about the model or the images moved; the base rate did, and every threshold tuned around the old rate is now wrong. The fix is to re-tune the threshold against the current rate before touching the model, and to notice that a defect rate moving is an operations finding worth raising on its own.
LexData takes the station 4 model through its whole life. You type what to look for, Lexi puts a box on every frame, and a person checks each label before anything trains on it. The model then watches the camera on the line, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime.
On a rare class, the doubted frames are where the next real examples come from, since the model doubts most on exactly the housings that look a little like a pinch. Our manufacturing work holds 99%+ accuracy in production on lines like this one, and on this line the number that proves it is recall on eleven housings and whatever joins them on the shelf.
See it on your own footage.
Start with your footageMore in Labeling

Labeling · 7 min read
AI data labeling workflows, three ways to label footage and when each one pays
A pipeline right-of-way survey labeled three ways: every box by hand, Lexi proposing and a person checking, and synthetic frames for the leak nobody has filmed.
Rajiya Sultana · Oct 2, 2026

Labeling · 6 min read
Annotation analytics, the numbers a labeling queue produces besides labels
Three labelers on a month of warehouse footage produce boxes, and also a throughput, an acceptance rate and a map of where the rejections cluster.
Rajiya Sultana · Oct 2, 2026

Labeling · 7 min read
Annotation format conversion between COCO, YOLO and CVAT without losing a box
Three years of line inspection labels from two tools arrive in three formats. The boxes that shift are the ones nobody draws on a frame before training.
Esdras Ntuyenabo · Oct 2, 2026