Computer vision · 6 min read
Reading a training run before trusting the model
The loss curve says the run finished. The gap to validation, the recall on the crack class and the frames behind the score say whether the line can trust it.
Summary
This post reads the output of a training run for a panel-inspection model on a stamping line, from the loss curves to the per-class precision and recall and the gap between training and validation. It concludes that the validation frames decide what the score means, that the rare class is the row to read first, and that the training result becomes the baseline the operator correction rate is compared with after go-live. It is for engineers deciding whether a run is ready to watch a camera.
Rob Hickey · Chief AI Officer · Sep 28, 2026

Stamping line with steel panels passing an inspection station, press and conveyor boxed, generated scene with detections from our model
The training run for the panel model on the stamping line finished at 4 pm on a Wednesday, and the dashboard showed what dashboards show: two loss curves sloping down, a mean average precision that had climbed and levelled, and a green tick. The line lead, who had been asked to approve the model for the night shift, looked at it and asked whether it was good.
The honest answer has several parts, and the green tick is the least informative of them.
The loss curve says the training ran and little else
Loss is the number the training procedure was minimising, and a curve that slopes down and flattens says the procedure did its job on the frames it was given. It does not say the model finds dents. A loss can fall steadily while the model learns that Wednesday's panels are mostly clean and the safest guess is "clean", which is a perfectly good way to lower a loss and a useless model for the line.
A curve that never flattens says the run stopped early. A curve that drops to almost nothing says something worse, which the next section is about. Beyond that, the loss curve is a health check on the run, and the sign-off decision needs the numbers that were measured on frames the run never touched.
The gap between training and validation is the tell
Every run reports two versions of its metrics: one on the frames it trained on, one on a validation set it was kept away from. The training figure is always the flattering one. The gap between them is the number that tells you whether the model learned dents or memorised Wednesday.
A small gap with both figures healthy is what a sign-off wants to see. A large gap, training near perfect and validation well below it, is a model that has learned the specific panels it was shown, their scratches, the exact position of the press's shadow at 4 pm, and will meet a new panel as a stranger. The usual cause on a line is a training set that is too small or too uniform: a single shift, a single coil of steel, a single lighting state. The cure is frames, from the other shifts and the other coils, rather than any setting on the run.
Watch the validation curve over the run as well as its final value. If it rose and then began to fall while the training curve kept climbing, the best model was the one from the middle of the run, and most training tools keep it.
Precision and recall are read per class, rarest first
The mean average precision on the dashboard is an average over the model's classes, and the stamping line's Wednesday run has three: panel, dent and crack. Panels appear on every frame and the model will find them. Dents appear a few times a shift and cracks a few times a week, and the mean is dominated by the class nobody is worried about.
So read precision and recall per class, starting with the rarest. For the crack class, recall is the number that matters: of the cracks in the validation set, how many did the model box. Precision says how often a flagged crack was really a scratch or a smear of drawing oil. A model with high precision and low recall on cracks is cautious, and cautious on a crack detector means the crack goes past. That is the lopsidedness to look for, and the mean hides it completely.
The press operator on days, incidentally, still chalks a mark on every panel he rejects by eye, and the chalk marks are the cheapest validation labels on the line.
Per-class recall on the rare class is the number the line signs off
For an imbalanced set, one row of one table decides the go-live: recall on the class that costs the most to miss, at the threshold the rule will run at. The number the model puts beside each detection is a ranking of that box against the others, and the lesson on what the number beside a detection means explains why the same threshold behaves differently on cracks than on panels. Sweep it on the validation frames, find the position where crack recall is high enough that the line lead accepts the misses, and write down what precision costs at that position. That pair of numbers is the sign-off.
If the validation set holds only nine cracks from the Wednesday run, the recall figure is a rumour and the fix is to find more cracked panels before trusting the run. On the manufacturing lines we run, the crack frames are the ones a person goes looking for in the archive before training starts, because the model will never be better at cracks than the set is.
The validation frames decide what the score is about
Every number above is a number on the validation set, and the set was assembled by someone. If it holds Wednesday's day shift, the score is about Wednesday's day shift. The night shift, when half the overhead lights are off and the press's shadow moves across the inspection station, is the condition the model will meet first and was tested on least.
Build the validation set the way the model will be used: every camera it will watch, every shift, a few weeks apart from the training frames so that near-identical neighbours of a training frame do not leak in and flatter the score. Lexi puts a first box on each validation frame and a person checks every one, and the person's care on the crack boxes is the difference between a recall figure and a guess.
My own view is that a training run whose validation set nobody in the room can describe should not be signed off, whatever its numbers. The numbers describe the set before they describe the model.
The training result is the baseline the correction rate is compared with
Once the model is on the line, the sign-off numbers stop being predictions and become a baseline. What is measured from then on is how often a person disagrees with the model: frames coming back for review, corrections landing, the operator marking a flagged crack as oil. When that rate climbs on the night shift and stays flat on days, the model has met a condition the validation set did not hold.
LexData takes the panel model through its whole life. You type what to look for, Lexi puts a box on every frame, and a person checks each label before anything trains on it. The model then watches the cameras the line already has, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime.
The alert written as a sentence carries the threshold the sign-off chose, and the override rate on that alert is the live version of the precision figure on the dashboard. The chalk marks on the reject rack are the live version of recall.
See it on your own footage.
Start with your footageMore in Computer vision

Computer vision · 6 min read
When the outline is the answer and a box will not do
A weld defect whose area decides the rework and a part a robot has to grasp both need a mask. Here is what the mask costs to label and when a box is enough.
Esdras Ntuyenabo · Sep 28, 2026

Computer vision · 6 min read
DETR explained, what the detection transformer changed about finding objects
Detection as set prediction, a fixed set of queries instead of anchors, no suppression step, and the small-object and slow-training problems later fixed.
Finn Ellingwood · Sep 28, 2026

Computer vision · 6 min read
What mAP measures and what it cannot tell you about your cameras
Mean average precision is built from per-class curves at an overlap line. A high score on the held-out set says nothing about the fence camera at 2 am.
Stephen Biswas · Sep 28, 2026