Operations · 6 min read
Underspecification in machine learning, and why two models with the same score do not both survive the second plant
Fifty training runs, one recipe, one validation score. The test set could not tell them apart, and the second plant's cameras can.
Summary
This post explains underspecification through fifty candidate weld inspection models with the same validation score and a second plant with newer cameras where only some of them work. It concludes that the test set could not separate them because it came from the first plant, that stress tests and per-site checks split them before rollout, and that a labeled window from the new site is what settles it. It is for the people choosing which trained model goes live.
Rob Hickey · Chief AI Officer · Sep 30, 2026

A robotic welding cell with a part on the fixture, generated scene with detections from our model
The weld inspection model at the first plant was trained fifty times. Same frames, same labels, same recipe, a different random seed each run, because the team wanted to know how much the seed mattered. The answer on the held-out set was that it did not: every one of the fifty scored within a hair of the others. They named the runs after the weather on the day and picked the one from a Tuesday in May, because it was fractionally top and someone had to choose.
The second plant went live in September with newer cameras. Tuesday in May was wrong on a third of its welds. Two of the other forty-nine, tried in a hurry, were fine.
The test set could not tell the fifty apart
The held-out set at the first plant was drawn from the same cameras, the same fixture, the same lighting as the training set. It is a fair test of how well a model learned that plant. It says nothing about which of many different ways of learning that plant each model found.
There are many. The frames from cell 3 contain the weld bead, and they also contain the fixture's clamp, a smear of spatter that lands in the same place every cycle, and the brightness of the torch reflection at the end of each pass. A model can find a bad weld by looking at the bead, or by noticing the spatter pattern that a bad weld tends to leave, or by the reflection. All three work at the first plant. Only the first works at the second, where the fixture is different and the newer cameras handle the torch glare differently. The validation score gave all three the same mark.
Overfitting is the wrong name for this
The natural reaction is to call it overfitting and add regularisation. Overfitting is when a model has learned the training set so closely that it fails on frames it has not seen from the same distribution, and the held-out score catches it. The fifty models did not overfit by that test. They generalised well to every frame from cell 3 at the first plant.
What they did was fit an underspecified problem. The frames did not contain enough variation to force the model to use the bead rather than the spatter, so the seed decided, and the seed is not something anyone reviews. The distinction matters because the fixes are different: more regularisation shrinks all fifty equally, and the one that learned the spatter stays as good as it was on the first plant's frames.
The team's own log, incidentally, recorded the seed for every run and the weather for none of them, which is why the names were the only way anyone could remember which was which.
Stress tests split the models the validation set could not
If the held-out set cannot separate the fifty, something else has to before rollout. The tests that do it all push the frames away from the first plant in a direction the second plant might. Perturb the held-out frames: blur, darken, shift the colour, re-encode at a lower bitrate. Mask the fixture and the spatter region and score again. Stratify the score by cell and by shift, cell 3 against cell 4, days against nights, and look at the spread rather than the mean.
Under those tests the fifty stop agreeing. The model that learned the spatter falls apart when the spatter region is masked. The one that learned the torch reflection collapses under a colour shift. The ones that learned the bead hold. The spread across the fifty seeds under each test is itself a number, and my own view is that it belongs next to every validation score a team reports, because a single score from a single seed is a report of luck.
The number beside a detection does not settle it either
A tempting shortcut, and the one the team reached for in September, is to trust the model that seems surest. It is the wrong shortcut. The number beside every detection is a ranking, an ordering of this detection against that one on frames like the training set. The field guide's lesson on what that number means is blunt about the rest: it is not a probability of being right, and it shifts with the class, the camera and the season. The model that learned the spatter is as sure of itself at the second plant as it was at the first. It has no way to know it is looking at the wrong thing.
The second plant is the test that counts
The drift catalog files September's failure under a new site came online. Same task, same definitions, a different input distribution arriving all at once, with the second plant's detections differing from the first plant's from the first day rather than degrading into it. The prescription is a labeled window from the new site before it goes live, and a per-site comparison rather than a fleet average, since the first plant's good numbers hide the second's.
That window does two jobs here. It is the training set's first taste of the newer cameras, so the bead becomes the only thing that works across both plants. And it is the test set that finally separates the candidates: score the fifty on the second plant's window before any of them is chosen, and the choice stops being a coin toss named after the weather.
LexData takes the weld model through its whole life. You type what to look for, Lexi puts a box on every frame, and a person checks each label before anything trains on it. The model then watches the cameras each plant already has, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime. The welds from cell 3 at the second plant that came back for review in the first week were the frames that told the team which of its fifty models had been looking at the bead all along.
See it on your own footage.
Start with your footageMore in Operations

Operations · 6 min read
Camera calibration for computer vision, and why the part at the edge of the frame measures wrong
A straight edge bows at the corner of the frame, so a part that passes in the centre fails at the edge. Calibrate once, and again the day the lens changes.
Stephen Biswas · Sep 30, 2026

Operations · 6 min read
Evaluating a computer vision platform once the pilot ends
The feature spreadsheet cannot tell you which platform survives the second year. Walk the lifecycle instead, from import to a retrain that keeps the old frames.
Ayman Quadir · Sep 30, 2026

Operations · 6 min read
Face blurring for privacy in computer vision, done before the frame is stored
A blur applied after the request arrives is a blur applied to a frame that has already been copied. The step belongs at ingestion, beside the recorder.
Esdras Ntuyenabo · Sep 30, 2026