Labeling · 6 min read
How much training data a computer vision model needs from your own cameras
A weld defect detector got useful on a few hundred chosen frames. A sixty-class shelf model needed far more. The difference is the task and the variation.
Summary
This post compares two first models, a weld defect detector that became useful on a few hundred chosen frames and a shelf model with sixty classes that needed far more, and explains the difference through learning curves, task complexity and coverage of the variation the camera will see. It concludes that the first set only has to be good enough to start the loop, because the loop adds the frames that matter. It is for teams sizing their first labeling run.
Sheikh Srijon · GTM Lead · Sep 28, 2026

Robotic welding cell, the part on its fixture and the arm boxed, generated scene with detections from our model
The welding cell at station 4 had a camera on the fixture for a year before anyone trained anything on it. When the quality team finally did, they labeled a few hundred frames, chosen rather than sampled, and the first version caught the porosity the inspector was catching and a few he was not. Two aisles over in a different business, a shelf model with sixty product classes trained on ten times as many frames and was still confusing the two sizes of the same yoghurt at launch.
Both teams asked the same question at the start, how many frames do we need, and both got the same unhelpful answer: it depends. What it depends on is worth being specific about.
A weld defect detector got useful on a few hundred chosen frames
The welding cell at station 4 is an easy case, and it is worth saying why. One camera, fixed. One part, on a fixture that puts it in the same place every time. Consistent lighting from the cell's own lamps. Two classes, a good weld and a porous one, and the porous one has a visual signature the inspector can point at. The variation the model has to learn is small, because the cell was designed to remove variation.
So a few hundred frames covered it. The frames were chosen to include every shift, the two operators who set the fixture slightly differently, the day after the lens was cleaned, and every porous weld in the scrap log. A random few hundred would have been mostly good welds from the day shift, and the first version would have been worse.
The footage lesson makes the point in one line: the unit is distinct conditions, and hours of the same condition count as one.
The learning curve flattens where the task lets it
Plot accuracy against the number of labeled frames and the curve rises steeply, then bends, then goes nearly flat. The place it bends is the number worth knowing, and it moves with the task. For the welding cell it bent early, because after a few hundred frames the model had seen every way a porous weld looks on that fixture and more frames of the same thing added little.
The useful practice is to train on a quarter of the set, then half, then all of it, and look at the three results. If the last doubling still moved the accuracy, more frames will help. If it barely moved, more frames of the same kind will not, and the next frames should be different rather than more.
At station 4 the curve from half to all was almost flat. The team stopped labeling and put the model on the cell.
Pose estimation and segmentation need more per frame than a box does
Task complexity is the second axis. A box on a weld at station 4 carries one decision: where. A polygon around the porous region carries the shape too, and the model needs more examples to learn a shape than a position. A pose estimation task, keypoints on a part so a robot can pick it, carries several decisions per instance, and each keypoint has its own learning curve.
The rule of thumb that holds up is that the frame count scales with how much each label says. Boxes first, and upgrade the classes that prove they need more. A team that starts with polygons on everything pays for shape on classes where position would have done.
An aside from the welding cell: the polygon the team eventually drew around porosity was never used by the model. The inspector wanted to see the region drawn in the review queue, and a box was enough for the decision.
A shelf model with sixty classes needs far more, and the classes are why
The shelf camera in aisle 6 is the hard case. Sixty products, some of which differ by a word on the label. Shoppers in the way. Light that changes with the time of day and the season through the front window. Products that move between shelves when the planogram changes. Each class needs its own few hundred instances, so sixty classes need sixty times the instances, and the small yoghurt and the large yoghurt need more than most because the model has to learn a difference the camera barely shows.
Instances, though, and not frames. A frame of a full shelf holds dozens of instances across many classes, so the frame count is lower than the multiplication suggests. It is still far above the welding cell, and the classes that are rare on the shelf, the ones that sell out or sit on the bottom row, are the ones that hold the whole model back.
Class imbalance is where the shelf team spent its second month. The common products had a great many instances and the rare ones had a handful, and the model was excellent on the front row and guessing on the bottom. Pulling extra frames of the rare products, and labeling every instance in them, did more than doubling the whole set would have.
Coverage of what the camera will see matters more than the count
Both models were trained for one camera, and the question that decides the frame count is what that camera will see over a year. The welding cell will see the same fixture under the same lamps. The shelf camera will see a promotion end cap in November, a new packaging design in March, and the afternoon sun for three weeks in June. The shelf set has to cover all of that, or the model will meet each one for the first time in production.
My own view is that a team should list the conditions before counting the frames. Write down every shift, season, lighting change, product change and camera event the camera will see, then make sure the set has frames from each. The count falls out of the list, and a set built from the list is smaller and better than one built from a target number.
The first set only has to be good enough to start the loop
Neither team's first set was the model's last. The welding cell model was on station 4 within days, and the frames it was unsure of came back to the inspector: a new wire batch that spattered differently, and the week the cell's lamp was dimming before anyone replaced it. His corrections retrained it, and the new version replaced the old one with no downtime.
The shelf model went the same way, with a longer queue. The new packaging in March arrived as a run of doubted frames, a person confirmed the new design was the same product, and the correction went into the next version rather than into a new labeling project.
That is what the frame count is really for: enough to get a first version onto the camera, where the loop does the rest. The set the model has seen by the end of the year is several times the one it launched with, and almost none of it was labeled up front.
See it on your own footage.
Start with your footageMore in Labeling

Labeling · 6 min read
AI labeled vs human labeled data, how much a model can do before a person has to look
On a shelf dataset the model's proposals are mostly right. On safety cones in a yard they are mostly wrong. Measure the gap, then keep a person on every label.
Finn Ellingwood · Sep 28, 2026

Labeling · 6 min read
Bounding boxes in computer vision, what a box teaches and what it reports
The same rectangle on a forklift is the lesson at training time and the answer at run time. Corner and centre formats, the shifted-box bug, floor in the box.
Esdras Ntuyenabo · Sep 28, 2026

Labeling · 6 min read
Which words find the forklift, measured instead of guessed
Five ways to say forklift, five sets of boxes on aisle 6. Score each phrase on a small labeled set before Lexi labels the whole archive with the winner.
Sheikh Srijon · Sep 28, 2026