Skip to content
LexDataLexData
PlatformIndustriesCustomers
DocsThe Field GuideBlogWhy models drift
AboutCareersSecurityContact
Log inStart now
← All posts

Labeling · 6 min read

What a model trained on ten frames is good for

Ten labeled frames from a new site camera give a first pass by the end of the afternoon. Trust it for the obvious, and let its doubts grow the set.

Summary

This post takes a new camera over a construction yard and the ten frames a person labels on the first afternoon, and asks what the model trained on them can and cannot be trusted for. It concludes that a ten-frame model is a labeling assistant rather than a watcher, that its doubted frames are the shortest path to a set that holds, and that a season is the thing ten frames can never cover. It is for teams standing up a new camera.

Sheikh Srijon · GTM Lead · Sep 29, 2026

Scaffold tower beside site containers, workers, ladders and plank decks boxed, from a customer site camera

The new camera over the site's container yard went up on a Tuesday morning, and by 2 pm a person had labeled ten frames from it: workers, hard hats, the scaffold tower against the containers, the ladder leaning where it should not. By the end of the afternoon there was a model. It found the workers in the next frame it was shown, drew a decent box on the tower, and boxed a stack of orange fencing as a worker in hi-vis.

That is roughly what ten frames buy, and it is more than it sounds. Knowing exactly what it buys is what stops a team from either dismissing the afternoon's model or trusting it with the gate.

Few shot learning starts from a model that has already seen a lot

A model trained from nothing on ten frames would learn nothing. Few shot learning works because the model does not start from nothing. It starts from weights that have already seen a great many pictures of the world, and the ten frames adjust that general sense of edges, people and shapes toward this camera's view of this yard. The ten Tuesday frames teach it what a worker looks like from this pole at this height, and the pretrained part supplies everything else.

Which is why the first afternoon's model is good at the things the pretraining already covered, people and vehicles, and weak on the things only this site has. Those are the yard's particular fencing, the way the tower's netting reads from above, the difference between a worker and a coat on a hook.

Ten frames can be trusted to find the obvious and nothing else

The afternoon model finds workers in plain view in daylight, because it had ten frames of exactly that. It misses the worker half behind a container, because none of the ten frames had one. It boxes the orange fencing because orange and vertical was the whole of what it learned about hi-vis. Its recall on the obvious cases is fine and its precision on anything unusual is poor, and both of those follow directly from what the ten frames contained.

A ten-frame model, in other words, is a summary of ten frames. What it is trusted for should be limited to the conditions those frames showed: the Tuesday afternoon, the daylight, the yard as it was arranged that week. Anything outside that is a guess wearing a box.

The first pass is a labeling assistant rather than a watcher

The right job for the afternoon model is to draw the first boxes on the next few hundred frames, from Wednesday onward, so that a person checks rather than draws. That is a large saving even when the boxes are rough, because moving a box is quicker than placing one, and the person's attention goes to the frames where the model was wrong. The quickstart covers the same idea from the platform's side. Lexi puts a first box on every frame from a phrase, and a person checks each one before it trains anything. The ten-frame model is that first pass with the site's own ten frames folded in.

The wrong job for it is the gate rule. A model that boxes fencing as a worker will fire a hazard alert on a stack of fencing, and the site will have turned the alerts off before the model has seen its eleventh frame.

The doubted frames are where the set grows

The afternoon model's most useful output is its doubts. The frames it was unsure of, the half-hidden worker, the coat on the hook, the tower's netting in the low sun at 4 pm, are exactly the frames the ten did not cover, and labeling those next grows the set where it is thinnest. In my experience the ten-frame model's main value is telling you which ten frames to label next. A team that labels its doubts in order builds a set that holds far sooner than a team that labels the next hundred frames in the order the camera recorded them.

The lesson on how much footage you actually need makes the same point in general: distinct conditions are the unit, and the doubted frames are a list of the conditions the set is missing.

The yard foreman, for what it is worth, stacks the containers by colour, and the afternoon model learned the colour order as if it were a law of the world.

The set stops being few-shot when the corrections cross the threshold

From the afternoon model onward the process is the ordinary one. LexData takes the yard model through its whole life. You type what to look for, Lexi puts a box on every frame, and a person checks each label before anything trains on it. The model then watches the cameras the site already has, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime.

The ten frames become a few hundred within the first week, drawn from the model's own doubts, and when the corrections cross the project's threshold a new version is trained on them. A camera can be live in days this way, with the first version doing the labeling and the second version doing the watching, and the moment it stops being a few-shot model is the moment nobody is counting the frames any more.

What ten frames cannot cover is a season

The one thing no amount of clever training gets from ten Tuesday frames is October. The yard in the rain, the tower under floodlights on the first dark afternoon, the frost that turns the containers grey: none of it was in the ten, and none of it can be inferred from them. A ten-frame model that has grown to a few hundred over a summer will still meet its first wet week as a stranger, and the correction rate that week is the signal to label the rain.

That is the honest limit. Few frames start a model. A season of doubted frames, checked by a person, is what makes it hold.

See it on your own footage.

Start with your footage

More in Labeling

Labeling · 6 min read

COCO as a format you will use and a benchmark you should not trust

COCO JSON is the file a bottling line's labels travel in. The COCO benchmark is a score on somebody else's eighty classes, and none of them is a missing cap.

Sheikh Srijon · Sep 29, 2026

Labeling · 6 min read

EXIF orientation, the photo that is sideways only to the model

A phone photo looks upright on every screen and arrives rotated in training, because the pixels never turned. The check to run at import, before the first box.

Stephen Biswas · Sep 29, 2026

Labeling · 7 min read

Choosing an annotation tool for a month of inspection footage

The demo set labels itself in an afternoon. A month of drone footage is where the tool has to propose boxes, take corrections and review every label.

Sheikh Srijon · Sep 29, 2026

LexData
LexData

Product

  • Platform
  • Industries
  • Use cases

Resources

  • Docs
  • The Field Guide
  • Blog
  • Why models drift

Industries

  • Energy & utilities
  • Oil & gas
  • Agriculture
  • Manufacturing
  • Insurance
  • Retail
  • Robotics

Company

  • About
  • Customers
  • Careers
  • Contact

Trust

  • Security
  • Privacy
  • Terms

Stay updated

What we learn running vision models in production.

See everything.
Miss nothing.

Stay updated

What we learn running vision models in production.

Terms of use & Privacy policy

© 2026 LexData Labs · All rights reserved