Skip to content
LexDataLexData
PlatformIndustriesCustomers
DocsThe Field GuideBlogWhy models drift
AboutCareersSecurityContact
Log inStart now
← All posts

Labeling · 6 min read

Video annotation, turning an hour of line footage into a training set

An hour from the bottling line is a hundred thousand near-identical frames. Sample every couple of seconds, label the ones that differ, track the rest.

Summary

This post takes one hour of footage from a bottling line camera and turns it into a training set, from sampling a frame every couple of seconds to choosing which frames to label, tracking an object across frames instead of boxing it fresh each time, and checking the result. It concludes that the labeled set is not finished when the model trains, because the same line keeps sending frames back. It is for teams with a recorder full of footage and no dataset.

Rajiya Sultana · Engineering Manager · Sep 28, 2026

Bottling line, bottles queued under the fill head and one missing its cap, generated scene with detections from our model

The recorder above bottling line 3 holds six weeks of footage, and the quality lead has been asked to turn it into a dataset for a cap detector. She opens the first hour. At thirty frames a second it is over a hundred thousand frames, almost all of them bottles moving under the fill head at the same speed under the same lights. The one thing she needs, a bottle with no cap, appears for about a second and a half every twenty minutes.

Labeling that hour frame by frame would take longer than the hour did to record, and most of the labels would teach the model the same thing.

An hour of footage is one scene sampled a hundred thousand times

Adjacent frames in video are near-duplicates. The bottle that is under the fill head at one frame is a few millimetres further along at the next, the lighting has not changed, and the cap is still on or still missing. A model trained on all hundred thousand frames learns line 3 very well and learns nothing it would not have learned from a few hundred of them.

The footage lesson puts it in terms of conditions: what a model needs is variety, and an hour of one camera on one line under one lighting condition is one condition, sampled many times. The sampling is the first decision, and it is the one that sets the cost of everything after it.

Sample a frame every couple of seconds and start there

On line 3 the practical rate is a frame every two seconds, which is also the rate at which the live model will look at the same camera later. That turns the hour into about eighteen hundred frames, a set a person can look through in an afternoon, and it keeps the frames that differ from each other while dropping the ones that do not.

A fixed interval misses the short events. A missing cap that is in view for a second and a half can fall between two samples. So the interval is the base. On top of it the quality lead pulls extra frames around every event she knows about: the line stops, the changeover to the smaller bottle at 2 pm, the ten minutes after the fill head was cleaned when the lens had a smear on it. Those frames are rare in the hour and common in the model's future.

She keeps the timestamps of every sampled frame in the filename. A year later, when someone asks which shift a strange frame came from, that is the only record.

Which frames to label and which to leave out

Of the eighteen hundred frames from line 3, some are worth a label and some are not. The ones worth labeling are the ones with a cap missing, the ones where a bottle is partly hidden by the guard rail, and the ones at the changeover where the smaller bottle looks different under the head. A good number of ordinary frames with every cap on go in too, because a model has to learn what correct looks like as much as what a defect looks like.

The ones to leave out are frames that are the same as a frame already in the set. The bottle at the same position, the same lighting, the same result. Labeling those costs time and adds nothing, and a set that is mostly duplicates of the easy case trains a model that is confident on the easy case and unsure everywhere else.

Empty frames, when the line is stopped and the belt is bare, stay in as frames with no bottle. They are the cheapest labels in the set and the model needs them, because a stopped line is a state the camera will see every day.

Tracking carries the bounding box across frames instead of drawing it fresh

A bottle that is in frame for twelve samples does not need twelve independent boxes. The quality lead draws a bounding box on the first frame, Lexi carries it forward as the bottle moves, and she corrects the frames where the carried box has slipped, which on a line at constant speed is very few. The bottle keeps its identity across the frames, so the set records that these twelve boxes are one bottle rather than twelve.

That identity matters later. A cap detector that only sees single frames flags a missing cap once per frame, and a bottle with no cap that passes through twelve samples becomes twelve alerts. A model trained on tracked labels can count bottles rather than frames, and the alert says one bottle, which is what the line lead wants to hear.

Per-frame labels still have their place. The changeover frames, where the smaller bottle first appears and nothing about it has been seen before, are labeled fresh, because a track carried from the old bottle would be wrong from its first frame.

The check is on the track as well as the frame

Every label in the set is checked by a second person before anything trains on it, and for video the check has a second part. A frame from line 3 can be correct on its own while the track it belongs to is wrong. The usual case is a box that jumped from one bottle to its neighbour at the point where they were closest, and carried the wrong identity for the rest of the run. The reviewer scrubs the track, and a jump shows up as a box that moves faster than a bottle can.

That is how the millions of annotations behind our labeling work were checked, and it is the part of the job that is easy to skip on a deadline. My own view, from running review queues, is that the track check finds more training problems per hour of reviewer time than the frame check does, and it is the one teams drop first.

A small aside from the same queues: the frames reviewers argue about most are never the defects. They are the bottles at the very edge of the frame, half in and half out, and the argument is settled by writing the rule down once.

The labeled set goes back to the line it came from

Once the set trains a version, the version goes onto the same camera. The labeling doc describes the verification pass that gets the first set right; the live line is where the set keeps growing. The model watches line 3 and samples a frame every couple of seconds. The frames it is unsure of, a new bottle shape, a lighting change after a lamp is replaced, come back to the quality lead as a short queue rather than an hour of footage.

Her corrections on those frames are the next training set. When they cross the project's threshold a new version trains and replaces the old one with no downtime, and the hour of footage she started with is the smallest part of what the model has seen by the end of the year.

See it on your own footage.

Start with your footage

More in Labeling

Labeling · 6 min read

AI labeled vs human labeled data, how much a model can do before a person has to look

On a shelf dataset the model's proposals are mostly right. On safety cones in a yard they are mostly wrong. Measure the gap, then keep a person on every label.

Finn Ellingwood · Sep 28, 2026

Labeling · 6 min read

Bounding boxes in computer vision, what a box teaches and what it reports

The same rectangle on a forklift is the lesson at training time and the answer at run time. Corner and centre formats, the shifted-box bug, floor in the box.

Esdras Ntuyenabo · Sep 28, 2026

Labeling · 6 min read

Which words find the forklift, measured instead of guessed

Five ways to say forklift, five sets of boxes on aisle 6. Score each phrase on a small labeled set before Lexi labels the whole archive with the winner.

Sheikh Srijon · Sep 28, 2026

LexData
LexData

Product

  • Platform
  • Industries
  • Use cases

Resources

  • Docs
  • The Field Guide
  • Blog
  • Why models drift

Industries

  • Energy & utilities
  • Oil & gas
  • Agriculture
  • Manufacturing
  • Insurance
  • Retail
  • Robotics

Company

  • About
  • Customers
  • Careers
  • Contact

Trust

  • Security
  • Privacy
  • Terms

Stay updated

What we learn running vision models in production.

See everything.
Miss nothing.

Stay updated

What we learn running vision models in production.

Terms of use & Privacy policy

© 2026 LexData Labs · All rights reserved