Skip to content
LexDataLexData
PlatformIndustriesCustomers
DocsThe Field GuideBlogWhy models drift
AboutCareersSecurityContact
Log inStart now
← All posts

Labeling · 7 min read

Uploading a week of images, videos and annotations without losing the rare frames

A week of the scanner camera is mostly the same carton. Import the originals, sample the routine, keep the jam at 3 am whole, bring old labels across intact.

Summary

This post takes a week of recorder footage from a packaging line and gets it into a dataset: the files imported from S3, Google Drive or upload with nothing re-encoded, frames sampled so near-identical ones do not inflate the labeling, the incident windows kept whole, and existing labels brought across as COCO JSON, YOLO TXT or CVAT XML with every box on the frame it was drawn on. It is for the person who has the footage and the old labels and needs a dataset by Monday.

Finn Ellingwood · Engineer · Oct 2, 2026

Packaging line, cartons past the scanner, generated scene with detections from our model

The recorder on packaging line 3 holds a week of the camera over the scanner. Cartons passing at a steady rate, a label on each, a jam at 3 am on Wednesday that stopped the line for twenty minutes, and a changeover to a smaller carton on Friday. The quality engineer has that week as a folder of video files, one per hour, plus a folder of labels from a project two years ago in a format the old tool exported.

He wants a dataset out of it by Monday. Most of the week is the same carton passing the same scanner, and the parts that matter are a few minutes long.

A week of the scanner camera is mostly the same carton

A fixed camera over a steady line produces the least varied footage there is. Hour after hour of cartons in the same place on line 3, under the same light, at the same speed. The lesson on how much footage you actually need is blunt about what that is worth: one condition sampled many times, which teaches the model almost nothing after the first few hundred frames.

What makes the week worth importing is what is not routine. The jam, when cartons piled up and one came through crushed. The changeover, when the smaller carton appeared and the model that was trained on the larger one had never seen it. The night shift, when the sodium lights are on and the bay door is closed. Those are the conditions, and they occupy a small fraction of the files.

The engineer knows when the jam happened because the line's downtime log says so. The log is the map of where the rare frames are.

Import the recorder's files as they are, nothing re-encoded

The files come in as they were recorded. Footage is imported from S3 or Google Drive, or uploaded, and nothing is re-encoded on the way, which matters on a line camera more than it sounds. The crushed corner from the jam is a few pixels of shadow on cardboard, and an export that re-compresses the hour to make it smaller smooths those pixels away before anyone has labeled them.

So the hour files go in whole, at the resolution the recorder kept, and the quickstart has the same instruction for a single file: upload the original, and let the platform split it into frames. Seven days of hour files is a long upload and a short decision.

One project for line 3, named after the operation rather than the week, because there will be more weeks.

Sample frames every few seconds, and keep the incident windows whole

A video file split into every frame is a labeling bill nobody wants. Adjacent frames of a carton on a belt are near-duplicates, and labeling all of them is the same box drawn thirty times. The default is to sample a frame every couple of seconds, which on the routine hours keeps every carton at least once and throws away the copies.

The incident windows are the exception, and they are the whole reason for the post. The twenty minutes around the jam should be sampled densely or taken whole, because the crushed carton is in a handful of frames and a sample every two seconds can step over it. The same goes for the first hour after the changeover on Friday, when the smaller carton is new. The downtime log gives the timestamps, and the sampling is set per window rather than per week.

My own view is that uniform sampling over a week of line footage is the most common way a rare event gets lost before labeling, and that the downtime log should be open in the other window every time someone sets a sampling rate.

Existing labels come across as COCO JSON, YOLO TXT or CVAT XML

The folder of labels from two years ago is worth bringing. The old project boxed cartons and labels on the same camera, and a few thousand checked boxes are a head start on the reference set even if the class list has changed since. Labels are imported as COCO JSON, as YOLO TXT or as CVAT XML, and each format carries its boxes differently. COCO uses absolute pixels with the top-left corner and a size. YOLO is normalised to the frame with the centre and a size. The CVAT XML format has corners and the frame's own dimensions alongside.

The import reads whichever it is given and keeps the boxes as they were drawn. The class names come across with them, and the engineer maps the old "box" class to the new "carton" once, in the class list, rather than editing a file.

What does not come across on its own is the judgment behind the old labels. The old project had no rule for a carton half out of frame, and the old boxes handle it three different ways. Those go through the checking pass like any proposal.

Every bounding box has to land on the frame it was drawn on

The quiet failure in a label import is a bounding box that arrives on the wrong frame or at the wrong scale. The old labels from line 3 were drawn on frames extracted at one resolution from one file; the new import extracts frames from the same files at the recorder's full resolution. If the old frames were downscaled, every old box is now too small by the same factor, sitting in the top-left quarter of the carton it used to hug.

The import matches each label file to its frame by name and checks the frame dimensions the label file claims against the frame it lands on, and a mismatch is reported rather than silently scaled. The engineer's old labels claimed a smaller frame than the recorder's, and the fix was to scale the boxes once, on import, by the ratio. A label that lands on the wrong frame entirely, because the old extraction started a second later than the new one, is a worse failure and the reason a sample of imported boxes gets opened and looked at before any of them train.

He opened forty. Thirty-eight sat on their cartons. Two were a frame late, from an hour file whose old extraction had skipped the first second.

The rare frames are the reason the week was worth importing

With the routine hours sampled, the incident windows kept whole, and the old labels scaled onto the right frames, the dataset for line 3 is a few thousand frames rather than a few million, and the jam and the changeover are in it. You type the classes, Lexi proposes a box on every carton and label in every frame, and a person checks each one, with the labeling guide covering what that pass does with the frames from the jam.

LexData takes the line 3 model through its whole life. You type what to look for, Lexi puts a box on every frame, and a person checks each label before anything trains on it. The model then watches the scanner camera, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime. The next jam will be in the review queue within minutes of happening, which is how the week after this one gets imported: as corrections rather than as hour files.

The dataset was ready on Sunday. The engineer spent Monday morning looking at the crushed carton from Wednesday, boxed, for the first time.

See it on your own footage.

Start with your footage

More in Labeling

Labeling · 7 min read

AI data labeling workflows, three ways to label footage and when each one pays

A pipeline right-of-way survey labeled three ways: every box by hand, Lexi proposing and a person checking, and synthetic frames for the leak nobody has filmed.

Rajiya Sultana · Oct 2, 2026

Labeling · 6 min read

Annotation analytics, the numbers a labeling queue produces besides labels

Three labelers on a month of warehouse footage produce boxes, and also a throughput, an acceptance rate and a map of where the rejections cluster.

Rajiya Sultana · Oct 2, 2026

Labeling · 7 min read

Annotation format conversion between COCO, YOLO and CVAT without losing a box

Three years of line inspection labels from two tools arrive in three formats. The boxes that shift are the ones nobody draws on a frame before training.

Esdras Ntuyenabo · Oct 2, 2026

LexData
LexData

Product

  • Platform
  • Industries
  • Use cases

Resources

  • Docs
  • The Field Guide
  • Blog
  • Why models drift

Industries

  • Energy & utilities
  • Oil & gas
  • Agriculture
  • Manufacturing
  • Insurance
  • Retail
  • Robotics

Company

  • About
  • Customers
  • Careers
  • Contact

Trust

  • Security
  • Privacy
  • Terms

Stay updated

What we learn running vision models in production.

See everything.
Miss nothing.

Stay updated

What we learn running vision models in production.

Terms of use & Privacy policy

© 2026 LexData Labs · All rights reserved