Skip to content
LexDataLexData
PlatformIndustriesCustomers
DocsThe Field GuideBlogWhy models drift
AboutCareersSecurityContact
Log inStart now
← All posts

Labeling · 6 min read

Passive image collection at the edge, letting the camera build its own training set

Leave the dock camera sampling over a weekend, drop near-duplicates at the box, pull the rare scenes with a typed description, and send the rest to review.

Summary

This post describes passive collection from a camera over a loading dock, sampled on a fixed interval by a box on site rather than watched by a person. It concludes that dropping near-duplicates before storage and pulling the rare scenes with a typed description turns a weekend of footage into a first batch worth reviewing. It is for teams starting a model with no labeled frames and a camera already in place.

Andreas Ohrvall · CTO · Oct 3, 2026

Loading dock from a mounted camera, trucks at the bays and a forklift with a pallet boxed, generated scene with detections from our model

The camera over dock door 3 has been there for four years, wired to a recorder that overwrites itself every thirty days. Nobody has ever pulled a frame from it for training, because pulling frames means someone sitting at the recorder scrubbing through a weekend of footage, and there is always something more urgent. The dataset for the dock model has been "next quarter" for three quarters.

The cheaper way is to let the camera collect for itself. A small box beside the recorder samples a frame at a set interval, keeps the ones that differ from the last, and by Monday morning there is a folder of a few thousand frames that a person can filter in an hour. That folder is the dataset the project has been waiting for, and the camera built it on its own.

A camera left alone for a weekend collects its own training set

Passive collection means the sampling happens without anyone watching. The box on the dock reads the RTSP feed, takes a frame every few seconds, and writes it to local storage. Nothing leaves the site during the weekend. On Monday the folder is reviewed and the useful part goes onward.

The weekend matters as a choice. Friday afternoon is busy, Saturday has the half shift with different drivers, Sunday is empty except for the cleaning crew and the one late delivery, and Monday at 6 am is the queue. Those are four situations the dock model will meet, and a person scrubbing the recorder would have grabbed a hundred frames of Friday and stopped.

The dock lights are on a timer that turns them off at 22:00 whether or not a truck is at the door, so the late Sunday delivery is unloaded in the light from the truck's own cab. That frame was in the weekend folder. It would not have been in anyone's plan.

Sampling on a fixed interval keeps the collection honest

There are two ways to decide when to sample: on a clock, or when something moves. Motion-triggered collection sounds efficient and produces a folder full of forklifts and no empty docks, and the model then learns that a dock always has a forklift in it. A fixed interval collects the empty dock at 3 am as faithfully as the queue at 6 am, and the empty frames are how the model learns what the absence of a forklift looks like.

My view is that the interval should be fixed and the filtering done afterwards, every time, even though it costs storage. Every decision made at capture time is a decision the dataset can never undo. A frame that was never taken cannot be labeled later, and a folder that is twice as big as it needs to be can be trimmed on Monday.

The interval itself depends on how fast the scene changes. A dock where a truck takes forty minutes to unload can be sampled every ten seconds and miss nothing. A conveyor needs a shorter one. The how much footage lesson gives the reasoning: what the model needs is coverage of situations, and the interval only has to be short enough that no situation slips between two samples.

Near-duplicates are dropped at the box before anything is stored

A frame every ten seconds over a weekend is around seventeen thousand frames, and most of them are the empty dock at night. The box compares each new frame to the last one it kept and drops it if nothing changed, so the folder holds one frame of the empty dock per lighting condition rather than five thousand. This is the same near-duplicate rule a dataset would apply later, applied at the source, and it is what makes local storage on a small box enough.

The threshold is set loose on purpose. Dropping a frame that was almost the same as the last one costs little; dropping a frame where a pallet had shifted an inch and fallen costs the one event the dock model exists for. So the box keeps anything it is unsure about, and the Monday review does the finer sorting.

With a runner beside the recorder this is the shape the deployment doc describes for the live model too: the footage stays on site, and only the frames that matter leave.

A typed description pulls the rare scenes out of the pile

Monday's folder has a few thousand frames after the box has thinned it, and the person reviewing does not want to page through all of them. They type a description of what they are looking for: a pallet on the floor beside the truck, a person on the dock leveller, a truck backed in at an angle. Lexi finds the frames that match the description and shows them first. The rare scenes, which are the ones a model trained on a busy Friday will get wrong, come to the top of the pile in a minute.

The description search is a way to look, and it is only a way to look. A frame it surfaces is a candidate, and a frame it misses is still in the folder. The reviewer scrolls the rest at speed to catch what the description did not describe, which on the dock included a delivery driver's dog sitting on the leveller for most of Saturday morning.

The filtered set goes to review as the first labeled batch

What leaves the site is the filtered set: a few hundred frames chosen for variety, with the empty dock, the queue, the late delivery in cab light and the dog. This is the first batch. You type what to look for, Lexi puts a box on every frame, and a person checks each label before anything trains on it. The labeling with Lexi doc covers that pass; the point here is that the batch it receives came from the dock itself, at the dock's own hours and weather, rather than from a stock set of loading bays.

The first model trained on this batch will be weak in the situations the weekend did not contain. Snow, a new truck livery, a second forklift. That is expected. The first batch is for getting a model onto the camera; the collection that follows is for the rest.

Collection keeps running after the model is live

Once the model is watching the dock, the box has a second job. The model returns the frames it is unsure of, and those go to a person, whose corrections retrain it, and the new version replaces the old one with no downtime. Passive collection continues underneath, on the same interval, because the model's doubt only covers things it half recognises. A situation it has never seen and confidently misreads produces no doubt, and the fixed-interval sample is where that situation is caught, when the reviewer types a new description on a Monday and finds the frames.

The first winter, the snow on the dock apron was in the passive folder weeks before the model started doubting it.

See it on your own footage.

Start with your footage

More in Labeling

Labeling · 7 min read

Aerial dataset augmentation for drone frames where there is no up

A tower seen straight down has no top or bottom, so rotations and both flips are safe. Scale for altitude, brightness for sun, move every box with its pixels.

Esdras Ntuyenabo · Oct 3, 2026

Labeling · 6 min read

A collaborative data annotation workflow run as a pipeline

Batch the bottling line's frames by camera and shift, assign so nothing is boxed twice, attach the guideline, review every label, train on the approved set.

Rajiya Sultana · Oct 3, 2026

Labeling · 7 min read

Dataset health check for computer vision, what to look at before anything trains

A scratch dataset where every scratch sits in the centre of the frame will train a model that looks in the centre. Five counts to read before the first epoch.

Rajiya Sultana · Oct 3, 2026

LexData
LexData

Product

  • Platform
  • Industries
  • Use cases

Resources

  • Docs
  • The Field Guide
  • Blog
  • Why models drift

Industries

  • Energy & utilities
  • Oil & gas
  • Agriculture
  • Manufacturing
  • Insurance
  • Retail
  • Robotics

Company

  • About
  • Customers
  • Careers
  • Contact

Trust

  • Security
  • Privacy
  • Terms

Stay updated

What we learn running vision models in production.

See everything.
Miss nothing.

Stay updated

What we learn running vision models in production.

Terms of use & Privacy policy

© 2026 LexData Labs · All rights reserved