Skip to content
LexDataLexData
PlatformIndustriesCustomers
DocsThe Field GuideBlogWhy models drift
AboutCareersSecurityContact
Log inStart now
← All posts

Labeling · 6 min read

Checking a public dataset's provenance before it trains your model

Who made it, what the licence allows, what a sample of its boxes shows, and how far its frames sit from your gate camera. Four checks before a download trains.

Summary

This post walks through the checks a team should make on a public dataset before training on it: who assembled it and why, whether the licence covers a commercial deployment, what a hand review of a random sample of its boxes reveals about its labeling convention, and how far its frames sit from the team's own cameras. It concludes that a public dataset is a warm start whose replacement date is set by that gap, and that it must never sit in the test split. It is for teams who found a promising dataset online and want to know what they are about to train on.

Sheikh Srijon · GTM Lead · Oct 3, 2026

Dashcam frame with traffic signs, pedestrians and vehicles boxed, from a customer perception run

The depot has a camera on its vehicle gate and wants a model that counts trucks in and people on foot. Nobody has labeled a frame yet. An engineer finds a public dataset of street footage with vehicles and pedestrians boxed, forty thousand frames, downloaded in an afternoon, and the first model trained on it looks reasonable on the gate feed by Friday.

Whether it stays reasonable depends on four things the download page did not say, and the time to find them out is before the model is on the gate rather than after.

Who made the dataset and what they wanted it for

A dataset is built to answer somebody's question, and its boxes answer that question rather than yours. A street dataset collected from a dashcam in Berlin for a driving project boxes vehicles because they are obstacles, and boxes them at the distances and angles a car sees them from: ahead, at road level, moving. The depot gate sees trucks from a pole, from above, stationary at the barrier. Same class name, different object.

So the first check is the readme, the paper if there is one, and the author. A university group building a benchmark, a company releasing a subset of its own product footage, and an anonymous upload with no description are three different provenances, and the amount of trust each earns before a single box is looked at is different. The how much footage lesson makes the point that variety is what the model learns from, and a dataset's origin tells you which variety it holds and which it never saw.

The readme on the depot's download lists a class spelled "bycicle", and forty thousand frames later the model has learned the class faithfully under that name. It is a small thing, and it is also the surest sign that nobody with a second pair of eyes ever read the class list.

The licence has to cover a commercial deployment on your cameras

Public datasets carry licences, and the licences differ in a way that matters at the gate. Some allow research use only. Some allow any use with attribution. Some allow use of the frames but say nothing about a model trained on them, which is the case that most often ends in a lawyer's email eighteen months later. A few are collections of frames scraped from elsewhere, where the person who assembled the dataset had no right to license it at all.

The check is dull and takes twenty minutes on a Monday morning: read the licence, confirm it names commercial use, confirm it covers derived models and not only the frames, and write down where it was read and when. A dataset with no licence file is a dataset with no licence, and a model trained on it is a model whose right to run on the depot gate cannot be shown.

A random sample of its boxes reviewed by hand tells you the convention

The download page says the dataset is labeled. It does not say how. Pull two hundred frames at random, open them in the labeling tool, and look at the boxes as a reviewer would. Within an hour the convention is clear. Does a truck's box include its mirrors? Is a pedestrian half behind a parked car boxed or skipped? Is a group of three people three boxes or one, and is a vehicle at the frame edge cut or excluded?

None of these are wrong. All of them are decisions the depot will inherit. If the gate model is meant to count people and the public dataset boxes groups as one, the count will be low from the first day. The reason will be invisible in the accuracy figure, because the model is faithfully reproducing the labels it learned.

My own view is that the sample review is the check most teams skip and the one that would have saved them the most, because a convention mismatch is the one fault a bigger download cannot fix. The labeling with Lexi doc describes the class definition as the first thing written; on a public dataset the class definition was written by someone else, and the sample is how you read it.

The gap to your own cameras decides how soon it is replaced

The last check is the one that sets the schedule. Lay a hundred frames from the public dataset beside a hundred from the gate camera on a Tuesday and list what differs. The viewpoint and the height. The lens. The time of day and the weather. The vehicle types and the uniforms on the people. The barrier and the signage that appear in every gate frame and in none of the street frames. The longer the list, the sooner the public dataset stops being useful.

This is the same gap the drift catalog describes as a new site: a model that learned one place is asked about another, and its accuracy there is unknown until measured. A public dataset is a site the model has never been to, standing in for the site it will watch.

A public dataset is a warm start and never the test set

With the four checks done, the honest use of the download is as version 1. The model trained on it gets onto the gate camera, produces its first detections, and starts returning the frames it is unsure of. Those frames come from the gate, at the gate's height and light, with the gate's barrier in them, and a person's corrections to them are the first labels the project owns.

What the public dataset must never do is stand in as the test set.

Scoring the gate model on street frames says how well it learned the street. The score that matters comes from held-out gate frames, labeled by the depot's own reviewer under the depot's own convention, and until those exist the project has a model and no number.

The depot's own frames replace it one review batch at a time

From there the loop does the replacing. Frames the model is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime. Each batch of corrected gate frames shifts the dataset toward the gate, and by the following March the public frames are a minority of what the model learned from.

At that point the question of whether to keep them is worth asking on purpose. Frames that still help, because they show a vehicle type the gate has not yet seen, stay. Frames that are pulling the model toward a road-level view it will never have can go, and the score on the held-out gate frames is what says which. The "bycicle" class went in the second retrain, once the depot had its own.

See it on your own footage.

Start with your footage

More in Labeling

Labeling · 7 min read

Aerial dataset augmentation for drone frames where there is no up

A tower seen straight down has no top or bottom, so rotations and both flips are safe. Scale for altitude, brightness for sun, move every box with its pixels.

Esdras Ntuyenabo · Oct 3, 2026

Labeling · 6 min read

A collaborative data annotation workflow run as a pipeline

Batch the bottling line's frames by camera and shift, assign so nothing is boxed twice, attach the guideline, review every label, train on the approved set.

Rajiya Sultana · Oct 3, 2026

Labeling · 7 min read

Dataset health check for computer vision, what to look at before anything trains

A scratch dataset where every scratch sits in the centre of the frame will train a model that looks in the centre. Five counts to read before the first epoch.

Rajiya Sultana · Oct 3, 2026

LexData
LexData

Product

  • Platform
  • Industries
  • Use cases

Resources

  • Docs
  • The Field Guide
  • Blog
  • Why models drift

Industries

  • Energy & utilities
  • Oil & gas
  • Agriculture
  • Manufacturing
  • Insurance
  • Retail
  • Robotics

Company

  • About
  • Customers
  • Careers
  • Contact

Trust

  • Security
  • Privacy
  • Terms

Stay updated

What we learn running vision models in production.

See everything.
Miss nothing.

Stay updated

What we learn running vision models in production.

Terms of use & Privacy policy

© 2026 LexData Labs · All rights reserved