Skip to content
LexDataLexData
PlatformIndustriesCustomers
DocsThe Field GuideBlogWhy models drift
AboutCareersSecurityContact
Log inStart now
← All posts

Labeling · 6 min read

YOLO knowledge distillation, let the big model label and the small model run

A large model boxes the warehouse frames once, a floor lead corrects them, and a compact detector trained on that set runs beside the recorder.

Summary

This post explains distillation as it is done on a warehouse camera: a large model proposes boxes on the frames, a floor lead corrects them, and a compact YOLO detector trained on the corrected set is what runs on the runner beside the recorder. It covers why the large model never runs in production, how the small model's size is set by the hardware, and how the live feed keeps the set growing. It is for teams who have a good large model and no way to run it at the camera.

Andreas Ohrvall · CTO · Sep 28, 2026

Edge runner beside the recorder in a control cabinet, generated scene with detections from our model

The warehouse has a camera over each of its twelve dock doors, door 1 to door 12, and a small computer in the cabinet beside the recorder, and the operations manager wants a count of pallets crossing each door per hour. There is a large open-vocabulary model that boxes pallets on a still frame very well, and it runs on a server with a GPU the size of a suitcase. Nothing like that is going into the cabinet.

Distillation is the way out of that corner. The large model does the labeling, once, offline. The small model does the watching, every couple of seconds, on the box in the cabinet. The two never meet on the live feed.

The large model sees the warehouse frames once and never again

The first step is to run the large model across a sample of frames from the dock cameras: a frame every couple of seconds from each door, across a week. That way the sample covers the night shift, the Monday rush and the afternoon when the sun comes through door 7. The prompt is "pallet", and the model returns a box on every pallet it can find in every frame.

Those boxes are the large model's knowledge, written down as labels. That is the whole of what distillation means here. The small model will never see the large model's weights or its internal scores. It will see the boxes, after a person has been through them.

The large model has done its job by the end of that run. It is not deployed anywhere, it is not called from the cabinet, and when the next sample is needed in six months it runs again, once.

The proposed boxes are drafts until the floor lead has been through them

A large model's boxes on a warehouse are good and they are not right. It boxes the pallet on the forklift and the pallet stack by the wall as one object. It misses the pallet under the shrink-wrap because the words "pallet" and "wrapped load" are far apart in its head. It boxes a wooden crate as a pallet because they share a texture. On door 7, in the afternoon, it boxes the sun.

The floor lead opens the frames in LexAnnotate and corrects them. Split the stack. Box the wrapped pallet. Delete the crate. Delete the sun. Every label that will train the small model is checked before it trains anything, which is the same verification pass any labeled set goes through, and it is the reason the small model can be better than the large one on this warehouse. The large model has seen pallets in general. The corrected set has seen these doors.

An aside from the same floor: the lead's most frequent correction was on the empty pallets stacked by the wall, which the large model boxed as a pallet on some frames and as furniture on others. The rule she wrote, stack counts as one, went into the guide before the second sample.

A compact YOLO detector trains on the corrected set

The corrected set trains a small single-pass detector, the kind that looks at the whole frame once and returns boxes and classes together. A model of that shape is built to run fast on modest hardware, and a compact YOLO variant trained on a few thousand corrected frames from twelve doors is a fraction of the large model's size and answers in a fraction of the time.

It also answers the same way every time. The large model's boxes moved with the prompt wording and with the light. The small model has one class, one set of weights and a fixed answer for a given frame, which is what an hourly count per door needs.

The evaluation is on frames from the dock cameras that the small model has never seen, held out from the corrected set. The number worth reporting is how many pallets it found and how many of its boxes were pallets, per door, because door 7 in the afternoon is going to be the weak one and an average across twelve doors hides it.

The size of the small model is set by the box in the cabinet

There is a family of these detectors from tiny to large, and the choice is made by the hardware beside the recorder rather than by a leaderboard. Twelve cameras, one frame every couple of seconds each, and a box with no GPU: that is a budget, and the model that fits it is the one to train. A bigger model that runs at half the rate the cabinet needs is worse on this warehouse than a smaller one that keeps up.

The deployment doc makes the same argument in the other direction. Plan for the enclosure. The cabinet by the recorder is sealed, it gets warm in July, and a model that fits at the bench may throttle in there. The model is exported for the device, described as the device rather than the architecture, and tested in the cabinet before anyone counts a pallet.

My own view is that most teams pick the model size a year too early, before they know the frame rate the question actually needs. An hourly pallet count does not need every frame. Deciding that first would have let several projects I have seen run on the box they already had.

Live in days, and the live feed keeps the set growing

The whole path, from the large model's first run to the small model on the cabinet, is a labeling job and a training job, and it goes live in days rather than the months a hand-labeled set would have taken.

The small model then watches the twelve doors from the cabinet, and the frames it is unsure of come back to the floor lead in LexAnnotate. In the first month those are the new supplier's blue plastic pallets, which neither the large model nor the corrected set had seen. She corrects them, the corrections retrain the small model, and the new version replaces the old one on the cabinet with no downtime.

The large model is not part of that loop. When the warehouse adds a class, it runs once more across a fresh sample and the floor lead corrects the drafts again, but the model that learns from her corrections, and the only one that ever runs beside the recorder, is the small one.

See it on your own footage.

Start with your footage

More in Labeling

Labeling · 6 min read

AI labeled vs human labeled data, how much a model can do before a person has to look

On a shelf dataset the model's proposals are mostly right. On safety cones in a yard they are mostly wrong. Measure the gap, then keep a person on every label.

Finn Ellingwood · Sep 28, 2026

Labeling · 6 min read

Bounding boxes in computer vision, what a box teaches and what it reports

The same rectangle on a forklift is the lesson at training time and the answer at run time. Corner and centre formats, the shifted-box bug, floor in the box.

Esdras Ntuyenabo · Sep 28, 2026

Labeling · 6 min read

Which words find the forklift, measured instead of guessed

Five ways to say forklift, five sets of boxes on aisle 6. Score each phrase on a small labeled set before Lexi labels the whole archive with the winner.

Sheikh Srijon · Sep 28, 2026

LexData
LexData

Product

  • Platform
  • Industries
  • Use cases

Resources

  • Docs
  • The Field Guide
  • Blog
  • Why models drift

Industries

  • Energy & utilities
  • Oil & gas
  • Agriculture
  • Manufacturing
  • Insurance
  • Retail
  • Robotics

Company

  • About
  • Customers
  • Careers
  • Contact

Trust

  • Security
  • Privacy
  • Terms

Stay updated

What we learn running vision models in production.

See everything.
Miss nothing.

Stay updated

What we learn running vision models in production.

Terms of use & Privacy policy

© 2026 LexData Labs · All rights reserved