Labeling · 6 min read
Which words find the forklift, measured instead of guessed
Five ways to say forklift, five sets of boxes on aisle 6. Score each phrase on a small labeled set before Lexi labels the whole archive with the winner.
Summary
This post takes a warehouse camera and five phrases that all mean forklift, and shows how a text-prompted detector draws different boxes for each. It scores every phrase against a small set of frames a person labeled, picks the winner by F1, and only then lets the winner label the archive, with a person checking each box. It is for teams starting a labeling project with a prompt and a large archive.
Sheikh Srijon · GTM Lead · Sep 28, 2026

Warehouse aisle from a high camera, a forklift and a racked pallet boxed, generated scene with detections from our model
The high camera over aisle 6 has recorded a year of forklifts, and the team wants boxes on all of them without drawing a year of boxes by hand. A text-prompted detector will draw them from a phrase. The first phrase anyone types is "forklift", and it does well on the counterbalance trucks and misses the narrow-aisle reach truck entirely. "Lift truck" finds the reach truck and starts boxing the pallet jack. "Industrial vehicle" boxes the floor scrubber. Five phrases, five different archives.
Whichever phrase is chosen will label a year of footage, and every box it draws wrong is a box a person has to fix or a mistake the model learns. So the phrase is worth choosing carefully, and choosing carefully means measuring rather than eyeballing three frames and picking the one that looked right.
The prompt is a class definition written in someone else's words
A text-prompted detector was trained on captions from the internet, and its idea of "forklift" is the internet's. That is a good idea for the yellow counterbalance truck in every stock photo and a vague one for a reach truck with its mast folded, which the internet mostly calls something else or nothing. The phrase is a class definition, and the class it defines is whatever the captions meant by those words.
That is why the site's own word matters less than one would hope. The warehouse calls every forklift a "truck" on the radio, which is a fine word for the radio and a hopeless prompt, because the detector's "truck" is the thing at the loading dock. The prompt has to be the phrase the detector understands, and the only way to know which one that is on aisle 6 is to test it there.
A small labeled set is the judge for every phrase
Take a few dozen frames from the aisle 6 camera, chosen to span the week: the Monday inbound rush, a quiet Wednesday afternoon, the night shift under half the lights, a frame with the reach truck and a frame with the scrubber. A person draws the forklift boxes on those frames by hand, with the site's own definition of what counts, which for this rule includes the reach truck and excludes the pallet jack. Those boxes are the truth every phrase is measured against.
The set is small on purpose. Its job is to rank phrases, and a few dozen frames rank five phrases reliably enough if they are the right few dozen. The frames that matter most are the ones where the phrases disagree, the reach truck and the scrubber, so the set should carry several of each.
The labeling doc describes how a prompt becomes a class list on the platform; this small set is what decides which prompt goes into it.
Grounding DINO gets scored per phrase on the same frames
Run each of the five phrases through Grounding DINO on the same frames at the same threshold, match its boxes against the hand-drawn ones at the usual overlap line, and count. Precision for the phrase is the share of its boxes that landed on a real forklift; recall is the share of the real forklifts it boxed. F1 folds the two into one figure per phrase, which is fine here because the goal is a ranking rather than a diagnosis.
On aisle 6 the ranking is rarely what the team guessed. "Forklift" has the best precision and poor recall because of the reach truck. "Lift truck" has the best recall and pays for it on the pallet jack. A compound phrase, "forklift or reach truck", often wins the F1 outright, and a phrase nobody would have typed, "warehouse forklift", sometimes beats them all because the word "warehouse" pulls the detector's attention to the right kind of scene.
Read the ranking per frame type as well as overall. A phrase that wins on the day shift and collapses at night has told you something about the archive, and the night frames may need their own phrase.
The winner labels the archive and a person still checks every box
The phrase with the best F1 becomes the prompt, and Lexi runs it across the year of aisle 6 footage. What comes back is a first pass, and the small set already says how good a first pass it is: if the winner scored nine tenths on recall, one forklift in ten is unboxed in the archive, and the verification pass has to find those. A person checks each label before anything trains on it, and the misses on the reach truck are where they look first.
That checking is what makes the phrase's mistakes harmless. A prompt that is right nine times in ten produces a training set that is right nine times in ten, unless a person fixes the tenth, and a model trained on the fixed set has never seen the prompt's opinion of a floor scrubber.
In my experience the plain word beats the technical one more often than engineers expect, and the compound phrase beats both. Nobody has ever guessed the winner for a site from their desk, which is the whole argument for the small set.
The phrase that won on aisle 6 may lose at the dock
A prompt is scored on one camera, and the ranking belongs to that camera. The dock camera sees forklifts from the side and at a distance, with trucks behind them, and "forklift" and "truck" become harder to keep apart. The night camera sees headlights and reflections. Each camera gets its own small labeled set and its own ranking, and it is normal for a different phrase to win.
The scrubber, as it happens, has been boxed as a forklift by every phrase the team tried, and the fix was a frame of it in the labeled set with no box, so that the phrase which left it alone got the credit.
After the archive is labeled and the model is trained, the prompt's job is over. LexData takes the forklift model through its whole life from there. You type what to look for, Lexi puts a box on every frame, and a person checks each label before anything trains on it. The model then watches the cameras the warehouse already has, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime.
The frames it doubts are the reach truck at night and the scrubber at the end of the aisle, the same frames the phrases disagreed on. The lesson on how much footage you actually need explains why those are the ones worth a person's minute.
See it on your own footage.
Start with your footageMore in Labeling

Labeling · 6 min read
AI labeled vs human labeled data, how much a model can do before a person has to look
On a shelf dataset the model's proposals are mostly right. On safety cones in a yard they are mostly wrong. Measure the gap, then keep a person on every label.
Finn Ellingwood · Sep 28, 2026

Labeling · 6 min read
Bounding boxes in computer vision, what a box teaches and what it reports
The same rectangle on a forklift is the lesson at training time and the answer at run time. Corner and centre formats, the shifted-box bug, floor in the box.
Esdras Ntuyenabo · Sep 28, 2026

Labeling · 6 min read
How much training data a computer vision model needs from your own cameras
A weld defect detector got useful on a few hundred chosen frames. A sixty-class shelf model needed far more. The difference is the task and the variation.
Sheikh Srijon · Sep 28, 2026