Operations · 7 min read
A model trained on your warehouse beats one trained on the internet
A general detector calls the forklift a truck and the pallet nothing at all. The held-out week from your own cameras is the only score that counts.
Summary
This post compares a general vision service with a detector trained on a few hundred labeled frames from a warehouse's own cameras. It concludes that the comparison only means something on a held-out set of the site's frames, read class by class, and that the score measures domain fit rather than model quality. It is for operations and engineering teams deciding whether to train at all.
Ayman Quadir · Head of Product · Sep 28, 2026

Loading dock with a forklift carrying a pallet, trucks at the bays, generated scene with detections from our model
At 6 am on the cross-dock a forklift comes through the end of bay 4 with a pallet raised to head height, which is the one thing the safety rule cares about. The general detector the team tried first drew a box on it and called it a truck. The pallet stack behind it was a box. The person walking the aisle beside it was found, which was something, and the raised load, the thing the rule is about, was nothing at all.
None of that is a flaw in the general detector. It was trained on photographs from the internet, where forklifts are rare, raised loads are rarer, and nobody labels "pallet left outside a bay". It does exactly what it was built to do, on the pictures it was built from.
The question for the site is different: which model is right about this floor, from these cameras, at this hour. That question has a method, and the method is the same whether the answer turns out to be the general service or a model trained on the site's own frames.
A general detector has never seen this floor
The classes a warehouse needs are specific. A forklift with the forks raised and one with them lowered are the same object to a general model and two different events to the shift lead. A pallet on the floor outside a marked bay is a hazard; the same pallet inside the bay is Tuesday. A person inside the red zone around a reversing truck is the alert; a person on the walkway is nobody.
A general detector has labels for "truck", "person" and "box". It has no label for "raised load" because nobody on the internet needed one. So the comparison starts by writing down the classes the site's rules actually use, and then noticing that the general model can only be scored on the subset of those it has a word for.
On the ones it has no word for, its score is zero, and no benchmark figure changes that.
The held-out set is a week the model never saw
The frames used to judge either model have to come from the site's own cameras, and they have to be frames neither model trained on. The simplest honest set is a week: every camera, every shift, sampled every couple of seconds, with the night shift and the Sunday inbound left in rather than trimmed out because they are quiet.
A person labels that week. Lexi puts a first box on every frame, the person checks each one, and the checked labels are the ground truth both models are measured against. The labeling matters more than the model choice, because a held-out set with missing boxes on the raised loads will make the model that finds them look wrong.
Keep the held-out week apart from anything used for training. If frames from Wednesday appear in both, the trained model scores on things it has memorised and the comparison tells you nothing. The lesson on how much footage you actually need makes the same point from the other side: distinct conditions are the unit, and a week across three shifts holds more of them than a month of the day shift.
Compare class by class or the forklift disappears into the average
A single overall number is what everyone asks for and the last thing to read. The warehouse has a few classes that appear on every frame, "person" and "pallet", and a few that appear once an hour, "raised load" and "pallet outside bay". An average across classes is dominated by the common ones, and the common ones are exactly where a general detector does fine.
Put the per-class table on the wall for the Monday briefing instead. On the forklift row, the general model finds most of them and calls them trucks, which a rule can be written around. On the raised-load row it has nothing. On the person row the two models are within a whisker of each other, and that row is the one a general service will quote at you.
The site's drivers, as it happens, have names for the forklifts. The oldest one is "the tractor" on the night shift and "number two" on days, and neither name is a class.
Precision and recall are read against what each mistake costs
For every class there are two ways to be wrong, and the site has an opinion about which is worse. A raised-load alert that fires on a lowered fork at 2 pm costs the shift lead a glance at the frame in Slack. A raised load that goes unflagged is the one that ends up in an incident report. So for that class the site wants recall, and will pay for it in false alarms up to the point where the shift lead stops looking.
For the pallet-outside-bay class the arithmetic flips. Nobody is hurt by a pallet the model missed for ten minutes, and a false alarm every time a pallet crosses a painted line trains everyone to ignore the channel. Precision and recall get set per class with that in mind, and the threshold that produces them is a dial the site turns, which is why a comparison at one fixed threshold says little. Read the number beside each detection as a ranking of that frame against the others, as the lesson on what the number beside a detection means explains, and compare the two models at the operating point the rule will actually run at.
The score measures domain fit rather than model quality
When the model trained on the site's frames wins, and on the site's own classes it will, the temptation is to conclude it is the better model. It is the better fit. A general detector that has seen a million photographs of streets is a fine piece of engineering that was never asked about this floor. The trained model would be lost on a street.
That distinction decides what to do when the numbers are close on the shared classes. Close on "person" means the general service is a reasonable choice for a rule that only needs people, and the Monday count of people on the walkway is such a rule. Anything that needs the site's own vocabulary needs frames from the site, and no amount of tuning the general model's threshold produces a class it does not have.
My own view is that a general service is the right answer on the first day and the wrong one by the second month. It shows the team what a box on their footage looks like before anyone has labeled anything, and it is worth that. The site's rules then outgrow its class list, one hazard at a time.
A few hundred frames start it and the corrections grow it
The trained model that won the comparison began as a few hundred labeled frames from those same cameras, and it did not stay that size. LexData takes a vision model through its whole life. You type what to look for, Lexi puts a box on every frame, and a person checks each label before anything trains on it. The model then watches the cameras the warehouse already has, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime.
The held-out week does not retire when the comparison is over. It becomes the fixed reference every later version is scored against, so the number on the raised-load row can be read the same way in March as it was on the day the team chose. When a new camera goes up over the outbound doors, its first week of frames joins the set the same way the first one did, checked by a person before it counts. The rest of the platform is built around that habit: the frames the model doubts are the ones a person sees next.
See it on your own footage.
Start with your footageMore in Operations

Operations · 7 min read
Build or buy the layer that keeps a vision model accurate
Two engineers and a pilot can build a detector in a month. The review queue, the versioning and the rollout are what they are still building a year on.
Ayman Quadir · Sep 28, 2026

Operations · 7 min read
Benchmark the model on your own cameras before you believe a score
Two candidate models, two published scores, and a packaging line that only cares which one finds the torn label. The held-out week from your cameras decides.
Rob Hickey · Sep 28, 2026

Operations · 6 min read
Computer vision on the industrial HMI, the camera's verdict on the operator screen without the noise
The verdict as a state, the frame one click away, an alarm the operator acknowledges like any other, and everything else kept off the screen.
Andreas Ohrvall · Sep 28, 2026