Skip to content
LexDataLexData
PlatformIndustriesCustomers
DocsThe Field GuideBlogWhy models drift
AboutCareersSecurityContact
Log inStart now
← All posts

Labeling · 6 min read

Zero-shot detection vs custom training, where a prompted model runs out

A prompted model boxes the forklift first try and has never seen a wafer scratch. Test on a small labeled set, then let corrections train the model that has.

Summary

This post puts a prompted, zero-shot detector on two cameras, a warehouse rack aisle and a wafer inspection line, and shows where it works from the first frame and where it has no prior to draw on. It names the two reasons it runs out, what the model was trained on and how far the camera's footage sits from that, and describes the handover to a trained model with a person checking every label along the way. It is for teams deciding whether to prompt or to train.

Andreas Ohrvall · CTO · Oct 3, 2026

Warehouse aisle with racked pallets from a high camera, forklift and pallet boxed, generated scene with detections from our model

Two cameras were set up in the same week in March by the same engineer. The first looks down a rack aisle in a distribution warehouse, and the question is where the forklifts and pallets are. The second looks at wafers on an inspection stage in a clean room, and the question is whether there is a scratch across the surface. On both, the first thing tried was a prompted model: type the class, get boxes, no training.

On the aisle it worked in the first hour. Forklift, pallet, person: the boxes landed where they should have, and the count down the aisle matched what a person saw. On the wafer stage it returned nothing useful. Asked for a scratch, it boxed the wafer. Asked for a defect, it boxed the fixture. Asked for a line, it boxed the edge of the stage.

The same model, the same engineer, the same week. The difference was in what the model had seen before it ever saw either camera.

A prompted bounding box works where the warehouse fits the prior

A zero-shot detector was trained on a very large collection of images with captions, and what it learned is a mapping from words to appearances across that collection. Forklifts are in it. So are pallets and people and trucks and shelves, the ordinary furniture of the world, photographed and described in words more times than anyone has counted. A warehouse aisle, the one on camera 4 at the distribution site, is that furniture seen from a slightly unusual angle. The model has a prior for every object in the frame.

That is why it works on the aisle from the first frame, and it is worth being honest that on a camera like that a prompted model may be all a team needs for a while. Counting forklifts down an aisle from a high mount is a question a general prior answers.

A bounding box from a prompted model is still a proposal rather than a label, and the aisle has its own surprises. A pallet jack is a small forklift to a model that has seen both. A shrink-wrapped pallet on a high rack is a white block. The first week's frames need a person looking at the proposals, and the proposals the person corrects are the beginning of a dataset for that aisle.

The wafer has no prior to draw on

Nobody on the internet captions wafer scratches. The images the model was trained on include a few of wafers, fewer of scratched ones, and none labeled at the scale and lighting of an inspection stage in a clean room. The scratch the line cares about is a faint line a few pixels wide on a reflective surface under ring lighting, and the words "scratch" and "defect" in the model's vocabulary point at something else entirely, a scratched car door, a cracked phone screen.

Two things are wrong at once. The first is the training collection's bias: the objects the model can find are the objects people photographed and described, and a wafer scratch is neither. The second is domain shift: even for a word the model has a prior for, the camera's footage sits far from the images the word was learned on. A "line" in the model's memory is a road marking or a queue. On the stage it is a defect.

There is no prompt that fixes this. A better sentence does not give the model a prior it never formed. The engineer spent a Tuesday on phrasings before concluding what the first ten frames had already said.

A small labeled set is the honest test

The way to know which camera you have is to label a small set and score the prompted model against it. A hundred frames from each camera, boxed by a person on a Friday against a written class definition, and the prompted model's proposals compared with them, per class. On the aisle the forklift and pallet classes scored well and the pallet jack scored poorly. On the wafer stage every class scored near nothing.

The test costs a day and settles the question that otherwise gets argued for a month. The footage lesson makes the point that the unit of a dataset is distinct conditions rather than hours, and the hundred frames should span those conditions: day and night on the aisle, both lighting presets on the stage. A prompted model that passes on the easy hundred and fails on the hard hundred has told you where its prior ends.

My own view is that this test should be run before any prompted model goes near a live feed, and the result kept. Six months later, when someone asks why the wafer line has a trained model and the aisle does not, the two tables are the answer.

The handover is the corrections becoming a dataset

Where the prompted model runs out, a trained one takes over, and the handover is less of a cliff than it sounds. The platform starts every project the same way: you type what to look for, Lexi proposes a box on every frame, and a person confirms, fixes or rejects each one before anything trains. On the aisle the proposals are mostly right and the person's job is the pallet jacks. On the wafer stage the proposals are mostly wrong and the person draws the scratches by hand, on a hundred frames, then a few hundred.

The trained model has the prior the prompted one lacked, built from those corrected frames. The corrections keep coming after it is live. Frames it is unsure of come back to a person, the person's fix is a label, and when the corrections cross the project's threshold a new version trains and replaces the old one with no downtime. The wafer model that launched on a few hundred scratches is, a year on, a model trained on every scratch the line produced that a person confirmed.

On the aisle, the same loop runs on a model that started from a stronger place. The pallet jack that the prompted model missed is a class the trained one learned in its first week.

Which model runs live is decided by where it runs

The aisle camera and the wafer stage differ in one more way that matters after the accuracy question is settled. A prompted model is large, and it runs in the cloud at a cost per frame. A trained model for one camera's classes is small enough to run on a runner beside the recorder, sampling the feed about every two seconds, with the footage staying on site and only the doubted frames leaving.

For the clean room on line 1 that was the deciding argument before accuracy was. The wafer footage does not leave the building, so the model that watches it has to live there, and a model that lives there has to be one the team trained. The engineer who set up both cameras in the same week now has a trained model on each, and the aisle's started as a prompted one.

See it on your own footage.

Start with your footage

More in Labeling

Labeling · 7 min read

Aerial dataset augmentation for drone frames where there is no up

A tower seen straight down has no top or bottom, so rotations and both flips are safe. Scale for altitude, brightness for sun, move every box with its pixels.

Esdras Ntuyenabo · Oct 3, 2026

Labeling · 6 min read

A collaborative data annotation workflow run as a pipeline

Batch the bottling line's frames by camera and shift, assign so nothing is boxed twice, attach the guideline, review every label, train on the approved set.

Rajiya Sultana · Oct 3, 2026

Labeling · 7 min read

Dataset health check for computer vision, what to look at before anything trains

A scratch dataset where every scratch sits in the centre of the frame will train a model that looks in the centre. Five counts to read before the first epoch.

Rajiya Sultana · Oct 3, 2026

LexData
LexData

Product

  • Platform
  • Industries
  • Use cases

Resources

  • Docs
  • The Field Guide
  • Blog
  • Why models drift

Industries

  • Energy & utilities
  • Oil & gas
  • Agriculture
  • Manufacturing
  • Insurance
  • Retail
  • Robotics

Company

  • About
  • Customers
  • Careers
  • Contact

Trust

  • Security
  • Privacy
  • Terms

Stay updated

What we learn running vision models in production.

See everything.
Miss nothing.

Stay updated

What we learn running vision models in production.

Terms of use & Privacy policy

© 2026 LexData Labs · All rights reserved