Computer vision · 6 min read
What CLIP is and why typing a description finds the frame
A shared space for pictures and words lets an operator type a sentence and get the matching frames from a recorder. A way to look, and the boxes come after.
Summary
This post explains contrastive language-image pretraining through an operator typing a description and getting matching frames back from a warehouse recorder, covering how the shared space between text and images is learned and what zero-shot means in practice. It concludes that a description search is a starting point rather than an inspector, and that the reviewed boxes and a trained detector are where it hands over. It is for operations teams who want to know what is happening when a sentence finds a frame.
Rob Hickey · Chief AI Officer · Oct 3, 2026

Warehouse aisle from a high camera, a forklift and a racked pallet boxed, generated scene with detections from our model
The warehouse has a month of recorder footage from the camera over aisle 9 and a question from the safety lead: how often does a forklift travel the aisle with its forks raised and a pallet on them. Nobody is going to scrub a month of footage for that. The operator types the sentence into the search box instead, "forklift in an aisle with a pallet raised", and a minute later has a page of frames, most of them showing exactly that.
Nothing in that model was ever trained on aisle 9, on forklifts, or on pallets. It was trained on pictures and the sentences people wrote about them, and the useful thing it learned is a way to put a picture and a sentence in the same place.
Typing a description finds the frame because image and text share a space
The model behind the search has two halves. One turns a frame into a list of numbers, a point in a space with hundreds of dimensions. The other turns a sentence into a point in the same space. The two halves were trained together so that a picture and a sentence describing it land close to each other, and a picture and an unrelated sentence land far apart.
Search is then arithmetic. Every frame from the aisle 9 recorder has already been turned into its point. The typed sentence becomes a point. The frames whose points sit closest to the sentence's point are the results, ranked by distance. No detector ran, no box was drawn, and the frames the operator sees are the ones whose overall appearance the model has learned to associate with those words.
The warehouse calls a forklift a truck, and the operator's first search for "truck with a pallet raised" returned the loading dock. The second attempt said forklift, and the page filled with aisle 9.
The model learned the space by matching captions to pictures
Training used an enormous number of image and caption pairs collected from the open web, none of them from aisle 9. In each training batch, the model was shown many pictures and many captions and asked to match them: for each picture, which of these captions is its own? Getting the matching right pulls a picture toward its caption in the space and pushes it away from the others. That push and pull, repeated over an enormous number of pairs, is what "contrastive" means, and it is the whole training signal. There was never a class list.
That is why the model can be asked about a raised pallet without ever having been taught the phrase. Somewhere in the pairs it learned from were pictures with captions mentioning forklifts, pallets and raised loads, and the space absorbed the association. The what a vision model can and cannot see lesson makes the point that a model learns the appearance of what it is shown; this model was shown the web, and its sense of what a forklift looks like is the web's.
Zero-shot means the class was never in a training set
The word for asking the model about a class it was never explicitly trained on is zero-shot. It works well for things the web has many pictures of, described in ordinary English: forklifts, pallets, high-vis jackets, wet floors. It works badly for things the web has few pictures of, or describes inconsistently: a particular kind of bearing wear, the difference between a good weld and a marginal one, the company's own tote design.
Zero-shot search also depends on the wording more than a person expects. "Forklift with pallet raised" and "forklift carrying a load high" may return different pages, and adding "in a warehouse aisle seen from above" can help or hurt. The operator learns the model's vocabulary by trying, which is fine for search and would be alarming for a safety rule.
A description finds the frame and does not inspect it
This is the line worth drawing clearly. The search over aisle 9 returned frames that look like the sentence. It did not say where the forklift is in the frame, whether the forks are above a certain height, or whether there is a person within two metres. The result is a page of candidates, ranked by resemblance, and a candidate is something a person looks at.
My own view is that the value of these models in operations is search and triage, and that the moment one is asked to give a verdict on a frame, it is doing a job it was not built for. A description search over a month of footage answers "show me the frames where this might be happening" in a minute, which is a question nobody could answer before. It does not answer "is this happening now, and where", and a rule that fires on a description match will fire on a poster of a forklift.
The reviewed boxes are where the description hands over
What the search gives the safety lead is the frames to start from. The page of raised-pallet candidates goes to labeling: you type what to look for, Lexi puts a box on the forklift and the pallet in every frame, and a person checks each label before anything trains on it. The labeling with Lexi doc describes that pass, and the point here is that the frames it starts from were found by a sentence rather than by an afternoon at the recorder.
The boxes are what the description never had. A box has a position and a size, so a rule can be written about it: forks above the second rack beam, a person inside the forklift's box plus a margin. Those are position questions on a fixed camera, and a detector trained on the reviewed boxes answers them on every frame in a way a resemblance score cannot.
The trained detector takes over on the camera and the description stays as the search
Once the detector is trained, it watches the camera over aisle 9 directly, in the cloud, on the warehouse's servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime. The description search does not go away. It is still the fastest way to ask the recorder a new question, "forklift reversing without a spotter", and to find the frames that become the next class.
The two models sit at different ends of the same platform: one for finding what a person did not know to look for, the other for watching what they now know matters. The raised-pallet rule ran on the detector by the end of the month. The search that found the first frames took a minute, and the operator has used it every week since for something else.
See it on your own footage.
Start with your footageMore in Computer vision

Computer vision · 6 min read
An end-to-end object detection workflow starts with decisions, and the first frame is labeled last
Tag, box, mask or keypoints. A class list with occlusion and size rules. A metric and an acceptance sentence. Then, and only then, the first camera's frames.
Ayman Quadir · Oct 3, 2026

Computer vision · 7 min read
Image recognition AI, explained on the cameras a site already owns
A tag, a box, a mask or keypoints, on a substation camera, a line camera and an aisle camera. How a model gets there, and what its score does not promise.
Esdras Ntuyenabo · Oct 3, 2026

Computer vision · 7 min read
Image segmentation and the questions on a weld bead that only a mask can answer
A box finds the weld. A mask measures it. Semantic, instance and panoptic told apart on a weld cell and a board camera, with what each costs to label.
Esdras Ntuyenabo · Oct 3, 2026