Skip to content
LexDataLexData
PlatformIndustriesCustomers
DocsThe Field GuideBlogWhy models drift
AboutCareersSecurityContact
Log inStart now
← All posts

Operations · 6 min read

Semantic search across a hundred live camera feeds with one sentence

An operator types what the rare scene looks like and the yard cameras return the frames that match. Search finds candidates; a person confirms them.

Summary

This post follows an operator who types a description of a rare scene and pulls the matching frames from a fleet of yard cameras sampled every couple of seconds. It explains how vision language models score a frame against a sentence, why that makes search the fastest way to find the scene worth labeling, and why search returns candidates that a person still has to confirm. It is for the person who has to find twelve frames of something in a month of footage.

Sheikh Srijon · GTM Lead · Oct 2, 2026

Thermal view of a yard at night, seven people boxed, from a customer site camera

A distribution yard has a hundred cameras and one operator on the night desk. Twice a month, usually after 2 am, somewhere in that yard, a driver climbs onto a trailer to fix a strap, which is the thing the safety team most wants a model to catch and the thing there is almost no footage of. The operator has been asked to find examples. He has a month of recordings and no idea which camera or which night.

Scrubbing through a hundred feeds is not a plan. Writing a detector for "person on trailer" needs the examples he does not have. What he can do is type a sentence.

The scene worth labeling happens twice a month on a hundred cameras

The problem has a shape that comes up on every site with a fleet of cameras. The common scenes, trucks at the gate, forklifts in the lanes, people at the office door, are everywhere in the footage and easy to label. The rare scene, the one the rule is really for, is a handful of frames scattered across weeks and cameras, and the detector cannot be trained until someone finds them.

The lesson on what a vision model can and cannot see puts it as a footage question: the timeline is governed by the least common class you promised to detect. On the yard, that class is a person on a trailer, and the least common class is also the hardest to go looking for by hand.

The operator's predecessor kept a notebook of dates and camera numbers for incidents, and the notebook is the only index the yard has ever had. It stops in March.

Vision language models score a frame against a sentence

The reason a sentence can search footage is that vision language models put a frame and a piece of text in the same space. A frame of a person climbing on a trailer and the words "a person climbing on a trailer" land close together there, and a frame of an empty lane lands far from both. Every sampled frame from every camera gets a position in that space once, when it is sampled. A query is the sentence's position, and search is a question of which frames sit nearest it.

That is the whole trick, and it works without any training on the yard's footage. The model has seen enough of the world to know roughly what "climbing" and "trailer" look like, which is nothing like enough to run a safety rule on and exactly enough to shortlist frames for a person to look at.

The operator types "a person standing on top of a trailer at night" into LexInsight and gets a page of frames from across the fleet, ranked. Some are right. Several are a driver standing on the fuel tank step. One is a bird.

Sampling every couple of seconds keeps the index the size of a day

A hundred cameras at full frame rate would be an index nobody could afford to build or search. Every camera on the platform is watched in parallel, one consumer per stream, sampled about every two seconds, and the search index is built on those samples. A person climbing a trailer takes longer than two seconds, so the scene appears in several samples and the index misses nothing that matters.

The frames stay where they were recorded. With a runner beside the recorder, the position of each frame in the shared space leaves the site and the footage does not, so the search runs over positions and returns pointers to frames the operator then opens locally. A month of the yard is a month of small vectors rather than a month of video, and the runner at gate 3 holds the video.

Search returns candidates, and a person confirms them

The honest limit of the approach is in the page of results. Search returns what is near the sentence, and near is not the same as correct. The driver on the fuel tank step is near. The bird on the trailer roof is near. The operator's job is to open each candidate, look at the frame, and say yes or no, and that job takes minutes per query rather than the days that scrubbing would.

This is LexInsight at its most useful: ask the footage a question, get the frames behind the answer, and decide for yourself. It is not a count and it is not a rule. The night the operator asked for people on trailers, he confirmed eleven frames out of forty candidates, from six cameras, across three weeks. Eleven frames is a start.

My own view is that search is the most underused tool on a site with more than a few dozen cameras, and the reason is that people expect it to be the detector. It is the thing that finds the detector's training set.

The confirmed frames are the rare class the detector needed

Eleven confirmed frames go into a project as the first examples of "person on trailer". The operator types the class, Lexi proposes a box on each person in each frame, and he checks each one. The labeling guide covers what happens on that verification pass; the short version is that a small set of correct labels on the rare class is worth more than a large set of anything else.

LexData takes the yard model through its whole life. You type what to look for, Lexi puts a box on every frame, and a person checks each label before anything trains on it. The model then watches the yard cameras, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime. The first version, trained on eleven frames and a lot of ordinary yard, is weak, and its doubted frames are the next search.

Because that is what happens next: the model starts returning frames it is unsure of, and the operator runs the sentence again every week to catch the ones it did not even doubt. Between the two, the rare class grows to something a rule can stand on.

What a sentence cannot find

Some scenes have no sentence. The safety team also wanted "a trailer whose landing gear is not fully down", and no description of that returned anything useful, because the difference is a few pixels of geometry that the shared space does not care about. That one needed a person walking row 4 with a camera for an afternoon, and a detector trained on what they brought back.

Search finds what looks like something. It does not find what is subtly wrong with something. Knowing which of the two a rule needs is the first question, and the sentence is the fastest way to find out.

See it on your own footage.

Start with your footage

More in Operations

Operations · 6 min read

Batch video analysis of archived drone survey footage without a notebook open

Three seasons of right-of-way flights in a cloud bucket. Import the originals, sample the frames, run the model, and review only what it doubted.

Stephen Biswas · Oct 2, 2026

Operations · 7 min read

Computer vision event logging that keeps the frame with the prediction

A missed bone fragment on the night shift can only be explained if the frame, the prediction, the model version and the lot were logged together.

Rob Hickey · Oct 2, 2026

Operations · 6 min read

How big a computer vision model the line camera actually needs

Nano to large on one weld cell camera. What a bigger model buys on look-alike defects, what it costs on the runner, and choosing by the device it runs on.

Rob Hickey · Oct 2, 2026

LexData
LexData

Product

  • Platform
  • Industries
  • Use cases

Resources

  • Docs
  • The Field Guide
  • Blog
  • Why models drift

Industries

  • Energy & utilities
  • Oil & gas
  • Agriculture
  • Manufacturing
  • Insurance
  • Retail
  • Robotics

Company

  • About
  • Customers
  • Careers
  • Contact

Trust

  • Security
  • Privacy
  • Terms

Stay updated

What we learn running vision models in production.

See everything.
Miss nothing.

Stay updated

What we learn running vision models in production.

Terms of use & Privacy policy

© 2026 LexData Labs · All rights reserved