Skip to content
LexDataLexData
PlatformIndustriesCustomers
DocsThe Field GuideBlogWhy models drift
AboutCareersSecurityContact
Log inStart now
← All posts

Computer vision · 6 min read

What the big vision models are for, and what runs on the camera

The big models find the frames, draw the first boxes and tighten them, and a person checks. The small model trained on those labels runs beside the recorder.

Summary

This post lays out where the large general vision models belong in a camera project: finding the frames worth labeling, drawing the first boxes from a phrase, and tightening those boxes into outlines, always with a person checking the result. It concludes that none of them run on the small box beside the recorder, and that the compact model trained on the checked labels is what watches the camera. It is for engineers deciding which model goes where.

Andreas Ohrvall · CTO · Sep 29, 2026

Small compute box beside a video recorder in a plant cabinet, recorder and cable boxed, generated scene with detections from our model

In the control cabinet at the back of a bottling plant there is a video recorder, and beside it a compute box the size of a paperback, with a fan that runs at 3 pm when the room warms up. The plant's engineer has read about the large vision models and asks the reasonable question: can one of those run on that box, so the plant gets the benefit without training anything.

It cannot, and the reason is the box. What the large models are for is the work that happens before anything runs on the box at all: finding the frames, drawing the first labels, and getting a training set together in days rather than months. The small model that eventually watches the fill head is a different animal, trained on the plant's own checked labels, and it is small on purpose.

The big models carry the internet's vocabulary

The general models were trained on a large slice of the internet's pictures and captions, and what they carry is the internet's idea of things. They can find a bottle in a frame from a plant they have never seen, at 3 pm or at midnight, because bottles are everywhere online. They have a much weaker idea of "cap seated crooked", because nobody captions that. The lesson on what a vision model can and cannot see draws the same boundary from the other direction: vision is good at what is visible and consistent, and the general models add a wide vocabulary to that without adding the plant's.

That vocabulary is the asset. It is what lets a team start from a phrase instead of from a thousand hand-drawn boxes, and it is used up in the first week, after which the plant's own frames take over.

An embedding model finds the frames worth labeling

The first job is choosing frames. The recorder holds months of the fill head, nearly all of it the same scene, and a labeling budget spent evenly across it labels the same bottle a thousand times. An embedding model turns each frame into a point where similar frames sit close. Clustering those points collapses the months into the few thousand frames that actually differ: the changeover on Monday, the shift with the glare from the loading door, the run where the caps came from a new supplier.

It also finds the rare scene by description. "A bottle on its side on the conveyor" typed as a search returns the handful of frames from the year when that happened, which are exactly the frames the training set needs and nobody remembered.

Grounding DINO draws the first box from a phrase

With the frames chosen, a text-prompted detector puts the first boxes on them. Grounding DINO takes a phrase, "bottle cap", and draws a box wherever the phrase applies, on frames it has never seen, without any training on the plant. That is the step that used to be a month of a person's time with a mouse, and it is now an afternoon, plus a small labeled set to check which phrasing of "cap" finds the crooked ones.

The boxes are a first pass. They are loose, they miss the bottles at the frame edge, and the phrase's idea of a cap differs from the plant's. The labeling doc covers how a phrase becomes a class list and what the verification pass looks for; the point here is that the phrase does the bulk labour and the checking makes it true.

A segmentation model tightens the box into an outline

When the rule needs an outline rather than a box, a promptable segmentation model takes each first-pass box and returns the pixels inside it that belong to the object. A box around a bottle becomes the bottle's silhouette; a box around a spill on the Monday changeover frames becomes its area. The person's job on the verification pass shifts from drawing the outline to nudging its edge, which is a fraction of the work.

That is the whole chain: embedding to find, detector to box, segmenter to outline. Each model is large, slow and general, and each is used once per frame, offline, in the labeling room, where slow is fine.

A person checks every label before it trains anything

None of it is a label until a person has looked at it. On every frame, Lexi puts the first box from the phrase, and the person checks each one, moves the loose ones, adds the missed bottle at the edge, and marks the crooked cap the phrase called a cap. The corrections are the difference between a training set that holds the internet's opinion of a bottle and one that holds the plant's.

The cabinet fan, incidentally, is loud enough that the engineer times his visits to the mornings, and the morning frames are the ones the first training set was heaviest in.

A compact trained model is what runs beside the recorder

The model that watches the fill head is trained on those checked labels, and it is small: a handful of classes, tuned to one camera's view, exported for the box in the cabinet through a device preset that describes the box rather than the architecture. It runs on every frame the recorder produces, sampled about every two seconds, on hardware that would take most of a minute to run one of the general models once.

That is the division of labour. LexData takes the compact model through its whole life from there. You type what to look for, Lexi puts a box on every frame, and a person checks each label before anything trains on it. The model then watches the cameras the plant already has, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime.

My own view is that the large models belong in the labeling room and nowhere else, and I would not put one on a runner even if the box could carry it. A general model on the camera is watching for the internet's bottle. The plant's bottle is the one in the checked labels, and only the small model has seen it.

See it on your own footage.

Start with your footage

More in Computer vision

Computer vision · 6 min read

Why a good detector can still be a bad counter

The queue camera finds every shopper on every frame and still gets the count wrong. Swaps, lost tracks and jitter are tracker mistakes, and review fixes one.

Stephen Biswas · Sep 29, 2026

Computer vision · 6 min read

What semantic segmentation labels and what it costs

Drivable path, corrosion area and crop rows are questions about pixels, and the answer is a class for every one. Painted labels cost more than boxes.

Esdras Ntuyenabo · Sep 29, 2026

Computer vision · 6 min read

What YOLO does in one pass and what that costs

One look at the frame is why a single-pass detector fits the box beside the recorder, and why it misses the small cluttered things a two-stage model finds.

Andreas Ohrvall · Sep 29, 2026

LexData
LexData

Product

  • Platform
  • Industries
  • Use cases

Resources

  • Docs
  • The Field Guide
  • Blog
  • Why models drift

Industries

  • Energy & utilities
  • Oil & gas
  • Agriculture
  • Manufacturing
  • Insurance
  • Retail
  • Robotics

Company

  • About
  • Customers
  • Careers
  • Contact

Trust

  • Security
  • Privacy
  • Terms

Stay updated

What we learn running vision models in production.

See everything.
Miss nothing.

Stay updated

What we learn running vision models in production.

Terms of use & Privacy policy

© 2026 LexData Labs · All rights reserved