Skip to content
LexDataLexData
PlatformIndustriesCustomers
DocsThe Field GuideBlogWhy models drift
AboutCareersSecurityContact
Log inStart now
← All posts

Computer vision · 7 min read

Image recognition AI, explained on the cameras a site already owns

A tag, a box, a mask or keypoints, on a substation camera, a line camera and an aisle camera. How a model gets there, and what its score does not promise.

Summary

This post explains what image recognition means in practice by taking three cameras, on a substation yard, a production line and a store aisle, and showing what each asks for: a tag, a bounding box, a mask or a set of keypoints. It covers how a model reaches those outputs from checked labels and why the score beside a detection is a ranking rather than a probability. It is for people who own cameras and want to know what recognition would give them.

Esdras Ntuyenabo · Engineer · Oct 3, 2026

Bread shelf with an empty slot flagged and the rack sections boxed, a shopper at the end of the aisle, from a customer store camera

Three cameras, three questions. The one on the substation yard is asked whether anyone is inside the fence after dark. The one over the line is asked whether the cap is on the bottle. The one over the bread aisle is asked which slots on the shelf are empty at 7 am. All three get filed under image recognition, and the phrase hides the fact that they want different answers in different shapes.

Recognition is the general word for a model that looks at a frame and says something about what is in it. What it says, and in what form, is the first decision on every camera, and it is worth making before anyone talks about models.

Recognition means four different outputs on three cameras

The outputs come in four shapes. A tag for the whole frame from the substation at 2 am: person present, no person. A box around each object: this person, here, this size. A mask that follows an object's outline: this much of the shelf is empty. A set of points on an object: this is where the wrist is. Each is a different model, a different label, and a different cost, and the question the camera is asked picks one.

The what a vision model can and cannot see lesson frames it as three questions being three different problems, and the shape of the output is where the difference first shows. A person asking "is anyone there" and a person asking "how many are there and where" have asked for different models, whichever camera they are looking at.

A tag answers whether and a bounding box answers where

The substation camera's question, is anyone inside the fence, could be answered with a tag. One label per frame, yes or no, and the model learns the difference between a frame with a person and a frame without. That is the cheapest kind of recognition to label and to train. It is also the least useful the moment the question grows, because a tag cannot say where the person is, and inside the fence is a where.

So the substation model uses a bounding box: a rectangle around each person, with a class and a position. The rule "inside the fence" becomes a comparison between the box and a zone drawn on the frame, and the same model answers "how many" and "near which transformer" without being retrained. On the line, the cap question is a box too, one around each bottle and one around each cap, and a bottle box with no cap box inside it is the miss.

The bread shelf is restocked at 7 am and the first frame after that is the only one all day where every slot is full, which is why the store's reviewer uses it as the reference when checking the empty-slot boxes.

A mask answers how much and keypoints answer which part

The question on the camera over aisle 3 is partly a box question, which slots are empty, and partly something else, because the store also wants to know how empty. A slot with one loaf left and a slot with none are both "low", and the difference between them is area. A mask, an outline that follows the empty space, gives the area. It costs more to label, because a person traces an edge rather than dragging two corners, and it is worth it only where an amount is the answer.

Keypoints are the fourth shape, points on an object rather than an outline around it, and none of the three cameras needs them. They belong to questions about a part of a body or a machine: where a hand is relative to a guard, whether an arm is raised. A site that asks those questions will add a fourth model. A site that does not should not pay for one.

The model gets there from checked labels and nothing else

However the output is shaped, the model learns it from examples. Frames from the camera it will watch, the yard at 2 am or the shelf at 7 am, each with the tag, boxes, masks or points a person would have wanted, and every one of those checked by a second person before training. The model then learns which patterns in the pixels go with which labels, and on a new frame it produces the same kind of output for the same kind of pattern.

There is no other route in. A model that has never seen a frame from the bread aisle with the empty slots boxed does not know what an empty slot on that shelf looks like, whatever it learned elsewhere. This is why the platform starts with labeling on the site's own footage and treats the checked label as the unit of everything that follows.

My own view is that the word "recognition" should mostly be retired in favour of the question. Nobody at the substation wants recognition. They want to know whether anyone is inside the fence, and once the question is written that plainly, the output shape and the labels follow from it.

A score of nine tenths ranks the detection and does not promise it

Every output arrives with a score. A person box on the substation yard at 2 am at nine tenths, a cap box on line 4 at six tenths. The natural reading is that nine tenths means a ninety in a hundred chance the box is right, and that reading is wrong. The score is the model's internal measure of how much the pattern under the box resembles the patterns it was trained to call a person, and modern models are systematically overconfident about that. The what the score beside a detection means lesson goes through it in detail.

What the score is good for is ordering. A box at nine tenths is more likely to be right than one at six tenths on the same camera, and a threshold set on that camera decides which boxes count. The threshold is an operating point, chosen by looking at what the camera produces above and below it, and it is different on every camera. The substation's threshold for a person at night is not the line's threshold for a cap, and neither is a probability.

The aisle camera's misses are the next version's labels

Recognition on a live camera is never finished, because the shelf changes, the bottles change and the fence gets a new gate. LexData takes each model through its whole life: you type what to look for, Lexi puts a box on every frame, and a person checks each label before anything trains on it. The model then watches the cameras the site already has, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime.

On the aisle camera those returned frames are the new bread packaging, the promotional stand that blocked the end of the shelf, the morning the lights came on late. Each one is a slot the model was unsure about and a person was not, and the person's box is a label. The empty-slot model on that shelf has been through several versions since the store's first one, and every version was trained on the frames the one before it got wrong.

See it on your own footage.

Start with your footage

More in Computer vision

Computer vision · 6 min read

What CLIP is and why typing a description finds the frame

A shared space for pictures and words lets an operator type a sentence and get the matching frames from a recorder. A way to look, and the boxes come after.

Rob Hickey · Oct 3, 2026

Computer vision · 6 min read

An end-to-end object detection workflow starts with decisions, and the first frame is labeled last

Tag, box, mask or keypoints. A class list with occlusion and size rules. A metric and an acceptance sentence. Then, and only then, the first camera's frames.

Ayman Quadir · Oct 3, 2026

Computer vision · 7 min read

Image segmentation and the questions on a weld bead that only a mask can answer

A box finds the weld. A mask measures it. Semantic, instance and panoptic told apart on a weld cell and a board camera, with what each costs to label.

Esdras Ntuyenabo · Oct 3, 2026

LexData
LexData

Product

  • Platform
  • Industries
  • Use cases

Resources

  • Docs
  • The Field Guide
  • Blog
  • Why models drift

Industries

  • Energy & utilities
  • Oil & gas
  • Agriculture
  • Manufacturing
  • Insurance
  • Retail
  • Robotics

Company

  • About
  • Customers
  • Careers
  • Contact

Trust

  • Security
  • Privacy
  • Terms

Stay updated

What we learn running vision models in production.

See everything.
Miss nothing.

Stay updated

What we learn running vision models in production.

Terms of use & Privacy policy

© 2026 LexData Labs · All rights reserved