Skip to content
LexDataLexData
PlatformIndustriesCustomers
DocsThe Field GuideBlogWhy models drift
AboutCareersSecurityContact
Log inStart now
← All posts

Computer vision · 6 min read

Image classification workflows, when a tag on the frame is enough and a box is too much

One panel per frame at the inspection station wants a tag rather than a box, and a tag costs a fraction of what boxes cost to label and to check.

Summary

This post walks through an image classification workflow on two cameras, an inspection station that sees one steel panel per frame and a shelf bay that is stocked, low or empty, and shows where a single tag per frame answers the question. It sets the labeling cost of a tag against the cost of boxes and argues that most plant problems that get labeled as boxes could have been tags. It is for teams sizing a labeling budget before a first model.

Sheikh Srijon · GTM Lead · Oct 4, 2026

Stamping line, steel panels on a conveyor passing an inspection station, conveyor belt, press and steel sheet boxed, generated scene with detections from our model

At the inspection station after press 3 the stamping line presents one steel panel per frame, centred under the light bar, at the same distance every time. The question the line asks about that frame is whether the panel passes. One frame, one panel, one answer. A model that returns a single tag per frame answers it, and a model that returns boxes answers it with more machinery than the question needs.

Knowing which of those two a camera needs is the first decision on any project, and it is the decision that sets the labeling bill.

A tag answers a question about the whole frame

Image classification returns one label for the whole frame: pass or reject, stocked or low or empty, door open or door closed. There is no position in the answer and no count. The workflow is correspondingly plain. A few hundred frames from the press 3 camera, a tag on each, a model trained on the tagged frames, and from then on a tag per frame as the frames arrive.

The tag can carry more than two states. The shelf camera on the dairy bay reports stocked, low or empty, and the panel station could report pass, rework or scrap. What it cannot carry is where, and the workflow only works when nobody downstream needs where.

One part per frame is the condition that makes a tag enough

The inspection station after press 3 makes the tag possible by how it presents the panel. One part, centred, the same scale on every frame, so that the frame and the part are the same thing and a question about one is a question about the other. Move the same camera to a wide view of the whole line with six panels in shot and the tag stops meaning anything, because a reject tag on a frame of six panels does not say which one.

The dairy bay works the same way. The camera is cropped to one bay, so the frame is the bay, and stocked, low or empty describes the whole of it. A store that wants to know which facing is empty has asked a different question and needs boxes.

Labeling a tag costs a click and a box costs a decision per object

Here is where the two workflows part, and it is the reason to decide early. Tagging a frame is one decision: a person looks at the panel and presses pass or reject. Boxing a frame is a decision per object, and each decision has edges to place, so a checker has to agree with the labeler on where the edges go. On the panel station a tag takes a second or two. A set of boxes on the same panel takes many times that, and the checking pass takes longer still.

Lexi proposes the labels either way, tags or boxes, and a person checks each one before anything trains on it. The labeling doc covers how a prompt becomes a clean class list and what the verification pass does. For a tag the verification pass is fast, and the budget that would have gone on box edges goes on more frames, which is usually the better spend on a station that sees the same part all day.

The inspector at press 3 chalks a circle on the panel where the defect is before it goes to the rework bench. A tag model will never do that, and if the rework bench needs the circle, the tag was the wrong choice from the start.

Object detection earns its cost when the question is where

Boxes are worth their price when the answer has a position in it. The store that needs the empty facing on the dairy bay, the panel station whose rework bench needs the location of the dent, the yard that needs to know a person is inside the fence at 3 am rather than somewhere in the frame. Each is a where question, and object detection is the workflow for where. When the answer is also a severity, size on the part, the lesson on what a vision model can and cannot see is the right read. Severity needs an outline rather than a box, and that is a third workflow with a third labeling bill.

The three questions, is it, where is it, how bad is it, are three different labeling jobs, and the frames do not tell you which one you are in. The person who has to act on the answer does.

A tag hides its reason and that is the price of it

A tag model that says reject does not say why. On the panel station that is fine while the model is right, and a problem the day it is wrong, because a rejected panel with no marked defect is a panel a person has to inspect again from scratch. The tag also gives the model room to learn the wrong thing. If the light bar over press 3 flickers on the frames that happened to be rejects in the training set, the model can learn the flicker, and nobody sees it in a tag.

The check for that is the same as for any model: a held-out week of frames from the station, evaluated by hand, with the wrong tags looked at one by one. Wrong tags that cluster around a time of day or a coil change are the model reading the background.

The doubted frames give the tag model its next version

LexData takes the station model through its whole life. You type what to look for, Lexi puts a tag on every frame, and a person checks each label before anything trains on it. The model then watches the cameras the plant already has, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime.

On the dairy bay the doubt lives at the boundary between low and empty, and that is where the corrections land. When staff override the tag more often, the override rate is the signal, the corrections cross the project's threshold, and a version trained on the boundary frames replaces the old one.

My own view is that more plant problems are tags than boxes, and that most of them get labeled as boxes because the first demo anyone saw was a box. A station that presents one part per frame is asking a yes or no question, and the cheapest correct label is the one to buy.

See it on your own footage.

Start with your footage

More in Computer vision

Computer vision · 7 min read

Five computer vision applications in production, on the cameras a site already owns

Detect, count, inspect, track and read are the five jobs a fixed camera can be given, and each is a different question with a different label behind it.

Ayman Quadir · Oct 4, 2026

Computer vision · 6 min read

Computer vision projects worth building on the cameras you already have

A plant, a utility, a grower, a store and a warehouse each have a project that starts on an existing camera, produces a decision, and has someone to act on it.

Ayman Quadir · Oct 4, 2026

Computer vision · 6 min read

How to choose an object detection model architecture for a camera on your own site

Start from the decision the plant has to make, then the box beside the recorder the model has to fit, and only then the family. The benchmark comes last.

Andreas Ohrvall · Oct 4, 2026

LexData
LexData

Product

  • Platform
  • Industries
  • Use cases

Resources

  • Docs
  • The Field Guide
  • Blog
  • Why models drift

Industries

  • Energy & utilities
  • Oil & gas
  • Agriculture
  • Manufacturing
  • Insurance
  • Retail
  • Robotics

Company

  • About
  • Customers
  • Careers
  • Contact

Trust

  • Security
  • Privacy
  • Terms

Stay updated

What we learn running vision models in production.

See everything.
Miss nothing.

Stay updated

What we learn running vision models in production.

Terms of use & Privacy policy

© 2026 LexData Labs · All rights reserved