Skip to content
LexDataLexData
PlatformIndustriesCustomers
DocsThe Field GuideBlogWhy models drift
AboutCareersSecurityContact
Log inStart now
← All posts

Computer vision · 6 min read

How to choose an object detection model architecture for a camera on your own site

Start from the decision the plant has to make, then the box beside the recorder the model has to fit, and only then the family. The benchmark comes last.

Summary

This post orders the decisions behind choosing a detection architecture for a weld cell camera, starting from the decision the plant needs, then the runner beside the recorder as the constraint, then the family, and only then the benchmark. It argues that the architecture question is usually a budget question in disguise and that the site's own frames settle it. It is for engineers who have been asked which model to use before anyone has said what for.

Andreas Ohrvall · CTO · Oct 4, 2026

Edge box beside a video recorder in a plant control cabinet, cable and recorder boxed, generated scene with detections from our model

The question arrived in the form it usually does: which detector should the weld cell on line 5 use. The camera on the cell looks down at a part on the fixture after the robot has finished the seam, and the plant wants to know whether the seam is complete. Nobody had yet written down what would happen with the answer, or what hardware would produce it, and both of those decide the architecture before any family name is spoken.

The order below is the one that produces a model still running a year later.

The decision the plant needs fixes the shape of the answer

Three answers are possible for the weld cell frame. A tag on the whole frame, seam complete or seam incomplete, is enough if the only action is to divert the part. A box around each gap in the seam is needed if the rework station wants to know where to re-weld. An outline of each pore run is needed if the disposition depends on how long the run is, because length is a severity and only a mask carries length.

That is classify, detect or segment, and the plant's process engineer decides it without ever hearing those words. The lesson on what a vision model can and cannot see opens on exactly this: three questions that sound alike and are three different problems with three different labels.

On line 5 the rework station needs the position of the gap. That settles it as detection.

A bounding box is the right shape when the decision needs a position

A bounding box carries a class and four numbers, and that is precisely what the rework station uses: this gap, here, on this seam. Detection is the workflow whenever the action downstream has a where in it, and the family choice, single pass or two stages, is a question that comes after that and never before.

The labeling follows the shape. Boxes on every gap in a few hundred frames from the cell camera, proposed by Lexi and checked by a person before anything trains. The welders' own word for porosity is "worm holes", and the class list uses their word because they are the ones who will read the alerts.

The runner beside the recorder is the constraint that cuts the list

Where the model runs is the second decision and it removes more candidates than any benchmark. The cell camera's frames can go to the cloud, to a server in the plant, or to a runner in the cabinet beside the recorder. On line 5 the answer is the cabinet, because the plant's network does not leave the building and the footage is not allowed to either. Only the doubted frames leave.

A runner in a cabinet has a fixed budget of compute, one consumer per stream, sampled about every two seconds, and it has to finish with one frame before the next arrives. That budget fits a single-pass detector with a modest backbone and does not fit a two-stage model that looks twice at every region. The deployment doc describes how the model leaves the platform sized for the device: exported as PT, ONNX or TorchScript through a preset that names a Jetson, a Raspberry Pi or a GPU server rather than an architecture.

When the constraint is a server with a GPU and a survey scored after the drone lands, the list is longer and the two-stage families are back on it. The constraint decides.

Single pass or two stages is the last real architectural choice

A single-pass detector looks at the frame once and predicts every box from that one look, at a cost per frame that is roughly fixed. A two-stage detector proposes regions first and looks again at each one, which does better on small objects packed close together and costs more on a crowded frame than on an empty one.

The weld seam is large in the line 5 frame and the gaps in it are a few dozen pixels, which a single pass handles. A tray of small fasteners touching each other is the case that argues for two stages, or for tiling the frame, or for moving the camera closer so the fasteners are no longer small. On the cabinet runner the last of those is usually the cheapest.

Evaluate on the cell's own frames rather than on the benchmark

A benchmark score says how a model did on a public set that contains no weld seams. The number the plant can use is precision and recall on line 5's own frames, per class, on a held-out week the model has never seen. Hold out the most recent week rather than a random slice, because a random slice hides frames from every day in the training set and the score flatters the model.

Then look at the misses one at a time. On the cell they cluster: the gap at the end of the seam where the fixture clamp casts a shadow, the frame where the torch's afterglow washes out the surface. Each cluster is a set of frames to collect and label before the next version, and a list of clusters is a more useful evaluation than any single figure.

The family matters less than the frames, and the frames keep changing

LexData takes the weld cell model through its whole life. You type what to look for, Lexi puts a box on every frame, and a person checks each label before anything trains on it. The model then watches the cameras the plant already has, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime.

The architecture is fixed on the day the model ships. The frames are not. A new fixture in March, a different filler wire from a new supplier, a camera housing replaced after a coolant leak: each changes what the cell frame looks like. The model's answer to each is a version trained on the frames that came back for review, exported through the same preset to the same cabinet.

My own view is that the architecture question is nearly always a budget question in disguise. When a team asks which family to use, what it is asking is what hardware it can afford in the cabinet and how many frames it can afford to label, and once those two are written down the family chooses itself.

See it on your own footage.

Start with your footage

More in Computer vision

Computer vision · 7 min read

Five computer vision applications in production, on the cameras a site already owns

Detect, count, inspect, track and read are the five jobs a fixed camera can be given, and each is a different question with a different label behind it.

Ayman Quadir · Oct 4, 2026

Computer vision · 6 min read

Computer vision projects worth building on the cameras you already have

A plant, a utility, a grower, a store and a warehouse each have a project that starts on an existing camera, produces a decision, and has someone to act on it.

Ayman Quadir · Oct 4, 2026

Computer vision · 6 min read

Image classification workflows, when a tag on the frame is enough and a box is too much

One panel per frame at the inspection station wants a tag rather than a box, and a tag costs a fraction of what boxes cost to label and to check.

Sheikh Srijon · Oct 4, 2026

LexData
LexData

Product

  • Platform
  • Industries
  • Use cases

Resources

  • Docs
  • The Field Guide
  • Blog
  • Why models drift

Industries

  • Energy & utilities
  • Oil & gas
  • Agriculture
  • Manufacturing
  • Insurance
  • Retail
  • Robotics

Company

  • About
  • Customers
  • Careers
  • Contact

Trust

  • Security
  • Privacy
  • Terms

Stay updated

What we learn running vision models in production.

See everything.
Miss nothing.

Stay updated

What we learn running vision models in production.

Terms of use & Privacy policy

© 2026 LexData Labs · All rights reserved