Skip to content
LexDataLexData
PlatformIndustriesCustomers
DocsThe Field GuideBlogWhy models drift
AboutCareersSecurityContact
Log inStart now
← All posts

Operations · 6 min read

How big a computer vision model the line camera actually needs

Nano to large on one weld cell camera. What a bigger model buys on look-alike defects, what it costs on the runner, and choosing by the device it runs on.

Summary

This post compares model sizes from nano to large on one fixed camera over a robotic weld cell, where the hard cases are defects that look alike. It concludes that a larger model earns its place on look-alikes and loses it on a small dataset and a sealed runner, and that the size should be chosen from the device preset and the doubted frames rather than from a benchmark table. It is for engineers deciding what to put on the box beside the recorder.

Rob Hickey · Chief AI Officer · Oct 2, 2026

Robotic welding cell, part on the fixture, generated scene with detections from our model

The camera over robotic weld cell 2 sees the same fixture, the same torch angle and the same seam on every part. The defects it has to find are the problem: a spot of spatter, a pore in the bead and a cold lap all look like a small dark mark on bright metal, and the difference between them decides whether the part is reworked or scrapped. The first model trained for the cell was the smallest on offer, because the runner beside the recorder is a fanless box in a cabinet that gets warm by lunch.

It found the seam every time and the spatter most of the time. It could not tell a pore from a tack.

Whether a bigger model fixes that, and what the bigger model costs, is a question with a specific answer for this camera and no general one.

The hard cases on the weld camera are look-alikes, and size helps there

A small model has fewer features to spend, and it spends them on what separates the common classes from the background. Seam against fixture is easy and it learns that first. Pore against spatter is a fine distinction in texture and edge, and on cell 2 a nano model on a dataset of a few thousand instances tends to collapse the two into "mark on bead" and split them by luck.

Stepping up to a medium model on the same frames, with the same labels, is where the two classes usually come apart. The extra capacity goes on exactly the fine distinction the small model could not afford. On the weld cell, that was the difference between a queue of doubted frames that were mostly pores and a queue that was mostly noise.

That is the honest case for a larger model, and it is narrower than the leaderboards suggest. The benefit lands on the look-alike pairs. On a class the small model already found cleanly, the large one finds it cleanly too, and the extra cost buys nothing.

A larger model costs memory and heat on the runner beside the recorder

The box in the cabinet at cell 2 has a fixed amount of memory and no fan. A medium model fits with room to spare. A large one fits, runs, and after an hour of sustained load in a sealed enclosure the device throttles and the frame budget quietly slips. Nobody sees a crash. The deployment guide puts it plainly: plan for the enclosure, and budget headroom for throttling, for a swapped camera, and for video decode, which is usually the bottleneck before the model is.

So the cost of size is paid in the cabinet, in degrees and megabytes, and it is paid every shift. The gain is paid once, on the confusion between two classes. A plant that wants the large model's pore detection has a choice between a bigger box, a cooler cabinet, or a medium model with better labels on the pores, and the last of those is cheaper than people expect.

Overfitting is what a large model does with little footage

There is a second cost, and it is the one that shows up after the first month rather than in the first hour. A large model has the capacity to memorise the training set, and a weld cell dataset of a few thousand frames from one camera is small enough to memorise. The training curve looks wonderful. The frames from the next production run on cell 2, with a fresh batch of parts and a slightly different torch tip, do less well, because the model learned that fixture and that tip rather than the defect.

Overfitting is a footage problem before it is a model problem. The question in how much footage you actually need is how many distinct conditions the frames cover, and a large model on a narrow set is the fastest way to find out the set was narrow. The fix is more conditions in the training set, sampled from weeks of the cell rather than one afternoon, and the smaller model is the safer choice until those conditions are there.

My own view is that most line cameras are better served by a medium model and another month of doubted frames than by a large model on the frames they have today. The extra capacity has nothing to learn from until the footage is wide enough.

Two seconds a frame is a generous budget, so latency rarely decides

Every camera on the platform is watched in parallel, one consumer per stream, sampled about every two seconds. That is the latency budget, and even a large model on a modest box comes in well under it on a single stream. The reason latency rarely decides the size on a line camera is that the frame rate is not the constraint people assume from a benchmark table.

Where latency does bite is on the number of streams. A box watching cell 2 alone has budget to spare; the same box watching six cells with a large model on each is a different arithmetic, and the honest answer is a smaller model or a second box. The size question and the camera-count question are the same question, asked per box.

The controls engineer on this cell keeps a strip of masking tape on the cabinet with the model version written on it in marker, updated at each rollout. Nobody asked him to. It is the fastest way anyone has found to answer "what is running in there".

Pick the size from the device preset, then let the doubted frames argue

The model exports as PT, ONNX or TorchScript through a device preset that describes the box rather than the architecture: a Jetson, a Raspberry Pi, or a GPU server. The preset already knows what fits, and choosing the size from it is choosing from what the cabinet can run at lunch on a warm day. A leaderboard knows nothing about the cabinet.

LexData takes the weld model through its whole life. You type what to look for, Lexi puts a box on every frame, and a person checks each label before anything trains on it. The model then watches the cell camera, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime. The doubted frames are where the size argument gets settled, because they are the frames the current size could not handle.

If the queue is mostly the pore against spatter confusion, the next size up is worth trying on the same labels. If the queue is mostly a new part or a new torch tip, size will not help and footage will. The queue tells you which, and it tells you every week, which a benchmark table never will. Our manufacturing work holds 99%+ accuracy in production on cells like this one, and the size of the model on the box is the least interesting part of how.

See it on your own footage.

Start with your footage

More in Operations

Operations · 6 min read

Batch video analysis of archived drone survey footage without a notebook open

Three seasons of right-of-way flights in a cloud bucket. Import the originals, sample the frames, run the model, and review only what it doubted.

Stephen Biswas · Oct 2, 2026

Operations · 6 min read

Semantic search across a hundred live camera feeds with one sentence

An operator types what the rare scene looks like and the yard cameras return the frames that match. Search finds candidates; a person confirms them.

Sheikh Srijon · Oct 2, 2026

Operations · 7 min read

Computer vision event logging that keeps the frame with the prediction

A missed bone fragment on the night shift can only be explained if the frame, the prediction, the model version and the lot were logged together.

Rob Hickey · Oct 2, 2026

LexData
LexData

Product

  • Platform
  • Industries
  • Use cases

Resources

  • Docs
  • The Field Guide
  • Blog
  • Why models drift

Industries

  • Energy & utilities
  • Oil & gas
  • Agriculture
  • Manufacturing
  • Insurance
  • Retail
  • Robotics

Company

  • About
  • Customers
  • Careers
  • Contact

Trust

  • Security
  • Privacy
  • Terms

Stay updated

What we learn running vision models in production.

See everything.
Miss nothing.

Stay updated

What we learn running vision models in production.

Terms of use & Privacy policy

© 2026 LexData Labs · All rights reserved