Skip to content
LexDataLexData
PlatformIndustriesCustomers
DocsThe Field GuideBlogWhy models drift
AboutCareersSecurityContact
Log inStart now
← All posts

Edge · 6 min read

Inference latency in computer vision, and the budget between the camera and the reject arm

A carton takes a fixed time to reach the reject arm from the lens. Everything from capture to the signal has to fit inside it, and most of it is not the model.

Summary

This post works through end-to-end latency for three plant scenes, a conveyor with a reject arm downstream, a robot reaching for a moving part, and a guarded zone beside a press, from capture and decode to the signal that acts. It concludes that decode and the round trip eat more of the budget than the model does, that the runner beside the recorder removes the round trip, and that a reject gate at line speed is a job for the exported model inside the plant's own control loop. It is for controls and vision engineers sizing a station.

Esdras Ntuyenabo · Engineer · Oct 1, 2026

Cartons and labels boxed on a packaging line, generated scene with detections from our model

The reject arm on packaging line 2 sits eight hundred millimetres downstream of the camera, and at the belt's speed a carton covers that distance in under a second. Everything the station does, from the moment the label passes the lens to the moment the solenoid fires, has to fit inside that second, with margin. The solenoid's click can be heard from the office when it fires, and on a bad afternoon the office hears it fire late, on the carton behind the one it was meant for.

Latency is the time between the frame and the act. The model is one item on the list, and rarely the largest.

The budget is set by the belt, not by the model

The budget for line 2 is a physical number: the distance to the arm divided by the belt's speed. It does not care how fast the model is. If the belt speeds up for a big order, the budget shrinks and the station either fits inside it or rejects the wrong carton.

So the station's design starts with that number written on the drawing, and every stage between capture and signal gets a share of it. What is left over is margin, and the margin is what absorbs the afternoon the recorder is busy or the box in the cabinet is warm.

Where the time goes between the frame and the signal

Follow one carton. The camera exposes the frame and hands it to the encoder. The stream reaches the box, which decodes it back into pixels. The pixels are resized to what the model expects. The model runs. Its output is filtered, boxes below the threshold dropped, overlapping ones merged, and the rule decides whether this carton fails. The decision becomes a signal to the PLC that fires the arm.

Decode is the surprise on most stations. A compressed stream is cheap to carry and expensive to unpack, and on a small box the unpacking is often more work than the model. The resize is small but not free. The forward pass is what everyone measured in the pilot. The post-processing is small until the frame holds many cartons. The signal to the PLC crosses whatever wiring the plant has between the box and the panel, and that path was probably not measured at all.

My own view is that decode should be measured before the model is chosen, because a station that spends most of its budget turning the stream into pixels does not need a faster model.

Throughput and latency are different numbers

The pilot report said the model handled many frames a second. That is throughput, and it was measured with frames fed in batches, which is how a model goes fastest. Latency for one carton is the time that one frame waits for its batch to fill, plus the time the batch takes. A bigger batch raises throughput and raises latency at the same time.

For line 2, the batch is one. Each carton is decided on its own frame as soon as the frame exists, and the throughput number from the pilot has nothing to say about whether the arm fires in time.

The zone beside the press has a different budget

The second scene at the plant is the guarded zone beside press 3. A person stepping into it has to be noticed, and the budget is the time a person takes to walk from the line on the floor to the press, which is seconds rather than fractions of one. The platform samples each camera about every two seconds, one consumer per stream, and for the hazard zone intrusion question that is inside the budget, with the alert delivered to Slack, email or a webhook.

The same rule about the zone living in image coordinates applies: nudge the mount and the red zone covers the walkway, the detections are unchanged, and every alert is wrong. The latency budget is not the fragile part of this station. The camera's mounting is.

The robot reaching for a moving part has the tightest budget

The third scene is the arm at cell 4 that picks parts off a moving belt. Here the model's answer is a position, and the position is stale by the time the arm gets it, because the part has moved. Every stage of latency becomes a distance error in the grasp. A station like this either predicts where the part will be from where it was, or slows the belt at the pick point, or accepts the misses. What it cannot do is add a round trip to a server somewhere and hope.

A runner beside the recorder removes the round trip

The deployment doc lists the three things to plan headroom for on edge hardware: sustained load in a sealed enclosure, a swapped camera, and decode as the real bottleneck rather than the network. On a runner in the line 2 cabinet, the stream travels a short cable, the decode happens beside the recorder, and there is no network leg in the budget at all. The footage stays on site, and the only frames that leave are the ones the model is unsure of, sent to a person to review.

LexData takes the label model through its whole life. You type what to look for, Lexi puts a box on every frame, and a person checks each label before anything trains on it. The model then watches the line 2 camera, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime.

For the reject arm itself, at belt speed, the model is exported as PT, ONNX or TorchScript with the device preset for the box in the cabinet, a Jetson or a GPU server. It runs inside the plant's own control loop, where the controls engineer can measure the whole path from lens to solenoid with a scope. The platform's sampled watching answers the zone by press 3. The arm's second belongs to the plant.

See it on your own footage.

Start with your footage

More in Edge

Edge · 6 min read

Computer vision on multiple video streams from one runner

Twenty cameras, one box beside the recorder. Fair sampling, a slow camera that drops its own frames, an unplugged one that stalls nobody, one heartbeat each.

Andreas Ohrvall · Oct 1, 2026

Edge · 7 min read

CPU vs GPU for computer vision inference, and when a CPU is enough

A cap station checked every couple of seconds and a shelf camera sampled every few minutes both run on a CPU. The GPU earns its keep when the load piles up.

Andreas Ohrvall · Oct 1, 2026

Edge · 6 min read

Deploying computer vision models to edge devices beside the camera

A robot cell that cannot wait for a round trip and an orchard with no uplink. Export by device, run beside the recorder, and get the next version out there.

Andreas Ohrvall · Oct 1, 2026

LexData
LexData

Product

  • Platform
  • Industries
  • Use cases

Resources

  • Docs
  • The Field Guide
  • Blog
  • Why models drift

Industries

  • Energy & utilities
  • Oil & gas
  • Agriculture
  • Manufacturing
  • Insurance
  • Retail
  • Robotics

Company

  • About
  • Customers
  • Careers
  • Contact

Trust

  • Security
  • Privacy
  • Terms

Stay updated

What we learn running vision models in production.

See everything.
Miss nothing.

Stay updated

What we learn running vision models in production.

Terms of use & Privacy policy

© 2026 LexData Labs · All rights reserved