Edge · 6 min read
Edge computer vision for industrial automation, where the camera is the heaviest sensor on the plant network
A plant network built for PLC tags was never built for video. The runner beside the recorder turns frames into events, and only events cross the network.
Summary
This post describes putting computer vision onto a plant network that was built for PLC tags and a historian, not for video. It concludes that the runner beside the recorder is what turns frames into events the PLC and MES can consume, that footage stays in the cabinet with only doubted frames leaving, and that a second sensor on the same moment is where the pairing quietly breaks. It is for controls and IT engineers who own the plant network.
Andreas Ohrvall · CTO · Oct 1, 2026

A compute box beside a video recorder in a plant cabinet, where frames become events, generated scene with detections from our model
The controls engineer at the plant has a naming convention for every tag on the network, and it has held for eleven years. A temperature is a tag. A pressure is a tag. A cycle count is a tag. Each one is a few bytes, a few times a second, from a few thousand places, and the switches in the cabinets were sized for that. Then the camera over station 4 on the packaging line went in, and its stream on its own is heavier than every tag on the line put together.
The network was never built for video, and the mistake is to make it carry any.
A frame is not a reading until something turns it into one
A PLC does not want a picture. It wants a value it can compare against a setpoint: label present, label absent, carton count, a person in the zone. The MES wants an event with a time and a station and a reason. The historian wants a tag. None of them can do anything with a frame, and a frame sent to any of them is a frame that will be stored and never read.
So the question for the camera over station 4 is where the frame becomes a value. If it happens in the cloud, the stream crosses the plant network and the firewall to get there, and the value comes back the same way. If it happens in the cabinet, the frame never leaves it, and what crosses the network is the value, which is the size of a tag.
Object detection on the runner turns frames into events
The runner is a small box beside the recorder in the station 4 cabinet. It takes the stream from the recorder, samples a frame about every two seconds, and runs the model on it. The model does object detection: a box on each carton, a box on each label, a box on any person in the guarded zone. From those boxes the runner produces events. Carton without a label at station 4, 14:07. Person in zone 2 for longer than the window. Count of cartons past the scanner this shift.
Each event is a few bytes. The alert is a rule written as a sentence, with a severity and a cooldown, approved before it goes live, and it is delivered to Slack, to email, or to a webhook, which is the door the MES listens at. The deployment doc describes the boundary the same way: footage processed where it is captured, and what leaves the site is the answer, a detection, a count, an alert, rather than the video.
My own view is that video should not be routed across the OT network at all, even where the switches could take it. The runner sits in the same cabinet as the recorder, on the same short cable, and the network sees events. That keeps the vision system inside the controls engineer's naming convention rather than outside it.
The PLC takes a signal, the MES takes an event
The two consumers want different things from the same detection. The PLC that runs the reject gate at station 4 needs a signal it can act on within the cycle, on a path that does not depend on anything outside the cabinet. That is the plant's integration to build, from the runner's output to the PLC's input, and it is the same shape as any other sensor's wiring loom into the panel.
The MES needs the record. A missing label at 14:07 on station 4, with the frame attached, so that when the shift report asks why the reject count rose after lunch there is a picture of the label reel that ran out. The webhook carries that, and the frame that comes with it is the evidence frame with the box drawn on it, the one frame out of the many the runner looked at that anyone needed to keep.
Footage stays in the cabinet and doubted frames are what leave
The recorder keeps recording, as it did before the runner arrived. The runner watches the stream beside it and the footage does not go anywhere else. The frames that leave the site are the ones the model is unsure of, sent to a person to review, and those are a small share of what the camera saw.
LexData takes the station 4 model through its whole life. You type what to look for, Lexi puts a box on every frame, and a person checks each label before anything trains on it. The model then watches the camera on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime. The new version travels to the cabinet as weights, in the other direction from the doubted frames, and the packaging line does not stop for it.
Two sensors on the same moment is where the pairing breaks
The plant's second camera on the line is a thermal one over the sealing head at station 5, and the question the maintenance team wants answered is what the thermal and visual streams say about the same carton. That is the multi-sensor monitoring use case, and the value is in the correspondence: a box in one stream matched to the same carton in the other, which depends on knowing which visual frame belongs with which thermal frame.
That pairing is fragile in a way neither picture shows. A firmware update to one camera shifts its timestamps, or a compression setting changes on one recorder, and both feeds look perfect while the fusion pairs the wrong frames. On the plant network this is an IT change, pushed overnight, that nobody associated with a vision model. The runner's own record of what it received from each stream, and when, is what turns a morning of wrong pairings into a diagnosis rather than a retrain.
The controls engineer added the runner's events to the naming convention in a week. The frames never got a tag, because they never needed one.
See it on your own footage.
Start with your footageMore in Edge

Edge · 6 min read
Computer vision on multiple video streams from one runner
Twenty cameras, one box beside the recorder. Fair sampling, a slow camera that drops its own frames, an unplugged one that stalls nobody, one heartbeat each.
Andreas Ohrvall · Oct 1, 2026

Edge · 7 min read
CPU vs GPU for computer vision inference, and when a CPU is enough
A cap station checked every couple of seconds and a shelf camera sampled every few minutes both run on a CPU. The GPU earns its keep when the load piles up.
Andreas Ohrvall · Oct 1, 2026

Edge · 6 min read
Deploying computer vision models to edge devices beside the camera
A robot cell that cannot wait for a round trip and an orchard with no uplink. Export by device, run beside the recorder, and get the next version out there.
Andreas Ohrvall · Oct 1, 2026