Edge · 6 min read
RTSP stream processing for computer vision, and the three ways a camera feed fails without an error
A feed that falls behind, a reboot that ends the frames with no error, a stream with the wrong clock. Sample every two seconds and give each camera a pulse.
Summary
This post explains how an IP camera's RTSP feed fails quietly, from a consumer that falls behind because it tries to process every frame, to a reboot that ends the frames with no exception, to a reconnected stream with the wrong clock. It concludes that sampling about every two seconds, one consumer per stream, and a heartbeat per camera in LexInsight are what make a feed something a person can trust. It is for the engineers attaching models to cameras that already exist.
Esdras Ntuyenabo · Engineer · Oct 1, 2026

A camera on a pole over an industrial yard at dusk, generated scene with detections from our model
The camera on the pole over the yard at the depot has been sending its RTSP stream to the recorder for four years without anyone thinking about it, because the recorder never complained. Then a model was attached to the same stream, and within a week it had been wrong three times in three different ways, and none of them produced an error. The frames were there, or looked like they were, and the alerts were quiet.
A stream is not a file. It keeps coming whether or not anyone is keeping up, and what it does when nobody is keeping up is the first failure.
A consumer that reads every frame falls behind and never catches up
The camera on the pole produces frames all day. The first version of the pipeline read each one and ran the model on it, and the model took longer per frame than the camera took to produce the next. The queue grew. By 11 am the model was looking at frames from 10:52, by lunch it was twenty minutes behind, and the alert for a vehicle at the gate arrived when the vehicle had long gone. No exception, no log line, just a stream that was always slightly in the past.
The fix is to stop trying. The platform samples each stream about every two seconds, one consumer per camera, every camera in parallel, and the frames between samples are never decoded. A vehicle at the gate is still at the gate two seconds later. The consumer is always looking at now, because it never has a backlog to look at instead.
The yard's original recorder, incidentally, never had this problem because it does nothing to the frames but write them. A recorder is a consumer that can keep up with anything.
A reboot ends the frames with no exception
At 2 am on a Sunday the power in the yard blipped and the camera on the pole rebooted. It came back in a minute. The stream it had been sending did not: the connection the consumer held was dead, and the consumer sat on the read call waiting for the next frame from a socket that would never deliver one. No exception, because nothing was wrong from the socket's point of view. The model was still running. It just had no frames, and the last frame it had seen was the one before the blip.
The monitor showed a picture. The picture was from 2 am.
Every consumer needs a watchdog: a timer that notices no frame has arrived in longer than the sampling interval allows, tears the connection down, and reconnects. That is the whole of it, and the pipelines that fail here are the ones written as if the stream were a file with an end. The deployment doc makes the related point about hardware, that decode rather than the network is often the real bottleneck on a runner, and the watchdog is what stops a blocked decode from becoming a stalled camera nobody notices.
A reconnected stream can carry the wrong time
The third failure is the quietest. The camera on the pole came back after the Sunday reboot with its clock reset, because the depot's time server was on the other side of the same blip. The frames in the reconnected stream carried timestamps from years ago. The pipeline that stamped each event with the frame's own time recorded a vehicle at the gate on a date nobody could find in the recorder, and the alert went to a channel where it was read as a glitch.
The fix is to stamp events with the consumer's clock at the moment the frame was sampled, and to treat the frame's own timestamp as a claim the camera is making rather than a fact. A stream whose timestamps jump is a stream that reconnected, and that is worth knowing on its own.
My own view is that the camera's timestamp should never be the primary time on any event, on any site, because the camera is the least trustworthy clock on the network and the only one that reboots in the yard.
Every camera gets a heartbeat a person can see
Three silent failures mean the pipeline needs a signal of its own health that does not depend on the model finding anything. For each stream: the time of the last sampled frame, the time of the last reconnect, and how far behind the consumer is, if at all. That heartbeat is the number LexInsight shows per camera, so the question "is the pole camera still being watched" has an answer that is separate from "has it seen a vehicle".
The alert on the heartbeat is the one to write first. A camera with no sampled frame for longer than the window, sent to maintenance, routine, with a cooldown, and approved before it goes live like any other rule. The monitoring and alerts doc covers where alerts land and how to route by severity, and the heartbeat alert belongs on the routine path, because a stalled camera is not an emergency until it is the camera that mattered.
LexData takes the yard model through its whole life. You type what to look for, Lexi puts a box on every frame, and a person checks each label before anything trains on it. The model then watches the pole camera, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime. None of that helps a model that has not seen a new frame since 2 am, which is why the heartbeat sits beside the review queue rather than behind it.
Test the pipeline on a stream you can break on purpose
None of the three failures shows up on a test with a video file. The file has an end, its timestamps are in order, and reading it never blocks. Before the yard camera goes live, the pipeline should be pointed at a stream that can be interrupted from a terminal: stopped mid-frame, restarted with a different clock, slowed until the consumer would fall behind if it were going to.
A stream server on a laptop, fed a recording of the yard, does the job. The pipeline that survives an afternoon of being unplugged on purpose is the one that survives Sunday at 2 am.
See it on your own footage.
Start with your footageMore in Edge

Edge · 6 min read
Computer vision on multiple video streams from one runner
Twenty cameras, one box beside the recorder. Fair sampling, a slow camera that drops its own frames, an unplugged one that stalls nobody, one heartbeat each.
Andreas Ohrvall · Oct 1, 2026

Edge · 7 min read
CPU vs GPU for computer vision inference, and when a CPU is enough
A cap station checked every couple of seconds and a shelf camera sampled every few minutes both run on a CPU. The GPU earns its keep when the load piles up.
Andreas Ohrvall · Oct 1, 2026

Edge · 6 min read
Deploying computer vision models to edge devices beside the camera
A robot cell that cannot wait for a round trip and an orchard with no uplink. Export by device, run beside the recorder, and get the next version out there.
Andreas Ohrvall · Oct 1, 2026