Skip to content
LexDataLexData
PlatformIndustriesCustomers
DocsThe Field GuideBlogWhy models drift
AboutCareersSecurityContact
Log inStart now
← All posts

Edge · 6 min read

Edge AI fleet management for fifty runners across five plants

Fifty small boxes beside fifty recorders, one model version, and the IT patch that changes what every camera sends. What a fleet needs before the second plant.

Summary

This post describes running a fleet of edge runners beside the recorders in five plants, from provisioning a new box to a health signal per runner, one model version rolled out and rolled back with no downtime, and the overnight patch that changes what the cameras send. It concludes that a step change across part of the fleet is a pipeline event to rule out before anyone retrains. It is for the engineers who own vision at more than one site.

Andreas Ohrvall · CTO · Sep 30, 2026

A robot cell on a plant floor, the kind of line a runner in the cabinet watches, generated scene with detections from our model

Plant 1 has had a runner beside its recorder for a year, and the person who set it up can find it in the cabinet with their eyes closed. Plant 2 through plant 5 came online over the spring, ten runners each, and nobody can picture all fifty. One is in a cabinet with a fan heater in the winter. One is on a shelf in a server room that is really a cupboard. The model version on each of them was correct on the day it was installed, and that day was different for every box.

A fleet is what a deployment becomes when the person who did it cannot visit every one.

A new runner registers itself before anyone logs in

The runner for line 3 at plant 4 arrives in a box. A technician mounts it in the cabinet, plugs it into the recorder's switch and the power, and walks away. That is the whole on-site procedure, and it has to be, because the technician is a maintenance electrician with a hundred other jobs.

Everything after that is the runner's own business: it comes up, reaches the platform, identifies itself, receives the model version the fleet is on and the streams it is meant to watch, and starts. Its footage stays in the cabinet. The frames it doubts are the only thing that leaves. My own view is that a runner that needs a login on site to be replaced is not fleet-ready, whatever else it can do, because the day it fails will be a Sunday and the person with the login will be elsewhere.

Each runner sends a health signal, and silence is a signal too

Fifty boxes are fifty things that can quietly stop. The runner in the hot cabinet throttles in July. The one in the cupboard loses its switch port when someone reorganises the patch panel. A recorder reboots at 2 am after a power blip and the stream comes back with a different clock.

So each runner sends a heartbeat: the time of the last frame it took from each stream, how busy the decode is, how warm the box is, which model version it is running. The deployment doc lists the three things to plan headroom for, sustained load in a sealed enclosure, a swapped camera, and decode as the real bottleneck, and the heartbeat is how the fleet finds out which of the three has arrived at which plant.

A runner that stops sending is an alert of its own, before anyone notices the shelf-gap alerts from plant 4 have gone quiet.

One model version reaches the fleet without stopping a line

A new version of the model has trained on the corrections from all five plants. Getting it to fifty runners is the operation that fleet management exists for, and the rule is that no line stops for it.

Each runner loads the new version beside the old, moves the next frame to the new one, and retires the old one when the new one is serving. Done fleet-wide, in the order the fleet chooses, plant 1 first because its people will notice anything wrong within an hour. Rolling back is the same motion in reverse, per runner or across the fleet, and a version that was rolled back stays on the runner as the one to return to.

LexData takes the model through its whole life. You type what to look for, Lexi puts a box on every frame, and a person checks each label before anything trains on it. The model then watches the cameras each plant already has, on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime. Versions keep what they were trained on, so when plant 3 asks why the model started flagging the new tote colour, the answer is a list of the frames that taught it.

The overnight IT patch is the step change to rule out first

On a Wednesday night the corporate IT team pushes a firmware update to every camera bought in the last two years, which is every camera at plant 4 and plant 5 and none at plant 1. White balance shifts. The bitrate drops. On Thursday morning the review queues at plant 4 and plant 5 fill up, plant 1 is unchanged, and the first instinct is that the model has drifted.

It has not. The drift catalog calls this a firmware or encoding change, and files it as a pipeline event rather than drift, because the world did not change and neither did the definitions. The signal is the shape: an abrupt cliff, to the hour, confined to the part of the fleet that got the patch. Drift from the world is a slope. A step means something discrete changed. The fleet's own records, which runner is on which version and which camera got which firmware, turn the step into a diagnosis in an hour rather than a week of retraining on frames that were never the problem.

The fix is often to reverse the patch, or to re-label a short window from the changed cameras if it has to stay. What is never the fix is retraining the whole fleet on plant 4's Thursday.

Replacing a runner is routine, so the fleet has to make it routine

The box in the hot cabinet at plant 2 dies in August. A spare goes in, registers, receives the fleet's current version and the three streams line 2 needs, and is watching again before the electrician has packed up. The dead runner's record stays: what it was running, when it last sent a heartbeat, which frames it doubted in its last week.

The fleet is the sum of those records. Fifty runners, five plants, one model version, and a page that shows, for every box, when it last saw a frame and which version it was running when it did. The person who set up plant 1 can still find that runner with their eyes closed, and no longer needs to.

See it on your own footage.

Start with your footage

More in Edge

Edge · 6 min read

What computer vision model deployment means after the first week

Deploying a vision model is a decision about where it runs, how a new version replaces the old, and what tells you it has started to be wrong.

Andreas Ohrvall · Sep 30, 2026

Edge · 6 min read

Reducing computer vision inference costs when a store watches every aisle all day

The bill for watching a camera is set by how often you look, what you decode, and where the model runs. Most of a store's frames need no model at all.

Rajiya Sultana · Sep 30, 2026

Edge · 8 min read

AI cameras vs IP cameras, and what changes when the model moves to the edge

The dock camera has streamed to a recorder for six years. A model can watch that stream in the cloud, on your servers or beside the recorder. No new camera.

Andreas Ohrvall · Sep 27, 2026

LexData
LexData

Product

  • Platform
  • Industries
  • Use cases

Resources

  • Docs
  • The Field Guide
  • Blog
  • Why models drift

Industries

  • Energy & utilities
  • Oil & gas
  • Agriculture
  • Manufacturing
  • Insurance
  • Retail
  • Robotics

Company

  • About
  • Customers
  • Careers
  • Contact

Trust

  • Security
  • Privacy
  • Terms

Stay updated

What we learn running vision models in production.

See everything.
Miss nothing.

Stay updated

What we learn running vision models in production.

Terms of use & Privacy policy

© 2026 LexData Labs · All rights reserved