Edge · 6 min read
Reducing computer vision inference costs when a store watches every aisle all day
The bill for watching a camera is set by how often you look, what you decode, and where the model runs. Most of a store's frames need no model at all.
Summary
This post works through where the cost of watching a store's aisle cameras goes, from the frames decoded to the frames a model sees to where the model runs. It concludes that sampling about every two seconds, skipping frames that have not changed, and running on a runner beside the recorder remove most of the bill without losing the answer. It is for retail technology and operations teams pricing a rollout beyond the pilot store.
Rajiya Sultana · Engineering Manager · Sep 30, 2026

A store aisle from a fixed camera, garment stacks boxed, from a customer store camera
The pilot store ran the shelf-gap model on aisle 4 for a month and the number was good. Then somebody multiplied the monthly bill by every aisle in the store and every store in the estate, and the rollout meeting became a budget meeting. The model had not changed. The cost of watching had been priced for one camera in one aisle, and it does not scale the way the store does.
Most of that bill is paid for frames nobody needed.
The bill is for hours watched, so count the hours first
A camera in aisle 4 records all day and all night whether or not anything is happening. The store is open from 7 am to 10 pm; the night stocking crew works two of the remaining hours; for the rest the aisle is dark and still. Watching the dark hours costs the same as watching the busy ones, and answers nothing.
So the first cut is the schedule. Which cameras are watched, during which hours, for which question. The shelf-gap question matters between the morning fill and closing. The queue question at the tills matters at lunch and after 5 pm. A camera on the loading bay matters when a delivery is due. Written as a schedule per camera, the hours watched fall a long way before any engineering starts, and the pricing conversation starts from that number rather than from a camera count.
Sampling about every two seconds is enough for an aisle
The camera produces far more frames each second than any aisle question needs. A shelf gap that appears at 11:04 and is still there at 11:06 is the same gap, and reading every frame in between finds it again and again at full price.
The platform samples each stream about every two seconds, one consumer per camera, every camera in parallel. For a gap on a shelf, a queue forming at checkout 3, a trolley left in the fire lane, two seconds is faster than any person who will act on the alert. The exceptions are questions that happen between samples, a hand reaching across a belt, and those are not questions for a store's aisle cameras.
Sampling is the single largest saving on the list and it needs no model change. The frames not sampled are not decoded, not sent anywhere, and not paid for.
A frame that looks like the last one needs no model
Between samples the aisle is often unchanged. At 3 pm on a Wednesday aisle 4 can go ten minutes without a person in it, and the sampled frames are the same picture ten minutes apart. A cheap comparison against the previous sample, a difference over a threshold in the region the rule cares about, decides whether the model runs at all.
This gating is where the night stocking crew earns a mention. They restock the aisle at 4 am, every shelf changes, and every sampled frame goes to the model for two hours, which is right: the shelf state at 6 am is what the day starts from. Then the aisle settles and the gate closes again.
The rule for the gate has to be set per camera and checked on the review queue. A gate that is too tight passes frames that changed only in the light; a gate that is too loose swallows a shopper taking the last item.
When a gap is missed, the frame that should have run and did not is the one to look at first.
Decoding the stream costs before the model does
The stream from the recorder is compressed. Turning it back into a frame the model can read is work, and on a small box in the cabinet, a Jetson or something like it, it is often more work than the model. A store that plans its budget around the model and forgets decode ends up with a box that is busy before it has looked at anything.
My own view is that the cheapest frame is the one nobody decodes, and the decode budget should be planned before the model budget. Sampling and gating both help here because a frame that is skipped is skipped before decode. The deployment doc says the same thing from the other side: on edge hardware, decode rather than the network is often the real bottleneck, and the enclosure throttles under sustained load.
A runner beside the recorder keeps the footage and the bill on site
Sending every sampled frame from every aisle camera to the cloud means paying to move it and paying to watch it there. A runner beside the recorder, a small box in the same cabinet, watches the streams locally. The footage stays in the store, the model runs where the frames already are, and only the frames it is unsure of leave, to be reviewed.
LexData takes the shelf model through its whole life. You type what to look for, Lexi puts a box on every frame, and a person checks each label before anything trains on it. The model then watches the aisle cameras the store already has, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime. With the runner, the alert for the gap on aisle 4 fires locally first, and nothing about the store's footage has crossed the firewall.
The model itself can be sized for the box. Export it with the device preset for the hardware in the cabinet, and the version that runs on the runner is the one trained on the store's own frames.
The review queue is the cost that is worth paying
There is one cost this post has not tried to cut. The frames the model is unsure of go to a person, and that person's time is the most valuable hour in the whole system. A model that sends fewer doubted frames because it was made more confident by fiat saves a reviewer's afternoon and loses the corrections that would have made the next version better.
So the doubted frames are the budget line that stays. The 4M+ annotations and validations behind our retail work were paid for one reviewed frame at a time. The estate that scales the shelf model to its fiftieth store does it by spending on those frames, and on nothing that is dark, still, or the same as two seconds ago.
See it on your own footage.
Start with your footageMore in Edge

Edge · 6 min read
Edge AI fleet management for fifty runners across five plants
Fifty small boxes beside fifty recorders, one model version, and the IT patch that changes what every camera sends. What a fleet needs before the second plant.
Andreas Ohrvall · Sep 30, 2026

Edge · 6 min read
What computer vision model deployment means after the first week
Deploying a vision model is a decision about where it runs, how a new version replaces the old, and what tells you it has started to be wrong.
Andreas Ohrvall · Sep 30, 2026

Edge · 8 min read
AI cameras vs IP cameras, and what changes when the model moves to the edge
The dock camera has streamed to a recorder for six years. A model can watch that stream in the cloud, on your servers or beside the recorder. No new camera.
Andreas Ohrvall · Sep 27, 2026