Skip to content
LexDataLexData
PlatformIndustriesCustomers
DocsThe Field GuideBlogWhy models drift
AboutCareersSecurityContact
Log inStart now
← All posts

Operations · 6 min read

How to reduce class flickering in object detection on video

A shelf facing that reads empty, stocked, empty every few frames is a per-frame model at a boundary. Vote over the track, then review the frames that flickered.

Summary

This post takes a store shelf camera whose label on one facing flips between empty and stocked every few frames and explains why a per-frame detector does that near a decision boundary. It works through voting over a window of sampled frames on the tracked object, and argues that the frames which flickered are the ones to send back for review. It is for anyone whose alerts fire on a label that will not hold still.

Rob Hickey · Chief AI Officer · Oct 1, 2026

Shopper at a dairy case with the basket and rack sections boxed, from a customer store camera

The camera over the dairy case at a supermarket sends an alert when a facing has been empty for ten minutes. On Thursday afternoon the yoghurt facing at the end of the second shelf sent eleven alerts in a minute, and the store manager, who had been fine with the alerts until then, asked for them to be turned off. On the recording the facing is half empty. Two pots remain at the back, a shopper's arm passes across it every so often, and the model's label on the facing reads empty, stocked, empty, empty, stocked, empty, at the rate the frames are sampled.

The model was not wrong on any single frame. It was undecided, and a per-frame detector has no way to stay undecided, so it picked a side on each frame and picked differently.

A per-frame model has no memory of the last frame

A detector looks at one frame and produces boxes, each with a class and a score. It does not know what it said two seconds ago about the same facing. When the yoghurt facing is clearly full or clearly empty the label holds on every frame because the evidence is strong. When two pots at the back of a deep shelf sit right at the edge of what the model has learned to call stocked, the evidence is weak. Small changes between frames, a reflection off the door, the shopper's sleeve, a half-pixel shift in the box, push the answer over the line and back, and on Thursday they did so all afternoon.

The flicker is a symptom of the boundary. It says the facing in front of the camera is a case the training set covered thinly, which is useful to know, and which is separate from making the alert stop.

The learn section has a lesson on what the number beside a detection actually means. The short version is that a score near the middle is a ranking signal that the model is at a boundary, which is what the yoghurt facing produced all afternoon.

The tracked bounding box is what the vote is taken over

The fix starts with identity. A bounding box on its own is a fact about one frame; a tracked box is the same object followed across frames, and it is the object that a store cares about, not any one frame of it. Instance tracking gives the yoghurt facing in aisle 4 an identity that persists through the shopper's arm and the reflection, and every per-frame label the detector produces is attached to that identity rather than left loose.

Once the labels are attached to a track, the class the track carries can be decided by a vote over a window of recent sampled frames instead of by the latest one. Six samples of empty and two of stocked over the last window is empty. The two stocked frames stay in the record, and the alert rule never sees them.

Voting over a window trades a few seconds for a steady label

A window is a choice about time. Sampling about every two seconds, a window of a dozen samples is under half a minute, and a facing that empties for real is called empty a few seconds later than a per-frame model would have called it. On the dairy case that delay is nothing. On a hand entering a robot cell at 3 pm it is everything, which is why the window is set per rule rather than once for a site.

Two refinements make the vote behave. A change of class should need a run of consecutive frames agreeing, so a single stray frame in the other direction does not start a switch. And the vote should be weighted toward the recent end of the window, so a facing that was full ten samples ago and empty for the last five reads as empty now rather than as a tie.

The store manager's own rule, when she restocks by eye, is that a facing is empty when it has been empty since she last walked past. That is a window with a human sample rate.

The flickering frames are the ones to send back for review

The vote makes the label hold still. It does not make the model better, and the frames that flickered are the most valuable frames the camera produced that afternoon, because each one is a case the model could not decide. Those are the frames that should come back to a person.

On our platform that is what the review queue is for. Frames the model is unsure of come back to a person, on Thursday afternoon or the morning after, and a facing that a person marks as stocked with two pots at the back is a label that the next version learns from. When corrections cross the project's threshold, a new version is trained on them, and the boundary that produced the flicker moves. My own view is that a team which only smooths the flicker, and never reviews the frames behind it, has hidden the one signal that tells them where the training set is thin.

Asking a question about the footage afterwards changes nothing in the model. The corrections at review are what do.

The alert rule reads the voted label

An alert is a rule written as a sentence, with a severity and a cooldown, approved before it goes live, and delivered to Slack, email or a webhook. "A facing is empty for ten minutes" is the sentence for the dairy case, and the ten minutes is itself a window. The rule reads the voted label, and it reads it for a duration, so a flicker that gets through the vote still cannot fire an alert on its own. The cooldown stops a facing that is restocked and emptied again in the same hour from paging twice. The monitoring and alerts guide covers writing the rule and deciding where it lands.

LexData takes the shelf model through its whole life. You type what to look for, Lexi puts a box on every facing in every frame, and a person checks each label before anything trains on it. The model then watches the aisle cameras the store already has, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime.

A window that is too long hides a real change

The mistake in the other direction is to lengthen the window until nothing flickers. A window that spans minutes will hold "stocked" on a facing that a shopper cleared in one go, and the alert that was meant to land in ten minutes lands in fifteen. Set the window to the shortest span that stops the flicker on the recorded afternoon, and review the frames that still flip.

The two yoghurt pots at the back of the shelf on Thursday were, according to the store manager, the ones nobody buys because they are the flavour that comes free in the multipack.

See it on your own footage.

Start with your footage

More in Operations

Operations · 7 min read

AGPL-3.0 licensing risk for computer vision teams serving a model

A camera streaming to a served detector is the network interaction the licence was written for. What a legal review will ask, and why to pick the weights first.

Ayman Quadir · Oct 1, 2026

Operations · 7 min read

Cloud vs owned GPU inference for computer vision, worked out per camera hour

A plant on three shifts and a retailer with cameras spread across stores get different answers from one sum, and footage leaving the building is a cost too.

Ayman Quadir · Oct 1, 2026

Operations · 7 min read

Computer vision heatmaps drawn from the aisle cameras a store already has

Footpoints from every tracked box, aggregated over a day and mapped onto the floor plan, show where footfall goes. A camera nudged in cleaning shifts the map.

Rajiya Sultana · Oct 1, 2026

LexData
LexData

Product

  • Platform
  • Industries
  • Use cases

Resources

  • Docs
  • The Field Guide
  • Blog
  • Why models drift

Industries

  • Energy & utilities
  • Oil & gas
  • Agriculture
  • Manufacturing
  • Insurance
  • Retail
  • Robotics

Company

  • About
  • Customers
  • Careers
  • Contact

Trust

  • Security
  • Privacy
  • Terms

Stay updated

What we learn running vision models in production.

See everything.
Miss nothing.

Stay updated

What we learn running vision models in production.

Terms of use & Privacy policy

© 2026 LexData Labs · All rights reserved