Skip to content
LexDataLexData
PlatformIndustriesCustomers
DocsThe Field GuideBlogWhy models drift
AboutCareersSecurityContact
Log inStart now
← All posts

Industries · 6 min read

Retail shelf item detection, a hundred products in one frame

A wide aisle camera sees a hundred boxes per frame, most of them identical neighbours. The rule is a tight box per facing, and the rare SKU stays in on purpose.

Summary

This post covers detecting a hundred products in one frame from a wide shelf camera: dense, small and mostly identical objects, a tight box per facing as the labeling rule, the far end of the shelf where the small objects live, and rare SKUs kept in the set on purpose. It concludes that a new packaging run is the drift the footage cannot show and the override rate can. It is for retail technology and store operations teams.

Sheikh Srijon · GTM Lead · Sep 29, 2026

Bread shelf with rack sections boxed and an empty slot flagged, from a customer store camera

The overnight fill on the cereal aisle finishes at 5 am and the wide camera at the end of the aisle gets its first clean look at the shelf. Six shelves high, three bays wide, and somewhere over a hundred boxes in the frame, most of them the same brand in three flavours whose packs differ by the colour of one band. At the far end of the bay, forty feet from the lens, each box is a smudge a few dozen pixels wide. A person walking the aisle sees a full shelf. The model has to see a hundred separate things.

Object detection on a shelf is dense, small and full of identical neighbours

Most detection work has a few objects per frame with space between them. The cereal shelf at 5 am has none of that. The boxes touch, they repeat, they sit at an angle to the camera that changes along the bay, and the far half of the frame is small-object territory where a general detector trained on ordinary scenes finds almost nothing.

Object detection here has to be tuned for density from the first label. That means training on frames from this camera at this distance rather than on a general set, and it means the labeling rule for the boxes matters more than the choice of architecture. A model learns what its boxes taught it, and on a shelf the boxes are the whole lesson.

A tight box per facing is the labeling rule that matters

The rule that decides whether the model works is how a person draws a box between two identical neighbours. One box per facing, hugging the front face of that one pack, ending where the next pack begins even though the two are the same colour and there is no visible gap. A box that spans two packs teaches the model that a facing is two packs wide, and the count on every shelf is then wrong by half.

You type "cereal box" once, Lexi proposes a box on every pack in every frame, and a person checks them. The checking on a shelf is mostly the seams: a box that slid onto the neighbour, a pair that merged, a pack turned sideways that got no box at all. The labeling guide covers why the ruling on each of these gets written down, because the next store's shelf will ask the same question.

My own view is that one box per facing is the only rule worth having, even when the run of identical packs makes it tedious. A box per run is faster to draw and useless for every question a store actually asks.

The far end of the shelf is where the small objects live

On the 5 am frame the pack nearest the camera is a few hundred pixels wide and the one at the far end of the bay is a few dozen. A frame resized down to a standard training size erases the far end entirely, and the model that results is confident near the lens and blind at the back. The fix is capture and tiling rather than weights: label at full resolution, cut the frame into tiles that keep the far packs at their native size, and run detection per tile.

Where the camera can be repositioned, one bay per camera beats one aisle per camera every time. Where it cannot, the far end of the frame is the section that comes back for review most, and that is the honest shape of the problem rather than a defect in the model.

Rare SKUs stay in the set on purpose

A shelf with a hundred packs of the leading brand and two packs of the one that sells a case a month produces a training set in the same proportion. The model then learns that the rare pack does not exist. The store cares about the rare pack more, because its facing is the one that goes empty for a week without anyone noticing.

So the rare SKUs are kept in the set deliberately, oversampled from the frames where they appear, and their boxes checked with more care than the leading brand's. The shelf and planogram compliance use case makes the point that the class count is enormous and near-identical, and that distinguishing two flavours of the same brand at shelf angle is the hard part. The rare one is the hardest, and it is the one a person has to make sure is in.

An aside from the stockroom: the planogram printed at head office is taped inside the stockroom door, and it is the reference the night fill crew works from, which makes it the reference the labels should agree with.

A new packaging run is the drift the footage cannot show

The frames on the Monday morning after a packaging refresh look like any other morning. The shelf is full, the light is the same, the camera has not moved. The leading brand has a new box, correct on the shelf and unknown to the model, and every facing of it reads as unrecognised. The drift catalog calls this a spec change: the pixels did not drift, the right answer did.

The signal is the staff overriding the alerts on that aisle in a consistent direction from the first morning. That override rate is what tells the model it needs the new pack, and the corrections at review retrain it. The refresh date was on the supplier's calendar for a quarter, and the cheapest version of this is a window of frames labeled the week the new box arrives in the stockroom rather than the week it reaches the shelf.

A person checks the labels before the model sees a shelf

LexData takes the shelf model through its whole life. You type what to look for, Lexi puts a box on every frame, and a person checks each label before anything trains on it. The model then watches the aisle cameras the store already has, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime. The retail work behind our figure, 4M+ annotations and validations, is mostly shelves like this one, boxed a facing at a time and checked by someone who knew which flavour was which.

The night fill crew finishes at 5 am. The camera looks at a hundred boxes. By the time the store opens at 8 am the two empty facings at the far end of the bay are on the morning list, and the pack that sells a case a month is one of them.

See it on your own footage.

Start with your footage

More in Industries

Industries · 7 min read

Aerial fire detection from a drone patrol, smoke before the flame reaches the line

On a right-of-way patrol, smoke is a few dozen grey pixels that look like haze. Two boxed classes, an alert with a cooldown, and the bad-weather days kept.

Rob Hickey · Sep 29, 2026

Industries · 6 min read

AI in robotics after the robot ships, what the warehouse cameras keep learning

The forward camera boxed pallets and people well at the pilot site. Then the racking moved, and the edge cases the planner never saw came back for review.

Andreas Ohrvall · Sep 29, 2026

Industries · 6 min read

Automated sorting with computer vision, from the camera over the conveyor to the diverter

A box on every apple, a grade from the box, and an air jet that acts on it before the belt moves on. The new cultivar is when the model needs the graders again.

Rob Hickey · Sep 29, 2026

LexData
LexData

Product

  • Platform
  • Industries
  • Use cases

Resources

  • Docs
  • The Field Guide
  • Blog
  • Why models drift

Industries

  • Energy & utilities
  • Oil & gas
  • Agriculture
  • Manufacturing
  • Insurance
  • Retail
  • Robotics

Company

  • About
  • Customers
  • Careers
  • Contact

Trust

  • Security
  • Privacy
  • Terms

Stay updated

What we learn running vision models in production.

See everything.
Miss nothing.

Stay updated

What we learn running vision models in production.

Terms of use & Privacy policy

© 2026 LexData Labs · All rights reserved