Skip to content
LexDataLexData
PlatformIndustriesCustomers
DocsThe Field GuideBlogWhy models drift
AboutCareersSecurityContact
Log inStart now
← All posts

Industries · 7 min read

Retail customer movement monitoring from the cameras already on the ceiling

People boxed and tracked, zones drawn as polygons, dwell and occupancy per zone, and the camera nudged during a Sunday reset as what moves every zone at once.

Summary

This post takes one ceiling camera over a grocery department and follows a shopper from the aisle end to the dairy case, with people boxed and tracked, zones drawn once as polygons, and dwell and occupancy per zone answered in LexInsight. It concludes that the model has to be tuned on the store's own frames, that the count is the track rather than the frame, and that a camera bumped during a Sunday reset moves every zone at once. It is written for store operations and retail analytics teams.

Ayman Quadir · Head of Product · Sep 24, 2026

Shopper at a dairy case with carts boxed, from a customer store camera

The dairy case at the back of the store has the highest-margin promotion in the building on its end cap, and the category manager has no idea how many people stop at it. The camera above the aisle has been there since the store opened, recording for security and watched by nobody. On a Saturday in June it records a shopper with a cart pausing at the end cap, walking on, coming back, standing for ninety seconds and leaving without the product. That is the sequence the category manager would pay to know about, and it is on a hard drive in the back office.

The camera does not need replacing. The frames need reading.

People are tracked and the count is the track

A frame from the aisle camera at 11 am on the Saturday holds two shoppers and a cart. Boxing the people on every frame is the first job, and it is not the count. The same shopper appears in hundreds of frames as she moves along the case, and a count of boxes is a count of frames rather than of people. The traffic counting and conversion use case makes the point that an identity has to persist across the frames or the number is a function of the frame rate.

So each person is boxed and then tracked as the same person from frame to frame, and the track is the unit everything else is built on. A track that enters the department, pauses at the end cap and leaves is one visit. Two tracks that turn out to be the same shopper, because she stepped behind a display and came out the other side, are the error the labeling pass spends most of its time on.

Object detection on the store's own frames, from the ceiling

The people detector that ships in every general model was trained on people seen from the front, at street level. The aisle camera sees the tops of heads, shoulders and the handle of a cart, from four metres up, under fluorescent light. A general model finds some of those people and misses the child, the shopper bending into the case and the staff member crouched at the bottom shelf.

That is why object detection here starts from the store's own frames. The class is typed once, person, and Lexi proposes a box on every frame from the aisle camera. A person checks them, and their attention goes to the crouched and the half-hidden, because a model that only finds upright shoppers in the aisle centre reports a quiet department that was full. The retail work behind our numbers, 4M+ annotations and validations, is mostly frames like these, ceiling views of stores where the cameras were already up.

Zones drawn as polygons turn one camera into many questions

On the frame, someone draws the zones once: the end cap, the run of the dairy case, the aisle itself, the till approach at the far end. Each zone is a polygon in the frame, and each track's entry into and exit from each polygon is stamped with the time. From those stamps come the numbers the category manager asked for. How many tracks entered the end cap zone on Saturday. How long each one stayed. How many were in the dairy run at 5 pm, when the case was being restocked and the aisle was half blocked by a cage.

The polygon has to be drawn in the store, on the frame, by someone who knows where the end cap actually ends. A zone that cuts through the middle of the fixture counts the shopper reading the top shelf as having left.

Dwell and occupancy per zone arrive with their frames

Dwell time is the number retailers ask for, and my view is that it should never be reported without the frames behind it. Ninety seconds at the end cap could be a shopper comparing two prices or a shopper waiting for her husband; the average hides both, and the frames tell them apart. LexInsight answers "how long did people stay at the end cap on Saturday" with the tracks and the frames, so the category manager can watch the ninety-second visit that ended without a purchase and see that the shelf-edge label was missing.

Occupancy per zone is the number stores should ask for more often. The dairy run holding six people at 5 pm with a restock cage in it is a queue forming at a case with no till, and that is a staffing finding rather than a merchandising one. The wider retail picture, from queue length to open doors, comes from the same tracks and the same polygons on other cameras.

Staff, trolleys and the man at the aisle end

Not every track is a shopper. The staff member restocking the case is in the dairy zone for forty minutes, and the trolley collector passes through six times an hour. On the Saturday in question a man in a hi-vis vest stands at the aisle end for most of the afternoon, because he is the store's security contractor. Left in the count, they make the department look busier than it is and the dwell longer.

The fix is a label. Staff by uniform, contractor by vest, trolley collector by the train of carts, each its own class checked by a person on the training frames, and each excluded from the shopper count while still visible in the frames. Nothing about the count is trustworthy until that filtering is done, and it is not visible in any single frame.

The camera nudged during a Sunday reset moves every zone

The seasonal reset happens on a Sunday night. The fixtures move, the end cap changes, and the person on the ladder hanging the new signage bumps the camera housing with an elbow. On Monday the aisle looks fine on the monitor, because a person recognises the aisle from any angle. The polygons do not move with the camera. The end cap zone now covers half the end cap and a strip of floor, the dairy run zone starts a metre late, and every dwell number from Monday is measured against the wrong shapes.

The drift catalog calls this a camera moved, and the symptom is a step on a maintenance date: the end cap count drops on a Monday when footfall did not. The fix is the cheapest on the list. Redraw the zones from the new framing, re-label a short window of frames, and retrain on footage the store already has.

LexData takes the department's model through its whole life. You type what to look for, Lexi puts a box on every person in every frame, and a person checks each label before anything trains on it. The model then watches the aisle camera, in the cloud, on your servers or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime. The crouched shoppers from the first week are the second version's evidence.

No faces leave the store, only a count

The tracks carry no identity. A box on a head from four metres up is not a face, the record is a count per zone per hour, and nothing about the June Saturday survives except that a shopper stood at the end cap for ninety seconds and left. With a runner beside the store's recorder the footage stays on site and only the frames the model doubts leave the building, which is the arrangement the store's privacy policy already describes for its security recording.

See it on your own footage.

Start with your footage

More in Industries

Industries · 6 min read

AI visual inspection as the nondestructive testing step a camera can take over

Visual testing is the first NDT gate, its acceptance criteria are already written, and a camera can apply them to every weld instead of one in twenty.

Rob Hickey · Sep 24, 2026

Industries · 6 min read

Appearance inspection systems that judge scratches, chips and burrs the same way on every shift

The station, the light and the written standard matter more than the model. The outlines carry the limit, and a tightened tolerance makes every label wrong.

Rajiya Sultana · Sep 24, 2026

Industries · 7 min read

Automated pallet accounting from the camera over the staging zone

A polygon on the frame, every pallet tracked so it is counted once, entries and exits as the ledger, and a wash-down that nudges the camera as the failure.

Andreas Ohrvall · Sep 24, 2026

LexData
LexData

Product

  • Platform
  • Industries
  • Use cases

Resources

  • Docs
  • The Field Guide
  • Blog
  • Why models drift

Industries

  • Energy & utilities
  • Oil & gas
  • Agriculture
  • Manufacturing
  • Insurance
  • Retail
  • Robotics

Company

  • About
  • Customers
  • Careers
  • Contact

Trust

  • Security
  • Privacy
  • Terms

Stay updated

What we learn running vision models in production.

See everything.
Miss nothing.

Stay updated

What we learn running vision models in production.

Terms of use & Privacy policy

© 2026 LexData Labs · All rights reserved