Skip to content
LexDataLexData
PlatformIndustriesCustomers
DocsThe Field GuideBlogWhy models drift
AboutCareersSecurityContact
Log inStart now
← All posts

Computer vision · 6 min read

Why a good detector can still be a bad counter

The queue camera finds every shopper on every frame and still gets the count wrong. Swaps, lost tracks and jitter are tracker mistakes, and review fixes one.

Summary

This post separates detection from tracking using a checkout queue camera whose detector finds every person and whose count is still wrong. It walks through the three ways a tracker fails, identities swapped when two shoppers cross, a track lost behind a shelf end, and jitter that splits one person into three, and says which of them a correction in review actually fixes. It is for anyone counting people or things from a camera over time.

Stephen Biswas · Engineer · Sep 29, 2026

Checkout lane from a ceiling camera, a person and the products on the belt boxed, generated scene with detections from our model

The camera over checkout 3 has a detector that finds every person in every frame, and a person checking its boxes on a Saturday afternoon would find little to correct. The queue count it produces is still wrong. It says seven when there are four, then two when there are five, and the store's rule for opening another till fires and clears in the same minute.

The detector is doing its job. Counting is a different job, and the part of the pipeline that does it, the tracker, is the part nobody labeled and nobody is looking at.

A count needs identity and a detector has none

A detector answers one question per frame: where are the people. It has no memory. The box it draws on a shopper in one frame and the box it draws on the same shopper two seconds later are, to the detector, two unrelated boxes. Counting how many people waited, or how long each waited, needs a second thing that links boxes across frames into a track per person and gives each track a lifetime.

That linking is the tracker, and every wrong count at checkout 3 comes from a link the tracker made wrongly or failed to make. A frame-by-frame count of boxes, with no tracker at all, is right on any single frame and useless for "how long did the fourth person wait".

Two shoppers crossing swap identities

The first failure is the swap. Two shoppers pass each other at the end of the queue on a Saturday afternoon, one leaving with a basket and one joining. For a few frames their boxes overlap, and when they separate the tracker has to decide which box continues which track. It guesses from position and speed, sometimes from appearance, and when both are similar it guesses wrong. The leaving shopper's track now continues on the joining shopper, and a person who has been in the queue for a minute inherits a wait time of nine.

Swaps are invisible in the boxes. Every frame has the right number of boxes in the right places, and the detector's held-out score is untouched. They are visible only in the tracks, as a wait time that jumps, and finding them means someone watching tracks rather than frames.

A track lost behind the shelf end comes back as a new person

The second failure is the gap. A shopper steps behind the end of the confectionery shelf to reach a magazine, is hidden for a few seconds, and steps back. The tracker, which kept her track alive for a while and then gave up, starts a new one when she reappears. One person has become two, the count is one too high, and her wait time has reset to zero.

How long a tracker waits before giving up on a hidden person is a setting, and the right setting for checkout 3 depends on how long people stand behind that shelf. Too short and every magazine browser is counted twice. Too long and a shopper who genuinely left keeps a ghost in the queue for a minute after the till has served the person behind her.

The camera over checkout 3, as it happens, also sees the lottery stand, and the lottery stand is where the tracks go to die on a Saturday.

Jitter splits one person into three

The third failure comes from the detector after all, in a way the detector's score does not show. A box that wobbles from frame to frame, a little wider here, shifted a few pixels there, dropping out for one frame under the fluorescent flicker, is still a correct detection on every frame. To the tracker at checkout 3 it can look like a person who vanished and a new person who appeared, twice in ten seconds. One shopper becomes three short tracks, and each of the three is counted.

This is the one failure the detector's training can fix, because steadier boxes mean fewer breaks, and steadier boxes come from more frames of that camera in the training set, checked by a person.

An R-CNN finds people per frame and the tracker links them

The usual pipeline is tracking by detection: a detector, single-pass or a two-stage R-CNN style model, draws the boxes on each frame, and a separate tracker links them. The detector is trained on labeled frames; the tracker is mostly a set of rules and settings about motion and how long to wait. That split is what makes the review question sharp. A correction a person makes in review is a correction to a box on a frame. It retrains the detector. It does not touch the tracker's settings.

So of the three failures, review fixes the jitter, because jitter is a box problem. A swap and a lost track are not box problems, and no number of corrected frames changes how long the tracker waits behind the confectionery shelf. Those are tuned by watching tracks on that camera and adjusting the settings until the count matches what a person sees, and then leaving them alone.

My own view is that a queue rule should be written around a line the shoppers cross rather than around the tracks themselves. Counting crossings of a line at the end of the queue survives most swaps and gaps, because a swap on the far side of the line does not change how many crossed it. The checkout queue monitoring use case is built that way for the same reason.

The count that matters is the one a person would agree with

The lesson on what a vision model can and cannot see describes "how many people crossed this line today" as a tracking problem wearing a detection costume, and the queue at checkout 3 is exactly that. The detector can be excellent and the costume can still be wrong.

LexData takes the queue model through its whole life. You type what to look for, Lexi puts a box on every frame, and a person checks each label before anything trains on it. The model then watches the cameras the store already has, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime. The boxes get steadier with every version, and the jitter that split one shopper into three on a Saturday becomes one track, which is the one part of the count that review was ever going to fix.

See it on your own footage.

Start with your footage

More in Computer vision

Computer vision · 6 min read

What the big vision models are for, and what runs on the camera

The big models find the frames, draw the first boxes and tighten them, and a person checks. The small model trained on those labels runs beside the recorder.

Andreas Ohrvall · Sep 29, 2026

Computer vision · 6 min read

What semantic segmentation labels and what it costs

Drivable path, corrosion area and crop rows are questions about pixels, and the answer is a class for every one. Painted labels cost more than boxes.

Esdras Ntuyenabo · Sep 29, 2026

Computer vision · 6 min read

What YOLO does in one pass and what that costs

One look at the frame is why a single-pass detector fits the box beside the recorder, and why it misses the small cluttered things a two-stage model finds.

Andreas Ohrvall · Sep 29, 2026

LexData
LexData

Product

  • Platform
  • Industries
  • Use cases

Resources

  • Docs
  • The Field Guide
  • Blog
  • Why models drift

Industries

  • Energy & utilities
  • Oil & gas
  • Agriculture
  • Manufacturing
  • Insurance
  • Retail
  • Robotics

Company

  • About
  • Customers
  • Careers
  • Contact

Trust

  • Security
  • Privacy
  • Terms

Stay updated

What we learn running vision models in production.

See everything.
Miss nothing.

Stay updated

What we learn running vision models in production.

Terms of use & Privacy policy

© 2026 LexData Labs · All rights reserved