Skip to content
LexDataLexData
PlatformIndustriesCustomers
DocsThe Field GuideBlogWhy models drift
AboutCareersSecurityContact
Log inStart now
← All posts

Computer vision · 7 min read

Object counting with computer vision, three kinds of count and why each needs its own pipeline

Forty boxes on a pallet, unique trucks through a gate over a shift, people in a checkout zone right now. Same word, three pipelines, and two ways to be wrong.

Summary

This post separates three things people mean by counting: how many boxes are on a pallet in one frame, how many distinct trucks passed a gate over a shift, and how many people are inside a checkout zone right now. It shows why each needs a different pipeline, where occlusion undercounts and double counting overcounts, and how the wrong counts come back to a person as frames. It is for teams who have been asked for a number and want to know which kind.

Finn Ellingwood · Engineer · Oct 3, 2026

Checkout lane from a ceiling camera, a person and the products on the belt boxed, no faces, generated scene with detections from our model

Three people at the same company ask for counting in the same week in March. The warehouse lead wants to know how many boxes are on each pallet as it leaves the wrap station. The yard manager wants to know how many trucks came through the gate on the night shift. The store manager wants to know how many people are in the checkout area right now, so a second till can be opened before the queue reaches the end of the aisle.

All three say counting. None of them is asking for the same pipeline, and a team that builds one and points it at all three will get one right.

Counting forty boxes on a pallet is object detection on one frame

The pallet count is the simplest case, and it is pure object detection. One frame from the camera at wrap station 2, a box drawn by the model around each carton, and the count is the number of boxes. There is no time involved. The pallet is stationary for the few seconds the wrap arm circles it, and one clean frame is enough.

What makes it hard is what makes any detection hard on a dense scene: cartons touching, cartons stacked so that the back row is only visible as a strip, a label facing the camera on one and not the next. Training on frames from that camera at that station, with the cartons the site actually ships, gets most of the way. The rest is a threshold set for the count the site cares about, which is usually a count that must be exactly right, so the operating point is chosen for that and the doubtful pallets go to a person.

Counting trucks through a gate over a shift needs tracking

The gate count is a different problem, because a truck at gate 1 is in the frame for thirty seconds and the model sees it on every sample. Detection alone counts that truck fifteen times. What the yard manager wants is the number of distinct trucks, and that means following each one across frames, giving it an identity, and counting the identity once when it crosses a line drawn across the gate.

Tracking is where the pipeline grows. A truck that is detected, lost behind the barrier for two frames and detected again can become two trucks. Two trucks nose to tail can become one. The line has to be placed where the crossing is unambiguous, and the count has to be directional, so a truck that reverses out is not counted in twice. The traffic counting and conversion use case has the same shape for people at a store entrance, and the what a vision model can and cannot see lesson names it: a counting question is a tracking problem wearing a detection costume.

The site's own water delivery truck goes out through the gate to turn around and comes back in, twice a day, and the first version counted it as four deliveries. The yard manager knew the number was wrong before anyone looked at the frames.

Counting people in a checkout zone right now is a zone rule

The store's question is different again. Nobody wants a total. They want a number that is true at this moment: how many people are inside the drawn zone in front of the four tills at the Kingsway store. That is detection on the current frame plus a zone comparison, and the answer changes every few seconds. It is the count the store acts on, because the action is opening a till, and the action is for now.

Because it is a live number, it needs smoothing. A count that reads six, then four, then seven, then five in consecutive samples, as people shift and the model flickers on a person half behind a display, will open and close the second till every minute. The rule waits for the count to hold above the limit across a few samples before it fires, and the cooldown stops it firing again while the first till is still opening. The monitoring and alerts doc describes the rule in that form, a sentence with a severity and a cooldown, and for a zone count the cooldown is doing most of the work.

Occlusion undercounts and double counting overcounts

The two errors run in opposite directions and show up in different pipelines. Occlusion undercounts: the back row of cartons the camera at wrap station 2 cannot see, the shopper behind the promotional stand, the second truck behind the first. The count is low, and the pallet leaves with the wrong number on its sheet.

Double counting overcounts: the truck that lost its track and became two, the shopper who left the zone and came back, the carton that the model boxed twice because the label and the box edge both looked like a carton. The count is high, and a till opens for a queue that is not there.

Neither error is visible in the count itself, and this is why the frames matter. A count of thirty-eight boxes is only wrong if someone knows the pallet holds forty, and a person looking at the frame with the boxes drawn on it sees at once that the back row was missed. My own view is that every counting pipeline should run beside a manual tally for at least a week before anyone trusts it. The tally should be kept by the person who will use the count, because they are the one who will stop believing it first.

The line, the zone and the cooldown are where the counting logic lives

The model detects. The counting logic is a separate layer, and it is where most of the tuning happens. For the pallet, the logic is a threshold and a rule for what to do when the count is below the expected number. For gate 1, it is the line's position, its direction, and how long a lost track is kept alive before it is dropped. For the checkout, it is the zone's outline, the number of samples the count has to hold, and the cooldown.

None of that logic is learned, and all of it is written down, which is the useful property. When the gate count is wrong, the fix is often the line, moved a metre further inside the gate so the turning water truck never crosses it, and no retraining is involved.

Wrong counts come back as frames a person can settle

Where the model itself is wrong, the correction is a frame. Frames the model is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime. On the camera at wrap station 2 those are the pallets where the count and the expected number disagree. At gate 1 they are the tracks that were lost and recovered. At the checkout they are the samples where the zone count jumped. In each case a person sees the frame with the boxes drawn and settles it, and the settled frame is a label.

The three counts run on three cameras the company already had, and each one was wrong in its own way for the first week. The pallet count missed back rows until the camera was raised. The gate count was high until the line moved. The checkout count flickered until the cooldown was set. None of those fixes was a better model, and all of them were found by someone who knew what the number should have been.

See it on your own footage.

Start with your footage

More in Computer vision

Computer vision · 6 min read

What CLIP is and why typing a description finds the frame

A shared space for pictures and words lets an operator type a sentence and get the matching frames from a recorder. A way to look, and the boxes come after.

Rob Hickey · Oct 3, 2026

Computer vision · 6 min read

An end-to-end object detection workflow starts with decisions, and the first frame is labeled last

Tag, box, mask or keypoints. A class list with occlusion and size rules. A metric and an acceptance sentence. Then, and only then, the first camera's frames.

Ayman Quadir · Oct 3, 2026

Computer vision · 7 min read

Image recognition AI, explained on the cameras a site already owns

A tag, a box, a mask or keypoints, on a substation camera, a line camera and an aisle camera. How a model gets there, and what its score does not promise.

Esdras Ntuyenabo · Oct 3, 2026

LexData
LexData

Product

  • Platform
  • Industries
  • Use cases

Resources

  • Docs
  • The Field Guide
  • Blog
  • Why models drift

Industries

  • Energy & utilities
  • Oil & gas
  • Agriculture
  • Manufacturing
  • Insurance
  • Retail
  • Robotics

Company

  • About
  • Customers
  • Careers
  • Contact

Trust

  • Security
  • Privacy
  • Terms

Stay updated

What we learn running vision models in production.

See everything.
Miss nothing.

Stay updated

What we learn running vision models in production.

Terms of use & Privacy policy

© 2026 LexData Labs · All rights reserved