Labeling · 6 min read
Annotation analytics, the numbers a labeling queue produces besides labels
Three labelers on a month of warehouse footage produce boxes, and also a throughput, an acceptance rate and a map of where the rejections cluster.
Summary
This post follows a three-person labeling queue on warehouse dock footage and reads the numbers it throws off: frames per hour, first-pass acceptance, how often a model proposal was kept, and where the reviewer's rejections cluster. It argues that the rejection map is the number that improves the dataset, and that a QA pass on every label is what gets labels to up to 99.9% accuracy. It is for the person running a labeling effort rather than drawing in it.
Rajiya Sultana · Engineering Manager · Oct 2, 2026

Loading dock with trucks at the bays and a forklift carrying a pallet, generated scene with detections from our model
Three labelers start on Monday with a month of footage from the camera over dock 4. The classes are forklift, pallet, person and truck, the frames are sampled every couple of seconds from the recorder, and by Friday the queue has produced a few thousand boxed frames and one reviewer's verdict on each. The labels are the deliverable. The numbers the week threw off alongside them are what this post is about, because they say whether the labels can be trusted before a single epoch runs.
Most of those numbers are cheap to collect and easy to misread.
Throughput is a cost figure and never a quality one
The first number anyone looks at is frames per hour, and it is the least useful one. Labeler A did sixty an hour on the dock frames, labeler B did forty, and the instinct is that A is the better labeler. What the number carries is how long each frame took, which depends on how many objects it held and how many the model had already proposed. The dock at 6 am is empty and a frame takes seconds. The dock at 2 pm has three trucks, two forklifts and a cluster of people at the bay, and a frame takes minutes.
Throughput budgets the job. It says nothing about whether the boxes are right, and a queue managed on throughput alone rewards whoever accepts the fastest. On a labeling job the fastest way to accept is to stop looking.
Count unique frames rather than sessions, too. A labeler who opened the same frame three times because the tool crashed has three sessions on one frame, and a session count triples the apparent output.
First-pass acceptance says how well the labelers read the guideline
The number that matters more is how often the reviewer accepted a frame the first time it came through. On the dock job it started at about seven in ten and climbed through the week. A low first-pass rate at the start of a job is normal, since nobody has read the guideline the same way yet. A low rate that does not climb means the guideline is ambiguous, and a rate that differs between labelers on the same frames means one of them is reading it differently.
The rule that caused most of the early rejections was the pallet. Is a pallet on a forklift's tines a pallet, or part of the forklift? Is a stack of six empty pallets one box or six? The guideline said one thing, two labelers assumed another, and the reviewer's rejections through Tuesday were mostly that one ruling. Written down and sent round on Wednesday morning, and the acceptance rate moved the same afternoon.
First-pass acceptance is still agreement with a reviewer rather than agreement with the truth. A reviewer who has the same blind spot as the labeler accepts the same mistake. The check on the reviewer is a small random sample of accepted frames redrawn from scratch each session, and the disagreement on that sample is the honest accuracy number for the week.
Kept proposals show where object detection already works
Every frame in the queue arrives with proposals already drawn. Lexi has boxed what it took to be the forklifts and the trucks, and the labeler's first act is to confirm, fix or reject each box. How often a proposal is kept as drawn, by class, is a number the queue produces for free, and it is a preview of the detection model before that model has trained on a single corrected frame.
On the dock 4 footage the trucks were kept almost every time, since a truck at a bay is large and there are three of them. The pallets were kept about half the time. People were kept least, because the camera looks down on the bay from a high mount and a person seen from above beside a truck is a small dark shape the proposal often missed or doubled.
That per-class number tells the person running the queue where the labelers' time is going, and it tells the person planning the dataset which class needs more frames. The pallet proposals will improve as corrections feed the next pass, since every fix trains the proposer on that footage, and the second batch should arrive cleaner than the first. When the kept rate for a class stops climbing, the class needs more examples rather than more corrections.
Rejections cluster and the cluster names the problem
My own view, from running queues, is that the rejection map is worth more than every other number here. A reviewer rejects a frame with a reason, and reasons cluster. On the dock job there were three clusters. The pallet ruling through Tuesday. Boxes on people cut off at the frame edge on the left-hand bay. And a run of frames from Thursday afternoon where the sun came through the dock door and a labeler had boxed the shadow of a forklift as a forklift.
Each cluster is a different fix. The first was a guideline. The second was one labeler's habit, corrected in a minute once someone showed them. The third was a lighting condition the guideline had never mentioned, and it went into the guideline with a frame attached. A rejection rate on its own would have reported that about a fifth of frames came back, and pointed at nothing.
Rejection events and rejected frames are different counts. A frame sent back twice is one frame and two events, and a queue where the same frames keep bouncing has a disagreement between the labeler and the reviewer that the guideline is not settling. Somebody has to rule on it, in writing, and the ruling becomes the class definition from then on.
The reviewer on the dock job kept the rejected frames in a folder by reason. It became the training set for the next labeler who joined, and it was better than the guideline.
The QA pass on every label is where the accuracy figure comes from
None of these numbers is the accuracy of the labels. They are the health of the process that made them, and the process ends in a pass where a person looks at every label before it trains. That pass, with the random re-check behind it, is what gets labels to up to 99.9% accuracy, and it is the reason the figure can be said at all. The lesson on how much footage a model needs makes the related point that a few hundred clean instances per class beat a much larger pile of doubtful ones, and the queue's numbers are how you know which kind you have.
What trains is the approved set. The first-pass rate, the kept-proposal share and the rejection clusters go into the report for the week, and the report goes to whoever is planning the next month of footage from dock 4.
See it on your own footage.
Start with your footageMore in Labeling

Labeling · 7 min read
AI data labeling workflows, three ways to label footage and when each one pays
A pipeline right-of-way survey labeled three ways: every box by hand, Lexi proposing and a person checking, and synthetic frames for the leak nobody has filmed.
Rajiya Sultana · Oct 2, 2026

Labeling · 7 min read
Annotation format conversion between COCO, YOLO and CVAT without losing a box
Three years of line inspection labels from two tools arrive in three formats. The boxes that shift are the ones nobody draws on a frame before training.
Esdras Ntuyenabo · Oct 2, 2026

Labeling · 6 min read
Blur augmentation in computer vision, training for the blur the camera will produce
A camera on a vibrating gantry, an autofocus that hunts, fog on the housing. Train on the blur the line makes, never on the class it would erase.
Rob Hickey · Oct 2, 2026