Labeling · 7 min read
Outsourced data labeling for computer vision, what to write down first
Before anyone else boxes your packing station frames, the guideline, the gold set and the checks on each delivery are what decide whether the labels are usable.
Summary
This post follows a fulfilment team sending a packing station's frames out for labeling, and lists what has to exist before the first batch leaves: a guideline that rules on occlusion, truncation, small objects and class edges, a gold set to score the first delivery against, and the spatial, class and coverage checks that run on every delivery after. It argues that a person checking each label is what separates a usable dataset from a pile of boxes. It is for anyone buying labels rather than drawing them.
Sheikh Srijon · GTM Lead · Oct 2, 2026

Overhead view of a parcel packing station with boxes, tape and a crushed carton boxed, generated scene with detections from our model
The overhead camera on packing station 6 sees a bench, a roll of tape, a stack of flat cartons and, several times an hour, a carton that has been crushed at one corner before it reaches the bench. The fulfilment team wants every carton boxed and every crushed one flagged, across two months of footage, and they do not have the people to do it. So the frames are going out to a labeling team who have never seen the station, do not know what a crushed carton looks like at that bench, and will label exactly what they are told and nothing more.
What they are told is the whole project. The guideline that leaves with the frames is the dataset's specification, and everything that is not in it will be decided by a labeler on their own, differently each time.
The guideline rules on the cases the labeler will meet in the first hour
A guideline that asks for every carton boxed and the damaged ones flagged, and stops there, will produce a different dataset from every labeler who reads it. The questions that come up in the first hour of any labeling job are the same ones, and they need a ruling each, with a picture.
Occlusion first. A carton half behind the tape dispenser is a carton, and the box goes around the visible part. The guideline says so and shows one. Truncation next: a carton cut off by the frame edge gets a box up to the edge, and the frame is kept. Small objects: a label sticker is boxed only if it is wider than a stated number of pixels, because below that the box is guesswork and the model cannot learn it anyway.
Class edges: a carton with a scuffed corner is intact, a carton with a corner folded in is crushed, and the difference is a picture of each side by side, because no sentence draws that line as well as two frames do.
Then the cases that are specific to station 6. A flat, unfolded carton on the stack is not a carton for this project. A carton on the floor beside the bench is. Somebody who has watched an hour of the footage knows both of these and the labeler does not, and the guideline is where that hour of watching is written down.
The labeling guide puts it as writing the prompt like a work order, and the guideline for an outside team is the same document at greater length. Say what the packing crew would say. Then rule on the edge cases the packing crew never had to name.
A gold set scores the first delivery before the second is labeled
Before the frames go out, someone on the fulfilment team labels a hundred of them by hand on a Friday, carefully, against the guideline, and keeps them back. That is the gold set. The first delivery from the outside team includes those hundred frames, unmarked, mixed in with the rest, and the delivery is scored against them before anyone looks at the other frames.
The score is per class and per rule. Cartons found against cartons in the gold set. Crushed cartons found against crushed in the gold set. Boxes that overlap the gold box well enough to count as the same object. And, separately, the rules: were the truncated cartons boxed to the edge, were the flat cartons on the stack left alone. A delivery that scores well on cartons and badly on crushed has a labeling team that has not understood the class edge, and the fix is a call and three more pictures in the guideline, before the second batch is started.
My own view is that the gold set is the one thing a buyer of labels cannot skip. It is also the one most often skipped, because the hundred frames it takes to build are the frames the team wanted to outsource in the first place. A hundred careful frames is a day. A month of labels that have to be redrawn is a month.
Every delivery gets spatial, class and coverage checks
The gold set catches misunderstanding. Three cheap checks on every delivery catch drift, fatigue and the second labeler who joined in week three and never read the guideline.
The spatial check draws every box's position and size onto one blank frame. On station 6 the cartons should cluster on the bench and the stack, and the crushed ones should be a subset of the same region. A cloud of boxes in the top left corner where the camera sees only wall is a labeler who has been drawing on the wrong monitor. A set of boxes all the same size is a labeler who has been copying one box.
The class check counts boxes per class per delivery and compares with the last one. The crushed rate at the station is a few an hour and should stay near that from batch to batch. A delivery with no crushed cartons in it is either a good week for the packing crew or a labeler who stopped flagging, and the gold set in that batch says which.
The coverage check lists every frame in the delivery with no boxes on it. Some are empty, the bench at 3 am. Most are frames with cartons on them and no labels, and every one of those trains the model that a carton is background. The empty frames the labeler judged empty and the frames nobody labeled have to be told apart, which means the guideline asks for a mark on each frame that was looked at and found empty.
A person checking each label is what computer vision projects need
None of the checks above look at a single box and ask if it is right. They look at the delivery as a whole and ask if it is plausible, which is a lower bar. The bar that matters is a person opening each frame and confirming, fixing or rejecting each label. It is the bar we hold our own labeling to: millions of annotations, checked by hand, and labels that come back at up to 99.9% accuracy because a reviewer has looked at every one.
For computer vision projects built on bought labels, the practical version is a reviewer on the buyer's side who opens a random sample from each delivery and redraws it from scratch. The disagreement rate on that sample is the honest accuracy number for the batch, and it is the number that decides whether the delivery trains or goes back.
Somebody on the fulfilment team keeps the first delivery's rejected frames in a folder called "before the pictures". It is the clearest argument for the guideline that anyone on the project has.
The corrections keep coming after the outside team is gone
The labels that pass go into the training set, the model trains, and station 6 is watched. Frames the model doubts, a carton in a colour the labelers never saw, a crushed corner facing away from the camera, come back to a person, and the correction on each is a label drawn against the same guideline. The guideline the fulfilment team wrote for an outside labeling team is the same document the reviewer uses a year later, and every ruling added to it since is a case the station produced that nobody had anticipated.
That is the part of the guideline that never finishes. Write the first version before the frames leave. Add to it every time a frame comes back that the current version cannot settle.
See it on your own footage.
Start with your footageMore in Labeling

Labeling · 7 min read
AI data labeling workflows, three ways to label footage and when each one pays
A pipeline right-of-way survey labeled three ways: every box by hand, Lexi proposing and a person checking, and synthetic frames for the leak nobody has filmed.
Rajiya Sultana · Oct 2, 2026

Labeling · 6 min read
Annotation analytics, the numbers a labeling queue produces besides labels
Three labelers on a month of warehouse footage produce boxes, and also a throughput, an acceptance rate and a map of where the rejections cluster.
Rajiya Sultana · Oct 2, 2026

Labeling · 7 min read
Annotation format conversion between COCO, YOLO and CVAT without losing a box
Three years of line inspection labels from two tools arrive in three formats. The boxes that shift are the ones nobody draws on a frame before training.
Esdras Ntuyenabo · Oct 2, 2026