Labeling · 6 min read
A dataset quality audit for object detection when mAP hides a weak class
The PPE camera at the rig site scored well on average and missed most bare heads. Audit the labels per class, fix the boxes in review, and retrain on the fixes.
Summary
This post takes a PPE camera at a rig site whose overall mAP looked healthy while the hard hat class alone was weak, and walks through auditing the dataset per class, finding the mislabeled and missing boxes behind the weak number, fixing them in review, and measuring the retrain. It argues the average is the wrong number to report for a safety camera. It is for teams whose model scores well and still misses the thing that matters.
Rob Hickey · Chief AI Officer · Oct 3, 2026

Crane lift over a fenced site with workers in hi-vis, hard hats and the load boxed, generated scene with detections from our model
The PPE model for the drill floor camera came back from training with a mean average precision that everyone in the Thursday review was pleased with. Four classes, person, hard hat, hi-vis vest and gloves, and one number across all of them that sat comfortably where the project plan said it should. The safety lead asked one question: how good is it at the hard hat on its own?
Nobody had looked. When they did, the hard hat class was the weakest of the four by a wide margin, and the average had been carried by the person class, which is large, common and easy, and the vest, which is bright orange on a grey deck.
The camera exists for the hard hat. A model that finds every person and every vest and misses a third of bare heads is a model that has failed at the one job the safety lead cares about, and the score that everyone approved said nothing about it.
An object detection average is carried by the easy classes
Mean average precision is a mean. On the drill floor's four-class dataset, where one class is hard, three easy classes at a high score and one hard class at a low one produce an average that looks fine. The PPE monitoring page says where the cost lands: the failure that costs you is a false pass, and false passes are the ones nobody reports.
The first move in the audit, agreed in the same Thursday review, is to stop reporting the mean and report the per-class numbers, precision and recall separately, on a held-out set that a person has checked by hand. Recall on the hard hat class is the number the safety lead asked for. Precision on it says how many of the model's hard hat calls were right. On the drill floor camera, precision was acceptable and recall was poor: when the model said hard hat it was usually right, and it was saying it far too rarely.
Low recall on a small object has a short list of causes, and most of them are in the labels rather than the model.
Missing boxes teach the model that the object is background
The audit started with the training frames for the hard hat class, opened one at a time with the labels drawn on. Within the first hundred frames from camera 2 the pattern was visible. Frames from the day shift, with two or three people on the floor, had every hard hat boxed. Frames from shift change, with eleven people on the floor and half of them partly behind someone else, had the hard hats in the front row boxed and the ones behind left blank.
Every one of those blanks is a hard hat the model was trained to ignore. A frame where a hat is visible and unlabeled is a frame that says, in training, that this shape on a deck is not a hard hat. The model learned it, and at shift change, when it was most needed, it did what it had been taught.
The second pattern was smaller and stranger. A run of frames from one week had hard hats boxed loosely, with the box including the shoulders, by one labeler who had joined that week. A loose box on a small object is mostly deck, and the model that trains on it learns a hard hat that is half background.
Neither pattern is visible from the score. Both are visible from opening frames, sorted by the model's own doubt, and the frames it doubted most were the shift change frames with the unlabeled back row.
Mislabels cluster where two classes look alike
The third finding was a class confusion. A white hard hat seen from above, on the drill floor camera's high mount, is a pale oval, and so is a bald head in the sun. A few dozen bare heads had been boxed as hard hats, mostly on bright afternoons in July, mostly by the same labeler. The model had learned that a pale oval is a hard hat, which is also the reason its precision on the class was only acceptable rather than good.
The way to find these is the same as for any confusion: list the frames where the model and the label disagree, and sort by class pair. The hard hat against bare head pair had the longest list, and the frames in it were the ones to reopen.
My own view is that the confusion pair list is the most useful page in any audit, more than any curve, because it says where the labeling guideline has a gap. The guideline for the drill floor had a picture of a hard hat and no picture of a bare head from above.
Fix the boxes in review and retrain on the fixes
The fixes went through the review queue over a Monday and a Tuesday rather than through a script. The shift change frames were reopened and the back row boxed. The loose boxes from the new labeler's week were redrawn tight, shell and brim. The pale ovals were relabeled, hat or head, one by one, and the guideline gained the picture it was missing. A person confirmed each change, because a batch relabel that is wrong is a second audit.
Then the model was retrained on the corrected set and scored again, per class, on the same held-out frames. Recall on the hard hat class rose, precision rose with it, and the mean moved by less than either, which is the whole point: the number everyone had been watching was the one least sensitive to the fix.
The confidence lesson makes a related point that applies here. The score the model attaches to a detection ranks that detection against others; it is not a probability that the detection is right, and it is not a measure of the dataset. A model that is confidently wrong about bare heads was confidently wrong because it was taught to be.
Somebody on the rig site printed the per-class table and taped it to the cabinet beside the recorder, with the hard hat row circled. It is still there.
The audit repeats every time corrections come back
The per-class audit is a thing to do before the first training run and again before every retrain, because the corrections that come back from the live camera are labels too and carry the same risks. Frames the model doubts on the drill floor come back to a person through the Awaiting your review queue, the person's correction becomes a label, and the retrain folds it in. Whether that correction was a tight box on a hard hat or a loose one, whether the bare heads were labeled as such, is the same question the audit asked the first time.
The drill floor camera now reports four numbers to the safety lead and no mean. When the hard hat row moves, somebody opens the frames.
See it on your own footage.
Start with your footageMore in Labeling

Labeling · 7 min read
Aerial dataset augmentation for drone frames where there is no up
A tower seen straight down has no top or bottom, so rotations and both flips are safe. Scale for altitude, brightness for sun, move every box with its pixels.
Esdras Ntuyenabo · Oct 3, 2026

Labeling · 6 min read
A collaborative data annotation workflow run as a pipeline
Batch the bottling line's frames by camera and shift, assign so nothing is boxed twice, attach the guideline, review every label, train on the approved set.
Rajiya Sultana · Oct 3, 2026

Labeling · 7 min read
Dataset health check for computer vision, what to look at before anything trains
A scratch dataset where every scratch sits in the centre of the frame will train a model that looks in the centre. Five counts to read before the first epoch.
Rajiya Sultana · Oct 3, 2026