Labeling · 6 min read
Hard hat detection datasets, helmet or head is the label that matters
A rig site camera labeled helmet, head and person. The head class is the one that makes a compliance rule possible, and the one most datasets leave out.
Summary
This post walks through labeling a rig site's PPE footage as helmet, head and person, and explains why the head class, the bare head with no helmet on it, is what turns detection into a compliance rule. It covers tight head boxes, crowded frames where workers hide each other, and checking the class balance before training. It is for safety and operations teams building or buying a hard hat dataset.
Ayman Quadir · Head of Product · Oct 2, 2026

Two workers under safety netting on a scaffold, hard hats boxed, from a customer site camera
The camera on the pipe rack at a rig site looks down on the mud pit walkway, and at the 6 am shift change there are eleven people in the frame. Nine wear helmets. One has his in his hand. One is bare-headed because he took it off to wipe his face and has not put it back on. The safety rule for that walkway is simple to state and the question for the model is exactly as simple: is anyone on this walkway without a helmet on their head?
Most hard hat datasets cannot answer it. They were labeled with one class, helmet, and a model trained on them can tell you where the helmets are. It cannot tell you where the heads are that have none.
Object detection with a helmet class alone never finds the violation
A model that boxes helmets sees nine helmets in the 6 am frame and reports nine. The man with the helmet in his hand is a tenth helmet, boxed at waist height, and the bare-headed man is nothing at all. He is the one the rule exists for and he produces no detection, because the class he belongs to was never labeled.
The frame needs three classes. Helmet, for a helmet on a head. Head, for a head with no helmet on it. And person, so that both can be tied to a body and a helmet held at waist height can be told apart from one being worn. The personnel PPE monitoring page puts it as per-person rather than per-frame: the question is whether each individual in the zone is wearing what the zone requires, and a helmet on a rail does not count as a helmet on a head.
The head class is the expensive one to label and the one that makes the dataset worth having. Every labeler wants to skip it, because bare heads are rare on a well-run site and each one takes a deliberate box. Every model trained without it will pass the violation.
Tight boxes on heads and a rule for the brim
The head and helmet boxes have to be tight, and tight means the same thing to every labeler. On the rig footage the ruling, written on a Monday and sent round, was that a helmet box covers the shell and the brim and stops at the chin strap, and a head box covers hair to chin. A box that includes the shoulders on some frames and stops at the ears on others teaches the model two different objects under one name.
Seen from the pipe rack camera, a helmet is a small object. At the far end of the walkway it is a dozen pixels across, orange or white against a grey deck, and the box around it has to be placed with the frame zoomed in. Labelers who draw at the default zoom draw boxes that are loose by a few pixels on each side, which on a dozen-pixel object is a box that is mostly deck.
The labeling guide has the line about writing the prompt like a work order, and it applies to the guideline too. "Helmet on head, box the shell and brim" is a work order. "Label PPE" is a request for eleven different datasets from eleven labelers.
Crowded frames hide heads behind shoulders
Shift change is the frame that matters and it is the frame that is hardest to label. Eleven people in a walkway two people wide means most of them are partly behind someone else. A head half hidden behind a shoulder is still a head and still gets a box, on the part that is visible. A helmet whose brim is visible above a colleague's hard hat gets a box on the brim.
The temptation is to skip the occluded ones because the box is a judgement call. Skipping them teaches the model that a partly hidden head is background, and on a crowded walkway most heads are partly hidden. The frames where the model is most needed are the frames with the most occlusion, and a dataset that only labels the clear view is a dataset for a quiet site.
My own view is that a site safety model should be trained on shift change and tested on shift change. A model that finds every bare head at 10 am when there are three people in the frame has been tested on the easy case, and compliance is judged on the worst moment.
Check the class balance before the first training run
Count the boxes per class before training and expect the count to be lopsided. On the rig footage after a month of labeling from camera 3 there were helmet boxes by the thousand, a comparable number of person boxes, and a few hundred head boxes, because bare heads are rare on a site that follows its own rules. A model trained on that split learns that the answer is almost always helmet, and it is right often enough that the score looks fine.
The score for the head class alone is the number to look at, and the score to distrust is the average across the three. A per-class recall on the head class in the low range means the model misses most bare heads, and no overall figure will show it. The fix is more head frames, which on a compliant site means going to the footage from the days it was hot, from the break area at the edge of the frame, from the contractor crews on their first week. Those are the frames where the head class lives.
Somebody on every rig site project ends up with a folder of frames called "no hat", collected by hand from the recorder, and it becomes the most valuable folder in the dataset.
The three classes make a rule a sentence can state
With helmet, head and person labeled, a rule can be written as a sentence: a head with no helmet in the pipe rack walkway, critical, sent to the shift supervisor's phone with the frame attached. The frame shows the person, the box on the bare head, and the walkway around them, and the supervisor can judge in one look from a phone whether it is a violation or a man wiping his face.
That is what the labeling was for. The oil and gas work behind our 12k+ precision image annotations delivered is footage of this kind, rig decks and walkways where the object is small and the rule is per person. The head class is the reason those annotations turn into a compliance rule at all.
Frames where the model is unsure, a head at the far end of the walkway in the glare off the mud pit, come back to a person, and the correction on that frame is a head box the next version learns from. The model that launched at shift change is the model that still works at shift change a year later, because the bare heads it doubted were the bare heads it was retrained on.
See it on your own footage.
Start with your footageMore in Labeling

Labeling · 7 min read
AI data labeling workflows, three ways to label footage and when each one pays
A pipeline right-of-way survey labeled three ways: every box by hand, Lexi proposing and a person checking, and synthetic frames for the leak nobody has filmed.
Rajiya Sultana · Oct 2, 2026

Labeling · 6 min read
Annotation analytics, the numbers a labeling queue produces besides labels
Three labelers on a month of warehouse footage produce boxes, and also a throughput, an acceptance rate and a map of where the rejections cluster.
Rajiya Sultana · Oct 2, 2026

Labeling · 7 min read
Annotation format conversion between COCO, YOLO and CVAT without losing a box
Three years of line inspection labels from two tools arrive in three formats. The boxes that shift are the ones nobody draws on a frame before training.
Esdras Ntuyenabo · Oct 2, 2026