Labeling · 7 min read
How to evaluate an image annotation partner before sending them a year of footage
Six questions for a labeling partner, from how accuracy is measured to who settles a disagreement, asked before the inspection footage leaves the building.
Summary
This post sets out the questions a team should ask a labeling partner before handing over a year of inspection footage: how accuracy is measured and audited, who resolves disagreements, how labelers learn the ontology, where the footage lives, whether every label is checked, and what happens after launch. It concludes that the answers to the audit and disagreement questions predict the dataset's quality better than any quoted accuracy figure. It is for teams choosing where their first dataset gets made.
Ayman Quadir · Head of Product · Oct 3, 2026

Lattice tower from a drone pass, insulator strings boxed and two corrosion spots flagged on the hardware, from a customer inspection run
A utility has a year of drone footage from its transmission line inspections sitting on a drive since last March: several hundred flights, each a few thousand frames of towers, insulators and the hardware between them. The team wants a model that flags corrosion and cracked insulators, and they do not have the people to label a year of footage themselves. So the footage is going out to a partner, and the question is which one.
The quoted accuracy figures on the brochures are all high and all mean different things. What separates a partner who will produce a dataset the model can learn from, and one who will produce a folder of boxes that look right and are not, is a handful of questions about method, asked before the first frame leaves.
How accuracy is measured and who audits it
Every partner quotes an accuracy figure. Ask what it is a figure of. Accuracy against what reference, on what sample, measured by whom? A figure measured by the labelers against their own work is a different thing from one measured by a separate reviewer against a gold set the client supplied. The first is a self-report. The second is an audit.
Ask, too, whether the figure is per label or per frame. A frame with six insulators and one missed is a correct frame at the frame level and five out of six at the label level, and the model trains on labels. On inspection footage where a tower carries a dozen strings, per-frame accuracy hides most of the misses.
Our own figure is labels at up to 99.9% accuracy, and it comes from a QA pass on every label before training, which is the practice the next questions are about. The figure without the practice behind it is a number on a page.
Who resolves a disagreement between two labelers
Two labelers looking at the same insulator on the same Wednesday flight will disagree about whether the discolouration at the cap is corrosion or shadow. That disagreement is the most valuable event in the whole labeling process, because it marks a frame where the class definition is unclear, and the partner's answer to "what happens then" tells you how the dataset will turn out.
The bad answer is that the label goes with whoever finished the frame. The adequate answer is that a senior labeler decides. The good answer is that the disagreement is logged, a domain reviewer with the client's own inspector decides, and the decision is written back into the labeling guide so the next labeler who meets a shadowed cap has a rule to follow. The corrosion that the utility's own linesman would call in is the corrosion the model should learn, and only the linesman knows which that is.
Whether the labelers learn the ontology before the first frame
A class list for transmission inspection is short and every entry hides a decision. Does "cracked insulator" include a chipped shed? Is surface rust on a bolt "corrosion" or is that reserved for pitting? Where does the box go on an insulator string that crosses the frame edge? A partner who starts labeling the day the footage arrives has answered those questions for themselves, differently on each shift.
Ask how the labelers are trained on the ontology. A written guide with example frames for each class and each edge case, a calibration batch labeled by everyone and compared, and a way to add a rule when a new case appears. The labeling with Lexi doc frames the class definition as the first thing typed, before any box, for the same reason: the model learns whatever the labelers agreed on, and if they never agreed, it learns the average of their disagreement.
The first frame of every flight in the utility's archive shows the pilot's hand over the lens, because the drone starts recording before the pre-flight check ends. A partner who boxes the hand as nothing and moves on has read the guide. One who asks what class a hand is has read it more carefully.
Where the footage lives while it is being labeled
A year of inspection footage shows every tower on the network, its access track and the substations at each end. Ask where it is stored, who can open it, whether labelers see it through a tool or download it to a laptop, and what happens to the copies when the job ends. Ask whether labelers sign in with named accounts and whether access can be limited by project.
On our side, the footage stays in the workspace the client owns, labelers reach it through the tool with role based access and SSO, and the security page carries what we do and do not claim. A partner who cannot answer the storage question in a sentence has not been asked it before, and that is its own answer.
Whether every label gets a second look or only a sample
Most partners check a sample of labels. Ask what fraction, chosen how, and by whom. A random sample of a tenth of the frames from the Wednesday flights, checked by a reviewer who was not the labeler, catches systematic errors: a labeler who boxes the whole string when the guide says the damaged shed. It does not catch the individual miss, the one insulator at the top of the frame that nobody boxed on one flight.
My own view is that a sample-based QA rate quoted without the sample size is a marketing number, and that on footage where a miss is the whole point, a check on every label is the only method whose figure means what it says. Our practice is that pass on every label, and the millions of annotations checked by hand behind that figure are why the accuracy number can be quoted as a ceiling rather than an average.
Sampling has a place afterwards, as a check on the checkers. A second reviewer sampling the first reviewer's approvals is how the QA pass itself is kept honest over a year.
What happens to the labels after the model is live
The footage from the year on the drive, March to March, is the first dataset. Once the model is watching the flights that come after, frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime. The question for the partner is whether they are part of that, or whether the relationship ends when the folder is delivered.
The useful answer is that the returned frames go to the same labelers, under the same guide, with the same review, so that a correction made in the second year is the same kind of label as one made in the first. The how much footage lesson argues that the second dataset, the one built from what the first model got wrong, matters more than the first. A partner who prices only the first has priced half the job.
On the utility's footage, the first model's returned frames were mostly towers photographed against low winter sun. Those frames went back to the same people who had drawn the first boxes, and that continuity is what the six questions were for.
See it on your own footage.
Start with your footageMore in Labeling

Labeling · 7 min read
Aerial dataset augmentation for drone frames where there is no up
A tower seen straight down has no top or bottom, so rotations and both flips are safe. Scale for altitude, brightness for sun, move every box with its pixels.
Esdras Ntuyenabo · Oct 3, 2026

Labeling · 6 min read
A collaborative data annotation workflow run as a pipeline
Batch the bottling line's frames by camera and shift, assign so nothing is boxed twice, attach the guideline, review every label, train on the approved set.
Rajiya Sultana · Oct 3, 2026

Labeling · 7 min read
Dataset health check for computer vision, what to look at before anything trains
A scratch dataset where every scratch sits in the centre of the frame will train a model that looks in the centre. Five counts to read before the first epoch.
Rajiya Sultana · Oct 3, 2026