Labeling · 6 min read
Instance segmentation data labeling where two parts touch
Two fillets overlap on a blue conveyor and a corroded patch fades into clean steel. The label is a decision about where one thing ends, made before the drawing.
Summary
This post takes two labeling jobs that a bounding box cannot do, fillets lying against each other on a processing line and a corroded patch with no clean edge, and works through where one instance ends and the next begins. It compares polygons with masks, describes the QA pass that checks every boundary, and settles on COCO as the format the labels leave in. It is for teams starting their first segmentation dataset.
Finn Ellingwood · Engineer · Oct 2, 2026

Fillets on a blue processing conveyor, each one boxed where they lie against each other, generated scene with detections from our model
Two fillets come down the blue conveyor on line 2 lying against each other, the tail of one under the shoulder of the next. The camera over the belt is there to count them and to measure each one, because the trim station downstream is set by size. A box around the pair says there is fish on the belt. Two masks, one per fillet, with a boundary drawn along the line where one slides under the other, say how big each one is.
That boundary is the whole labeling job, and it is a decision rather than an observation.
Instance segmentation asks a labeler to say, for every pixel, which object it belongs to. Most pixels are easy. The pixels along the line where two objects meet are the ones the model will learn the task from, and they are the ones the labeler is guessing at.
A bounding box answers where and a mask answers how much
The labeling guide puts it plainly: boxes answer where and how many, and they are the default. A mask is for the question where the exact shape is the answer, and it costs more per label because it carries more. On line 2 the question is the size of each fillet, which is an area, and an area is what a box discards. On a pipe rack the question is how far the rust has spread since the last inspection, and the corrosion use case says why a box will not do: the label has to be the measurement.
Start with boxes anyway, on any class where the shape is not yet known to matter. The classes that prove they need a mask get upgraded. A dataset where every class was masked from the first day has spent its labeling budget on shapes nobody will ever measure.
Where one instance ends is a rule, not a guess
On the fillet pair the question is whether the hidden tail belongs to the fillet it is part of or to the one covering it. Both answers are defensible. What is not defensible is one labeler choosing the first and the next labeler the second, because a model trained on both learns a boundary that wanders. The ruling on line 2 was that a mask covers only the visible part of its instance, and the hidden tail belongs to whatever is on top. Written down, with a picture, before the second labeler started.
The corroded patch is harder because there is no second object to hand the pixels to. Rust on the pipe rack fades from deep orange through staining to steel that is clean or nearly so, and two qualified inspectors will draw the edge in two different places. The guideline has to say where the patch ends. Ours, for the rack projects, says the mask stops where the surface stops being pitted, and staining without pitting is outside it. Another site could rule the other way. What matters is that the ruling exists before the drawing does, because a model cannot be more consistent than the labels that taught it.
Somebody on the corrosion job kept a printed ruler in the frame of every inspection photo. It turned out to be the only way to check the area a mask enclosed against anything real.
Polygons are cheaper and masks are truer
A polygon is a list of vertices around the object. A mask is a value for every pixel.
For a fillet on line 2, whose edge is smooth, a polygon of thirty vertices is a fine approximation and takes a minute. For a corroded patch whose edge is ragged, the polygon either has hundreds of vertices or it smooths the edge and loses the area the report needs.
The choice follows from the edge. Smooth objects, polygons. Ragged objects where the area is the number, masks, usually drawn by clicking and correcting a proposal rather than by hand. Either can be converted to the other for export, and the labeling guide notes that shape type drives what the labeling costs, so tell whoever is pricing the work which classes are which.
Holes are the case that catches people. A patch of rust with a bolt head in the middle is a mask with a hole in it, and a polygon has to be drawn as an outer ring and an inner one. A tool that does not support holes will fill the bolt with rust, and the area is wrong by the size of a bolt head on every frame that has one.
Every boundary gets a second pair of eyes
My own view is that segmentation is the one label type where a QA pass on every label is not optional. A box that is loose by a few pixels is a slightly worse box. A mask boundary that is drawn by a different rule on Friday's frames than on Monday's is two different objects under one name, and the model learns the disagreement rather than the object.
The reviewer opens each mask at high zoom along its edge, checks it against the ruling for that class, and checks what the mask missed. On line 2 that was usually the thin edge of a fillet against the belt, where the blue and the pale flesh had almost the same brightness. A random sample from each session is redrawn from scratch and compared with what was accepted, and the disagreement on that sample is the honest number for the week. Labels that go through that pass come back at up to 99.9% accuracy, and on a segmentation job that figure is carried entirely by the boundaries.
Merged instances are the reviewer's most common find. Two fillets under one mask because the labeler did not see the join, or two rust patches joined by a thin bridge of staining that the ruling says is outside both. The count is wrong by one either way, and on a line set by count that is a trim station running on the wrong size.
The labels leave as COCO with the mask reduced to a polygon
The dataset exports as COCO JSON, with each instance carrying its class, its polygon, and the box the polygon fits inside. The box comes free from the polygon, so a segmentation dataset is also a detection dataset, and a model that only needs to count the fillets can train on the same file as the one that measures them.
LexData takes that model through its whole life. You type what to look for, Lexi puts a mask on every frame, and a person checks each label before anything trains on it. The model then watches the camera over line 2, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime. The frames it doubts are, almost always, two fillets lying against each other at an angle the labelers had not drawn yet.
See it on your own footage.
Start with your footageMore in Labeling

Labeling · 7 min read
AI data labeling workflows, three ways to label footage and when each one pays
A pipeline right-of-way survey labeled three ways: every box by hand, Lexi proposing and a person checking, and synthetic frames for the leak nobody has filmed.
Rajiya Sultana · Oct 2, 2026

Labeling · 6 min read
Annotation analytics, the numbers a labeling queue produces besides labels
Three labelers on a month of warehouse footage produce boxes, and also a throughput, an acceptance rate and a map of where the rejections cluster.
Rajiya Sultana · Oct 2, 2026

Labeling · 7 min read
Annotation format conversion between COCO, YOLO and CVAT without losing a box
Three years of line inspection labels from two tools arrive in three formats. The boxes that shift are the ones nobody draws on a frame before training.
Esdras Ntuyenabo · Oct 2, 2026