Labeling · 6 min read
Polygon annotations for object detection when the model only draws boxes
A part lying diagonal on the bench gets a box that is mostly bench. Rotate a polygon and the box stays tight, which is why the extra labeling minute pays back.
Summary
This post takes a long, thin part lying at an angle on an assembly bench and shows what happens to its bounding box under rotation and the other spatial augmentations: the box grows to cover bench, while a polygon around the part turns with it and the box drawn from the polygon afterwards stays tight. It explains why that matters for a model that only ever predicts boxes, and when the extra time per label pays back. It is for teams deciding whether polygons are worth drawing on a detection dataset.
Finn Ellingwood · Engineer · Oct 3, 2026

Assembly station with a housing, fasteners and a cable laid out, the cable boxed, generated scene with detections from our model
The cable on the assembly bench at station 4 lies wherever the operator dropped it, which on most frames is diagonal, from the top left of the tray toward the housing at the bottom right. The camera over the bench is there to confirm the cable is present before the housing is closed. A box around a diagonal cable is a large rectangle whose corners are bench and whose diagonal is cable, and the model that trains on it learns that a cable is a rectangle that is mostly tray.
That is the ordinary cost of a bounding box on a long thin thing at an angle, and on the frame as the camera took it there is not much to be done about it. The cost gets worse in training, when the frame is turned.
A bounding box grows every time the frame is rotated
Rotation is a common augmentation on a bench camera, because the tray can be placed at any angle and the model should not care. Turn the frame by twenty degrees and the cable turns with it. The box around the cable cannot turn, since a detection box is axis-aligned by definition, so the pipeline draws a new upright box around the corners of the old one. The old box was already larger than the cable. The new one is larger than the old box.
Turn it again for the next epoch, by a different angle, and the box grows again, from an already grown box. After a few augmented copies the label on the cable is a rectangle that covers most of the tray, and the model is being told that the tray is the cable. It learns a loose box, and at station 4 it draws loose boxes, and the rule that checks the cable is near the housing starts firing on frames where the cable is a hand's width away.
The same thing happens with shear and with any crop that is then resized unevenly. Every spatial augmentation that changes the angle of the object grows an axis-aligned box, and the growth compounds.
A polygon turns with the object and the box is drawn afterwards
Label the same cable at station 4 with a polygon, a dozen vertices along its length, and rotate the frame. The polygon's vertices turn with the pixels, exactly, and the polygon after the turn is as tight to the cable as it was before. Draw an upright box around the turned polygon and it is the smallest box that contains the cable at its new angle. It is not the box that contained the old box.
That is the whole mechanism. The model still trains on boxes and still predicts boxes. What changes is that the box it trains on is derived from the polygon after every augmentation rather than carried through from the original, and a box derived after the turn does not carry the previous turn's growth. Over any number of augmented copies the label stays as tight as a box can be.
The labeling guide puts polygons in the category of labels for questions where the exact shape is the answer, and that is still true. This post is about a second use: polygons as the source from which tight boxes are generated, on a dataset where the question is only where.
My own view is that a detection dataset with heavy rotation augmentation and long thin classes should be labeled with polygons from the start, and that most teams find this out after training a model that draws loose boxes and wondering why.
The extra minute per label pays back on some classes and not others
A polygon takes longer to draw than a box, and on a bench dataset of a few thousand frames the difference is days of labeling. Whether it pays back depends on two things about the class.
The first is shape. A cable, a bracket, a screwdriver, a pipe: long and thin, and at most angles the box around them is mostly background. A housing, a fastener, a washer: roughly as wide as tall, and the box is tight at any angle, and the polygon buys almost nothing. On the station 4 dataset the cable and the bracket got polygons and the housing and the fasteners kept boxes.
The second is the augmentation. A dataset trained without rotation, on a camera where the tray is always square to the frame, gets the same box from a polygon as from a box, because nothing ever turns. The polygon pays back in proportion to how much the pipeline rotates and shears, which is a number the training configuration holds.
Polygons come with an extra augmentation of their own, worth having where it applies. A polygon can be cut out of one frame and pasted onto another, cable onto a different tray, and the box for the pasted copy is exact. A box cannot be pasted without pasting the bench around it.
The polygon is reduced to a box at export and the model sees only the box
The dataset exports as COCO JSON with each cable carrying its polygon and the box that fits it. A training pipeline that only reads boxes reads the box. One that reads polygons rotates the polygon and derives the box itself. The choice is in the loader, and the label file supports both, so a dataset labeled with polygons can be trained either way and a dataset labeled with boxes can be trained only one way.
The footage lesson says a few hundred labeled instances per class is the practical floor for detection. On a long thin class with rotation in the pipeline, a few hundred tight boxes derived from polygons are worth more than a few hundred boxes that have grown through every epoch, and that is the arithmetic behind the extra minute.
Somebody at station 4 keeps the first loose-box model's output pinned beside the bench, a frame with a box that covers the whole tray. It is the argument for the polygons, and it is a better one than this post.
The corrections come back as polygons on the classes that need them
The live model at station 4 watches the bench and doubts the frames it cannot read, a cable coiled rather than laid out, a bracket half under the housing. Those frames come back to a person, and the correction on a cable is a polygon, because the project's ruling for that class says so, and the correction on a fastener is a box. When the corrections cross the project's threshold the model retrains, the polygons are rotated and reduced to tight boxes in the pipeline, and the new version replaces the old one with no downtime. The rule that checks the cable is near the housing fires when the cable is near the housing, which is what it was for.
See it on your own footage.
Start with your footageMore in Labeling

Labeling · 7 min read
Aerial dataset augmentation for drone frames where there is no up
A tower seen straight down has no top or bottom, so rotations and both flips are safe. Scale for altitude, brightness for sun, move every box with its pixels.
Esdras Ntuyenabo · Oct 3, 2026

Labeling · 6 min read
A collaborative data annotation workflow run as a pipeline
Batch the bottling line's frames by camera and shift, assign so nothing is boxed twice, attach the guideline, review every label, train on the approved set.
Rajiya Sultana · Oct 3, 2026

Labeling · 7 min read
Dataset health check for computer vision, what to look at before anything trains
A scratch dataset where every scratch sits in the centre of the frame will train a model that looks in the centre. Five counts to read before the first epoch.
Rajiya Sultana · Oct 3, 2026