Computer vision · 6 min read
What semantic segmentation labels and what it costs
Drivable path, corrosion area and crop rows are questions about pixels, and the answer is a class for every one. Painted labels cost more than boxes.
Summary
This post explains semantic segmentation as a class for every pixel, using a tractor camera in an orchard, a corroded pipe rack and a field boundary as the questions it answers. It covers why the background class swamps the training loss and how focal loss rebalances it, what painted labels cost compared with boxes, and where instance masks are needed instead. It is for teams whose question is about an area rather than an object.
Esdras Ntuyenabo · Engineer · Sep 29, 2026

Orchard tree rows with three drivable paths marked from the tractor camera, from a customer tractor run
The camera on the front of the tractor in the orchard is not looking for an object. It is looking for the strip of ground between the tree rows that the tractor can drive on, which has no edges a box could hold, and which changes shape with every metre. A drone over a pipe rack in March is asking a similar question about rust: where is the corroded area, and how much of the pipe does it cover. A field-boundary survey asks it about hedges.
Those are questions about pixels. The answer is a class for every pixel in the frame, drivable or not, rust or sound metal, crop or hedge or track, and that is what semantic segmentation produces.
Every pixel gets a class and nothing gets an identity
A semantic segmentation model returns a frame-sized map in which each pixel carries one class. Two tree rows are both "tree row" with no line between them; every patch of rust is "rust" whether it is one patch or twelve. The model has no notion of instances, and for the tractor on the March run that is fine, because the tractor does not care how many drivable regions there are, only where the drivable pixels are in the frame in front of it.
That is the difference from detection in one sentence. A detector reports that a thing is here and draws a rectangle round it. A segmentation model assigns each pixel to a class, and the shape of the region is the answer rather than a rectangle around it.
Drivable ground, corrosion area and crop rows are pixel questions
The tractor camera's classes are drivable path, tree row and everything else, and the useful output is the drivable region's shape twenty metres ahead. The pipe-rack drone's classes on the March survey are rust, sound pipe and background, and the useful output is the rust area as a fraction of the pipe, which a maintenance rule can be written against and a box cannot supply. The field boundary segmentation use case asks for the outline of a field from an aerial frame, and the outline is the region's edge.
What the three share is that the rule downstream is about an area or an edge. A box would have to be corrected into a region by hand on every frame, and the correction would take longer than painting the region did.
Background dominates the loss until focal loss rebalances it
Here is the training problem that every segmentation set has and few teams expect. In the orchard frame the drivable path is a modest strip and the tree rows are large, but most of the pixels are sky, trunk, grass and the tractor's own bonnet, all labeled "background". On the pipe rack the rust is a few small patches on a frame that is mostly pipe and gravel. When the training loss is summed over every pixel, the background pixels, being easy and numerous, supply most of the total, and the model can lower the loss a long way by getting the background right and the rust roughly wrong.
Focal loss is the usual correction. It scales down the loss from pixels the model already predicts confidently, which are overwhelmingly the easy background ones, so that the hard pixels, the edge of the drivable strip, the small rust patch, keep their weight in the total. Weighting the rare class by hand does a similar job more crudely. Either way the model is made to attend to the pixels the rule cares about rather than the pixels there are most of, and without one of them a rust model on the March survey learns gravel.
Painted labels cost more than boxes and less than you would fear
A segmentation label is a painted mask: a person marks the region, pixel by pixel in principle and with a brush or a polygon in practice. It takes longer than a box, it is more tiring, and the edges are where two labelers disagree. Where does the drivable path stop and the grass begin, on a frame where the transition is a metre wide? The labeling rule has to say, and the labeling doc covers how a prompt becomes a class list and what the verification pass checks. On a segmentation job the verification pass is almost entirely at the edges.
On each frame, Lexi paints the first mask, and a person's job becomes correcting an edge rather than drawing one, which is the cheaper end of the work. My own view is that most teams label at a finer resolution than the rule needs. The tractor does not require the drivable edge to the pixel; it requires it to about the width of a tyre, and a label painted at that coarseness is faster to make, easier to check, and no worse for the model.
The orchard's rows, for what it is worth, are numbered from the far end, because the person who numbered them started at the gate that has since been moved.
Instance masks are needed when the count matters
When the question changes from "where is the rust" to "how many rust patches, and which is the largest", semantic segmentation runs out. Two patches touching are one region to it. Instance segmentation is needed from there: an outline per patch, each with its own identity, so that the largest one can be found and the count reported. It costs more to label, because each instance is its own polygon, and it is the right tool only when the rule is written against instances.
A good test is whether the downstream rule uses a number of things or a fraction of the frame. A fraction is semantic. A number is instance. The pipe-rack rule on the March survey, "rust covering more than a set share of any pipe section", is a fraction, and semantic segmentation answers it.
The doubted regions come back to a person
After training, the tractor model returns a mask on every frame and a doubted region on some. The doubtful ones are the edge of the drivable strip at the row end where the ground turns to mud, the shadow the trunk throws at 4 pm that looks like a path edge. Those regions go to a person, who moves the edge to where the ground actually is.
LexData takes the segmentation model through its whole life. You type what to look for, Lexi paints a mask on every frame, and a person checks each label before anything trains on it. The model then watches the cameras the farm already has, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime. The row ends in the mud are the frames the next version learns from, and the edge the person moved is the label.
See it on your own footage.
Start with your footageMore in Computer vision

Computer vision · 6 min read
What the big vision models are for, and what runs on the camera
The big models find the frames, draw the first boxes and tighten them, and a person checks. The small model trained on those labels runs beside the recorder.
Andreas Ohrvall · Sep 29, 2026

Computer vision · 6 min read
Why a good detector can still be a bad counter
The queue camera finds every shopper on every frame and still gets the count wrong. Swaps, lost tracks and jitter are tracker mistakes, and review fixes one.
Stephen Biswas · Sep 29, 2026

Computer vision · 6 min read
What YOLO does in one pass and what that costs
One look at the frame is why a single-pass detector fits the box beside the recorder, and why it misses the small cluttered things a two-stage model finds.
Andreas Ohrvall · Sep 29, 2026