Skip to content
LexDataLexData
PlatformIndustriesCustomers
DocsThe Field GuideBlogWhy models drift
AboutCareersSecurityContact
Log inStart now
← All posts

Computer vision · 7 min read

Image segmentation and the questions on a weld bead that only a mask can answer

A box finds the weld. A mask measures it. Semantic, instance and panoptic told apart on a weld cell and a board camera, with what each costs to label.

Summary

This post explains image segmentation through a weld bead in a robot cell and a cracked solder joint on a board camera, separating semantic, instance and panoptic segmentation by the question each answers. It concludes that a mask is worth its labeling cost only when the line needs a measurement a box cannot give, and that COCO polygons are how the mask survives export. It is for manufacturing teams deciding whether their inspection needs outlines or boxes.

Esdras Ntuyenabo · Engineer · Oct 3, 2026

Robotic welding cell with the torch and the robot boxed over the fixture, generated scene with detections from our model

The camera over the robot weld cell on line 3 sees every bead a few seconds after the torch lifts. The question the line asks is whether the bead is wide enough, and whether it is wide enough all the way along. A box around the bead answers neither. It says a bead is there, which the robot already knew. The width, the taper at the end, the undercut along one edge: those live in the outline of the bead, and getting the outline is what segmentation is for.

Across the plant, a camera over a circuit board sees a solder joint with a hairline crack through it. A box around the joint says a joint is there. The crack is a line of pixels inside it.

A bounding box says where the weld is and a mask says how wide it is

Detection puts a rectangle around an object. Segmentation labels each pixel of the bead on line 3, so the result is a mask that follows the object's edge. For a weld bead, the difference is the difference between knowing the bead runs from here to there and knowing its width at every point along that run. The weld quality inspection use case is a shape question, and shape needs the outline.

The rectangle is still the right answer for most questions on the line. Whether a part is present in the fixture, whether the torch is over the seam, whether a person is inside the cell. Each of those is a position question, and a bounding box answers it at a fraction of the labeling cost. Segmentation earns its keep when the answer is a measurement or an area, and nowhere else.

Semantic, instance and panoptic answer three different questions

Semantic segmentation gives every pixel a class: weld, base metal, spatter, background. It does not distinguish one bead from the next, so two beads that touch become one region of weld. On a part with a single seam that is fine. On a fixture holding four parts, it is not, because the line wants to know which bead is narrow.

Instance segmentation gives each object its own mask. Four beads, four masks, each with its own width profile, and each can be flagged on its own. It is what the weld cell needs, and it is what the board camera needs when three joints on the same pad each have to be judged separately.

Panoptic segmentation does both at once: countable things get instance masks, and the uncountable background, the base metal and the fixture, gets a semantic label. It is the most complete answer and the most expensive to label, and on a weld cell it is rarely worth it, because nobody is asking a question about the fixture.

The inspector on line 3 measures bead width with a comb gauge held against the part, and the mask is doing in pixels what the comb does in steel.

A mask lets the line measure the bead instead of finding it

With an instance mask for each bead on line 3, the measurements fall out of geometry. Width at each point along the bead's length, from the mask's boundary. Area, from the pixel count. Taper, from how the width changes toward the end. Undercut, from a notch in the boundary where the bead meets the base metal. Each is a number the line can put a limit on, in the same units the weld procedure specification uses, once the camera has been calibrated against a known part.

That calibration is the step that turns pixels into millimetres, and it belongs to the camera and its mount rather than to the model. A camera that is moved, or a lens that is swapped, changes the number of pixels per millimetre and every width the mask reports is off by the same factor until someone re-measures the known part.

On the board camera the measurement is different in kind. The crack has almost no area, so the useful output is its length and whether it crosses the joint from one side to the other, and a mask that follows a hairline is a hard mask to draw.

The mask costs more to draw and the reviewer pays for it

A box is two clicks. A mask is a polygon traced around the bead, and a bead's edge is irregular, so a careful outline is thirty or forty points. Lexi proposes the outline and a person adjusts it, which brings the time down. A reviewer still has to check each edge rather than each corner, and a QA pass on a segmentation dataset takes several times longer per frame than on a detection dataset.

My own view is that most segmentation projects should have been detection projects, and that the mask was chosen because it looked more precise rather than because anyone needed the measurement. Before the first polygon is drawn, the question to answer is which number the line will act on. If it is a count or a position, the labeling with Lexi doc will steer the project to boxes and the dataset will be built in a fraction of the time.

Where the measurement is needed, the mask is worth every point. The weld cell on line 3 is that case.

COCO carries the mask as a polygon so it survives the export

A mask has to be stored somehow, and the format decides whether it can leave the tool. COCO JSON stores each instance as a polygon, a list of vertices in frame coordinates, alongside the class and the bounding box that encloses it. That is the format the dataset is exported in, and it is the format the next tool, or the next version of the model, reads it back from.

Storing masks as images, one per instance, is the alternative, and it is the one that becomes unmanageable at a few thousand frames. Polygons stay small, keep the instance identity, and can be edited by moving a vertex. The choice of format is made once, at the start, and a project that starts with polygons never has to convert.

The cracked joint is the frame that comes back for review

LexData takes the weld model through its whole life. You type what to look for, Lexi draws the outline on every bead in every frame, and a person checks each mask before anything trains on it. The model then watches the cell camera the plant already has, in the cloud, on the plant's servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime.

On a segmentation model the frames that come back are the ones where the outline is uncertain: a bead under spatter, a joint where the crack is a pixel wide, a part in the fixture at an angle the dataset did not have. A person fixes the outline and the fix is a label. On the manufacturing lines we run, that review is how a mask model holds 99%+ accuracy in production, because a bead that was mis-measured on Tuesday is a bead the model measures correctly by the next version.

See it on your own footage.

Start with your footage

More in Computer vision

Computer vision · 6 min read

What CLIP is and why typing a description finds the frame

A shared space for pictures and words lets an operator type a sentence and get the matching frames from a recorder. A way to look, and the boxes come after.

Rob Hickey · Oct 3, 2026

Computer vision · 6 min read

An end-to-end object detection workflow starts with decisions, and the first frame is labeled last

Tag, box, mask or keypoints. A class list with occlusion and size rules. A metric and an acceptance sentence. Then, and only then, the first camera's frames.

Ayman Quadir · Oct 3, 2026

Computer vision · 7 min read

Image recognition AI, explained on the cameras a site already owns

A tag, a box, a mask or keypoints, on a substation camera, a line camera and an aisle camera. How a model gets there, and what its score does not promise.

Esdras Ntuyenabo · Oct 3, 2026

LexData
LexData

Product

  • Platform
  • Industries
  • Use cases

Resources

  • Docs
  • The Field Guide
  • Blog
  • Why models drift

Industries

  • Energy & utilities
  • Oil & gas
  • Agriculture
  • Manufacturing
  • Insurance
  • Retail
  • Robotics

Company

  • About
  • Customers
  • Careers
  • Contact

Trust

  • Security
  • Privacy
  • Terms

Stay updated

What we learn running vision models in production.

See everything.
Miss nothing.

Stay updated

What we learn running vision models in production.

Terms of use & Privacy policy

© 2026 LexData Labs · All rights reserved