Skip to content
LexDataLexData
PlatformIndustriesCustomers
DocsThe Field GuideBlogWhy models drift
AboutCareersSecurityContact
Log inStart now
← All posts

Computer vision · 6 min read

When the outline is the answer and a box will not do

A weld defect whose area decides the rework and a part a robot has to grasp both need a mask. Here is what the mask costs to label and when a box is enough.

Summary

This post separates the jobs where a bounding box is enough from the ones where only a per-instance outline will do, using a weld defect whose area decides rework and a part a robot must grasp. It explains what a Mask R-CNN style model returns, what polygon labels cost compared with boxes, and how COCO polygons carry them. It is for engineers deciding whether a job is a detection job or a segmentation job.

Esdras Ntuyenabo · Engineer · Sep 28, 2026

Robotic welding cell with a part on the fixture, metal part and robot arm boxed, generated scene with detections from our model

Two cameras in the same welding cell, on the Monday shift, want two different things from the same frame. The first watches the seam after the torch has passed and the question is whether the porosity along the weld is big enough to send the part back, which is a question about area. The second looks down at the fixture where the robot picks the finished part up, and the question is where the edge of the part is, to the millimetre, so the gripper closes on metal and not on air.

A box answers neither. A box around the porosity says a defect exists and gives its extent as a rectangle, which overstates the area of a thin crescent by several times. A box around the part tells the robot roughly where the part is, and roughly is how parts get dropped.

A box is enough when the answer is whether and where

Most jobs on a line are answered by a box. Is there a person in the cell? Did a part go past? Is the cap on the bottle? For those, the box gives a location and a class and that is the whole answer. Boxes are quick to draw, quick to check, and a person can verify a few hundred in an hour. A detector trained on boxes runs fast enough for a runner beside the recorder.

The temptation is to use boxes for everything because they are cheap, and to fix the shape problem afterwards with geometry. That works for round things and fails for the crescent of porosity, the L-shaped bracket, and the part lying at an angle on the fixture, which a box describes with a rectangle that is mostly fixture.

A Mask R-CNN outline answers area and edge questions

Instance segmentation returns, for each object found, an outline: the set of pixels that belong to that one instance and to no other. A Mask R-CNN style model does this by finding the object as a detector would and then predicting a mask inside the box. The mask is the part of the output the welding cell wants. The porosity's area in pixels is the count of pixels in its mask, and that count, through the camera's scale, is a figure in square millimetres the rework rule can be written against.

For the gripper the edge is the answer, and the mask's boundary is the edge. Two parts touching on the fixture are two masks, which is the difference between instance segmentation and the semantic kind that would paint both as "part" with no line between them. The weld quality inspection use case has the same shape: the defect is a region with an area, and the rule is about the area.

Polygon labels cost several times what boxes cost

A box is two clicks. A polygon around a crescent of porosity is a dozen points placed by a person who has to decide, at each one, where the defect stops and the sound weld begins. That decision is the label. It takes longer, it is more tiring, and two people will draw it slightly differently, which is why the labeling rule for a segmentation job has to say how tight the outline should be and how a fuzzy edge is handled.

On every frame, Lexi draws the first outline, and the person's job becomes moving points rather than placing them, which is the cheaper end of the work. The verification pass on a segmentation set is mostly at the edges: a point dragged onto the weld bead, a corner of the part cut off. The labeling doc describes what the verification pass checks and how the corrections compound into the next version.

Budget for it honestly. A set of a few hundred boxed frames becomes a set of a few hundred outlined frames at a multiple of the labeling time, and the multiple is the price of getting an area rather than a rectangle.

COCO polygons are the format the masks travel in

The outlines are stored as polygons in COCO JSON: for each instance, a class, a list of point coordinates around the boundary, and the box that encloses them. That is the format the dataset is imported in and exported as, and the same file carries the boxes for the classes that only need boxes, so a project can mix the two. A camera that needs masks for porosity and boxes for "torch on" writes both into one file.

The welder on the day shift, for what it is worth, calls the porosity "the wormholes", and the class list uses his word because it was the word already on the reject tags.

The porosity outlines the model doubts are the ones a person sees

After training, the model watches the seam camera and returns an outline on every weld it doubts as well as every weld it is sure of, and the doubted ones go to a person. A crescent whose edge the model drew through the middle of the bead, an outline that merged two pores into one, a part on the fixture whose corner was hidden by the gripper: those are the frames that come back.

LexData takes the segmentation model through its whole life. You type what to look for, Lexi draws an outline on every instance, and a person checks each label before anything trains on it. The model then watches the cameras the cell already has, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime.

My own view is that most teams pick segmentation too early, when a box would have answered the question, and the few who need it pick it too late, after a season of fighting rectangles with geometry. The test is a sentence: if the rule is written against an area or an edge, label the outline, and if it is written against whether or where, a box will do.

See it on your own footage.

Start with your footage

More in Computer vision

Computer vision · 6 min read

DETR explained, what the detection transformer changed about finding objects

Detection as set prediction, a fixed set of queries instead of anchors, no suppression step, and the small-object and slow-training problems later fixed.

Finn Ellingwood · Sep 28, 2026

Computer vision · 6 min read

Reading a training run before trusting the model

The loss curve says the run finished. The gap to validation, the recall on the crack class and the frames behind the score say whether the line can trust it.

Rob Hickey · Sep 28, 2026

Computer vision · 6 min read

What mAP measures and what it cannot tell you about your cameras

Mean average precision is built from per-class curves at an overlap line. A high score on the held-out set says nothing about the fence camera at 2 am.

Stephen Biswas · Sep 28, 2026

LexData
LexData

Product

  • Platform
  • Industries
  • Use cases

Resources

  • Docs
  • The Field Guide
  • Blog
  • Why models drift

Industries

  • Energy & utilities
  • Oil & gas
  • Agriculture
  • Manufacturing
  • Insurance
  • Retail
  • Robotics

Company

  • About
  • Customers
  • Careers
  • Contact

Trust

  • Security
  • Privacy
  • Terms

Stay updated

What we learn running vision models in production.

See everything.
Miss nothing.

Stay updated

What we learn running vision models in production.

Terms of use & Privacy policy

© 2026 LexData Labs · All rights reserved