Computer vision · 6 min read
Object detection vs image classification vs keypoint detection, and how to choose before you label
Is the frame defective, where are the defects, where are the joints of the arm. Three questions on one cell camera, and the answer fixes the label format.
Summary
This post takes one camera over a welding cell and asks the three questions three teams asked of it, whether the frame is defective, where the defects are, and where the joints of the arm are, and shows which task answers each. It explains what each choice means for the labels, from a tag per frame to a box per thing to points per instance, and argues for boxes when in doubt because a box can be collapsed to a tag and a tag cannot be expanded. It is for teams about to open a labeling project.
Sheikh Srijon · GTM Lead · Oct 4, 2026

Robotic welding cell with a part on the fixture, metal part and robot arm boxed, generated scene with detections from our model
The camera over the welding cell on line 6 looks down at a part on the fixture with the robot arm poised above it, and in the space of a month three people asked it three different questions. The QA lead wanted to know whether the weld on each part was acceptable. The rework station wanted to know where on the part the spatter and undercut were. The robot programmer wanted to know where the arm's joints were on every frame, to check the taught path against the real one.
One camera and three tasks, and each task fixes the label format before the first frame is uploaded.
Image classification answers whether and costs a tag per frame
The QA lead's question has a yes or no answer for the whole frame. Image classification returns one tag per frame, acceptable or reject, and nothing about where. For a cell like line 6 that presents one part per frame, centred on the fixture, the frame and the part are the same thing, and a tag on one is a verdict on the other.
The label is a tag. A person looks at the frame and presses one of two keys, and the verification pass is as quick as the labeling. What the tag cannot carry is the reason, so a rejected part with no marked defect is a part the rework station has to inspect from scratch, which is the price of the cheapest label.
Object detection answers where and costs a box per thing
The rework station's question has a position in it. Object detection returns a list of boxes, each with a class, spatter or undercut or porosity, and the rework welder reads the boxes as a map of what to grind and re-weld. On the line 6 frame a part with three spatter clusters and one run of undercut returns four boxes, and the station acts on each.
The label is a box per defect, drawn to a written rule so that every labeler's edge lands in the same place, and checked by a person before anything trains. The cost scales with the number of defects per frame and with the care taken on the edges, which is several times the cost of a tag on the same frame. It buys the where, and on line 6 the where is what the rework station was paying for.
The welders call spatter "the sparkle", and the class in the project is named sparkle, because the rework welder is the person reading the boxes.
Keypoint detection answers how a thing is arranged
The programmer's question is about neither whether nor where. It is about arrangement: where each named joint of the arm sits on each line 6 frame, from the base through the elbow to the torch tip, so that the taught path can be compared with the path the arm actually swept. Keypoint detection returns a set of named points per instance, joined into a skeleton, and for a robot arm the skeleton is the arm's pose.
The label is a box around the arm and then a point on each named joint, with each point marked visible or hidden. A joint behind the part is hidden and stays hidden in the label; a point placed by guesswork on a hidden joint teaches the model to guess. On a person the same layout gives head, shoulders and wrists, which is why the site safety use cases label box plus keypoints. Keypoints are the most expensive of the three labels per instance, and they are the only one that answers the programmer's question.
The person who acts on the answer chooses the task
The frames do not say which task they are for. The QA lead, the rework welder and the programmer do, because each of them has an action waiting on the answer, and the shape of the action is the shape of the label. Divert the part needs a tag. Grind here needs a box. Compare this pose with the taught one needs points.
The lesson on what a vision model can and cannot see opens on the same idea: questions that sound alike are different problems with different labels, and teams routinely scope one and discover in month four that the stakeholder wanted another. On line 6 the fourth question, whether the weld will hold, was asked as well, and it is not a vision question at all. A weld that will fail and a weld that will hold can be the same photograph.
The label format is fixed before the first frame is uploaded
In LexAnnotate the task is set by what you type. Describe what to find as a tag per frame, a box per defect, or points on the arm, and Lexi proposes labels of that shape on every frame for a person to check. The labeling doc covers how the prompt becomes a class list and what the verification pass does with it. The dataset leaves as COCO, which carries tags, boxes and keypoints in one format, and the model leaves as PT, ONNX or TorchScript.
What cannot be done afterwards is the thing to know before starting. A set of boxes can be collapsed into tags for free: any frame with a defect box is a reject frame. A set of tags cannot be expanded into boxes without labeling every frame again, and a set of boxes cannot grow points without a second pass on every instance.
My own view, for a team that cannot decide, is to label boxes. It costs more than tags on the day, and it is the only one of the three that can be turned into the cheaper label later without touching a frame.
Two tasks on one camera is common and the loop treats them alike
Line 6 ended up with two models on the same camera: a detector for the rework station and a keypoint model for the programmer, with the QA lead's verdict derived from the detector's boxes. That is a normal outcome. A camera has one view and a plant has several questions about it, and the labels for each question are separate projects on the same footage.
LexData takes each of them through its whole life. You type what to look for, Lexi puts a tag, a box or the points on every frame, and a person checks each label before anything trains on it. The model then watches the cell camera, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime. The doubted frames from the detector and from the keypoint model land in the same review queue, and the same welder settles both.
See it on your own footage.
Start with your footageMore in Computer vision

Computer vision · 7 min read
Five computer vision applications in production, on the cameras a site already owns
Detect, count, inspect, track and read are the five jobs a fixed camera can be given, and each is a different question with a different label behind it.
Ayman Quadir · Oct 4, 2026

Computer vision · 6 min read
Computer vision projects worth building on the cameras you already have
A plant, a utility, a grower, a store and a warehouse each have a project that starts on an existing camera, produces a decision, and has someone to act on it.
Ayman Quadir · Oct 4, 2026

Computer vision · 6 min read
How to choose an object detection model architecture for a camera on your own site
Start from the decision the plant has to make, then the box beside the recorder the model has to fit, and only then the family. The benchmark comes last.
Andreas Ohrvall · Oct 4, 2026