Skip to content
LexDataLexData
PlatformIndustriesCustomers
DocsThe Field GuideBlogWhy models drift
AboutCareersSecurityContact
Log inStart now
← All posts

Computer vision · 7 min read

Pose estimation algorithms and what keypoints give a safety camera that boxes do not

A box says a person is beside the press. Keypoints say where the wrist is. How joints are learned as heatmaps, and which questions are worth the cost of asking.

Summary

This post explains pose estimation through two questions a box cannot answer: whether a worker's wrist has crossed into a press zone, and whether a shopper is reaching for a shelf or walking past it. It covers how keypoints are learned as heatmaps, why the rule is written on a joint rather than a box, and which frames come back for review. It concludes that most safety questions are answered by a box and a zone, and that keypoints are worth their cost only when the question is about a limb.

Rob Hickey · Chief AI Officer · Oct 3, 2026

Robot arm loading a machine behind a safety fence, the fence and the arm boxed, generated scene with detections from our model

The camera over the press on line 6 sees an operator at the loading position all shift. A box around the operator is on every frame, and it tells the safety system nothing it does not know, because the operator is supposed to be there. The question is narrower than "is a person near the press". It is "is a hand inside the guard line while the ram is down", and the box has no hand in it. It has a rectangle.

That is the gap pose estimation fills. Rather than a rectangle around a person, the model returns a set of points: shoulders, elbows, wrists, hips, knees, ankles, seventeen or so on a common layout. A wrist is a point with a position, and a position can be compared with a line.

A box says a person is near the press and keypoints say where the wrist is

A detector answers where the person is. A pose model answers where the parts of the person are, and it is the parts that safety rules are written about. The guard line on the press is a line on the floor and in the frame; the rule is that no wrist crosses it while the ram cycles. With keypoints, that rule is a comparison between a point and a line on every frame. With a box, the best the system can do is ask whether the box's edge crosses the line, and the box's edge is the operator's elbow, or their apron, or the tool in their hand.

The same shape appears wherever the question is about a limb rather than a body. Whether a technician's hand is on the rail when the lift moves. Whether a worker on a scaffold is holding on. The hazard zone intrusion detection use case is a box-and-zone question for most sites, and it becomes a keypoint question at the moment the zone is smaller than a person.

The press operator on line 6 wears a glove on the left hand only, for the hot side of the part, and the pose model has to find the wrist under the glove as reliably as the bare one.

Keypoints are learned as heatmaps with one map per joint

Modern pose models do not predict a wrist's coordinates directly. For each joint they produce a heatmap over the frame, a blurred spot of high values where the joint most likely is, and the keypoint is the peak of that map. Training the model means showing it frames where a person has marked each joint, and the model learns to light up the map at the marked spot.

The heatmap has a property that matters on the press on line 6. Its peak has a height as well as a position, and the height falls when the joint is hidden or ambiguous. A wrist behind the part being loaded produces a low, flat map rather than a confident wrong point, and the safety rule can treat a low peak as unknown instead of as safe. The earlier approach of regressing coordinates directly gave a number every time, whether or not the wrist was visible, and a number for an invisible wrist is exactly the wrong output for a guard line.

Multi-person scenes add a step. Either the model finds people first and estimates a pose inside each box, or it finds all the joints in the frame and then groups them into people by which joints connect. The first is more accurate per person and slower with a crowd; the second is the reverse. On a press with one operator, the choice barely matters. On a packing floor with twelve people it does.

The wrist crossing the guard line is a rule a box cannot write

With keypoints on every frame, the safety rule is a sentence. A wrist inside the guard zone while the press on line 6 is in its cycle, critical, sent to the line supervisor in Slack and to the press controller, with a cooldown so a single reach does not produce ten alerts. The monitoring and alerts doc describes the alert as a rule written that way, approved before it goes live, and the point of the keypoint is that the rule can name the wrist rather than the person.

The zone itself is drawn once, on the frame, as the guard line's projection into the camera's view. If the camera is moved, the line moves in the frame and the zone is wrong until it is redrawn, which is the same silent failure a moved camera causes for any fixed-camera rule.

A shopper reaching and a shopper passing have the same box

The other question is from a store, and it is not about safety. A camera over aisle 4 sees people walk past a shelf all day, and the store wants to know how many reach for the promotional display versus walk by. From a box, the two are the same: a person-sized rectangle beside the shelf. From keypoints, a reach is an arm extended toward the shelf with the wrist near the shelf plane, and a pass is arms at the sides and hips moving along the aisle.

This is a different kind of question from the press, a behaviour rather than a boundary, and it is answered by comparing joints to each other and to the shelf over a few frames rather than to a line on one frame. It also shows why the keypoints matter for the store's own privacy stance: the model that counts reaches never needs a face, and a skeleton is what leaves the camera.

Angles, occlusion and blur are the frames that come back for review

Pose models fail in predictable places. A person seen from directly above, which is how the camera over the press on line 6 sees them, has their shoulders and hips stacked, and the model, trained mostly on people seen from the side or front, mislabels left for right. A wrist behind a part, a leg behind a pallet, a torso behind another person: the hidden joints produce weak maps, and the grouping step can attach a visible arm to the wrong body. Motion blur on a fast reach smears the wrist into a streak, and the peak is somewhere along it.

Those are the frames that come back for review. Frames the model is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime. On a pose model the correction is a person dragging a keypoint to the right joint, which is slower than moving a box and matters more, because a rule on the wrist is only as good as the wrist.

Pose estimation earns its cost only where the question is about a limb

My own view is that pose estimation is used far more often than it is needed, and that most of the safety questions we are asked about are answered by a box and a zone. Whether anyone is inside the fence, whether a forklift and a person share a bay, whether a hard hat is on a head. Each of those is a position question, the what a vision model can and cannot see lesson would call it a detection problem, and keypoints add labeling cost and review time for an answer the box already gave.

The press on line 6 is the exception, and it is a real one. The zone is the size of a hand, the rule is about a wrist, and no box will ever say where the wrist is. For that question, the seventeen points are the cheapest honest answer there is.

See it on your own footage.

Start with your footage

More in Computer vision

Computer vision · 6 min read

What CLIP is and why typing a description finds the frame

A shared space for pictures and words lets an operator type a sentence and get the matching frames from a recorder. A way to look, and the boxes come after.

Rob Hickey · Oct 3, 2026

Computer vision · 6 min read

An end-to-end object detection workflow starts with decisions, and the first frame is labeled last

Tag, box, mask or keypoints. A class list with occlusion and size rules. A metric and an acceptance sentence. Then, and only then, the first camera's frames.

Ayman Quadir · Oct 3, 2026

Computer vision · 7 min read

Image recognition AI, explained on the cameras a site already owns

A tag, a box, a mask or keypoints, on a substation camera, a line camera and an aisle camera. How a model gets there, and what its score does not promise.

Esdras Ntuyenabo · Oct 3, 2026

LexData
LexData

Product

  • Platform
  • Industries
  • Use cases

Resources

  • Docs
  • The Field Guide
  • Blog
  • Why models drift

Industries

  • Energy & utilities
  • Oil & gas
  • Agriculture
  • Manufacturing
  • Insurance
  • Retail
  • Robotics

Company

  • About
  • Customers
  • Careers
  • Contact

Trust

  • Security
  • Privacy
  • Terms

Stay updated

What we learn running vision models in production.

See everything.
Miss nothing.

Stay updated

What we learn running vision models in production.

Terms of use & Privacy policy

© 2026 LexData Labs · All rights reserved