Skip to content
LexDataLexData
PlatformIndustriesCustomers
DocsThe Field GuideBlogWhy models drift
AboutCareersSecurityContact
Log inStart now
← All posts

Industries · 7 min read

Zero-shot pose estimation for a collaborative robot, and what makes it hold

A generalist pose model reads a worker's reach on day one, keypoints are smoothed across sampled frames, and the site's own reviewed frames make it hold.

Summary

This post describes reading a worker's posture and reach from a collaborative robot's camera, starting from a generalist pose model that needs no training and works on day one. It concludes that the generalist model is a good first week and a poor first year, that keypoints have to be smoothed across sampled frames before a rule reads them, and that the site's own reviewed frames are what make the model hold on a new floor. It is for automation and safety engineering teams.

Rob Hickey · Chief AI Officer · Sep 30, 2026

Robotic welding cell, a part on the fixture and the robot arm boxed, generated scene with detections from our model

At station 6 a collaborative robot hands brackets to a worker who fits them to a housing. The two share a bench. When the worker leans in to seat a bracket, a hand crosses into the space the robot's arm sweeps, and the robot is supposed to slow before that happens rather than after. The camera above the bench sees the worker and the arm together, and at 1 pm on a Wednesday the worker leans further than usual because the bracket is tight.

Whether the robot slowed in time depends on a skeleton drawn on that frame and on what was done with it.

Zero-shot pose estimation is the start, not the finish

A generalist pose model returns a skeleton of keypoints on any person in any frame: shoulders, elbows, wrists, hips, knees, and a few more. It was trained on people in general, and it works on the worker at station 6 on the day the camera is switched on, with no labels from the plant. That is what zero-shot pose estimation means in practice, and it is the right way to start, because the plant gets a working skeleton in an afternoon and can begin writing rules against it.

It is not the finish. The generalist model has never seen this bench, this light from the skylight at 1 pm, the robot arm crossing in front of the worker's torso, or the welding jacket that hides the shape of an elbow. On the frames where those things happen, the skeleton wobbles, and the rule that reads it wobbles too.

My own view is that a zero-shot model is a good first week and a poor first year. The first week it does everything. By the end of the year every frame it doubted has either been reviewed and folded in, or it has been ignored and the skeleton is still wobbling on the same frames it wobbled on in week one.

Posture and reach are two numbers read from the skeleton

The rule does not want a skeleton. It wants two numbers. Reach is the distance from the worker's shoulder to the wrist, projected onto the bench, and the direction it points. Posture is the angle of the torso from vertical, read from the shoulders and hips. From those, the rule at station 6 is a sentence: a wrist inside the robot's sweep zone at high severity, and the robot slows; a torso leaning past a threshold for a run of frames at routine severity, and the ergonomics lead gets the frame.

The pedestrian and obstacle detection use case sets the bar for the first of those, and it is the same bar for a person at a bench as for a person in a vehicle's path. It is the safety-critical class, the acceptable miss rate is effectively zero, and it has to hold for people who are partly hidden, crouching, or moving fast at the edge of the frame. A wrist behind the robot's arm is exactly the partly hidden case, and it happens every cycle.

Keypoints are smoothed across sampled frames before a rule reads them

A single frame's skeleton is not to be trusted. A wrist point jumps to the bracket the worker is holding. An elbow vanishes behind the arm and reappears a few pixels off. A skeleton for one frame puts the worker's hip on the bench. If the rule reads each frame on its own, the robot slows for phantoms and the worker learns to ignore it.

So the keypoints are read across a run of sampled frames, about every two seconds each, and the rule fires on the smoothed reach rather than on any one frame's. A wrist that appears inside the sweep zone on one frame and outside on the next is held as doubtful rather than acted on. The frame goes to a person. On station 6 that is a handful of frames a shift, and each one is a label.

The worker at station 6 wears a hi-vis vest with reflective tape across the shoulders, and at certain angles under the skylight the tape reads as a wrist. Nobody would have predicted that, and it is on the list of doubted frames from the first week.

The site's own reviewed frames are what make it hold

This is where the generalist model becomes the plant's model. The frames it doubted, the arm across the torso, the tape on the vest, the 1 pm skylight, come back to a person, who confirms or corrects the keypoints. You type the joints you care about once, Lexi proposes them on the doubted frames, and a person checks each one before anything trains. The corrections train a version that has seen station 6.

LexData takes the pose model through its whole life. You type what to look for, Lexi puts a skeleton on every person in every frame, and a person checks each label before anything trains on it. The model then watches the camera above the bench, on a runner beside the cell's controller. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime. The skylight frames from the first week are in the version that watches the bench in the second month.

In our robotics work the figure we hold to is 99%+ safety-critical accuracy, and at a shared bench it is held by that mechanism: the generalist skeleton as the start, and the plant's own reviewed frames as what keeps it right.

A new floor is the rollout drift

The model that holds at station 6 is moved to the second plant, which has the same robot, the same bracket and the same bench drawing. The camera is mounted a little higher, the light comes from fluorescent tubes rather than a skylight, and the workers wear a different jacket. The skeleton wobbles on the first day exactly as the generalist one did at station 6 in week one.

The drift catalog files this as a new site came online: the model was trained at one site and deployed at another, and the second site does not look like the first. The fix is not to start over. The second floor labels a window of its own frames from its first shift, with the same joints and the same rule, and folds them into the version both plants run. The station 6 frames stay in, since station 6 is still running.

The rule at the second plant is the same sentence. The frames behind it are the second plant's own.

The override count is the signal

The measure of whether the pose model still holds is the people at the bench. Every time the robot slows for a phantom wrist and the worker waves it on, that is an override, and the rate of overrides per station per week is what says the model has started reading something new. A rise on station 6 under a new lighting fit-out is the light. A rise at the second plant in its first week is the rollout.

Neither needs a metric from the model to see, because the worker is already counting, and the frames behind the overrides say what to label next.

See it on your own footage.

Start with your footage

More in Industries

Industries · 6 min read

Perimeter security with fixed cameras, object detection and a drone sent to look

A frame every two seconds is enough to catch a person at the fence, a CPU is enough to run it, and the drone is the second look rather than the detector.

Andreas Ohrvall · Sep 30, 2026

Industries · 7 min read

Food service QA with a camera over the tray packing line

Every component on the tray gets a box, the missing one is flagged before the sealer, and the alert count is read against the line's own history.

Ayman Quadir · Sep 30, 2026

Industries · 7 min read

Railway safety with trackside cameras, zones and a signaller who can live with the alerts

People and vehicles boxed, the track bed and crossing drawn as zones, the frame sent to the control room, and a false alarm rate a signaller will keep reading.

Rajiya Sultana · Sep 30, 2026

LexData
LexData

Product

  • Platform
  • Industries
  • Use cases

Resources

  • Docs
  • The Field Guide
  • Blog
  • Why models drift

Industries

  • Energy & utilities
  • Oil & gas
  • Agriculture
  • Manufacturing
  • Insurance
  • Retail
  • Robotics

Company

  • About
  • Customers
  • Careers
  • Contact

Trust

  • Security
  • Privacy
  • Terms

Stay updated

What we learn running vision models in production.

See everything.
Miss nothing.

Stay updated

What we learn running vision models in production.

Terms of use & Privacy policy

© 2026 LexData Labs · All rights reserved