Industries · 6 min read
Object orientation from two keypoints, on a bottle line and a parcel sorter
Cap and base as keypoints, the angle between them as the answer, and consistent placement as the labeling rule that makes the whole thing hold.
Summary
This post describes reading an object's orientation with two keypoints instead of a box: cap and base on a bottle, the two ends of the long axis on a parcel, and the angle between them as the answer a robot or a sorter needs. It concludes that the labeling rule for where each point goes is the whole of the work, and that an angle should be smoothed across sampled frames rather than trusted from one. It is for robotics and packaging automation teams.
Esdras Ntuyenabo · Engineer · Sep 30, 2026

Bottling line, bottles queued under a fill head with one cap missing, bottle and cap boxed, generated scene with detections from our model
The unscrambler on line 3 tips a tote of empty bottles onto a rotating disc and is supposed to send every one of them down the track upright. Most of the time it does. A few times a shift a bottle arrives on its side, or upside down, and the filler either fills the floor or jams. The operators call an inverted bottle a diver. A robot cell at the other end of the plant has the same question about parcels in a tote: which way is this one facing, so the gripper can take it.
Both questions are an angle. A box does not contain one.
A box cannot tell you which way up a bottle is
A detector on line 3 returns a rectangle aligned with the image. A bottle standing upright has a tall narrow box. A bottle on its side has a wide flat one. A bottle upside down has exactly the same box as one the right way up, and a bottle tilted at some angle has a box whose shape says "tilted" and nothing about which way.
The width-to-height ratio of the box is sometimes used as a proxy, and it is a poor one. It cannot distinguish a diver from an upright bottle, it says nothing about direction, and on a parcel, which is a rectangle at any angle, the box's ratio changes with the parcel's rotation in a way that cannot be undone.
Pose estimation for a bottle is two keypoints and an angle
Two points solve it. A keypoint at the centre of the cap and a keypoint at the centre of the base, on every bottle in every frame, and the line between them is the bottle's axis. The angle of that line from vertical is the tilt, and the sign of it, whether the cap point is above or below the base point, says whether the bottle is a diver.
That is the whole of pose estimation for a bottle on a track, and it is smaller than the phrase suggests. No skeleton, no dozens of joints. Two named points and a bit of trigonometry, and the answer is a number in degrees that a rule can read: any bottle on line 3 whose axis is more than a few degrees off vertical, and the reject arm fires before the filler.
A parcel on a sorter is the same shape of problem with a different pair of points. The two ends of the parcel's long axis, or two corners of the label, give the angle the sorter needs to rotate the parcel so the label faces the scanner.
The labeling rule is where the model is made
Everything above depends on the two points being placed the same way on every bottle. That is the labeling rule, written down before the first frame is labeled and kept for the life of the project. The cap point goes on the centre of the cap's top face. The base point goes on the centre of the base as seen. When the cap is hidden behind a neighbouring bottle, the point is marked as not visible rather than guessed, because a guessed point teaches the model to guess.
You type the two points once, Lexi proposes them on every bottle in every frame, and a person checks their placement before anything trains. The labeling doc covers the verification pass. On a keypoint job the person is looking at nothing except whether each point sits where the rule says. A model trained on points that wander by a few pixels returns angles that wander by a few degrees.
My own view is that a team asking for rotated boxes should try two keypoints first. A rotated box is a harder label to draw consistently, and the angle it gives is the same angle two points give with less arguing about where the corners go.
Grasping needs the orientation before the gripper closes
The parcel in the tote is the harder version. The manipulation and grasping use case names why: objects arrive in bins, overlapping and at arbitrary angles, often shiny or transparent, and the pose has to be read from a partial view of a thing that is partly under other things. Two keypoints on the visible long axis are the start of that answer. A gripper that needs the full position and rotation of the parcel in space needs more than two points, and that is where the six degrees of freedom the use case describes come in.
For most of what happens on line 3, two are enough. The bottle is on a track, the parcel is on a belt, and the plane they lie in is known. The angle in that plane is the question, and two points answer it.
An angle from one frame is smoothed across several
A single frame can lie. Glare on the cap moves the cap point a few pixels and the angle jumps. A bottle behind another loses its base point for one frame and the angle is undefined. A rule that fires on one frame's angle rejects bottles that were fine.
So the angle is read across a run of sampled frames rather than from one, and a rule fires on the smoothed angle. A bottle whose angle jumps between frames in a way no bottle on a track can is held as doubtful rather than rejected, and the frame goes to a person. On line 3 that is a handful of frames a shift, and each one is a label the next version learns from.
Hidden caps and glare are the frames that come back
LexData takes the orientation model through its whole life. You type what to look for, Lexi puts two keypoints on every bottle in every frame, and a person checks each label before anything trains on it. The model then watches the camera over line 3, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime.
The frames that come back are the cap hidden behind a neighbour, glare off the cap at the hour the sun reaches the track, a bottle with no cap yet, and the diver itself, which the first version had seen only a few times. Each is a point corrected by a person, and when the corrections cross the project's threshold a new version trains on line 3's own footage. The operators still call it a diver, and the rule that rejects it is written in those words.
See it on your own footage.
Start with your footageMore in Industries

Industries · 6 min read
Perimeter security with fixed cameras, object detection and a drone sent to look
A frame every two seconds is enough to catch a person at the fence, a CPU is enough to run it, and the drone is the second look rather than the detector.
Andreas Ohrvall · Sep 30, 2026

Industries · 7 min read
Food service QA with a camera over the tray packing line
Every component on the tray gets a box, the missing one is flagged before the sealer, and the alert count is read against the line's own history.
Ayman Quadir · Sep 30, 2026

Industries · 7 min read
Railway safety with trackside cameras, zones and a signaller who can live with the alerts
People and vehicles boxed, the track bed and crossing drawn as zones, the frame sent to the control room, and a false alarm rate a signaller will keep reading.
Rajiya Sultana · Sep 30, 2026