Computer vision · 7 min read
Monocular vs stereo depth estimation for a robot that has to stop in time
A stop rule needs a distance in metres. Monocular depth gives a ranking, stereo gives metres and loses them on a blank wall. The rule decides which one.
Summary
This post compares monocular and stereo depth for a mobile robot in a warehouse aisle whose stop rule is written in metres. It concludes that monocular depth is a ranking that cannot carry a safety stop on its own, that stereo gives metric distance and fails on textureless walls, and that the choice follows from the stop rule rather than from a benchmark. It is for robotics teams deciding what the obstacle model should feed the brake.
Rob Hickey · Chief AI Officer · Oct 3, 2026

Forward camera at a junction, vehicles, pedestrians and signs boxed, from a customer perception run
The robot in aisle 14 of the distribution centre carries a pallet at walking pace and has one rule that matters more than the rest: stop before a person is closer than a metre and a half. The obstacle model finds the person. The depth model says how far away they are. And the difference between a depth model whose answer is closer than the shelving, and one whose answer is one point four metres, is the difference between a rule that can be written and one that cannot.
That is the whole choice between monocular and stereo depth, seen from the brake.
A stop rule needs a distance in metres and monocular gives a ranking
A monocular depth model takes one frame and produces a depth for every pixel. It learns to do this from cues a person also uses: the size of things it knows the size of, where the floor meets the wall, how texture compresses with distance, what sits in front of what. On a frame from aisle 14 it will correctly say that the person is nearer than the racking behind them and further than the pallet on the forks.
What it produces is an ordering. The nearest things get small values, the furthest get large ones, and the values are consistent with each other within the frame. They are not, on their own, metres. A model that has learned aisles in general does not know whether this aisle is three metres wide or four, and so it cannot know whether the person who fills a third of the frame is two metres away or three. The pedestrian and obstacle detection use case needs the number, because the stop rule is written in it.
Scale ambiguity means a small box close and a large box far look alike
The reason is geometric. A person one and a half metres tall at three metres and a child half that height at a metre and a half project to the same size on the sensor. One frame cannot tell them apart, and a monocular model that appears to is using a learned assumption about how tall people are. In a warehouse where the people are adults in the same hi-vis, that assumption mostly holds. On the day a pallet of mannequins for the retail display comes through aisle 14, it does not.
Monocular models can be given scale. A known camera height above a flat floor, a known object in the frame, a calibration against a tape measure at commissioning. Each one pins the ranking to metres for that camera, on that floor, and each one breaks when the assumption behind it does: the robot tilts on a ramp, the floor is not flat, the known object is not there. The result is a distance that is usually right and occasionally wrong by a factor, with nothing in the frame to say which.
My own view is that monocular depth alone should not be the input to a safety stop on a robot that moves among people, whatever a benchmark says about its accuracy. A ranking is a fine thing for planning a path and a poor thing for deciding when to brake.
Stereo gets metric depth from two cameras and loses it on a blank wall
Two cameras a known distance apart on the front of the robot in aisle 14 see the same person from slightly different positions, and the shift between the two views, the disparity, gives distance directly through triangulation. The baseline is measured once, the disparity is measured on every frame, and the output is metres with no learned assumption about how tall people are. For the stop rule, that is what is wanted.
The cost is that stereo needs something to match. Disparity is found by locating the same feature in both views, and a flat grey wall, a shiny floor, a plain pallet wrap under a strip light, offers nothing to locate. In those regions the depth is noise or absent. The warehouse's end-of-aisle walls are painted the same flat grey as the floor, at the fire officer's request, and the stereo pair on the robot returns almost nothing for them until the robot is close enough to see the texture in the paint.
Stereo also degrades with distance. Disparity shrinks as objects recede, so the metre-level precision at three metres becomes something coarser at fifteen, and the useful range is set by the baseline, which on a small robot is limited by the width of the robot.
Depth has to hold from one frame to the next or the robot brakes on noise
Both kinds of model produce a depth per frame, and neither promises that the depth on the next frame agrees. A monocular model can re-estimate the scale of the scene from one frame to the next and have the whole of aisle 14 breathe. A stereo pair can lose a match on one frame and recover it on the next, and the person jumps a metre closer and back.
A robot that brakes on the raw output brakes on every one of those jumps. The fix is the same in both cases: filter the depth of a tracked object over time, treat a single-frame jump as suspect, and decide to stop on a distance that has held for more than one frame. The what a vision model can and cannot see lesson makes a related point, that a tracking question and a detection question are different problems, and a stop rule is a tracking question wearing a depth model.
The choice is set by the stop rule, and most robots end up with both
Written down, the decision is short. If the rule is "stop when a person is within one and a half metres", the robot needs metres, and stereo or a range sensor gives them in a way monocular does not. If the rule is "slow down when something is ahead", a ranking is enough, and monocular depth from the camera the robot already has is cheaper and works on the grey wall.
Most robots on the robotics programmes we work with end up with both, used for different things: stereo or a range sensor feeding the stop, monocular from the same forward camera feeding the planner and filling the blank-wall regions where stereo has nothing. The 99%+ safety-critical accuracy those programmes hold is measured on the stop, and the stop is fed by the sensor that gives metres.
The depth that matters is checked at review against a tape measure
Whichever model feeds the brake, its errors show up at the same place as everyone else's: in the frames the obstacle model doubts. A person half hidden by a pallet, a mannequin, a reflection in a polished floor. Those frames come back to a person, the corrections retrain the obstacle model, and the new version replaces the old with no downtime.
Depth needs one more thing the box does not, which is a reference. At commissioning, and again whenever a camera is moved or a lens replaced, the robot is stopped in aisle 14 with a cone at a measured distance, and the depth the model reports for the cone is compared with the tape. A stereo pair whose baseline shifted in a collision, or a monocular model whose scale assumption no longer holds, is found by that cone before the stop rule finds out the harder way. The 95%+ navigation reliability that a fleet holds in a changing warehouse comes from the cone as much as from the model.
See it on your own footage.
Start with your footageMore in Computer vision

Computer vision · 6 min read
What CLIP is and why typing a description finds the frame
A shared space for pictures and words lets an operator type a sentence and get the matching frames from a recorder. A way to look, and the boxes come after.
Rob Hickey · Oct 3, 2026

Computer vision · 6 min read
An end-to-end object detection workflow starts with decisions, and the first frame is labeled last
Tag, box, mask or keypoints. A class list with occlusion and size rules. A metric and an acceptance sentence. Then, and only then, the first camera's frames.
Ayman Quadir · Oct 3, 2026

Computer vision · 7 min read
Image recognition AI, explained on the cameras a site already owns
A tag, a box, a mask or keypoints, on a substation camera, a line camera and an aisle camera. How a model gets there, and what its score does not promise.
Esdras Ntuyenabo · Oct 3, 2026