Skip to content
LexDataLexData
PlatformIndustriesCustomers
DocsThe Field GuideBlogWhy models drift
AboutCareersSecurityContact
Log inStart now
← All posts

Industries · 6 min read

AI in robotics after the robot ships, what the warehouse cameras keep learning

The forward camera boxed pallets and people well at the pilot site. Then the racking moved, and the edge cases the planner never saw came back for review.

Summary

This post follows a warehouse robot's forward camera from the day it ships, through a racking change, to the second site, covering boxes on pallets, forklifts and people, pose estimation for the person about to step out, and edge cases mined from the model's own doubt. It concludes that the runner beside the recorder is where this model belongs and that each new site is a new distribution. It is for robotics and fleet teams running mobile robots in changing buildings.

Andreas Ohrvall · CTO · Sep 29, 2026

Warehouse aisle from a high camera, forklift and racked pallets boxed, generated scene with detections from our model

The robot that pulls pallets between receiving and the pick faces shipped in May with a forward camera and a perception model that had seen the pilot warehouse from every aisle. On the second Monday of June the night shift moved two rows of racking to make room for a new mezzanine stair. The robot's map was updated by 6 am. Its model was not, because nobody thinks of a model as something that needs updating when the building changes.

By 8 am the robot had stopped four times in aisle 12 for a shadow under the new stair that looked, to a model trained on the old floor, like the edge of a pallet.

The forward camera boxes pallets, forklifts and people on every frame

The model's job on each frame is a short list: pallets, forklifts, people, and the free floor between them. Each is a box, and the box is what the planner consumes, with the bottom edge of a person's box telling it where that person's feet are on the floor. There is no measurement of extent here, so boxes do the job and masks would add labeling cost without changing a single stop.

What makes the list hard is the fourth item. Free floor is the absence of everything else, and a model that has learned what a pallet looks like from one warehouse has learned that warehouse's floor as the background. The new stair's shadow in aisle 12 is neither pallet nor floor as the model met them, and the planner's cautious answer to an unfamiliar patch is to stop.

Pose estimation tells the planner where a person is about to step

A box on a person says where they are. It does not say which way they are facing or whether their weight has shifted onto the foot that is about to leave the aisle. Pose estimation adds keypoints inside the box, shoulders, hips, knees, and the line through the hips gives the planner a heading before the person has moved. On a floor where forklifts and robots share aisles with pickers, that heading is the difference between a robot that slows early and one that brakes late.

The labeling is heavier than boxes and it is worth confining to the class it helps. Keypoints on people, boxes on everything else, and the keypoints checked by a person who knows what a picker turning to lift looks like from a camera at knee height in aisle 12. Lexi proposes the keypoints on every frame and the person confirms or moves them, which on a forward camera mostly means the frames where a picker is half behind a pallet.

The layout changed and the model was fitted to the old one

Warehouses are rearranged more often than any planner is told. Seasonal stock moves the pick faces, a new customer adds a staging zone, the mezzanine stair goes in. Each change is a small shift in what the forward camera sees, and the model has weak evidence for exactly the patch of the frame that changed.

The signal is the correction rate. The robot stops, the frame it stopped on comes back for review, and a person marks the stair shadow as free floor. Four stops in aisle 12 on one morning is four corrections in the same direction, and that is the signal that the model needs to see the new floor. The corrections retrain it; when they cross the project's threshold a new version is trained on the frames the person already judged, and the robot in aisle 12 the following week has seen the stair.

Edge cases are mined from the robot's own doubt

The frames that matter after shipping cannot be collected by sampling, because they are rare by definition. A reflective spill on the floor, a pallet wrapped in black film that reads as a hole, a picker kneeling behind a low cart. The navigation edge cases use case makes the point that these have to be mined from fleet footage by looking for the model's own uncertainty, which is a pipeline problem before it is a labeling one.

That is what the review queue is. The model returns the frames it doubts, a person rules on them, and the rulings are the training set for the situations the pilot warehouse never produced. The fleet's own operation is the collection strategy, and the reliability figure in our robotics work, 95%+ navigation reliability in changing environments, is held that way rather than by a larger first dataset.

An aside from the floor: the night shift names the robots, and the one that stopped four times under the stair was known for a week as the one that was afraid of the dark.

The second warehouse is a new site and the fleet average hides it

The pilot site's model goes to the second building with the second robot. The second building has different racking vintages, a lower camera height on the older trucks, sodium lighting in the far bays, and a floor painted a different grey. The model's detections at the new site differ from the pilot's from the first Monday rather than degrading into it, and the drift catalog calls this a new site came online.

Compare each site against its own history. A fleet-wide average of stops per shift will look fine, because the pilot site carries the new one. Label a window of frames from the second building and fold it in before the robot goes live there, and the model that has seen two buildings will reach the third with more to go on than the one that saw the first very thoroughly.

The model runs beside the recorder and only the doubted frames leave

LexData takes the perception model through its whole life. You type what to look for, Lexi puts a box on every frame, and a person checks each label before anything trains on it. The model then watches the robot's camera on a runner beside the recorder, so the frames stay in the building and only the frames the model doubted leave for review. The corrections retrain it, and the new version replaces the old one with no downtime. The deployment guide covers the runner, the export presets, and why the enclosure matters more than the benchmark.

My own view, which comes from the architecture side, is that a planner's perception model should never depend on a round trip out of the building, and that the cloud's job in a robot fleet is the review and the retraining rather than the inference. The robotics work we do is built on that split: the model runs where the robot is, and what leaves the site is a question about a frame.

The mezzanine stair is finished by July. The racking will move again in October for the peak season stock, and the robot in aisle 12 will stop under something new, and the frame will come back.

See it on your own footage.

Start with your footage

More in Industries

Industries · 7 min read

Aerial fire detection from a drone patrol, smoke before the flame reaches the line

On a right-of-way patrol, smoke is a few dozen grey pixels that look like haze. Two boxed classes, an alert with a cooldown, and the bad-weather days kept.

Rob Hickey · Sep 29, 2026

Industries · 6 min read

Automated sorting with computer vision, from the camera over the conveyor to the diverter

A box on every apple, a grade from the box, and an air jet that acts on it before the belt moves on. The new cultivar is when the model needs the graders again.

Rob Hickey · Sep 29, 2026

Industries · 6 min read

Infrastructure asset management with computer vision from a truck at road speed

Signs, guardrails and poles pass the truck camera in a fraction of a second. The model finds them, the register says what should be there, a crew gets the gap.

Rob Hickey · Sep 29, 2026

LexData
LexData

Product

  • Platform
  • Industries
  • Use cases

Resources

  • Docs
  • The Field Guide
  • Blog
  • Why models drift

Industries

  • Energy & utilities
  • Oil & gas
  • Agriculture
  • Manufacturing
  • Insurance
  • Retail
  • Robotics

Company

  • About
  • Customers
  • Careers
  • Contact

Trust

  • Security
  • Privacy
  • Terms

Stay updated

What we learn running vision models in production.

See everything.
Miss nothing.

Stay updated

What we learn running vision models in production.

Terms of use & Privacy policy

© 2026 LexData Labs · All rights reserved