Skip to content
LexDataLexData
PlatformIndustriesCustomers
DocsThe Field GuideBlogWhy models drift
AboutCareersSecurityContact
Log inStart now
← All posts

Labeling · 7 min read

Occlusion in computer vision, training for the half the camera can see

A forklift behind a rack, a shopper behind a display, a part under the robot arm. On a fixed camera the partial view is the normal one, and the labels say so.

Summary

This post takes three fixed cameras, a warehouse aisle, a store floor and a robot cell, where the object is half hidden on most frames, and explains why a model trained on clear views fails on the partial ones. It sets out labeling the visible part as the policy, keeping the occluded frames in the training set rather than filtering them out, and telling a blocked view from a hard one before retraining. It is for teams whose model works on the demo frames and misses on the live feed.

Stephen Biswas · Engineer · Oct 3, 2026

Shopper at an open dairy case, partly behind the door, rack sections and two carts boxed, from a customer store camera

The forklift on the warehouse camera is half behind the end of rack 12 for most of the time it is in view. The shopper on the store camera is behind the promotional display for the four seconds it takes to reach the dairy case, and then behind the case door. The part in the robot cell is under the arm from the moment it is picked up until the moment it is set down. On all three cameras, the clear, unobstructed view of the object, the one the demo was built on, is the rare frame.

The models were trained on the rare frame. A team picking training frames picks the ones where the object is plainly visible, because those are the ones that are easy to label and easy to agree on, and the model that results has learned the whole forklift, the whole shopper, the whole part. Shown half of any of them, it has weak evidence and a low score, or no detection at all.

The drift catalog has a page on the day something new blocks the view. This post is about the ordinary case that was there from the start: on a fixed camera, the partial view is the view.

Overfitting to the whole object is learning the wrong thing

A detector learns which arrangement of pixels goes with a label. Train it only on whole forklifts and it learns that a forklift is a mast plus a cab plus a counterweight plus two forks, together, in that arrangement. Show it a mast and a cab with the rest behind rack 12 and the arrangement is broken. The model has no rule that says a mast and a cab are enough, because no training frame ever said so.

This is overfitting in a specific and common form. The model has fit the clear view so well that it has stopped generalising to the partial one, and the clear view was the only view in the set. The labeling guide meets it at the class definition. The score on the validation split, drawn from the same clear frames, stays high. The live feed, where the forklift is behind the rack, disagrees.

The fix is in the frames, and it starts with a labeling decision.

Label the visible part, every time, as policy

On the warehouse camera the ruling was written on the first Monday: an object that is partly hidden gets a box around the part that can be seen, and the box does not extend to where the hidden part would be. A forklift with its forks behind a pallet is boxed as a mast and a cab. A shopper behind the display is boxed from the shoulders up. A part under the arm is boxed as the part of it beside the gripper.

The alternative, guessing the full extent, teaches the model to put boxes on racks and displays. The other alternative, skipping the frame because the box is a judgement call, teaches it that a partly hidden object is background, which is the failure this post opened with. Boxing the visible part is the only ruling that produces a model that finds the object it can see.

There is a threshold to write down as well. A forklift that is visible as one fork tip through a gap in the rack is, for most questions, not a detectable forklift, and a box on eight pixels of fork is a label the model cannot learn. The ruling on the warehouse camera was that an object is boxed if a person could name it from the visible part alone, and a picture of a borderline case went beside the sentence.

My own view is that this one ruling, visible part only with a named threshold, does more for a fixed-camera model than any amount of augmentation. The augmentation libraries have a dozen ways to paste a grey square over a training frame, and they are useful. None of them is as useful as a few hundred real frames of the forklift behind the real rack, labeled the same way every time.

Keep the occluded frames in the set instead of filtering them out

The second decision is which frames train. A labeling batch picked by hand from the recorder is biased toward clear frames because clear frames are quicker, and a batch sampled at a fixed interval is honest about what the camera sees. On the warehouse camera, sampled every couple of seconds across the Tuesday day shift, most frames with a forklift in them had it partly behind something. That proportion is the one the training set should carry.

Somebody usually proposes filtering the hard frames out to make the numbers cleaner. The numbers do get cleaner, and the model gets worse on the feed by exactly the proportion that was removed.

The occluded frames are also where the labels most need the review pass. Two labelers will draw the visible part of a half-hidden shopper differently, one stopping at the display's edge and one including a sliver of the display, and the reviewer checking every label against the ruling is what keeps the boxes consistent enough to learn from. Frames from the store camera at the dairy case, where the shopper is behind the door, were the batch with the most rejections in week one, and the most useful batch in the set by week three.

A blocked view is a maintenance ticket and a hard view is a labeling job

The live model doubts a frame and it comes back to a person, and the person has to make a distinction the model cannot. Is the forklift hard to see, or is it gone? If a new stack of pallets has been left at the end of rack 12 and the forklift now passes entirely behind it, no label makes the model see through pallets. The occlusion page is blunt about this: check the frame before you retrain, because retraining around a blocked view teaches the model to stop expecting the thing.

The tell is spatial. Detections fall in one region of the frame and hold everywhere else, and they fall on a date that matches something being moved. That is a ticket for whoever moved the pallets. Detections that are low across the frame, on objects that are partly visible in the ordinary way, are the labeling job this post is about, and the doubted frames are the batch.

In the robot cell the distinction had a third answer. The part under the arm was never going to be visible from the cell camera during the pick, so the question was moved. The model checks the part on the fixture before the pick and on the outfeed after, and the frames from the pick itself are not asked about. Sometimes the fix for occlusion is a different question.

The model that watches the feed learned the feed's view

LexData takes the warehouse model through its whole life. You type what to look for, Lexi puts a box on every frame, and a person checks each label before anything trains on it. The model then watches the aisle camera, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime.

The frames it is unsure of are, on a fixed camera, almost always the half-hidden ones, and the correction on each is a box on the visible part. That is the policy applied one more time, by the person the doubt came back to, and it is the reason the second version is better behind rack 12 than the first.

See it on your own footage.

Start with your footage

More in Labeling

Labeling · 7 min read

Aerial dataset augmentation for drone frames where there is no up

A tower seen straight down has no top or bottom, so rotations and both flips are safe. Scale for altitude, brightness for sun, move every box with its pixels.

Esdras Ntuyenabo · Oct 3, 2026

Labeling · 6 min read

A collaborative data annotation workflow run as a pipeline

Batch the bottling line's frames by camera and shift, assign so nothing is boxed twice, attach the guideline, review every label, train on the approved set.

Rajiya Sultana · Oct 3, 2026

Labeling · 7 min read

Dataset health check for computer vision, what to look at before anything trains

A scratch dataset where every scratch sits in the centre of the frame will train a model that looks in the centre. Five counts to read before the first epoch.

Rajiya Sultana · Oct 3, 2026

LexData
LexData

Product

  • Platform
  • Industries
  • Use cases

Resources

  • Docs
  • The Field Guide
  • Blog
  • Why models drift

Industries

  • Energy & utilities
  • Oil & gas
  • Agriculture
  • Manufacturing
  • Insurance
  • Retail
  • Robotics

Company

  • About
  • Customers
  • Careers
  • Contact

Trust

  • Security
  • Privacy
  • Terms

Stay updated

What we learn running vision models in production.

See everything.
Miss nothing.

Stay updated

What we learn running vision models in production.

Terms of use & Privacy policy

© 2026 LexData Labs · All rights reserved