Skip to content
LexDataLexData
PlatformIndustriesCustomers
DocsThe Field GuideBlogWhy models drift
AboutCareersSecurityContact
Log inStart now
← All posts

Computer vision · 6 min read

Overfitting in computer vision, or how a model learns the warehouse instead of the pallet

A pallet detector that shines at its own warehouse and fails at the next one learned the lighting, and the training versus validation gap is the tell.

Summary

This post explains overfitting through a pallet detector on a warehouse camera that learned the building's lighting and racking rather than the pallet. It shows the training versus validation gap as the tell, the second site as the moment the problem stops being theoretical, and frames from several cameras and sites as the fix that augmentation only patches. It is for fleet and operations teams whose first model worked at the pilot site.

Rob Hickey · Chief AI Officer · Oct 4, 2026

Warehouse aisle with racked pallets from a high camera, forklift and pallet boxed, generated scene with detections from our model

The pallet detector on the high camera over aisle 12 scored almost perfectly on the frames it was tested on. For the first month it earned that score: every pallet on the racking boxed, every empty slot found, the count on the shift report matching the count on the stock system. Then the same model went to the second warehouse, and on the first morning it missed a third of the pallets and boxed a fire extinguisher.

The model had not broken. It had learned the first warehouse, and it turned out the pallet was only a small part of what it learned.

Overfitting is a model that memorised its training frames

A model overfits when it fits the particular frames it was trained on more closely than it fits the thing those frames were meant to teach. The training score keeps climbing, because the model is getting better at those exact frames, while the score on frames it has never seen stalls or falls. The model has stopped learning the pattern and started memorising the examples.

In vision the memorised part is rarely the object. It is everything around the object that happened to be constant in the training set: the background, the light, the position in the frame, the camera height. Those are constant on a fixed camera by construction, which is why overfitting on a fixed camera like the one over aisle 12 is the normal outcome rather than a surprise.

The warehouse's lighting is easier to learn than the pallet

On aisle 12 the pallets sat under the same sodium lights on the same blue racking, at the same distance from the same camera, with the skylight throwing the same bar of daylight across the floor at 2 pm. A pallet in that setting has a dozen cues that are not the pallet: a blue edge above it, a yellow cast on it, a bright patch beside it in the afternoon. The model learned the cues because they were always there and easier to find than the timber.

There was also a stencil. The first warehouse's pallets carried the company's stencil on the leading board, and the model learned it as part of what a pallet is. The second warehouse's pallets were plain, and the model had no idea what to make of them.

The gap between training and validation is the tell

The tell is a gap that opens between the training score and the validation score as training goes on. Both rise at first; then the validation score stops while the training score continues, and the distance between them is the amount of memorising the model is doing. Stop training at the point where validation stops improving and the model is as good as those frames can make it.

The gap only shows if the validation frames are genuinely unseen. A random tenth of the frames, held out from the same footage, is not unseen. Adjacent frames from a fixed camera are near duplicates of each other, so a frame in the validation set has a near twin in the training set and the score is flattered. The lesson on how much footage you need makes the point that thirty frames a second is one scene sampled thirty times, and it applies to validation as much as to training. Hold out a different day on aisle 12, and better still a different camera.

The new site is the moment the problem stops being theoretical

The second warehouse had LED lights, grey racking, a camera two metres lower and no stencil. Every cue the model had leaned on was gone at once, and it was left with what it had learned about pallets themselves, which was less than anyone had assumed. The drift catalog describes this as a new site came online, and its note from the floor is the honest one: the model is doing exactly what a model trained on one site does.

Overfitting and site drift are the same fact seen from two ends. In training it is a gap on a chart; in operations it is a fleet rollout that stalls at the second building.

Frames from several cameras and sites are the fix

The fix is variety in the training set, and the unit of variety is distinct conditions rather than frames. A few hundred checked frames from each of three or four cameras, at different heights, under sodium and LED lights, on different racking, teach the model that the pallet is the constant and the surroundings are not. Two warehouses in the training set generalise to a third far better than one warehouse with twice the frames.

Augmentation, the random brightness shifts and crops applied during training, helps at the margin. It teaches the model that a pallet under slightly different light is still a pallet, which is true, and it cannot teach the model about grey racking it has never seen. A smaller model, early stopping and dropout are the same kind of tool: they limit how much the model can memorise, and none of them adds a condition the frames do not contain.

My own view is that augmentation is oversold as a fix for overfitting on fixed cameras. Every hour spent tuning augmentation on a one-camera set would have been better spent getting a few hundred frames from a second camera.

A model that memorises a station is fine until the station changes

There is a case where a tightly fitted model is what you want. A single inspection station that presents one part per frame, under one light, from one camera, will never see a second warehouse, and a model fitted closely to that station is accurate there for as long as the station stays the same. The trouble is that stations change: a lamp is replaced, a camera is nudged, a new part variant arrives, and the tightly fitted model has no margin.

LexData takes the pallet model through its whole life. You type what to look for, Lexi puts a box on every frame, and a person checks each label before anything trains on it. The model then watches the cameras the warehouse already has, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime.

At the second warehouse that is how the first morning ended. The missed pallets came back as doubted frames, the corrections crossed the project's threshold within the week, and the version that replaced the old one had grey racking in its training set. The fire extinguisher never got boxed again.

See it on your own footage.

Start with your footage

More in Computer vision

Computer vision · 7 min read

Five computer vision applications in production, on the cameras a site already owns

Detect, count, inspect, track and read are the five jobs a fixed camera can be given, and each is a different question with a different label behind it.

Ayman Quadir · Oct 4, 2026

Computer vision · 6 min read

Computer vision projects worth building on the cameras you already have

A plant, a utility, a grower, a store and a warehouse each have a project that starts on an existing camera, produces a decision, and has someone to act on it.

Ayman Quadir · Oct 4, 2026

Computer vision · 6 min read

How to choose an object detection model architecture for a camera on your own site

Start from the decision the plant has to make, then the box beside the recorder the model has to fit, and only then the family. The benchmark comes last.

Andreas Ohrvall · Oct 4, 2026

LexData
LexData

Product

  • Platform
  • Industries
  • Use cases

Resources

  • Docs
  • The Field Guide
  • Blog
  • Why models drift

Industries

  • Energy & utilities
  • Oil & gas
  • Agriculture
  • Manufacturing
  • Insurance
  • Retail
  • Robotics

Company

  • About
  • Customers
  • Careers
  • Contact

Trust

  • Security
  • Privacy
  • Terms

Stay updated

What we learn running vision models in production.

See everything.
Miss nothing.

Stay updated

What we learn running vision models in production.

Terms of use & Privacy policy

© 2026 LexData Labs · All rights reserved