Skip to content
LexDataLexData
PlatformIndustriesCustomers
DocsThe Field GuideBlogWhy models drift
AboutCareersSecurityContact
Log inStart now
← All posts

Labeling · 7 min read

Stop sign detection edge cases and the long tail of a simple object

A stop sign is the easiest class there is until it is snowed on, held by a crossing guard or printed on a bus. The tail is covered one doubted frame at a time.

Summary

This post uses the stop sign, the simplest object a detector is ever asked to find, to show how long the tail of a simple class really is: snow, glare, a hand-held sign, a sign on the side of a bus, a sign seen from behind. It explains why no training set anticipates the tail and argues that the frames a live model doubts, sent back for review, are how the tail gets covered. It is for anyone whose simple class is failing in the field.

Stephen Biswas · Engineer · Oct 3, 2026

Dashcam frame with traffic signs, pedestrians and vehicles boxed, from a customer perception run

A stop sign is a red octagon with four white letters on it. There is no simpler object in the world to describe and, on a sunny afternoon on a suburban street, no simpler object to detect. A model trained on a few thousand of them from a dashcam finds every one on the validation set, and the team moves on to harder things.

Then it is January. The sign at the end of the street has a crust of snow on its top edge and the octagon is a red semicircle. The low sun is directly behind the next one and the sign is a black shape with a halo.

The crossing guard at the school holds a hand-held one, at chest height, moving. A school bus pulls out with its own stop sign folded flat against the side, and then swings it open. And at the junction the model has passed a hundred times, there is a stop sign facing the other way, seen from behind as a grey octagon on a pole, which is not a stop sign for this car at all.

The simple object has a tail, and the tail is where the model lives once it is deployed.

Every simple object detection class has the same long tail

Swap the sign for a hard hat on a site camera, a pallet in a warehouse aisle, the gate on a yard fence at dock 2, and the shape of the story does not change. The class is easy to describe and easy to find in the common case, and the common case is nearly all of the training set. The rest of the world, the conditions that happen a few times a month, is the tail, and the tail is long because it is made of every combination of weather, angle, occlusion and human improvisation the camera will ever see.

The hard hat is on the ground beside the worker. The pallet is stood on end against the rack. The gate is open, at an angle the training frames never showed, with a truck halfway through it. A detection model asked to find any of these has to answer a question the training set never posed, and it answers from the nearest thing it has seen, which is often wrong and often confident.

My own view is that the length of the tail is the single most underestimated quantity in a vision project. Teams estimate it from the training set, which is the one place it does not appear.

No training set anticipates the tail

The instinct is to collect more frames before launch, and it helps less than it should. Frames collected before launch are collected from the conditions that exist before launch, by people who choose the frames, and people choose the frames they can imagine. A team building the stop sign set can imagine snow and can go and find some. They cannot imagine the bus, the guard, the sign printed on a delivery van's rear door as an advert, the sign reflected in a shop window, until each one happens.

That is the whole difficulty. The tail is defined as the cases nobody anticipated. A dataset that anticipates it is a contradiction, and a team that believes it has covered the tail has covered the part of it they could think of on a Tuesday.

What a training set can do is span the conditions it can reach. The drift lesson separates the ways a live model diverges from its training, and the tail is the one that arrives one frame at a time rather than as a season or a moved camera. It cannot be collected in advance. It can be collected as it happens.

The doubted frames are the tail arriving

A model watching the street has, on every frame, a detection and a score for it. On the common case the score is high and the detection is right. On the snow-crusted semicircle the score is low, or there is no detection at all where a person would expect one, and that frame is exactly the frame the training set lacked.

So the frame comes back. The model returns what it is unsure of to a person. The person draws the box on the snowy sign from January and confirms it is a stop sign, or boxes the bus sign and confirms it is one too, or looks at the grey octagon from behind and marks that it is not. Each of those is a label from the tail, chosen by the one process that can find the tail, which is the live model failing on it.

A correction on a frame the model got wrong teaches it what it does not know. A label on a frame chosen at random teaches it something it may already know. The doubted frames are the highest-value labels a project ever gets, and they arrive for free, in the order the world produces them.

The person reviewing them is the check that keeps the tail honest. The sign seen from behind is the case that matters here: a model that is corrected to call every grey octagon a stop sign has learned the tail wrong, and the reviewer who marks it as not-a-sign is the reason it learns the tail right.

The test set is curated from the tail, separately

Corrections retrain the model. The other half of the job is knowing whether the retrain helped, and the validation split drawn from the original training set cannot say, because it has no tail in it. A model can improve on every snowy sign in the world and score exactly the same on a sunny-afternoon validation set.

The test set for a simple class has to be built on purpose, from the tail, and kept out of training. One frame of each tail case the review queue has produced since January, the snowed-on sign and the glare and the hand-held one and the bus and the one seen from behind, with a person's label on it, held aside. Every new version is scored on that set, per case, and the cases where it still fails are the cases the next batch of corrections should weight toward. When a new tail case appears in review, one example goes into the test set before the rest go into training.

That set grows slowly and is the most valuable file in the project. Somebody on every project ends up naming it something like "the hard ones".

The tail never ends and the loop is built for that

There is no version of the stop sign model that has seen the last edge case. Next winter there will be a new one, a sign wrapped in a plastic bag by a road crew, a sign knocked to face the sky. The question is never whether the model will meet a frame it cannot read. It is what happens to that frame.

LexData takes the model through its whole life on that basis. You type what to look for, Lexi puts a box on every frame, and a person checks each label before anything trains on it. The model then watches the cameras you already have, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime. The snowy sign that came back in January is in the training set by February, and the model that meets it next winter has seen one before.

See it on your own footage.

Start with your footage

More in Labeling

Labeling · 7 min read

Aerial dataset augmentation for drone frames where there is no up

A tower seen straight down has no top or bottom, so rotations and both flips are safe. Scale for altitude, brightness for sun, move every box with its pixels.

Esdras Ntuyenabo · Oct 3, 2026

Labeling · 6 min read

A collaborative data annotation workflow run as a pipeline

Batch the bottling line's frames by camera and shift, assign so nothing is boxed twice, attach the guideline, review every label, train on the approved set.

Rajiya Sultana · Oct 3, 2026

Labeling · 7 min read

Dataset health check for computer vision, what to look at before anything trains

A scratch dataset where every scratch sits in the centre of the frame will train a model that looks in the centre. Five counts to read before the first epoch.

Rajiya Sultana · Oct 3, 2026

LexData
LexData

Product

  • Platform
  • Industries
  • Use cases

Resources

  • Docs
  • The Field Guide
  • Blog
  • Why models drift

Industries

  • Energy & utilities
  • Oil & gas
  • Agriculture
  • Manufacturing
  • Insurance
  • Retail
  • Robotics

Company

  • About
  • Customers
  • Careers
  • Contact

Trust

  • Security
  • Privacy
  • Terms

Stay updated

What we learn running vision models in production.

See everything.
Miss nothing.

Stay updated

What we learn running vision models in production.

Terms of use & Privacy policy

© 2026 LexData Labs · All rights reserved