Operations · 7 min read
OCR on video, reading a container number from a moving truck
A yard camera tracks each container as it passes, reads the ID panel on the sampled frames, and takes a vote. One blurry frame never becomes a wrong record.
Summary
This post follows a container through a yard gate camera and shows why reading its number is a pipeline rather than a single read: a detector finds the ID panel, a track carries the container across the frames it is visible in, the reader runs on the sampled frames, and a vote with the check digit decides the record. It concludes that a doubtful read goes to a person with the frames rather than into the log. It is for yard and depot teams replacing a clipboard.
Rajiya Sultana · Engineering Manager · Oct 1, 2026

Yard camera over a fence line with a vehicle boxed, generated scene with detections from our model
The gate at the container yard on the Humber has a camera on a pole, and a clerk in the weighbridge hut with a clipboard. Every truck that comes through carries a container with an eleven-character number painted on its side and its end, and the clerk copies the number as the truck slows for the barrier. On a busy morning the trucks do not slow enough, the number on the side is half hidden by the cab, and the clerk writes down what she saw and checks it against the driver's paperwork later.
The camera has been recording the same containers for years. Reading the number off the recording is a different job from reading it off a photograph, and the difference is what this post is about.
Running the reader on every frame is the wrong default
The obvious pipeline runs a text reader on every frame the camera produces. It is also the slowest and the least accurate, because most frames of a passing truck contain no readable number at all. The cab is in the way, the container is at the edge of the frame, or the number is at an angle that turns the characters into slivers, and at 8 am the trucks are nose to tail. A reader fed every frame spends its time on those, produces a stream of partial reads, and somebody then has to decide which of the partial reads was the truck.
The better pipeline reads rarely and deliberately. It finds the number first, follows the container while it is in view, reads the frames where the number is actually legible, and combines those reads into one answer per container. Each step exists to make the next one cheaper.
Object detection finds the ID panel on the container side
The first step is a detector, and its class is the number panel, the rectangle of painted characters on the side and end of the container, rather than the characters themselves. Object detection on the yard camera's frames puts a box on that panel wherever it is, at whatever angle, on a blue container or a rusted one. It has learned what the panel looks like from this pole rather than what the letter A looks like.
The box is the crop the reader will get. It also tells the pipeline whether there is anything to read on this frame. A frame with no panel box is skipped, which on the gate camera is most frames, and the reader never sees them.
Labeling the panels is a short job: you type "container number panel", Lexi proposes the boxes, and a person checks that each box takes the whole panel with a margin rather than clipping the check digit at the end.
The track carries the container across the frames it is visible in
A truck at the barrier is in view for many sampled frames, and the container's panel is boxed on some of them. Instance tracking joins those boxes into one track. The pipeline then knows that the panel it saw at the edge of the frame, the one it saw square-on as the truck slowed for the barrier at 8:14, and the one it saw disappearing behind the arm were the same container.
Without the track, three reads of one container are three containers. With it, they are three attempts at one number, which is what the vote needs.
The track also settles when the record is written. A container gets one record, when its track ends, rather than one per frame, and the record carries the time it entered the frame and the time it left, which the weighbridge clerk's clipboard never had.
The read is a vote across frames rather than one frame's guess
Each legible frame in the track gets a read, and the reads disagree. The frame at 8:14 says the fourth character is a B; the next sample, square-on, says it is an eight; the one after, in shadow, says B again. A single-frame pipeline would have logged whichever frame it happened to run on. A vote across the track takes the character each position agrees on most often, weighted toward the frames where the panel was largest and squarest, and settles on the eight.
The vote is also where the container's own check digit earns its place. The last character of the number is computed from the ten before it, so a read whose check digit does not match is wrong somewhere, and it is dropped from the vote before it can outvote a correct read. Most wrong reads fail the check. The ones that pass it and are still wrong are rare enough to be worth a person's time.
My own view is that a pipeline reading anything with a check digit, container numbers, vehicle identification numbers, some barcodes, should be reading the check before it does anything clever, because the check is free and the clever part is not.
The hard frames are blur, glare and a cab in the way
The frames that produce the wrong reads sort into piles. Motion blur, when the truck did not slow: the characters smear along the direction of travel and a three becomes an eight. Glare off a wet container at 4 pm: a stripe of white across the panel takes out two characters. The cab, on the side panel: the front of the number is hidden until the truck has passed, and only the end panel shows the whole thing. Low light at the start of the night shift, when the yard lamps have not yet come on.
Each pile has a fix that is mostly on the yard side. A speed hump before the camera slows the trucks. A hood on the housing cuts the glare. The end panel is the one to prefer, because nothing hides it. And the frames the detector was unsure of come back to a person, which is where the night-shift frames go until the model has seen enough of them.
The hazard zone intrusion detection use case runs on the same gate camera in most yards, a zone drawn once in the frame and a person box crossing it, and it shares the camera's failures: a moved pole moves the zone and the panel alike.
A doubtful read goes to a person with the frames
When the vote is close, or the check digit fails on every read, the pipeline should not log its best guess. It holds the record as unresolved and sends the frames, the three or four legible ones with the panel boxed, to a person. On the yard that is the weighbridge clerk, who now looks at three crops on a screen instead of at a truck through a window, and types the number in a few seconds. Her verdict is the record, and it is also a label the next version of the detector learns from.
An alert is a rule written as a sentence, with a severity and a cooldown, approved before it goes live: a container number that could not be resolved at the gate, routine, sent to the weighbridge screen with the frames attached. The monitoring and alerts guide covers writing it. A container number that resolves cleanly needs no alert at all, and most do.
LexData takes the gate model through its whole life. You type what to look for, Lexi puts a box on every number panel in every frame, and a person checks each label before anything trains on it. The model then watches the gate camera the yard already has, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime.
The clerk reads only the last four characters of a number when she talks to a driver on the radio, and the drivers answer with the same four. The full number is on the screen now, and the radio habit has not changed.
See it on your own footage.
Start with your footageMore in Operations

Operations · 7 min read
AGPL-3.0 licensing risk for computer vision teams serving a model
A camera streaming to a served detector is the network interaction the licence was written for. What a legal review will ask, and why to pick the weights first.
Ayman Quadir · Oct 1, 2026

Operations · 7 min read
Cloud vs owned GPU inference for computer vision, worked out per camera hour
A plant on three shifts and a retailer with cameras spread across stores get different answers from one sum, and footage leaving the building is a cost too.
Ayman Quadir · Oct 1, 2026

Operations · 7 min read
Computer vision heatmaps drawn from the aisle cameras a store already has
Footpoints from every tracked box, aggregated over a day and mapped onto the floor plan, show where footfall goes. A camera nudged in cleaning shifts the map.
Rajiya Sultana · Oct 1, 2026