Skip to content
LexDataLexData
PlatformIndustriesCustomers
DocsThe Field GuideBlogWhy models drift
AboutCareersSecurityContact
Log inStart now
← All posts

Computer vision · 7 min read

What optical character recognition is and where OCR breaks on a real camera

A packaging line camera is not a scanner. Acquisition, detection, recognition and validation on a lot-code label, and the doubtful read goes to a person.

Summary

This post takes OCR through a lot-code label on a packaging line, stage by stage: how the frame is captured, how the text is found and read, and how the read is validated against what the code is allowed to be. It concludes that a camera on a line is a different problem from a scanner page, that a checksum catches more misreads than a better model, and that the low-confidence read belongs with a person. It is for teams reading printed text from footage rather than from documents.

Finn Ellingwood · Engineer · Oct 3, 2026

Cartons with labels moving past a scanner on a packaging line, carton and label boxed, generated scene with detections from our model

The camera over packaging line 2 sees each carton for about a second as it passes the scanner, and on the side of every carton is a label with a lot code, a best-before date and a product name. The line needs the lot code read and matched against the batch that is supposed to be running. If the wrong lot goes into a pallet, the pallet gets pulled at the customer's dock, and the customer does not send it back quietly.

Reading printed text sounds like a solved problem, and on a scanned page it mostly is. On a carton moving past a camera under a strip light, with the label at a slight angle and the last hundred cartons of every ribbon printed faint, it is a different job with the same name.

The label camera sees a label the way a scanner never does

A scanner produces a flat, evenly lit, high-resolution page with black text on white. The camera on line 2 produces a frame in which the label is a small region, at an angle, partly in the shadow of the carton flap, with the text a few dozen pixels tall and motion blur along the belt's direction. Everything a scanner does for free, the pipeline on the line has to do on purpose.

That is the reason a pipeline has stages at all. Each one takes a specific fault out of the frame before the next stage sees it, and the package label inspection use case is mostly a list of those faults: glare, skew, a torn corner, a smudged digit.

Acquisition decides more than any model downstream

The first stage is the camera and its mount, and it sets the ceiling for everything after. A lot code printed at a few millimetres tall needs enough pixels per character to be read at all. The what a vision model can and cannot see lesson gives the arithmetic: a character that is six pixels tall is a smudge, whatever the model. The fixes are optical. A lens with a narrower field, a camera closer to the label, a shutter fast enough that the belt's motion does not streak the digits, a light that does not reflect off the label's gloss.

Preprocessing follows and it is specific to the line. Straightening the label from its angle, cropping to the label region, lifting the contrast on the faint prints, sharpening the blur if it is small enough to be recoverable. The same steps on a different line with different labels would be wrong, and a pipeline copied from one line to another usually fails here first.

The lot code printer's ribbon fades over the last hundred cartons of every roll, so the faint prints arrive in a predictable run at the end of each ribbon. The operator on line 2 can tell from the reads alone when a ribbon change is due.

Detection finds the text and recognition reads it

With a clean crop, the pipeline splits into two models. Detection finds where the text is: a box around the lot code, another around the date, another around the product name. Recognition takes each box and turns the pixels into characters. They fail differently. Detection misses a line of text when it is too faint or too small, or merges two lines into one box. Recognition reads the characters wrong: an eight for a B, a zero for an O, a one for an I, especially in the fonts label printers use.

Those confusions are the same on every line and they are where most of the errors live. A recognition model trained on the label printer's own font, from frames of this line's cartons, confuses them less than a general one, and a few hundred labeled crops from the line are enough to make the difference.

Validation is where a misread lot code is caught

The read comes out as a string, and the string is checked against what a lot code is allowed to be. Ten characters, letters in the first two positions, digits after, a checksum in the last. A date that exists and lies in the future. A product name that matches the batch running on the line. Every rule is a chance to catch a misread before it becomes a mismatch.

My own view is that a good validation stage catches more errors than a better recognition model, and for less money. A model that reads an eight as a B is a hard thing to fix. A checksum that rejects the result is an afternoon's work, and it rejects the error every time rather than most of the time.

Validation also decides what happens next. A string that passes every rule and matches the running batch is a good carton. A string that fails a rule is a doubtful read, and a doubtful read on a lot code is a decision a person should make, with the frame in front of them.

Vision language models read the label in context and invent the rest

A newer approach reads the whole label at once. Vision language models take the frame and a question, "what is the lot code on this label", and answer in words, using the context of the label to resolve what a character-by-character reader cannot. They are good at that: a faint digit between two clear ones is read from the pattern, and a date in an unusual layout is found without a template.

The cost is that they will answer confidently when the label is unreadable. Asked for a lot code on a carton from line 2 whose label is torn off, a character reader returns nothing and a language model returns a plausible lot code. On a packaging line that is the worst possible failure, because a plausible wrong code passes the format check. Vision language models fit best as a second opinion on the doubtful reads, where a person is looking anyway, and least well as the only reader with nobody checking.

The doubtful read goes to a person and the correction stays

The reads that fail validation, and the ones the recognition model returns with low certainty, go to a queue with the frame attached. A person looks at the crop, types the code as it actually reads, and the carton is either released or held. That typed correction is a label, and it is exactly the kind of frame the recognition model was weakest on: the faint print at the end of the ribbon, the label under the flap shadow.

LexData takes the label model through its whole life. You type what to look for, Lexi puts a box on every label in every frame, and a person checks each label before anything trains on it. The model then watches the line camera the plant already has, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime.

The alert for a lot mismatch is a rule written as a sentence, with a severity and a cooldown, approved before it goes live, and the monitoring and alerts doc describes where it lands. What arrives is the frame with the label boxed, the code the camera read, and the code the batch expected. The person with that in hand pulls one carton. The person with a mismatch count and no frame pulls the pallet.

See it on your own footage.

Start with your footage

More in Computer vision

Computer vision · 6 min read

What CLIP is and why typing a description finds the frame

A shared space for pictures and words lets an operator type a sentence and get the matching frames from a recorder. A way to look, and the boxes come after.

Rob Hickey · Oct 3, 2026

Computer vision · 6 min read

An end-to-end object detection workflow starts with decisions, and the first frame is labeled last

Tag, box, mask or keypoints. A class list with occlusion and size rules. A metric and an acceptance sentence. Then, and only then, the first camera's frames.

Ayman Quadir · Oct 3, 2026

Computer vision · 7 min read

Image recognition AI, explained on the cameras a site already owns

A tag, a box, a mask or keypoints, on a substation camera, a line camera and an aisle camera. How a model gets there, and what its score does not promise.

Esdras Ntuyenabo · Oct 3, 2026

LexData
LexData

Product

  • Platform
  • Industries
  • Use cases

Resources

  • Docs
  • The Field Guide
  • Blog
  • Why models drift

Industries

  • Energy & utilities
  • Oil & gas
  • Agriculture
  • Manufacturing
  • Insurance
  • Retail
  • Robotics

Company

  • About
  • Customers
  • Careers
  • Contact

Trust

  • Security
  • Privacy
  • Terms

Stay updated

What we learn running vision models in production.

See everything.
Miss nothing.

Stay updated

What we learn running vision models in production.

Terms of use & Privacy policy

© 2026 LexData Labs · All rights reserved