Skip to content
LexDataLexData
PlatformIndustriesCustomers
DocsThe Field GuideBlogWhy models drift
AboutCareersSecurityContact
Log inStart now
← All posts

Operations · 7 min read

Ten questions that tell you whether a vision AI pilot will ever leave the pilot line

A pilot with no owner, no definition of done and no path for the frames it gets wrong stays a pilot. Ten questions a plant committee can score in one sitting.

Summary

This post gives a plant operating committee ten questions about the system around a vision model, from a written definition of done to the cost of the next plant, and asks for an honest score on each. It concludes that the low number is the useful one, and that the questions about the frames the model gets wrong decide whether the pilot ever scales. It is for the people who fund vision projects at manufacturers.

Ayman Quadir · Head of Product · Sep 27, 2026

Stamping line inspection station, conveyor and press boxed, generated scene with detections from our model

The surface defect pilot on line 2 has been a pilot for fourteen months. It runs. The camera over the inspection station finds scratches and pitting on the stamped panels, the inspector on the shift agrees with it most of the time, and the slide in the quarterly review says the accuracy is good. Nobody can say what would have to be true for it to go to line 5, or to the sister plant, and so it does not.

The model is rarely the reason a pilot stays a pilot. The system around it is, and the system can be examined with ten questions a plant's operating committee can answer in one sitting, scored honestly, from zero to two each. The score is less useful than the argument over each one.

Two questions about who owns it and what done means

Does a written definition of done for a production deployment exist? Written means a document, with an accuracy target per defect class from the line's own frames, an alert path, a review cadence and a named person who signs it. A pilot with no such document cannot finish, because finishing has not been defined, and fourteen months is what that looks like.

Is every live use case owned by a named person? A named person, with the line in their remit, who is asked about it in the weekly meeting. A committee is not an owner. A vendor is not an owner. The question the pilot on line 2 could not answer was who would be blamed if it was switched off, and a pilot nobody would be blamed for losing is one nobody will fight to scale.

Of the ten, the owner is the question I weight most. A named owner with a poor model will get a better model. A good model with no owner is a slide.

The word pilot comes from ships, where it names the person who takes a vessel into a harbour they know and steps off once it is docked. A pilot that never steps off was never that kind of pilot.

Two questions about whether the model has met the real line

Was the model tested on frames from the production line before it was called ready? From the camera at its mounted height, in the light of the night shift as well as the day, on the panels the line actually runs and the coolant mist that sits over the station by the end of a shift. A model evaluated on a held-out slice of its own training set has been tested on the frames it was built from. The surface defect detection use case describes the same gap: the pilot station is clean and the line is not.

Can a second plant pick the solution up without a new project? The sister plant has a different press, a different camera height and different lighting, and a model trained on line 2 will do visibly worse there on day one. That is expected, and the drift catalog files it under a new site came online: the fix is a window of labeled frames from the new line folded in before go-live. The question is whether that fold-in is a known procedure with a known cost, or a fresh statement of work. If it is the second, every plant is a first plant.

Three questions about what happens to the frames it gets wrong

Is there a path for the frames the model gets wrong to reach the training set? The inspector on line 2 sees a false pass on Tuesday, a scratch the model missed under the mist. What happens to that frame? On most pilots it goes nowhere. The inspector remembers it, mentions it in the review, and it is gone. A production system has a place for that frame to be flagged, corrected and queued.

Is there a review cadence for the doubted frames? The model returns frames it is unsure of. Someone looks at them, on a schedule, and their corrections are counted. If the review queue is opened when someone remembers, the corrections arrive in bursts and the retrain never has enough to work with.

Are the labeling standards written down and enforced? What counts as a scratch, at what length, on which surface; whether a box hugs the defect or the panel; who checks. A model trained on labels drawn by three inspectors with three private definitions of pitting has learned an average nobody agreed to. On the lines we run this is a written guideline, and every label passes a QA check against it before anything trains, which is how the manufacturing work behind our numbers holds 99%+ accuracy in production.

Two questions about monitoring and retraining without a restart

Does anything alert when the model fails in production? A model on line 2 that quietly stops finding pitting on the amber-coated panels produces no error. What it produces is fewer detections on one line than that line's own history, more doubted frames, and an inspector overriding it more often in one direction. The correction rate is the signal, and something has to watch it and tell a person. The alert written as a sentence is one shape of that; a weekly number on a wall is another. No number on any wall is a zero.

Can the model be retrained without restarting the project? On the pilot, retraining meant the vendor, a quote and six weeks. On a production system, corrections cross a threshold, a new version trains on the original set plus the corrections, it is checked per defect class against the current version, and it replaces the old one with no downtime. That is the loop, and it is the difference between a model that is maintained and a model that is re-bought.

One question about the cost of the next plant

Is the cost of each new deployment known, and falling? Line 2 cost what it cost. Line 5 should cost less, because the classes, the guideline, the alert rules and the review cadence exist. The fifth should cost a window of frames and a week. If nobody can put a number on the second deployment, the first one has not produced a capability, only a result.

A committee that reaches this question with a high score on the other nine will usually find the answer is yes without having measured it. A committee with a low score elsewhere will find the cost of the second plant equals the cost of the first, and that is the honest diagnosis of a pilot that will not scale.

A vision AI pilot scales when the system around the model exists

Score the ten and add them up. A total in the teens, with the owner and the definition of done at two, is a programme that will scale with ordinary effort. A total under ten is a pilot, and the low scores say exactly where the missing system is: usually questions five through seven, the path from a wrong frame to the training set, because those are the questions nobody thought to ask a model vendor.

The score is not a judgement on the model on line 2, which may be excellent. It is a judgement on whether anything around it would notice if it stopped being excellent, and whether anyone could fix it if it did.

LexData takes the panel model through its whole life. You type what to look for, Lexi puts a box on every frame, and a person checks each label before anything trains on it. The model then watches the camera over the inspection station, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime. Questions four through nine are what that loop is, which is why a committee scoring itself against them is also scoring the platform it would have to build or buy.

See it on your own footage.

Start with your footage

More in Operations

Operations · 7 min read

Active learning for computer vision on a line camera that never stops

The weld camera runs three shifts. The model returns the frames it doubts, the inspector corrects them, and past the threshold a new version trains and ships.

Rob Hickey · Sep 27, 2026

Operations · 7 min read

Camera focus measurement for a fixed camera that slowly goes soft

A lens loosened by vibration fails over weeks, and the model suffers before anyone sees blur. A sharpness score against the camera's own history catches it.

Rajiya Sultana · Sep 27, 2026

Operations · 6 min read

Danger zone monitoring with object detection on a site camera

A polygon over the crane swing radius on a site camera. People and vehicles as classes, the bottom of the box as the test, the alert with the frame attached.

Esdras Ntuyenabo · Sep 27, 2026

LexData
LexData

Product

  • Platform
  • Industries
  • Use cases

Resources

  • Docs
  • The Field Guide
  • Blog
  • Why models drift

Industries

  • Energy & utilities
  • Oil & gas
  • Agriculture
  • Manufacturing
  • Insurance
  • Retail
  • Robotics

Company

  • About
  • Customers
  • Careers
  • Contact

Trust

  • Security
  • Privacy
  • Terms

Stay updated

What we learn running vision models in production.

See everything.
Miss nothing.

Stay updated

What we learn running vision models in production.

Terms of use & Privacy policy

© 2026 LexData Labs · All rights reserved