Computer vision · 6 min read
How to train, evaluate and deploy a computer vision model on a camera you already own
From the first labeled frame off a dock camera to a model watching that dock, in the order the work really happens, with the retrain built in from day one.
Summary
This post follows one dock camera from the question it is asked to the model watching it, through labeling, checking, training, evaluating on the camera's own frames, deploying and the retrain that follows. It argues that the order matters more than any single step and that the first fortnight of corrections is worth more than the evaluation everyone agonises over. It is for a team about to put its first model on a camera it already has.
Rob Hickey · Chief AI Officer · Oct 4, 2026

Loading dock from a mounted camera, trucks at the bays and a forklift with a pallet, forklift, pallet, person and truck boxed, generated scene with detections from our model
The camera over bay 4 at the distribution centre was fitted for insurance. It records the dock, the trucks backing in, the forklifts crossing between the trailer and the racking, and it has never been asked a question. The question the site wants answered is whether a person is standing in the forklift lane while a forklift is moving, because that is the near miss the safety log keeps recording.
Getting from that sentence to a model watching bay 4 is a sequence, and the order is the part that first projects get wrong. Here it is in the order the work actually happens.
Write the question as a sentence before any frame is labeled
The question has to be one a camera can answer from the pixels. "Is a person in the forklift lane while a forklift is moving" is answerable: a person is a thing the camera can see, the lane is a zone drawn on the frame, and a moving forklift is a box that changes position between samples. "Is the dock safe" is not answerable, and a project that starts there spends a month discovering what it meant.
Write the sentence down, with the camera it applies to and the hours it matters. On bay 4 the hours are the night shift, when the lane is worst lit and the supervisor is on the far side of the building. That sentence becomes the acceptance test, the alert rule and the class list, so it is worth an hour of argument before anything else starts.
Label a few hundred frames and check every bounding box
The class list for bay 4 is short: person, forklift, pallet. Lexi puts a bounding box on each of them in every frame you upload, and a person checks each box before anything trains on it. The checking is the job. A box that hangs loose around a forklift on every frame teaches the model that a forklift includes a metre of floor, and a person box that stops at the shoulders because a pallet hid the legs teaches the model that people are short.
The frames come from bay 4 itself, at the hours the question matters, which means a set of night frames labeled on purpose rather than whatever the morning recorded. The quickstart covers the mechanics, from the first upload to the checked labels. Labels checked this way come back at up to 99.9% accuracy, and the accuracy of the labels is the ceiling the model will train up to.
The dock's own vocabulary goes into the class list. The site calls them trucks, so the class is truck, and nobody has to translate when the alert arrives.
Train on the checked frames and hold out the most recent week
The first version of the bay 4 model trains on the checked frames, and the held-out set is the most recent week rather than a random tenth of the whole. A random slice hides frames from every day in the training set, so the model is tested on days it has already seen, and the score flatters it. Holding out the last week tests the model on a stretch of time it has never seen, which is what it will face from the moment it goes live.
The first version is not the last, which is the point of the whole sequence. Sites that treat the first training run as the product spend the next quarter trying to make it perfect before anyone sees it. A model that is live in days and corrected in the second week beats one that is polished for a quarter and then meets the night shift for the first time.
Evaluate on the dock's own frames at the hours it will run
The score that matters is precision and recall on bay 4's frames, at night, per class. A public benchmark says nothing about this dock. The held-out week says whether the model finds the person in the lane when a forklift is moving, and whether it flags an empty lane as occupied often enough that the supervisor stops reading the alerts.
Look at the misses one by one rather than at the average. On a dock the misses cluster: the person half behind a pallet, the forklift with its mast raised so the box is a different shape, the frame at 4 am where the lane light is out. Each cluster is a set of frames to collect and label before the next version, and the list of clusters is a better evaluation report than a single number.
Put the model on the camera and write the alert as a sentence
The model can run in the cloud, on the site's own servers, or on a runner beside the recorder in the dock office, where the footage stays on site and only the doubted frames leave. On a dock with a thin connection the runner is the usual answer, and the alert then fires locally first.
The alert is a rule written as a sentence, with a severity and a cooldown, approved before it goes live: a person inside the forklift lane while a forklift is moving, high severity, sent to the shift supervisor's phone through Slack. The frame arrives with the boxes drawn on it, so the supervisor can judge from the far side of the building whether to walk over. The monitoring and alerts doc covers where alerts land and how the review queue is built from them.
Send the doubted frames to a person and let the corrections retrain it
LexData takes the bay 4 model through its whole life. You type what to look for, Lexi puts a box on every frame, and a person checks each label before anything trains on it. The model then watches the camera the site already has, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime.
The signal that the model needs that next version is the correction rate. When the supervisor overrides the alert more often, or the doubted frames pile up around one shape, a mast-up forklift say, the corrections cross the project's threshold and a new version is trained on the frames that came back. Versions keep what they were trained on, so the night frames that fixed the mast-up forklift are still in the set a year later.
My own view, and it is not the consensus, is that the evaluation step above matters less than the first fortnight of corrections. A model evaluated on a held-out week is a model tested on the recent past. A model corrected for two weeks by the supervisor who reads its alerts has been tested on the present by the person it has to satisfy. How the whole sequence hangs together, from the first frame to the version that replaces the old one, is on the platform page.
See it on your own footage.
Start with your footageMore in Computer vision

Computer vision · 7 min read
Five computer vision applications in production, on the cameras a site already owns
Detect, count, inspect, track and read are the five jobs a fixed camera can be given, and each is a different question with a different label behind it.
Ayman Quadir · Oct 4, 2026

Computer vision · 6 min read
Computer vision projects worth building on the cameras you already have
A plant, a utility, a grower, a store and a warehouse each have a project that starts on an existing camera, produces a decision, and has someone to act on it.
Ayman Quadir · Oct 4, 2026

Computer vision · 6 min read
How to choose an object detection model architecture for a camera on your own site
Start from the decision the plant has to make, then the box beside the recorder the model has to fit, and only then the family. The benchmark comes last.
Andreas Ohrvall · Oct 4, 2026