Labeling · 6 min read
What an open-source annotation tool gives you and where it stops
Two years of fillet-line labels in a self-hosted tool, exported as CVAT XML, and no model watching the line. The tool was never the bottleneck. The loop was.
Summary
This post looks at a food processing team with two years of labels on line footage in a self-hosted open-source tool, and asks what the tool gave them and what it never could. It concludes that the boxes and the database were the easy part, that nothing in the tool watched the line, reviewed doubted frames or retrained, and that the CVAT XML export comes across into a loop that does. It is for teams with a labeled archive and no model in production.
Ayman Quadir · Head of Product · Sep 29, 2026

Food processing line with fillets on a blue conveyor, fillets and the rail boxed, generated scene with detections from our model
The food processing line has had a camera over the fillet conveyor for two years, and for two years a QA technician has spent part of every Friday in a self-hosted labeling tool, drawing boxes on fillets: whole, ragged edge, bone in, foreign object. The archive is large and carefully done. The export is a folder of CVAT XML files. And on the line, at 6 am on a Monday, nothing is watching the conveyor except the technician who drew the boxes.
That is the shape of a lot of teams' first two years. The tool did what it was for. What it was for turned out to be a fraction of the job.
An open-source annotation tool gives you boxes and a database
What the tool provides is real and worth having. A place to define classes, so "ragged edge" means the same thing on Friday as it did last March. A drawing surface for boxes, polygons and tags, with instance tracking across frames of a clip. A database of who drew what and when. Some assistance on the first pass, a general model proposing boxes the person accepts or moves, which on fillets under the line's blue conveyor works about half the time.
For a team with a technician's Friday to spend, that is enough to build an archive, and this team built a good one. The class list is sane, the boxes are tight, and the ragged-edge class, which the QA lead named after the word already on the reject sheet, is consistent across two years because one person drew nearly all of it.
The tool stops at the export and nothing watches the line
Then the archive sits there. The tool has no notion of the camera over the conveyor as a live thing; it knows the clips the technician uploaded. There is no model in it watching the line, no rule that fires when a bone-in fillet passes at 6 am, no path by which a frame the model doubts comes back to the technician. Training happens elsewhere, if it happens, by someone exporting the folder and running a script on a machine under a desk, and the model that results is deployed nowhere in particular.
The thing that was missing for two years was never the labeling. It was everything after it: the watching, the review of what the model doubted, the retraining on the corrections, and the new version replacing the old with no downtime. The platform exists for that part, and the labeling is the first step of it rather than the whole.
Two years of CVAT XML come across without losing a box
The archive is the asset, and it moves intact. The CVAT XML files import as they are: each clip's frames matched to the footage, the classes becoming the class list, the boxes and polygons becoming labels with the technician's two years of care attached. Nothing about the footage is re-encoded on the way in. COCO JSON and YOLO TXT come across the same way, so an archive split across tools over the years arrives as one dataset. The quickstart covers the first import.
What arrives is checked before it trains. Lexi draws its own first pass on a sample of the imported frames and the two are compared. That surfaces the places where the class list drifted over two years, a "ragged edge" that meant something slightly different in the first six months, before anything is trained on it.
The first pass is the model's and the review is the person's
From the first training run the technician's Friday changes shape. Lexi puts a box on every new frame from the conveyor camera, and the technician checks the boxes rather than drawing them, which is quicker and, on a line of near-identical fillets, far less tiring. Every label is still checked by a person before it trains anything. The difference is that the person's time goes to the frames the model was unsure of, the fillet half under the rail, the foreign object that might be a glove tip, rather than to the thousandth whole fillet.
My own view is that the labeling tool was never the bottleneck on this line, and that teams who go looking for a better tool are usually looking for a better loop. The boxes were fine. The two years without a model watching the conveyor were the cost.
Doubted frames and retraining are what the tool never had
Once the model watches the conveyor, it produces the stream the tool had no place for: the frames it doubts. LexData takes the fillet model through its whole life. You type what to look for, Lexi puts a box on every frame, and a person checks each label before anything trains on it. The model then watches the cameras the line already has, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime.
The corrections are the same boxes the technician was drawing on Fridays, drawn now on the frames that matter, and when they cross the project's threshold a new version trains on them. A supplier change in November, when the fillets arrive a shade paler, shows up as a rise in the technician's corrections on the Monday, and the version trained that week has seen the paler fillets. The archive from the self-hosted tool is the foundation of every one of those versions.
Self-hosting was the hidden second job
The tool ran on a server in the plant's IT cupboard, and somebody kept it running: upgrades, backups, the disk filling up with clips, the login that broke after a password policy change. That was never in anyone's job description, and it is the part of "free" that teams forget to cost. Moving the archive out did not just add the loop; it retired a server that a person had been quietly nursing for two years.
The folder of XML files is still kept, because the technician does not fully trust anything that cannot be opened in a text editor, and that is a reasonable position for the person who drew every box in it.
See it on your own footage.
Start with your footageMore in Labeling

Labeling · 6 min read
COCO as a format you will use and a benchmark you should not trust
COCO JSON is the file a bottling line's labels travel in. The COCO benchmark is a score on somebody else's eighty classes, and none of them is a missing cap.
Sheikh Srijon · Sep 29, 2026

Labeling · 6 min read
EXIF orientation, the photo that is sideways only to the model
A phone photo looks upright on every screen and arrives rotated in training, because the pixels never turned. The check to run at import, before the first box.
Stephen Biswas · Sep 29, 2026

Labeling · 6 min read
What a model trained on ten frames is good for
Ten labeled frames from a new site camera give a first pass by the end of the afternoon. Trust it for the obvious, and let its doubts grow the set.
Sheikh Srijon · Sep 29, 2026