Labeling · 7 min read
Few-shot image labeling on a vial tray, one box drawn and the rest proposed
Draw one box on a vial, let the tray fill with proposals, and give a person the last word on each before the batch commits. The fastener bin is the hard case.
Summary
This post follows few-shot labeling across a tray of vials and a bin of loose fasteners, from the first drawn box to the proposals it produces and the slider that governs them. It concludes that the proposals are worth their speed only because a person confirms or rejects each one before the batch commits. It is for teams labeling dense, repetitive frames from a production line.
Rajiya Sultana · Engineering Manager · Oct 3, 2026

Assembly station with a housing, fasteners and a cable laid out, the cable boxed, generated scene with detections from our model
A tray comes off the fill line at a pharmaceutical plant holding ninety-six vials in a twelve by eight grid, and the camera above the outfeed sees every one of them. Labeling that frame by hand means ninety-six boxes, and the next tray is identical, and there are four hundred trays on a night shift that ends at 6 am. Nobody draws ninety-six boxes four hundred times. They draw one.
Few-shot labeling is the name for what happens after that one box. The labeler boxes a single vial, the tool searches the frame for everything that looks like it, and the tray fills with proposals. What decides whether those proposals become a dataset or a mess is the part that follows, where a person looks at each one.
One drawn box on the vial tray proposes the rest
The first box is the example, and Lexi takes the pixels inside it, looks for regions of the frame with a similar appearance at a similar scale, and places a candidate box on each. On a vial tray under fixed lighting this works almost embarrassingly well: the vials are the same shape, the same glass, the same spacing, and the ninety-five proposals land within a pixel or two of where a patient labeler would have put them.
The same first box works across frames, too. The next tray, and the tray after it, get proposals from the example the labeler drew on the first, so the drawn boxes for a whole shift can be counted on one hand. This is the case the method was made for: many instances per frame, a controlled background, and a class whose members look alike.
The labeler at the outfeed names the trays by the colour of the tape on the corner rather than by the batch number printed on the side, because the tape is visible from the camera and the number is not.
A bounding box proposal is a suggestion until a person confirms it
Every proposal is drawn in a different colour from a confirmed label, and stays that colour until someone clicks it. That distinction is the whole discipline. A bounding box the tool placed has not been checked, and a model trained on unchecked boxes learns whatever the tool got wrong, at scale, on every tray.
The review is quick when the proposals are good. The person scans the tray, sees ninety-five boxes where ninety-five vials are, confirms the lot with one action, and moves to the next frame. It is slower when a proposal sits on the reflection of a vial in the tray lid, or on the gap where a vial is missing, because those need a rejection each. Rejections matter as much as confirmations: a missing vial is exactly the frame the line wants to catch later, and a box confirmed on the empty slot teaches the model that empty slots are vials.
The labeling with Lexi doc puts it as a QA pass on every label before training. On a vial tray, the pass is the difference between ninety-six labels and ninety-six guesses.
The slider trades misses against wrong proposals
Each proposal carries a score, and a slider sets how low a score is still shown. Slide it down and the tray fills with more proposals, including the reflections and the shadows between rows. Slide it up and only the cleanest matches remain, and the vial at the edge of the frame, half cut by the tray rim, disappears from the proposals and has to be drawn by hand.
Neither end is right for every frame. A tray under the bright outfeed lamp on line 2 can take a high setting; the same tray photographed on the packing bench at the end of the line, under mixed light, needs a lower one and more rejections. My own view is that the slider should start high and be lowered by the reviewer as they learn the frame, never the other way round. A reviewer who starts low spends the first hour rejecting shadows and stops looking carefully at what is left.
Fasteners in a bin are where the proposals go wrong
Across the plant, at the assembly station on line 4, a bin of loose fasteners is a different problem in the same tool. The fasteners are one class, but they lie at every angle, half on top of each other, some head up and some thread up, some under the shadow of the bin wall. The first drawn box shows one head-up screw, and the proposals find the other head-up screws and miss most of the rest.
This is the case few-shot labeling is weakest at: high variation inside one class, heavy occlusion, and a background that is other members of the same class. Two or three examples at different angles help. Proposals still miss the screw lying on its side under two others, and the person reviewing has to draw it, because a model that never sees the occluded screw will never count it.
The fine distinctions are worse. If the bin holds two screw lengths that differ by a few millimetres, the proposals cannot tell them apart from one example, and a reviewer confirming them as a single class has quietly decided that length does not matter. That decision should be made on purpose, in the class list, before anyone opens the bin frame.
The batch commits only after every proposal has a verdict
A batch of trays goes into the dataset together, and it goes in only when every proposal in it has been confirmed or rejected. Proposals left in their undecided colour hold the batch. That rule is what makes a fast method safe: the speed comes from proposing, the safety from the fact that nothing proposed can train until a person has said so.
The same rule catches a drift in the labeler's own habits. If the first hundred trays are confirmed with the vial box hugging the glass and the next hundred with the box including the cap, the model sees two conventions. A sampled second look at confirmed labels, a few frames from each batch, finds that before training does. Labels that come back at up to 99.9% accuracy are the product of that second look, applied to every batch rather than to the first one.
Confirmed labels are what the line model learns from
LexData takes the tray model through its whole life. You type what to look for, Lexi puts a box on every vial in every frame, and a person checks each label before anything trains on it. The model then watches the outfeed camera the plant already has, in the cloud, on the plant's own servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime.
Few-shot proposals and the frames the model doubts are the same review, seen from two ends. Before launch, the proposals fill the trays and a person decides each one. After launch, the model fills the trays itself, and the frames where it hesitates, a vial at a new angle, a tray under a lamp that has been changed, come back to the same person for the same verdict. On the manufacturing lines we run, that is how 99%+ accuracy is maintained in production rather than measured once at launch.
The tool that placed the ninety-five boxes was never the point of the method. The person who rejected the one on the empty slot was.
See it on your own footage.
Start with your footageMore in Labeling

Labeling · 7 min read
Aerial dataset augmentation for drone frames where there is no up
A tower seen straight down has no top or bottom, so rotations and both flips are safe. Scale for altitude, brightness for sun, move every box with its pixels.
Esdras Ntuyenabo · Oct 3, 2026

Labeling · 6 min read
A collaborative data annotation workflow run as a pipeline
Batch the bottling line's frames by camera and shift, assign so nothing is boxed twice, attach the guideline, review every label, train on the approved set.
Rajiya Sultana · Oct 3, 2026

Labeling · 7 min read
Dataset health check for computer vision, what to look at before anything trains
A scratch dataset where every scratch sits in the centre of the frame will train a model that looks in the centre. Five counts to read before the first epoch.
Rajiya Sultana · Oct 3, 2026