Skip to content
LexDataLexData
PlatformIndustriesCustomers
DocsThe Field GuideBlogWhy models drift
AboutCareersSecurityContact
Log inStart now
← All posts

Computer vision · 6 min read

Improving a custom object detection model, the tactics that still move the number

Two thousand frames from the real line beat ten thousand from a bench, a written box rule beats a talented labeler, and curation is a job that never finishes.

Summary

This post lists the training tactics that still improve a custom detector on a food line camera, from frames off the real line instead of a bench to a written box rule, rare cases collected on purpose, resolution chosen for the smallest thing, and pretrained weights. It argues that curation is a standing job fed by the doubted frames and should have a name on the line. It is for teams whose first detector scored well and then disappointed.

Rajiya Sultana · Engineering Manager · Oct 4, 2026

Food processing line, fillets on a blue conveyor with the conveyor, fillet and rail boxed, generated scene with detections from our model

The first detector for the fillet line on line 4 was trained on ten thousand frames photographed on a bench in the QA room: fillets laid out under white light, still, dry, one per frame. On the bench it scored beautifully. On the line, where the fillets are wet, moving, overlapping on a blue conveyor under a rail that casts a shadow, it missed the thing it was built to find, a fragment of blue glove lying across a fillet, on its first shift.

The second detector was trained on two thousand frames from the line camera itself, and it found the glove. What follows are the tactics that produced the second one, none of which involve a newer architecture.

Frames from the real line beat more frames from anywhere else

Two thousand frames from the line 4 camera carry the conditions the model will meet: the motion blur at belt speed, the wet sheen, the blue background that the glove fragment nearly matches, the rail's shadow crossing every fillet at the same angle. Ten thousand bench frames carry none of them, however many there are, because a bench is one condition sampled ten thousand times.

The lesson on how much footage you need puts the unit as distinct conditions rather than frames, and the line has more of them than it looks. There is the morning shift under full lights and the evening shift under half, the belt after a wash-down and the belt an hour later, the winter frames when the room runs colder and the fillets fog the lens. A set that samples each of those is worth more than any quantity of the same shot.

A written box rule beats a talented labeler

Two labelers given the same wet fillet will box it differently unless told how. One includes the sheen of water around the edge, one stops at the flesh; one boxes a fragment of glove tightly, one includes the fillet it is lying on. Neither is wrong, and a model trained on both learns that the edge of a fillet is somewhere in a band a few pixels wide, which comes back as loose boxes on the line.

The rule is written before the first frame is labeled and it is short. The box hugs the visible flesh and excludes the water sheen. A foreign object is boxed on its own extent, never with what it lies on. An occluded fillet is boxed on what is visible and never guessed. Lexi proposes the boxes on every frame, and a person checks each one against that rule, and the checking pass is where the rule is enforced rather than the labeling pass.

The line's QA sheet has called foreign matter "FM" for years, and the class list adopted the abbreviation because that is what the supervisor reads at the end of the shift.

Rare cases are collected on purpose or not at all

A fragment of blue glove appears on the line 4 belt perhaps once a week. Two thousand frames sampled at random from the line camera contain none of them, and the model trained on those frames has never seen the thing it exists to find. The rare class has to be collected deliberately: every doubted frame with a possible fragment kept, the frames from the weeks a fragment was found pulled from the recorder, and a set of fragments placed on the belt during a cleaning break and filmed.

Recall on the rare class is the number to watch, class by class, and never the average across classes. A model that finds every fillet and half the glove fragments has an excellent average and is failing at its job.

Resolution is chosen for the smallest thing that has to be found

A bone fragment on a fillet is a few dozen pixels on the line 4 camera at its native size. A training pipeline that shrinks every frame to a small square before it reaches the network shrinks the fragment to a few pixels, below what any detector resolves. The input size is chosen for the smallest class that matters, and if that size is too heavy for the runner beside the recorder, the frame is tiled or the camera is moved closer.

The same choice runs the other way for a class that is large in the frame. The fillet itself does not need the native resolution, and a model that only had to find fillets could run smaller and faster. The bone fragment sets the floor, and the runner in the cabinet has to be sized to it rather than to the fillet.

Pretrained weights help until the domain gets too far away

Starting from a checkpoint trained on a large general set gives the model its edges, corners and textures before it has seen a single fillet, and on a food line camera that is a fair start: the general set contains food, conveyors and gloves. Fine-tuning from it reaches a usable model on far fewer frames than training from nothing. The help thins out as the domain gets stranger, and a thermal camera or a microscope stage is far enough from the general set that the pretrained edges are worth little.

One tactic that goes with it: freeze the backbone for the first part of training and let only the head learn, then release the backbone once the head has settled. On line 4 that kept the general features intact while the head learned what a glove fragment on blue looks like.

Curation is the standing job most computer vision projects never staff

LexData takes the fillet model through its whole life. You type what to look for, Lexi puts a box on every frame, and a person checks each label before anything trains on it. The model then watches the cameras the plant already has, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime.

That loop is what makes curation a standing job rather than a launch task. The doubted frames are a steady supply of exactly the frames the training set lacked, and the corrections on them are the highest-value labels the line will ever get, as the lesson on retraining without starting over argues. Versions keep what they were trained on, so the set only grows, and the held-out week stays the most recent one so the score is honest about the present.

My own view is that curation should be a named responsibility on the line, on the same footing as the weekly knife check, with a person who owns the doubted frames for that camera and an hour in the week to review them. Most computer vision projects budget for the labeling that happens before launch and nothing for the labeling that happens after it, and that is the reason the second version so often never arrives.

See it on your own footage.

Start with your footage

More in Computer vision

Computer vision · 7 min read

Five computer vision applications in production, on the cameras a site already owns

Detect, count, inspect, track and read are the five jobs a fixed camera can be given, and each is a different question with a different label behind it.

Ayman Quadir · Oct 4, 2026

Computer vision · 6 min read

Computer vision projects worth building on the cameras you already have

A plant, a utility, a grower, a store and a warehouse each have a project that starts on an existing camera, produces a decision, and has someone to act on it.

Ayman Quadir · Oct 4, 2026

Computer vision · 6 min read

How to choose an object detection model architecture for a camera on your own site

Start from the decision the plant has to make, then the box beside the recorder the model has to fit, and only then the family. The benchmark comes last.

Andreas Ohrvall · Oct 4, 2026

LexData
LexData

Product

  • Platform
  • Industries
  • Use cases

Resources

  • Docs
  • The Field Guide
  • Blog
  • Why models drift

Industries

  • Energy & utilities
  • Oil & gas
  • Agriculture
  • Manufacturing
  • Insurance
  • Retail
  • Robotics

Company

  • About
  • Customers
  • Careers
  • Contact

Trust

  • Security
  • Privacy
  • Terms

Stay updated

What we learn running vision models in production.

See everything.
Miss nothing.

Stay updated

What we learn running vision models in production.

Terms of use & Privacy policy

© 2026 LexData Labs · All rights reserved