Skip to content
LexDataLexData
PlatformIndustriesCustomers
DocsThe Field GuideBlogWhy models drift
AboutCareersSecurityContact
Log inStart now
← All posts

Labeling · 6 min read

COCO as a format you will use and a benchmark you should not trust

COCO JSON is the file a bottling line's labels travel in. The COCO benchmark is a score on somebody else's eighty classes, and none of them is a missing cap.

Summary

This post separates the two things called COCO, a JSON format that holds images, categories, boxes and polygons, and a benchmark dataset of eighty everyday classes. It walks a bottling line's labels into and out of the file and concludes that the format is the right passport for a dataset while the benchmark says nothing about a line's own objects. It is for teams who have been told to use COCO and want to know what that means.

Sheikh Srijon · GTM Lead · Sep 29, 2026

Bottling line with bottles queued under the fill head, one missing its cap, generated scene with detections from our model

The bottling line's class list is short: bottle, cap, and missing cap. The team labeling it is told by three different people to "use COCO", and each of the three means something different. One means the file format the labels should be saved in. One means a model that was pretrained on the COCO pictures. One means the benchmark score in a paper, and is about to be disappointed.

They are three uses of one name, and only the first one is unambiguously good advice.

The COCO dataset is two things wearing one name

COCO is a public dataset: a large collection of everyday photographs, labeled with boxes and outlines for eighty classes of common things, such as people, cars and chairs. It is also the name of the JSON layout those labels were published in, which became the layout everyone else used because the tools already read it. The COCO dataset and the COCO format share a name and nothing else that matters to a bottling line.

The format is the part the line will use every day. The dataset is the part it will hear about in benchmarks and should treat with care.

The JSON holds images, categories and one record per box or outline

A COCO file is one JSON document with three lists. The images list has one entry per frame: a file name, a width and a height, and an id. The categories list has one entry per class, with an id and a name, "bottle", "cap", "missing cap". The annotations list is the long one: one entry per labeled thing, carrying the image it belongs to, the category, a box as four numbers, and, for an outline, a polygon as a list of points. That is nearly all of it. There are a few more fields, and none of them is needed to read a file back.

The structure is why it travels. Any tool that reads the three lists can read a file from any other tool, and a labeler who wants to check a file by eye can open it and find the bottle with the missing cap on frame twelve by following two ids.

Frame twelve, on the Monday run, was the one where the cap was on the conveyor beside the bottle, and the labeling rule had to decide whether that was a missing cap or a loose one.

A line's labels travel in and out of the file without re-encoding

The bottling line's dataset arrives as COCO JSON alongside the frames, from an older tool or a contractor, and it imports as it is. The images list is matched to the footage, the categories become the class list, the boxes and polygons become labels a person can check. Nothing about the frames is re-encoded on the way in, so a frame that was sharp in the old tool is sharp here. The quickstart covers the first import, and the same file shape goes back out: the dataset exports as COCO, with the checked labels and the corrections that came from review, so it can be trained elsewhere or handed to the next contractor.

A team with a year of line labels in the format loses nothing by moving them, which is the strongest argument for keeping labels in it from the start.

The eighty classes contain none of yours

Now the benchmark. The COCO dataset's eighty classes are the things that appear in everyday photographs, chosen a decade ago, and a bottling line's classes are not among them. There is, as it happens, a "bottle" class, and the pictures behind it are wine bottles on dinner tables and water bottles in gym bags, photographed by people rather than by a fixed camera over a conveyor. There is no cap. There is no missing cap. A score on the COCO benchmark is a score on dinner tables.

That is why a published figure of any size tells the line nothing about how the model will do on the fill head. The only figure that does is one measured on frames from the line's own camera, labeled by a person. The lesson on how much footage you actually need makes the case that the line needs a spread of its own conditions far more than it needs a large public set.

I would not put a COCO mAP figure in a deck meant for a plant manager. It answers a question nobody on the line asked, and it is read as a promise about the line.

A pretrained backbone helps and a benchmark score does not

The second person's advice, start from a model pretrained on the COCO pictures, is sound, and it is sound for a reason that has nothing to do with the score. A model that has seen a million photographs has learned edges, textures and shapes that transfer, and it reaches a usable state on the line's few hundred labeled frames far sooner than one starting from nothing. The pretraining is a head start on seeing. It is no head start on caps.

The line's own frames do the rest. Lexi puts the first box on each, a person checks every label, and the model that results carries the line's three classes and none of the eighty. LexData takes that model through its whole life. You type what to look for, Lexi puts a box on every frame, and a person checks each label before anything trains on it. The model then watches the cameras the plant already has, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime.

The file is the dataset's passport and the checks happen before training

Before a COCO file trains anything, a few things are worth checking by hand, because the format does not check them for you. Every annotation's image id should point at a frame that exists. Every category id should appear in the categories list, with the name the line uses rather than the contractor's shorthand. A box should sit inside its frame's width and height, and a polygon should close. A file that fails any of those trains a model on labels that point at nothing, and the training tool will rarely say so.

The Monday frame twelve, with the cap on the conveyor, ended up labeled as a missing cap, because the rule the line wrote says the class is about the bottle and the bottle had no cap on it. That rule lives in a paragraph beside the file, and the file cannot carry it.

See it on your own footage.

Start with your footage

More in Labeling

Labeling · 6 min read

EXIF orientation, the photo that is sideways only to the model

A phone photo looks upright on every screen and arrives rotated in training, because the pixels never turned. The check to run at import, before the first box.

Stephen Biswas · Sep 29, 2026

Labeling · 6 min read

What a model trained on ten frames is good for

Ten labeled frames from a new site camera give a first pass by the end of the afternoon. Trust it for the obvious, and let its doubts grow the set.

Sheikh Srijon · Sep 29, 2026

Labeling · 7 min read

Choosing an annotation tool for a month of inspection footage

The demo set labels itself in an afternoon. A month of drone footage is where the tool has to propose boxes, take corrections and review every label.

Sheikh Srijon · Sep 29, 2026

LexData
LexData

Product

  • Platform
  • Industries
  • Use cases

Resources

  • Docs
  • The Field Guide
  • Blog
  • Why models drift

Industries

  • Energy & utilities
  • Oil & gas
  • Agriculture
  • Manufacturing
  • Insurance
  • Retail
  • Robotics

Company

  • About
  • Customers
  • Careers
  • Contact

Trust

  • Security
  • Privacy
  • Terms

Stay updated

What we learn running vision models in production.

See everything.
Miss nothing.

Stay updated

What we learn running vision models in production.

Terms of use & Privacy policy

© 2026 LexData Labs · All rights reserved