Skip to content
LexDataLexData
PlatformIndustriesCustomers
DocsThe Field GuideBlogWhy models drift
AboutCareersSecurityContact
Log inStart now
← All posts

Industries · 6 min read

Why the PlantDoc dataset is not your field, and what it is good for

A public leaf disease set is a photo of somebody else's field, lab backgrounds and box errors included. Use it as a baseline, then label the grower's frames.

Summary

This post explains why a model trained on a public plant disease dataset does well on its own test split and badly on a grower's row camera: the public frames are picked leaves on tables and web photos, with label errors of their own. It concludes that the public set is a baseline and a vocabulary, and that the model the grower runs has to be trained on the grower's own frames with a QA pass. It is for agronomy and crop technology teams.

Sheikh Srijon · GTM Lead · Sep 30, 2026

Orchard tree rows and drivable paths boxed, from a customer tractor camera

An agronomy team downloads a public plant disease dataset on a Monday, trains a detector on it by Wednesday, and gets a number on the held-out split that looks like a result. On Thursday they point the model at the camera over row 6 of the grower's orchard. It boxes the sky as a lesion, misses the leaf the scout pegged that morning, and finds disease on the tractor's mudguard.

Nothing went wrong with the training. The model learned the dataset it was given, and the dataset had never seen row 6.

A public dataset is a photo of somebody else's field

Open the PlantDoc dataset, or any public leaf disease set, and look at the frames rather than the class list. Many are a single leaf, picked and laid on a table or held against a plain wall, photographed at arm's length under indoor light. Others are photos found on the web, at whatever resolution and crop the page had, with a watermark in the corner. The classes span crops: apple, tomato, potato, grape, corn, each with a few diseases, so a class has a few hundred examples and no two of them from the same camera.

The file names give the source away. A run of frames named the way a search engine names its downloads is a run of frames from a search engine.

None of this makes the set useless. It makes it a set of leaves, and a grower does not need a model that finds disease on a leaf on a table. The grower needs one that finds it on row 6, from three metres up, with the sun behind the tree and a hundred other leaves in the frame.

Object detection trained on lab leaves fails on a row camera

The failure is specific. Object detection learns from the whole frame, and the whole frame in the public set is a leaf and a table. The model learns that the leaf is the large green thing in the middle and that the disease is the discoloured patch on it. On the camera over row 6 there is no large green thing in the middle. There are hundreds of small leaves, overlapping, at every angle, some in sun and some in shade, and a lesion the size of a fingernail is a handful of pixels.

The scale is wrong, the background is wrong, the occlusion is wrong, and the light is wrong, and any one of those would be enough. A model that was right on the test split was right about tables.

The crop health and disease detection use case names the hard part on a real camera. Early symptoms are exactly what you want to catch and exactly what looks like ordinary variation, sun scald or dust, and the distinction is a few shades of colour under uncontrolled light. No frame in the public set has that light.

Errors in the labels are part of the set

PlantDoc and the other public sets have label errors of their own, and the well-known ones have been catalogued. Boxes that extend past the edge of the image. Boxes with no area, a single point where a label should be. Images in the set with no label at all, which a trainer reads as "nothing here" when there is a diseased leaf in the middle of the frame. Later releases of some sets clean those up, and the cleaning is worth having.

The lesson is not about one dataset. Every set of labels has errors, and the errors in a set you did not make are ones you have not seen. The QA pass a grower's own labels get, a person checking every box before anything trains, is the step a downloaded set skipped on your behalf.

What a public set is good for

My own view is that a public set is worth downloading and never worth deploying. It is a baseline: train on it, run it on the frames from row 6, and the number you get is the floor the grower's own labels have to beat. They will beat it by a margin that ends the conversation about whether labeling is worth it.

It is also a vocabulary. The class list in a public set is a reasonable first draft of the grower's, and the frames are a way to check that two labelers agree on what leaf mould looks like before they are shown a real row. And it is a warm start: a model that has seen a few thousand leaves, even on tables, has less to learn from the first week of orchard frames than one that has seen nothing.

What it is not is the model. The model is trained on the camera it will run on.

The grower's own frames, with a QA pass

The switch is not complicated. The camera over row 6 has been recording since it was fitted. A window of frames from it, across the hours of the day and the weeks of the season, is the training set the model was always going to need. The how much footage you actually need lesson covers the question everyone asks first, and the honest answer is that a week from each camera, labeled and checked, is worth more than any number of downloaded leaves.

You type the classes once, using the public set's list as the draft, and Lexi puts a box on every leaf in every frame from row 6. A person checks each box before anything trains, and on a crop the person is someone who has walked the rows, because the label that matters is the one the agronomist would agree with. The labeling doc covers the verification pass.

The agriculture work behind our numbers, 80k+ image annotations in production, was built this way: frames from the grower's own cameras, at the light and the scale the cameras actually have, checked by people who know the crop. None of it came from a table.

The model the grower runs is the one trained on row 6

LexData takes the disease model through its whole life. You type what to look for, Lexi puts a box on every leaf in every frame, and a person checks each label before anything trains on it. The model then watches the cameras over the rows, in the cloud or on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime.

The frames that come back from row 6 in the first month are the ones the public set could never have prepared it for. A lesion in the shade of a branch, a leaf half hidden by fruit, dust on the lower canopy after the sprayer passed. Each one is a correction, and when the corrections cross the project's threshold a new version trains on the orchard's own footage. The public set stays on the disk as the baseline it was, and the scout keeps pegging the plants the model should have found.

See it on your own footage.

Start with your footage

More in Industries

Industries · 6 min read

Perimeter security with fixed cameras, object detection and a drone sent to look

A frame every two seconds is enough to catch a person at the fence, a CPU is enough to run it, and the drone is the second look rather than the detector.

Andreas Ohrvall · Sep 30, 2026

Industries · 7 min read

Food service QA with a camera over the tray packing line

Every component on the tray gets a box, the missing one is flagged before the sealer, and the alert count is read against the line's own history.

Ayman Quadir · Sep 30, 2026

Industries · 7 min read

Railway safety with trackside cameras, zones and a signaller who can live with the alerts

People and vehicles boxed, the track bed and crossing drawn as zones, the frame sent to the control room, and a false alarm rate a signaller will keep reading.

Rajiya Sultana · Sep 30, 2026

LexData
LexData

Product

  • Platform
  • Industries
  • Use cases

Resources

  • Docs
  • The Field Guide
  • Blog
  • Why models drift

Industries

  • Energy & utilities
  • Oil & gas
  • Agriculture
  • Manufacturing
  • Insurance
  • Retail
  • Robotics

Company

  • About
  • Customers
  • Careers
  • Contact

Trust

  • Security
  • Privacy
  • Terms

Stay updated

What we learn running vision models in production.

See everything.
Miss nothing.

Stay updated

What we learn running vision models in production.

Terms of use & Privacy policy

© 2026 LexData Labs · All rights reserved