Industries · 6 min read
Training computer vision models on aerial imagery where a person is fifteen pixels tall
Over a corridor a crew member is a few dozen pixels. Tile instead of resizing, check orientation at import, label at full resolution, QA pass on every box.
Summary
This post follows a survey campaign over a transmission corridor where people and vehicles on the access track are a few dozen pixels tall, and sets out the pipeline decisions that make the frames trainable: tiling instead of resizing, orientation metadata checked at import, a bounding box at full resolution with a QA pass on every one, and negatives that look like the target. It concludes that a re-planned flight is the framing that moves and the model was fitted to the old one. It is for survey, inspection and labeling teams working from drones and aircraft.
Rajiya Sultana · Engineering Manager · Sep 29, 2026

Wide drone view of a transmission tower over a country road, insulator strings and a corrosion mark boxed, from a customer inspection run
The September campaign flies the transmission corridor south of the substation at survey altitude. The footage that comes back has everything the asset team asked for: towers, insulator strings, the access track, and the crew's two trucks parked at the base of tower 14 with three people beside them. The people are fifteen pixels tall. The trucks are perhaps forty. The safety lead wants a model that finds people and vehicles inside the corridor on every future flight, and the first version, trained on frames resized to the size a general detector expects, finds the towers and nothing else.
At the size the training pipeline made them, the people were not there to find.
Tiling instead of resizing keeps a person the size the camera saw
The survey camera on the September campaign produces large frames, and a training loader that resizes them to a standard input shrinks a fifteen-pixel person to a smudge of two or three. The first decision in the pipeline is to stop doing that. Cut each frame into tiles at native resolution, so a person is fifteen pixels on the tile as they were on the frame, and train and run detection on the tiles. A person straddling a tile edge gets a box in each tile and the two are merged by their overlap.
Tiling multiplies the frames and the labeling, and there is no way round that. The vision lesson sets the floor: detection accuracy on objects under a few dozen pixels a side runs at a fraction of the headline number for every architecture, and the fixes are in capture and tiling rather than in the weights. A campaign flown lower would make the person forty pixels and the tiling lighter, and that is a conversation with the pilot rather than with the model.
Orientation metadata is checked at import before a box is drawn
Survey frames arrive with a note in the header saying which way was up when the shutter fired, and some loaders honour it and some do not. A frame that a person labeled upright and a loader read raw puts every box on the wrong part of the picture, and nothing errors. On a campaign that mixes the aircraft's camera with a crew member's handheld shots of the base of tower 14, the mixture is where this shows up.
The check runs at import on every batch: read a sample of files with the note honoured and again with it ignored, and count the ones whose width and height swap. Those had a live flag. The rotation is applied once, the upright pixels are saved, and the flag is stripped, so there is nothing left for a downstream loader to interpret differently. The quickstart has the companion rule, upload originals rather than compressed exports, because a re-encoded frame has changed the pixels a fifteen-pixel person is made of.
A bounding box at full resolution, with a QA pass on every one
The labeling is boxes. A bounding box on each person and each vehicle, drawn on the tile at native resolution, hugging the figure rather than the shadow beside it. At fifteen pixels a box that is three pixels wide of the person is a box that is a fifth wrong, and the error does not average out; it becomes a fixed offset the model learns.
You type the classes once, person and vehicle, Lexi proposes the boxes on every tile, and a person checks each one. On this campaign the checking is the whole job. A proposed box on a fence post that reads as a person, a truck under a tree that got no box, a crew member half behind a tailgate. Every label passes a QA pass before it trains, and on small objects that pass is where the dataset is made or lost.
My own view, from running these queues, is that every full-resolution box on a small object should get a second look from a different person than the one who confirmed it. The first reviewer anchors on the draft. The second is looking for what the draft missed, and on a corridor tile that is the person standing still beside the truck.
An aside from the flight log: the pilot records altitude on every leg, and the leg where the altitude reads lower is the one where the labels come back cleanest, which the labeling team noticed before the pilots did.
Small objects need the negatives that look like them
At fifteen pixels a person and a fence post and a survey marker are near neighbours, and a model trained only on tiles with people in it has never been shown the post. The negatives go into the set on purpose: tiles of the access track with nothing on it, the base of tower 14 with the trucks gone, the markers, the posts, the cattle in the field beside the corridor. A few hundred of them, tagged, so the model trains on the empty corridor as often as on the occupied one.
A re-planned flight is the framing that moves
The October campaign is flown on a different plan. Airspace, weather and battery range move the legs, the aircraft is higher over the northern towers and lower over the southern ones, and the same tower arrives at a different angle and scale from the one the model was fitted to. The power line and grid inspection use case names this as the condition the whole task lives under: the asset has not changed, the angle and scale it arrives at have. The drift catalog covers it as a camera moved.
The signal is a step on the flight date rather than a slope: detections on the northern legs diverge from the September history while the southern legs hold. The fix is the cheapest on the catalog. Re-label a short window of tiles from the new framing and retrain on frames already flown.
The campaign's doubted tiles are the next version
LexData takes the corridor model through its whole life. You type what to look for, Lexi puts a box on every tile, and a person checks each label before anything trains on it. The model then watches the campaign footage in the cloud, on your servers or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime.
The alert path across grid assets built the same way is what produces the 21,000+ hazard detections per month in our energy work, and a person inside the corridor is the one detection on that list that goes to the safety lead first.
The three people beside the trucks at tower 14 are fifteen pixels tall in September and, on the lower southern legs, about twenty-five in October. The model that flies in November has seen both.
See it on your own footage.
Start with your footageMore in Industries

Industries · 7 min read
Aerial fire detection from a drone patrol, smoke before the flame reaches the line
On a right-of-way patrol, smoke is a few dozen grey pixels that look like haze. Two boxed classes, an alert with a cooldown, and the bad-weather days kept.
Rob Hickey · Sep 29, 2026

Industries · 6 min read
AI in robotics after the robot ships, what the warehouse cameras keep learning
The forward camera boxed pallets and people well at the pilot site. Then the racking moved, and the edge cases the planner never saw came back for review.
Andreas Ohrvall · Sep 29, 2026

Industries · 6 min read
Automated sorting with computer vision, from the camera over the conveyor to the diverter
A box on every apple, a grade from the box, and an air jet that acts on it before the belt moves on. The new cultivar is when the model needs the graders again.
Rob Hickey · Sep 29, 2026