Labeling · 7 min read
Aerial dataset augmentation for drone frames where there is no up
A tower seen straight down has no top or bottom, so rotations and both flips are safe. Scale for altitude, brightness for sun, move every box with its pixels.
Summary
This post takes a season of drone survey frames along a transmission corridor and works out which augmentations are honest for footage with no ground orientation: quarter turns and both flips, scale for the altitude the pilot flew at, brightness for sun angle and cloud. It shows how each transform has to move the labels with the pixels and where a transform invents a frame the drone never took. It is for teams training on aerial footage from inspection flights.
Esdras Ntuyenabo · Engineer · Oct 3, 2026

Transmission tower over a country road from a drone, six insulator strings and one corrosion mark boxed, from a customer inspection run
The drone flew the corridor in April, straight down the line at a fixed height. Every frame it brought back shows a tower from directly above: the crossarms as a cross, the insulator strings as short dark dashes at the ends, the shadow of the whole thing thrown to one side across the field. The team has a few hundred of those frames labeled, insulators boxed, corrosion flagged, and the model they train on them has to work on the flight in October, from a different height, with the sun in a different place.
A few hundred frames is thin for a class as small as an insulator seen from altitude. The usual answer is to make more, by transforming the ones there are. The question is which transforms produce a frame the drone could have taken and which produce one it never could.
On a ground camera the answer is constrained by gravity. A person is upright, a truck's wheels are at the bottom, and flipping the frame top to bottom makes a scene the camera will never see. Straight down from a drone, there is no up. That single fact decides most of what follows.
Rotations and both flips are safe because the ground has no top
Turn a straight-down frame by a quarter turn and it is a frame the drone would have taken on a heading ninety degrees off. Turn it by a half turn, flip it left to right, flip it top to bottom: every one of these is a real view of the same tower from a real heading. The tower does not care which way the drone was pointing, and the model should not either.
So the safe set for nadir footage is the four quarter turns and the two flips, and their combinations, which multiplies a few hundred frames by eight without inventing a pixel. Rotation by an arbitrary angle is also honest for the scene, and dishonest for the labels, because a box turned by thirty degrees is no longer axis-aligned. The usual fix draws a new box around the turned corners, and that box is larger than the insulator by a margin that grows with the angle. On an object a few dozen pixels across, the margin is most of the box.
There is an exception the corridor team hit in the first week. Some of the April frames were shot oblique, the drone looking forward and down at the tower rather than straight at it, and on those the sky is at the top and a vertical flip puts it at the bottom. The rule is per frame, or per flight: straight-down frames take every turn and flip, oblique frames take the horizontal flip only.
Scale stands in for the altitude the pilot chose
The pilot flew April at one height and will fly October at another, because the airspace notice changed and the battery plan with it. The same insulator string is a different number of pixels across on the two flights. The power line inspection page has this as its drift case: flight paths get re-planned, the asset has not changed, and the angle and scale it arrives at have.
Random scale during training is how the model sees both heights in April. Crop a region of the frame and resize it up, and the insulator is larger. Shrink the frame and pad it, and the insulator is smaller. The range should cover the heights the pilots actually fly, which is a number the flight plans hold, rather than a default range from a library. A model trained to find insulators from a wider band of heights than the corridor is ever flown at has spent its capacity on flights that will not happen.
The boxes scale with the pixels, exactly. A crop that cuts through a box keeps the part inside the crop, and if what is left is a sliver, the box is dropped rather than kept as a label for a few pixels of insulator.
Brightness covers the sun angle and the cloud
The April flight was midday under thin cloud. October will be late morning and lower sun, with the tower's shadow longer and the hardware on the sunlit side blown out. Brightness and contrast shifts, applied across the whole frame, give the model a version of April under October's light.
This is the augmentation where the weather drift page is worth reading first. Rain, fog and dust arrive for a week and detection falls off before recovering, and the fix it gives is to keep the bad-weather footage, because those days are rare in a training set and worth a great deal in it. A brightness shift stands in for weather the corridor has not been flown in yet. It does not replace a real flight under cloud, and when that flight happens, its frames go into the set as themselves.
Hue shifts are the one to be careful with here. Corrosion is found partly by colour, the orange of rust against the grey of galvanised steel, and a hue shift that turns the rust grey has produced a frame with a rust label and no rust in it. Brightness moves every colour together. Hue moves them apart.
Data augmentation moves the label with the pixels or it is wrong
My own view is that label handling is where aerial augmentation actually fails, and the transforms themselves almost never are. A quarter turn is a quarter turn. The box that goes with it has to have its corners turned by the same quarter and its new corners read off, and the polygon around a corroded patch has to have every vertex turned. A pipeline that turns the frame and forgets the polygon has a mask on the wrong part of the tower on every augmented copy.
The check is the one every labeling post ends up recommending: draw ten augmented frames with their transformed labels and look. Pick the ones with a box against the frame edge, the ones with a polygon, the ones that were scaled hardest. If the box sits on the insulator after the turn, the pipeline is right. On the corridor set the first draw showed every polygon sitting on the original coordinates while the frame beneath had been flipped, and it took a minute to find and a day to have avoided.
The evaluation frames are never augmented. They are a flight, labeled, held aside, and the October flight is the real test of whether April's augmentation did its job.
The October flight is the frames the set was missing
Whatever the augmentations covered, the October flight will show what they did not. Frames the model doubts, an insulator at a height the scale range did not reach, a tower in low sun with a shadow across the crossarm, come back to a person. The correction on each is a label from a real frame of the real corridor. When the corrections cross the project's threshold the model retrains, on April's frames and their turned and flipped copies and October's frames as themselves, and the new version replaces the old one with no downtime.
Somebody on the corridor team keeps the first frame from every flight pinned in the project, with the height and the time of day written on it. It is the list of conditions the set has actually seen, and the augmentations are only ever standing in for the ones it has not.
See it on your own footage.
Start with your footageMore in Labeling

Labeling · 6 min read
A collaborative data annotation workflow run as a pipeline
Batch the bottling line's frames by camera and shift, assign so nothing is boxed twice, attach the guideline, review every label, train on the approved set.
Rajiya Sultana · Oct 3, 2026

Labeling · 7 min read
Dataset health check for computer vision, what to look at before anything trains
A scratch dataset where every scratch sits in the centre of the frame will train a model that looks in the centre. Five counts to read before the first epoch.
Rajiya Sultana · Oct 3, 2026

Labeling · 6 min read
A dataset quality audit for object detection when mAP hides a weak class
The PPE camera at the rig site scored well on average and missed most bare heads. Audit the labels per class, fix the boxes in review, and retrain on the fixes.
Rob Hickey · Oct 3, 2026