Skip to content
LexDataLexData
PlatformIndustriesCustomers
DocsThe Field GuideBlogWhy models drift
AboutCareersSecurityContact
Log inStart now
← All posts

Computer vision · 6 min read

Oriented bounding box detection, when the box needs to turn with the object

A straight box on a panel tilted at forty-five degrees is mostly background, and on a packed array the boxes overlap. A rotated box fixes that, at a price.

Summary

This post explains oriented bounding boxes through a solar array seen at an angle from a drone and a part rotated on a belt, showing where a straight box fills with background and overlaps its neighbours. It covers what the rotated label adds, what it costs to draw, and where the straight box remains the right answer. It is for teams whose objects are long, thin, packed and rotated in the frame.

Stephen Biswas · Engineer · Oct 3, 2026

Drone pass along a solar farm, panels boxed with straight boxes across the rows, from a customer drone survey

The drone turns at the end of row 12 on the solar farm, and for a few seconds the panels below it sit at forty-five degrees in the frame. Each panel is a long rectangle, and the straight box the detector draws around it is a square whose corners are grass and whose middle is mostly the neighbouring panel. On the straight run along the row, the boxes were tight. On the turn, they are a mess of overlapping squares that a person can barely read.

Across town, a part on a conveyor lands on the belt at whatever angle it fell, and the same thing happens at a smaller scale. A straight box around a long bracket lying diagonally is a box that is mostly belt.

A straight bounding box on a tilted panel is mostly background

A bounding box is axis-aligned by default: its sides run parallel to the frame's edges. For an object that is also aligned with the frame, a panel in row 12 seen square-on from directly above, the box hugs it. For an object that is long and rotated, the box is the smallest upright rectangle that contains it, and for a thin object at forty-five degrees that rectangle can be mostly empty. The model is then trained to call that whole square "panel", and it learns the grass in the corners as part of what a panel looks like.

For a single object that is tolerable. The detector still finds the panel and the count is still right. What suffers is anything that uses the box's contents afterwards: a crop for a closer look, an area estimate, a comparison against the panel's expected size. All of those get the corners too.

A rotated box adds one angle and drops most of the background

An oriented bounding box carries the same four sides plus an angle, and the box turns to lie along the object. On the tilted panel at the end of row 12 it becomes a tight rectangle again, with the grass outside it and the neighbouring panel in its own box. The label is one number longer than a straight box, and the model that predicts it has one extra output per detection.

That one number is the whole difference. The box drawn on the tilted panel now contains the panel, the crop is the panel, and the size estimate is the panel's size. On the belt, the diagonal bracket gets a box that is bracket rather than belt, and a downstream check on whether the bracket is the right length has something to measure.

The drone pilot flies the array along the rows, so panels appear at an angle only on the turns, and the straight boxes are fine for most of every flight. The turn frames were the ones the reviewer kept sending back.

Overlap between neighbours is what breaks the straight box first

The clearest case for the rotated box is a packed scene. On the turn over the array at the end of row 12, the straight boxes of adjacent panels overlap heavily, because each square includes the corners of the panels beside it. A detector uses overlap to decide whether two boxes are the same object, and when every box overlaps its neighbour by half, the detector starts merging neighbouring panels into one and dropping the rest.

Rotated boxes on the same frame barely overlap at all, because each one lies along its own panel. The same packed array that produced a merged smear with straight boxes produces a clean row of separate detections with rotated ones. The power line and grid inspection use case has a related problem with insulator strings that hang at an angle from the crossarm, and the reasoning is the same: long, thin, rotated and close together is the shape that breaks the straight box.

Labeling a rotated box costs a third click and a convention

A straight box is two clicks, one corner and the opposite one. A rotated box is those two plus an angle, usually drawn as a line along the object's long side and then a width. It is not much slower per box. What costs time is the convention, and it has to be decided before the first frame is labeled.

Which way does the angle run? A panel at forty-five degrees and one at two hundred and twenty-five degrees are the same rectangle, and if one labeler measures the angle from the short side and another from the long side, the model is trained on two conventions and learns neither. The labeling with Lexi doc frames the class definition as the first thing written, and for rotated boxes the angle convention belongs in that definition, with an example frame from the array and one from the belt.

The QA pass on every label is where a convention slip shows up. A reviewer who sees a box lying across a panel rather than along it has found a labeler measuring from the wrong side, and one afternoon of that is easier to fix than a model that learned it.

Where the straight box is still the right answer

My own view is that oriented boxes are reached for more often than they are needed, usually because a frame from a drone looked untidy. If the objects in the frame are roughly square, or appear at one angle on most frames, the straight box is cheaper to label, simpler to train and easier to review, and the untidy corners cost nothing. A person on a site camera, a truck at dock 3, a bottle under a fill head: none of these turn in the frame in a way that matters.

The rotated box earns its place when three things hold at once: the objects are long relative to their width, they appear at many angles, and they sit close enough together that straight boxes overlap. The array on the turn has all three. The bracket on the belt has the first two, and whether it needs the third depends on whether a second bracket ever lands beside it. The what a vision model can and cannot see lesson asks what question the model is answering, and for a rotated box the question is nearly always a measurement or a count in a packed scene.

The rotated model comes back for review on the panels at odd angles

Once the model is on the drone footage, the frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime. For a rotated-box model on the Tuesday flight those doubted frames cluster where the angle is hardest: panels seen nearly end-on where the long side is short, a bracket standing on its edge, a frame where the drone's own roll has tilted the whole array.

The reviewer's correction on those frames is a dragged angle, and it is worth more than a dragged corner, because it is the number the straight box never had. The turn at the end of the row, which produced the mess that started the project, was clean by the third version.

See it on your own footage.

Start with your footage

More in Computer vision

Computer vision · 6 min read

What CLIP is and why typing a description finds the frame

A shared space for pictures and words lets an operator type a sentence and get the matching frames from a recorder. A way to look, and the boxes come after.

Rob Hickey · Oct 3, 2026

Computer vision · 6 min read

An end-to-end object detection workflow starts with decisions, and the first frame is labeled last

Tag, box, mask or keypoints. A class list with occlusion and size rules. A metric and an acceptance sentence. Then, and only then, the first camera's frames.

Ayman Quadir · Oct 3, 2026

Computer vision · 7 min read

Image recognition AI, explained on the cameras a site already owns

A tag, a box, a mask or keypoints, on a substation camera, a line camera and an aisle camera. How a model gets there, and what its score does not promise.

Esdras Ntuyenabo · Oct 3, 2026

LexData
LexData

Product

  • Platform
  • Industries
  • Use cases

Resources

  • Docs
  • The Field Guide
  • Blog
  • Why models drift

Industries

  • Energy & utilities
  • Oil & gas
  • Agriculture
  • Manufacturing
  • Insurance
  • Retail
  • Robotics

Company

  • About
  • Customers
  • Careers
  • Contact

Trust

  • Security
  • Privacy
  • Terms

Stay updated

What we learn running vision models in production.

See everything.
Miss nothing.

Stay updated

What we learn running vision models in production.

Terms of use & Privacy policy

© 2026 LexData Labs · All rights reserved