Skip to content
LexDataLexData
PlatformIndustriesCustomers
DocsThe Field GuideBlogWhy models drift
AboutCareersSecurityContact
Log inStart now
← All posts

Industries · 7 min read

Object dimension measurement from a fixed camera, using the mask rather than the bounding box

A rotated carton needs a polygon, scale comes from a printed reference or depth, the carrier's bands decide pass or fail, and a nudged mount breaks it.

Summary

This post describes measuring length, width and height of cartons at a dimensioning station from a fixed overhead camera: a mask polygon instead of an axis-aligned box, a printed reference or a depth sensor for scale, and the carrier's tolerance bands as the pass or fail. It concludes that a reference printed on the station surface catches a nudged mount, and that a tilt does not get caught without re-verifying both ends. It is for packing and dispatch operations teams.

Stephen Biswas · Engineer · Sep 30, 2026

Parcel packing station from overhead, cartons, a crushed carton and a label boxed, generated scene with detections from our model

At 3 pm the dimensioning station at the end of packing line 2 has a queue. Each carton is placed on the mat, the operator runs a tape along three edges, reads the carrier's rate card taped to the bench, writes a band on the carton in marker, and lifts it onto the pallet for that band. The tape has its first few centimetres worn to the fabric. A carton set down at an angle gets measured at an angle, and a carton that lands on a band edge gets the band the operator thinks the carrier will accept.

An overhead camera over the mat can measure every carton without the tape. Whether its numbers can be trusted comes down to which shape it draws around the carton, and to what it uses as a ruler.

The bounding box is the wrong shape for a measurement

A detector returns a bounding box: a rectangle aligned with the image's edges. For a carton placed square to the camera the box is the carton's outline and its sides are the length and width in pixels. For a carton set down at an angle, which on line 2 is most of them, the box is the rectangle that encloses the rotated carton, and both its sides are longer than the carton's. The error grows with the angle, and at forty-five degrees a square carton's box is a much larger square.

The shape that survives rotation is the mask: a polygon traced around the carton's outline. From the polygon, the smallest rectangle that encloses it, at whatever angle it sits, gives the true length and width in pixels. You type "carton" once, Lexi proposes the polygon on every frame, and a person checks the edges before anything trains, since a polygon that cuts a corner or includes a tape flap is a measurement error on every carton afterwards.

A bounding box still has a job here. It is what finds the carton on the mat; the polygon is what measures it.

Scale comes from a reference printed on the station

A polygon in pixels on line 2 is not yet a length. The conversion needs something in the frame of known size, and the cheapest thing is a square printed on the mat itself, a known number of centimetres on each side, in view of the camera on every frame. Its size in pixels on each frame gives the pixels per centimetre for that frame, and the carton's polygon is converted through it.

My own view is that the reference should be printed on the mat and never on a card. A card gets moved to make room for a big carton, gets a coffee ring, and ends up in a drawer. A square printed on the mat is read on every frame whether anyone remembers it or not.

A depth camera is the other route. It returns a distance per pixel, so the carton's top face has a height above the mat directly, and the scale at the top face is known rather than estimated. It costs more than a printed square and it solves the next problem outright.

Height is the dimension a top camera cannot see

An overhead camera sees the carton's top face. It does not see how tall the carton is, and worse, the top face of a tall carton is closer to the lens than the mat is, so it covers more pixels per centimetre than the printed square does. A tall carton measured through the mat's scale reads longer and wider than it is, by an amount that depends on the height nobody has measured yet.

There are two honest fixes. The depth camera gives the height and the corrected scale at the top face in one pass. The single-camera fix is a second view: a low camera at the side of the mat that sees the carton's height against a printed scale on the back wall of the station, with the height then used to correct the top view's scale. Either way, the three dimensions are not independent, and a station that measures two and guesses the third is guessing the band.

The operator on line 2 knows this without the geometry. The tape goes along the top face for length and width and then down the side for height, and the marker band is written only after the third measurement.

The rate card's bands decide pass or fail rather than a number

The carrier serving line 2 does not want a length. It wants a band, and the bands have edges. A carton whose long side lands well inside a band is that band. A carton whose long side lands within the camera's own uncertainty of a band edge is not decided by the camera. It goes to a person with the frame, the three numbers, and the edge it is near, and the person puts the tape on it.

The package and label inspection use case names the constraint this shares with any decision on a line: it has to be made in the time between cartons, and a false reject costs as much as a false accept. On a dimensioning station the false reject is a carton sent to the person with the tape when the camera could have called it. The tolerance band for "near the edge" is tuned on a week of real cartons so that the person sees the ones that need them and not the rest.

A nudged mount breaks the calibration and the reference catches most of it

The failure to expect on line 2 is the bracket. The overhead camera is on an arm above a bench where people lift cartons all day, and one day the arm is knocked, or re-tensioned, and the camera is a few centimetres lower or a few degrees off. The polygons still find every carton. Every measurement is now off by the same fraction, in a direction nobody can see on the frame.

The drift catalog files this as a camera moved: someone nudged a lens and the model has been looking slightly past the thing ever since. The printed square absorbs the part of this that is a plain shift, since a lower camera makes the square larger in pixels and the scale corrects itself on the next frame.

What it cannot absorb is a tilt, because a tilted camera stretches the far side of the mat more than the near side, and a single square in one corner does not see that. Two squares, one at each end of the mat, do. The fix is to re-verify the scale at both squares and relabel a short window of frames from the new framing.

Tape flaps and overhangs are the frames that come back

LexData takes the carton model through its whole life. You type what to look for, Lexi puts a polygon on every carton in every frame, and a person checks each label before anything trains on it. The model then watches the camera over the mat on line 2, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime.

The frames that come back in the first weeks are a tape flap standing proud of the outline, a carton overhanging the mat's edge, two cartons placed together, and the operator's hand still on the box. Each is a polygon corrected by a person, and when the corrections cross the project's threshold a new version trains on line 2's own frames.

The tape stays at the station. When a carton is near a band edge, it is the tape the carrier's rate card was written for, and the camera's job is to make sure the tape only comes out for those.

See it on your own footage.

Start with your footage

More in Industries

Industries · 6 min read

Perimeter security with fixed cameras, object detection and a drone sent to look

A frame every two seconds is enough to catch a person at the fence, a CPU is enough to run it, and the drone is the second look rather than the detector.

Andreas Ohrvall · Sep 30, 2026

Industries · 7 min read

Food service QA with a camera over the tray packing line

Every component on the tray gets a box, the missing one is flagged before the sealer, and the alert count is read against the line's own history.

Ayman Quadir · Sep 30, 2026

Industries · 7 min read

Railway safety with trackside cameras, zones and a signaller who can live with the alerts

People and vehicles boxed, the track bed and crossing drawn as zones, the frame sent to the control room, and a false alarm rate a signaller will keep reading.

Rajiya Sultana · Sep 30, 2026

LexData
LexData

Product

  • Platform
  • Industries
  • Use cases

Resources

  • Docs
  • The Field Guide
  • Blog
  • Why models drift

Industries

  • Energy & utilities
  • Oil & gas
  • Agriculture
  • Manufacturing
  • Insurance
  • Retail
  • Robotics

Company

  • About
  • Customers
  • Careers
  • Contact

Trust

  • Security
  • Privacy
  • Terms

Stay updated

What we learn running vision models in production.

See everything.
Miss nothing.

Stay updated

What we learn running vision models in production.

Terms of use & Privacy policy

© 2026 LexData Labs · All rights reserved