Computer vision · 7 min read
What a vision transformer does with a frame, and what changes in monitoring when you use one
The frame becomes patches, every patch looks at every other, and a tower over a hedge gets easier. The cost grows with resolution. The review queue stays.
Summary
This post explains what a vision transformer does with a frame from an aerial survey of a transmission line: how the frame is cut into patches, how attention lets every patch see every other, why that helps in cluttered scenes and what it costs at high resolution. It concludes that the architecture changes the training start and the compute budget, and changes nothing about how the model is monitored or retrained. It is for teams deciding whether the newer architecture is worth the switch.
Rob Hickey · Chief AI Officer · Oct 3, 2026

Three transmission towers across a field, insulators boxed, defects flagged and vegetation marked near the line, from a customer aerial survey
The survey drone's frame over the transmission line north of the Dee crossing is busy. Three towers at different distances, a dozen insulator strings, a conductor crossing the whole width, a hedge that has grown up under the line, a road, a field with tractor lines that look, from the air, like conductors.
A detector built on convolutions looks at that frame through a series of small windows, building up from edges to shapes to objects, and it does well on the insulators, which are compact and local. It does less well on the question of whether that hedge is under the conductor, because the hedge is at the bottom of the frame and the conductor is a thin line across the top, and a small window never sees both.
A vision transformer is a different way of looking at the same frame, and the difference is in what each part of the frame gets to see.
A vision transformer cuts the frame into patches and treats them as words
The transformer was built for sentences, where a word's meaning depends on the words around it, however far away they are. To use it on a frame, the frame is cut into a grid of small square patches, each patch is flattened into a list of numbers, and each list is given a marker for where in the grid it came from. The result is a sequence of patches, in the same shape as a sequence of words, and the model processes it the way a language model processes a sentence.
That framing sounds like a trick and it mostly is not. A patch from the Dee crossing frame is a small square of pixels, sixteen on a side or thereabouts, and by itself it is a piece of sky, or a piece of hedge, or a piece of conductor. What the model does with the sequence is what matters.
Attention lets every patch look at every other patch
The operation at the centre of a transformer is attention. For each patch, the model computes how much every other patch in the frame is relevant to it, and then builds the patch's new representation as a weighted mix of all of them. A patch of conductor at the top of the Dee crossing frame can attend to a patch of hedge at the bottom, in one step, without anything in between. That is the thing a convolution cannot do until its windows have grown, layer by layer, to cover the whole frame.
Stacked a dozen or more times, with several attention patterns running side by side in each layer, this gives every patch a representation that knows about the whole frame. The insulator patch knows it is on a tower; the tower patch knows there are two more behind it; the hedge patch knows the conductor is above it.
The survey drone's gimbal drifts over a long flight, so the horizon tilts a few degrees by the end of each pass, and the position markers on the patches are what let the model keep track of which way is up.
Whole-frame attention helps in cluttered scenes like a tower over a hedge
Where this pays off is the cluttered frame. Vegetation under the line at the Dee crossing is a relationship between two things far apart in the frame, and a model that can relate them in a single step finds it more reliably than one that has to assemble the relationship from local evidence. The same holds for a tower partly behind another tower, a conductor that disappears into the sky's brightness and reappears, a defect whose meaning depends on where on the tower it sits.
It is less of an advantage on a frame with one thing in it. A cap on a bottle under a fill head is a local question, and a convolutional detector answers it as well as a transformer does, faster and on smaller hardware. The what a vision model can and cannot see lesson makes the point that the question comes first, and whether the question involves things far apart in the frame is a fair guide to whether the architecture matters.
The cost grows with resolution and shows on a small runner
Attention compares every patch with every other, so its cost grows with the square of the number of patches. Double the frame's width and height and the patch count goes up four times, and the attention cost goes up sixteen. A Dee crossing survey frame at full resolution has a great many patches, and the transformer that handles it is a large model that needs a capable GPU.
That matters for where the model runs. On a runner beside the recorder, the kind of box the deployment doc describes for a site that wants its footage to stay on site, a large transformer at full resolution is a tight fit. The usual answers are to run at a reduced resolution, to tile the frame and run on the tiles, or to run the transformer in the cloud and a smaller model at the edge. None of those is wrong. Each is a decision the site makes with the compute it has.
A transformer needs more frames to start and a checkpoint solves that
Convolutions carry an assumption that nearby pixels matter to each other, and that assumption is a head start when frames are few. A transformer starts with no such assumption and has to learn from the frames that nearby patches are related, which takes a great many frames when training from nothing. That was the reputation the architecture earned early, and it was accurate for training from scratch.
Almost nobody trains from scratch. A transformer checkpoint pretrained on a large general set, COCO or something broader, has already learned the spatial assumptions and a great deal else. Adapting it to the survey frames takes a few hundred reviewed frames, the same way it does for any other backbone. The retraining without starting over lesson applies unchanged: each version begins from the last, with the site's corrections on top.
Monitoring and retraining do not change with the architecture
My own view is that the architecture question gets more attention than it deserves from teams who have not yet built their review queue, and that the queue matters more than the model. Whatever produces the boxes, the boxes arrive with a score, the doubted frames come back to a person, and the person's corrections are the signal. A transformer that is drifting because the survey moved to a new region shows it the same way a convolutional model does: frames coming back, corrections landing, the override rate rising.
LexData takes the survey model through its whole life without caring which backbone it has. You type what to look for, Lexi puts a box on every insulator and every encroaching hedge in every frame, and a person checks each label before anything trains on it. The model then watches the footage, in the cloud, on your servers, or on a runner beside the recorder. Frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime. The survey team switched backbones between the third and fourth versions, and the review queue on the Monday after looked the same as it had the Monday before.
See it on your own footage.
Start with your footageMore in Computer vision

Computer vision · 6 min read
What CLIP is and why typing a description finds the frame
A shared space for pictures and words lets an operator type a sentence and get the matching frames from a recorder. A way to look, and the boxes come after.
Rob Hickey · Oct 3, 2026

Computer vision · 6 min read
An end-to-end object detection workflow starts with decisions, and the first frame is labeled last
Tag, box, mask or keypoints. A class list with occlusion and size rules. A metric and an acceptance sentence. Then, and only then, the first camera's frames.
Ayman Quadir · Oct 3, 2026

Computer vision · 7 min read
Image recognition AI, explained on the cameras a site already owns
A tag, a box, a mask or keypoints, on a substation camera, a line camera and an aisle camera. How a model gets there, and what its score does not promise.
Esdras Ntuyenabo · Oct 3, 2026