Computer vision · 6 min read
What a convolutional neural network does with a frame from a line camera
One small filter slides across every position in the frame, and that single design decision is why the network fits on a box beside the recorder.
Summary
This post explains a convolutional neural network through one frame from a packaging line camera, from the filter that slides across the frame to the coarse map the last layer hands to the head. It argues that weight sharing is the reason these networks learn from a few hundred frames and the reason a compact one still runs on a runner beside the recorder. It is for engineers who meet the term in a config file and want to know what it is doing to their frames.
Finn Ellingwood · Engineer · Oct 4, 2026

Packaging line, cartons with labels passing a scanner, carton and label boxed, generated scene with detections from our model
The camera over the packaging line on line 2 sends a frame every couple of seconds: a carton on the belt, a label on the carton, the scanner arch just out of shot. The model that decides whether the label is there and straight is a convolutional neural network, which is the name printed at the top of its config file and, for most people on the line, the last thing they read about it.
What the network does with that frame is simpler than the name suggests, and one decision explains most of it.
A fully connected network has to learn every pixel position on its own
The obvious way to wire a network to a frame is to connect every pixel to every unit in the first layer. A frame from the line camera has close to a million pixels, so the first layer alone would carry more weights than the line will produce frames in a year, and every one of those weights has to be learned from examples.
The deeper problem is position. To a fully connected network the carton in the top left of the frame and the same carton in the bottom right are unrelated inputs, arriving on different wires, and a label learned in one place is not recognised in the other. The network would need to see the label at every position it might ever occupy, which the belt on line 2 never provides, because cartons arrive along one path.
A convolution slides one small filter across every position
A convolution keeps a small grid of weights, three pixels by three is common, and slides it across the frame from line 2, one step at a time, producing at each position a single number that says how strongly the pattern under the filter matches. The result is a map of the frame: bright where the pattern was found, dark where it was not.
The same nine weights are used at every position. That is weight sharing, and it is the whole trick. An edge of a label in the corner of the frame and an edge in the centre are the same pattern, found by the same filter, so the network learns it once. A layer has dozens of such filters, each learning its own pattern, and all of them are cheap to store and cheap to run compared with a wire per pixel.
The first layer's filters, when someone bothers to look at them after training, are almost always edges and colour steps. Nobody tells the network to start with edges; it arrives there because edges are what a frame is made of.
Each layer adds meaning and gives up position
Stack the layers and the patterns compound, and on the line 2 frame the compounding is easy to follow. The first finds edges, the second finds corners and the bar pattern of a barcode from the edges, a deeper one finds the rectangle of a label from the corners. Between layers a stride or a pooling step halves the map in width and height, so the deeper maps are coarser and each position in them describes a larger patch of the original frame.
By the last convolution the frame has become a map a few dozen positions across, where each position holds a vector describing what is in that patch: carton, label, belt, the gap between cartons. The network has traded precise position for meaning. That is the right trade for a question like "is there a label on this carton" and the wrong one for "which pixel is the corner of the label", and the detector families spend most of their design on getting some of that position back.
The head reads the final map and turns it into an answer
Everything above is the backbone. A head sits on top and decides what the answer looks like. A classification head pools the last map into one vector and produces a tag for the whole frame, label present or label missing, which is enough for line 2. A detection head reads maps at several depths, so it can use the early ones that still know where an edge is, and proposes boxes with a class on each.
The two heads share the backbone. That is why a network pretrained to classify a large general set can be turned into a detector on the packaging line with a few hundred checked frames. The filters that find edges and corners are already there, and only the head and the last layers have to learn what a label on this belt looks like.
Weight sharing generalises across position and not across appearance
The filter that finds the label in the corner also finds it in the centre, on a frame the network never saw in training. That is what convolution buys. It does not buy a new label design. When packaging switches to a matte stock with a smaller barcode in November, the patterns the deeper layers learned are simply absent, and the network's answer on those frames is a guess dressed as a tag.
The lesson on what a vision model can and cannot see puts the boundary plainly: a model is strong on what is visible and consistent, and consistent means consistent in the frames it trained on. On line 2 the doubted frames after the packaging change went back to a person, the corrections retrained the model, and the new version took over with no downtime. The filters at the bottom of the network barely moved; the ones near the top learned the new stock.
An aside from the same line: the scanner arch still reads every barcode it can, and the camera was never meant to replace it. The camera's job is the carton the scanner reports as no read, because the label is skewed, missing or covered by tape, and the scanner cannot say which.
Compact convolutional models still fit on the runner beside the recorder
A convolution is a regular, repeated operation, and small chips have been built around exactly that shape. A compact convolutional network on a Jetson in the line cabinet takes a frame about every two seconds and returns its answer before the next one arrives, with no round trip to anywhere. The model leaves the platform as PT, ONNX or TorchScript through a preset that names the device, as the deployment doc describes, and the footage stays in the cabinet; only the doubted frames leave.
My own view is that the attention-based models get the conference talks and the convolutional ones get the cabinets, and that this will still be true for a while. A network with a few million weights that finds a label in a frame on a box that costs less than the camera is a very good answer to the question line 2 asked.
See it on your own footage.
Start with your footageMore in Computer vision

Computer vision · 7 min read
Five computer vision applications in production, on the cameras a site already owns
Detect, count, inspect, track and read are the five jobs a fixed camera can be given, and each is a different question with a different label behind it.
Ayman Quadir · Oct 4, 2026

Computer vision · 6 min read
Computer vision projects worth building on the cameras you already have
A plant, a utility, a grower, a store and a warehouse each have a project that starts on an existing camera, produces a decision, and has someone to act on it.
Ayman Quadir · Oct 4, 2026

Computer vision · 6 min read
How to choose an object detection model architecture for a camera on your own site
Start from the decision the plant has to make, then the box beside the recorder the model has to fit, and only then the family. The benchmark comes last.
Andreas Ohrvall · Oct 4, 2026