Computer vision · 6 min read
ResNet-18 explained, what a residual backbone does underneath your detector
Why the shortcut connection made deep networks trainable, where the eighteen layers sit inside a detector on a plant camera, and why the frames matter more.
Summary
This post explains ResNet-18 by the problem it solved, deep networks that got worse as they got deeper, and the residual shortcut that fixed it. It walks through the eighteen layers as they appear inside a detector on a food line camera, from the stem to the feature maps the detection head reads, and it argues that the backbone matters less than the frames it is fine-tuned on. It is for engineers who see the name in a config file and want to know what it is doing.
Stephen Biswas · Engineer · Sep 28, 2026

Food processing line, fillets on a blue conveyor with the conveyor and rail boxed, generated scene with detections from our model
The config file for the fillet detector on line 4 has one line that names the backbone: a ResNet with eighteen layers. Nobody on the food line chose it. It was the default, it was small enough to run on the box beside the recorder, and it worked. What it is doing, between the frame of fillets on the blue conveyor and the boxes the detector returns, is worth understanding before the next decision about it is made.
Deep networks got worse as they got deeper, and that was the problem
Before ResNet, adding layers to an image network helped up to a point and then hurt. A network with twenty layers beat one with ten, and one with fifty was worse than either, on the training set as well as the test set. That last part was the puzzle. A deeper network can in principle do anything a shallower one can, by making its extra layers do nothing, and yet in practice the training could not find that solution. The extra layers made the network harder to train, and harder in a way that no amount of training time fixed.
The shortcut lets a layer learn the difference rather than the whole thing
The residual idea in ResNet is a wire around each pair of layers. Instead of asking a block of layers to produce its output from scratch, the block's output is added to its own input. The block now only has to learn what to change, the residual, and if the right answer is to change nothing, it can learn to output zero, which is easy.
That is the whole trick. The shortcut makes doing nothing the default, so depth stops costing anything, and the training signal has a clean path back through every block instead of fading as it passes through each one. Networks of a hundred and more layers became trainable overnight, and the eighteen-layer version became the small, fast member of the family that fits where the big ones do not.
Each block is two convolutions, a ReLU and a shortcut
Inside the eighteen layers the pattern repeats. A stem at the front, a single wide convolution and a pooling step, brings the frame down to a quarter of its size. Then four stages, each made of two blocks. A block is two small convolutions, three pixels by three, each followed by a normalisation and a ReLU, the function that passes positive values through and sets negative ones to zero, with the shortcut added in before the second ReLU. At each new stage the feature map halves in width and height and doubles in channels, so the network trades position for meaning as it goes deeper.
Count them up and the convolutions plus the final layer come to eighteen. The last of the four stages produces a coarse map of the frame, a few dozen positions across, where each position holds a vector describing what is there: fillet, conveyor, rail, the gap between fillets. On its own, with a pooling step and one more layer, that is a classifier. Inside a detector, it is the backbone.
An aside from line 4: the ReLU is the reason the feature maps are mostly zeros. Whole regions of the map go dark on a frame of empty conveyor, and the detection head learns to read the dark as absence.
The backbone sits under the neck and the head of the detector
A detector is three parts. The backbone reads the frame and produces feature maps at several scales. A neck combines the maps, so that the fine, early features that know where an edge is are joined with the coarse, late features that know what an object is. A head reads the combined maps and proposes boxes and classes. The ResNet is only the first of the three, and it is the part that is usually pretrained on a large general set before anyone on the food line touches it.
That pretraining is what makes fine-tuning work. The backbone arrives already knowing edges, textures and shapes from the world's pictures, and the training on line 4 teaches the neck and the head what a fillet looks like on this conveyor, while adjusting the backbone a little. A backbone trained from scratch on a few thousand fillet frames would learn far less.
The frames it is fine-tuned on matter more than the backbone
The eighteen-layer version is the smallest of the family. Deeper versions read a frame more carefully and cost more to run, and on a bench they score a little higher. On line 4 the difference between the small backbone and a deeper one was smaller than the difference between the first labeled set and the second.
The second set added the frames from the night shift and the week the lighting was changed.
My own view is that the backbone is the last thing to change when a detector underperforms. The frames it trained on, the tightness of their boxes, and whether the reviewer saw the night shift, each moves the result more than a deeper network does, and each is cheaper. The lesson on what vision can see is the right place to start when a model is failing, before the config file.
What matters on the box beside the recorder
The small backbone fits on modest hardware, which is why it was the default on line 4. The frame is sampled every couple of seconds, the model answers in a fraction of that, and the box in the cabinet stays cool. Input resolution is the setting to watch: a backbone that halves the frame four times has lost a small fillet by the last stage if the frame started small, and raising the resolution costs more than swapping backbones.
The deployment doc covers the export for the device and the test in the enclosure. The model is described by the device it runs on rather than by the network inside it. On the live line the frames it is unsure of come back to a person, the corrections retrain it, and the new version replaces the old one with no downtime. The eighteen layers are still there in the config file. What changed is what they were shown.
See it on your own footage.
Start with your footageMore in Computer vision

Computer vision · 6 min read
When the outline is the answer and a box will not do
A weld defect whose area decides the rework and a part a robot has to grasp both need a mask. Here is what the mask costs to label and when a box is enough.
Esdras Ntuyenabo · Sep 28, 2026

Computer vision · 6 min read
DETR explained, what the detection transformer changed about finding objects
Detection as set prediction, a fixed set of queries instead of anchors, no suppression step, and the small-object and slow-training problems later fixed.
Finn Ellingwood · Sep 28, 2026

Computer vision · 6 min read
Reading a training run before trusting the model
The loss curve says the run finished. The gap to validation, the recall on the crack class and the frames behind the score say whether the line can trust it.
Rob Hickey · Sep 28, 2026