This post follows fastai’s content on building, training and fine-tuning Neural Networks. It leverages one of the examples provided in their course involving classifying the breeds of dogs and cats.
The code for this example is stored on my GitHub.
Connecting to the data and pre-sizing the images
I connected to fastai’s PETS dataset, and used regex to extract the label names from the file names.
Full details on the Oxford-IIIT Pet dataset here.
Some of the code applied, using fastai’s DataBlock:
pets = DataBlock(blocks = (ImageBlock, CategoryBlock),
# telling the DataBlock that the data is image based
get_items=get_image_files,
# set random seed
splitter=RandomSplitter(seed=42),
# use regex to extract y values from names
get_y=using_attr(RegexLabeller(r'(.+)_\d+.jpg$'), 'name'),
# resize and transfom as part of presizing
item_tfms=Resize(460),
batch_tfms=aug_transforms(size=224, min_scale=0.75))
# pointing the DataBlock to the images directory
dls = pets.dataloaders(path/"images")
One of the lines above (item_tfms=Resize(460)) pre-sizes the images.
Pre-sizing is image preprocessing that does several things: give images same dimensions, so they can be transformed to tensors for GPU-based operations. Keeping image sizes consistent reduces the amount of distinct augmentation computations that are required, and makes for more efficient processing
Many common data augmentation transforms introduce spurious empty patches or degrade data. For example, simple rotation will introduce empty pixels into the image. Other techniques may interpolate pixels.
The solution works in two steps:
- Resize to a fairly large size, so all images are consistent.
- Combine the individual augmentations into one and perform the combined operation on the GPU once, instead of performing operations individually.
The resize creates images large enough to leave margins for further augmentation.
The GPU is used for all data augmentation, and all operations are done together, with a single interpolation at the end.
The fastai course content shows an interesting comparison between their pre-sizing operations (left) and the same transformations applied with a more traditional approach (right).

Here are a few examples from the datablock, to make sure the data is loaded correctly.

Fastai often recommends training to a simple model initially, rather than overengineering an elaborate model from the start. This helps establish a baseline and gives an idea of the data can train a model at all.
Below, applying the loaded data to a neural net, with error rate used as the metric, and resnet34 used for transfer learning.
learn = cnn_learner(dls, resnet34, metrics=error_rate)
learn.fine_tune(2)

That loss is the function chosen to optimize the parameters of the model. Fastai tries to select an appropriate loss function based on what kind of data and model being used. In this case, with image data and a categorical outcome, fastai defaults to using cross-entropy loss.
Cross-entropy loss is a loss function works efficiently with multiple categories. I have already written a blog post on it, here: https://frankiecoughlan.data.blog/2021/12/10/cross-entropy-loss/
Model Interpretation
A confusion matrix will help see where the model is doing well, and where it’s doing badly:
interp = ClassificationInterpretation.from_learner(learn)
interp.plot_confusion_matrix(figsize=(12,12), dpi=60)

That confusion matrix is a little awkward to read. The most_confused method will show the cells of the confusion matrix with the most incorrect predictions.
interp.most_confused(min_val=5)

Taking this as the baseline, the next section focuses on further improvements.
Improving the Model
The Learning Rate Finder
Finding the right learning rate is important:
Too low and it will take a long time to train. This wastes time, but can also cause overfitting.
Too high and it will overstep, overshooting the minimum loss. Repeating that can move the model away from the optimum solution, not closer.
This happens below with a learning rate of 10%
learn = cnn_learner(dls, resnet34, metrics=error_rate)
learn.fine_tune(1, base_lr=0.1)

One approach to find the optimum learning rate (from Leslie Smith in 2015) is called the learning rate finder. Start with a tiny learning rate, use it for one mini-batch, and measure the loss. Double the rate, use it for one mini-batch, and measure the loss again. The loss will improve because we’ve taken a slightly bigger step in the right direction. Repeat until the loss gets worse. At that point, the learning rate is too large. It’s common to select a learning rate slightly lower than that which produced the worsening loss. Two recommendations for choosing that point are (a) one order of magnitude lower that that where the minimum loss was achieved, or (b) the last point where the loss was clearly decreasing. (a) and (b) are often very similar.
The default learning rate for fastai is 1e-3.
learn = cnn_learner(dls, resnet34, metrics=error_rate)
lr_steep = learn.lr_find()

In the graph above, learning rates lower than 1e-4 won’t allow the model to train – the slope is flat.
Rates higher than 1e-1 will cause the model to diverge – positive slope.
The point that produces the minimum loss is tempting, but the slope is flat here too.
The sweet spot is the point in the graph with the steepest negative slope.
learn = cnn_learner(dls, resnet34, metrics=error_rate)
learn.fine_tune(2, base_lr=3e-3)

Unfreezing and Transfer Learning
Transfer learning in a nutshell:
We take a pretrained model that is finetuned for a particular task. It has linear layers, with a nonlinear activation function between each linear pair, and an activation function like softmax at the end. The final linear layer uses a matrix with enough columns that the output size = number of classes in the classification problem.
This final linear won’t help when we are transfer learning, because it is specifically designed to classify the categories in the original context. So when transfer learning that layer is discarded and replaced with a layer that provides the correct number of outputs for the new problem. This new linear layer is random, but the output of the entire model isn’t random because this final random layer sits on top of all the prior carefully trained layers.
The aim is to retain all the useful things learned in early layers and then apply them to our new problem. The trick is to freeze the weights for the prior layers, and tell the optimizer to only update the weights in the later randomly added layers.
Fastai automatically freezes pretrained layers when creating a model from a pretrained network. The fine_tune method makes fastai do two things:
Train the randomly added layers for one epoch, with all other layers frozen
Unfreeze all the layers, and train them for all the epochs required
The method has parameters that can change its behaviour, or we can call the underlying methods directly to get custom behaviour. fit_one_cycle is fastai’s method for training without fine_tune.
learn = cnn_learner(dls, resnet34, metrics=error_rate)
learn.fit_one_cycle(3, 3e-3)
Then we’ll unfreeze the model, and run lr_find again, because having more layers to train, and weights that have already been trained for three epochs, means our previously found learning rate isn’t appropriate any more.
The graph looks different this time – the flat line followed by the sharp increase is due to the fact that the model has been trained already. Looking for the point with maximum gradient won’t help here, and instead it makes sense to choose a point well before the sharp increase, like 1e-5.

Training at that learning rate improved the model a bit, but there’s more that can be done.

Discriminative Learning Rates
After unfreezing, all of the weights can be changed by further training. However the quality or usefulness of the pretrained weights is higher than that of the randomly added parameters. The pretrained weights have been trained over a high number of epochs with a lot of data. The implication is that those pretrained weights would be better off with a lower learning rate, since less drastic change is required.
Also, earlier layers typically have learned generally useful things – like detecting edges and gradients – which is useful for any task. Later layers are far more specific, and so may be less relevant to the new task. It seems reasonable to let the later layers fine tune more quickly than the earlier ones.
Fastai’s default approach is to use discriminative learning rates. It uses lower learning rates in the early layers and higher learning rate for the later layers.
In Fastai you can pass a first and last learning rate to the model (via a python slice object) and it will start training with the first rate, and finish with the last rate, making incremental steps in between.
See below for previous training reapplied with discriminative learning rates.
learn = cnn_learner(dls, resnet34, metrics=error_rate)
learn.fit_one_cycle(3, 3e-3)
learn.unfreeze()
learn.fit_one_cycle(12, lr_max=slice(1e-6,1e-4))


The training loss inmproves throughout, but the validation loss slows and plateaus. This is indicating the model is starting to overfit, and becoming overconfident in its predictions.
That doesn’t mean it’s getting less accurate. Accuracy continues to improve even as validation loss gets worse. The end goal is optimising the chosen metrics, not necessarily the loss.
Number of Epochs
If time is a limiting factor, then an initial approach could be to train for the number of epochs that you’re willing to wait for. Then looking at the training/validation loss plots and the metrics will help. If they are still improving in the final epochs then it might be worth training for longer.
If the validation loss get worse towards the end of training, then the model is first getting overconfident, then incorrectly memorizing the data. The latter is the real issue. If the validation loss is deteriorating while the other metrics are improving, then the model can still be performing better.
It’s tempting to run the model for 100 epochs, choose the epoch with the best metrics, and pick that model as the winner. However, this approach doesn’t allow the model to use the smallest rates on the epochs with the best metrics, so the model can’t really fine-tune around the correct area. A better approach is to retrain the model from scratch with the number of epochs adjusted based on where the results were previously good.
Deeper Architectures
In general (but this is a big generalisation), more parameters will lead to more accurate results. Deeper models also use more memory, so deeper models should use smaller batch sizxes to avoid ‘out of memory’ issues on the GPU. Training on deeper architectures can be sped up by using ‘mixed-precision training’, where floating points are reduced from 32 to 16 places where possible.
Trying a deeper architecture with mixed precision:
from fastai.callback.fp16 import *
learn = cnn_learner(dls, resnet50, metrics=error_rate).to_fp16()
learn.fine_tune(6, freeze_epochs=3)

In this case the deeper model doesn’t provide a massive improvement, but the runtime for each epoch has gone from 50 seconds to 5 minutes.



























































