Collaborative Filtering

Collaborative filtering is a general solution to the following problem: given behavioural features for items, what are the underlying factors that describe or connect them. This sounds very vague, but think of it this way: Netflix uses collaborative filtering to understand the preferences (underlying factors) that its users (items) have, so that it can make recommendations on what to watch.

From Wikipedia:

collaborative filtering is a method of making automatic predictions (filtering) about the interests of a user by collecting preferences or taste information from many users (collaborating). The underlying assumption of the collaborative filtering approach is that if a person A has the same opinion as a person B on an issue, A is more likely to have B’s opinion on a different issue than that of a randomly chosen person.

https://en.wikipedia.org/wiki/Collaborative_filtering

A Simple Dataset

Here’s a portion of a randomly generated dataset containing User IDs, Book IDs, and the users’ book ratings. It’s not in a format that is particularly easy to digest or interpret.

Here’s the same dataset in a format that makes it easier to see patterns within and between users.

The empty cells in the table above lack ratings. These imply that the user has not read this book. The challenge then, is to understand which books should be recommended to which users, and in what order. If a model can fill the blanks on the table accurately, we can use that to know what books to recommend (or not recommend).

One way to do this is to understand the factors that describe each book and whether the user liked that factor. So if we had information related to genre, writing style and tone and a user’s preference score for each of these factors, then we could derive an overall weighted preference score for the book by multiplying the score.

Assume the scores and factors range from -1 to 1, it might look like this:

Very_like_a_thriller: 0.9
Casual_writing_style: -0.7
Dark_tone: 0.6

A user that loves dark, dense, thrillers might have respective scores of 0.8, -0.7, and 0.8.
Their preference score for this book would be:
0.9×0.8 + (-0.7)x(-0.7) + 0.8×0.6 = 1.69
This calculation is the dot product.

Latent Factors

The underlying factors like genre and tone mentioned above are latent factors. Instead of specifying the latent factors, we can learn them using a general gradient descent approach.

We start by randomly initialising some parameters which represent a set of latent factors for each user. We’ll have to decide how many to use, but for this example I use 3. Each user has a set of factors and each book has a set of factors. In the image below, the latent factors for each user are in yellow and those for each book are in orange. They are random and need to be optimised.

The cells in white above are the predictions, based on the dot product of each user’s factors with each book’s factors. If the prediction is high then it implies the user will like the book, and if it low it implies the user will dislike the book. The max possible score is 10, and the min is 0, to align with the scoring system used in the original dataset.

The next step is to calculate the loss. Mean Squared Error is a good loss function to start with. The MSE for two matrices is calculated as follows:
Get the difference between the actual and predicted values
Square each value
Sum over the matrix
Divide by the number of elements in the matrix

We will move to python and another similar dataset to make this easier to work through.

MovieLens Data and Embeddings

This example uses a dataset called MovieLens. It contains millions of movie rankings. We will use a subset of 100,000 here. The data is available through a fastai function.

from fastai.collab import *
from fastai.tabular.all import *
path = untar_data(URLs.ML_100k)

ratings = pd.read_csv(path/'u.data', delimiter='\t', header=None,
                      names=['user','movie','rating','timestamp'])
ratings.head()

The table u.item contains the mapping of IDs to titles:

movies = pd.read_csv(path/'u.item',  delimiter='|', encoding='latin-1',
                     usecols=(0,1), names=('movie','title'), header=None)
movies.head()
ratings = ratings.merge(movies)
ratings.head()

The ratings table contains the data we need to model. We can build a DataLoaders object from it. By default, dataloaders takes the first column for the user, the second column for the item (here the movies), and the third column for the ratings. The ‘item_name’ option overrides this default.

dls = CollabDataLoaders.from_df(ratings, item_name='title', bs=64)
dls.show_batch()

With N factors, we need a matrix of N x N_Users for the users’ latent factors, and a matrix of N_Movies for the movies’ latent factors:

n_users  = len(dls.classes['user'])
n_movies = len(dls.classes['title'])
n_factors = 5

user_factors = torch.randn(n_users, n_factors)
movie_factors = torch.randn(n_movies, n_factors)

The result or score for a single user/movie combination is the dot product of the user’s latent factors and the movie’s latent factors. Mentally this is easy to conceptualise, and in something like Excel you might use lookups to select a user and a movie and perform the dot product multiplication, but that mechanism doesn’t really transfer to deep learning frameworks. Deep learning models are based on matrix products and activation functions, not lookups.

There is a trick to represent ‘lookups’ as a matrix product. We replace our indices with one-hot-encoded vectors. Here’s an example for index 3:

one_hot_3 = one_hot(3, n_users).float()

user_factors.t() @ one_hot_3

If we index directly into the matrix, we get the same vector.

user_factors[3]

So one-hot-encoding a matrix and multiplying it by the matrix of latent factors is the same as doing a lookup directly.

This works, but will use a lot of memory and time. PyTorch has a special layer that indexes into a vector using an integer, but has its derivative calculated in such a way that it’s identical to what it would have been if it had done a matrix multiplication with a one-hot-encoded vector. This is an embedding.

For some problems (like image recognition) the number of latent factors might seem obvious – for example images are represented in RGB form by three numbers.
For something like movie recommendations, how many factors are relevant? We don’t instruct the model on this, we let it learn for itself. Our model can analyse the relationships between users and movies to figure out which features seem important and which don’t.

So each user and each movie gets a random vector of a certain length, and those are learnable parameters. At each step the loss is calculated by comparing predictions to targets, and then the gradients of the loss with respect to the embedding vectors allow for parameter update using Stochastic Gradient Descent.

Initially the parameters are randomly generated and mean nothing, but they will be trained to represent underlying characteristics that drive movie preferences.

Collaborative Filtering from Scratch

The forward method is called when the module is called, and is passed any parameters included in the call. So the first attempt at the dot product model looks like this:

class DotProduct(Module):
    def __init__(self, n_users, n_movies, n_factors):
        self.user_factors = Embedding(n_users, n_factors)
        self.movie_factors = Embedding(n_movies, n_factors)
        
    def forward(self, x):
        users = self.user_factors(x[:,0])
        movies = self.movie_factors(x[:,1])
        return (users * movies).sum(dim=1)

The input of the model is a tensor of shape batch_size x 2, where the first column (x[:, 0]) contains the user IDs and the second column (x[:, 1]) contains the movie IDs.

x,y = dls.one_batch()
x.shape
x[:10,:]

Now we can use a plain Learner class:

model = DotProduct(n_users, n_movies, 50)
learn = Learner(dls, model, loss_func=MSELossFlat())

And fit our model:

learn.fit_one_cycle(5, 5e-3)

Before getting into the results or looking for improvements, it’s worth looking at the step-by-step architecture in place here.

The dataloaders object contains the input data, and the true scores for each user/movie pair (y).

The model takes the number of users, movies, and factors. Then it transforms them into matrices of user factors and movie factors. Then it dot-products each combination of those. These are the predicted scores.

The learner compares the true scores to the predicted ones, and calculates the loss. It increments the parameters (aka the factors) based on Gradient Descent. Then it recalculates the loss.

The first improvement: force the outputs to range from 0-5 using the sigmoid function. Empirical evidence will denomstrate that setting the range from 0 to 5.5 is more effective.

class DotProduct(Module):
    def __init__(self, n_users, n_movies, n_factors, y_range=(0,5.5)):
        self.user_factors = Embedding(n_users, n_factors)
        self.movie_factors = Embedding(n_movies, n_factors)
        self.y_range = y_range
        
    def forward(self, x):
        users = self.user_factors(x[:,0])
        movies = self.movie_factors(x[:,1])
        return sigmoid_range((users * movies).sum(dim=1), *self.y_range)

model = DotProduct(n_users, n_movies, 50)
learn = Learner(dls, model, loss_func=MSELossFlat())
learn.fit_one_cycle(5, 5e-3)

At this point we treat each user as equal – but in reality some users are more positive (or negative) than others. We also treat each movie as equal – but in reality some movies are better (or worse) than others.

We have weights, but we need biases. Biases will be a single number for each user and movie that represents the inherent positiveness or goodness of the respective user and movie.

Adjusting our model architecture:

class DotProductBias(Module):
    def __init__(self, n_users, n_movies, n_factors, y_range=(0,5.5)):
        self.user_factors = Embedding(n_users, n_factors)
        self.user_bias = Embedding(n_users, 1)
        self.movie_factors = Embedding(n_movies, n_factors)
        self.movie_bias = Embedding(n_movies, 1)
        self.y_range = y_range
        
    def forward(self, x):
        users = self.user_factors(x[:,0])
        movies = self.movie_factors(x[:,1])
        res = (users * movies).sum(dim=1, keepdim=True)
        res += self.user_bias(x[:,0]) + self.movie_bias(x[:,1])
        return sigmoid_range(res, *self.y_range)

To run through this architecture in detail again…

We have a matrix of N factors for U users, so that’s a U x N matrix. We have a matrix of N factors for M movies, so that’s a M x N matrix. We have a 1 bias each for U users, so that’s a U x 1 vector. We have a 1 bias each for M movies, so that’s a M x 1 vector.

The function retuns a single score, which is a dot product of the user factors by the movie factors, plus a user bias, plus a movie bias.

The learner then trains on this model:

model = DotProductBias(n_users, n_movies, 50)
learn = Learner(dls, model, loss_func=MSELossFlat())
learn.fit_one_cycle(5, 5e-3)

Validation loss gets worse after a few epochs – a sign of overfitting. One way to tackle this is data augmentation, but that’s not possible here. Another is weight decay:

Weight Decay

Weight decay, also called L2 regularlisation, means adding the sum of all the weights squared to your loss function. This will affect the gradients by including a contribution that encourages them to be as small.

One way to think of overfitting is this: With unrestrained coefficients, the model is free to find extreme values that allow it to jump from observed value to observed value. This makes it fit well on the training data but not so much on the untrained data. This is overfitting in a nutshell. Forcing the model to restrict the values of the coefficients makes for a smoother (and more generalised) fit.

model = DotProductBias(n_users, n_movies, 50)
learn = Learner(dls, model, loss_func=MSELossFlat())
learn.fit_one_cycle(5, 5e-3, wd=0.1)

Creating an Embedding Module

Recreating DotProductBias without using the Embedding class:

Optimisers require that they get all the parameters of a module from the module’s paramter class. If we add a tensor as an attribute to a Module, it will not automatically be included in parameters.

class T(Module):
    def __init__(self): self.a = torch.ones(3)

L(T().parameters())

Wrapping a tensor in the nn.Parameter class tells Module to treat the tensor as a parameter. It will then automatically call requires_grad_.

class T(Module):
    def __init__(self): self.a = nn.Parameter(torch.ones(3))

L(T().parameters())

We can create a tensor as a parameter, with random initialization, like so:

def create_params(size):
    return nn.Parameter(torch.zeros(*size).normal_(0, 0.01))

DotProductBias without Embedding looks like this:

class DotProductBias(Module):
    def __init__(self, n_users, n_movies, n_factors, y_range=(0,5.5)):
        self.user_factors = create_params([n_users, n_factors])
        self.user_bias = create_params([n_users])
        self.movie_factors = create_params([n_movies, n_factors])
        self.movie_bias = create_params([n_movies])
        self.y_range = y_range
        
    def forward(self, x):
        users = self.user_factors[x[:,0]]
        movies = self.movie_factors[x[:,1]]
        res = (users*movies).sum(dim=1)
        res += self.user_bias[x[:,0]] + self.movie_bias[x[:,1]]
        return sigmoid_range(res, *self.y_range)

Training again:

model = DotProductBias(n_users, n_movies, 50)
learn = Learner(dls, model, loss_func=MSELossFlat())
learn.fit_one_cycle(5, 5e-3, wd=0.1)

How to Interpret Embeddings and Biases

Here are the movies with the lowest values in the bias vector:

movie_bias = learn.model.movie_bias.squeeze()
idxs = movie_bias.argsort()[:5]
[dls.classes['title'][i] for i in idxs]

Here’s how to interpret this: These movies are generally not liked, even by people that like similar kinds of movies. We could have sorted movies by their average rating, but that would just tell us if the movie was generally liked. This is telling us that people don’t like watching these movies even if they otherwise enjoy these kinds of movies.

Same now for movies with the highest biases:

idxs = movie_bias.argsort(descending=True)[:5]
[dls.classes['title'][i] for i in idxs]

So these are the movies you might enjoy, even if you don’t normally like these kinds of movies.

The embedding matrices have a lot of factors, and it makes it hard to interpret them. Principal Component Analysis can identify the set of vectors that transforms the matrix of embedding factors to a smaller set of factors, while capturing much of the variance of the original matrix. This helps by aggregating correlated embeddings (and therefore similar embeddings).

#caption Representation of movies based on two strongest PCA components
#alt Representation of movies based on two strongest PCA components
g = ratings.groupby('title')['rating'].count()
top_movies = g.sort_values(ascending=False).index.values[:1000]
top_idxs = tensor([learn.dls.classes['title'].o2i[m] for m in top_movies])
movie_w = learn.model.movie_factors[top_idxs].cpu().detach()
movie_pca = movie_w.pca(3)
fac0,fac1,fac2 = movie_pca.t()
idxs = list(range(50))
X = fac0[idxs]
Y = fac2[idxs]
plt.figure(figsize=(12,12))
plt.scatter(X, Y)
for i, x, y in zip(top_movies[idxs], X, Y):
    plt.text(x,y,i, color=np.random.rand(3)*0.7, fontsize=11)
plt.show()

The model is roughly classifying some concept of classic and pop culture.

Using fastai.collab

Fastai’s collab_learner allows us to create and train a collaborative filtering model.

learn = collab_learner(dls, n_factors=50, y_range=(0, 5.5))
learn.fit_one_cycle(5, 5e-3, wd=0.1)

You can print the model to see the names of the layers.

learn.model

Recreating previous analysis

movie_bias = learn.model.i_bias.weight.squeeze()
idxs = movie_bias.argsort(descending=True)[:5]
[dls.classes['title'][i] for i in idxs]

Embedding Distance

We can calculate the distance between embeddings to measure similarity (or lack of it) between movies. The metric is pythagorean distance. In this sense movies are similar if the users that like the movies are similar. The most similar movie to Silence of the Lambs is:

movie_factors = learn.model.i_weight.weight
# idx = dls.classes['title'].o2i['Silence of the Lambs, The (1991)']
# idx = dls.classes['title'].o2i['Forrest Gump (1994)']
idx = dls.classes['title'].o2i['Terminator, The (1984)']
distances = nn.CosineSimilarity(dim=1)(movie_factors, movie_factors[idx][None])
idx = distances.argsort(descending=True)[1]
dls.classes['title'][idx]

Bootstrapping a Collaborative Filtering Model

Great, so you have a trained model. Given a user with a reasonable watch history, you can make recommendations based on similar users’ preferences.

Now, what do you do with a new user? They have no watch history to compare to other users. You could give a new user the mean of all the other users’ embedding vectors. But the mean of the other users’ vectors won’t necessarily make sense. If half the users like horror, and half the users like action, your new user will be recommended half-horror half-action movies.

Another option is to ask a set of questions, then assign an embedding vector based on the answers. This will give some control over the possible enbeddings.

The entire concept of collaborative filtering here is based on users’ ratings. If there are users that rate a lot, they can tilt the system towards their preferences. They can create feedback loops, where their tastes are overrepresented. It’s hard to quantify this overrepresentation though.

Deep Learning for Collaborative Filtering

For deep learning, we can concatenate the results of the embedding lookup. The resulting matrix can be passed through linear layers and nonlinear layers as normal.

Instead of getting the dot product we will concatenate the embeddings. Because of this the matrices can have different sizes. Fastais get_emb_sz returns recommended sizes – aka numbers of latent factors.

embs = get_emb_sz(dls)
embs

Let’s implement this class

class CollabNN(Module):
    def __init__(self, user_sz, item_sz, y_range=(0,5.5), n_act=100):
        self.user_factors = Embedding(*user_sz)
        self.item_factors = Embedding(*item_sz)
        self.layers = nn.Sequential(
            nn.Linear(user_sz[1]+item_sz[1], n_act),
            nn.ReLU(),
            nn.Linear(n_act, 1))
        self.y_range = y_range
        
    def forward(self, x):
        embs = self.user_factors(x[:,0]),self.item_factors(x[:,1])
        x = self.layers(torch.cat(embs, dim=1))
        return sigmoid_range(x, *self.y_range)

And use it to create a model:

model = CollabNN(*embs)

The Embedding layers are created in CollabNN. embs infers the sizes. self.layers is our mini-Neural Net and in forward we apply the embeddings, concatenate results, and pass through the mini neural net. Then sigmoid_range is applied.

learn = Learner(dls, model, loss_func=MSELossFlat())
learn.fit_one_cycle(5, 5e-3, wd=0.01)

Within fastai.collab, if you use collab_learner and pass use_nn=True it will apply the neural network as part of the collaborative filtering model. The code below makes two hidden layers, of size 100 and 50 respectively:

learn = collab_learner(dls, use_nn=True, y_range=(0, 5.5), layers=[100,50])
learn.fit_one_cycle(5, 5e-3, wd=0.1)

Conclusion

This application shows how we can learn the underlying factors within a dataset, and quantify aspects of the data. This can give us information that is otherwise unmeasured.

Classifying Pet Breeds and Fine-Tuning a Neural Net

This post follows fastai’s content on building, training and fine-tuning Neural Networks. It leverages one of the examples provided in their course involving classifying the breeds of dogs and cats.

The code for this example is stored on my GitHub.

Connecting to the data and pre-sizing the images

I connected to fastai’s PETS dataset, and used regex to extract the label names from the file names.
Full details on the Oxford-IIIT Pet dataset here.

Some of the code applied, using fastai’s DataBlock:

pets = DataBlock(blocks = (ImageBlock, CategoryBlock),
#                  telling the DataBlock that the data is image based
                 get_items=get_image_files, 
#                  set random seed
                 splitter=RandomSplitter(seed=42),
#                  use regex to extract y values from names
                 get_y=using_attr(RegexLabeller(r'(.+)_\d+.jpg$'), 'name'),
#                  resize and transfom as part of presizing
                 item_tfms=Resize(460),
                 batch_tfms=aug_transforms(size=224, min_scale=0.75))

# pointing the DataBlock to the images directory
dls = pets.dataloaders(path/"images")

One of the lines above (item_tfms=Resize(460)) pre-sizes the images.

Pre-sizing is image preprocessing that does several things: give images same dimensions, so they can be transformed to tensors for GPU-based operations. Keeping image sizes consistent reduces the amount of distinct augmentation computations that are required, and makes for more efficient processing

Many common data augmentation transforms introduce spurious empty patches or degrade data. For example, simple rotation will introduce empty pixels into the image. Other techniques may interpolate pixels.

The solution works in two steps:

  1. Resize to a fairly large size, so all images are consistent.
  2. Combine the individual augmentations into one and perform the combined operation on the GPU once, instead of performing operations individually.

The resize creates images large enough to leave margins for further augmentation.

The GPU is used for all data augmentation, and all operations are done together, with a single interpolation at the end.

The fastai course content shows an interesting comparison between their pre-sizing operations (left) and the same transformations applied with a more traditional approach (right).

Image from fastai’s course documentation.


Here are a few examples from the datablock, to make sure the data is loaded correctly.

A few examples of the loaded and prepared images

Fastai often recommends training to a simple model initially, rather than overengineering an elaborate model from the start. This helps establish a baseline and gives an idea of the data can train a model at all.

Below, applying the loaded data to a neural net, with error rate used as the metric, and resnet34 used for transfer learning.

learn = cnn_learner(dls, resnet34, metrics=error_rate)
learn.fine_tune(2)
A baseline run.

That loss is the function chosen to optimize the parameters of the model. Fastai tries to select an appropriate loss function based on what kind of data and model being used. In this case, with image data and a categorical outcome, fastai defaults to using cross-entropy loss.

Cross-entropy loss is a loss function works efficiently with multiple categories. I have already written a blog post on it, here: https://frankiecoughlan.data.blog/2021/12/10/cross-entropy-loss/

Model Interpretation

A confusion matrix will help see where the model is doing well, and where it’s doing badly:

interp = ClassificationInterpretation.from_learner(learn)
interp.plot_confusion_matrix(figsize=(12,12), dpi=60)
The confusion matrix with 37 categories.

That confusion matrix is a little awkward to read. The most_confused method will show the cells of the confusion matrix with the most incorrect predictions.

interp.most_confused(min_val=5)
The breeds that the model got most confused over.

Taking this as the baseline, the next section focuses on further improvements.

Improving the Model

The Learning Rate Finder

Finding the right learning rate is important:
Too low and it will take a long time to train. This wastes time, but can also cause overfitting.
Too high and it will overstep, overshooting the minimum loss. Repeating that can move the model away from the optimum solution, not closer.

This happens below with a learning rate of 10%

learn = cnn_learner(dls, resnet34, metrics=error_rate)
learn.fine_tune(1, base_lr=0.1)
Poor performance with a very high learning rate.

One approach to find the optimum learning rate (from Leslie Smith in 2015) is called the learning rate finder. Start with a tiny learning rate, use it for one mini-batch, and measure the loss. Double the rate, use it for one mini-batch, and measure the loss again. The loss will improve because we’ve taken a slightly bigger step in the right direction. Repeat until the loss gets worse. At that point, the learning rate is too large. It’s common to select a learning rate slightly lower than that which produced the worsening loss. Two recommendations for choosing that point are (a) one order of magnitude lower that that where the minimum loss was achieved, or (b) the last point where the loss was clearly decreasing. (a) and (b) are often very similar.

The default learning rate for fastai is 1e-3.

learn = cnn_learner(dls, resnet34, metrics=error_rate)
lr_steep = learn.lr_find()

In the graph above, learning rates lower than 1e-4 won’t allow the model to train – the slope is flat.
Rates higher than 1e-1 will cause the model to diverge – positive slope.
The point that produces the minimum loss is tempting, but the slope is flat here too.
The sweet spot is the point in the graph with the steepest negative slope.

learn = cnn_learner(dls, resnet34, metrics=error_rate)
learn.fine_tune(2, base_lr=3e-3)
Better performance with a single chosen learning rate.

Unfreezing and Transfer Learning

Transfer learning in a nutshell:

We take a pretrained model that is finetuned for a particular task. It has linear layers, with a nonlinear activation function between each linear pair, and an activation function like softmax at the end. The final linear layer uses a matrix with enough columns that the output size = number of classes in the classification problem.

This final linear won’t help when we are transfer learning, because it is specifically designed to classify the categories in the original context. So when transfer learning that layer is discarded and replaced with a layer that provides the correct number of outputs for the new problem. This new linear layer is random, but the output of the entire model isn’t random because this final random layer sits on top of all the prior carefully trained layers.

The aim is to retain all the useful things learned in early layers and then apply them to our new problem. The trick is to freeze the weights for the prior layers, and tell the optimizer to only update the weights in the later randomly added layers.

Fastai automatically freezes pretrained layers when creating a model from a pretrained network. The fine_tune method makes fastai do two things:
Train the randomly added layers for one epoch, with all other layers frozen
Unfreeze all the layers, and train them for all the epochs required

The method has parameters that can change its behaviour, or we can call the underlying methods directly to get custom behaviour. fit_one_cycle is fastai’s method for training without fine_tune.

learn = cnn_learner(dls, resnet34, metrics=error_rate)
learn.fit_one_cycle(3, 3e-3)

Then we’ll unfreeze the model, and run lr_find again, because having more layers to train, and weights that have already been trained for three epochs, means our previously found learning rate isn’t appropriate any more.
The graph looks different this time – the flat line followed by the sharp increase is due to the fact that the model has been trained already. Looking for the point with maximum gradient won’t help here, and instead it makes sense to choose a point well before the sharp increase, like 1e-5.

The learning rate finder on a transfer-learning model.

Training at that learning rate improved the model a bit, but there’s more that can be done.

Discriminative Learning Rates

After unfreezing, all of the weights can be changed by further training. However the quality or usefulness of the pretrained weights is higher than that of the randomly added parameters. The pretrained weights have been trained over a high number of epochs with a lot of data. The implication is that those pretrained weights would be better off with a lower learning rate, since less drastic change is required.

Also, earlier layers typically have learned generally useful things – like detecting edges and gradients – which is useful for any task. Later layers are far more specific, and so may be less relevant to the new task. It seems reasonable to let the later layers fine tune more quickly than the earlier ones.

Fastai’s default approach is to use discriminative learning rates. It uses lower learning rates in the early layers and higher learning rate for the later layers.

In Fastai you can pass a first and last learning rate to the model (via a python slice object) and it will start training with the first rate, and finish with the last rate, making incremental steps in between.

See below for previous training reapplied with discriminative learning rates.

learn = cnn_learner(dls, resnet34, metrics=error_rate)
learn.fit_one_cycle(3, 3e-3)
learn.unfreeze()
learn.fit_one_cycle(12, lr_max=slice(1e-6,1e-4))
Comparing training loss to validation loss.

The training loss inmproves throughout, but the validation loss slows and plateaus. This is indicating the model is starting to overfit, and becoming overconfident in its predictions.

That doesn’t mean it’s getting less accurate. Accuracy continues to improve even as validation loss gets worse. The end goal is optimising the chosen metrics, not necessarily the loss.

Number of Epochs

If time is a limiting factor, then an initial approach could be to train for the number of epochs that you’re willing to wait for. Then looking at the training/validation loss plots and the metrics will help. If they are still improving in the final epochs then it might be worth training for longer.

If the validation loss get worse towards the end of training, then the model is first getting overconfident, then incorrectly memorizing the data. The latter is the real issue. If the validation loss is deteriorating while the other metrics are improving, then the model can still be performing better.

It’s tempting to run the model for 100 epochs, choose the epoch with the best metrics, and pick that model as the winner. However, this approach doesn’t allow the model to use the smallest rates on the epochs with the best metrics, so the model can’t really fine-tune around the correct area. A better approach is to retrain the model from scratch with the number of epochs adjusted based on where the results were previously good.

Deeper Architectures

In general (but this is a big generalisation), more parameters will lead to more accurate results. Deeper models also use more memory, so deeper models should use smaller batch sizxes to avoid ‘out of memory’ issues on the GPU. Training on deeper architectures can be sped up by using ‘mixed-precision training’, where floating points are reduced from 32 to 16 places where possible.

Trying a deeper architecture with mixed precision:

from fastai.callback.fp16 import *
learn = cnn_learner(dls, resnet50, metrics=error_rate).to_fp16()
learn.fine_tune(6, freeze_epochs=3)
Results from a deeper architecture.

In this case the deeper model doesn’t provide a massive improvement, but the runtime for each epoch has gone from 50 seconds to 5 minutes.

Redundant Encoding in Data Visualisations

Redundant Encoding is the practice of adding multiple visual elements to a visualisation, to enhance effectiveness and ease-of-understanding. It is also referred to as redundancy. According to displayr.com, it can ‘improve the chances of a reader interpreting a visualization quickly and correctly’.

Redundant encoding applies to many elements of a visualisation, including colours, shapes, labels, and sizes.

Here is a simple barplot with some countries and their populations. All the information needed to interpret the graph is included, but it is not necessarily easy to do so.

# create dataset
height = [45, 7, 51, 6, 1]
bars = ('Argentina', 'Bulgaria', 'Colombia', 'Denmark', 'Estonia')
y_pos = np.arange(len(bars))
 
# Create horizontal bars
plt.barh(y_pos, height)
 
# Create names on the x-axis
plt.yticks(y_pos, bars)

plt.title('5 Countries and their Populations in Millions')

# Show graphic
plt.show()

Firstly, we’ll add some axis labels and reorder the data based on the value of the bar, not its place in the dataset.

# Create a data frame
df = pd.DataFrame ({
        'Group':  ['Argentina', 'Bulgaria', 'Colombia', 'Denmark', 'Estonia'],
        'Value': [45, 7, 51, 6, 1]
})

# Sort the table
df = df.sort_values(by=['Value'])

y_pos = np.arange(len(bars))

# Create horizontal bars
plt.barh(y=df.Group, width=df.Value);
 
# Create names on the x-axis
plt.yticks(y_pos, bars)

plt.title('5 Countries and their Populations in Millions')
plt.xlabel('Population (Millions)')
plt.ylabel('Countries')

# Show graphic
plt.show()

This simple change makes the graph easier to take in.

Another option is using colour to deliver a message. Here is an example of the same graph with colour being used to emphasize bar length. This might be an example of ‘less is more’. I am not sure in this case if the greyscale is adding to the graph or taking away from it.

# Create a data frame
df = pd.DataFrame ({
        'Group':  ['Argentina', 'Bulgaria', 'Colombia', 'Denmark', 'Estonia'],
        'Value': [45, 7, 51, 6, 1]
})

height = [45, 7, 51, 6, 1]

totalheight = sum(height)
# color = height/totalheight
color2 = [number / totalheight for number in height]
color2.sort(reverse=True)
color = [str(a) for a in color2]

# Sort the table
df = df.sort_values(by=['Value'])

y_pos = np.arange(len(bars))

# Create horizontal bars
plt.barh(y=df.Group, width=df.Value, color = color);
 
# Create names on the x-axis
plt.yticks(y_pos, bars)

plt.title('5 Countries and their Populations in Millions')
plt.xlabel('Population (Millions)')
plt.ylabel('Countries')


# Show graphic
plt.show()

Examples

Here are some examples from other sources of great examples of redundant encoding. From displayr.com, showing use of colours and labels.

Before
And After!

From cnothelfer.com, showing redundant encoding in a scatterplot:

Great example.

clauswilke.com is a wealth of great visualisations. Here’s a lovely example of redundanct encoding – a few seconds of examination provides effortless transfer of information.