Collaborative Filtering

Collaborative filtering is a general solution to the following problem: given behavioural features for items, what are the underlying factors that describe or connect them. This sounds very vague, but think of it this way: Netflix uses collaborative filtering to understand the preferences (underlying factors) that its users (items) have, so that it can make recommendations on what to watch.

From Wikipedia:

collaborative filtering is a method of making automatic predictions (filtering) about the interests of a user by collecting preferences or taste information from many users (collaborating). The underlying assumption of the collaborative filtering approach is that if a person A has the same opinion as a person B on an issue, A is more likely to have B’s opinion on a different issue than that of a randomly chosen person.

https://en.wikipedia.org/wiki/Collaborative_filtering

A Simple Dataset

Here’s a portion of a randomly generated dataset containing User IDs, Book IDs, and the users’ book ratings. It’s not in a format that is particularly easy to digest or interpret.

Here’s the same dataset in a format that makes it easier to see patterns within and between users.

The empty cells in the table above lack ratings. These imply that the user has not read this book. The challenge then, is to understand which books should be recommended to which users, and in what order. If a model can fill the blanks on the table accurately, we can use that to know what books to recommend (or not recommend).

One way to do this is to understand the factors that describe each book and whether the user liked that factor. So if we had information related to genre, writing style and tone and a user’s preference score for each of these factors, then we could derive an overall weighted preference score for the book by multiplying the score.

Assume the scores and factors range from -1 to 1, it might look like this:

Very_like_a_thriller: 0.9
Casual_writing_style: -0.7
Dark_tone: 0.6

A user that loves dark, dense, thrillers might have respective scores of 0.8, -0.7, and 0.8.
Their preference score for this book would be:
0.9×0.8 + (-0.7)x(-0.7) + 0.8×0.6 = 1.69
This calculation is the dot product.

Latent Factors

The underlying factors like genre and tone mentioned above are latent factors. Instead of specifying the latent factors, we can learn them using a general gradient descent approach.

We start by randomly initialising some parameters which represent a set of latent factors for each user. We’ll have to decide how many to use, but for this example I use 3. Each user has a set of factors and each book has a set of factors. In the image below, the latent factors for each user are in yellow and those for each book are in orange. They are random and need to be optimised.

The cells in white above are the predictions, based on the dot product of each user’s factors with each book’s factors. If the prediction is high then it implies the user will like the book, and if it low it implies the user will dislike the book. The max possible score is 10, and the min is 0, to align with the scoring system used in the original dataset.

The next step is to calculate the loss. Mean Squared Error is a good loss function to start with. The MSE for two matrices is calculated as follows:
Get the difference between the actual and predicted values
Square each value
Sum over the matrix
Divide by the number of elements in the matrix

We will move to python and another similar dataset to make this easier to work through.

MovieLens Data and Embeddings

This example uses a dataset called MovieLens. It contains millions of movie rankings. We will use a subset of 100,000 here. The data is available through a fastai function.

from fastai.collab import *
from fastai.tabular.all import *
path = untar_data(URLs.ML_100k)

ratings = pd.read_csv(path/'u.data', delimiter='\t', header=None,
                      names=['user','movie','rating','timestamp'])
ratings.head()

The table u.item contains the mapping of IDs to titles:

movies = pd.read_csv(path/'u.item',  delimiter='|', encoding='latin-1',
                     usecols=(0,1), names=('movie','title'), header=None)
movies.head()
ratings = ratings.merge(movies)
ratings.head()

The ratings table contains the data we need to model. We can build a DataLoaders object from it. By default, dataloaders takes the first column for the user, the second column for the item (here the movies), and the third column for the ratings. The ‘item_name’ option overrides this default.

dls = CollabDataLoaders.from_df(ratings, item_name='title', bs=64)
dls.show_batch()

With N factors, we need a matrix of N x N_Users for the users’ latent factors, and a matrix of N_Movies for the movies’ latent factors:

n_users  = len(dls.classes['user'])
n_movies = len(dls.classes['title'])
n_factors = 5

user_factors = torch.randn(n_users, n_factors)
movie_factors = torch.randn(n_movies, n_factors)

The result or score for a single user/movie combination is the dot product of the user’s latent factors and the movie’s latent factors. Mentally this is easy to conceptualise, and in something like Excel you might use lookups to select a user and a movie and perform the dot product multiplication, but that mechanism doesn’t really transfer to deep learning frameworks. Deep learning models are based on matrix products and activation functions, not lookups.

There is a trick to represent ‘lookups’ as a matrix product. We replace our indices with one-hot-encoded vectors. Here’s an example for index 3:

one_hot_3 = one_hot(3, n_users).float()

user_factors.t() @ one_hot_3

If we index directly into the matrix, we get the same vector.

user_factors[3]

So one-hot-encoding a matrix and multiplying it by the matrix of latent factors is the same as doing a lookup directly.

This works, but will use a lot of memory and time. PyTorch has a special layer that indexes into a vector using an integer, but has its derivative calculated in such a way that it’s identical to what it would have been if it had done a matrix multiplication with a one-hot-encoded vector. This is an embedding.

For some problems (like image recognition) the number of latent factors might seem obvious – for example images are represented in RGB form by three numbers.
For something like movie recommendations, how many factors are relevant? We don’t instruct the model on this, we let it learn for itself. Our model can analyse the relationships between users and movies to figure out which features seem important and which don’t.

So each user and each movie gets a random vector of a certain length, and those are learnable parameters. At each step the loss is calculated by comparing predictions to targets, and then the gradients of the loss with respect to the embedding vectors allow for parameter update using Stochastic Gradient Descent.

Initially the parameters are randomly generated and mean nothing, but they will be trained to represent underlying characteristics that drive movie preferences.

Collaborative Filtering from Scratch

The forward method is called when the module is called, and is passed any parameters included in the call. So the first attempt at the dot product model looks like this:

class DotProduct(Module):
    def __init__(self, n_users, n_movies, n_factors):
        self.user_factors = Embedding(n_users, n_factors)
        self.movie_factors = Embedding(n_movies, n_factors)
        
    def forward(self, x):
        users = self.user_factors(x[:,0])
        movies = self.movie_factors(x[:,1])
        return (users * movies).sum(dim=1)

The input of the model is a tensor of shape batch_size x 2, where the first column (x[:, 0]) contains the user IDs and the second column (x[:, 1]) contains the movie IDs.

x,y = dls.one_batch()
x.shape
x[:10,:]

Now we can use a plain Learner class:

model = DotProduct(n_users, n_movies, 50)
learn = Learner(dls, model, loss_func=MSELossFlat())

And fit our model:

learn.fit_one_cycle(5, 5e-3)

Before getting into the results or looking for improvements, it’s worth looking at the step-by-step architecture in place here.

The dataloaders object contains the input data, and the true scores for each user/movie pair (y).

The model takes the number of users, movies, and factors. Then it transforms them into matrices of user factors and movie factors. Then it dot-products each combination of those. These are the predicted scores.

The learner compares the true scores to the predicted ones, and calculates the loss. It increments the parameters (aka the factors) based on Gradient Descent. Then it recalculates the loss.

The first improvement: force the outputs to range from 0-5 using the sigmoid function. Empirical evidence will denomstrate that setting the range from 0 to 5.5 is more effective.

class DotProduct(Module):
    def __init__(self, n_users, n_movies, n_factors, y_range=(0,5.5)):
        self.user_factors = Embedding(n_users, n_factors)
        self.movie_factors = Embedding(n_movies, n_factors)
        self.y_range = y_range
        
    def forward(self, x):
        users = self.user_factors(x[:,0])
        movies = self.movie_factors(x[:,1])
        return sigmoid_range((users * movies).sum(dim=1), *self.y_range)

model = DotProduct(n_users, n_movies, 50)
learn = Learner(dls, model, loss_func=MSELossFlat())
learn.fit_one_cycle(5, 5e-3)

At this point we treat each user as equal – but in reality some users are more positive (or negative) than others. We also treat each movie as equal – but in reality some movies are better (or worse) than others.

We have weights, but we need biases. Biases will be a single number for each user and movie that represents the inherent positiveness or goodness of the respective user and movie.

Adjusting our model architecture:

class DotProductBias(Module):
    def __init__(self, n_users, n_movies, n_factors, y_range=(0,5.5)):
        self.user_factors = Embedding(n_users, n_factors)
        self.user_bias = Embedding(n_users, 1)
        self.movie_factors = Embedding(n_movies, n_factors)
        self.movie_bias = Embedding(n_movies, 1)
        self.y_range = y_range
        
    def forward(self, x):
        users = self.user_factors(x[:,0])
        movies = self.movie_factors(x[:,1])
        res = (users * movies).sum(dim=1, keepdim=True)
        res += self.user_bias(x[:,0]) + self.movie_bias(x[:,1])
        return sigmoid_range(res, *self.y_range)

To run through this architecture in detail again…

We have a matrix of N factors for U users, so that’s a U x N matrix. We have a matrix of N factors for M movies, so that’s a M x N matrix. We have a 1 bias each for U users, so that’s a U x 1 vector. We have a 1 bias each for M movies, so that’s a M x 1 vector.

The function retuns a single score, which is a dot product of the user factors by the movie factors, plus a user bias, plus a movie bias.

The learner then trains on this model:

model = DotProductBias(n_users, n_movies, 50)
learn = Learner(dls, model, loss_func=MSELossFlat())
learn.fit_one_cycle(5, 5e-3)

Validation loss gets worse after a few epochs – a sign of overfitting. One way to tackle this is data augmentation, but that’s not possible here. Another is weight decay:

Weight Decay

Weight decay, also called L2 regularlisation, means adding the sum of all the weights squared to your loss function. This will affect the gradients by including a contribution that encourages them to be as small.

One way to think of overfitting is this: With unrestrained coefficients, the model is free to find extreme values that allow it to jump from observed value to observed value. This makes it fit well on the training data but not so much on the untrained data. This is overfitting in a nutshell. Forcing the model to restrict the values of the coefficients makes for a smoother (and more generalised) fit.

model = DotProductBias(n_users, n_movies, 50)
learn = Learner(dls, model, loss_func=MSELossFlat())
learn.fit_one_cycle(5, 5e-3, wd=0.1)

Creating an Embedding Module

Recreating DotProductBias without using the Embedding class:

Optimisers require that they get all the parameters of a module from the module’s paramter class. If we add a tensor as an attribute to a Module, it will not automatically be included in parameters.

class T(Module):
    def __init__(self): self.a = torch.ones(3)

L(T().parameters())

Wrapping a tensor in the nn.Parameter class tells Module to treat the tensor as a parameter. It will then automatically call requires_grad_.

class T(Module):
    def __init__(self): self.a = nn.Parameter(torch.ones(3))

L(T().parameters())

We can create a tensor as a parameter, with random initialization, like so:

def create_params(size):
    return nn.Parameter(torch.zeros(*size).normal_(0, 0.01))

DotProductBias without Embedding looks like this:

class DotProductBias(Module):
    def __init__(self, n_users, n_movies, n_factors, y_range=(0,5.5)):
        self.user_factors = create_params([n_users, n_factors])
        self.user_bias = create_params([n_users])
        self.movie_factors = create_params([n_movies, n_factors])
        self.movie_bias = create_params([n_movies])
        self.y_range = y_range
        
    def forward(self, x):
        users = self.user_factors[x[:,0]]
        movies = self.movie_factors[x[:,1]]
        res = (users*movies).sum(dim=1)
        res += self.user_bias[x[:,0]] + self.movie_bias[x[:,1]]
        return sigmoid_range(res, *self.y_range)

Training again:

model = DotProductBias(n_users, n_movies, 50)
learn = Learner(dls, model, loss_func=MSELossFlat())
learn.fit_one_cycle(5, 5e-3, wd=0.1)

How to Interpret Embeddings and Biases

Here are the movies with the lowest values in the bias vector:

movie_bias = learn.model.movie_bias.squeeze()
idxs = movie_bias.argsort()[:5]
[dls.classes['title'][i] for i in idxs]

Here’s how to interpret this: These movies are generally not liked, even by people that like similar kinds of movies. We could have sorted movies by their average rating, but that would just tell us if the movie was generally liked. This is telling us that people don’t like watching these movies even if they otherwise enjoy these kinds of movies.

Same now for movies with the highest biases:

idxs = movie_bias.argsort(descending=True)[:5]
[dls.classes['title'][i] for i in idxs]

So these are the movies you might enjoy, even if you don’t normally like these kinds of movies.

The embedding matrices have a lot of factors, and it makes it hard to interpret them. Principal Component Analysis can identify the set of vectors that transforms the matrix of embedding factors to a smaller set of factors, while capturing much of the variance of the original matrix. This helps by aggregating correlated embeddings (and therefore similar embeddings).

#caption Representation of movies based on two strongest PCA components
#alt Representation of movies based on two strongest PCA components
g = ratings.groupby('title')['rating'].count()
top_movies = g.sort_values(ascending=False).index.values[:1000]
top_idxs = tensor([learn.dls.classes['title'].o2i[m] for m in top_movies])
movie_w = learn.model.movie_factors[top_idxs].cpu().detach()
movie_pca = movie_w.pca(3)
fac0,fac1,fac2 = movie_pca.t()
idxs = list(range(50))
X = fac0[idxs]
Y = fac2[idxs]
plt.figure(figsize=(12,12))
plt.scatter(X, Y)
for i, x, y in zip(top_movies[idxs], X, Y):
    plt.text(x,y,i, color=np.random.rand(3)*0.7, fontsize=11)
plt.show()

The model is roughly classifying some concept of classic and pop culture.

Using fastai.collab

Fastai’s collab_learner allows us to create and train a collaborative filtering model.

learn = collab_learner(dls, n_factors=50, y_range=(0, 5.5))
learn.fit_one_cycle(5, 5e-3, wd=0.1)

You can print the model to see the names of the layers.

learn.model

Recreating previous analysis

movie_bias = learn.model.i_bias.weight.squeeze()
idxs = movie_bias.argsort(descending=True)[:5]
[dls.classes['title'][i] for i in idxs]

Embedding Distance

We can calculate the distance between embeddings to measure similarity (or lack of it) between movies. The metric is pythagorean distance. In this sense movies are similar if the users that like the movies are similar. The most similar movie to Silence of the Lambs is:

movie_factors = learn.model.i_weight.weight
# idx = dls.classes['title'].o2i['Silence of the Lambs, The (1991)']
# idx = dls.classes['title'].o2i['Forrest Gump (1994)']
idx = dls.classes['title'].o2i['Terminator, The (1984)']
distances = nn.CosineSimilarity(dim=1)(movie_factors, movie_factors[idx][None])
idx = distances.argsort(descending=True)[1]
dls.classes['title'][idx]

Bootstrapping a Collaborative Filtering Model

Great, so you have a trained model. Given a user with a reasonable watch history, you can make recommendations based on similar users’ preferences.

Now, what do you do with a new user? They have no watch history to compare to other users. You could give a new user the mean of all the other users’ embedding vectors. But the mean of the other users’ vectors won’t necessarily make sense. If half the users like horror, and half the users like action, your new user will be recommended half-horror half-action movies.

Another option is to ask a set of questions, then assign an embedding vector based on the answers. This will give some control over the possible enbeddings.

The entire concept of collaborative filtering here is based on users’ ratings. If there are users that rate a lot, they can tilt the system towards their preferences. They can create feedback loops, where their tastes are overrepresented. It’s hard to quantify this overrepresentation though.

Deep Learning for Collaborative Filtering

For deep learning, we can concatenate the results of the embedding lookup. The resulting matrix can be passed through linear layers and nonlinear layers as normal.

Instead of getting the dot product we will concatenate the embeddings. Because of this the matrices can have different sizes. Fastais get_emb_sz returns recommended sizes – aka numbers of latent factors.

embs = get_emb_sz(dls)
embs

Let’s implement this class

class CollabNN(Module):
    def __init__(self, user_sz, item_sz, y_range=(0,5.5), n_act=100):
        self.user_factors = Embedding(*user_sz)
        self.item_factors = Embedding(*item_sz)
        self.layers = nn.Sequential(
            nn.Linear(user_sz[1]+item_sz[1], n_act),
            nn.ReLU(),
            nn.Linear(n_act, 1))
        self.y_range = y_range
        
    def forward(self, x):
        embs = self.user_factors(x[:,0]),self.item_factors(x[:,1])
        x = self.layers(torch.cat(embs, dim=1))
        return sigmoid_range(x, *self.y_range)

And use it to create a model:

model = CollabNN(*embs)

The Embedding layers are created in CollabNN. embs infers the sizes. self.layers is our mini-Neural Net and in forward we apply the embeddings, concatenate results, and pass through the mini neural net. Then sigmoid_range is applied.

learn = Learner(dls, model, loss_func=MSELossFlat())
learn.fit_one_cycle(5, 5e-3, wd=0.01)

Within fastai.collab, if you use collab_learner and pass use_nn=True it will apply the neural network as part of the collaborative filtering model. The code below makes two hidden layers, of size 100 and 50 respectively:

learn = collab_learner(dls, use_nn=True, y_range=(0, 5.5), layers=[100,50])
learn.fit_one_cycle(5, 5e-3, wd=0.1)

Conclusion

This application shows how we can learn the underlying factors within a dataset, and quantify aspects of the data. This can give us information that is otherwise unmeasured.

Classifying Pet Breeds and Fine-Tuning a Neural Net

This post follows fastai’s content on building, training and fine-tuning Neural Networks. It leverages one of the examples provided in their course involving classifying the breeds of dogs and cats.

The code for this example is stored on my GitHub.

Connecting to the data and pre-sizing the images

I connected to fastai’s PETS dataset, and used regex to extract the label names from the file names.
Full details on the Oxford-IIIT Pet dataset here.

Some of the code applied, using fastai’s DataBlock:

pets = DataBlock(blocks = (ImageBlock, CategoryBlock),
#                  telling the DataBlock that the data is image based
                 get_items=get_image_files, 
#                  set random seed
                 splitter=RandomSplitter(seed=42),
#                  use regex to extract y values from names
                 get_y=using_attr(RegexLabeller(r'(.+)_\d+.jpg$'), 'name'),
#                  resize and transfom as part of presizing
                 item_tfms=Resize(460),
                 batch_tfms=aug_transforms(size=224, min_scale=0.75))

# pointing the DataBlock to the images directory
dls = pets.dataloaders(path/"images")

One of the lines above (item_tfms=Resize(460)) pre-sizes the images.

Pre-sizing is image preprocessing that does several things: give images same dimensions, so they can be transformed to tensors for GPU-based operations. Keeping image sizes consistent reduces the amount of distinct augmentation computations that are required, and makes for more efficient processing

Many common data augmentation transforms introduce spurious empty patches or degrade data. For example, simple rotation will introduce empty pixels into the image. Other techniques may interpolate pixels.

The solution works in two steps:

  1. Resize to a fairly large size, so all images are consistent.
  2. Combine the individual augmentations into one and perform the combined operation on the GPU once, instead of performing operations individually.

The resize creates images large enough to leave margins for further augmentation.

The GPU is used for all data augmentation, and all operations are done together, with a single interpolation at the end.

The fastai course content shows an interesting comparison between their pre-sizing operations (left) and the same transformations applied with a more traditional approach (right).

Image from fastai’s course documentation.


Here are a few examples from the datablock, to make sure the data is loaded correctly.

A few examples of the loaded and prepared images

Fastai often recommends training to a simple model initially, rather than overengineering an elaborate model from the start. This helps establish a baseline and gives an idea of the data can train a model at all.

Below, applying the loaded data to a neural net, with error rate used as the metric, and resnet34 used for transfer learning.

learn = cnn_learner(dls, resnet34, metrics=error_rate)
learn.fine_tune(2)
A baseline run.

That loss is the function chosen to optimize the parameters of the model. Fastai tries to select an appropriate loss function based on what kind of data and model being used. In this case, with image data and a categorical outcome, fastai defaults to using cross-entropy loss.

Cross-entropy loss is a loss function works efficiently with multiple categories. I have already written a blog post on it, here: https://frankiecoughlan.data.blog/2021/12/10/cross-entropy-loss/

Model Interpretation

A confusion matrix will help see where the model is doing well, and where it’s doing badly:

interp = ClassificationInterpretation.from_learner(learn)
interp.plot_confusion_matrix(figsize=(12,12), dpi=60)
The confusion matrix with 37 categories.

That confusion matrix is a little awkward to read. The most_confused method will show the cells of the confusion matrix with the most incorrect predictions.

interp.most_confused(min_val=5)
The breeds that the model got most confused over.

Taking this as the baseline, the next section focuses on further improvements.

Improving the Model

The Learning Rate Finder

Finding the right learning rate is important:
Too low and it will take a long time to train. This wastes time, but can also cause overfitting.
Too high and it will overstep, overshooting the minimum loss. Repeating that can move the model away from the optimum solution, not closer.

This happens below with a learning rate of 10%

learn = cnn_learner(dls, resnet34, metrics=error_rate)
learn.fine_tune(1, base_lr=0.1)
Poor performance with a very high learning rate.

One approach to find the optimum learning rate (from Leslie Smith in 2015) is called the learning rate finder. Start with a tiny learning rate, use it for one mini-batch, and measure the loss. Double the rate, use it for one mini-batch, and measure the loss again. The loss will improve because we’ve taken a slightly bigger step in the right direction. Repeat until the loss gets worse. At that point, the learning rate is too large. It’s common to select a learning rate slightly lower than that which produced the worsening loss. Two recommendations for choosing that point are (a) one order of magnitude lower that that where the minimum loss was achieved, or (b) the last point where the loss was clearly decreasing. (a) and (b) are often very similar.

The default learning rate for fastai is 1e-3.

learn = cnn_learner(dls, resnet34, metrics=error_rate)
lr_steep = learn.lr_find()

In the graph above, learning rates lower than 1e-4 won’t allow the model to train – the slope is flat.
Rates higher than 1e-1 will cause the model to diverge – positive slope.
The point that produces the minimum loss is tempting, but the slope is flat here too.
The sweet spot is the point in the graph with the steepest negative slope.

learn = cnn_learner(dls, resnet34, metrics=error_rate)
learn.fine_tune(2, base_lr=3e-3)
Better performance with a single chosen learning rate.

Unfreezing and Transfer Learning

Transfer learning in a nutshell:

We take a pretrained model that is finetuned for a particular task. It has linear layers, with a nonlinear activation function between each linear pair, and an activation function like softmax at the end. The final linear layer uses a matrix with enough columns that the output size = number of classes in the classification problem.

This final linear won’t help when we are transfer learning, because it is specifically designed to classify the categories in the original context. So when transfer learning that layer is discarded and replaced with a layer that provides the correct number of outputs for the new problem. This new linear layer is random, but the output of the entire model isn’t random because this final random layer sits on top of all the prior carefully trained layers.

The aim is to retain all the useful things learned in early layers and then apply them to our new problem. The trick is to freeze the weights for the prior layers, and tell the optimizer to only update the weights in the later randomly added layers.

Fastai automatically freezes pretrained layers when creating a model from a pretrained network. The fine_tune method makes fastai do two things:
Train the randomly added layers for one epoch, with all other layers frozen
Unfreeze all the layers, and train them for all the epochs required

The method has parameters that can change its behaviour, or we can call the underlying methods directly to get custom behaviour. fit_one_cycle is fastai’s method for training without fine_tune.

learn = cnn_learner(dls, resnet34, metrics=error_rate)
learn.fit_one_cycle(3, 3e-3)

Then we’ll unfreeze the model, and run lr_find again, because having more layers to train, and weights that have already been trained for three epochs, means our previously found learning rate isn’t appropriate any more.
The graph looks different this time – the flat line followed by the sharp increase is due to the fact that the model has been trained already. Looking for the point with maximum gradient won’t help here, and instead it makes sense to choose a point well before the sharp increase, like 1e-5.

The learning rate finder on a transfer-learning model.

Training at that learning rate improved the model a bit, but there’s more that can be done.

Discriminative Learning Rates

After unfreezing, all of the weights can be changed by further training. However the quality or usefulness of the pretrained weights is higher than that of the randomly added parameters. The pretrained weights have been trained over a high number of epochs with a lot of data. The implication is that those pretrained weights would be better off with a lower learning rate, since less drastic change is required.

Also, earlier layers typically have learned generally useful things – like detecting edges and gradients – which is useful for any task. Later layers are far more specific, and so may be less relevant to the new task. It seems reasonable to let the later layers fine tune more quickly than the earlier ones.

Fastai’s default approach is to use discriminative learning rates. It uses lower learning rates in the early layers and higher learning rate for the later layers.

In Fastai you can pass a first and last learning rate to the model (via a python slice object) and it will start training with the first rate, and finish with the last rate, making incremental steps in between.

See below for previous training reapplied with discriminative learning rates.

learn = cnn_learner(dls, resnet34, metrics=error_rate)
learn.fit_one_cycle(3, 3e-3)
learn.unfreeze()
learn.fit_one_cycle(12, lr_max=slice(1e-6,1e-4))
Comparing training loss to validation loss.

The training loss inmproves throughout, but the validation loss slows and plateaus. This is indicating the model is starting to overfit, and becoming overconfident in its predictions.

That doesn’t mean it’s getting less accurate. Accuracy continues to improve even as validation loss gets worse. The end goal is optimising the chosen metrics, not necessarily the loss.

Number of Epochs

If time is a limiting factor, then an initial approach could be to train for the number of epochs that you’re willing to wait for. Then looking at the training/validation loss plots and the metrics will help. If they are still improving in the final epochs then it might be worth training for longer.

If the validation loss get worse towards the end of training, then the model is first getting overconfident, then incorrectly memorizing the data. The latter is the real issue. If the validation loss is deteriorating while the other metrics are improving, then the model can still be performing better.

It’s tempting to run the model for 100 epochs, choose the epoch with the best metrics, and pick that model as the winner. However, this approach doesn’t allow the model to use the smallest rates on the epochs with the best metrics, so the model can’t really fine-tune around the correct area. A better approach is to retrain the model from scratch with the number of epochs adjusted based on where the results were previously good.

Deeper Architectures

In general (but this is a big generalisation), more parameters will lead to more accurate results. Deeper models also use more memory, so deeper models should use smaller batch sizxes to avoid ‘out of memory’ issues on the GPU. Training on deeper architectures can be sped up by using ‘mixed-precision training’, where floating points are reduced from 32 to 16 places where possible.

Trying a deeper architecture with mixed precision:

from fastai.callback.fp16 import *
learn = cnn_learner(dls, resnet50, metrics=error_rate).to_fp16()
learn.fine_tune(6, freeze_epochs=3)
Results from a deeper architecture.

In this case the deeper model doesn’t provide a massive improvement, but the runtime for each epoch has gone from 50 seconds to 5 minutes.

Implementing k Nearest Neighbours from scratch in Python

This post is a follow up from the previous post. I will work through an implementation from scratch in python.

The code is saved at my github.

The data

The data is a well known dataset related to features of Iris flowers, from the Iris Plants Database. It might be the best known dataset in classification literature, related to several classic papers. The data is here, and more info on the dataset is here.

The code

First, importing the dataset and having a look at the data.

# import required packages
from csv import reader
from math import sqrt

#create a simple function for loading csvs

def load_csv(filename):
    dataset = list()
    with open(filename, 'r') as file:
        csv_read = reader(file)
        for row in csv_read:
            if not row:
                continue
            dataset.append(row)
    return dataset

#identify the file name and load the data

filename = 'iris.csv'
dataset = load_csv(filename)

#view the data

#note it is all in string format and needs to be reformatted

dataset
First view of the data.

Then cleaning the data, converting strings to the correct formats.

# function to convert string column to float

#strip removed trailing and leading blanks
#float function converts to float

def str_to_float(dataset, col):
    for row in dataset:
        row[col] = float(row[col].strip())

#apply the formula to all but the last column

for i in range(len(dataset[0])-1):
    str_to_float(dataset, i)

#now need to clean up the categorical variable, convert it to an integer.

def str_to_int(dataset, col):
    #identify classes
    classes = [row[col] for row in dataset]
    #get distinct classes
    distinct = set(classes)
    #initialise dictionary of class values
    lookup = dict()
    #fill dict with integer for each class value
    for i, value in enumerate(distinct):
        lookup[value] = i
        print('[%s] = %d' % (value, i))
    #lookup into the dictionary, based on the string
    for row in dataset:
        row[col] = lookup[row[col]]
    return lookup 

 #apply formula to last column
    
str_to_int(dataset, len(dataset[0])-1)

Cleaner!

The algorithm

Functions to calculate euclidean distance, identify neighbours, and apply the algorithm to new data.


# Calculate the Euclidean distance
def euclid_dist(r1, r2):
    #initialise distance
    distance = 0.0
    #loop through rows
    for i in range(len(r1)-1):
        distance += (r1[i] - r2[i])**2
    return sqrt(distance)

# Locate the most similar neighbors
def get_neighbours(train, test_row, num_neighbours):
    #train is the training dataset
    #test_row is a single observation
    #num_neighbours is the predefined k value
    
    #initialise list
    distances = list()
    #loop through train rows
    for train_row in train:
        #measure distance betweeen test row and train row
        dist = euclid_dist(test_row, train_row)
        #add train row and distance to distances dataset
        distances.append((train_row, dist))
    #sort by distance
    distances.sort(key=lambda tup: tup[1])
    neighbours = list()
    #identify top k neighbours
    for i in range(num_neighbours):
        neighbours.append(distances[i][0])
    return neighbours
 

# Make a prediction

def predict_class(train, test_row, num_neighbours):
    #identify neighbours using neighbour function
    neighbours = get_neighbours(train, test_row, num_neighbours)
    #keep the class variable
    output = [row[-1] for row in neighbours]
    #assign the most common category as the prediction
    prediction = max(set(output), key=output.count)
    return prediction
 

and then applying it to a single new input!

# define k
num_neighbours = 5
# define a new record
row = [4.3,1.1,2.2,1.1]
# predict the class
label = predict_class(dataset, row, num_neighbours)
print('Data=%s, Predicted: %s' % (row, label))
It runs!

Testing the accuracy of the algorithm

import random

random.shuffle(dataset)

train_data = dataset[:120]
test_data = dataset[120:]

num_neighbours = 5

correct = 0
for i in range(len(test_data)):
    y = predict_class(train_data, test_data[i][0:4], num_neighbours)
    if y == test_data[i][4]:
        correct += 1
    print(y, test_data[i][4])

29 out of 30 correct classifications in the test dataset. Not bad!

Algorithm Overview: k-Nearest Neighbours

k-Nearest Neighbours (k-NN) is a non-parametric classification algorithm that has been around since the 50’s. It is typically used for classification, but can also be used for evaluation. With classification, a data point is compared to the k observations closest to it, and the classes of those observations infer the class of the data point. With regression, the same approach is taken, but the mean of the regression variable is used as the inferred value for the data point. Classification is the more common use.

The measure of distance used to identify neighbours is important, as is the scales of the variables involved – if they vary significantly then standardisation can help with accuracy. Finally, weights are sometimes applied to the neighbours, with nearer data points carrying more significance.

How The Algorithm Works

The training dataset consists of multidimensional vectors, each with a class label.

The classification dataset is an unlabeled vector. With a user-defined k, the unlabeled vector is compared to its k nearest neighbours. For classification, the unlabeled vector gets the mode label from the k nearest neighbours. For regression, it gets the average of the k nearest labels.

Distance Metrics

The most commonly used distance metric for continuous variables is Euclidean Distance. In n-dimensional space the formula is:

Formula from linked Wikipedia page.

For discrete variables, including text classification, metrics like the Hamming Distance can work.

Choosing K

Choosing k depends on the data. Larger values of k can reduce noise, but the boundaries between classes can become complicated. Choice of features can also introduce significant noise: having irrelevant or unimportant features in the data degrades the algorithm’s performance. Feature selection and feature scaling can help with optimisation.

Running the algorithm at different values of k will allow for error rate inspection and optimal parameter choice.

Examples

Here’s a simple example that does calculations by hand.

Here’s a blog post with a from-scratch implementation in python.

Here’s a great blog post from analyticsvidhya.com with python and R implementation, and some nice graphical examples.

Graphs from analyticsvidhya’s post.

Cross Entropy Loss

In a previous blog post I described the process of training a neural network for classifying the quality of Guinness. That was a binary classification problem (the result was ‘good’ or ‘bad’) and a key step in it was using the sigmoid function to ensure that the model’s activations were between 0 and 1. This blog post looks at the methods used for classification with neural net loss functions in binary and multi-class scenarios.

To make the scenarios easier to work through, I will use simple numerical examples where possible.

The Binary Case

Consider a neural net that has to classify into two categories – identifying an image as a cat or a dog. The way the classification works in practice is this – with 6 images that are run through our model, and we get a single activation score out for each image:

torch.random.manual_seed(42);

acts = torch.randn((6,1))*2
acts
Our activation scores.

These activation scores aren’t automatically interpretable, and the values don’t have a relatable meaning. To help with this, they are fed through the sigmoid function. The sigmoid function translates any value to a value between 0 and 1. This means our activations have an interpretable, common scale, and they can be considered synonymous with probabilities.

After our activation scores are fed through the sigmoid function we get:

acts.sigmoid()
Output of the sigmoid function are always between 0 and 1.

We can consider these activations as probabilities that the input image is classified as one of our binary categories. For example, if the model is categorising cats and dogs, then our model is 98.81% sure that the first image is a cat, and 21.82% sure that the second image is a cat (which means it is 78.18% sure that it is a dog).

The actual measure of loss depends on whether the model is correct or not. In the example above, if the first image is a cat, then the loss is (1-0.9881). If the first image is a dog, then the loss is 0.9881. The benefit of this approach is that the loss is dependent on the model’s confidence in its classifications, not the amount of classifications it gets correct or incorrect. This means a small change of parameters (e.g. from Gradient Descent) will always cause a change in loss, even if the classification decision has not changed.

Another approach to the Binary case

Here’s another way of looking at the binary case: what if we create two activations, one for the ‘cat’ and one for the ‘dog’?

acts = torch.randn((6,2))*2
acts
Two activations for the binary case.

So here we have two activations for each image – one for the ‘cat’ and one for the ‘dog’. One of the complexities this introduces is that these activations are independent from eachother. In the first row above, 2.2206 and -3.3796 are not directly dependent on eachother. This means that when the activations are fed through the sigma function the probabilities don’t really make sense. The probabilities in the rows below don’t sum to 1 as we would expect them to.

Outputs of the sigmoid function don’t really make sense as probabilities here.

What’s happening is that we are not really getting a true probability. We getting the model’s confidence in relation to each category. Whether the numbers are high or low doesn’t matter, what matters is which is higher or lower in comparison to the other.

To convert this into something that we can solve with the sigmoid function we do the following: get the difference between the activations and put that through the sigmoid function. The difference between the activations represents how much more sure the model is about category A vs B, or cat vs dog.

(acts[:,0]-acts[:,1]).sigmoid()
The sigmoid output of the difference between the activation scores.

This ‘trick’ to utilise the sigmoid with two activation values already has a name – the softmax function.

Softmax

From wikipedia:

The softmax function takes as input a vector z of K real numbers, and normalizes it into a probability distribution consisting of K probabilities proportional to the exponentials of the input numbers. That is, prior to applying softmax, some vector components could be negative, or greater than one; and might not sum to 1; but after applying softmax, each component will be in the interval {\displaystyle [0,1]}[0,1], and the components will add up to 1, so that they can be interpreted as probabilities. Furthermore, the larger input components will correspond to larger probabilities.

https://en.wikipedia.org/wiki/Softmax_function

And in function form:

The softmax function.

In short, the function takes a list of real numbers, and transforms them into a set of probabilities. It can take negative numbers as input. The output will always sum to 1. This works because the exponential function always outputs a positive number. Using the exponential function also means that inputs that are slightly larger will be assigned much larger probabilities. For example exp(4) = 55 and exp(8) = 2981. This means the softmax is likely to assign a single category as the definitive ‘winner’ – which is helpful for training.

Exponential function from 0 to 4.

Applying this to our example gives the same result that the alternative application of the sigmoid function did!

sm_acts = torch.softmax(acts, dim=1)
sm_acts
Softmax output.

The second part of cross entropy loss is the log likelihood.

Log Likelihood

We could calculate Loss directly from the softmax output above. To do this we would compare our labels to the probabilities assigned, and pull the values from the softmax output based on the labels. For example, if the first column represents the probability that the image is a cat, and we know that the first image is a cat, then we’d take the 0.6025. If the label for the second row indicates that it’s a dog, then we’d take 0.8668, and so on. This generalises nicely to a situation with three, ten or a hundred categories.

One shortfall of this is that the loss is then being expressed using probabilities. This means that the model would think of 0.99 and 0.999 as very similar, when in reality 0.99 is incorrect 1/100 times, and 0.999 is incorrect 1/1000 times, and that might be a massive difference in accuracy for a particular problem.

The solution to this is to run these values through the log function. The log function will convert the probabilities (range 0:1) to a log scale with range (-infinity: 0). Then it takes the negative value of that. So a correct classification with a high probability gets a value close to 0, and an incorrect classification with a high probability gets a very large value. Since the aim is to minimise loss, this penalises the model for incorrect classifications,

Natural log from 0 to 4. All probabilities will fall in the range of 0-1, and will get values assigned from -inf to 0.

All these transformations are neatly named the Negative Log Likehood. In practice then we get the softmax, calculate its log, and calculate the negative log likelihood based on that. This is Cross Entropy Loss! PyTorch does that in the following function:

loss_func = nn.CrossEntropyLoss()

loss_func(acts, targ)

This give our example an output of 1.8045. We can also split that loss value out into the individual values for each row:

nn.CrossEntropyLoss(reduction='none')(acts, targ)

1.8045 is just the mean of these values.

All-in-all this functions as an effective loss function because:
1. It works for multi-category problems
2. Softmax produces probabalistic activations
3. Softmax wants to pick a winning categorisation
4. The log functions allows for differentiation of small probability differences

Final benefit of Logs

One final note on the benefit of logs:

Log(a x b) = Log(a) + Log(b)

This simple expression is very useful when dealing with very small and very large numbers. Being able to replace multiplication with addition reduces risk of computational inaccuracies from floating point errors. Computers are then a lot less likely to end up with scales of numbers that they can’t handle.

Redundant Encoding in Data Visualisations

Redundant Encoding is the practice of adding multiple visual elements to a visualisation, to enhance effectiveness and ease-of-understanding. It is also referred to as redundancy. According to displayr.com, it can ‘improve the chances of a reader interpreting a visualization quickly and correctly’.

Redundant encoding applies to many elements of a visualisation, including colours, shapes, labels, and sizes.

Here is a simple barplot with some countries and their populations. All the information needed to interpret the graph is included, but it is not necessarily easy to do so.

# create dataset
height = [45, 7, 51, 6, 1]
bars = ('Argentina', 'Bulgaria', 'Colombia', 'Denmark', 'Estonia')
y_pos = np.arange(len(bars))
 
# Create horizontal bars
plt.barh(y_pos, height)
 
# Create names on the x-axis
plt.yticks(y_pos, bars)

plt.title('5 Countries and their Populations in Millions')

# Show graphic
plt.show()

Firstly, we’ll add some axis labels and reorder the data based on the value of the bar, not its place in the dataset.

# Create a data frame
df = pd.DataFrame ({
        'Group':  ['Argentina', 'Bulgaria', 'Colombia', 'Denmark', 'Estonia'],
        'Value': [45, 7, 51, 6, 1]
})

# Sort the table
df = df.sort_values(by=['Value'])

y_pos = np.arange(len(bars))

# Create horizontal bars
plt.barh(y=df.Group, width=df.Value);
 
# Create names on the x-axis
plt.yticks(y_pos, bars)

plt.title('5 Countries and their Populations in Millions')
plt.xlabel('Population (Millions)')
plt.ylabel('Countries')

# Show graphic
plt.show()

This simple change makes the graph easier to take in.

Another option is using colour to deliver a message. Here is an example of the same graph with colour being used to emphasize bar length. This might be an example of ‘less is more’. I am not sure in this case if the greyscale is adding to the graph or taking away from it.

# Create a data frame
df = pd.DataFrame ({
        'Group':  ['Argentina', 'Bulgaria', 'Colombia', 'Denmark', 'Estonia'],
        'Value': [45, 7, 51, 6, 1]
})

height = [45, 7, 51, 6, 1]

totalheight = sum(height)
# color = height/totalheight
color2 = [number / totalheight for number in height]
color2.sort(reverse=True)
color = [str(a) for a in color2]

# Sort the table
df = df.sort_values(by=['Value'])

y_pos = np.arange(len(bars))

# Create horizontal bars
plt.barh(y=df.Group, width=df.Value, color = color);
 
# Create names on the x-axis
plt.yticks(y_pos, bars)

plt.title('5 Countries and their Populations in Millions')
plt.xlabel('Population (Millions)')
plt.ylabel('Countries')


# Show graphic
plt.show()

Examples

Here are some examples from other sources of great examples of redundant encoding. From displayr.com, showing use of colours and labels.

Before
And After!

From cnothelfer.com, showing redundant encoding in a scatterplot:

Great example.

clauswilke.com is a wealth of great visualisations. Here’s a lovely example of redundanct encoding – a few seconds of examination provides effortless transfer of information.

Building a Neural Network from scratch using PyTorch and FastAI

This post follows content from fastai to build a two layer Neural Network from scratch using python. The code is saved on my github.

The data, and data prep

I used a sample dataset from FastAI based on a famous computer vision dataset, MNIST. The sample dataset uses 3s and 7s only, to make the overall problem simpler.

After importing the required packages and connecting to the data, we can see some of the images in the dataset. There are over 6000 ‘3s’ in the dataset. Here are a few of them:

They can also be viewed as arrays, where each pixel has a value between 0 (white) and 255 (black). This is the top portion of one.

It is easier to see if the array is shaded.

Using python we can stack the images into a single tensor, and then see what the ‘average’ 3 or 7 looks like:

# creating lists of the images in the 7 and 3 folders
seven_tensors = [tensor(Image.open(o)) for o in sevens]
three_tensors = [tensor(Image.open(o)) for o in threes]
len(three_tensors),len(seven_tensors)
# using torch.stack to stack the tensors into a 3 dimensional tensor
stacked_sevens = torch.stack(seven_tensors).float()/255
stacked_threes = torch.stack(three_tensors).float()/255
stacked_threes.shape
# the average 3
mean3 = stacked_threes.mean(0)
show_image(mean3);
# the average 7
mean7 = stacked_sevens.mean(0)
show_image(mean7);
The average 3
The average 7.

Structuring the data for the Neural Net

Using PyTorch we get the data into a format and shape that works for our calculations. Our images are transformed into a tensor with 12396 rows and 784 columns. The 784 columns represent the 28 x 28 pixels, flattened. Our labels are a vector with 12396 rows and 1 column.

train_x = torch.cat([stacked_threes, stacked_sevens]).view(-1, 28*28)
train_y = tensor([1]*len(threes) + [0]*len(sevens)).unsqueeze(1)
train_x.shape,train_y.shape
The data in the correct shape and format.

The training data is put into a dataset where the x,y variables are available as a tuple. x is the 784 pixels of an image and y is the label. The same prep is applied to the validation data.

dset = list(zip(train_x,train_y))
x,y = dset[0]
x.shape,y
The prepared dataset, in tuple form.

With a set of randomly generated weights and biases, a simple linear model of the form y = ax + b is created, and when the training data is applied to it, a score for every image is output:

def linear1(xb): return xb@weights + bias
preds = linear1(train_x)
preds
The outputted scores.

Turns out this has a 57% accuracy, so not much better than guessing. This is to be expected since the weights were randomly generated.

The next step should be to update the weights in a way that improves the accuracy slightly. One issue with this is that a small change in the weights might not change the accuracy at all, causing the algorithm to get stuck. Introducing the sigmoid function; It is a smooth continuous function that will always give a different output for small changes in weights. This allows us to make and measure the effect of small changes in the parameters, so that the model moves in the right direction.
Any input to the function always gives an output between 0 and 1.

Putting the model together

All in all then, we do the following:
1. Initialise a set of random weights and biases
2. Apply our training data to the model
3. Calculate the loss of the predictions compared to the correct answers
4. Calculated the gradients of the parameters and adjust the parameters
5. Go back to step 2

# This function:
#     Takes the input data
#     Makes predictions based on the model ax + b
#     Calculates the loss using the sigmoid functionality when comparing to the dependent values
#     Calculated the gradients using the .backward() method
def calc_grad(xb, yb, model):
    preds = model(xb)
    loss = mnist_loss(preds, yb)
    loss.backward()
def train_epoch(model, lr, params):
    for xb,yb in dl:
        calc_grad(xb, yb, model)
        for p in params:
            p.data -= p.grad*lr
            p.grad.zero_()
def batch_accuracy(xb, yb):
    preds = xb.sigmoid()
    correct = (preds>0.5) == yb
    return correct.float().mean()
def validate_epoch(model):
    accs = [batch_accuracy(model(xb), yb) for xb,yb in valid_dl]
    return round(torch.stack(accs).mean().item(), 4)

Running the code for 20 epochs shows increasing accuracy

for i in range(20):
    train_epoch(linear1, lr, params)
    print(validate_epoch(linear1), end=' ')
Accuracy after each epoch.

Adding Nonlinearity

Adding a nonlinear layer gives us a true neural net. Here is a tw0-layer neural net, with a RELU layer in between.

def simple_net(xb): 
    res = xb@w1 + b1
    res = res.max(tensor(0.0))
    res = res@w2 + b2
    return res

Here’s the nonlinear aspect introduced by the RELU function, visualised:

RELU

Much of this can be replaced with ready-made components from either PyTorch or fastai, making it easier to implement,

simple_net = nn.Sequential(
    nn.Linear(28*28,30),
    nn.ReLU(),
    nn.Linear(30,1)
)
learn = Learner(dls, simple_net, opt_func=SGD,
                loss_func=mnist_loss, metrics=batch_accuracy)
learn.fit(40, 0.1)
Output from the fastai learner.
Accuracy, visualised.

The accuracy after 40 epochs is 98.2%. A quick comparison with more mature methods shows that there is still lots more to do: a single epoch of fastai’s resnet18 model gives accuracy of 99.7%!

dls = ImageDataLoaders.from_folder(path)
learn = cnn_learner(dls, resnet18, pretrained=False,
                    loss_func=F.cross_entropy, metrics=accuracy)
learn.fit_one_cycle(1, 0.1)
Output from one epoch of Resnet18

A Simple Demonstration of Gradient Descent in Python

This post looks at the concept of gradient descent, and applies it to a simple function. It is prompted by and takes some code from the fastai lesson content.
Put simply, gradient descent is used to give feedback to a deep learning model so it can adjust its parameters. The intention is to prompt the model to adjust the parameters so that the model will improve slightly, and make repeated improvements so that the model gets iteratively better.

From the wikipedia link above.

Mathematically, it is an algorithm for finding the local minimum of a differentiable function. The differentiable function in this case is the loss function of the model, and the local minimum represents the (or one of the) best set(s) of parameters for the model.

The full code for this is saved here. I’ll include highlights and graphs in this post.

First, we generate a simple dataset to train our model to.

# Define a tensor of time values
time = torch.arange(0,20).float(); time
# Define a tensor of speed values, with a quadratic element built in
speed = torch.randn(20)*2 + 0.75*(time-9.5)**2 + 1
plt.scatter(time,speed);
A plot of the ‘dataset’.

Then, define a function as the basis for our model (in this case it is a quadratic function), and a loss function (mean square error).

# We base our model on a quadratic function, based on the shape of the plot
# Need to solve for a, b and c

def f(t, params):
    a,b,c = params
    return a*(t**2) + (b*t) + c

# Also define a function to measure loss. The mean squared error works.

def mse(preds, targets): return ((preds-targets)**2).mean().sqrt()

Initialising the parameters, making the predictions, and plotting them:

params = torch.randn(3).requires_grad_()
orig_params = params.clone()

preds = f(time, params)
First fit with the actual data in blue and the fitted model in red.

We want to calculate the gradients of the parameters, and then adjust the parameters by some small fraction of those gradients. Taking small steps allows for incremental improvements in the parameters and avoids divergence.

In this case we will iterate through the process 25 times. 25 times is an arbitrary choice and in practice observation of the loss would help indicate when to stop.

# Create function to run through this process on repeat

def apply_step(params, prn=True):
    preds = f(time, params)
    loss = mse(preds, speed)
    loss.backward()
    params.data -= lr * params.grad.data
    params.grad = None
    if prn: print(loss.item())
    return preds

# Repeat it 25 times, outputting the loss

for i in range(25): apply_step(params)
The output of our arbitrary 25 iterations

A plot of the first few graphs shows small changes to the model each time, as it moves slowly towards a better fit.

And here is the plot after 100 iterations.

More resources:
Builtin.com
3Blue1Brown on Youtube

Deploying a Python Neural Net using Binder

The previous blog post trained a neural network to classify images of Guinness. This post will demonstrate how to host the model on Binder so that people can use it themselves.

What is Binder?

The Binder Project helps you create one-click, sharable, live code environments from public code repositories that runs entirely in the cloud. You put your files in a github repository, and point Binder to it. Binder builds a docker image of the repository and hosts it on a JupyterHub server, allowing you to or others to interact with your code in a browser.

Create a github repository and add the required files to it. For this example, I will be adding:
a. A python script that uses widgets to create upload and classify buttons
b. The model trained in the previous post, exported to a .pkl file
c. A requirements text file that specifies the packages that Binder needs

Here is some of the script that will be uploaded. It defines the buttons that will be displayed and the output shown after classification.

# function for image classification
def on_click_classify(change):
    #sets img as an image object based on button upload
    img = PILImage.create(btn_upload.data[-1])
    #clears existing output
    out_pl.clear_output()
    #sets the output as 128x128 and creates a prediction based on learn_inf
    with out_pl: display(img.to_thumb(128,128))
    pred,pred_idx,probs = learn_inf.predict(img)
    if {pred} == 'good':
        lbl_pred.value = f'The model is {probs[pred_idx]:.0%} sure that this is a LOVELY pint. Enjoy!'
    else:
        lbl_pred.value = f'The model is {probs[pred_idx]:.0%} sure that this is a shocking bad pint. Unlucky.'

See the rest of the code at the github repository.

Step 2

Tell Binder where to look! Enter the repository name and a specific file to look for. In this case my file path includes ‘voila/render’ indicating that I want the output to show the widgets and text only – not the code. It gives me a link that I can use to share with others.

This image has an empty alt attribute; its file name is image.png
The Binder Interface.

When I click launch a window opens that establishes the docker image and install the packages in my requirements.txt file. It can take a few minutes to finish.

Step 3

Try it out!

And here is the link to the binder page!

Guinness Image Classifier: Training the model in Python using fastai

I have been learning about Neural Networks and Image Classifiers through fastai recently, and wanted to try apply what I had learned to a problem outside the scope of that material.

The Problem

Any Irish person that drinks Guinness will probably have strong opinions on the quality of the drink in just about any pub they’ve been in (and non-Guinness drinkers might claim that it is all nonsense). The opinions range from ‘beautiful’ to slightly more offensive at the other end of the scale. Recently social media has filled with examples of lovingly-crafted and not-so-lovingly-crafted beverages.

I wanted to train a Neural Network to be able to distinguish between a good and a bad pint, and have documented the process of training the model below.

The Data

I collected approximately 300 images of pints of Guinness. Sourcing these images from social media helped because they were pre-identified as good or bad pints, making the labelling process easier. I didn’t do any cleaning on the images before trying to train the model – instead after training I used fastai’s built-in cleaning functionality to easily identify images that were not appropriate for the model.

The process to scrape and organise the images will become a post of its own at some point.

Examining the Data

Checking a single image, to make sure file path works and that the filetype is actually an image.

Image 1 in the ‘good pints’ dataset.

Pulling a set of images in, with their labels. Here we see that a few of the images are not pictures of pints, and these will need to be cleaned. fastai makes identifying them and deleting them fairly convenient, and will happen after the first round of model training.

Some good pints, some bad pints, and some non-pints.

Model Training Attempt 1

Without further ado, we get into it. Below is the first output of the model fitting. The neural network applies transfer learning based on fastai’s resnet18, over 8 epochs. Loss rates on the training and validation datasets decrease fairly steadily, while the classification error rate levels off after 4 or 5 epochs.

And the confusion matrix gives us a visual depiction of how the validation date is classified. It also tells us whether bad pints are being misclassified as good or vice-versa. The first thing I learned here is that there is something off with the data structures that I am feeding into the model. As well as good and bad, the matrix has a category of ‘6.jpg’ which obviously shouldn’t be there. ‘number.jpg’ is the format I would expect from the images I am using, so something has gone wrong there. Apart from that things don’t look too terrible. This iteration of the model classified 9 good pints as bad, and 4 bad pints as good.

Confusion Matrix for the first model training attempt.

Data Clean 1

Fastai’s functionality for cleaning images displays the images that the model is least sure about, and then allows the user to make quick decisions on deleting or recategorising them. You can do this across the training and validation set, and the categories within the data (in this case ‘good’ and ‘bad’).

Here’s how that looks. Clearly there are some odd images in here.

Before making cleaning decisions.
After making cleaning decisions. Delete delete delete.

I applied the same process to the ‘bad’ validation set, and both datasets for the ‘good’ group.

Model Training Attempt 2 + 3

Immediately the results are better.

Attempt 2
Attempt 2

Oops. Still hadn’t fixed that ‘6.jpg’. A quick check of the data showed that there was a folder in with the source images called ‘6.jpg’. This folder was being picked up as a category (with a single image in it), so I corrected that.

Another couple of cleans of the data gives incrementally better results:

And a classification matrix that actually makes sense.

But there is still some image cleaning to be done:

The final (for now) model

After a final clean, the last run of the model gives us slightly better results again, with an error rate of just over 10%.

The final results from the model.

And the classification matrix shows 58 correct classifications, with 7 incorrect classifications.

The final classification matrix.

Here are some of the image that the model is classifying incorrectly. The tags above the images show prediction/actual, so good/bad means the model thought it was good and that it was actually bad. The final figure identifies the probability the model assigned to it – aka how confident the model was.
The images that confused the mode are interesting. It doesn’t seem to do well with images of multiple pints. Also, the first image shows pints that look decent in a non-Guinness glass. Feeding the model a lot more data might allow it to learn to differentiate between glass types or brand logos.

These confused the model. I’m a bit confused myself.

What next?

I exported the model and will publish it using Binder. I’ll also try set up an embedded model in this blog.

Finally, I’ll look at what improvements I can make to the model to improve its accuracy.

The code is on github.