In a previous blog post I described the process of training a neural network for classifying the quality of Guinness. That was a binary classification problem (the result was ‘good’ or ‘bad’) and a key step in it was using the sigmoid function to ensure that the model’s activations were between 0 and 1. This blog post looks at the methods used for classification with neural net loss functions in binary and multi-class scenarios.
To make the scenarios easier to work through, I will use simple numerical examples where possible.
The Binary Case
Consider a neural net that has to classify into two categories – identifying an image as a cat or a dog. The way the classification works in practice is this – with 6 images that are run through our model, and we get a single activation score out for each image:
torch.random.manual_seed(42);
acts = torch.randn((6,1))*2
acts

These activation scores aren’t automatically interpretable, and the values don’t have a relatable meaning. To help with this, they are fed through the sigmoid function. The sigmoid function translates any value to a value between 0 and 1. This means our activations have an interpretable, common scale, and they can be considered synonymous with probabilities.

After our activation scores are fed through the sigmoid function we get:
acts.sigmoid()

We can consider these activations as probabilities that the input image is classified as one of our binary categories. For example, if the model is categorising cats and dogs, then our model is 98.81% sure that the first image is a cat, and 21.82% sure that the second image is a cat (which means it is 78.18% sure that it is a dog).
The actual measure of loss depends on whether the model is correct or not. In the example above, if the first image is a cat, then the loss is (1-0.9881). If the first image is a dog, then the loss is 0.9881. The benefit of this approach is that the loss is dependent on the model’s confidence in its classifications, not the amount of classifications it gets correct or incorrect. This means a small change of parameters (e.g. from Gradient Descent) will always cause a change in loss, even if the classification decision has not changed.
Another approach to the Binary case
Here’s another way of looking at the binary case: what if we create two activations, one for the ‘cat’ and one for the ‘dog’?
acts = torch.randn((6,2))*2
acts

So here we have two activations for each image – one for the ‘cat’ and one for the ‘dog’. One of the complexities this introduces is that these activations are independent from eachother. In the first row above, 2.2206 and -3.3796 are not directly dependent on eachother. This means that when the activations are fed through the sigma function the probabilities don’t really make sense. The probabilities in the rows below don’t sum to 1 as we would expect them to.

What’s happening is that we are not really getting a true probability. We getting the model’s confidence in relation to each category. Whether the numbers are high or low doesn’t matter, what matters is which is higher or lower in comparison to the other.
To convert this into something that we can solve with the sigmoid function we do the following: get the difference between the activations and put that through the sigmoid function. The difference between the activations represents how much more sure the model is about category A vs B, or cat vs dog.
(acts[:,0]-acts[:,1]).sigmoid()

This ‘trick’ to utilise the sigmoid with two activation values already has a name – the softmax function.
Softmax
From wikipedia:
The softmax function takes as input a vector z of K real numbers, and normalizes it into a probability distribution consisting of K probabilities proportional to the exponentials of the input numbers. That is, prior to applying softmax, some vector components could be negative, or greater than one; and might not sum to 1; but after applying softmax, each component will be in the interval {\displaystyle [0,1]}
https://en.wikipedia.org/wiki/Softmax_function, and the components will add up to 1, so that they can be interpreted as probabilities. Furthermore, the larger input components will correspond to larger probabilities.
And in function form:

In short, the function takes a list of real numbers, and transforms them into a set of probabilities. It can take negative numbers as input. The output will always sum to 1. This works because the exponential function always outputs a positive number. Using the exponential function also means that inputs that are slightly larger will be assigned much larger probabilities. For example exp(4) = 55 and exp(8) = 2981. This means the softmax is likely to assign a single category as the definitive ‘winner’ – which is helpful for training.

Applying this to our example gives the same result that the alternative application of the sigmoid function did!
sm_acts = torch.softmax(acts, dim=1)
sm_acts

The second part of cross entropy loss is the log likelihood.
Log Likelihood
We could calculate Loss directly from the softmax output above. To do this we would compare our labels to the probabilities assigned, and pull the values from the softmax output based on the labels. For example, if the first column represents the probability that the image is a cat, and we know that the first image is a cat, then we’d take the 0.6025. If the label for the second row indicates that it’s a dog, then we’d take 0.8668, and so on. This generalises nicely to a situation with three, ten or a hundred categories.
One shortfall of this is that the loss is then being expressed using probabilities. This means that the model would think of 0.99 and 0.999 as very similar, when in reality 0.99 is incorrect 1/100 times, and 0.999 is incorrect 1/1000 times, and that might be a massive difference in accuracy for a particular problem.
The solution to this is to run these values through the log function. The log function will convert the probabilities (range 0:1) to a log scale with range (-infinity: 0). Then it takes the negative value of that. So a correct classification with a high probability gets a value close to 0, and an incorrect classification with a high probability gets a very large value. Since the aim is to minimise loss, this penalises the model for incorrect classifications,

All these transformations are neatly named the Negative Log Likehood. In practice then we get the softmax, calculate its log, and calculate the negative log likelihood based on that. This is Cross Entropy Loss! PyTorch does that in the following function:
loss_func = nn.CrossEntropyLoss()
loss_func(acts, targ)
This give our example an output of 1.8045. We can also split that loss value out into the individual values for each row:
nn.CrossEntropyLoss(reduction='none')(acts, targ)

1.8045 is just the mean of these values.
All-in-all this functions as an effective loss function because:
1. It works for multi-category problems
2. Softmax produces probabalistic activations
3. Softmax wants to pick a winning categorisation
4. The log functions allows for differentiation of small probability differences
Final benefit of Logs
One final note on the benefit of logs:
Log(a x b) = Log(a) + Log(b)
This simple expression is very useful when dealing with very small and very large numbers. Being able to replace multiplication with addition reduces risk of computational inaccuracies from floating point errors. Computers are then a lot less likely to end up with scales of numbers that they can’t handle.


























