IntroKeras
Simplest Introduction to Neural Networks with Keras
This notebook is a part of AI for Beginners Curricula. Visit the repository for complete set of learning materials.
Neural Frameworks
There are several frameworks for training neural networks. However, if you want to get started fast and not go into much detail on how things work internally - you should consider using Keras. This short tutorial will get you started, and if you want to get deeper into understanding how things work - look into Introduction to Tensorflow and Keras notebook.
Getting things ready
Keras is a part of Tensorflow 2.x framework. Let's make sure we have version 2.x.x of Tensorflow installed:
pip install tensorflow
or
conda install tensorflow
Tensorflow version = 2.7.0 Keras version = 2.7.0
Basic Concepts: Tensor
Tensor is a multi-dimensional array. It is very convenient to use tensors to represent different types of data:
- 400x400 - black-and-white picture
- 400x400x3 - color picture
- 16x400x400x3 - minibatch of 16 color pictures
- 25x400x400x3 - one second of 25-fps video
- 8x25x400x400x3 - minibatch of 8 1-second videos
Tensors give us a convenient way to represent input/output data, as well we weights inside the neural network.
Sample Problem
Let's consider binary classification problem. A good example of such a problem would be a tumour classification between malignant and benign based on it's size and age. Let's start by generating some sample data:
C:\Users\dmitryso\AppData\Local\Temp/ipykernel_103052/2721537645.py:17: UserWarning: Matplotlib is currently using module://matplotlib_inline.backend_inline, which is a non-GUI backend, so cannot show the figure. fig.show()
Normalizing Data
Before training, it is common to bring our input features to the standard range of [0,1] (or [-1,1]). The exact reasons for that we will discuss later in the course, but in short the reason is the following. We want to avoid values that flow through our network getting too big or too small, and we normally agree to keep all values in the small range close to 0. Thus we initialize the weights with small random numbers, and we keep signals in the same range.
When normalizing data, we need to subtract min value and divide by range. We compute min value and range using training data, and then normalize test/validation dataset using the same min/range values from the training set. This is because in real life we will only know the training set, and not all incoming new values that the network would be asked to predict. Occasionally, the new value may fall out of the [0,1] range, but that's not crucial.
Training One-Layer Network (Perceptron)
In many cases, a neural network would be a sequence of layers. It can be defined in Keras using Sequential model in the following manner:
Model: "sequential_2"
_________________________________________________________________
Layer (type) Output Shape Param #
=================================================================
dense_2 (Dense) (None, 1) 3
activation_1 (Activation) (None, 1) 0
=================================================================
Total params: 3
Trainable params: 3
Non-trainable params: 0
_________________________________________________________________
Here, we first create the model, and then add layers to it:
- First
Inputlayer (which is not strictly speaking a layer) contains the specification of network's input size Denselayer is the actual perceptron that contains trainable weights- Finally, there is a layer with sigmoid
Activationfunction to bring the result of the network into 0-1 range (to make it a probability).
Input size, as well as activation function, can also be specified directly in the Dense layer for brevity:
Model: "sequential_9"
_________________________________________________________________
Layer (type) Output Shape Param #
=================================================================
dense_9 (Dense) (None, 1) 3
=================================================================
Total params: 3
Trainable params: 3
Non-trainable params: 0
_________________________________________________________________
Before training the model, we need to compile it, which essentially mean specifying:
- Loss function, which defines how loss is calculated. Because we have two-class classification problem, we will use binary cross-entropy loss.
- Optimizer to use. The simplest option would be to use
sgdfor stochastic gradient descent, or you can use more sophisticated optimizers such asadam. - Metrics that we want to use to measure success of our training. Since it is classification task, a good metrics would be
Accuracy(oraccfor short)
We can specify loss, metrics and optimizer either as strings, or by providing some objects from Keras framework. In our example, we need to specify learning_rate parameter to fine-tune learning speed of our model, and thus we provide full name of Keras SGD optimizer.
After compiling the model, we can do the actual training by calling fit method. The most important parameters are:
xandyspecify training data, features and labels respectively- If we want validation to be performed on each epoch, we can specify
validation_dataparameter, which would be a tuple of features and labels epochsspecified the number of epochs- If we want training to happen in minibatches, we can specify
batch_sizeparameter. You can also pre-batch the data manually before passing it tox/y/validation_data, in which case you do not needbatch_size
Epoch 1/10 70/70 [==============================] - 0s 4ms/step - loss: 0.3379 - acc: 0.9000 - val_loss: 0.3282 - val_acc: 0.9000 Epoch 2/10 70/70 [==============================] - 0s 2ms/step - loss: 0.3270 - acc: 0.9429 - val_loss: 0.3336 - val_acc: 0.9000 Epoch 3/10 70/70 [==============================] - 0s 2ms/step - loss: 0.3195 - acc: 0.9143 - val_loss: 0.3137 - val_acc: 0.9000 Epoch 4/10 70/70 [==============================] - 0s 2ms/step - loss: 0.3087 - acc: 0.9286 - val_loss: 0.2970 - val_acc: 0.9333 Epoch 5/10 70/70 [==============================] - 0s 3ms/step - loss: 0.3006 - acc: 0.9429 - val_loss: 0.3210 - val_acc: 0.9000 Epoch 6/10 70/70 [==============================] - 0s 3ms/step - loss: 0.3003 - acc: 0.9000 - val_loss: 0.2985 - val_acc: 0.9000 Epoch 7/10 70/70 [==============================] - 0s 3ms/step - loss: 0.2956 - acc: 0.9286 - val_loss: 0.3037 - val_acc: 0.9000 Epoch 8/10 70/70 [==============================] - 0s 3ms/step - loss: 0.2891 - acc: 0.9429 - val_loss: 0.3035 - val_acc: 0.9000 Epoch 9/10 70/70 [==============================] - 0s 3ms/step - loss: 0.2809 - acc: 0.9000 - val_loss: 0.2815 - val_acc: 0.9000 Epoch 10/10 70/70 [==============================] - 0s 3ms/step - loss: 0.2809 - acc: 0.9286 - val_loss: 0.2907 - val_acc: 0.9000
<keras.callbacks.History at 0x2c420b89910>
You can try to experiment with different training parameters to see how they affect the training:
- Setting
batch_sizeto be too large (or not specifying it at all) may result in less stable training, because with low-dimensional data small batch sizes provide more precise direction of the gradient for each specific case - Too high
learning_ratemay result in overfitting, or in less stable results, while too low learning rate means it will take more epochs to achieve the result
Note that you can call
fitfunction several times in a row to further train the network. If you want to start training from scratch - you need to re-run the cell with the model definition.
To make sure our training worked, let's plot the line that separates two classes. Separation line is defined by the equation
C:\Users\dmitryso\AppData\Local\Temp/ipykernel_103052/2721537645.py:17: UserWarning: Matplotlib is currently using module://matplotlib_inline.backend_inline, which is a non-GUI backend, so cannot show the figure. fig.show()
Plotting the training graphs
fit function returns history object as a result, which can be used to observe loss and metrics on each epoch. In the example below, we will re-start the training with small learning rate, and will observe how the loss and accuracy behave.
Note that we are using slightly different syntax for defining
Sequentialmodel. Instead ofadd-ing layers one by one, we can also specify the list of layers right when creating the model in the first place - this is a bit shorter syntax, and you may prefer to use it.
Epoch 1/10 70/70 [==============================] - 1s 5ms/step - loss: 0.6600 - acc: 0.6143 - val_loss: 0.6351 - val_acc: 0.8000 Epoch 2/10 70/70 [==============================] - 0s 2ms/step - loss: 0.6384 - acc: 0.7143 - val_loss: 0.6187 - val_acc: 0.8333 Epoch 3/10 70/70 [==============================] - 0s 2ms/step - loss: 0.6188 - acc: 0.7571 - val_loss: 0.6001 - val_acc: 0.8667 Epoch 4/10 70/70 [==============================] - 0s 3ms/step - loss: 0.6022 - acc: 0.7714 - val_loss: 0.5837 - val_acc: 0.9000 Epoch 5/10 70/70 [==============================] - 0s 2ms/step - loss: 0.5860 - acc: 0.8571 - val_loss: 0.5673 - val_acc: 0.9000 Epoch 6/10 70/70 [==============================] - 0s 2ms/step - loss: 0.5702 - acc: 0.8571 - val_loss: 0.5597 - val_acc: 0.8667 Epoch 7/10 70/70 [==============================] - 0s 2ms/step - loss: 0.5568 - acc: 0.8286 - val_loss: 0.5458 - val_acc: 0.9000 Epoch 8/10 70/70 [==============================] - 0s 2ms/step - loss: 0.5430 - acc: 0.8714 - val_loss: 0.5325 - val_acc: 0.9000 Epoch 9/10 70/70 [==============================] - 0s 2ms/step - loss: 0.5308 - acc: 0.8714 - val_loss: 0.5234 - val_acc: 0.9000 Epoch 10/10 70/70 [==============================] - 0s 3ms/step - loss: 0.5175 - acc: 0.9143 - val_loss: 0.5170 - val_acc: 0.8667
[<matplotlib.lines.Line2D at 0x2c41a32fe80>]
Multi-Class Classification
If you need to solve a problem of multi-class classification, your network would have more that one output - corresponding to the number of classes . Each output will contain the probability of a given class.
Note that you can also use a network with two outputs to perform binary classification in the same manner. That is exactly what we will demonstrate now.
When you expect a network to output a set of probabilities , we need all of them to add up to 1. To ensure this, we use softmax as a final activation function on the last layer. Softmax takes a vector input, and makes sure that all components of that vector are transformed into probabilities.
Also, since the output of the network is a -dimensional vector, we need labels to have the same form. This can be achieved by using one-hot encoding, when the number of a class is converted to a vector of zeroes, with 1 at the -th position.
To compare the probability output of the neural network with expected one-hot-encoded label, we use cross-entropy loss function. It takes two probability distributions, and outputs a value of how different they are.
So, to summarize what we need to do for multi-class classification with classes:
- The network should have neurons in the last layer
- Last activation function should be softmax
- Loss should be cross-entropy loss
- Labels should be converted to one-hot encoding (this can be done using
numpy, or using Keras utilsto_categorical)
Epoch 1/10 70/70 [==============================] - 1s 6ms/step - loss: 0.6524 - acc: 0.7000 - val_loss: 0.5936 - val_acc: 0.9000 Epoch 2/10 70/70 [==============================] - 0s 2ms/step - loss: 0.5715 - acc: 0.8286 - val_loss: 0.5255 - val_acc: 0.8333 Epoch 3/10 70/70 [==============================] - 0s 3ms/step - loss: 0.4820 - acc: 0.8714 - val_loss: 0.4213 - val_acc: 0.9000 Epoch 4/10 70/70 [==============================] - 0s 3ms/step - loss: 0.4426 - acc: 0.9000 - val_loss: 0.3694 - val_acc: 0.9333 Epoch 5/10 70/70 [==============================] - 0s 3ms/step - loss: 0.3602 - acc: 0.9000 - val_loss: 0.3454 - val_acc: 0.9000 Epoch 6/10 70/70 [==============================] - 0s 3ms/step - loss: 0.3209 - acc: 0.8857 - val_loss: 0.2862 - val_acc: 0.9333 Epoch 7/10 70/70 [==============================] - 0s 3ms/step - loss: 0.2905 - acc: 0.9286 - val_loss: 0.2787 - val_acc: 0.9000 Epoch 8/10 70/70 [==============================] - 0s 3ms/step - loss: 0.2698 - acc: 0.9000 - val_loss: 0.2381 - val_acc: 0.9333 Epoch 9/10 70/70 [==============================] - 0s 3ms/step - loss: 0.2639 - acc: 0.8857 - val_loss: 0.2217 - val_acc: 0.9667 Epoch 10/10 70/70 [==============================] - 0s 2ms/step - loss: 0.2592 - acc: 0.9286 - val_loss: 0.2391 - val_acc: 0.9000
Sparse Categorical Cross-Entropy
Often labels in multi-class classification are represented by class numbers. Keras also supports another kind of loss function called sparse categorical crossentropy, which expects class number to be integers, and not one-hot vectors. Using this kind of loss function, we can simplify our training code:
Epoch 1/10 70/70 [==============================] - 1s 6ms/step - loss: 0.2353 - acc: 0.9143 - val_loss: 0.2190 - val_acc: 0.9000 Epoch 2/10 70/70 [==============================] - 0s 3ms/step - loss: 0.2243 - acc: 0.9286 - val_loss: 0.1886 - val_acc: 0.9333 Epoch 3/10 70/70 [==============================] - 0s 2ms/step - loss: 0.2366 - acc: 0.9143 - val_loss: 0.2262 - val_acc: 0.9000 Epoch 4/10 70/70 [==============================] - 0s 2ms/step - loss: 0.2259 - acc: 0.9429 - val_loss: 0.2124 - val_acc: 0.9000 Epoch 5/10 70/70 [==============================] - 0s 2ms/step - loss: 0.2061 - acc: 0.9429 - val_loss: 0.2691 - val_acc: 0.9000 Epoch 6/10 70/70 [==============================] - 0s 2ms/step - loss: 0.2200 - acc: 0.9286 - val_loss: 0.2344 - val_acc: 0.9000 Epoch 7/10 70/70 [==============================] - 0s 3ms/step - loss: 0.2133 - acc: 0.9286 - val_loss: 0.1973 - val_acc: 0.9000 Epoch 8/10 70/70 [==============================] - 0s 3ms/step - loss: 0.2062 - acc: 0.9429 - val_loss: 0.1893 - val_acc: 0.9000 Epoch 9/10 70/70 [==============================] - 0s 3ms/step - loss: 0.2060 - acc: 0.9571 - val_loss: 0.2719 - val_acc: 0.9000 Epoch 10/10 70/70 [==============================] - 0s 3ms/step - loss: 0.2021 - acc: 0.9571 - val_loss: 0.2293 - val_acc: 0.9000
<keras.callbacks.History at 0x2c42267de80>
Multi-Label Classification
Sometime we have cases when our objects can belong to two classes at once. As an example, suppose we want to develop a classifier for cats and dogs on the picture, but we also want to allow cases when both cats and dogs are present.
With multi-label classification, instead of one-hot encoded vector, we will have a vector that has 1 in position corresponding to all classes relevant to the input sample. Thus, output of the network should not have normalized probabilities for all classes, but rather for each class individually - which corresponds to using sigmoid activation function. Cross-entropy loss can still be used as a loss function.
Note that this is very similar to using different neural networks to do binary classification for each particular class - only the initial part of the network (up to final classification layer) is shared for all classes.
Summary of Classification Loss Functions
We have seen that binary, multi-class and multi-label classification differ by the type of loss function and activation function on the last layer of the network. It may all be a little bit confusing if you are just starting to learn, but here are a few rules to keep in mind:
- If the network has one output (binary classification), we use sigmoid activation function, for multiclass classification - softmax
- If the output class is represented as one-hot-encoding, the loss function will be cross entropy loss (categorical cross-entropy), if the output contains class number - sparse categorical cross-entropy. For binary classification - use binary cross-entropy (same as log loss)
- Multi-label classification is when we can have an object belonging to several classes at the same time. In this case, we need to encode labels using one-hot encoding, and use sigmoid as activation function, so that each class probability is between 0 and 1.
| Classification | Label Format | Activation Function | Loss |
|---|---|---|---|
| Binary | Probability of 1st class | sigmoid | binary crossentropy |
| Binary | One-hot encoding (2 outputs) | softmax | categorical crossentropy |
| Multiclass | One-hot encoding | softmax | categorical crossentropy |
| Multiclass | Class Number | softmax | sparse categorical crossentropy |
| Multilabel | One-hot encoding | sigmoid | categorical crossentropy |
Task: Use Keras to train a classifier for MNIST handwritten digits:
- Notice that Keras contains some standard datasets, including MNIST. To use MNIST from Keras, you only need a couple of lines of code (more information here)
- Try several network configuration, with different number of layers/neurons, activation functions.
What is the best accuracy you were able to achieve?
Takeaways
- Keras is really recommended for beginners, because it allows to construct networks from layers quite easily, and then train it with just a couple of lines of code
- If non-standard architecture is needed, you would need to learn a bit deeper into Tensorflow. Or you can ask someone to implement custom logic as a Keras layer, and then use it in Keras models
- It is a good idea to look at PyTorch as well and compare approaches.
A good sample notebook from the creator of Keras on Keras and Tensorflow 2.0 can be found here.