This article will show a complete Convolutional Neural Network development and how to improve it in detail.
From TensorFlow library cifar10 dataset is used for this purpose:
import tensorflow as tf
(X_train, y_train), (X_test, y_test) = tf.keras.datasets.cifar10.load_data()Getting good accuracy in this dataset is pretty hard. If you ever used this dataset and tried a regular Sequential or functional Dense Neural Network, you already know that.
Let’s experiment with this dataset today and learn to use a CNN model.
The shape of that training features:
X_train.shapeOutput:
(50000, 32, 32, 3)Here is one image from the dataset:
import matplotlib.pyplot as plt
image=X_train[3]
plt.imshow(image)
plt.show()
The image quality is very poor. That is the reason this is hard to train a model on this. So, we will try to train a model on these poor-quality images today.
Normalizing the features:
X_train = X_train/255
X_test = X_test/255Here is the model. Explanation will follow:
model = tf.keras.Sequential([
tf.keras.layers.Conv2D(32, (3, 3), padding="valid",
activation="relu", input_shape=(32, 32, 3)),
tf.keras.layers.MaxPooling2D((2, 2), strides=2),
tf.keras.layers.Conv2D(48, (3, 3), padding="valid", activation="relu"),
tf.keras.layers.MaxPooling2D((2, 2), strides=2),
tf.keras.layers.Conv2D(48, (3, 3), padding="valid", activation="relu"),
tf.keras.layers.MaxPooling2D((2, 2), strides=2),
tf.keras.layers.Flatten(),
tf.keras.layers.Dense(100, activation="relu"),
tf.keras.layers.Dense(100, activation="relu"),
tf.keras.layers.Dense(100, activation="relu"),
tf.keras.layers.Dense(10, activation="softmax")]
)The model has two convolution layers. Let’s look at the hyperparameters of the first convolution layer:
32 is the number of filters.
(3, 3) is the size of the convolution window
Padding “valid” means no padding to the data. You can try with padding “same” as well that provides padding evenly to the left/right or up/down.
MaxPooling2D takes the maximum value of each window. There is the option to use min or avg as well.
(2, 2) is the stride length of the convolution.
Here is a tutorial that can give you a much more detailed understanding of how a Convolutional Neural network calculates the output:
Training the model:
from tensorflow.keras.callbacks import EarlyStopping
callbacks = [EarlyStopping(patience=5)]
model.compile(optimizer="adam",
loss=tf.keras.losses.SparseCategoricalCrossentropy(),
metrics=['accuracy'])
history = model.fit(X_train, y_train, epochs = 50,
validation_data = (X_test, y_test),
callbacks=callbacks)Output:
Epoch 1/50
1563/1563 ━━━━━━━━━━━━━━━━━━━━ 70s 43ms/step - accuracy: 0.3188 - loss: 1.7970 - val_accuracy: 0.5137 - val_loss: 1.3438
Epoch 2/50
1563/1563 ━━━━━━━━━━━━━━━━━━━━ 81s 43ms/step - accuracy: 0.5554 - loss: 1.2425 - val_accuracy: 0.5958 - val_loss: 1.1289
Epoch 3/50
1563/1563 ━━━━━━━━━━━━━━━━━━━━ 83s 43ms/step - accuracy: 0.6199 - loss: 1.0723 - val_accuracy: 0.6359 - val_loss: 1.0221
Epoch 4/50
1563/1563 ━━━━━━━━━━━━━━━━━━━━ 68s 44ms/step - accuracy: 0.6595 - loss: 0.9657 - val_accuracy: 0.6409 - val_loss: 1.0213
Epoch 5/50
1563/1563 ━━━━━━━━━━━━━━━━━━━━ 79s 42ms/step - accuracy: 0.6907 - loss: 0.8766 - val_accuracy: 0.6586 - val_loss: 0.9833
Epoch 6/50
1563/1563 ━━━━━━━━━━━━━━━━━━━━ 83s 42ms/step - accuracy: 0.7178 - loss: 0.8101 - val_accuracy: 0.6697 - val_loss: 0.9645
Epoch 7/50
1563/1563 ━━━━━━━━━━━━━━━━━━━━ 65s 42ms/step - accuracy: 0.7295 - loss: 0.7695 - val_accuracy: 0.6893 - val_loss: 0.9025
Epoch 8/50
1563/1563 ━━━━━━━━━━━━━━━━━━━━ 84s 43ms/step - accuracy: 0.7508 - loss: 0.7103 - val_accuracy: 0.6967 - val_loss: 0.8824
Epoch 9/50
1563/1563 ━━━━━━━━━━━━━━━━━━━━ 81s 43ms/step - accuracy: 0.7649 - loss: 0.6689 - val_accuracy: 0.7048 - val_loss: 0.8752
Epoch 10/50
1563/1563 ━━━━━━━━━━━━━━━━━━━━ 81s 42ms/step - accuracy: 0.7705 - loss: 0.6458 - val_accuracy: 0.6949 - val_loss: 0.9075
Epoch 11/50
1563/1563 ━━━━━━━━━━━━━━━━━━━━ 64s 41ms/step - accuracy: 0.7811 - loss: 0.6234 - val_accuracy: 0.7085 - val_loss: 0.8869
Epoch 12/50
1563/1563 ━━━━━━━━━━━━━━━━━━━━ 84s 42ms/step - accuracy: 0.7930 - loss: 0.5849 - val_accuracy: 0.6977 - val_loss: 0.9042
Epoch 13/50
1563/1563 ━━━━━━━━━━━━━━━━━━━━ 64s 41ms/step - accuracy: 0.8062 - loss: 0.5519 - val_accuracy: 0.7011 - val_loss: 0.9429
Epoch 14/50
1563/1563 ━━━━━━━━━━━━━━━━━━━━ 66s 42ms/step - accuracy: 0.8106 - loss: 0.5349 - val_accuracy: 0.7051 - val_loss: 0.9212Training accuracy is 81.06% and validation accuracy is 70.51%. Serious overfitting, right?
Hyperparameter Tuning
Here we will see how to use keras tuner for hyperparameter tuning.
You may have to install it. I installed it in the Google colab environment:
!pip install keras-tunerThis time we will define the model in a function like this:
from tensorflow import keras
from tensorflow.keras.layers import Conv2D, MaxPooling2D, Dense, Flatten, Activation
from kerastuner.tuners import RandomSearch
def buil_model(hp):
model = keras.models.Sequential()
model.add(tf.keras.layers.Conv2D(hp.Int('input_units',
min_value=64,
max_value=216,
step=32), (3, 3), input_shape=X_train.shape[1:]))
model.add(Conv2D(hp.Int('input_units1',
min_value=64,
max_value=216,
step=32), (3, 3), input_shape=X_train.shape[1:]))
model.add(Activation('relu'))
model.add(MaxPooling2D(pool_size=(2, 2)))
model.add(Flatten())
hp_dense1 = hp.Int('l1', min_value = 128, max_value = 512, step = 32)
model.add(keras.layers.Dense(units = hp_dense1, activation='relu'))
hp_dense2 = hp.Int('l2', min_value = 128, max_value = 512, step = 32)
model.add(keras.layers.Dense(units = hp_dense2, activation='relu'))
hp_lr = hp.Choice('learning_rate', values=[0.001, 0.0001])
model.add(keras.layers.Dense(10, activation='softmax'))
model.compile(optimizer="adam",
loss = "sparse_categorical_crossentropy",
metrics=["accuracy"])
return modelPlease notice the difference. Here we are providing a name of the layer in each layer. For example, the name of the first convolution layer is ‘input_units’,
Instead of using one concrete input of the number of kernels, here we have min_value and max_value so that the Keras tuner can use a range to choose the best number of kernels from.
Define the tuner:
tuner = RandomSearch(
buil_model,
objective = kerastuner.Objective('val_accuracy', direction = 'max'),
max_trials=5,
executions_per_trial=2,
directory="/content/drive.MyDrive/Colab Notebooks/TensorFlow",
project_name="cnn_tuner"
)I used RandomSearch here. Please feel free to try Hyperband and BayesianOptimization tuner as well.
Here is the documentation:
https://www.tensorflow.org/tutorials/keras/keras_tuner
I used max_trial and execution_per_trial 3 but you can use more. But the more trials you use the better it is to choose the right hyperparameters in between but it takes a lot of time. So for tutorial purposes, I picked only 3.
Model training:
tuner.search(X_train, y_train, epochs=5, validation_data = (X_test, y_test))Output:
Trial 5 Complete [00h 02m 43s]
val_accuracy: 0.6708999872207642
Best val_accuracy So Far: 0.6921499967575073
Total elapsed time: 00h 17m 34sThe conditions of all five trials:
tuner.search_space_summary()Output:
Search space summary
Default search space size: 5
input_units (Int)
{'default': None, 'conditions': [], 'min_value': 64, 'max_value': 216, 'step': 32, 'sampling': 'linear'}
input_units1 (Int)
{'default': None, 'conditions': [], 'min_value': 64, 'max_value': 216, 'step': 32, 'sampling': 'linear'}
l1 (Int)
{'default': None, 'conditions': [], 'min_value': 128, 'max_value': 512, 'step': 32, 'sampling': 'linear'}
l2 (Int)
{'default': None, 'conditions': [], 'min_value': 128, 'max_value': 512, 'step': 32, 'sampling': 'linear'}
learning_rate (Choice)
{'default': 0.001, 'conditions': [], 'values': [0.001, 0.0001], 'ordered': True}The results of all the trials:
tuner.results_summary()Output:
Results summary
Results in /content/drive.MyDrive/Colab Notebooks/TensorFlow/cnn_tuner
Showing 10 best trials
Objective(name="val_accuracy", direction="max")
Trial 0 summary
Hyperparameters:
input_units: 128
input_units1: 128
l1: 288
l2: 416
learning_rate: 0.0001
Score: 0.6921499967575073
Trial 1 summary
Hyperparameters:
input_units: 64
input_units1: 160
l1: 128
l2: 160
learning_rate: 0.001
Score: 0.6796999871730804
Trial 3 summary
Hyperparameters:
input_units: 128
input_units1: 160
l1: 448
l2: 352
learning_rate: 0.0001
Score: 0.6796500086784363
Trial 2 summary
Hyperparameters:
input_units: 192
input_units1: 160
l1: 352
l2: 448
learning_rate: 0.0001
Score: 0.6732999980449677
Trial 4 summary
Hyperparameters:
input_units: 192
input_units1: 64
l1: 224
l2: 192
learning_rate: 0.001
Score: 0.6708999872207642Finding the best model:
final_model = tuner.get_best_models(num_models=1)
best_model = final_model[0]
best_model.summary()Output:
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━┓
┃ Layer (type) ┃ Output Shape ┃ Param # ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━┩
│ conv2d (Conv2D) │ (None, 30, 30, 128) │ 3,584 │
├──────────────────────────────────────┼─────────────────────────────┼─────────────────┤
│ conv2d_1 (Conv2D) │ (None, 28, 28, 128) │ 147,584 │
├──────────────────────────────────────┼─────────────────────────────┼─────────────────┤
│ activation (Activation) │ (None, 28, 28, 128) │ 0 │
├──────────────────────────────────────┼─────────────────────────────┼─────────────────┤
│ max_pooling2d (MaxPooling2D) │ (None, 14, 14, 128) │ 0 │
├──────────────────────────────────────┼─────────────────────────────┼─────────────────┤
│ flatten (Flatten) │ (None, 25088) │ 0 │
├──────────────────────────────────────┼─────────────────────────────┼─────────────────┤
│ dense (Dense) │ (None, 288) │ 7,225,632 │
├──────────────────────────────────────┼─────────────────────────────┼─────────────────┤
│ dense_1 (Dense) │ (None, 416) │ 120,224 │
├──────────────────────────────────────┼─────────────────────────────┼─────────────────┤
│ dense_2 (Dense) │ (None, 10) │ 4,170 │
└──────────────────────────────────────┴─────────────────────────────┴─────────────────┘
Total params: 7,501,194 (28.61 MB)
Trainable params: 7,501,194 (28.61 MB)
Non-trainable params: 0 (0.00 B)Here we have our best models. I tried with only 5 epochs to save time. Please feel free to try more epochs.