Convolutional Neural Networks (CNNs) are typically used in image recognition tasks, though their applications aren't only limited to that. For the purpose of understanding how they work, we will simply focus on image recognition.
This is the layer that is essentially responsible for finding and identifying the edges within an image, this helps the model learn the shape of an object like a dog or cat on an image
A filter slides over each portion of the image, until it reaches the end.
Let's look at the green section to see what is happening. Each element in the image correspoding with the filter's element. Since the number 3 is at the top left, it will be multiplied with 1 at the top left at the convolution filter. Since 2 is at the right bottom of the image, it is multiplied with the bottom right of the convolution filter. All elements are then added up together in the end. For the green filter the result is -5. To cover the entire image, we slide this filter by one, so -4 is the result of the second section and so on.
Here is a gif to better visualize this process. Just as we described earlier, the filter slides by one column until it reaches the end, then it goes down by one row and repeats until it reaches the end of the image.
The previous example uses a simple 2D image that can only be represented in black and white. However most images today use colours by encoding them in RGB. To effectively represent this, we have to use multiple channels. Furthermore, channels also help the model to find complex patterns and relationships in the image, such as the shape of an object, in order for it to make a better prediction.
You can think of batches as essentially the number of 2D grids stacked together. For an RGB image, we can use one 2D grid for red, one for green and another for blue. In this case, we effectively have 3 channels.
When we actually apply a convolution layer to this, it will be somewhat different to our previous one channel example. The convolution filter we use also uses three channels.
The reason why our result isn't three channels as well, is because that orange dot is the sum of all the outputs of each convolution on each channel. So red had convolutions with the first channel of the convolutional filter, then green also had a convolution operation with the second layer and same for blue with the third channel. These resulting three numbers are added together, which is what that orange dot represents. The mechanism of sliding by one is the same as the previous example. However this isn't what we actually want for a full convolutional neural network, we actually want more channels than the first RGB picture.
To actually have more channels for the output, we need to apply multiple different convolution filters multiple times like so:
We used a second convolutional filter but with different numbers from the first one, as a result we now have two channels in our resulting output. We can apply multiple channels for an arbitrary number of times. But using thi ssame principle, we can actually create a full convolutional neural network
This is Alex Net, one of the first large scale neural networks created. As we can see, each layer has changes in the number of channels jumping from 3 to 96. That jump as can be induced from the previous example, happened because 96 convolutional filters were applied to the image. This process then happens again across different layers.