RNNs used to be the state of the art architecture for natural language processing tasks or text based tasks. However, because of their sequential nature, their architecture's capabilities was limited to taking each token sequentially, meaning it was hard for them to take the entire text sequence as an input all at once and find relationships between the words. In 2017, google researchers thought, "what if a text based architecture could take the entire text sequence at once (like CNNs did with images) instead of sequentially?". Based on this intuition and knowledge from other architectures, google researchers developed the famous transformer architecture, the primary architecture used to develop Large Language models like ChatGPT or Gemini.
This is the general standard architecture. The inputs are the tokens (note: tokens could actually represent sub words in a word or parts of a word, not the entire word, this is another concept known as tokenization which won't be explored in this section) which are first embedded and transformed into vectors. Remember, these models can only find relationships between words through numbers, not through their grammar. This main block containing the add&Norm, position-wise feed forward and Multi-Head Attention modules is repeated by an arbitrary number of times. It could be 10 times, 20 times and for bigger models like Deepseek roughly 61 times.
The X input which is the piece of text, which could be a sentence, containing words encoded as vectors, with each vector representing a token in the sentence. For the sake of simplicity, let's assume each token represents a word. The sentences X are transformed through the Weights into three separate Q, K and Vs matrices, which are different representations of the sentence.
Now why did we just do that? Well after turning the sentence into three separate Q, V and K matrices, we then perform this computation. We will break this computation down piece by piece
First the computation between QK^T, this is responsible for helping the model understand relationships between words. In this example, the model finds relationships between the word American to red and white This allows the model to understand the relationship between words and how relevant each word is to another.
Source: Welch Labs
Now multiplying the numbers together usually makes them bigger. This may cause numerical instability, so we apply a normalizing factor by dividing the values by the model's dimension length. After this, the model then applies the softmax function, which is what we will look at next. After the softmax function is applied, the multiplication of the V matrix. The V matrix is essentially responsible for helping the model what each word means in the context of the sentence. For example the word might recognize that bolt might mean the object in one context or running away fast in another context.
Why do we apply softmax in multi head attention? The numbers we just normalized are still too large and they are simply raw scores. If we want to turn these raw numbers into weighted scores, we need to turn them into probabilites (meaning that each number is from 0 to 1).
So let's break down how this works, the z input is actually a vector and i is the index of an element. Now let's consider a z row vector as an example
When we apply the softmax function to this entire row vector, we get this as the result:
Let's focus on what happened to the first element in order to fully understand the computation behind this function:
First we compute the exponents, for the first element, we only need e^z1 for the numerator
To find the denominator, we take the sum of all the numbers above:
Now we simply divide the numerator by the denomiantor and we get the same result as the first element of the vector
While this example was a simple row vector, this idea can be extended for a matrix, by applying the function for each row independantly
This is the only other important module in the transformer which can simply be thought of as a standard neural network. The input x is passed through one linear layer, then the ReLU activation functions is applied. The output of that ReLU function is then passed to another layer. The first layer expands the model input's embedding dimension length, in order to expand the model's representations for each words in higher dimension and to find new patterns. After ReLU is applied to find non linear patterns, the second layer projects the embedding dimension length back to its original length in order to condense the information.
If you didn't understand the section above, then you may have to review how neural networks work in the intro to deep learning section