Before the famous transformer architecture was invented, the state of the art networks for natural language processing were Recurrent Neural Networks (RNNs). So before Transformers, why did we use RNNs instead of standard neural networks? Well the issue with texts is that the outputs or inputs could vary in lengths, meaning they aren't always fixed. Another issue with standard neural networks is also because they don't share learned features across different positions in texts. Meaning no learned information of certain words are passed on to the next words. This is where RNNs solve this issue. They pass on learned information from the last word or piece of the sentence to the next layer.
RNNs actually have multiple different variants, which we will go through in this section
Many-to-one:
This architecture is typically used for sentiment analysis. So when you are given a text like the one below, the model simply gives out a 0 or 1 label output, so either good sentiment (1) or bad sentiment (0).
In music generation, you usually give the model a single input and let the model generate the song or music on its own.
Such encoder-decoder variants are typically used for translation tasks. This structure was actually used by google translate, before Transformers were introduced.
There is another many to many variant of RNNs. This version can used for entity recognition. For example for the sentence "Harry Potter and Hermione Granger invented a new spell", it might flag names like Harry and Potter as part of a person or name.
If you are really curious for how one cell actually works under the hood, here it is. You can refer to the diagram above to try and understand the computation in one cell. Intuitively, one cell of an RNN recieves information about the sentence previously through a_t-1 and gets the new word input x_t. It then performs a computation resulting in a_t. Then a_t is used to make a prediction on a given word, so for the Harry Potter example it will flag 1 if the word is a name and 0 if it isn't.
Right now, Transformers have pretty much replaced most RNN use cases in Natural language processing (NLP). The examples I gave you were simply for understanding, even if the architecture is less relevant in NLP nowadays. So why do we still use it? Well since RNNs are good sequential models, they are currently used for time series and forecasting tasks like stock market predictions.
However the earlier standard RNN celI just showed earlier isn't usually used for stock market prediction. Because it is actually very sensitive to noise in past data and it thus may struggle on longer sequence lengths.
If you don't care about going to deep into RNNs, you can totally skip this part, but I will leave the mathematics here for anyone who is curious or if they want to try implementing these architectures themselves.
GRUs have the same inputs and outputs as a normal RNN cell. But what is different about a GRU cell is an increase in complexity in computation. First the update gate determines how much information should be carried into the future. The reset gate then determines what information should be forgotten in the past. This makes the GRU cell much more resistant to noise than standard RNNs.
The maths behind LSTMs is a bit more complex than the previous cells. So the major difference between LSTMs are the inputs and outputs. Alongside the x_t and a_t-1 input, it also receives a cell state c_t-1 that it itself outputs. The input gate of an LSTM decides which information from the previous cell is relevant. The forget gate decides what information to discard from the previous cell state. The output gate simply decides what part of the short term input (a_t-1) should be kept for the prediction of the current cell. The cells (c_t-1) are actually the LSTMs long term memory, which is why it is passed alongside a_t-1. This computation is done with the cell and the output gate, for the model's final output prediction.