github links:
PPO Lunar Lander
https://github.com/Thomasche69/PPO-Lunar-Lander
PPO Bipedal Walker
https://github.com/Thomasche69/PPO-Bipedal-Walker
PPO racing car
https://github.com/Thomasche69/PPO-racing-car
Before I dive into this project, let's go over the basic idea behind reinforcement learning.
Reinforcement learning works completely different to the other projects I have explored. My other projects utilized supervised learning which is based on the idea of giving a model an input and predicting a label or outcome. In reinforcement learning, the goal of the AI model is to take action in an environment given a state and reward. A good example could be the mazegame above. The state could be where the AI is within the maze and the reward given to the model depends on where it is. If the model ends up in a dead end, the reward could be -10 however if the model makes it to the end, the model gets +10 reward.
To sum it up, the model gets a state, it then takes action within that state and as a result of its actions it gets a reward.
The library I used for this project is Open AI's gym library, which is often used to test algorithms. It provides simple multiple simple environments like car racing, lunar lander and bipedal walking. The general setup for these environments when doing my projects is mostly the same.
The algorithm I used for all setups was PPO. I won't go into how exactly it works, but essentially it consists of an actor and critic network. The actor network takes actions in an environment and the critic network evaluates the actor network's actions. Based on the evaluation it updates itself and the actor network so that the model can make better decisions in an environment.
This is quite a lot of code but I will go through the general idea. The hyperparameters are the parameters I have to set by myself (meaning the model can't learn it) which will determine the model's overall performance.
First the hyperparameters (learning_rate, n_steps, etc.) will be randomly initialized within their defined ranges. What the model will first do is go through multiple loops of trial and error in the environment, at around 50000 timesteps. Then the model will be evaluated 10 times, and its performance will be determined by the reward it gets from the environment.
The objective function will be repeated for 200 trials, in order to find the best combination of hyperparameters for the best possible performance
After the best hyperparameters are found, the model is then trained through a million steps in order to perform most optimally within the envrionment.
After training the algorithm on different environments, here are the results which are shown in video below