The goal of Lab 05 is to train a neural locomotion policy for the Pupper robot that allows you to command the robot to walk at velocities that you provide using a PS3 controller! This requires setting up the cloud training environment and tuning some reward scales (the reward functions are implemented for you; you are responsible for tuning their scales). You will also test the robot's default policy and compare your controller against it.
You will write your lab report in Overleaf using the LaTeX template provided in the Lab5 directory of the course repository. Upload the template file, make sure the document compiles, and fill in your team names. The report you'll upload is the final compiled PDF from Overleaf.
Important: We want you to learn the basics of reward definition and reward tuning on your own, so we do not allow AI to solve these tasks for you! We won't grade groups that use AI, and accept no excuses.
In this section, you will set up the cloud computing environment required to train and run your models. Because this lab involves heavy computation, you will use Google Colab Pro and Weights & Biases (W&B) to track your experiments. YOU CAN DO THIS SECTION OF WORK ON YOUR LAPTOP.
Step 1: Navigate to the Google Colab education signup page at https://colab.research.google.com/signup. Register for a student Google Colab Pro account and verify it using your official student email address.
Step 2: Download the base Jupyter Notebook for this lab from here. Once downloaded, upload this .ipynb file directly to the Google Drive associated with the student account you just created.
Step 3: Create an account on Weights & Biases by visiting https://wandb.ai/site. Once logged in, navigate to your user settings (by clicking your profile icon) to generate an API key. Copy this API key and store it somewhere accessible, as you will need it to authenticate your notebook in the next step.
Step 4: Open your Google Drive and double-click the uploaded Jupyter Notebook to open it in Google Colab.
Step 5: Configure your notebook to use a high-performance GPU. In the top menu of Colab, click on Runtime -> Change runtime type. Under the Hardware accelerator dropdown, select A100 GPU and save your settings.
Step 6: In the first code cell of your notebook, you will see a designated spot for your Weights & Biases credentials. Paste your copied W&B API key there and run the cell to authenticate your session.
In this section, you will train a walking policy! Your main objective is to experiment with and properly adjust training parameters (reward, domain randomization, etc) to make a natural walking behavior. YOU CAN DO THIS SECTION OF WORK ON YOUR LAPTOP. You will start this section in the lab, but you must continue working on it during the week and before your next lab.
The provided training notebook is organized into several key sections and leverages JAX’s powerful GPU acceleration to train Pupper in thousands of parallel environments. This massive parallelization dramatically speeds up the training process, allowing us to collect large amounts of experience data efficiently. Each environment runs an independent simulation of Pupper, enabling rapid exploration of different walking strategies and faster convergence to optimal policies. The key sections are (YOU CAN READ THESE PARTS AT HOME AND GO TO STEP1):
The reward function is crucial for training Pupper to walk effectively. This is the part of the notebook that you will be tuning for this lab. The notebook provides several reward terms that you can tune:
Velocity Tracking: Encourages Pupper to match desired linear and angular velocities
Effort Minimization: Penalizes excessive motor torques and energy consumption
Stability: Rewards maintaining a stable orientation and penalizes falling
Smoothness: Encourages smooth joint movements and penalizes jerky motions
Height: Rewards maintaining a desired body height
Foot Contact: Encourages proper foot placement and contact timing
Refer to the rewards.py file for reward definitions. You need to understand the exact implementation of each reward term to determine what coefficients to use on these rewards.
The MJX (MuJoCo + Jax) configuration controls the physics simulation environment:
Robot Parameters: Mass, inertia, and dimensions of Pupper’s components (you should not change these)
Control: Motor dynamics, PID gains, and actuation limits (you don’t need to change these for this lab)
PPO Configs: Proximal Policy Optimization (PPO) is a popular RL algorithm for training robot policies–you can think of it as an optimizer to train RL models for maximizing rewards. The PPO configuration controls the learning process: Network Architecture: Size and structure of the policy and value networks (you should not change these for this lab)
Training Parameters: Learning rate, batch size, number of epochs Policy Clipping: Limits on policy updates to ensure stable learning (you should not change these for this lab)
Value Function: Parameters for the value function estimation (you should not change these for this lab)
Entropy Bonus: Encourages exploration during training (you should not change these for this lab)
Command Sampling: Controls how velocity commands are generated during training: Linear Velocity: Range for forward/backward and lateral movement Angular Velocity: Range for turning commands Zero Command Probability: Chance of receiving a zero-velocity command (you should not change these for this lab)
Stand Still Threshold: Velocity threshold below which commands are considered “standing still” (you should not change these for this lab)
Defines when an episode should end:
Body Height: Episode ends if the body center goes below a certain height (you should not change these)
Body Angle: Episode ends if body angle exceeds a threshold (you should not change these)
Early Termination: Allows episodes to end before reaching maximum length (you should not change these)
Parameters that add variability to the simulation to improve robustness:
Perturbations: Random kicks, angular velocity noise, and gravity variations
Motor Properties: Random variations in position control gains
Starting Position: Random initial positions for training (you should not change these)
Latency: Simulated delays in action execution and IMU readings (you should not change these)
Body Properties: Random variations in mass, inertia, and center of mass
Friction: Random variations in ground friction
Environment features to test Pupper’s capabilities:
Heightfield Types: Random terrain or steps
Heightfield Parameters: Grid size, radius, and elevation
Obstacles: Number, position, and dimensions of obstacles
Understanding and tuning these parameters is key to training an effective walking policy. We will start with the basic velocity-tracking reward and gradually add additional terms to improve Pupper’s walking behavior.
DELIVERABLE: Before actually training a policy, what do you think will be the most important rewards to tune when training Pupper to walk forward? What about making Pupper walk stably? Could these two factors have a combined effect and interfere? Write a few sentences in your lab report.
First, let's implement a naive reward function for Pupper velocity tracking. In the Reward Configuration section, change the tracking_lin_vel and tracking_ang_vel values to some nonzero values to get Pupper to follow a velocity command. In practice, the linear velocity tracking coefficient should be around double the angular velocity tracking. Run the entire notebook, which loads in all the training and MJX configs, initializes Pupper in a flat environment, and trains Pupper to follow a desired velocity.
DELIVERABLE: Visualize Pupper’s progress during training. How does Pupper look in the first 20 million env steps? How does it look after 200 million env steps? You need to upload the videos of the final trained policy and call it step1_sim.mp4.
DELIVERABLE: What is the mathematical expression of the reward term? How does it encourage tracking linear velocity? Should the reward scale for this term be positive or negative?
DELIVERABLE: Download a copy of your notebook, rename it to step1.ipynb, and include it in your submission.
Now that you have a basic walking policy, let's make it walk more efficiently (the robot should be a bit lazy and avoid wasting energy :-) ). Edit the Reward Configuration section to tune a reward function that helps Pupper conserve effort. When tuning the reward function, think about which reward coefficients should be nonzero to encourage Pupper to conserve energy. Should the coefficients be positive or negative? Rerun the entire notebook to initialize Pupper in a flat environment and train Pupper to walk forward more efficiently.
DELIVERABLE: What is your reward function (in math, don’t just take a screenshot from the notebook!)? Why did you choose this function? DELIVERABLE: Qualitatively, how does this Pupper policy compare to the previous one? You need to upload a video of this policy in your submission and call it step2_sim.mp4.
DELIVERABLE: Download a copy of your notebook, rename it to step2.ipynb, and include it in your submission.
Tune the config to make Pupper smoothly follow velocities with a natural gait. Feel free to use any rewards you like (to reduce your search space, don’t try tuning other parameters yet!). Increase the training_config.ppo.num_timesteps to at least 300 million. Rerun the entire notebook, and train Pupper to walk in simulation.
DELIVERABLE: What terms are included in your reward functions? What coefficients did you use? How did you come up with these terms, and what was their desired effect? Do you think this policy will perform well on the physical robot?
DELIVERABLE: Visualize Pupper’s progress during training. How does Pupper look in the first 20 million env steps? How does it look after 200 million env steps?
DELIVERABLE: Record a video of Pupper walking in simulation. You need to upload a video of this policy in your submission and call it step3_sim.mp4.
DELIVERABLE: Download a copy of your final notebook after you are happy with the final policy, rename it to step3.ipynb, and include it in your submission.
Now that you have all the policies trained and verified in simulation, in this section, you will try them on the robot and compare them to the default walking policy. Based on what you see on the robot, you are free to return to the previous section to improve your reward functions. YOU MUST DO THIS SECTION OF WORK ON THE ROBOT.
Step 1: On your Pupper, open a terminal, clone the course repository, and navigate to the Lab5 directory.
Step 2: Follow the instructions here to connect your remote controller to the Pupper via Bluetooth so you can send velocity commands. Note that because there are multiple groups in the class, pair the remotes each group at a time to avoid confusion and pairing with other teams' remotes! You will be able to control the Pupper and switch between policies using the remote controller, as shown in the image below:
Step 3: First, configure your WandB access key so Popper can download the policies you trained from your account. Note that each group member should train the policy on their own once, and here, provide the API key for the person who trained the best policy! You can configure the key by typing the following command in the Popper's terminal (SSH into the robot):
export WANDB_API_KEY=Paste_Your_API_Key
Step 4: Configure your Pupper to work with the neural controller settings for this lab so it can download your trained policies. Navigate to the lab directory and rebuild the neural controller:
python3 rebuild_neural_controller.py
Step 5: Start the neural controller policy and state machines. This state machine will read the joystick and run different policies based on the keys you press, then command it to move in different directions using the analog stick:
python deploy.py
The script will prompt you to enter the policy run number you want to deploy. You can find the run number in your Wandb dashboard from your Colab training session (in the format of “random adjective-random noun-run number”). Enter the run number, or press Enter to keep the current policy.
Before deploying your own policy, first check the default walking policy. Press the “x” button on the remote controller. You should see the Pupper begin to walk using the default policy (it will stand up and wait for you to command it to move). Use the left joystick to control the Pupper’s walking direction and speed, and the right joystick to control its turning.
DELIVERABLE: Upload a video of the default policy working on the robot, naming it real_default.mp4.
Step 7: When you're ready, press the square button on your remote controller to switch from the default policy to your newly trained policy and take a video of all three policies you trained.
DELIVERABLE: In what ways is this policy different on the physical robot (compared to simulation)? This difference is called the “sim2real gap.”
DELIVERABLE: Record a video of the final policy you trained (step 3), name it step3_real.mp4, and upload it along with your submission.
One way to make the policy much more robust and make it walk much more naturally is to train it on a randomized rough terrain. You can also consider randomizing other parameters such as Pupper mass or PID gains. This way, the policy is forced to learn a more robust and generalizable locomotion skill because it won't be able to hone itself to the ideally flat surface of the simulator or its specific body parameters during training.
Edit the environment config to accurately reflect all the situations Pupper might encounter in the real world. Try several magnitudes of the domain randomization terms to see what works. Iterate many times, tuning domain randomization and rewards to find the best possible policy! An agile policy should be fast, efficient, stable, and robust to disturbances. Train the best policy you can and bring it to the office hour to try on the real robot! The team with the best policy will get a bonus!
DELIVERABLE: Comment on what might happen if you add too much domain randomization.
DELIVERABLE: Describe your approach to training an agile Pupper policy. What parameters were key? Did you use a heightfield? Why/why not?