Reinforcement learning
Bipedal walking agents
I trained PPO, TD3 and SAC agents to walk in BipedalWalker-v3 and compared how each one learned.
Python
PyTorch
Techniques: OpenAI Gym, Box2D, PPO, TD3, SAC, Behaviour cloning, Reward shaping
View the codehover the chart
The problem
BipedalWalker-v3 asks a two-legged robot to learn balance and gait from nothing. Different algorithms get there at very different speeds, and some never get there cleanly.
How I built it
- 01Trained PPO, TD3 and SAC on the same environment and tracked the running average reward for each.
- 02Used behaviour cloning for a head start and reward shaping to push agents towards stable walking.
- 03Compared where each algorithm stalled or collapsed, and tuned the reward structure from that.
Result
SAC crossed a running average reward of 300 in about 330 episodes. TD3 got there in about 1,240, after collapsing once on the way. PPO peaked near 286 after more than 5,000.