All work

Reinforcement learning

Bipedal walking agents

I trained PPO, TD3 and SAC agents to walk in BipedalWalker-v3 and compared how each one learned.

  • Python
  • PyTorch

Techniques: OpenAI Gym, Box2D, PPO, TD3, SAC, Behaviour cloning, Reward shaping

View the code
hover the chart
-100010020030001k2k3k4k5k300 = solved

The problem

BipedalWalker-v3 asks a two-legged robot to learn balance and gait from nothing. Different algorithms get there at very different speeds, and some never get there cleanly.

How I built it

  1. 01Trained PPO, TD3 and SAC on the same environment and tracked the running average reward for each.
  2. 02Used behaviour cloning for a head start and reward shaping to push agents towards stable walking.
  3. 03Compared where each algorithm stalled or collapsed, and tuned the reward structure from that.

Result

SAC crossed a running average reward of 300 in about 330 episodes. TD3 got there in about 1,240, after collapsing once on the way. PPO peaked near 286 after more than 5,000.

Next projectDocQAOpen