Contents
Ant and HalfCheetah get you comfortable with continuous control, but there's something specifically satisfying about BipedalWalker that those quadrupeds don't quite deliver — watching a two-legged stick figure go from faceplanting instantly to genuinely striding across the terrain feels like the closest hobbyist RL gets to the "robot learns to walk" videos that got you into this field in the first place.
I gravitated toward this project specifically because it sits in a genuinely useful sweet spot: real continuous control, real balance challenges, but running on lightweight Box2D physics instead of MuJoCo's heavier simulation — meaning you get locomotion-learning satisfaction without the multi-hour training commitment that comes with more complex physics engines.
By the end of this tutorial, you'll understand what makes BipedalWalker genuinely tricky, how PPO handles its continuous action space, and you'll have a trained agent that can actually walk. IMO, this is one of the most rewarding "watch it get good" projects in accessible RL :)
What BipedalWalker Actually Is
BipedalWalker-v3 simulates a simple two-legged robot on procedurally generated terrain, and the goal is straightforward to state but genuinely hard to achieve: walk as far as possible without falling over, while minimizing wasted motor effort.
Action space: 4 continuous values, controlling torque at each of the walker's four joints (two hips, two knees). Observation space: 24 continuous values covering hull angle, angular velocity, horizontal/vertical speed, joint positions and speeds, leg contact with ground, and 10 lidar readings sensing upcoming terrain. Reward: positive for moving forward, a penalty for falling, and a small penalty for motor torque usage — encouraging efficient movement rather than flailing wildly forward.
Unlike Ant's four-legged stability, a bipedal robot has a fundamentally harder balance problem — fewer points of ground contact means less inherent stability, and the agent has to actively manage balance rather than getting some of it "for free" from having four legs planted on the ground.
If you're coming from our MuJoCo tutorials, our Ant and locomotion guide covers continuous control fundamentals that apply here — the same PPO algorithm, the same action-space handling, just applied to a harder balance problem.
Figure 1: Training a bipedal walker from random torques to an actual walking gait using PPO
Setting Up the Environment
BipedalWalker runs on Box2D, a lightweight 2D physics engine, which is genuinely part of what makes this project more approachable than MuJoCo-based locomotion.
pip install "gymnasium[box2d]" stable-baselines3 torch
Let's confirm the environment loads and take a look at what we're working with.
import gymnasium as gym
env = gym.make("BipedalWalker-v3", render_mode="human")
print("Observation space:", env.observation_space)
print("Action space:", env.action_space)
observation, info = env.reset(seed=42)
This gives us a physics-based 2D bipedal robot that needs to learn how to walk from nothing more than reward feedback — no hardcoded gait patterns, no motion capture data, just trial and error against physics.
Watching Random Actions Fail Spectacularly
Before training anything, let's establish the baseline — and with BipedalWalker, this baseline is genuinely comedic.
observation, info = env.reset()
total_reward = 0
for _ in range(500):
action = env.action_space.sample()
observation, reward, terminated, truncated, info = env.step(action)
total_reward += reward
if terminated or truncated:
break
print(f"Random agent total reward: {total_reward}")
env.close()
Expect a strongly negative reward here — random torques applied to four joints simultaneously produce something closer to a seizure than an attempt to walk, and the walker typically collapses within the first few steps. This is your "before" picture; everything from here is about closing that gap.
For a deeper understanding of how PPO compares to other algorithms, our evaluating reinforcement learning algorithms guide covers the metrics and methods used to assess continuous-control performance.
Why PPO Is the Right Algorithm Here
DQN, which worked fine for CartPole and Snake, fundamentally can't handle this environment — DQN expects a discrete set of actions, and BipedalWalker's joints need continuous torque values. This is exactly the kind of problem PPO was built for.
from stable_baselines3 import PPO
env = gym.make("BipedalWalker-v3")
model = PPO(
"MlpPolicy",
env,
verbose=1,
learning_rate=3e-4,
n_steps=2048,
batch_size=64,
n_epochs=10,
gamma=0.99,
)
model.learn(total_timesteps=1_000_000)
model.save("bipedal_walker_ppo")
That million-timestep figure isn't arbitrary padding — BipedalWalker genuinely needs substantial training to go from random flailing to competent walking. Expect real training time here, particularly on CPU; patience is a real part of this project, not just a throwaway line in the README.
Understanding What's Actually Happening During Training
Training progress on BipedalWalker follows a genuinely recognizable arc, and knowing what to expect at each stage helps you avoid panicking over what looks like a stalled training run.
Early training (first ~50k-100k steps): the agent falls almost immediately, every episode. Reward stays deeply negative. Middle training: the agent starts staying upright longer, often through strategies that look nothing like walking — bracing, shuffling, odd asymmetric gaits. Later training: something resembling an actual walking gait emerges, and reward climbs steadily as the agent covers more distance before failing or completing the course.
Watching the "so close, but so weird" middle phase is honestly one of the most entertaining parts of this whole project. The agent's early solutions to "don't fall over" are rarely what a human would call walking, and that's genuinely fine — it's still learning.
Watching Your Trained Agent
model = PPO.load("bipedal_walker_ppo")
env = gym.make("BipedalWalker-v3", render_mode="human")
observation, info = env.reset()
for _ in range(1600):
action, _ = model.predict(observation, deterministic=True)
observation, reward, terminated, truncated, info = env.step(action)
if terminated or truncated:
observation, info = env.reset()
env.close()
deterministic=True matters just like it did for the Fetch arm — you want to see the policy's actual learned behavior, not exploration noise layered on top.
Recording Your Training Progress
Watching live is fun, but recording video lets you actually compare an early checkpoint against a later one side by side.
from gymnasium.wrappers import RecordVideo
env = gym.make("BipedalWalker-v3", render_mode="rgb_array")
env = RecordVideo(env, "./videos")
This is genuinely worth doing. Side-by-side footage of a 50k-step checkpoint against a 1M-step checkpoint makes the improvement dramatically more visible than reward numbers alone ever will.
The Reward Shaping Story
Reward design in BipedalWalker is more subtle than it first appears, and it's worth understanding the built-in structure before you consider modifying it.
Forward progress is rewarded, encouraging the agent to actually move rather than just standing still indefinitely. A motor torque penalty discourages wasteful, high-energy flailing — without this, agents can find degenerate solutions that technically move forward through violent, inefficient motion. Falling incurs a penalty, giving the agent a clear signal to avoid collapsing, though not so severe that it discourages the exploration needed to learn balance in the first place.
Reward shaping significantly impacts how the agent learns — this is exactly the kind of environment where you can tell the reward function's fingerprints are all over the resulting gait style.
The Hardcore Variant: A Genuine Next Challenge
Once standard BipedalWalker is solved, BipedalWalkerHardcore exists specifically to punish overfit solutions.
env = gym.make("BipedalWalker-v3", hardcore=True, render_mode="human")
Hardcore mode adds stairs, stumps, and pitfalls to the terrain — obstacles that a policy trained only on flat, gently varying ground has never encountered. Many policies that walk beautifully on standard terrain fall apart completely here, which is a genuinely useful lesson about the difference between solving a specific environment and learning something that actually generalizes.
Our Gymnasium tutorial covers the wrapper system and environment families that BipedalWalker belongs to — understanding that foundation makes the Box2D environment setup here feel familiar.
Exploring Alternative Algorithms
PPO is the standard starting point, but BipedalWalker is popular enough as a benchmark that it's worth knowing what else gets used here.
SAC (Soft Actor-Critic) — an off-policy alternative, sometimes more sample-efficient than PPO, at the cost of more hyperparameter sensitivity. TD3 (Twin Delayed DDPG) — another actor-critic method commonly benchmarked against PPO on this exact environment. Curiosity-driven exploration (ICM) — an add-on technique that rewards the agent for encountering novel states, sometimes helping escape mediocre local gait patterns that vanilla PPO can get stuck in.
Start with PPO. It's the best-documented, most forgiving choice for a first attempt, and the alternatives above are genuinely worth exploring only once PPO's baseline behavior feels understood.
If you want to formalize your algorithm knowledge, our best courses for RL in robotics and game AI covers Hugging Face's Deep RL Course, NVIDIA's Physical AI path, and Stanford CS234 — each suited to different learning goals.
Common Mistakes People Make
I've hit a few of these myself, and seen the rest repeated across community implementations of this project.
Expecting fast results. BipedalWalker genuinely needs substantial training time — judging progress after only 50k steps leads to premature, discouraged conclusions. Jumping straight to Hardcore mode. Solve standard BipedalWalker first; Hardcore is specifically designed to break policies that haven't learned genuinely robust balance. Ignoring the torque penalty's role. If you modify the reward function and remove the energy penalty, expect bizarre, high-energy gaits that "solve" the task without looking anything like efficient walking. Judging progress purely from reward numbers. Actually watching rendered episodes (or recorded video) reveals qualitative gait improvements that raw reward curves can hide or obscure.
Wrapping This Up
BipedalWalker takes everything you've learned about continuous control from Ant and HalfCheetah and applies it to a genuinely harder balance problem — fewer ground contact points, real torque-efficiency tradeoffs, and a training arc that visibly progresses from comedic collapse to actual walking. PPO handles the continuous 4-dimensional action space naturally, the same way it did for MuJoCo locomotion tasks.
Remember that patience is a real part of this project — training runs genuinely take time, and the "weird but functional" middle phase of training is completely normal, not a sign something's broken. FYI, once standard BipedalWalker is solved, Hardcore mode is a genuinely honest test of whether your agent learned real balance or just memorized flat terrain :)
Now go let it train for a while and check back periodically — watching the gait evolve from "seizure" to "shuffle" to "actual walk" across checkpoints is genuinely one of the more satisfying progressions in accessible RL.