Contents
Every RL tutorial name-drops CartPole like you already know why it matters, then dives into code without ever showing you the "aha" moment of watching an agent go from falling over instantly to balancing like it's got nothing better to do. That moment is genuinely the best part of learning RL, and today you're building toward it from scratch.
I remember running my first successful CartPole training loop and just... watching it balance for way longer than necessary, like showing off a magic trick to nobody. It's a small problem, but it's the exact right size to actually understand every piece of what's happening instead of trusting a black box.
By the end of this tutorial, you'll have a working DQN agent that learns to balance a pole through trial and error, and you'll understand every line of what's making that happen. IMO, this is the single best first project in all of reinforcement learning :)
Figure 1: Building a DQN agent from scratch to solve the CartPole environment with Gymnasium and PyTorch
What CartPole Actually Is
Picture a cart that can slide left or right along a track, with a pole hinged on top that wants to fall over. Your job is training an agent to push the cart left or right at exactly the right moments to keep that pole balanced upright.
That's genuinely the entire problem. No graphics to render, no complex physics to reason about — just four numbers describing the situation and two possible actions to choose between.
State space: cart position, cart velocity, pole angle, and pole angular velocity — four numbers, nothing more. Action space: push left, or push right. That's the complete list of choices. Reward: +1 for every timestep the pole stays upright; the episode ends once it tips too far or the cart drifts off the track.
Ever wondered why this specific toy problem shows up in literally every RL course? It's small enough to train in minutes, but genuinely demonstrates the full learning loop without hiding any complexity behind scale.
If you want to understand the broader context of why CartPole matters, our beginner's guide to RL for games and robotics explains the full learning loop in plain language before diving into code like this.
Setting Up Your Environment
Let's get the dependencies installed. Gymnasium is the actively maintained successor to OpenAI's original Gym library, so make sure you're installing the current package, not the deprecated one.
python -m venv venv
source venv/bin/activate # Windows: venv\Scripts\activate
pip install gymnasium torch numpy
FYI, if you find an older tutorial importing import gym instead of import gymnasium as gym, that's referencing the unmaintained original package. The API is nearly identical, but Gymnasium is where active development and bug fixes actually happen now.
If you're curious about the hardware side of things — like what GPU you'd need for more complex environments beyond CartPole — our robotics simulation hardware guide covers exactly what to buy and when.
Exploring the Environment First
Before writing any learning logic, let's just look at what we're working with.
import gymnasium as gym
env = gym.make("CartPole-v1")
print("Observation space:", env.observation_space)
print("Action space:", env.action_space)
observation, info = env.reset()
print("Initial observation:", observation)
Run this, and you'll see the observation space is a 4-element array, and the action space is Discrete(2) — exactly the two choices we described above. Getting comfortable reading these shapes matters — every environment you touch later will describe itself this same way.
Watching Random Actions Fail
Let's confirm the baseline: an agent taking completely random actions should fail almost immediately.
observation, info = env.reset()
total_reward = 0
for _ in range(200):
action = env.action_space.sample()
observation, reward, terminated, truncated, info = env.step(action)
total_reward += reward
if terminated or truncated:
break
print(f"Random agent survived {total_reward} steps")
env.close()
You'll typically see somewhere between 10 and 30 steps before the pole tips over. That number is your baseline — everything we build from here should meaningfully beat it.
Building the DQN Agent
Now for the actual learning part. We're using Deep Q-Learning (DQN), which trains a neural network to estimate the value of taking each action in a given state.
import torch
import torch.nn as nn
import torch.optim as optim
import random
from collections import deque
class QNetwork(nn.Module):
def __init__(self, state_dim, action_dim):
super().__init__()
self.net = nn.Sequential(
nn.Linear(state_dim, 128),
nn.ReLU(),
nn.Linear(128, 128),
nn.ReLU(),
nn.Linear(128, action_dim)
)
def forward(self, x):
return self.net(x)
This network takes in the four state numbers and outputs two values — one estimating how good "push left" is right now, and one for "push right." The action with the higher predicted value is the one the agent picks.
For a deeper understanding of how DQN fits alongside other RL algorithms, our evaluating reinforcement learning algorithms guide breaks down the metrics and methods used to compare approaches like this.
Why We Need a Replay Buffer
Here's a detail that trips up a lot of beginners: you can't just train on each experience the instant it happens. Consecutive game states are highly correlated, and training on them in sequence badly destabilizes learning.
class ReplayBuffer:
def __init__(self, capacity=10000):
self.buffer = deque(maxlen=capacity)
def push(self, state, action, reward, next_state, done):
self.buffer.append((state, action, reward, next_state, done))
def sample(self, batch_size):
return random.sample(self.buffer, batch_size)
def __len__(self):
return len(self.buffer)
Storing experiences and sampling randomly later breaks that harmful correlation. The agent learns from a shuffled mix of past experiences instead of whatever just happened a moment ago — genuinely one of the more important tricks that makes DQN training actually stable.
Putting the Training Loop Together
Here's where everything connects: the agent explores, stores what happened, and periodically learns from a random batch of past experiences.
import numpy as np
env = gym.make("CartPole-v1")
state_dim = env.observation_space.shape[0]
action_dim = env.action_space.n
q_network = QNetwork(state_dim, action_dim)
optimizer = optim.Adam(q_network.parameters(), lr=0.001)
buffer = ReplayBuffer()
epsilon = 1.0
epsilon_decay = 0.995
epsilon_min = 0.01
gamma = 0.99
batch_size = 64
def select_action(state):
if random.random() < epsilon:
return env.action_space.sample()
with torch.no_grad():
state_tensor = torch.FloatTensor(state).unsqueeze(0)
q_values = q_network(state_tensor)
return q_values.argmax().item()
Understanding Epsilon-Greedy Exploration
That epsilon variable controls a genuinely important tradeoff. Early in training, the agent should explore randomly since it doesn't know anything useful yet. As training progresses, it should increasingly trust its learned values instead of guessing randomly.
epsilon starts at 1.0 — pure random action, since the network hasn't learned anything worth trusting yet.
epsilon decays gradually toward epsilon_min (0.01) as training progresses.
Skip this decay, and your agent either never explores enough to discover good strategies, or never stops acting randomly even once it's learned something useful.
The Learning Step
This is the actual "learning" in reinforcement learning — updating the network based on a batch of past experiences.
def train_step():
if len(buffer) < batch_size:
return
batch = buffer.sample(batch_size)
states, actions, rewards, next_states, dones = zip(*batch)
states = torch.FloatTensor(np.array(states))
actions = torch.LongTensor(actions)
rewards = torch.FloatTensor(rewards)
next_states = torch.FloatTensor(np.array(next_states))
dones = torch.FloatTensor(dones)
current_q = q_network(states).gather(1, actions.unsqueeze(1)).squeeze()
with torch.no_grad():
next_q = q_network(next_states).max(1)[0]
target_q = rewards + gamma * next_q * (1 - dones)
loss = nn.MSELoss()(current_q, target_q)
optimizer.zero_grad()
loss.backward()
optimizer.step()
That target_q calculation is the mathematical heart of Q-learning: the value of an action should equal the immediate reward plus the discounted value of the best action available next. The (1 - dones) term zeroes out that future value once an episode ends — there's no "next state" to consider once the pole's already fallen.
Running the Full Training Loop
With everything defined, let's actually train the agent across multiple episodes.
num_episodes = 300
for episode in range(num_episodes):
state, info = env.reset()
episode_reward = 0
for step in range(500):
action = select_action(state)
next_state, reward, terminated, truncated, info = env.step(action)
done = terminated or truncated
buffer.push(state, action, reward, next_state, done)
train_step()
state = next_state
episode_reward += reward
if done:
break
epsilon = max(epsilon_min, epsilon * epsilon_decay)
if episode % 20 == 0:
print(f"Episode {episode}, Reward: {episode_reward}, Epsilon: {epsilon:.3f}")
env.close()
Run this, and you'll watch something genuinely satisfying happen: early episodes score in the 10-30 range, matching our random baseline, then gradually climb toward 200+ as training progresses. CartPole-v1 caps episodes at 500 steps, so consistently hitting that ceiling means your agent has essentially solved the problem.
Common Mistakes When Building This Yourself
I've hit most of these while learning DQN myself, so treat this as a shortcut past my own debugging sessions.
Forgetting to decay epsilon. Without decay, the agent either never stops acting randomly or never explores enough to find good strategies in the first place.
Training before the buffer has enough samples. That if len(buffer) < batch_size: return check exists specifically to prevent training on a nearly-empty buffer, which produces garbage gradients.
Using a learning rate that's too high. DQN training can genuinely diverge — reward suddenly crashing back to near-zero after looking good — and an overly aggressive learning rate is the usual culprit.
Expecting smooth, monotonic improvement. RL training is noisy by nature; expect the reward curve to bounce around significantly even while the overall trend improves.
Visualizing Your Trained Agent
Once training's done, it's worth actually watching your agent perform instead of just trusting the numbers.
env = gym.make("CartPole-v1", render_mode="human")
state, info = env.reset()
epsilon = 0 # no more random exploration
for _ in range(500):
action = select_action(state)
state, reward, terminated, truncated, info = env.step(action)
if terminated or truncated:
break
env.close()
Setting epsilon = 0 here matters — you want to see the agent's actual learned policy, not random exploration mixed in. Watching a trained agent balance the pole smoothly after seeing it fail instantly at the start is genuinely one of the more satisfying moments in this whole field.
Where to Go From Here
CartPole solved, a few natural next steps make sense before jumping into anything dramatically harder.
Try LunarLander-v2 next — a larger state space and more nuanced reward shaping, while still staying in Gymnasium's built-in environments. Experiment with PPO using Stable-Baselines3 instead of hand-rolling DQN — you'll see how a more modern algorithm handles the same problem with less manual tuning. Tweak the reward structure yourself — try modifying CartPole's reward to penalize large cart movements, and watch how that changes the learned behavior.
If you want to go deeper on algorithms after this, our best courses for RL in robotics and game AI covers Hugging Face's Deep RL Course, NVIDIA's Physical AI path, and Stanford CS234 — each suited to different learning goals.
Wrapping This Up
Training an agent on CartPole boils down to a repeating loop: act, observe the reward, store the experience, periodically learn from a random batch of past experiences, and gradually shift from random exploration to trusting the learned policy. Every more advanced RL algorithm builds on this exact same core structure.
Remember that epsilon decay and the replay buffer aren't optional extras — they're the two tricks that make DQN training actually stable instead of chaotic. FYI, once this clicks, moving to more complex Gymnasium environments feels like changing the input size, not learning a new discipline from scratch :)
Now go tweak the hyperparameters and watch what breaks. Bumping the learning rate too high or shrinking the replay buffer too small will teach you more about why these pieces matter than reading about them ever could.