Sam Austin AI

Stable-Baselines3 Tutorial: Train RL Agents in Minutes (2026)

September 5, 2026 16 min read Sam Austin
Contents

You've now hand-built DQN for CartPole, DQN for Snake, and touched PPO for BipedalWalker and MuJoCo. Notice a pattern? Replay buffers, epsilon decay, target networks — you've been rewriting variations of the same infrastructure over and over. Stable-Baselines3 exists specifically so you stop doing that.

I held off introducing this earlier in this whole series on purpose — building DQN by hand at least once genuinely matters for understanding what's happening under the hood. But now that you get it, there's no medal for continuing to hand-roll replay buffers forever. This is the library nearly every serious RL project actually uses once the learning phase is done.

By the end of this tutorial, you'll be training PPO, SAC, and DQN agents in a handful of lines each, and you'll understand exactly which algorithm to reach for and when. IMO, this is the tool that turns "I understand RL conceptually" into "I can actually ship RL projects quickly" :)

What Stable-Baselines3 Actually Is

Stable-Baselines3 (SB3) is a PyTorch-based library providing reliable, well-tested implementations of the major RL algorithms — PPO, DQN, SAC, TD3, A2C, and more — behind one consistent API. It's the direct successor to the original TensorFlow-based Stable Baselines, rewritten from scratch in PyTorch.

Training an agent genuinely takes just a few lines of code, letting you focus on your environment and problem instead of implementation details. The library deliberately focuses on model-free, single-agent algorithms, leaning on companion projects for things like imitation learning or offline RL rather than trying to do everything itself. It prioritizes stable, correct implementations over chasing every new paper — the whole point is giving you a trustworthy baseline, not the bleeding edge of research.

Ever wondered why so many of the tutorials in this whole series eventually reached for from stable_baselines3 import PPO? This is exactly why. Once you understand the concepts, SB3 removes the repetitive plumbing without hiding what's actually happening.

Our Gymnasium tutorial covers the environment API that SB3 plugs into — understanding reset(), step(), and action/observation spaces first makes SB3 feel like a natural extension rather than a new system.

Stable-Baselines3 Tutorial Train RL Agents in Minutes PyTorch

Figure 1: Stable-Baselines3 removes the infrastructure overhead — replay buffers, target networks, training loops — so you focus on your environment and problem

Installation

pip install stable-baselines3[extra]

That [extra] tag pulls in genuinely useful optional dependencies — Tensorboard for training visualization, OpenCV and ale-py for Atari, plus pandas and matplotlib for analyzing results afterward. SB3 requires Python 3.10 or newer, so double-check your environment if you're working from an older setup.

Your First Agent in Four Lines

Let's prove the "minutes" part of this tutorial's title isn't an exaggeration.

import gymnasium as gym
from stable_baselines3 import PPO

env = gym.make("CartPole-v1")
model = PPO("MlpPolicy", env, verbose=1)
model.learn(total_timesteps=25_000)

That's genuinely the entire training loop — no replay buffer to write, no epsilon decay schedule to tune, no manual gradient step. Compare this to the CartPole DQN implementation from earlier in this series, and you'll immediately see why SB3 exists: everything you wrote by hand there is happening automatically here, tested and debugged by people who do this full-time.

Using Your Trained Model

observation, info = env.reset()
for _ in range(1000):
    action, _states = model.predict(observation, deterministic=True)
    observation, reward, terminated, truncated, info = env.step(action)
    if terminated or truncated:
        observation, info = env.reset()

Notice model.predict() replaces the manual "run the state through the network, take argmax" logic you wrote by hand for CartPole's DQN. Same underlying concept, dramatically less code to maintain.

Our CartPole tutorial shows the hand-rolled DQN version — comparing that implementation to these four lines makes the value of SB3 immediately obvious.

Choosing the Right Algorithm

This is genuinely the part beginners get wrong most often — picking an algorithm without matching it to the actual problem shape. Here's SB3's own guidance, distilled.

Discrete action spaces (CartPole, Snake, Atari) — DQN with extensions is a solid, sample-efficient choice, though it trains slower in wall-clock time. PPO or A2C are reasonable alternatives worth trying first for simplicity. Continuous action spaces (BipedalWalker, MuJoCo locomotion, robot arms) — PPO is the dependable, forgiving default. Current state-of-the-art options for continuous control specifically include SAC, TD3, CrossQ, and TQC. Need extreme sample efficiency — if every environment interaction is expensive (think real robotics, not simulation), a DroQ configuration is specifically recommended for wringing maximum learning out of minimal data.

My honest take: start with PPO regardless of your action space unless you have a specific reason not to. It handles both discrete and continuous actions reasonably well, it's forgiving of imperfect hyperparameters, and it's genuinely the least likely algorithm to leave you confused about why training isn't working.

Our evaluating reinforcement learning algorithms guide covers how to assess whether your chosen algorithm is actually performing well — useful for deciding when to switch from PPO to something more specialized.

Training on Continuous Control: SAC Example

Let's train something with continuous actions, mirroring what you'd need for BipedalWalker or a MuJoCo task.

import gymnasium as gym
from stable_baselines3 import SAC

env = gym.make("Pendulum-v1")
model = SAC("MlpPolicy", env, verbose=1, learning_rate=3e-4)
model.learn(total_timesteps=50_000)
model.save("sac_pendulum")

SAC is currently one of the recommended state-of-the-art choices for continuous control in SB3's own guidance, alongside TD3 and TQC. Swapping algorithms in your code is genuinely this simple — same MlpPolicy, same .learn() call, just a different imported class.

For a deeper comparison of when SAC outperforms PPO, our MuJoCo tutorial covers continuous control fundamentals that apply directly here.

Saving, Loading, and the RL Zoo

Rebuilding a trained agent from scratch every session wastes real time. SB3 makes persistence trivial.

model.save("ppo_cartpole")

del model
model = PPO.load("ppo_cartpole", env=env)

Don't skip this in any real project. Beyond individual save/load, SB3's companion project — RL Baselines3 Zoo — provides over 100 pre-trained agents plus scripts for training, hyperparameter tuning, and evaluation, so you're not always starting from a blank slate.

Vectorized Environments: Training Faster

Just like you saw with Gymnasium's make_vec(), SB3 supports running multiple environment copies in parallel to speed up data collection.

from stable_baselines3.common.env_util import make_vec_env

vec_env = make_vec_env("CartPole-v1", n_envs=4)
model = PPO("MlpPolicy", vec_env, verbose=1)
model.learn(total_timesteps=100_000)

This is essentially free speed for on-policy algorithms like PPO and A2C, which can genuinely take advantage of batched experience collection across multiple parallel environments.

Monitoring Training with Callbacks

Watching a bare verbose=1 scroll of numbers gets old fast. SB3's callback system lets you hook into training for logging, early stopping, or custom evaluation.

from stable_baselines3.common.callbacks import EvalCallback

eval_callback = EvalCallback(
    env,
    best_model_save_path="./best_model/",
    eval_freq=5000,
    deterministic=True,
)

model.learn(total_timesteps=50_000, callback=eval_callback)

That EvalCallback periodically pauses training to evaluate the current policy, automatically saving the best-performing checkpoint rather than just whatever the model looked like at the very end of training. This matters more than it sounds — RL training is noisy, and the final checkpoint isn't always the best one.

TensorBoard Integration

model = PPO("MlpPolicy", env, verbose=1, tensorboard_log="./ppo_tensorboard/")
model.learn(total_timesteps=50_000)

Run tensorboard --logdir ./ppo_tensorboard/ afterward, and you get genuinely useful visualizations of reward curves, loss values, and training statistics over time — far easier to interpret than scrolling print statements.

Using Custom Environments

Everything you've built throughout this series — Snake, a custom grasping task — works with SB3 as long as it follows the Gymnasium API you already know.

from stable_baselines3.common.env_checker import check_env

env = YourCustomEnv()
check_env(env)  # verifies Gymnasium compliance before training

model = PPO("MlpPolicy", env, verbose=1)
model.learn(total_timesteps=100_000)

That check_env() call is genuinely worth running before training anything custom. It catches common issues — wrong observation shapes, missing reset/step signatures, malformed action spaces — before you waste an hour training against a broken environment.

Handling Dictionary Observations (Like Fetch's Grasping Task)

Remember the robot arm grasping tutorial's dictionary-based observation space? SB3 handles that natively too.

from stable_baselines3 import SAC
from stable_baselines3.her import HerReplayBuffer

model = SAC(
    "MultiInputPolicy",
    env,
    replay_buffer_class=HerReplayBuffer,
    verbose=1,
)

Notice MultiInputPolicy instead of MlpPolicy — this is exactly the same distinction from the Fetch grasping tutorial, and it's worth recognizing that pattern now shows up as a general SB3 convention, not a one-off exception for that specific task.

Our robot arm grasping tutorial covers HER in detail — understanding that technique makes the MultiInputPolicy choice here feel natural rather than arbitrary.

SB3 Contrib: Where the Newer Algorithms Live

The core SB3 library deliberately stays conservative about adding new algorithms, prioritizing stability. For more recent techniques, there's a companion "contrib" package.

pip install sb3-contrib
from sb3_contrib import QRDQN, TQC

model = QRDQN("MlpPolicy", "CartPole-v1", verbose=1)

QR-DQN (Quantile Regression DQN) and TQC (Truncated Quantile Critics) live here specifically because they're genuinely useful but newer or more experimental than what the core library commits to maintaining long-term. If the base algorithms don't quite fit your problem, checking contrib before hand-rolling something custom is usually the better move.

Common Mistakes Beginners Make

I've watched these trip up plenty of people moving from hand-rolled implementations into SB3 specifically.

Using DQN for continuous action spaces. DQN fundamentally expects discrete actions — for BipedalWalker or MuJoCo tasks, that's an immediate, confusing error, not a training problem. Forgetting deterministic=True during evaluation. Without it, you're seeing exploration noise mixed into your "final" policy's behavior, not its actual best performance. Skipping check_env() on custom environments. A malformed observation or action space produces genuinely confusing training failures that are dramatically easier to catch upfront. Not using vectorized environments when training on-policy algorithms. PPO and A2C genuinely benefit from parallel environments — running single-environment training when you didn't need to just wastes wall-clock time. Assuming SB3 handles multi-agent RL. It deliberately doesn't — for PettingZoo-style multi-agent problems, you'd reach for RLlib or a dedicated MARL library instead.

Our multi-agent RL guide covers the PettingZoo-based alternatives for multi-agent training that SB3 doesn't cover.

When Should You Go Back to Hand-Rolling?

Genuinely rare, but worth naming: if you're doing actual RL research and need a novel algorithm SB3 doesn't implement, or you need to modify the core training loop in ways the library's API doesn't expose, hand-rolling is still the right call. For literally everything else — prototyping, benchmarking, applied robotics or game AI projects — SB3 is the correct default, not a shortcut you should feel guilty about taking.

Wrapping This Up

Stable-Baselines3 takes every piece of infrastructure you've built by hand across this series — replay buffers, target networks, epsilon-greedy exploration, continuous-action policy gradients — and packages it behind one consistent, well-tested API. Understanding those fundamentals first is exactly what makes SB3 genuinely useful now, rather than a black box you're trusting blindly.

Remember to match your algorithm to your action space (PPO as a safe general default, SAC/TD3/TQC for continuous control specifically), always run check_env() on custom environments, and lean on callbacks for evaluation rather than just trusting the final checkpoint. FYI, RL Baselines3 Zoo's 100+ pre-trained agents are genuinely worth browsing before you train anything from scratch — someone may have already solved a very similar problem to yours :)

Now go retrain CartPole, Snake, or BipedalWalker through SB3 instead of your hand-rolled version, and time how much faster you get to a working result. That comparison is genuinely the best way to appreciate what this library is actually saving you from.

Share this article X Facebook LinkedIn Reddit WhatsApp

Related Articles