Sam Austin AI

Introduction to Gymnasium: OpenAI's RL Environment Toolkit (2026)

September 4, 2026 14 min read Sam Austin
Contents

Quick correction before we even start: Gymnasium isn't actually OpenAI's toolkit anymore, and getting that straight upfront will save you a genuinely confusing afternoon later. OpenAI built the original Gym library, then stopped maintaining it back in 2020. The Farama Foundation picked it up, forked it, and Gymnasium is what came out the other side — same spirit, different maintainers, actively developed instead of frozen in time.

I bring this up first because I've watched people follow outdated tutorials that still reference the old gym package, hit confusing errors, and assume they did something wrong. You didn't. The tutorial is just old. Every RL project I've touched in the last couple of years uses Gymnasium, and this guide is going to make sure you start with the right foundation from day one.

By the end of this, you'll understand what Gymnasium actually is, how its core API works, and why it's become the standard interface practically every RL library builds around. IMO, understanding this one library well pays off across basically every RL project you'll ever touch :)

What Gymnasium Actually Is

Gymnasium is a standardized API for reinforcement learning environments — not an algorithm, not a training library, just a consistent interface for the "world" your agent interacts with. Think of it as the universal adapter that lets any RL algorithm plug into any environment without needing custom glue code for each combination.

It provides a standard set of methods — reset(), step(), render() — that every environment implements the same way, regardless of whether you're balancing a pole or landing a lunar module. It ships with a library of built-in reference environments, spanning classic control, Box2D physics, toy text problems, MuJoCo robotics simulations, and Atari games. It's become genuinely widely adopted — reportedly over 18 million downloads since its initial release through mid-2025, with hundreds of contributors actively maintaining and extending it.

Ever wondered why so many different RL tutorials — CartPole, Snake, Atari — all use nearly identical setup code? This standardization is exactly why. Learn Gymnasium's API once, and every environment you encounter afterward follows the same basic pattern.

Introduction to Gymnasium RL Environment Toolkit

Figure 1: Gymnasium provides the standardized interface that powers CartPole, LunarLander, Atari, and hundreds of other RL environments

Why the Fork Happened (And Why It's a Good Thing)

Here's the honest history: OpenAI staff maintained the original Gym library for years, but stopped updating it around October 2022. No bug fixes, no new features, just a slowly aging codebase everyone still depended on.

The Farama Foundation stepped in, created Gymnasium as a maintained fork, and it's now the actual place where development happens. Gymnasium is designed as a drop-in replacement — if you know Gym's API, Gymnasium will feel immediately familiar, just with quality-of-life improvements and features like vectorized environments layered on top.

FYI, if you're reading a tutorial that imports import gym instead of import gymnasium as gym, that's referencing the abandoned original. It might still technically work, but you're building on something nobody's fixing bugs in anymore.

If you're just getting started with RL, our beginner's guide to RL for games and robotics explains the core concepts before diving into environment APIs like this.

Installing Gymnasium

Gymnasium is modular by design — you install just the environment families you actually need, keeping your dependency footprint small.

pip install gymnasium

That's the core package alone. For specific environment families, install the relevant extra:

pip install "gymnasium[classic-control]"  # CartPole, MountainCar, etc.
pip install "gymnasium[box2d]"            # LunarLander, BipedalWalker
pip install "gymnasium[atari,accept-rom-license]"  # Atari games
pip install "gymnasium[mujoco]"           # Robotics simulation

Don't install every extra just in case. Each one pulls in its own dependencies, and Atari or MuJoCo installs are noticeably heavier than classic control. Install what your current project actually needs.

The Core API: Four Methods You'll Use Constantly

Everything in Gymnasium boils down to a small set of methods, consistent across every single environment.

Creating and Resetting an Environment

import gymnasium as gym

env = gym.make("CartPole-v1", render_mode="human")
observation, info = env.reset(seed=42)

That reset() call starts a fresh episode and returns the initial observation. Notice it returns two values — the observation itself, and an info dictionary carrying auxiliary diagnostic data that varies by environment. Seeding it once at the start, rather than repeatedly mid-training, is what actually gives you reproducible results.

Stepping Through the Environment

action = env.action_space.sample()
observation, reward, terminated, truncated, info = env.step(action)

This is where the actual interaction happens — you pick an action, the environment responds with a new observation and a reward. Notice there are five return values, not four. That's a genuine change from the original Gym API worth calling out specifically.

Where Did done Go?

If you've seen older RL code checking a single done flag, that's outdated Gym syntax. Gymnasium splits that single flag into two more meaningful signals:

terminated — the episode ended because of the actual task logic (the pole fell, the agent reached a goal). truncated — the episode ended because of an external limit (a max step count was hit), not because anything meaningful happened in the environment itself.

done = terminated or truncated
if done:
    observation, info = env.reset()

This distinction genuinely matters for training correctness — an episode ending because it hit a time limit shouldn't necessarily be treated identically to one ending because the agent actually failed the task. Algorithms that bootstrap value estimates need to know the difference.

For a deeper dive into how this API is used in practice, our CartPole tutorial walks through building a complete DQN agent using these exact methods.

Understanding Observation and Action Spaces

Every environment describes what it expects and returns through spaces — and reading these correctly is genuinely one of the first skills worth building.

print(env.observation_space)  # Box(4,) for CartPole
print(env.action_space)       # Discrete(2) for CartPole

Discrete spaces represent a fixed set of distinct choices — CartPole's two actions (push left, push right), or Snake's four directions. Box spaces represent continuous ranges — a robot arm's joint angles, or CartPole's four continuous state values (position, velocity, angle, angular velocity). More exotic spaces exist too — MultiDiscrete, Dict, Tuple — for environments with more structured inputs and outputs.

Get comfortable reading these shapes early. Every neural network you build afterward needs to match these dimensions exactly, and mismatched shapes are one of the most common early debugging headaches in RL.

Our evaluating reinforcement learning algorithms guide covers how these space definitions affect algorithm choice and evaluation methodology.

Wrappers: Modifying Environments Without Touching Their Code

Here's a genuinely elegant part of Gymnasium's design: wrappers let you modify an environment's behavior without editing the underlying environment code at all.

from gymnasium.wrappers import RecordVideo, TimeLimit

env = gym.make("CartPole-v1", render_mode="rgb_array")
env = RecordVideo(env, video_folder="videos")

Most environments created through gymnasium.make() are automatically wrapped with a few defaults — TimeLimit, OrderEnforcing, and PassiveEnvChecker — handling episode length limits and catching common usage mistakes before they become confusing bugs. You can chain additional wrappers on top for things like frame stacking, reward clipping, or observation normalization, all without modifying a single line of the base environment.

This is exactly how Atari preprocessing pipelines get built — frame stacking, grayscale conversion, and frame skipping are typically implemented as chained wrappers rather than baked directly into training code. Our Atari tutorial shows this in practice.

The Environment Families Worth Knowing

Gymnasium organizes its built-in environments into a few broad categories, each suited to different learning goals.

Classic Control — CartPole, MountainCar, Pendulum. Small state spaces, fast training, ideal for learning fundamentals. Box2D — LunarLander, BipedalWalker. Slightly more complex physics, still fast to train, a natural step up from classic control. Toy Text — FrozenLake, Taxi. Discrete grid-based problems, useful for understanding tabular methods before moving to deep RL. MuJoCo — HalfCheetah, Humanoid, Ant. Physics-based robotics simulations, generally significantly harder to solve than other environment families. Atari — the full Atari 2600 library, accessed through the Arcade Learning Environment, requiring visual processing and convolutional networks.

Progressing through these roughly in order — classic control, then Box2D, then Atari or MuJoCo — genuinely mirrors the natural difficulty curve most RL learners benefit from following.

If robotics simulation interests you, our robotics simulation hardware guide covers the physical kits you'd need to bridge from MuJoCo simulations to real hardware.

Vectorized Environments: Training Faster by Running Many at Once

One of Gymnasium's actual improvements over the original Gym is first-class support for vectorized environments — running many copies of an environment in parallel to speed up training.

envs = gym.make_vec("CartPole-v1", num_envs=4)
observations, infos = envs.reset()

Instead of collecting one experience at a time, you're now collecting four (or more) simultaneously, which meaningfully speeds up training for algorithms that can take advantage of batched experience collection. This becomes genuinely important once you move past toy problems into anything where sample efficiency actually matters.

Common Mistakes Beginners Make

I've either made these myself or watched them trip up others repeatedly.

Importing gym instead of gymnasium. They're not interchangeable long-term — Gym is unmaintained, and subtle API differences will eventually cause confusing bugs. Checking a single done flag. That pattern is outdated; use terminated or truncated, and understand why the distinction matters for training correctness. Forgetting render_mode="rgb_array" before wrapping with RecordVideo. Without it, video recording silently fails to produce any output. Reseeding the environment on every reset during training. Seed once at the start for reproducibility; reseeding repeatedly defeats the purpose and can introduce subtle bias. Installing every environment extra "just in case." This bloats your environment unnecessarily — install only what your current project actually touches.

What Gymnasium Doesn't Do (And Why That's Fine)

Worth being clear about scope: Gymnasium is purely an environment API — it doesn't include any training algorithms itself. You still need a separate library like Stable-Baselines3, or your own hand-rolled DQN, to actually train an agent.

This separation is intentional and genuinely useful — it means Gymnasium environments work across an enormous ecosystem of compatible training libraries, rather than locking you into one specific algorithm implementation. Think of it as the standardized socket, not the appliance plugged into it.

If you want to explore which training libraries pair best with Gymnasium, our best RL frameworks guide covers Stable-Baselines3, CleanRL, RLlib, and more.

Wrapping This Up

Gymnasium gives reinforcement learning a genuinely standardized foundation: consistent reset() and step() methods, clearly defined observation and action spaces, and a wrapper system for extending environments without touching their internals. Understanding this API well transfers directly across every RL project you'll build afterward.

Remember that Gymnasium is the actively maintained successor to OpenAI's original Gym, not OpenAI's current toolkit itself, and that the terminated/truncated split replaced the old single done flag for good reason. FYI, once this core API feels familiar, jumping between CartPole, LunarLander, and Atari environments feels like changing a few lines of setup code rather than learning a new system each time :)

Now go run env.action_space.sample() in a loop against a few different environments and just watch how differently each one behaves. That's genuinely the fastest way to get comfortable with the variety this library actually offers.

Share this article X Facebook LinkedIn Reddit WhatsApp

Related Articles