Contents
Figure 1: Your environment, your rules — the contract is four parts: action space, observation space, reset(), step()
Want to train an RL agent on your own game, simulator, or weird little idea instead of another ready-made benchmark? Building a custom Gymnasium environment gives you that freedom.
You define the rules. You decide what the agent sees, what actions it can take, what counts as success, and how badly it gets punished for doing something ridiculous. That sounds powerful because it is powerful — and because reward design can turn into a tiny chaos factory if you rush it.
Gymnasium provides the standard interface that lets reinforcement learning libraries interact with environments consistently. Your environment needs an action space, an observation space, a reset() method, and a step() method. Once you follow that contract, you can connect your environment to agents from libraries such as Stable-Baselines3, RLlib, CleanRL, or your own code.
What a Custom Gymnasium Environment Does
A Gymnasium environment acts as the bridge between your agent and your problem. The agent chooses an action, the environment applies that action, updates its internal state, calculates a reward, and returns new information.
The loop looks like this:
- The environment returns an observation.
- The agent selects an action.
- The environment processes the action.
- The environment returns the next observation, reward, episode status, and extra details.
- The agent learns from that transition.
Gymnasium calls one pass through this action-observation exchange a timestep. Its standard step() method returns:
observation, reward, terminated, truncated, info
The terminated flag signals a natural end to the task, such as reaching a goal or crashing. The truncated flag signals an artificial cutoff, usually a time limit. When either becomes True, reset the environment before the next episode.
Plan the Task Before Writing Code
Start with the environment design, not the Python file. A clean specification saves you from fixing confusing RL behavior later.
Ask yourself these questions:
- What does the agent need to accomplish?
- What information does it need to make a good decision?
- Which actions can it choose?
- What should reward progress?
- What should end an episode?
- What prevents the agent from exploiting your reward function?
For this tutorial, let's build a small GridWorld environment. The agent starts somewhere on a grid and must reach a target square. It can move up, right, down, or left.
Simple? Yes. Useless? Not at all. GridWorld forces you to define every key Gymnasium concept without hiding logic behind a giant simulator.
Define the Action Space
Your action space describes every legal action the agent can select. Our GridWorld agent gets four discrete movement actions:
| Action ID | Movement |
|---|---|
| 0 | Right |
| 1 | Up |
| 2 | Left |
| 3 | Down |
Gymnasium provides spaces.Discrete(n) for a fixed list of integer actions. In our case, spaces.Discrete(4) means the agent can choose values from 0 through 3.
from gymnasium import spaces
self.action_space = spaces.Discrete(4)
Use Discrete when the agent picks one action from a menu. Use Box for continuous actions, such as steering, acceleration, force, or joystick values. Gymnasium's action-space definitions tell an agent exactly what it may do, while observation spaces tell it what it may see.
Match the Space to the Problem
Choose an action space that matches the actual task:
spaces.Discrete(4)for four movement choices.spaces.MultiDiscrete([3, 3])for separate horizontal and vertical controls.spaces.MultiBinary(6)for six independent buttons.spaces.Box(low=-1, high=1, shape=(2,))for continuous two-dimensional controls.
Do not make actions more complicated than necessary. If your agent only needs to move in four directions, do not give it 500 action combinations because you enjoy debugging at midnight.
Define the Observation Space
The observation space tells the agent which information it receives after every step. For GridWorld, the agent needs to know its location and the target location.
We can represent that state as a dictionary:
{
"agent": np.array([x_agent, y_agent], dtype=np.int32),
"target": np.array([x_target, y_target], dtype=np.int32),
}
A Dict space makes this structure explicit:
self.observation_space = spaces.Dict(
{
"agent": spaces.Box(
low=0,
high=size - 1,
shape=(2,),
dtype=np.int32,
),
"target": spaces.Box(
low=0,
high=size - 1,
shape=(2,),
dtype=np.int32,
),
}
)
Gymnasium expects every observation your environment returns to fit inside this declared space. That requirement catches many bugs early, so treat it as a helpful guardrail rather than annoying paperwork.
Choose Observations Carefully
The agent can only learn from what it observes. If you hide critical information, it may face an impossible task.
For example, a driving agent needs enough information to infer its position, velocity, track direction, and nearby obstacles. A Flappy Bird agent may need its height, vertical velocity, the next pipe distance, and the pipe gap position.
Give the agent the information it needs, but avoid handing it the answer directly. A navigation agent should see distances and positions, not a precomputed "move right now" signal. Otherwise, your RL project becomes a very expensive if statement.
Create the Environment Class
A custom Gymnasium environment subclasses gymnasium.Env. At a minimum, you define the action space, observation space, reset(), and step() methods.
Here is a complete GridWorld example:
import gymnasium as gym
from gymnasium import spaces
import numpy as np
class GridWorldEnv(gym.Env):
metadata = {"render_modes": ["human", "ansi"], "render_fps": 4}
def __init__(self, size=5, render_mode=None):
super().__init__()
self.size = size
self.render_mode = render_mode
self.observation_space = spaces.Dict(
{
"agent": spaces.Box(
low=0,
high=size - 1,
shape=(2,),
dtype=np.int32,
),
"target": spaces.Box(
low=0,
high=size - 1,
shape=(2,),
dtype=np.int32,
),
}
)
self.action_space = spaces.Discrete(4)
self._action_to_direction = {
0: np.array([1, 0]), # right
1: np.array([0, 1]), # up
2: np.array([-1, 0]), # left
3: np.array([0, -1]), # down
}
self.agent_location = None
self.target_location = None
def _get_obs(self):
return {
"agent": self.agent_location.copy(),
"target": self.target_location.copy(),
}
def _get_info(self):
distance = np.abs(
self.agent_location - self.target_location
).sum()
return {"distance": int(distance)}
def reset(self, seed=None, options=None):
super().reset(seed=seed)
self.agent_location = self.np_random.integers(
0,
self.size,
size=2,
dtype=np.int32,
)
self.target_location = self.agent_location.copy()
while np.array_equal(
self.agent_location,
self.target_location,
):
self.target_location = self.np_random.integers(
0,
self.size,
size=2,
dtype=np.int32,
)
observation = self._get_obs()
info = self._get_info()
return observation, info
def step(self, action):
direction = self._action_to_direction[action]
self.agent_location = np.clip(
self.agent_location + direction,
0,
self.size - 1,
)
terminated = np.array_equal(
self.agent_location,
self.target_location,
)
reward = 1.0 if terminated else -0.01
truncated = False
observation = self._get_obs()
info = self._get_info()
return observation, reward, terminated, truncated, info
def render(self):
if self.render_mode == "ansi":
grid = np.full(
(self.size, self.size),
".",
dtype="<U1",
)
agent_x, agent_y = self.agent_location
target_x, target_y = self.target_location
grid[target_y, target_x] = "T"
grid[agent_y, agent_x] = "A"
return "\n".join(
" ".join(row) for row in grid
)
if self.render_mode == "human":
print(self.render())
def close(self):
pass
This environment uses random starting positions, keeps the agent inside grid boundaries, gives a reward for reaching the target, and ends the episode when the agent succeeds.
Understand reset()
Gymnasium calls reset() at the beginning of every episode. Your method must restore the environment to a valid starting state and return:
observation, info
Gymnasium recommends calling super().reset(seed=seed) inside your method. That call initializes the environment's random-number generator, which helps you reproduce experiments by using the same seed.
Our reset() method performs four jobs:
- It initializes random-number handling.
- It selects a random agent position.
- It selects a different random target position.
- It returns the initial observation and diagnostic information.
The info dictionary does not drive learning directly. It gives you optional debugging details, statistics, or environment-specific measurements. We include Manhattan distance between the agent and target.
Why Seeding Matters
Reproducibility matters in reinforcement learning because randomness affects almost everything: initial states, exploration, neural-network initialization, and action sampling.
Use a seed when you want to compare experiments fairly:
env = GridWorldEnv(size=5)
observation, info = env.reset(seed=42)
A seed will not make every part of your deep learning pipeline perfectly identical across all hardware, but it gives you a far more reliable starting point than pure randomness. The evaluation walkthrough shows why seeds matter when you compare algorithms.
Understand step()
The step(action) method contains the heart of your environment. It receives an action, changes the state, calculates reward, checks episode endings, and returns the next transition.
Every step() call must return five values:
observation, reward, terminated, truncated, info
That five-value API matters. Older Gym code often returned a single done boolean, but Gymnasium separates natural task endings from externally imposed cutoffs.
Terminated vs Truncated
Use terminated=True when the task itself ends:
- The agent reaches the target.
- The player loses all health.
- A robot falls over.
- The bird crashes into a pipe.
- The game reaches a defined win condition.
Use truncated=True when an external limit ends the episode:
- The agent reaches the maximum number of steps.
- A simulator timeout occurs.
- Your experiment enforces a safety or time limit.
This distinction helps RL algorithms calculate targets correctly. A natural terminal state should stop future-return bootstrapping, while a time-limit cutoff often should not mean "the world truly ended." Tiny flag, large consequences.
Design Rewards That Encourage the Right Behavior
Reward design makes or breaks a custom Gymnasium environment. Your agent optimizes the reward you provide, not the goal you describe in your README.
Our GridWorld reward system uses:
- +1.0 when the agent reaches the target.
- -0.01 for every other step.
The small step penalty encourages short paths. The agent has a reason to reach the target quickly instead of wandering around the grid forever.
Avoid Reward Hacking
Reward hacking happens when the agent finds a way to maximize reward without solving the intended task.
Imagine you reward a driving agent for speed but forget to reward staying on the road. The agent may accelerate directly into a wall because it technically collected speed reward for a few glorious milliseconds.
Before training, ask:
- Can the agent repeat a reward-generating action forever?
- Does the reward favor actual progress?
- Does the agent need a penalty for invalid or unsafe behavior?
- Does the goal reward dominate smaller shaping rewards?
Keep rewards simple at first. Add shaping only when you can explain exactly how it supports the real objective. The reward function design guide goes deeper on the usual traps.
Add a Time Limit
Our example sets truncated = False, so a lost agent could theoretically wander forever. You can solve that problem by tracking steps yourself or registering the environment with max_episode_steps.
Gymnasium's registration system supports max_episode_steps, which applies a time-limit wrapper and truncates episodes that exceed the limit.
Add a counter in __init__ and reset() if you want full control:
self.max_steps = 100
self.steps = 0
Then update it inside step():
self.steps += 1
truncated = self.steps >= self.max_steps
For most projects, a time limit protects training from endless episodes and improves throughput. An agent that cannot reach the target within a sensible number of moves needs a reset, not an infinite travel allowance.
Register Your Custom Environment
You can instantiate your class directly:
env = GridWorldEnv(size=5)
But registration lets you create it through gymnasium.make(), just like a built-in environment.
from gymnasium.envs.registration import register
register(
id="CustomGridWorld-v0",
entry_point=GridWorldEnv,
max_episode_steps=100,
)
Then create it like this:
import gymnasium as gym
env = gym.make(
"CustomGridWorld-v0",
size=5,
render_mode="ansi",
)
Gymnasium environment IDs include a required name and an optional namespace and version. Registration also lets Gymnasium track specifications and apply wrappers such as time limits. If you are new to the ecosystem, start with the Gymnasium introduction.
Test Before You Train
Never throw PPO, DQN, or any other RL algorithm at a fresh environment before you test it manually — and before you pick an algorithm, the DQN vs PPO guide helps you choose. You should confirm that actions work, observations match the declared space, rewards make sense, and episodes end correctly.
Run a random policy first:
env = GridWorldEnv(size=5, render_mode="ansi")
observation, info = env.reset(seed=42)
for _ in range(20):
action = env.action_space.sample()
observation, reward, terminated, truncated, info = env.step(
action
)
print(env.render())
print(
{
"action": action,
"reward": reward,
"terminated": terminated,
"truncated": truncated,
"info": info,
}
)
if terminated or truncated:
observation, info = env.reset()
Check these items before training:
- Every observation fits
observation_space. - Every sampled action fits
action_space. - The agent cannot move beyond valid boundaries.
- The target never starts on the agent.
- Rewards match your intended rules.
terminatedsignals real task endings.truncatedsignals time limits or external cutoffs.reset()returns a valid starting observation.
The Gymnasium documentation also recommends an environment checker for validating custom environments. It can catch API mistakes before your training logs become a mystery novel.
Connect Your Environment to an RL Agent
Once your environment passes manual testing, connect it to an algorithm. For GridWorld, a DQN-style method fits the discrete four-action space well.
Here is an example using Stable-Baselines3:
from stable_baselines3 import DQN
env = GridWorldEnv(size=5)
model = DQN(
policy="MultiInputPolicy",
env=env,
learning_rate=1e-3,
buffer_size=50_000,
learning_starts=1_000,
batch_size=64,
verbose=1,
)
model.learn(total_timesteps=100_000)
Because this environment uses a dictionary observation, the model uses MultiInputPolicy. For a plain vector observation, you would typically use MlpPolicy. The Stable-Baselines3 tutorial covers training setups in detail.
Start with a simple task. If the agent cannot solve a 5 × 5 grid with a clear target, do not immediately add obstacles, moving enemies, procedural maps, weather effects, and an emotionally complex boss fight.
Common Custom Environment Mistakes
Returning the Wrong API Format
Gymnasium expects reset() to return (observation, info) and step() to return five values. Do not return the older four-value done format unless you intentionally use legacy compatibility code.
Declaring Spaces That Do Not Match Data
If you declare dtype=np.float32 but return int64 observations, wrappers and RL libraries may complain — or worse, quietly behave unexpectedly.
Keep shapes, ranges, and data types consistent. Your observation_space should accurately describe every observation you return.
Using the Wrong Episode Flag
Do not mark a time-limit exit as terminated=True unless the task genuinely ends at that point. Use truncated=True for external time limits.
Creating Unlearnable Rewards
A reward only at the final goal can work, but agents may struggle when success happens rarely. Consider light reward shaping, shorter curricula, or easier starting states if learning stalls.
Forgetting Deterministic Seeding
Use self.np_random after super().reset(seed=seed) instead of relying on unmanaged global random calls. This approach gives your environment better reproducibility.
Recommended Books
- Reinforcement Learning: An Introduction by Richard S. Sutton and Andrew G. Barto — the formal treatment of states, actions, rewards, and termination that your environment's design implements in code.
- Modeling and Simulation in Python by Allen B. Downey — a practical guide to turning real-world rules into clean, testable simulations, exactly the skill environment building demands.
- Deep Reinforcement Learning Hands-On by Maxim Lapan — shows how custom environments plug into real training loops, wrappers, and agents.
Want to see your environment trained end to end? Grab the GPTAstra full course at https://cutt.ly/5yviN6qd — it walks from a blank gymnasium.Env file to a tuned agent on your own task.
Frequently Asked Questions
What are the required parts of a Gymnasium environment?
A custom environment subclasses gymnasium.Env and defines an action space, an observation space, a reset() method returning (observation, info), and a step() method returning (observation, reward, terminated, truncated, info). Once you follow that contract, any Gymnasium-compatible library can train on it.
What is the difference between terminated and truncated?
terminated=True signals a natural end to the task, such as reaching a goal or crashing. truncated=True signals an external cutoff such as a time limit. The distinction matters because natural terminal states stop return bootstrapping while time-limit cutoffs usually should not.
When should I use Discrete vs Box action spaces?
Use spaces.Discrete(n) when the agent picks one option from a fixed menu, such as four movement directions. Use spaces.Box for continuous controls such as steering, force, or joystick values. MultiDiscrete and MultiBinary cover separate axes and independent buttons.
Why does my custom environment fail in Stable-Baselines3 or RLlib?
Usually the declared spaces do not match the data: wrong shapes, ranges, or dtypes, or reset()/step() returning the wrong number of values. Keep every observation inside observation_space and return the full five-value step tuple.
How do I register my environment so gymnasium.make() can find it?
Call gymnasium.envs.registration.register with an id like CustomGridWorld-v0, an entry point, and optionally max_episode_steps. Then create it with gym.make('CustomGridWorld-v0', ...), exactly like a built-in environment.
How should I test a new environment before training?
Run a random policy first and inspect observations, rewards, and episode endings. Confirm every observation fits observation_space, sampled actions fit action_space, boundaries hold, and terminated/truncated behave correctly. Gymnasium's environment checker validates the API automatically.
Final Thoughts
Building a custom Gymnasium environment from scratch means turning an idea into a clean, trainable reinforcement learning problem. Define a clear action space, provide useful observations, create honest rewards, separate termination from truncation, and test every detail before training.
Start with a tiny environment like GridWorld. Once you understand the reset() and step() contract, you can scale up to custom games, financial simulations, robotics tasks, scheduling problems, or any other challenge where an agent can learn through actions and feedback.
Your first environment may include bugs. That is normal. Just make sure your agent loses because the task challenges it — not because your target accidentally spawned outside the grid.