Contents
Every RL project you've built so far had exactly one agent making decisions in a world that just... sat there and responded. Multi-agent RL blows that assumption up entirely — now the "environment" includes other learning agents, each one changing its behavior in response to yours, which is changing in response to theirs. The ground you're standing on is actively moving.
I got interested in MARL specifically after noticing how many of the most famous RL achievements — AlphaGo Zero, OpenAI Five, AlphaStar — weren't single-agent problems at all. They involved agents competing or cooperating with other learning agents, and that's a genuinely different flavor of hard than anything CartPole or BipedalWalker prepared me for. Ever wondered why training a single agent against a fixed opponent feels satisfying but training two agents against each other feels like chasing a moving target? That's the entire subject of this article.
By the end, you'll understand what actually makes multi-agent RL different from everything you've built before, the major coordination strategies, and which frameworks to actually use. IMO, this is where RL stops being about "solving an environment" and starts being about navigating genuinely social dynamics :)
What Actually Changes With Multiple Agents
In every environment you've trained on — CartPole, Snake, BipedalWalker — the environment's rules were fixed. The pole's physics didn't change because your agent got better. Multi-agent RL removes that guarantee entirely.
The environment becomes non-stationary from any single agent's perspective. Other agents are learning and changing their policies simultaneously, so "the same action in the same state" can start producing different outcomes purely because someone else's strategy shifted. Agents can be cooperative, competitive, or a genuine mix of both — think a soccer team (cooperate with teammates, compete with the other team) rather than a clean cooperative-or-adversarial split. Credit assignment gets genuinely harder. If a team of agents succeeds, which agent's actions actually deserve the reward? This problem barely exists in single-agent RL and dominates a lot of MARL research.
Research in this area has motivated serious attention precisely because of results like AlphaGo Zero, OpenAI Five, and AlphaStar — MARL is behind some of the most publicized achievements in the field, not a niche corner of it.
If you're coming from single-agent projects, our beginner's guide to RL covers the core concepts — agent, environment, rewards — that MARL extends to multiple interacting agents.
Figure 1: Multi-agent RL introduces non-stationary environments where other learning agents change the rules as you play
The Three Coordination Strategies
Multi-agent approaches broadly split into three camps, based on how much agents actually know about or coordinate with each other.
Independent Learners
Each agent runs its own single-agent algorithm — PPO, Q-learning, SAC — and simply treats every other agent as part of the environment, with no explicit awareness that they're being trained too.
IPPO, IQL, ISAC are independent adaptations of their familiar single-agent counterparts (PPO, Q-learning, and SAC respectively). Genuinely simple to implement — you're mostly reusing single-agent code with minimal modification. The tradeoff: since every other agent is also learning and changing, the "environment" each individual agent sees is constantly shifting, which can make training unstable or slow to converge.
Parameter Sharing Approaches
Instead of treating agents as fully independent, these methods have agents share components — often a critic network or value function — while sometimes still maintaining separate policies.
MAPPO, MASAC, MADDPG are common examples in this category. Sharing a critic gives agents access to more global information during training, even if they act on local observations at execution time. This genuinely helps with the credit assignment problem — a shared critic can evaluate joint actions rather than each agent guessing at its individual contribution in isolation.
Centralized Training, Decentralized Execution (CTDE)
This has become one of the dominant paradigms in modern MARL specifically because it resolves a real tension: training benefits from full information, but deployment often can't have it.
During training, agents (or a shared critic) get access to the full global state — what every agent is doing, not just what one agent can locally observe. During execution, each agent acts using only its own local observations, since that's genuinely all it would have access to in a real deployed scenario. MADDPG and MAPPO both fundamentally rely on this training/execution split.
This distinction matters enormously in practice. A group of real robots coordinating on a warehouse floor can't all instantly share perfect global information at execution time the way they can during a training simulation — CTDE bakes that real-world constraint directly into the training process itself.
For a deeper understanding of how these coordination strategies affect algorithm choice, our evaluating reinforcement learning algorithms guide covers the metrics and methods used to assess multi-agent performance.
Cooperative vs. Competitive vs. Mixed Settings
The nature of the reward structure fundamentally shapes which techniques actually work well.
Fully cooperative — all agents share a common goal and typically a shared (or highly correlated) reward. Benchmarks like Pistonball and Cooperative Pong from PettingZoo test exactly this setting. Fully competitive — agents have directly opposing goals, most classically represented by two-player zero-sum games. Mixed-motive — agents have partially aligned, partially conflicting incentives, which is genuinely the most realistic setting for anything resembling real-world multi-agent interaction. Frameworks like Melting Pot specifically target these scenarios.
Mixed-motive settings are where MARL research has been trending, precisely because pure cooperation and pure competition are comparatively rare in genuinely realistic multi-agent scenarios — most real coordination problems involve some blend of shared and competing interests.
PettingZoo: The Gymnasium Equivalent for Multi-Agent RL
If you've gotten comfortable with Gymnasium's API, PettingZoo will feel immediately familiar — that's intentional. It was built specifically to extend Gymnasium's standardized interface into multi-agent settings.
pip install pettingzoo
from pettingzoo.butterfly import pistonball_v6
env = pistonball_v6.env(render_mode="human")
env.reset(seed=42)
for agent in env.agent_iter():
observation, reward, termination, truncation, info = env.last()
if termination or truncation:
action = None
else:
action = env.action_space(agent).sample()
env.step(action)
env.close()
Notice the agent_iter() pattern — this is genuinely new compared to single-agent Gymnasium, and it's how PettingZoo handles the fact that multiple agents need turns to act, whether sequentially or simultaneously. PettingZoo supports both turn-based and simultaneous-action games, and follows strict environment versioning so results stay comparable across research.
Our Gymnasium tutorial covers the core API that PettingZoo extends — reset(), step(), action and observation spaces — so the actual new learning here is about the multi-agent iteration pattern, not new environment mechanics.
Why PettingZoo Deliberately Leaves Out Preprocessing
Unlike Gymnasium's built-in wrappers, PettingZoo intentionally keeps preprocessing logic separate — that's handled by a companion library called SuperSuit, which provides common wrappers like frame stacking for giving agents temporal context across observations. This separation keeps the core library focused purely on the multi-agent environment interface itself.
PettingZoo has become genuinely foundational in this space — other libraries like CleanRL, Tianshou, and AgileRL use it as their standard MARL benchmark suite, similar to how Gymnasium became the shared reference point for single-agent work.
Other Frameworks Worth Knowing
Beyond PettingZoo, a handful of other tools solve different pieces of the MARL puzzle.
Melting Pot — focused specifically on modeling social dilemmas: free-rider problems, mixed motives, and coordination challenges that mirror genuinely social dynamics rather than clean game-theoretic setups. EPyMARL / PyMARLzoo+ — research-oriented frameworks integrating standard baseline algorithms (MAA2C, MAPPO, QMIX) across multiple benchmark suites, useful if you want to rigorously compare algorithms rather than just get one working. BenchMARL — built specifically to address fragmentation and reproducibility problems across the MARL research landscape, offering standardized baselines researchers can actually compare against each other. RLlib — a broader distributed RL library that includes solid multi-agent training support, useful once you're scaling beyond a single machine.
Don't feel obligated to explore all of these immediately. PettingZoo plus Stable-Baselines3 or RLlib covers the overwhelming majority of what a beginner project actually needs.
Our best RL frameworks guide covers Stable-Baselines3, CleanRL, RLlib, and more — including their multi-agent capabilities.
A Practical Starting Project
If you want to get hands-on rather than just reading about coordination strategies, here's a sensible entry point.
Start with a fully cooperative PettingZoo environment like Pistonball or Cooperative Pong — shared reward structures are genuinely easier to reason about than competitive or mixed settings. Use independent PPO first (essentially running single-agent PPO per agent) before attempting anything more sophisticated like MAPPO. Once independent learners work reasonably, experiment with a shared critic (moving toward MAPPO) and compare training stability and final performance. Move to a competitive two-agent environment afterward, and notice how differently training dynamics behave when agents are working against each other rather than together.
For hands-on projects that build toward MARL, our Snake AI tutorial and Atari tutorial cover the single-agent DQN fundamentals that independent MARL agents build directly on.
Common Mistakes People Make
I've seen these repeated across enough MARL discussions and projects to flag them as genuine patterns.
Assuming single-agent algorithms transfer cleanly. Independent learners technically work, but the non-stationarity problem is real — expect more training instability than single-agent RL prepared you for. Jumping straight into competitive settings. Cooperative environments have clearer reward structures and are genuinely easier to debug when something isn't working. Ignoring the training/execution information gap. If your agents need to act on local observations at deployment, training with full global state access (and no plan to transition away from it) sets up an unrealistic expectation for what execution will actually look like. Underestimating credit assignment difficulty. In cooperative settings especially, figuring out which agent's actions actually earned a shared reward is a genuinely hard, unsolved-in-general problem — don't expect a trivial fix.
Wrapping This Up
Multi-agent RL trades the comparatively stable world of single-agent problems for something genuinely more dynamic: other learning agents whose changing behavior makes the environment non-stationary from any individual agent's perspective. Independent learners, parameter sharing, and centralized-training-decentralized-execution represent three increasingly sophisticated ways of handling that instability.
Remember that PettingZoo extends the exact Gymnasium API you already know into multi-agent settings, and that cooperative environments are genuinely the more approachable starting point before competitive or mixed-motive scenarios. FYI, a huge share of RL's most publicized achievements — AlphaGo, OpenAI Five, AlphaStar — are fundamentally multi-agent problems, so this genuinely isn't a niche corner of the field you can skip :)
Now go run independent PPO agents against each other in a simple PettingZoo environment and watch what happens when neither one's strategy stays still long enough for the other to fully adapt. That moving-target dynamic is genuinely the heart of what makes this subfield different.