Contents
Everything you've built in this series so far has run in a plain Python environment — Gymnasium, MuJoCo, PettingZoo. Unity ML-Agents flips that setup entirely: you build the world yourself in an actual game engine, then train agents to live inside it. If you've ever wanted your RL agent to inhabit an environment you designed rather than one someone else built for you, this is genuinely the tool that gets you there.
I find this project satisfying for a specific reason: everywhere else in this series, the environment came pre-made. Here, you're the one deciding what the agent sees, what it can do, and what it gets rewarded for — genuinely game-design thinking wrapped around RL fundamentals you already understand.
By the end of this tutorial, you'll have Unity and Python talking to each other, a basic agent set up in a scene, and a trained policy actually controlling it. IMO, watching an agent you designed the rules for actually learn something is a different kind of satisfying than training on someone else's benchmark environment :)
What Unity ML-Agents Actually Is
ML-Agents (Machine Learning Agents Toolkit) is Unity's open-source package that turns any Unity scene into a training environment for reinforcement learning and imitation learning agents. It bridges Unity's engine — physics, rendering, game logic — with a Python-side training backend built on PyTorch.
Supports single-agent, cooperative multi-agent, and competitive multi-agent training scenarios, all within the same framework. Ships with PPO, SAC, and MA-POCA (a multi-agent variant) as built-in training algorithms, plus imitation learning through Behavioral Cloning and GAIL. Includes self-play support natively — genuinely the same self-play concept from AlphaGo, built directly into the toolkit for adversarial game scenarios. The current package (4.0.3 as of mid-2026) targets modern Unity Editor versions, so make sure you're not following documentation written for the older 0.x releases, which used a substantially different API and TensorFlow instead of PyTorch.
FYI, if you find a tutorial referencing version 0.15.0 or earlier, that's genuinely outdated — the toolkit went through a major rewrite since then, and a lot of the setup steps have changed.
Our Gymnasium tutorial covers the core RL API concepts — reset(), step(), action and observation spaces — that ML-Agents maps directly onto through Unity's C# scripting interface.
Figure 1: Unity ML-Agents bridges game engine physics with Python-side RL training, letting you design environments and train agents inside them
Setting Up the Environment
You'll need both Unity Editor and a Python environment configured to talk to each other.
python -m venv venv
source venv/bin/activate # Windows: venv\Scripts\activate
pip install mlagents
On the Unity side, install the ML-Agents package through the Unity Package Manager (Window → Package Manager → search "ML Agents"). Match your Python mlagents package version to your Unity package version — mismatched versions between the two sides is genuinely the single most common source of confusing connection errors beginners hit here.
The Core Building Blocks in Unity
Before writing any training code, you need to understand three Unity-side concepts that map directly onto the RL vocabulary you already know.
Agent — a Unity GameObject with an attached Agent script, representing the thing being trained. This is your RL agent, living inside the actual 3D (or 2D) scene. Behavior Parameters — a component defining the agent's observation size, action space (discrete or continuous), and which neural network model controls it. Academy — the singleton managing the overall training process, synchronizing Unity's simulation with the Python training backend.
Writing Your First Agent Script
Here's a minimal C# agent script — genuinely the Unity-side equivalent of defining reset(), step(), and reward logic in Gymnasium.
using Unity.MLAgents;
using Unity.MLAgents.Sensors;
using Unity.MLAgents.Actuators;
public class BallAgent : Agent
{
public Transform target;
Rigidbody rBody;
public override void OnEpisodeBegin()
{
transform.localPosition = new Vector3(0, 0.5f, 0);
rBody.velocity = Vector3.zero;
target.localPosition = new Vector3(Random.Range(-4f, 4f), 0.5f, Random.Range(-4f, 4f));
}
public override void CollectObservations(VectorSensor sensor)
{
sensor.AddObservation(target.localPosition);
sensor.AddObservation(transform.localPosition);
sensor.AddObservation(rBody.velocity.x);
sensor.AddObservation(rBody.velocity.z);
}
public override void OnActionReceived(ActionBuffers actions)
{
float moveX = actions.ContinuousActions[0];
float moveZ = actions.ContinuousActions[1];
rBody.AddForce(new Vector3(moveX, 0, moveZ) * 10f);
float distanceToTarget = Vector3.Distance(transform.localPosition, target.localPosition);
if (distanceToTarget < 1.42f)
{
SetReward(1.0f);
EndEpisode();
}
else if (transform.localPosition.y < 0)
{
EndEpisode();
}
}
}
Notice the structural parallel to Gymnasium: OnEpisodeBegin() mirrors reset(), CollectObservations() builds your state vector, and OnActionReceived() is exactly your step() function, complete with reward assignment. The concepts transfer directly — only the syntax and the fact that you're inside an actual game engine change.
Our CartPole tutorial shows the Python equivalent — comparing that implementation to this C# version makes the conceptual mapping immediately clear.
Configuring the Training Run
Training hyperparameters live in a YAML config file, separate from your Unity scene code.
behaviors:
BallAgent:
trainer_type: ppo
hyperparameters:
batch_size: 64
buffer_size: 12000
learning_rate: 3.0e-4
network_settings:
normalize: true
hidden_units: 128
num_layers: 2
max_steps: 500000
time_horizon: 64
This separation is genuinely convenient — you can experiment with different algorithms and hyperparameters without touching a single line of C# code. Swap trainer_type from ppo to sac, and you're training a different algorithm entirely with no scene changes required.
Running Training
With your scene built and config file ready, training kicks off from the command line — Unity itself just plays the role of the environment.
mlagents-learn config/ball_agent_config.yaml --run-id=first_run
Press play in the Unity Editor when prompted, and training happens live, visibly, in the actual game view — genuinely one of the more satisfying differences from watching print statements scroll by during a Gymnasium training run. You're watching your agent physically attempt the task, fail, and gradually improve, rendered in real time.
Speeding Up Training with Multiple Environments
Just like vectorized environments in Gymnasium and SB3, ML-Agents supports running multiple concurrent Unity environment instances to collect experience faster.
mlagents-learn config/ball_agent_config.yaml --run-id=first_run --num-envs=4
This is genuinely the same speed-up principle you saw with SB3's vectorized environments — more parallel experience collection, faster wall-clock training, no change to the underlying algorithm.
Our Stable-Baselines3 tutorial covers vectorized environments in detail — understanding that concept makes ML-Agents' --num-envs flag feel familiar rather than new.
Self-Play for Competitive Games
Remember the self-play concept from AlphaGo? ML-Agents bakes it directly into the toolkit for training agents in adversarial scenarios — think a simple 1v1 game where both sides are learning simultaneously.
behaviors:
SoccerTwos:
trainer_type: poca
self_play:
save_steps: 50000
team_change: 200000
swap_steps: 10000
window: 10
play_against_latest_model_ratio: 0.5
initial_elo: 1200.0
That play_against_latest_model_ratio setting genuinely mirrors the curriculum-learning idea from earlier in this series — it controls how often the agent faces its most recent self versus older versions, balancing training stability against the risk of overfitting to one specific past opponent.
Our self-play guide covers AlphaGo through AlphaZero in detail — understanding that progression makes ML-Agents' built-in self-play configuration feel like a natural application rather than a new concept.
Curriculum Learning, Also Built In
Remember the curriculum learning discussion — easy tasks first, harder tasks later? ML-Agents has this as a first-class feature, letting you define difficulty progression declaratively rather than hand-coding it into your environment logic.
environment_parameters:
wall_height:
curriculum:
- name: Lesson0
completion_criteria:
measure: progress
behavior: BigWallJump
threshold: 0.1
value: 0.0
- name: Lesson1
completion_criteria:
measure: progress
behavior: BigWallJump
threshold: 0.3
value: 4.0
This is genuinely the exact concept from our curriculum learning article, expressed as toolkit configuration — wall height (task difficulty) increases automatically once the agent hits a specified competence threshold, no manual intervention required.
Environment Randomization for Generalization
To prevent agents from overfitting to one specific scene layout, ML-Agents supports randomizing environment parameters across training episodes — genuinely the same domain randomization concept from the MuJoCo sim-to-real discussion, applied here to game environments instead of robotics hardware.
environment_parameters:
gravity:
sampler_type: uniform
sampler_parameters:
min_value: 7.0
max_value: 12.0
Training against varied gravity, lighting, or obstacle placement produces agents that generalize better to slightly different conditions, rather than ones that only work in the exact scene configuration they trained on.
Our MuJoCo tutorial covers domain randomization for sim-to-real robotics — the same principle applies here to game environments.
Deploying the Trained Model Back Into Unity
Once training completes, ML-Agents exports a trained model (.onnx format) that plugs directly back into your Behavior Parameters component in Unity — no Python runtime needed at inference time.
Drag the exported .onnx file into the Behavior Parameters component's model slot.
Switch the Behavior Type to Inference Only, and the agent now runs its trained policy natively inside Unity.
This is genuinely how you'd ship a trained agent as actual NPC behavior in a real, shippable game — not just a training demo.
Common Mistakes People Make
I've hit a few of these myself while getting comfortable with the toolkit, and seen others repeated in community projects.
Mismatched Python and Unity package versions. This produces confusing connection errors that look like a code bug but are actually a version compatibility issue.
Forgetting to normalize observations for continuous state spaces. The normalize: true setting in your config genuinely matters for training stability, similar to input normalization in any neural network.
Overcomplicating the reward function on a first project. Start with a simple, clear reward signal (like the distance-based reward in the ball example) before attempting more nuanced reward shaping.
Not using curriculum learning for genuinely hard tasks. If your custom environment has a hard-exploration problem, ML-Agents' built-in curriculum system solves exactly the same issue this whole series covered earlier — don't skip it just because it's config-based instead of code-based.
Following outdated 0.x documentation. The toolkit changed substantially since those versions — always confirm you're reading current Unity Package documentation, not legacy GitHub wiki pages.
Where This Fits Compared to Everything Else in This Series
Genuinely worth being direct about when to reach for this versus Gymnasium plus SB3.
Use ML-Agents when you're building an actual game and want trained NPC behavior shipped inside it, or when you want full control over environment design through a real game engine's tools (physics, rendering, level design). Use Gymnasium plus SB3 when you're doing research, benchmarking against standard tasks, or don't need Unity's visual/game-engine capabilities specifically. The underlying RL concepts transfer completely either way — PPO is PPO, curriculum learning is curriculum learning, self-play is self-play. What changes is which tool actually renders and simulates your environment.
Wrapping This Up
Unity ML-Agents takes everything you've learned across this series — PPO, self-play, curriculum learning, environment randomization — and lets you apply it inside an actual game engine you control, rather than a pre-built Python environment. The RL concepts don't change; what changes is that you're now the one designing the world your agent lives in.
Remember to keep your Python and Unity package versions matched, start with simple reward functions before layering in complexity, and lean on the toolkit's built-in curriculum and self-play systems rather than reimplementing concepts you already understand from scratch. FYI, a trained model exports to ONNX and runs natively in Unity with no Python dependency at inference time — meaning what you build here can genuinely ship inside a real game, not just live as a training demo :)
Now go build the simplest possible scene — one agent, one target, one reward signal — and watch it learn before adding any real complexity. That minimal version teaches you the Unity-specific plumbing without also debugging a complicated reward function at the same time.