Sam Austin AI

Training AI Agents in Minecraft with Reinforcement Learning

September 26, 2026 17 min read Sam Austin
Contents

Game console and controller on a green backdrop — Minecraft is the sandbox where game AI research meets real gameplay

Figure 1: The reward is one diamond at the end of a ten-step crafting chain — and that single number is why this game broke naive RL

Recall the curriculum learning article's central example from much earlier in this series — a robot that can't learn rough-terrain locomotion because random exploration never stumbles into a single success. Minecraft's "obtain diamond" task is genuinely the canonical, most-cited version of that exact problem in all of RL research. To mine a diamond, an agent must chop wood, craft planks, build a crafting table, make a wooden pickaxe, mine stone, craft a stone pickaxe, mine iron, smelt it, craft an iron pickaxe, and only then can it mine diamond ore at all. A single reward at the end of that chain, with nothing in between, is about as sparse and hard-exploration as an RL benchmark gets — which is exactly why Minecraft became one of the field's most important open-ended testbeds rather than just a fun demo.

MineRL — the standard Gym-compatible environment for this research, currently at v1.0 — was built specifically to study this. But the genuinely interesting story of the last several years isn't "researchers found a clever reward shaping trick." It's that pure RL from scratch essentially failed on this task, and the field's actual breakthroughs came from two very different directions: pretraining on human behavior at internet scale (VPT), and abandoning gradient-based RL almost entirely in favor of LLM-driven planning (Voyager).

By the end of this guide, you'll understand why vanilla RL struggles here, how VPT's video-pretraining approach connects directly to this series' behavior cloning article, and how Voyager represents a genuinely different paradigm worth knowing about even if you never train a single neural network. IMO, "recent works found pretrained LLMs could serve as a strong mind that provides planning ability" is the sentence that best captures how much this specific sub-field's center of gravity has shifted :)

Why Minecraft Broke Pure RL

A 2026 study investigating hierarchical deep RL for Minecraft states the problem directly: open-world games pose significant challenges for RL due to their long-horizon objectives, sparse rewards, and requirement for compositional skill learning. Recall the reward design article's sparse-versus-shaped tradeoff from earlier in this series directly — Minecraft sits at the genuinely extreme end of that spectrum.

  • Early works focused on pure RL or pure imitation learning without satisfactory performance — worth being direct about this rather than glossing over it; this isn't a case where a clever hyperparameter fix eventually cracked it, it's a case where the field's initial approaches genuinely didn't scale to the task's difficulty.
  • The 2019, 2020, and 2021 MineRL competitions were explicitly framed around "sample efficient reinforcement learning using human priors" — the competition's own framing acknowledges directly that pure trial-and-error exploration wasn't going to get an agent to a diamond within any reasonable compute budget.
  • The clearest single datapoint comes from OpenAI's own VPT writeup: an RL policy trained from random initialization barely achieves any reward and never learns to reliably collect logs, the very first step of the chain. Everything downstream of that first step is unreachable by random exploration alone.

That last point is worth pausing on, because it reframes the whole topic. The problem was never that the reward function was badly shaped — it's that the exploration problem is combinatorially hopeless before reward design even enters the picture.

Getting Started: The MineRL Environment

pip install git+https://github.com/minerllabs/minerl

Requires Java JDK 8 specifically — a genuinely real, common installation friction point, since Minecraft's underlying Java version dependency doesn't play well with newer JDK releases on many systems.

import minerl
import gym

env = gym.make("MineRLObtainDiamondShovel-v0")
obs = env.reset()

action = env.action_space.sample()
obs, reward, done, info = env.step(action)

Notice this is genuinely the same Gym API shape from the very first Gymnasium tutorial in this series' RL arc — reset(), step(), an action space and observation space — Minecraft's complexity lives entirely in what those spaces actually contain (a full RGB frame observation, a compound action space covering movement, camera control, and crafting), not in a different fundamental interface.

  • v1.0 is the current version, needed for OpenAI's VPT models and the MineRL BASALT 2022 competition — recall the exact same "check which fork or version you're actually using" discipline from the Gym/Gymnasium and TFLite/LiteRT transitions covered earlier in this series; MineRL v0.3 and v0.4 remain around for older competition compatibility but genuinely aren't where current work happens.
  • The environment ID above (MineRLObtainDiamondShovel-v0) is a shovel-first variant of the diamond chain — useful because it exercises the identical crafting dependency graph with a shorter horizon, which matters when you're debugging an exploration problem rather than chasing a benchmark number.
  • Budget real setup time. Between the JDK pin, the Minecraft launcher handshake, and the Python package build, environment setup is the most common place newcomers quietly give up before training has even started.

A Concrete Starting Project: DQN on a Simple Maze

Before tackling the full diamond challenge, a genuinely sensible entry point exists: a documented open-source project (vincentberaud/Minecraft-Reinforcement-Learning, published alongside arXiv 1903.04311) trains a Double DQN agent to solve a maze using the Malmo platform — Microsoft's original Minecraft-as-RL-environment project — observing the current frame plus the last three frames stacked together. Recall the Atari tutorial's frame-stacking technique from much earlier in this series directly; solving the exact same "a single frame can't convey motion" problem in this genuinely different visual domain.

  • TensorBoard tracks the training curves throughout a run — recall the Stable-Baselines3 tutorial's tensorboard_log parameter directly; the same monitoring discipline, applied here to a genuinely harder visual environment than CartPole or even Atari.
  • Checkpointing lets training pause and resume — recall the model registries and DVC articles' versioning philosophy directly; a Minecraft training run can genuinely take long enough that resumable checkpoints aren't a nicety, they're required infrastructure.
  • The project's own comparison of stacked versus recurrent versus dueling variants is a useful bonus: frame stacking measurably beat a single-frame baseline on the simple mission, which is the same lesson Atari taught, restated in a 3D first-person world.

The honest scope check: a maze is a maze, and the diamond chain is not. What this project buys you is a working Malmo-plus-Gym-plus-DQN loop with all the boring infrastructure already wired up — the part you would otherwise burn a weekend on before learning anything about RL.

VPT: The Behavior Cloning Article's Pattern, at Internet Scale

Recall the behavior cloning article's core argument from earlier in this series directly: "pretrain a policy with behavior cloning on demonstration data first, then fine-tune with reinforcement learning." OpenAI's Video PreTraining (VPT) is genuinely the most ambitious real-world instance of exactly that pattern this series has covered.

  • VPT trains an inverse dynamics model on a small set of labeled data (video paired with the actual keyboard and mouse actions that produced it), then uses that model to label millions of hours of unlabeled Minecraft videos scraped from the internet — turning raw YouTube footage into a genuinely massive behavior-cloning dataset with no manual annotation required at that scale. The filtered training set came to roughly 70,000 hours of gameplay.
  • This behavioral prior is then used as a starting point for RL fine-tuning — recall the exact BC-then-RL sequence from the behavior cloning article directly, just with "expert demonstrations" meaning "millions of hours of anonymous human Minecraft footage" rather than a small, deliberately-collected teleoperation dataset.
  • VPT scaled semi-supervised behavior learning from YouTube videos to human-level capability on early tech-tree tasks — genuinely the concrete, large-scale validation that this series' behavior cloning article's central recommendation works even at internet scale, not just in a controlled robotics lab setting. The paper reports human-level performance on many tasks and, notably, human-level success rates at collecting everything leading up to a diamond pickaxe once BC and RL fine-tuning are stacked in sequence.

The three-phase result is the number worth remembering: pretraining, then BC fine-tuning, then RL fine-tuning produced over 80% reliability on iron pickaxes and around 20% on collecting diamonds, where random-init RL produced effectively nothing. Human players given the same objective hit 57% and 15% on those same two milestones — which is what "human-level on early tech tree" actually means in practice.

Hierarchical RL: Recall the Curriculum Learning Article Directly

A genuinely direct implementation of the curriculum learning and hierarchical decomposition principles from earlier in this series — Hierarchical Deep Reinforcement Learning (HDRL) approaches decompose Minecraft's long-horizon task into a high-level planner choosing sub-goals ("get wood," "craft a pickaxe") and a low-level controller executing the primitive actions needed to achieve each sub-goal.

  • This mirrors exactly the "give the agent a foothold of existing skill before tackling the full hard task" principle from the curriculum learning article — rather than training one flat policy against the sparse diamond reward, hierarchical decomposition breaks the problem into a sequence of genuinely tractable sub-tasks, each with its own denser, more learnable signal.
  • JueWu-MC and SEIHAI, both named in the research literature as MineRL-competition entries, use exactly this sample-efficient hierarchical approach directly — genuinely established, competition-validated technique, not a theoretical proposal. SEIHAI took first place in the 2020 competition by scheduling across multiple sub-task experts; JueWu-MC took the 2021 championship on the same sample-efficiency framing.
  • The 2026 HDRL study's own conclusion states this directly: hierarchical reinforcement learning provides a scalable and interpretable framework for developing agents capable of long-term reasoning and adaptive skill composition — worth reading as direct, current confirmation that the curriculum learning article's principles generalize to this genuinely harder domain.

The structural insight is the same one the path planning article reached through a different door: long-horizon problems get solved by decomposition, not by a bigger network.

Voyager: The Genuinely Different Paradigm

This is worth treating as a real departure from everything else in this article, not a variant of the same approach. Voyager builds an open-ended embodied agent using GPT-4 as the planner, not a trained neural network policy at all — no gradient descent, no reward function, no environment interaction loop in the traditional RL sense.

Three components make up the system: an automatic curriculum maximizing exploration, an ever-growing skill library of executable code storing and retrieving complex behaviors, and an iterative prompting mechanism incorporating environment feedback and execution errors to refine its own generated code.

  • Recall this connecting directly to the RAG and agent-tooling arc from much earlier in this series — Voyager's skill library, genuinely persisted and retrievable across sessions, is conceptually a specialized vector store-adjacent memory system, just storing executable Minecraft skills instead of document embeddings, and the MCP fundamentals article's tool-calling framing is the closest match to how those skills get invoked.
  • The paper's own framing is genuinely important: classical RL and imitation learning operate on primitive actions, challenging for systematic exploration, interpretability, and generalization — while LLM-based agents harness pretrained world knowledge to generate consistent action plans directly, without needing millions of environment interactions to discover that "chop a tree to get wood" is a sensible first step.
  • Voyager is explicitly described as orthogonal to and combinable with VPT — worth knowing this isn't a strict either-or choice; the paper's own condition is that the low-level controller provides a code API, so a system could genuinely use VPT-style motor control alongside Voyager-style high-level planning, each handling the part of the problem it's actually suited for.
Approach What it learns Where it gets its knowledge Failure mode
Pure RL from scratch Policy over primitive actions Trial and error against a sparse terminal reward Never finds the first success to reinforce
Hierarchical RL (SEIHAI, JueWu-MC) Sub-goal policies plus a scheduler Shaped sub-task rewards and human priors Needs the right decomposition designed up front
VPT (BC then RL) Low-level motor prior, then fine-tuned policy 70,000 hours of labeled internet video Still needs a reward to fine-tune against
Voyager (LLM planning) Executable code skills, no gradients GPT-4's pretrained world knowledge Depends on the planner's grounding and API

The table is the article's thesis in one view: the field moved away from asking one policy to absorb the entire crafting chain, and toward supplying prior knowledge — from video, from decomposition, or from a language model.

Distributed Training: Ray for Minecraft Agents

Current, real-world Minecraft agent tooling explicitly recommends Ray for distributed training and inference — recall the reinforcement learning frameworks guide's in-depth RLlib coverage from earlier in this series, and the Ray Tune article for the hyperparameter half of the same problem. A documented multi-agent evaluation setup used MineStudio's distributed inference to enable real-time multi-agent evaluation, genuinely the same "training or evaluating many agents in parallel, potentially across machines" problem RLlib's EnvRunner architecture was built to solve.

  • Minecraft rollouts are slow in a way CartPole rollouts never are — each step drives a running game instance — so parallelizing environment interaction is where the wall-clock wins actually come from, not from a faster optimizer.
  • Recall the multi-agent RL article's evaluation problem directly: measuring many independently-trained agents against each other is itself a distributed workload, which is exactly what MineStudio's Ray pipeline is doing in that setup.
  • If your training run will outlive your laptop session — and at Minecraft scale it will — the checkpointing discipline from the maze project section stops being optional. Pair it with the GPU sizing guidance before you commit to a long run.

A Practical Decision Framework

  1. Are you learning RL fundamentals, or trying to build a genuinely capable agent quickly? Recall the maze-DQN starting project directly — for learning, a simplified Minecraft task with DDQN and frame stacking teaches the same lessons as Atari, in a genuinely more visually rich environment; don't start with the full diamond challenge.
  2. Do you have access to human demonstration data or footage? Recall the VPT pattern directly — if yes, a behavior-cloning-then-RL-fine-tuning pipeline is genuinely the evidenced, scalable path, not a fallback option.
  3. Is your task genuinely long-horizon and compositional (multi-step crafting chains)? Recall the hierarchical RL approach directly — decompose it explicitly into sub-goals rather than training one flat policy against a single sparse terminal reward.
  4. Do you want an agent that plans and adapts without gradient-based training at all? Recall Voyager directly — an LLM-driven planner with a persistent skill library is a genuinely viable, increasingly dominant alternative, especially for open-ended tasks without a single well-defined objective.
  5. Are you training or evaluating multiple agents, or need distributed compute? Recall the RL frameworks guide's RLlib coverage and MineStudio's Ray-based distributed inference directly — this is exactly the scaling problem that framework was built to solve.

Common Mistakes People Make

Attempting the full ObtainDiamond task with vanilla PPO or DQN from scratch

Recall the field's own documented history directly — pure RL without human priors was explicitly the approach that motivated the MineRL competitions' entire "sample efficient learning" framing in the first place, precisely because it doesn't work well alone.

Ignoring VPT-style pretraining when demonstration or video data is available

Recall the behavior cloning article's core argument directly — this is genuinely the evidenced, scalable path for exactly this kind of long-horizon, sparse-reward task.

Treating Voyager and traditional RL as competing, mutually exclusive approaches

Recall this being explicitly described as orthogonal and combinable in the original paper — the field's current direction increasingly blends LLM-based high-level planning with lower-level learned or hard-coded control.

Installing the wrong MineRL version for your intended use case

Recall the explicit version distinctions directly — v1.0 for VPT and BASALT-era work, v0.3/v0.4 only for older competition reproduction.

Underestimating the JDK and environment setup friction

Recall this being a genuinely common, documented pain point — budget real setup time before assuming training can start immediately.

Want the MineRL-to-VPT pipeline in runnable code? Grab the GPTAstra full course at https://cutt.ly/5yviN6qd — it walks from installing the environment and running random actions to building the pretrain-then-fine-tune loop this article describes.

Frequently Asked Questions

What is MineRL?

MineRL is an open-source, Gym-compatible Python environment for Minecraft that lets reinforcement learning algorithms control a Minecraft agent through the same reset() and step() interface used by Gymnasium. It was built specifically to drive research on sample-efficient learning in a sparse-reward, long-horizon sandbox world.

Why is the ObtainDiamond task so hard for reinforcement learning?

The crafting chain to a diamond — wood, planks, crafting table, wooden pickaxe, stone, stone pickaxe, iron, smelting, iron pickaxe — pays out only at the end, so random exploration essentially never stumbles into success. Open-world games pose long-horizon objectives, sparse rewards, and a requirement for compositional skill learning.

What is VPT and why does it matter?

Video PreTraining (VPT) trains an inverse dynamics model on a small set of labeled gameplay, uses it to pseudo-label roughly 70,000 hours of unlabeled Minecraft video, and trains a behavioral cloning prior from those labels. That prior can then be fine-tuned with reinforcement learning on tasks that are impossible to learn from scratch.

How is Voyager different from a traditional RL agent?

Voyager uses GPT-4 as its planner instead of a trained policy — no gradient descent and no reward optimization loop. It combines an automatic curriculum, an ever-growing skill library of executable code, and an iterative prompting mechanism that feeds environment feedback and execution errors back into code generation.

Which MineRL version should I install?

Use v1.0 for VPT models and the MineRL BASALT 2022 setup. The older v0.3 and v0.4 releases remain only for reproducing the 2019 to 2021 competition era. MineRL also requires Java JDK 8 specifically, which is the single most common setup failure point.

Do I need a GPU to start with Minecraft reinforcement learning?

Not for a small starting project — a Double DQN maze agent trains on modest hardware. Internet-scale pretraining and RL fine-tuning on raw pixels are GPU workloads measured in days, so treat the maze-scale project as your CPU-friendly entry point.

Wrapping This Up

Minecraft genuinely earned its place as one of RL's most important open-ended benchmarks precisely because its long-horizon, sparse-reward, compositional-skill structure broke naive approaches this series' earlier CartPole-through-robotics content never had to confront at this scale — VPT's internet-scale video pretraining is the behavior cloning article's BC-then-RL pattern taken to its most ambitious realization, hierarchical RL is the curriculum learning article's decomposition principle validated in a genuinely harder domain, and Voyager represents current research's growing recognition that an LLM's pretrained world knowledge can substitute for millions of environment interactions entirely.

Remember that pure RL from scratch was documented as insufficient for tasks like ObtainDiamond, which is exactly why the field's actual progress came from combining RL with either large-scale imitation learning or LLM-based planning rather than better reward engineering alone. FYI, this article genuinely closes an interesting loop across this entire series — the curriculum learning and behavior cloning articles' principles, developed in the context of robot locomotion and manipulation, turn out to generalize directly to an entirely different domain (a sandbox video game), which is itself a fairly strong signal that those aren't robotics-specific tricks but genuine, general RL principles :)

Now go run the MineRL MineRLObtainDiamondShovel-v0 environment with pure random actions for a few episodes, exactly like the sim-to-real transfer article's "watch random actions fail" baseline from much earlier in this series, and count how many crafting-chain steps it manages to stumble into by chance. That number, almost certainly zero, is the single clearest demonstration of why this specific benchmark forced the field toward VPT and Voyager rather than a bigger PPO run.

Share this article X Facebook LinkedIn Reddit WhatsApp

Related Articles