Contents
Locomotion tasks like Ant or HalfCheetah reward you for something simple: keep moving forward. Grasping is a genuinely different beast — the reward for "successfully picked up the object" is either 1 or 0, and for most of training, your agent has literally no idea what a successful grasp even looks like. That sparse reward problem is the single biggest obstacle standing between you and a working grasping policy.
I hit this wall directly the first time I tried extending MuJoCo locomotion skills into manipulation — the agent flailed the arm around for thousands of episodes without ever once accidentally succeeding, and without a single success, there's nothing for the algorithm to learn from. Ever wondered why grasping demos always seem to involve some extra trick beyond "just run PPO on it"? This is exactly why, and the fix is genuinely clever once you see it.
By the end of this tutorial, you'll understand why naive RL fails at grasping, what Hindsight Experience Replay actually does to fix it, and you'll have a working training pipeline for a real pick-and-place task. IMO, this is the project where RL for robotics stops feeling like locomotion tutorials and starts feeling like the actual hard, interesting problem :)
Why Grasping Is Harder Than Locomotion
Ant and HalfCheetah give you continuous feedback — moving forward a little bit earns a little bit of reward, so the agent can gradually improve through small increments. Grasping doesn't offer that luxury.
The natural reward signal is sparse and binary: did the arm successfully grasp and move the object to the target, or not? There's no meaningful partial credit in the raw task definition. With random exploration, the odds of an untrained arm accidentally achieving a full successful grasp are vanishingly small, especially with a 7-DOF arm plus a gripper joint to coordinate. Without ever experiencing success, there's genuinely nothing for standard RL algorithms to reinforce — you're stuck at zero, indefinitely.
This is the exact problem that motivated one of the more elegant tricks in modern RL: Hindsight Experience Replay (HER).
If you're coming from locomotion tasks, our MuJoCo tutorial covers the physics simulation foundation that Fetch environments build on — the same contact dynamics, joint controls, and continuous action spaces apply here.
Figure 2: Training a Fetch robot arm to grasp and place objects using SAC with Hindsight Experience Replay
The Task: Fetch Pick-and-Place
The standard entry point for this problem is the Fetch environment suite in Gymnasium-Robotics — a 7-DOF robotic arm with a two-fingered parallel gripper, tasked with manipulating objects like blocks, a pen, or an egg.
FetchReach — the simplest task, just moving the end-effector to a target position, no object involved. FetchPush — the arm needs to push an object along the floor to a target location. FetchSlide — similar to push, but on a low-friction surface where the arm has to strike the object rather than continuously push. FetchPickAndPlace — the hardest of the four: pick up an object and move it to a target position, which has some probability of being above the floor entirely, requiring an actual lift.
These are ordered by genuine difficulty for good reason. Starting directly on PickAndPlace without first solving Reach or Push means debugging your grasping logic and your locomotion logic simultaneously — resist that temptation.
Setting Up the Environment
pip install gymnasium gymnasium-robotics stable-baselines3[extra] mujoco
Let's confirm the environment loads and inspect what we're working with.
import gymnasium as gym
import gymnasium_robotics
gym.register_envs(gymnasium_robotics)
env = gym.make("FetchPickAndPlace-v4", render_mode="human")
observation, info = env.reset(seed=42)
print(observation.keys())
Understanding the Multi-Goal Observation Structure
Notice that observation here isn't a flat array like CartPole's — it's a dictionary with distinct keys, following what's called the Multi-Goal Reinforcement Learning framework.
observation — the actual state: arm position, gripper state, object position, velocities.
achieved_goal — where the object currently is right now.
desired_goal — where you actually want the object to end up.
This structure isn't incidental — it's the exact scaffolding HER needs to work, since the algorithm depends on being able to compare what was achieved against what was desired for any given experience.
For a deeper understanding of how observation spaces affect algorithm choice, our evaluating reinforcement learning algorithms guide covers the metrics and methods used to assess multi-goal performance.
The Core Problem: Sparse Rewards in Action
Let's see the sparse reward problem directly before fixing it.
observation, info = env.reset()
for _ in range(50):
action = env.action_space.sample()
observation, reward, terminated, truncated, info = env.step(action)
print(reward) # almost always -1.0
You'll see -1.0 printed almost every single step, with 0.0 only on the rare timestep the goal is actually achieved. With random actions, that success basically never happens. This is the wall every naive RL approach hits immediately on this task.
The Fix: Hindsight Experience Replay (HER)
Here's the genuinely clever idea behind HER: even a "failed" episode contains useful information — the arm ended up somewhere, even if it wasn't where you wanted. So what if you just... pretended that's what you wanted all along?
After an episode ends, HER takes the trajectory the agent actually experienced and relabels it, substituting the achieved goal in place of the originally desired goal. From the algorithm's perspective, that "failed" episode now looks like a successful episode for a slightly different goal — one it actually accomplished. This gives the agent genuine learning signal from every episode, not just the rare ones that happened to succeed at the original task.
It's a bit like learning to throw darts by declaring wherever your dart landed to be the actual target, then studying what motion got you there. Do that enough times across a variety of "accidental" targets, and you start to understand the relationship between motion and outcome — which transfers directly to hitting your real intended target later.
Training with SAC + HER
Soft Actor-Critic (SAC) pairs naturally with HER for this kind of task — it's an off-policy algorithm, meaning it can reuse and relabel past experience efficiently, which is exactly what HER needs.
from stable_baselines3 import SAC
from stable_baselines3.her import HerReplayBuffer
env = gym.make("FetchPickAndPlace-v4")
model = SAC(
"MultiInputPolicy",
env,
replay_buffer_class=HerReplayBuffer,
replay_buffer_kwargs=dict(
n_sampled_goal=4,
goal_selection_strategy="future",
),
verbose=1,
learning_rate=1e-3,
buffer_size=1_000_000,
)
model.learn(total_timesteps=1_000_000)
model.save("sac_her_fetch_pick_and_place")
Notice MultiInputPolicy, not the MlpPolicy you'd use for a flat observation — that's required specifically because Fetch's dictionary-based observation needs a policy architecture that can handle multiple named inputs. That goal_selection_strategy="future" setting tells HER specifically to relabel goals using states the agent visited later in the same episode — generally the most effective relabeling strategy in practice.
A genuine warning worth taking seriously: this is computationally intensive. Training this exact setup can take several hours on a CPU alone. Don't expect Snake-level training times here — budget accordingly, or use a GPU-accelerated setup if you have one available.
For GPU recommendations if you plan to train at scale, our best GPUs for deep learning guide covers what actually matters for MuJoCo and physics simulation workloads.
Watching Your Trained Agent
model = SAC.load("sac_her_fetch_pick_and_place")
env = gym.make("FetchPickAndPlace-v4", render_mode="human")
observation, info = env.reset()
for _ in range(1000):
action, _ = model.predict(observation, deterministic=True)
observation, reward, terminated, truncated, info = env.step(action)
if terminated or truncated:
observation, info = env.reset()
env.close()
deterministic=True matters here — you want to see the policy's actual best behavior, not exploration noise mixed into the action selection.
Understanding the Action Space
For FetchPickAndPlace specifically, the action space is genuinely compact given the complexity of the task: just 4 dimensions — 3 for the end-effector's movement in space, and 1 for gripper actuation (open/close). The orientation of the gripper is fixed to a preset grasp pose, which meaningfully simplifies the learning problem compared to controlling full 7-DOF joint angles directly.
This simplification is deliberate and worth appreciating. Full joint-space control would be dramatically harder to learn; constraining the action space to end-effector position plus gripper state removes a huge amount of unnecessary complexity while still capturing the actual task.
Progressing Through Task Difficulty
Given how much harder PickAndPlace is than pure locomotion, a staged approach genuinely pays off.
FetchReach — confirms your setup works and gives you a fast training loop to validate hyperparameters before committing to anything harder. FetchPush — introduces object manipulation without requiring an actual grasp, a useful intermediate step. FetchSlide — adds low-friction dynamics, forcing the policy to reason about momentum rather than continuous contact. FetchPickAndPlace — the full task, requiring an actual successful grasp and lift.
Don't skip straight to PickAndPlace to save time. If training fails, you won't know whether the problem is your HER configuration, your reward setup, or something environment-specific — solving the easier tasks first isolates those variables.
Our Gymnasium tutorial covers the wrapper system and observation spaces that Fetch environments use — understanding that foundation makes the multi-goal structure here feel familiar rather than foreign.
Beyond Fetch: The Franka Arm and Other Platforms
Once Fetch's simplified setup feels comfortable, more realistic manipulator platforms exist for further exploration.
Franka Emika Panda — a 7+1 DOF arm with a parallel gripper, available through open-source MuJoCo environments implementing push, slide, and pick-and-place tasks analogous to Fetch's, but modeling a real commercially available robot. Adroit Hand — a genuinely dexterous manipulation platform with up to 30 degrees of freedom, tackling much harder tasks like twirling a pen or opening a latched door. Robosuite — a broader benchmark supporting seven different robot arms and eight gripper types, useful once you want to compare how a trained approach generalizes across hardware.
For physical robot arms you can buy for real-world grasping experiments, our robotics simulation hardware guide covers the MyCobot Pro 630 and other platforms with ROS compatibility for bridging simulation to physical hardware.
Common Mistakes People Make
I've hit a few of these myself while getting this working, and seen others repeated across community implementations.
Skipping HER and expecting vanilla SAC or PPO to solve sparse-reward grasping. Without goal relabeling, the agent essentially never experiences success often enough to learn from it.
Using MlpPolicy instead of MultiInputPolicy. Fetch's dictionary observation space needs a policy architecture built to handle multiple named inputs, not a flat vector.
Jumping straight to PickAndPlace as a first attempt. Solve Reach and Push first to validate your setup before tackling the hardest task in the suite.
Underestimating training time. Several hours on CPU is a realistic expectation for PickAndPlace with HER — don't judge progress against Snake or CartPole's timeline.
Using dense rewards without understanding the tradeoff. Fetch environments support both sparse and dense reward variants — dense rewards train faster but can create unintended local optima where the agent hovers near the object without actually completing the task.
Wrapping This Up
Robot arm grasping exposes a genuinely different challenge than locomotion: sparse, binary rewards that naive RL algorithms simply can't learn from through random exploration alone. Hindsight Experience Replay solves this by relabeling failed episodes as successes toward different goals, giving the agent learning signal from essentially every episode rather than only the rare accidental successes.
Remember to progress through Reach, Push, Slide, and PickAndPlace in that order, and budget real compute time — this genuinely isn't a five-minute training run like CartPole was. FYI, once HER clicks conceptually, it's a technique you'll recognize showing up across a surprising range of sparse-reward RL problems well beyond just robotic grasping :)
Now go watch your untrained arm flail uselessly at a block for a while before HER kicks in and things start clicking. That flailing phase is genuinely normal, not a sign something's broken.