Sam Austin AI

Hindsight Experience Replay: Learning from Failed Attempts

September 25, 2026 13 min read Sam Austin
Contents

An engineer working with a robotic arm, the classic setting for hindsight experience replay

Figure 1: The arm missed the target — with HER, that miss is now training data for somewhere it did reach

A robot arm misses a block. A navigation agent stops near the wrong doorway. A game character runs confidently in the exact opposite direction of the treasure. Standard reinforcement learning often looks at those episodes and says, "No reward, no lesson."

Hindsight Experience Replay (HER) takes a smarter view. It asks: What did the agent actually achieve? Then it turns that supposedly failed attempt into useful training data.

HER helps goal-conditioned reinforcement learning agents learn from sparse rewards by relabeling failed trajectories with goals the agent did achieve. Instead of treating a missed target as worthless, HER reframes it as successful experience for a different, hindsight goal.

Why Sparse Rewards Cause Trouble

Sparse rewards sound clean on paper. You give the agent a reward only when it completes the task:

  • A robot receives +1 when it grasps an object.
  • A drone receives +1 when it reaches a waypoint.
  • A navigation agent receives +1 when it reaches the final destination.
  • Every other action earns 0.

Simple, right? Unfortunately, this setup often gives the agent almost no learning signal.

Imagine a robot arm that must place a cube at one exact location. The arm may spend thousands of episodes waving around, touching the cube, nudging it, or dropping it. If it never lands the cube precisely on target, it receives no success reward.

The agent does not know whether it nearly succeeded. It only knows that it failed. That feedback makes exploration painfully inefficient, especially in high-dimensional robotics tasks.

The Big HER Idea

HER changes the question from:

"Did the agent reach the original goal?"

to:

"Which goals did the agent reach during the attempt?"

Suppose you tell a robot to move a block to position A. It fails to reach A but moves the block to position B.

Normal replay stores the transition with goal A:

(s_t, a_t, r_t, s_{t+1}, g = A)

HER creates additional replay data by replacing the desired goal with B:

(s_t, a_t, r'_t, s_{t+1}, g' = B)

Under that new goal, the robot may count as successful. The transition now teaches the policy how to reach B, even though it failed at A.

That simple reframing gives the agent useful reward signals from experience that would otherwise look useless. HER does not erase failure; it extracts a lesson from it.

A Simple Analogy

Think about learning basketball. You aim for the hoop, miss, and hit the backboard.

A harsh coach might say, "You failed. Nothing happened." HER acts more like a useful coach: "You did not score, but you successfully hit the backboard. Now you know something about your throwing direction and force."

The original goal still matters. HER does not pretend the backboard equals the basket. It simply turns the attempt into training evidence for another reachable goal.

That distinction matters. HER improves exploration and replay efficiency, but it does not magically solve every difficult task.

How HER Works Step by Step

HER works with goal-conditioned policies. These policies receive both the current state and a desired goal.

A goal-conditioned policy looks like this:

pi(a | s, g)

The agent chooses an action a based on state s and goal g.

A goal-conditioned value function may look like:

Q(s, a, g)

The agent learns how valuable an action is for a particular goal.

Step 1: Run an Episode

The agent starts with a goal, such as moving a block to a target location.

During the episode, it gathers transitions:

(s_t, a_t, r_t, s_{t+1}, g)

The agent may fail completely according to the original task reward.

Step 2: Store the Original Experience

The replay buffer stores the normal transition with the original desired goal.

This part matters because the agent still needs to learn the actual task. HER supplements original data; it does not replace the intended goal with random excuses.

Step 3: Pick an Achieved Goal

HER selects an achieved state from the same episode. A common choice uses a future state from later in the trajectory.

For example, if a robot moved the block to coordinates [0.35, -0.10, 0.05], HER can use that position as a new goal.

Step 4: Relabel the Goal

HER replaces the old desired goal with the achieved goal:

g' = achieved_goal_{t'}

where t' often represents a future time step after t.

Step 5: Recompute the Reward

The environment calculates a new reward based on the relabeled goal.

If the state satisfies that new goal, the transition may receive a success reward. The agent now learns from an experience that looked like a failure under the original objective.

Step 6: Train an Off-Policy Agent

HER adds the original and relabeled transitions to replay. An off-policy RL algorithm samples both kinds of experience and improves its policy.

HER works with off-policy algorithms because those algorithms can learn from replayed transitions generated under older behavior policies. The original HER paper notes that HER can combine with any off-policy RL algorithm.

The "Future" Strategy

The most common HER relabeling strategy uses a future achieved goal. For each transition, HER chooses a state that the agent reaches later in the same episode and treats its achieved goal as the desired one.

Why choose a future goal instead of a past one?

Because the current action sequence actually helped lead toward that later achieved state. The relabeled training example stays more meaningful.

Imagine a robot starts at point X, moves toward Y, and ends at point Z. If HER relabels an early step with Z as the desired goal, the later actions in the trajectory demonstrate a path that reaches Z.

That creates a positive learning signal without inventing new physics or fake transitions. Clever, right?

Common HER strategies include:

Strategy Relabeled goal source Best use
Future A later achieved goal in the same episode Most common default
Final The final achieved goal Simple, lower diversity
Episode Any achieved goal from the same episode More variety
Random An achieved goal from another episode Broader coverage, less trajectory relevance

The original HER work explored replay strategies that draw goals from episode outcomes, and modern implementations commonly expose a future strategy as the default choice.

Why HER Works So Well

HER helps because many sparse-reward tasks contain useful progress even when the agent misses the final goal.

A robot may fail to grasp an object, but it may still learn how to:

  • Reach the object.
  • Move its gripper near the object.
  • Push the object.
  • Lift the object slightly.
  • Move the object across the table.

Without HER, the agent may receive zero reward for all of that behavior. With HER, it can treat reached states as temporary goals and learn controllable skills.

This process makes HER especially useful for multi-goal reinforcement learning. The agent learns a broader relationship between actions, states, and goals rather than memorizing one narrow target configuration.

Researchers describe HER as a method for off-policy goal-conditioned RL under sparse rewards, where failed rollouts become relabeled successes for goals that the agent actually achieved.

HER and Robot Manipulation

Robot manipulation gives HER a natural home because the desired goal often has a clear geometric form. Our robot arm grasping tutorial shows the task itself.

For a pick-and-place task:

  • Observation: robot joints, gripper position, object position, velocity.
  • Achieved goal: current object position.
  • Desired goal: target object position.
  • Reward: success when achieved and desired positions fall close enough together.

A goal-aware observation can look like this:

observation = {
    "observation": robot_state,
    "achieved_goal": object_position,
    "desired_goal": target_position,
}

The agent might initially fail to place the object at the target. However, it probably moves the object somewhere. HER treats those reached positions as valid hindsight goals and recomputes rewards accordingly.

Stable-Baselines3 expects precisely this structure for HER-compatible environments: a dictionary observation with observation, achieved_goal, and desired_goal keys, along with a vectorized compute_reward() implementation.

HER with DDPG, TD3, and SAC

HER does not replace an RL algorithm. It changes what goes into the replay buffer.

You typically pair HER with an off-policy algorithm such as:

  • DDPG: a classic choice for continuous control.
  • TD3: often more stable than DDPG because it reduces value overestimation.
  • SAC: a popular option that combines strong exploration with replay-buffer learning.
  • DQN: possible for discrete goal-conditioned tasks.
  • QR-DQN or other replay-based methods: possible if the goal-conditioned environment and action space fit.

For continuous robot manipulation, I usually prefer SAC + HER or TD3 + HER over vanilla DDPG. SAC's stochastic policy can help early exploration, while TD3 gives a strong deterministic-control baseline. Our SAC tutorial covers the algorithm in depth.

HER pairs poorly with standard on-policy methods such as PPO because PPO does not rely on a replay buffer in the same way. PPO learns from recent policy rollouts, while HER creates extra replay transitions by relabeling past goals — the DQN vs PPO comparison explains that off-policy/on-policy split.

Stable-Baselines3 documents HER as a replay-buffer class rather than a separate algorithm and lists compatibility with off-policy methods such as DQN, SAC, TD3, and DDPG.

Build a HER-Compatible Gymnasium Environment

A HER environment needs three observation components:

{
    "observation": ...,
    "achieved_goal": ...,
    "desired_goal": ...,
}

Here is a compact environment outline for a 2D point-reaching task — the same Gymnasium environment contract, plus the goal dictionary:

import gymnasium as gym
from gymnasium import spaces
import numpy as np


class PointReachEnv(gym.Env):
    def __init__(self):
        super().__init__()

        self.action_space = spaces.Box(
            low=-1.0,
            high=1.0,
            shape=(2,),
            dtype=np.float32,
        )

        self.observation_space = spaces.Dict(
            {
                "observation": spaces.Box(
                    low=-np.inf,
                    high=np.inf,
                    shape=(2,),
                    dtype=np.float32,
                ),
                "achieved_goal": spaces.Box(
                    low=-1.0,
                    high=1.0,
                    shape=(2,),
                    dtype=np.float32,
                ),
                "desired_goal": spaces.Box(
                    low=-1.0,
                    high=1.0,
                    shape=(2,),
                    dtype=np.float32,
                ),
            }
        )

        self.max_steps = 50
        self.step_count = 0

    def reset(self, seed=None, options=None):
        super().reset(seed=seed)

        self.position = self.np_random.uniform(
            -0.5,
            0.5,
            size=2,
        ).astype(np.float32)

        self.goal = self.np_random.uniform(
            -1.0,
            1.0,
            size=2,
        ).astype(np.float32)

        self.step_count = 0

        return self._get_obs(), {}

    def step(self, action):
        action = np.clip(
            action,
            -1.0,
            1.0,
        )

        self.position = np.clip(
            self.position + 0.05 * action,
            -1.0,
            1.0,
        )

        self.step_count += 1

        reward = self.compute_reward(
            self.position,
            self.goal,
            {},
        )

        terminated = np.linalg.norm(
            self.position - self.goal
        ) < 0.05

        truncated = self.step_count >= self.max_steps

        return (
            self._get_obs(),
            float(reward),
            terminated,
            truncated,
            {},
        )

    def compute_reward(
        self,
        achieved_goal,
        desired_goal,
        info,
    ):
        distance = np.linalg.norm(
            achieved_goal - desired_goal,
            axis=-1,
        )

        return -(distance > 0.05).astype(
            np.float32
        )

    def _get_obs(self):
        return {
            "observation": self.position.copy(),
            "achieved_goal": self.position.copy(),
            "desired_goal": self.goal.copy(),
        }

This environment uses sparse rewards:

  • 0 when the agent reaches the goal.
  • -1 when it misses.

That reward style mirrors common goal-based robotics environments. HER can transform failed point-reaching trajectories into successful examples for positions the agent did reach.

Train with Stable-Baselines3 HER

Stable-Baselines3 treats HER as HerReplayBuffer, not as a standalone algorithm. You attach that buffer to an off-policy agent.

Here is an example with SAC:

from stable_baselines3 import SAC
from stable_baselines3.her import HerReplayBuffer

env = PointReachEnv()

model = SAC(
    policy="MultiInputPolicy",
    env=env,
    replay_buffer_class=HerReplayBuffer,
    replay_buffer_kwargs={
        "n_sampled_goal": 4,
        "goal_selection_strategy": "future",
    },
    learning_rate=3e-4,
    buffer_size=100_000,
    batch_size=256,
    learning_starts=1_000,
    gamma=0.95,
    verbose=1,
)

model.learn(
    total_timesteps=200_000,
)

This setup creates four hindsight goal samples for each original transition, using future achieved goals from the same episode.

Use MultiInputPolicy because the observation uses a dictionary. Make sure the environment's compute_reward() handles batches of goals correctly, since HER recomputes rewards for many relabeled transitions at once. The Stable-Baselines3 tutorial covers more of the training workflow.

Choose HER Hyperparameters

HER adds a few important choices on top of your base RL algorithm.

Setting Good starting point Why it matters
n_sampled_goal 4 Sets how many relabeled goals HER samples
Goal strategy future Uses future achieved goals from the same trajectory
Base algorithm SAC or TD3 Strong off-policy choices for continuous control
Replay buffer size 100,000+ Stores enough full episodes for relabeling
Batch size 128–512 Supports stable off-policy updates
Success threshold Task-specific Defines when achieved and desired goals match
Episode length Enough for progress Gives the agent a chance to reach varied states

More hindsight goals can improve the density of useful learning signals, but too many can drown out original-goal experience. Start with four and compare against two or eight.

The right success threshold also matters. If your threshold is too strict, almost nothing counts as success. If it is too loose, the agent may earn success rewards while missing the task in any meaningful sense.

Common HER Mistakes

You Forget Goal Conditioning

HER cannot work if your policy does not receive the desired goal. The observation must include the goal, and the value function or policy must learn goal-dependent behavior.

A replay buffer alone cannot teach an agent to pursue different targets if the network never sees the target.

Your compute_reward() Cannot Handle Batches

HER recomputes rewards for multiple relabeled samples. Your reward function should handle inputs shaped like:

(batch_size, goal_dimension)

Avoid writing reward code that only works for one single state vector. Vectorize it with NumPy or PyTorch operations.

You Treat Sparse Rewards as Dense Rewards

HER works especially well when success has a clear binary or sparse definition. You can use dense rewards too, but reward recomputation must stay valid after you replace the desired goal.

If your reward depends on hidden task history, time-dependent rules, or external events, relabeling may become invalid. The reward function design guide details more shaping pitfalls.

You Use HER with PPO

PPO does not use replay data in the normal off-policy sense. HER depends on replaying and relabeling stored episode transitions, so pair it with DQN, DDPG, TD3, SAC, or another compatible off-policy method.

Your Agent Never Reaches Interesting States

HER cannot relabel goals that the agent never achieves. If the agent only sits still, all hindsight goals look like "stay still."

Improve initial exploration, simplify the task, use action noise or SAC entropy, shorten distances at first, or add a curriculum. HER learns from failures, but it still needs failures that contain some meaningful movement.

When HER Will Not Save You

HER offers a powerful solution for many sparse-reward goal tasks, but it has limits.

It may struggle when:

  • Goals lack a clear measurable representation.
  • Rewards depend heavily on hidden state or complex history.
  • The agent never reaches diverse states.
  • The environment requires long, precise action sequences.
  • Goal relabeling changes the meaning of the task.
  • Your problem lacks naturally interchangeable goals.

For example, consider a game where an agent must collect keys in a strict order before opening a door. Relabeling a later location as the goal may not preserve the task's actual rules. HER works best when goal achievement depends mainly on state similarity, such as distance between an object's current and desired position.

Want your robot to learn from every miss? Grab the GPTAstra full course at https://cutt.ly/5yviN6qd — it implements HER with SAC end to end, from goal dictionaries to relabeled replay.

Frequently Asked Questions

What is Hindsight Experience Replay (HER)?

HER is a relabeling technique for goal-conditioned reinforcement learning. It takes trajectories that failed under the original goal and rewrites them with goals the agent actually achieved, so failed attempts become successful training examples for a different goal.

Why do sparse rewards need HER?

With sparse rewards the agent receives signal only on success, so thousands of near-misses teach nothing. HER converts those near-misses into reward-bearing transitions by relabeling achieved states as goals, dramatically improving learning efficiency.

Which relabeling strategies does HER use?

Future (use a later achieved goal from the same episode — the default and most common), Final (use the episode's final state), Episode (any achieved goal from the episode), and Random (an achieved goal from a different episode).

Which RL algorithms work with HER?

Off-policy, replay-based methods: DDPG, TD3, SAC, DQN, and other replay algorithms. HER does not pair with on-policy methods like PPO, which learn from fresh rollouts instead of relabeled replay data.

What observation format does a HER environment need?

A dictionary observation with observation, achieved_goal, and desired_goal keys, plus a vectorized compute_reward(achieved_goal, desired_goal, info) that works on batches — HER recomputes rewards for many relabeled samples at once.

When does HER not help?

When goals lack a measurable representation, rewards depend on hidden history, the agent never reaches diverse states, tasks need long precise sequences, or relabeling changes the task's meaning — such as games that require collecting keys in a fixed order.

Final Thoughts

Hindsight Experience Replay turns failed attempts into useful lessons by relabeling achieved states as alternative goals. It gives sparse-reward agents a way to learn from movement, contact, and partial progress even when they miss the original objective.

Use HER for goal-conditioned tasks with clear achieved-goal and desired-goal representations. Pair it with an off-policy method such as SAC, TD3, or DDPG. Start with the future strategy, four sampled goals, and a clean vectorized reward function.

Most importantly, do not view an unsuccessful trajectory as wasted data. Your robot may fail to place the block where you asked — but it still placed it somewhere. With HER, that "somewhere" becomes the next lesson.

Share this article X Facebook LinkedIn Reddit WhatsApp

Related Articles