Contents
Here's a sentence that should genuinely terrify you a little: your robot will do exactly what you told it to, not what you meant. Every "robot arm technically satisfies its reward function while doing something bizarre" story you've ever heard traces back to this exact gap, and reward design is the entire discipline of closing it.
I've referenced this problem obliquely across half the tutorials in this series — the grasping arm's sparse rewards, BipedalWalker's torque penalty, Snake's death-versus-eating balance. This article is where we actually stop and treat reward design as its own subject instead of a footnote inside every other project. Ever wondered why so many RL horror stories involve an agent "cheating" its way to high reward? This is exactly the topic that explains why, and more importantly, how to avoid it.
By the end of this guide, you'll understand the core tensions in reward design, the technique that lets you shape rewards without breaking your task's actual definition of success, and the failure patterns worth checking for before you ever hit "train." IMO, reward design is genuinely the highest-leverage skill in applied RL — better than any hyperparameter tweak :)
The Central Tension: Sparse vs. Shaped Rewards
Every reward function design decision ultimately traces back to one tradeoff. Sparse rewards give infrequent, unambiguous feedback — a clean +1 for success, 0 otherwise. Shaped rewards add intermediate feedback throughout the task, guiding the agent step by step rather than only at the end.
Sparse rewards are honest. There's no ambiguity about what "success" means — the agent either did the thing or it didn't. But as you saw with robot grasping, sparse rewards can make learning nearly impossible if success is rare enough that random exploration never stumbles into it. Shaped rewards accelerate learning by rewarding progress toward the goal, not just goal completion itself. In grasping, that might mean rewarding a decreasing distance between gripper and object, not just the final successful grasp. The tradeoff is real: shaped rewards risk teaching the agent something subtly different from what you actually wanted, while pure sparse rewards risk teaching it nothing at all within a reasonable training budget.
Striking the right balance between the two is genuinely the core skill here — not eliminating one in favor of the other, but understanding which situations call for more of each.
Our robot arm grasping tutorial demonstrates the sparse reward problem directly — HER exists because sparse rewards made grasping unlearnable through standard approaches.
Figure 1: Reward design is where robotics RL lives or dies — more training time rarely fixes a reward function that's rewarding the wrong thing
Reward Hacking: When Your Agent Outsmarts Your Intent
This is the failure mode that makes reward design genuinely hard rather than just tedious. Reward hacking happens when an agent finds unintended shortcuts that maximize reward without accomplishing the actual goal you had in mind.
Remember BipedalWalker's torque penalty from earlier in this series? Remove it, and agents can find degenerate, high-energy solutions that technically move forward through violent, inefficient motion rather than anything resembling walking. In grasping tasks, an agent rewarded purely for "gripper close to object" might learn to hover permanently near the object without ever actually completing a grasp — technically satisfying a proxy for success, while dodging the actual task. This isn't a bug in the algorithm — it's the algorithm working exactly as designed. RL agents are relentless optimizers of whatever number you hand them; they have no innate concept of what you "actually meant."
The fix is rarely "add more reward terms." It's usually going back and asking whether your reward function has a loophole a sufficiently persistent optimizer could exploit — because if one exists, a well-trained agent will eventually find it.
Our BipedalWalker tutorial demonstrates this directly — the torque penalty exists specifically to prevent the degenerate high-energy solutions that reward hacking produces.
Potential-Based Reward Shaping: Getting Shaped Rewards Without Breaking Things
Here's the genuinely elegant piece of theory that makes reward shaping safe rather than a total gamble: potential-based reward shaping, which provides intermediate feedback while provably preserving the original task's optimal policy.
The core idea: define a potential function Φ(s) that measures how "good" or "close to goal" a given state is. Then, instead of just using the raw sparse reward, add a shaping term based on the change in potential between states:
shaped_reward = original_reward + γ * Φ(next_state) - Φ(state)
In a robot soccer example, Φ(s) might be the negative distance from the ball to the goal — the agent gets a small positive nudge every time it moves the ball closer, not just when it finally scores. The mathematical guarantee here matters enormously: because the shaping term is defined as a difference in potential rather than an arbitrary bonus, it doesn't change which policy is actually optimal — it only changes how fast the agent finds it. This is genuinely different from naively bolting on extra reward terms, which can (and often does) shift what the "best" policy actually looks like, sometimes toward something you didn't intend at all.
This is the technique worth reaching for whenever you're tempted to add ad-hoc bonus rewards. It gives you the training-speed benefits of dense feedback without the reward-hacking risk that comes from guessing at arbitrary bonus terms.
For a deeper understanding of how reward design affects algorithm performance, our evaluating reinforcement learning algorithms guide covers the metrics used to assess whether your reward function is actually producing the behavior you intended.
A Practical Framework: Composite Reward Design
Real robotics reward functions are rarely a single term — they're typically a weighted combination of several signals, each capturing a different aspect of what "good" behavior looks like.
For a robotic grasping task, a genuinely reasonable composite reward might combine:
Task success (sparse): a large positive reward for the actual completed grasp — this stays sparse and unambiguous, exactly as it should. Progress shaping (potential-based): smaller rewards for decreasing gripper-to-object distance, using the potential-based approach above to avoid distorting the actual objective. Effort penalty: a small negative term proportional to force or torque used, encouraging efficient movement rather than flailing — directly analogous to BipedalWalker's motor torque penalty. Safety constraints: penalties for excessive force that could damage the object or the robot itself — genuinely more important here than in any purely simulated game environment, since real hardware and real objects are at stake.
The composite weights matter enormously, and getting them wrong is exactly how you end up with an agent that's technically optimizing your function while doing something you didn't want. Expert-designed composite rewards typically combine these signals through weighted sums, though more sophisticated aggregation methods exist for cases where a simple weighted sum doesn't capture the right tradeoffs.
Dealing With Sparse and Delayed Rewards
A specific version of the sparse reward problem worth naming directly: delayed rewards, where feedback only arrives after a long sequence of actions — reaching the end of a maze, completing a multi-step assembly task.
The agent has to correctly attribute credit across many timesteps to whichever earlier actions actually contributed to eventual success — this is the credit assignment problem, and it gets genuinely harder as delay increases. This is exactly the scenario Hindsight Experience Replay (from the robot grasping tutorial) and curriculum learning (from that dedicated article) both exist to address — different angles on the same underlying "the agent can't learn from a reward signal it almost never encounters" problem. Reward shaping and these structural techniques aren't competing solutions — they're complementary. A well-shaped reward reduces how badly you need HER or curriculum learning; a well-designed curriculum reduces how much shaping you need to hand-craft.
Our curriculum learning guide covers the complementary approach of starting with easier tasks — understanding both techniques gives you two powerful tools for the same underlying sparse-reward problem.
Dynamic and Adaptive Reward Shaping
Static reward functions, hand-tuned once at the start, aren't always the best fit for training that spans a huge range of agent competence.
Dynamic potential-based shaping allows the shaping function itself to change over time during training, rather than staying fixed — though this genuinely still requires domain knowledge to design the potential function correctly. More recent approaches use adaptive auxiliary rewards that scale based on the agent's current capability — essentially, providing more scaffolding early in training when the agent needs it most, and gradually reducing that scaffolding as competence grows. This mirrors the "training wheels" intuition directly — useful when learning to balance, unnecessary and potentially limiting once you've actually learned to ride. Intrinsically motivated approaches let auxiliary rewards emerge from the agent's own learning progress, rather than requiring a human to hand-specify every shaping term upfront — a genuinely different philosophy from manually engineering every reward component.
The Divide-and-Conquer Problem in Reward Tuning
Here's a genuinely frustrating pattern worth naming explicitly, because you will hit it: tuning a reward function to work correctly across a representative set of training environments often means fixing behavior in one scenario while accidentally breaking it in another.
You adjust a term to fix an issue in scenario A, only to discover scenario B now produces unwanted behavior that wasn't a problem before your change. This iterative "fix one thing, break another" cycle is a widely recognized frustration in reward engineering, not a sign you're doing something uniquely wrong. The practical takeaway: budget real time for this iteration cycle, and test your reward function against a genuinely diverse set of scenarios before considering it stable — a reward function that looks perfect in one test case can hide problems that only surface elsewhere.
Common Mistakes People Make
I've referenced several of these across other tutorials in this series, but worth collecting them here directly since they're specifically reward-design failures.
Rewarding a proxy instead of the actual goal. "Distance traveled" without a stability or efficiency penalty invites exactly the kind of degenerate, technically-correct-but-wrong solutions BipedalWalker demonstrated. Adding ad-hoc bonus terms without checking for exploitable loopholes. Any reward with an unconsidered edge case is an invitation for a sufficiently trained agent to find and exploit it. Confusing "more shaping" with "better shaping." A densely shaped reward with poorly considered terms can genuinely slow learning compared to a cleaner, sparser signal — density alone isn't the goal. Ignoring safety and effort terms in real robotics tasks. Unlike simulated games, real hardware faces real wear, real cost, and real risk of damage — reward functions that ignore this optimize for outcomes that could be actively harmful to run on physical hardware. Treating reward design as a one-time task. Given how often fixing one scenario breaks another, expect genuine iteration — a reward function is rarely right on the first attempt.
A Practical Design Checklist
If you're building a reward function for a new robotics task, here's a sensible sequence to actually follow.
Start by defining the sparse, unambiguous success signal first — what does genuine task completion actually look like, expressed as cleanly as possible? Identify where sparsity will genuinely block learning — if success is rare enough that random exploration won't find it in a reasonable budget, that's your signal to add shaping. Use potential-based shaping wherever possible, specifically because it preserves your original optimal policy rather than risking a subtly different one. Add effort and safety penalties appropriate to real hardware constraints, not just simulation convenience. Test against a genuinely diverse set of scenarios, watching specifically for reward hacking — behavior that scores well but looks wrong when you actually watch it.
Our MuJoCo tutorial covers continuous control reward design in practice — understanding how Ant and HalfCheetah rewards work makes the principles here feel concrete rather than abstract.
Wrapping This Up
Reward function design is where robotics RL genuinely lives or dies — more training time or a fancier algorithm rarely fixes a reward function that's rewarding the wrong thing. Sparse rewards give unambiguous signal but can block learning entirely; shaped rewards accelerate learning but risk teaching something subtly different from your actual intent, unless you lean on techniques like potential-based shaping that mathematically preserve the original objective.
Remember that reward hacking isn't a bug — it's the expected behavior of a relentless optimizer given an imperfect proxy for what you actually want, and that composite reward functions blending task success, progress shaping, and effort/safety penalties genuinely need iterative testing across diverse scenarios before you can trust them. FYI, if you've ever watched a trained agent do something technically reward-maximizing but visibly wrong, you've witnessed reward hacking firsthand — and going back to fix the reward function, not adding more training time, is almost always the actual fix :)
Now go back through the grasping, BipedalWalker, or Snake reward functions from earlier in this series and ask what a maximally uncooperative agent could exploit in each one. That adversarial mindset is genuinely the fastest way to spot reward design flaws before they cost you a multi-hour training run.