Contents
Want an RL agent that learns to steer, balance, reach, or control a robot arm without face-planting into the same bad action for a million steps? Soft Actor-Critic (SAC) gives you one of the most practical tools for continuous control.
I like SAC because it takes exploration seriously. Instead of forcing an agent to choose between "play it safe" and "try something weird," SAC says: why not reward both smart behavior and healthy curiosity? Revolutionary, right?
The short version: SAC is the algorithm I reach for whenever the action space is continuous and every environment step costs real time — and by the end of this tutorial you'll know exactly why it works, how to configure it, and what to check when it doesn't.
Figure 1: SAC — twin critics, an entropy bonus, and a replay buffer that remembers everything
What Is Soft Actor-Critic?
Soft Actor-Critic, usually shortened to SAC, is a model-free, off-policy reinforcement learning algorithm built for continuous action spaces. Think actions such as a robot's joint torque, a car's steering angle, or a drone's throttle — not just "left" or "right."
SAC combines two familiar ideas:
- An actor that selects actions.
- A critic that judges how useful those actions look.
The "soft" part comes from SAC's maximum-entropy objective. SAC wants high reward, obviously, but it also wants the policy to retain some randomness. That extra randomness encourages exploration and helps the agent avoid locking itself into a mediocre strategy too early.
Ever watched an agent discover one barely functional trick and repeat it forever? SAC tries hard to prevent that charming little disaster.
Why SAC Fits Continuous Control
Continuous-control problems ask an agent to choose from infinitely many possible actions. A robotic arm can apply 0.1, 0.11, or 0.111 units of torque, for example. That flexibility makes these environments powerful — and annoying to solve.
SAC handles this setting well because it learns a stochastic policy, commonly represented by a Gaussian distribution. Rather than outputting one rigid action, the actor predicts a distribution and samples actions from it during training.
That design gives SAC several practical strengths:
- Efficient replay usage: SAC learns off-policy, so it can reuse past experience from a replay buffer.
- Strong exploration: entropy rewards policy randomness instead of relying only on manually added noise.
- More stable value learning: SAC uses two critics to reduce overly optimistic value estimates.
- Useful for high-dimensional actions: it works well when an agent controls several joints, motors, or continuous parameters.
IMO, that combination explains why SAC often becomes a default first choice for MuJoCo-style robotics benchmarks and many continuous-control experiments.
The Core SAC Objective
Traditional reinforcement learning asks the agent to maximize cumulative reward:
Standard RL: J = E[ Σ_t r(s_t, a_t) ]
SAC adds an entropy bonus:
SAC: J(π) = E[ Σ_t ( r(s_t, a_t) + α · H(π(·|s_t)) ) ]
Here, r(s_t, a_t) means the reward after the agent takes action a_t in state s_t. The entropy term H measures how random or spread out the policy remains, while α controls how much SAC values that randomness.
Why Entropy Matters
Imagine you train a robot to push a block. Early on, the robot finds one direction that earns a tiny reward. A standard algorithm might cling to that move like it discovered the meaning of life.
SAC keeps the policy flexible for longer. The agent still chases reward, but it also tests nearby actions and alternative approaches. That process often helps it find better strategies instead of camping at the first local optimum.
Entropy does not mean "act randomly forever." It means SAC balances exploration and exploitation in a deliberate, mathematical way.
The Four Main SAC Pieces
SAC sounds intimidating until you break it into parts. You need four major components:
1. The Actor
The actor receives a state s and produces an action distribution π(a|s). In common continuous-control implementations, it outputs:
- A mean action vector.
- A log standard deviation vector.
- A sampled action, often transformed with tanh to stay inside the environment's action limits.
The actor learns to choose actions that earn high value according to the critics while preserving enough entropy for exploration.
2. Twin Critics
SAC uses two Q-functions, Q₁(s,a) and Q₂(s,a). Each critic estimates the expected long-term return for a state-action pair.
Why two critics? Because one Q-network can become overly optimistic. SAC uses the smaller of the two estimates in its targets, which limits overestimation bias. No algorithm enjoys confidence without evidence — except maybe that one teammate who says "it works on my machine."
3. Target Critics
SAC maintains slowly updated copies of its critics. These target networks generate more stable learning targets and stop the agent from chasing a value estimate that changes wildly every gradient step.
Most implementations use a soft update:
θ̄ ← τ · θ + (1 − τ) · θ̄ # τ ≈ 0.005
A small τ, such as 0.005, makes the target network move gradually rather than lurch around.
4. Replay Buffer
The replay buffer stores transitions:
(s, a, r, s′, d)
These terms represent the current state, action, reward, next state, and terminal flag. SAC samples random batches from this memory during training, which improves data efficiency and reduces correlations between sequential experiences.
How SAC Training Works
Here's the usual SAC training loop in plain English.
- Initialize the actor, two critics, target critics, and replay buffer.
- Let the actor interact with the environment and collect transitions.
- Add each transition to the replay buffer.
- Sample a random batch of past transitions.
- Update both critics using entropy-adjusted target values.
- Update the actor so it selects high-value, sufficiently diverse actions.
- Update the entropy temperature if you use automatic tuning.
- Soft-update the target critics.
- Repeat until your reward curve stops looking like abstract art.
SAC usually starts with random actions for a short warm-up phase. Those initial experiences populate the replay buffer before the policy starts making confident choices based on approximately zero useful knowledge. FYI, this warm-up can save you from unstable early learning.
The Critic Update
The critic learns from a Bellman target that includes both future value and entropy. SAC typically computes a target like this:
y = r + γ(1 − d) · [ min(Q̄₁(s′, a′), Q̄₂(s′, a′)) − α · log π(a′|s′) ]
The critic then minimizes the difference between its prediction and that target.
A few pieces matter here:
γdiscounts future rewards.dprevents bootstrapping after terminal states.min(Q̄₁, Q̄₂)keeps value estimates conservative — the pessimism is the point.−α · log π(a′|s′)adds the entropy incentive.
SAC uses this entropy-regularized backup because the algorithm optimizes reward and policy randomness together, not as two unrelated tricks glued together at 2 a.m.
The Actor Update
The actor wants actions that score well under the critics while retaining entropy. Its loss commonly looks like:
J_actor = E_{s∼D, a∼π}[ α · log π(a|s) − min(Q₁(s, a), Q₂(s, a)) ]
The actor minimizes this loss. In practical terms, it increases the probability of actions with high Q-values and avoids collapsing into a completely deterministic policy too early.
Most SAC implementations use the reparameterization trick to backpropagate through sampled actions. The actor samples Gaussian noise, combines it with its predicted mean and standard deviation, and lets gradients flow through the resulting action.
Automatic Entropy Tuning
Choosing α by hand can feel annoying because it is annoying. Set it too high and your agent stays chaotic; set it too low and it becomes overly cautious before it has learned much.
Modern SAC implementations often use automatic temperature tuning. The algorithm learns α and aims for a target entropy, commonly related to the number of action dimensions.
Automatic tuning offers a major convenience:
- Early training: SAC often values exploration more.
- Later training: SAC can reduce unnecessary randomness as useful behavior emerges.
- Across tasks: you avoid retuning entropy weight for every environment.
For most projects, start with automatic tuning. You can always take manual control later if your experiment demands it.
Practical SAC Hyperparameters
You do not need to worship a magic hyperparameter list, but solid defaults make life easier.
| Hyperparameter | Common starting point | Why it matters |
|---|---|---|
| Learning rate | 3e-4 | Controls update size for actor and critics |
| Discount factor γ | 0.99 | Values future rewards |
| Replay buffer size | 1,000,000 | Stores diverse experience |
| Batch size | 256 | Balances stable gradients and compute cost |
| Target update rate τ | 0.005 | Smooths target critic updates |
| Initial random steps | 5,000–10,000 | Fills the replay buffer |
| Entropy coefficient | Automatic | Adapts exploration pressure |
These defaults resemble common SAC setups, but your environment still gets the final vote. Sparse rewards, action scaling, episode length, and observation noise can all change the answer.
Normalize What You Can
For continuous-control tasks, normalize observations when their scales vary wildly. Also ensure that your action outputs match the environment's valid range.
A common setup samples an action through tanh, which keeps values between −1 and 1, then rescales them to the environment's action bounds. If you skip that rescaling step, your agent may spend hours issuing nonsense commands. Very scientific.
Treat Time Limits Carefully
Some environments end an episode because the agent truly failed. Others end because a fixed time horizon expired. Those cases need different treatment in the critic target.
If the environment only hits a time limit, you often want SAC to bootstrap from the next state rather than treat it as a true terminal state. This small implementation detail can noticeably affect learning quality.
SAC vs. TD3 vs. PPO
SAC does not win every scenario, but it gives you an excellent baseline for continuous control.
| Algorithm | Policy type | Data usage | Exploration style | Best fit |
|---|---|---|---|---|
| SAC | Stochastic | Off-policy | Entropy-driven | Sample-efficient continuous control |
| TD3 | Deterministic | Off-policy | Added action noise | Simpler continuous-control setups |
| PPO | Stochastic | On-policy | Policy sampling | Stable, parallel data collection |
SAC and TD3 often perform strongly in MuJoCo-style continuous-control experiments, while PPO can show less predictable performance in those comparisons.
I usually choose SAC when environment interaction costs matter. Replay-buffer reuse gives SAC a huge practical edge over on-policy methods such as PPO, which discard most collected data after a small number of updates. On the other hand, PPO can feel easier to debug in distributed setups and can perform well with the right environment and tuning.
Common SAC Problems
The Reward Never Improves
First, verify the basics �� most "mysterious failures" are reward design bugs in disguise:
- Check reward signs and scales.
- Confirm action ranges match the environment.
- Plot action values and critic losses.
- Ensure the replay buffer actually receives transitions.
- Confirm terminal flags behave correctly.
Many "algorithm failures" turn out to be a reward bug or action-scaling mistake wearing a fake mustache.
Q-Values Explode
Exploding Q-values often point to an unstable reward scale, an overly high learning rate, incorrect target computation, or numerical issues in the log-probability calculation.
Try reward scaling, lower learning rates, gradient clipping, and careful handling of tanh-squashed action log probabilities. Also check that you use target critics — not online critics — for bootstrapped targets.
The Policy Stays Too Random
If your agent keeps wobbling after learning, inspect entropy tuning. A high α or an overly ambitious target entropy can keep exploration too strong.
During evaluation, choose deterministic actions when you want a clean measure of performance. Stable-Baselines3 specifically recommends deterministic inference for continuous-control SAC tasks.
A Simple SAC Workflow
Want a practical starting plan? Use this:
- Pick a standard continuous-control environment, such as Pendulum or a MuJoCo locomotion task.
- Normalize observations and confirm action bounds.
- Start with standard SAC defaults and automatic entropy tuning.
- Run several random seeds instead of trusting one lucky curve.
- Evaluate the learned policy deterministically every so often.
- Log rewards, actor loss, critic loss, entropy, and Q-values.
- Change one variable at a time when results look weird.
That final rule matters more than people admit. If you modify architecture, reward shaping, learning rate, and exploration settings at once, you will learn exactly nothing from the experiment.
Recommended Books
- Foundations of Deep Reinforcement Learning by Laura Graesser and Waail Wahba — the clearest book-length treatment of the actor-critic family this tutorial compresses, with SAC and TD3 worked through side by side.
- Algorithms for Reinforcement Learning by Csaba Szepesvári — mathematical rigor for the entropy-regularized ideas underneath SAC, short enough to read in a weekend.
- Reinforcement Learning: An Introduction by Richard S. Sutton and Andrew G. Barto — the canonical backup; Chapter 13's entropy-regularized RL is where SAC's objective comes from.
Unlock AI That Actually Works
Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.
Click here to get GPTAstra Max now — one-time payment, lifetime access.
Frequently Asked Questions
What is Soft Actor-Critic (SAC)?
SAC is a model-free, off-policy reinforcement learning algorithm built for continuous action spaces such as robot joint torques, steering angles, or drone throttles. It combines an actor that selects actions with twin critics that judge them, and it maximizes reward plus an entropy bonus so the policy keeps exploring instead of locking into one mediocre trick.
What does the 'soft' in Soft Actor-Critic mean?
It refers to the maximum-entropy objective. Alongside cumulative reward, SAC maximizes the entropy H of the policy — how random or spread out its action distribution stays — weighted by a temperature α. Entropy doesn't mean acting randomly forever; it means balancing exploration and exploitation deliberately.
Why does SAC use two critics?
A single Q-network can become overly optimistic about state-action pairs. SAC keeps two Q-functions and uses the smaller of the two estimates in its targets, which limits overestimation bias and makes value learning far more stable.
How do you set the entropy coefficient alpha?
Don't set it by hand. Modern SAC implementations use automatic entropy tuning: the algorithm learns α toward a target entropy related to the number of action dimensions, so exploration pressure adapts early in training and relaxes as useful behavior emerges — with no retuning across tasks.
When should I use SAC vs TD3 vs PPO?
SAC when environment interaction is expensive — its replay buffer reuses past experience for sample-efficient continuous control. TD3 for simpler deterministic continuous-control setups. PPO when you want on-policy stability and easy parallel data collection, and debugging simplicity matters more than sample efficiency.
Why is my SAC agent not learning?
Check the fundamentals first: reward signs and scales, action ranges matching the environment, transitions actually landing in the replay buffer, and correct terminal flags. If Q-values explode, lower the learning rate, scale rewards, clip gradients, and verify you bootstrap from target critics with properly computed tanh-squashed log probabilities.
Final Thoughts
Soft Actor-Critic gives continuous-control agents a strong mix of exploration, stability, and sample efficiency. Its stochastic actor searches broadly, its twin critics curb overconfidence, and its replay buffer squeezes more learning from every interaction.
Start simple, inspect your action scaling, and let automatic entropy tuning handle the exploration balance. Once SAC clicks, watching an agent move from chaotic flailing to smooth control feels oddly satisfying — like teaching a toddler calculus, except the toddler owns two neural networks. That first smooth walk across the sandbox is worth every debugging session.
Related Articles
- PPO Explained: Train Robust Game AI with Proximal Policy Optimization
- Stable-Baselines3 Tutorial: Train RL Agents in Minutes (2026)
- MuJoCo Tutorial: Physics Simulation for Robotics RL (2026)
- Training a Robot Arm to Grasp Objects with RL (2026): Complete Tutorial
- Reward Function Design for Robotics RL (2026): Complete Guide