Contents
Somewhere between Snake's grid-based simplicity and BipedalWalker's balance problem sits a genuinely different flavor of RL challenge: teaching an agent to drive fast without spinning out, cutting corners in ways that "technically" satisfy the reward function, or driving in nervous little circles forever because that felt safer than committing to the track. Racing car RL is where reward design, continuous control, and visual perception all show up in the same project at once.
I find this project genuinely satisfying for a specific reason — unlike BipedalWalker's slow crawl toward "not falling over," a racing agent's progress is immediately, viscerally obvious. Watching it go from spinning off the first corner to smoothly clipping apexes lap after lap is a completely different kind of feedback loop than watching a reward curve inch upward.
By the end of this tutorial, you'll understand the specific design decisions that make racing genuinely different from other continuous-control RL problems, and you'll have a trained agent actually completing laps. IMO, this is one of the more genuinely fun applications of everything you've learned across BipedalWalker, Atari, and MuJoCo combined into one project :)
What Makes Racing Genuinely Different
Racing sits at an interesting intersection of problems you've already tackled separately in this series. It borrows visual perception from Atari, continuous control from BipedalWalker, and reward-hacking risk from the reward design article — all in one environment.
Continuous steering and throttle control — like BipedalWalker's joint torques, but now mapped to a car's actual handling dynamics rather than abstract joint angles. Visual or lidar-style state representation — the agent needs to perceive the track ahead, whether through raw pixels (à la Atari) or distance sensors (à la BipedalWalker's lidar readings). A genuinely rich reward-hacking surface — reward "distance traveled" carelessly, and you'll get an agent that drives in tight, low-risk circles rather than actually racing forward, exactly the kind of proxy-gaming behavior the reward design article warned about.
Ever wondered why racing AI demos often look either genuinely impressive or comically broken, with nothing in between? Reward design is almost always the deciding factor, more than the underlying algorithm choice.
Setting Up: CarRacing-v3
Gymnasium's built-in CarRacing-v3 environment is the standard entry point — a top-down, procedurally generated race track rendered as pixels, genuinely similar in spirit to a simplified racing game.
pip install "gymnasium[box2d]" stable-baselines3 torch
import gymnasium as gym
env = gym.make("CarRacing-v3", render_mode="human")
print("Observation space:", env.observation_space)
print("Action space:", env.action_space)
observation, info = env.reset(seed=42)
Observation space: a 96×96 RGB image — a top-down view of the car and the track ahead, genuinely comparable to the preprocessed frames from the Atari tutorial. Action space: 3 continuous values — steering (-1 to 1), gas (0 to 1), and brake (0 to 1). Reward: a small negative reward per frame (encouraging speed), plus positive rewards for visiting new track tiles, and a penalty for driving off the track entirely.
Notice this reward structure is already doing real work to prevent reward hacking — a per-frame penalty discourages just sitting still, while rewarding new tiles (not just raw distance or speed) discourages spinning in circles on already-visited track sections.
Preprocessing: Borrowing From Atari
Since CarRacing's observations are raw pixels, the same preprocessing pipeline from the Atari DQN tutorial applies almost directly.
import cv2
import numpy as np
def preprocess_frame(frame):
gray = cv2.cvtColor(frame, cv2.COLOR_RGB2GRAY)
normalized = gray / 255.0
return normalized
Grayscale conversion works here for the same reason it worked for Atari — color rarely carries meaningful information for judging track position and upcoming curves, and stripping it out reduces the network's workload without meaningfully hurting performance.
Frame Stacking for Velocity Perception
Just like Atari, a single static frame can't convey how fast the car is moving or in which direction.
from gymnasium.wrappers import FrameStackObservation
env = gym.make("CarRacing-v3", render_mode="rgb_array")
env = FrameStackObservation(env, stack_size=4)
This is genuinely the identical motion-perception problem from the Atari tutorial, solved the identical way — stacking consecutive frames so velocity and direction become visible through comparison across the stack, rather than needing to be inferred from a single frozen image.
Our Atari DQN tutorial covers the original CNN-based visual RL foundation — understanding frame stacking and grayscale preprocessing there makes this racing setup feel familiar rather than new.
Choosing the Algorithm: PPO for Continuous Control
Since steering, gas, and brake are all continuous values, DQN is immediately off the table — same reasoning as BipedalWalker and MuJoCo. PPO is the standard, forgiving choice here.
from stable_baselines3 import PPO
from stable_baselines3.common.env_util import make_vec_env
env = make_vec_env("CarRacing-v3", n_envs=4)
model = PPO(
"CnnPolicy",
env,
verbose=1,
learning_rate=3e-4,
n_steps=2048,
batch_size=64,
n_epochs=10,
gamma=0.99,
)
model.learn(total_timesteps=1_000_000)
model.save("car_racing_ppo")
Notice "CnnPolicy" instead of "MlpPolicy" — this is genuinely important distinction from BipedalWalker's flat 24-value state. Since observations here are images, the policy needs convolutional layers to actually extract useful spatial features, exactly like the Atari DQN's architecture, just wrapped inside PPO's continuous-action framework instead.
Our BipedalWalker tutorial explains why continuous-action algorithms like PPO and SAC are necessary when control signals are continuous — the same logic applies here, with steering and throttle instead of joint torques.
Watching Training Progress
Racing agents follow a recognizable arc, similar in spirit to BipedalWalker's progression but with its own specific failure modes worth watching for.
Early training: the car drives off the track almost immediately, often spinning out on the very first turn or failing to accelerate meaningfully at all. Middle training: the agent learns to stay on track through overly cautious driving — creeping along at low speed, over-braking into every corner. Later training: genuine racing lines emerge — smoother cornering, appropriate speed modulation, and consistent lap completion.
That "overly cautious" middle phase is genuinely the most common sticking point. If your agent gets permanently stuck driving slowly and safely rather than progressing toward genuine speed, that's usually a signal your reward function is over-penalizing risk relative to rewarding progress — worth revisiting the per-frame penalty and tile-completion reward balance.
Reward Shaping Specific to Racing
Building directly on the reward design principles from earlier in this series, a few racing-specific considerations are worth calling out.
Penalize going off-track meaningfully, but not so harshly that the agent becomes pathologically risk-averse. Too severe a penalty, and you get an agent that crawls at minimum viable speed to avoid any chance of leaving the track. Reward genuine forward progress along the track, not just raw velocity. An agent rewarded for pure speed might floor the throttle into a wall — rewarding progress along the actual track centerline avoids this specific failure mode. Consider a small penalty for excessive steering changes, similar in spirit to BipedalWalker's torque penalty — this discourages jittery, unstable steering behavior in favor of smoother, more realistic driving lines.
This is potential-based shaping in practice — rewarding progress toward "completing more of the track" rather than an arbitrary proxy, following exactly the principle from the reward design article about avoiding proxies a clever agent could exploit.
Our reward function design guide covers the full taxonomy of reward shaping mistakes — understanding those pitfalls makes racing-specific reward design feel like a concrete application rather than guesswork.
Figure 1: Racing car RL combines CNN-based visual perception, continuous control, and careful reward design — one of the most genuinely fun multi-discipline projects in this series
Watching Your Trained Agent
model = PPO.load("car_racing_ppo")
env = gym.make("CarRacing-v3", render_mode="human")
env = FrameStackObservation(env, stack_size=4)
observation, info = env.reset()
for _ in range(1000):
action, _ = model.predict(observation, deterministic=True)
observation, reward, terminated, truncated, info = env.step(action)
if terminated or truncated:
observation, info = env.reset()
env.close()
deterministic=True matters here just like every other tutorial in this series — you want to see the actual learned racing line, not exploration noise mixed into the steering input.
Our CartPole DQN tutorial explains why deterministic evaluation matters for understanding what an agent actually learned versus what exploration noise adds — the same logic applies directly to watching your trained racer.
Alternative Environments Worth Knowing
CarRacing-v3 is the standard beginner entry point, but a few other options exist depending on what specifically interests you.
TORCS (The Open Racing Car Simulator) — a more realistic, full 3D racing simulator with genuinely complex vehicle dynamics, popular in academic RL-for-racing research. AWS DeepRacer — a physical (and simulated) 1/18th-scale racing platform specifically built around RL, including real sim-to-real transfer for hobbyist racing competitions. CARLA — a full autonomous driving simulator, considerably more complex than pure racing, useful if your interest extends toward self-driving rather than competitive lap times specifically.
CarRacing-v3 remains the right starting point regardless of which of these interests you long-term — the fundamentals (continuous control, visual perception, reward shaping against exploits) transfer directly to any of the more complex platforms.
If you want to deploy trained policies onto real robot hardware, the MyCobot Pro 630 offers 6-DOF with ROS compatibility — the same kind of hardware used in physical robotics competitions that benefit from simulation-first training.
Common Mistakes People Make
I've hit a few of these myself, and seen the rest repeated across community racing RL projects.
Using DQN or another discrete-action algorithm. Steering and throttle are continuous — this needs PPO, SAC, or similar continuous-action algorithms, not DQN. Forgetting frame stacking. A single frame can't convey the car's velocity or heading — expect erratic, uncertain-looking driving behavior without it. Over-penalizing off-track excursions. This is the racing-specific version of reward hacking — an agent can "solve" an overly harsh off-track penalty by simply never risking speed at all. Judging progress purely from reward numbers instead of watching actual episodes. Just like BipedalWalker, watching rendered laps reveals whether the driving style is genuinely improving or just finding a different way to satisfy the reward function. Expecting fast training. Vision-based continuous control genuinely takes real compute time — treat this more like the Atari or MuJoCo timelines than CartPole's five-minute demo.
Our self-play article covers the broader concept of agents finding unexpected strategies that satisfy reward functions without genuinely solving the intended task — racing reward hacking is a concrete instance of that exact phenomenon.
Wrapping This Up
Training a racing car AI combines nearly every technique covered across this series into one project: CNN-based visual perception from Atari, continuous control via PPO from BipedalWalker and MuJoCo, and careful reward design to avoid the exact proxy-gaming failures the reward design article warned about. Getting all three genuinely right, together, is what separates an agent that races from one that just avoids crashing.
Remember that the "overly cautious" middle-training phase is completely normal, and that a racing agent stuck driving slowly forever is almost always a reward-balance problem rather than an algorithm problem. FYI, once CarRacing-v3 clicks, the jump to something like TORCS or AWS DeepRacer feels like an extension of what you already understand rather than an entirely new discipline :)
Now go watch your early training checkpoints spin out on the very first corner, over and over, before that first genuinely clean lap emerges. That contrast is genuinely one of the most satisfying "before and after" moments in this whole series.