Contents
Figure 1: From arcade cabinets to Q-networks — Pong taught generations of agents to see
Teaching an AI to play Pong from pixels sounds almost too simple: see the ball, move the paddle, win the point. Then you actually start training, and your agent watches the ball sail past for thousands of frames like it has tickets to a very boring tennis match.
That challenge makes Pong such a classic reinforcement learning project. Your AI receives only game images, not hand-crafted coordinates for the ball or paddle. It must learn visual features, timing, movement, and long-term strategy from raw pixels. Training an AI to play Pong from pixels gives you a practical introduction to deep reinforcement learning, convolutional neural networks, and value-based learning in one compact game.
Why Pong Still Matters for Reinforcement Learning
Pong might look primitive, but it includes nearly every ingredient that makes reinforcement learning interesting. The agent sees a changing environment, selects actions, receives rewards, and gradually improves through experience.
The game also keeps the action space manageable. An agent usually chooses among a few actions:
- Move paddle up.
- Move paddle down.
- Stay still.
- Serve or fire, depending on the environment.
That small set makes Pong a strong fit for Deep Q-Networks (DQN). DQN can look at a stack of recent frames, estimate the value of each paddle action, and choose the move that should lead to the best future reward.
Pong also offers a harsh but useful reward signal. The agent earns a positive reward when it scores and a negative reward when the opponent scores. Everything between those points may feel quiet, but the agent still needs to learn how to keep the paddle aligned with the ball.
Easy game, difficult learning problem. Classic reinforcement learning behavior.
The Goal: Learn Directly From the Screen
A pixel-based Pong agent does not receive a convenient state such as:
[ball_x, ball_y, paddle_y, opponent_y]
Instead, it receives image frames from the screen. The neural network must identify important visual patterns by itself.
It needs to learn things like:
- Where the ball appears.
- Which direction the ball moves.
- Where its own paddle sits.
- How quickly the paddle responds.
- Whether moving now helps or hurts later.
- How to return the ball rather than merely chase it.
That last point matters. A naive agent may follow the ball perfectly but still miss because it moves too late. Pong rewards prediction, not panic.
Pick an Algorithm: Why DQN Fits Pong
For standard Pong, DQN is usually the natural starting point. Pong uses a small, discrete action space, and DQN performs well when an agent picks one action from a fixed list.
DQN estimates a Q-value for each action:
Q(s, a)
The Q-value predicts how much future reward an action should produce from the current state. If the ball approaches the lower half of the paddle, the network may assign a higher Q-value to moving down.
Why Not Start With PPO?
You can train Pong with Proximal Policy Optimization, or PPO. PPO handles discrete actions and often produces stable training, especially when you run many environments in parallel. Our PPO explained guide covers the algorithm behind that option.
Still, DQN offers a better first match for a raw-pixel Pong project. It uses replay buffers, reuses old frames, and directly connects game states to a small action set. You can learn a lot about deep reinforcement learning without introducing extra policy-gradient machinery. The DQN vs PPO comparison lays out the full trade-off.
Think of it this way: DQN gives you a practical wrench for a clear bolt. PPO brings an excellent multi-tool. Both can work, but one feels more direct.
Prepare the Pong Environment
You need an RL environment that can reset, accept actions, return frames, and report rewards. Gymnasium-compatible Atari environments work well because they follow a consistent interface.
At each step, your environment should return:
- A new observation frame.
- A reward.
- A termination signal.
- A truncation signal.
- Extra diagnostic information.
A typical loop looks like this:
observation, info = env.reset()
while not done:
action = agent.select_action(observation)
next_observation, reward, terminated, truncated, info = env.step(action)
done = terminated or truncated
agent.store(observation, action, reward, next_observation, done)
agent.learn()
observation = next_observation
The loop looks harmless. The real work happens inside the observation preprocessing, replay memory, network updates, and exploration schedule.
Preprocess the Raw Pixels
Raw Pong frames contain a lot of information that does not help much. They may include colors, score displays, borders, flickering sprites, and a large image resolution. Feeding every raw pixel directly into a small network wastes compute and slows learning.
Good preprocessing turns the screen into a cleaner learning signal.
Convert Frames to Grayscale
Pong does not need full color perception. The ball and paddles remain visible in grayscale, so converting RGB frames to one channel reduces input size.
This step gives the network less visual clutter and fewer values to process. The agent does not care whether the background uses a particular shade of blue. It cares about not embarrassing itself in front of the ball.
Resize the Frame
Many Atari DQN projects resize each frame to 84 × 84 pixels. That resolution keeps the visual structure of the game while reducing training cost.
You can use other sizes, but 84 × 84 offers a proven starting point. Tiny frames may hide the ball, while huge frames force the network to work harder for no clear payoff.
Stack Multiple Frames
One image does not show motion. A single Pong frame can tell the agent where the ball sits, but it cannot reveal whether the ball travels upward, downward, left, or right.
Stack four recent frames together:
s_t = [f_{t-3}, f_{t-2}, f_{t-1}, f_t]
The stacked observation gives the network a short motion history. It can infer the ball's direction and speed from changes across frames.
Without frame stacking, the agent tries to play Pong with no sense of momentum. That strategy works about as well as catching a cricket ball with your eyes closed.
Skip Frames
Atari games often run at a frame rate that produces many nearly identical observations. You can repeat each selected action for several frames, such as four frames, before asking the agent to choose again.
Frame skipping speeds up training and helps the agent focus on meaningful movement. It also creates a trade-off: skip too many frames and the ball moves too far between decisions.
A common setup uses:
- Grayscale observations.
- 84 × 84 frame size.
- Four stacked frames.
- Four-frame action repeats.
Build a Convolutional DQN
A fully connected network struggles with images because it ignores spatial structure. A convolutional neural network (CNN) handles pixels much better.
The CNN scans local parts of each frame, detects patterns, and gradually combines them into higher-level features. In Pong, early layers may notice edges or moving dots, while later layers may represent paddles, ball trajectories, and useful game situations.
A classic DQN-style network might include:
| Layer | Example configuration | Purpose |
|---|---|---|
| Input | 4 × 84 × 84 frames | Captures recent visual history |
| Conv 1 | 32 filters, 8 × 8 kernel, stride 4 | Detects broad visual patterns |
| Conv 2 | 64 filters, 4 × 4 kernel, stride 2 | Refines motion and object features |
| Conv 3 | 64 filters, 3 × 3 kernel, stride 1 | Learns detailed spatial features |
| Dense | 512 units | Combines learned features |
| Output | One value per action | Estimates Q-values |
The output layer contains one Q-value for each legal Pong action. If the environment allows three choices — up, down, and no-op — the network outputs three values.
Use Experience Replay
DQN learns from stored gameplay transitions:
(s_t, a_t, r_t, s_{t+1}, d_t)
A transition records:
- Current stack of frames.
- Chosen action.
- Reward.
- Next stack of frames.
- Episode-end flag.
The replay buffer stores many of these transitions. During training, the agent samples random batches instead of learning only from the newest sequence of frames.
Why does that help? Consecutive game frames look extremely similar. If the agent trained only on the latest sequence, it could overfit to a short run and make unstable updates.
Replay sampling mixes wins, losses, near-misses, and ordinary paddle movements. It gives the network a richer training diet.
Add a Target Network
DQN uses two networks:
- The online network chooses actions and learns from batches.
- The target network supplies stable Q-learning targets.
The target network copies the online network's weights only at scheduled intervals. That delay prevents the Q-value target from changing at every update.
The DQN target looks like this:
y = r + γ·(1 − d)·max_a' Q_target(s', a')
The online network then learns to make its current Q-value match that target.
Without the target network, your agent tries to predict values using another set of values that it changes constantly. Training can wobble, collapse, or produce Q-values with the emotional stability of a shopping cart on a hill.
Balance Exploration and Exploitation
Your Pong agent needs to explore before it can play well. DQN typically uses epsilon-greedy exploration.
With probability ε, the agent chooses a random action. With probability 1 - ε, it selects the action with the largest Q-value.
A useful schedule might look like this:
| Training phase | Epsilon value | Agent behavior |
|---|---|---|
| Early training | 1.0 | Mostly random exploration |
| Mid-training | 0.3 to 0.1 | Mixes learned actions with exploration |
| Late training | 0.05 to 0.01 | Mostly uses learned policy |
| Evaluation | 0.0 | Always chooses the best predicted action |
Start high, lower epsilon gradually, and evaluate with no random actions. Otherwise, you may blame the model for an avoidable miss when the exploration setting literally told it to press a random button. FYI, RL logs can save you from plenty of false conclusions.
Design Rewards Without Overthinking Them
Pong already offers a simple reward structure:
- Positive reward when the agent scores.
- Negative reward when the opponent scores.
- Zero reward during most rallies.
This sparse reward can make learning slow, but it also keeps the objective honest. The agent learns to win points rather than merely move in visually convincing ways.
You can experiment with small shaping rewards, such as rewarding paddle-ball contact. Be careful, though. If you reward contact too heavily, the agent may learn to return the ball without developing a scoring strategy. The reward function design guide walks through more shaping traps.
I prefer starting with the environment's original rewards. Once the agent learns basic play, you can add shaping only if training stalls.
Train the Agent Step by Step
Here is a practical sequence for training a Pong agent from pixels:
- Create a Pong environment with discrete actions.
- Apply grayscale conversion, resizing, frame stacking, and frame skipping.
- Initialize a CNN-based Q-network and matching target network.
- Create a replay buffer with room for hundreds of thousands of transitions.
- Collect random actions during an initial warm-up period.
- Store every transition in the replay buffer.
- Sample mini-batches after the warm-up phase.
- Train the online network against target Q-values.
- Update the target network at fixed intervals.
- Reduce epsilon gradually as training progresses.
- Evaluate the agent with epsilon set to zero.
- Save checkpoints, videos, and training logs.
The random warm-up matters. A replay buffer filled with a little variety gives DQN a better start than a buffer containing one repeated action and several immediate losses.
Useful Hyperparameters for Pong DQN
Start with sensible defaults before you start tweaking everything in sight.
| Hyperparameter | Starting value | Why it matters |
|---|---|---|
| Learning rate | 0.0001 | Controls how quickly network weights update |
| Discount factor (gamma) | 0.99 | Values future points and rallies |
| Replay buffer size | 100,000 to 1,000,000 | Stores diverse gameplay |
| Batch size | 32 | Balances memory use and stable updates |
| Initial replay warm-up | 10,000 to 50,000 steps | Builds varied experience first |
| Target update interval | 5,000 to 10,000 steps | Stabilizes Q-learning targets |
| Initial epsilon | 1.0 | Encourages early exploration |
| Final epsilon | 0.01 | Retains a small amount of exploration |
| Frame stack | 4 | Reveals motion direction |
| Frame skip | 4 | Reduces redundant decisions |
Do not treat these values as sacred. They give you a baseline, not a prophecy.
Change one element at a time. If you alter the network, reward scheme, learning rate, exploration schedule, and preprocessing in one run, you will learn almost nothing from the result.
Common Problems When Training Pong From Pixels
The Agent Never Scores
Check whether the agent actually sees a useful state. Incorrect preprocessing can crop out the paddle or make the ball nearly invisible.
Also inspect rewards. If your wrapper accidentally drops rewards, the agent has no reason to improve. It will happily wander forever because you accidentally removed the entire point of the game.
The Agent Moves but Misses the Ball
The agent may lack motion information. Confirm that frame stacking works and that the frame order stays correct.
Also reduce overly aggressive frame skipping. If your agent acts only after the ball crosses half the screen, it cannot make subtle paddle adjustments in time.
Q-Values Explode
Large or unstable Q-values often point to a learning rate that is too high, incorrect terminal-state handling, or a missing target network.
Try a lower learning rate, gradient clipping, reward clipping, or Huber loss. You can also use Double DQN to reduce overly optimistic action-value estimates.
Training Improves Then Collapses
This problem often appears when the policy overfits to a narrow strategy or the Q-network becomes unstable.
Try these fixes:
- Increase replay-buffer diversity.
- Slow down epsilon decay.
- Update the target network more regularly.
- Train with multiple random seeds.
- Use Double DQN or Dueling DQN.
- Evaluate with a fixed, deterministic policy.
Do not trust one spectacular game. A Pong agent can score a few lucky points and still have no clue how to play consistently. When you compare runs, the evaluation method here keeps you honest.
Improve Your Pong Agent
Once the basic DQN agent works, you can upgrade it.
Double DQN
Double DQN reduces Q-value overestimation. The online network selects the best next action, while the target network evaluates it.
That small change often makes training more stable:
y = r + γ·Q_target(s', argmax_a' Q_online(s', a'))
Dueling DQN
Dueling DQN separates state value from action advantage. It helps the network learn which states matter even when several actions produce similar results.
This can help in Pong when the ball sits far away and multiple paddle movements seem similarly reasonable.
Prioritized Experience Replay
Prioritized replay samples more informative transitions more often. Important moments, such as scoring, losing, or nearly missing the ball, can teach the agent more than another ordinary frame in the middle of a rally.
Use it carefully. You still need diverse replay data, not just a collection of dramatic disasters.
Recommended Books
- Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow by Aurélien Géron — the clearest CNN foundations available, covering exactly the convolution, stride, and feature-stack choices this architecture table makes.
- Deep Reinforcement Learning Hands-On by Maxim Lapan — walks through the full Atari-from-pixels pipeline: preprocessing wrappers, frame stacking, replay, and target networks.
- Reinforcement Learning: An Introduction by Richard S. Sutton and Andrew G. Barto — the theory of Q-learning and exploration that every pixel-based agent still runs on.
Want your agent to return aces instead of air? Grab the GPTAstra full course at https://cutt.ly/5yviN6qd and follow the full Atari pipeline with runnable preprocessing and training code.
Frequently Asked Questions
Why is Pong a good reinforcement learning project?
Pong pairs a tiny discrete action space with a hard visual learning problem: the agent sees only raw frames and must infer ball position, direction, and timing. You get a complete deep RL pipeline in one compact game.
Should I use DQN or PPO for pixel-based Pong?
Start with DQN. Its replay buffer reuses old frames, and the small fixed action set matches value-based learning well. PPO also works and stabilizes training with parallel environments, but adds policy-gradient machinery you do not need for a first pass.
Why stack four frames instead of using one?
A single frame shows position but not motion. Stacking recent frames gives the network a short motion history so it can infer the ball's direction and speed — without it, the agent plays with no sense of momentum.
Why grayscale and 84 x 84 resizing?
Pong does not need color to track the ball and paddles. Grayscale and an 84 x 84 frame cut input size and visual clutter while keeping the structure of the game — a proven starting point from the original DQN work.
Why does my Pong agent never score?
Check preprocessing first: a bad crop can hide the paddle or ball. Then check rewards — if a wrapper drops the reward signal, the agent has no reason to improve and will simply wander.
What should I try after vanilla DQN works?
Double DQN to reduce Q-value overestimation, Dueling DQN to separate state value from action advantage, and prioritized experience replay to sample the moments that teach the most — in that order, one change at a time.
Final Thoughts
Training an AI to play Pong from pixels teaches you the full deep reinforcement learning pipeline: image preprocessing, frame stacking, convolutional networks, replay buffers, target networks, exploration, and evaluation.
Start with a DQN agent, grayscale 84 × 84 frames, four-frame stacks, and a clean reward signal. Train patiently, log everything, and do not panic when your first agent behaves like it has never seen a paddle before. If pixel projects feel heavier than expected, the numerical-state Flappy Bird tutorial shows the lighter alternative.
Pong may look simple, but it forces your AI to learn perception, motion prediction, timing, and strategy from scratch. When your agent finally returns a fast ball and scores a point from raw pixels alone, that moment feels genuinely brilliant. Then you will probably spend another hour trying to figure out why it suddenly forgot how to move down.