Sam Austin AI

Train an AI to Play Pong from Pixels

September 25, 2026 14 min read Sam Austin
Contents

Retro arcade machines glowing in the dark, the lineage of pixel-based Pong agents

Figure 1: From arcade cabinets to Q-networks — Pong taught generations of agents to see

Teaching an AI to play Pong from pixels sounds almost too simple: see the ball, move the paddle, win the point. Then you actually start training, and your agent watches the ball sail past for thousands of frames like it has tickets to a very boring tennis match.

That challenge makes Pong such a classic reinforcement learning project. Your AI receives only game images, not hand-crafted coordinates for the ball or paddle. It must learn visual features, timing, movement, and long-term strategy from raw pixels. Training an AI to play Pong from pixels gives you a practical introduction to deep reinforcement learning, convolutional neural networks, and value-based learning in one compact game.

Why Pong Still Matters for Reinforcement Learning

Pong might look primitive, but it includes nearly every ingredient that makes reinforcement learning interesting. The agent sees a changing environment, selects actions, receives rewards, and gradually improves through experience.

The game also keeps the action space manageable. An agent usually chooses among a few actions:

  • Move paddle up.
  • Move paddle down.
  • Stay still.
  • Serve or fire, depending on the environment.

That small set makes Pong a strong fit for Deep Q-Networks (DQN). DQN can look at a stack of recent frames, estimate the value of each paddle action, and choose the move that should lead to the best future reward.

Pong also offers a harsh but useful reward signal. The agent earns a positive reward when it scores and a negative reward when the opponent scores. Everything between those points may feel quiet, but the agent still needs to learn how to keep the paddle aligned with the ball.

Easy game, difficult learning problem. Classic reinforcement learning behavior.

The Goal: Learn Directly From the Screen

A pixel-based Pong agent does not receive a convenient state such as:

[ball_x, ball_y, paddle_y, opponent_y]

Instead, it receives image frames from the screen. The neural network must identify important visual patterns by itself.

It needs to learn things like:

  • Where the ball appears.
  • Which direction the ball moves.
  • Where its own paddle sits.
  • How quickly the paddle responds.
  • Whether moving now helps or hurts later.
  • How to return the ball rather than merely chase it.

That last point matters. A naive agent may follow the ball perfectly but still miss because it moves too late. Pong rewards prediction, not panic.

Pick an Algorithm: Why DQN Fits Pong

For standard Pong, DQN is usually the natural starting point. Pong uses a small, discrete action space, and DQN performs well when an agent picks one action from a fixed list.

DQN estimates a Q-value for each action:

Q(s, a)

The Q-value predicts how much future reward an action should produce from the current state. If the ball approaches the lower half of the paddle, the network may assign a higher Q-value to moving down.

Why Not Start With PPO?

You can train Pong with Proximal Policy Optimization, or PPO. PPO handles discrete actions and often produces stable training, especially when you run many environments in parallel. Our PPO explained guide covers the algorithm behind that option.

Still, DQN offers a better first match for a raw-pixel Pong project. It uses replay buffers, reuses old frames, and directly connects game states to a small action set. You can learn a lot about deep reinforcement learning without introducing extra policy-gradient machinery. The DQN vs PPO comparison lays out the full trade-off.

Think of it this way: DQN gives you a practical wrench for a clear bolt. PPO brings an excellent multi-tool. Both can work, but one feels more direct.

Prepare the Pong Environment

You need an RL environment that can reset, accept actions, return frames, and report rewards. Gymnasium-compatible Atari environments work well because they follow a consistent interface.

At each step, your environment should return:

  • A new observation frame.
  • A reward.
  • A termination signal.
  • A truncation signal.
  • Extra diagnostic information.

A typical loop looks like this:

observation, info = env.reset()

while not done:
    action = agent.select_action(observation)
    next_observation, reward, terminated, truncated, info = env.step(action)
    done = terminated or truncated

    agent.store(observation, action, reward, next_observation, done)
    agent.learn()

    observation = next_observation

The loop looks harmless. The real work happens inside the observation preprocessing, replay memory, network updates, and exploration schedule.

Preprocess the Raw Pixels

Raw Pong frames contain a lot of information that does not help much. They may include colors, score displays, borders, flickering sprites, and a large image resolution. Feeding every raw pixel directly into a small network wastes compute and slows learning.

Good preprocessing turns the screen into a cleaner learning signal.

Convert Frames to Grayscale

Pong does not need full color perception. The ball and paddles remain visible in grayscale, so converting RGB frames to one channel reduces input size.

This step gives the network less visual clutter and fewer values to process. The agent does not care whether the background uses a particular shade of blue. It cares about not embarrassing itself in front of the ball.

Resize the Frame

Many Atari DQN projects resize each frame to 84 × 84 pixels. That resolution keeps the visual structure of the game while reducing training cost.

You can use other sizes, but 84 × 84 offers a proven starting point. Tiny frames may hide the ball, while huge frames force the network to work harder for no clear payoff.

Stack Multiple Frames

One image does not show motion. A single Pong frame can tell the agent where the ball sits, but it cannot reveal whether the ball travels upward, downward, left, or right.

Stack four recent frames together:

s_t = [f_{t-3}, f_{t-2}, f_{t-1}, f_t]

The stacked observation gives the network a short motion history. It can infer the ball's direction and speed from changes across frames.

Without frame stacking, the agent tries to play Pong with no sense of momentum. That strategy works about as well as catching a cricket ball with your eyes closed.

Skip Frames

Atari games often run at a frame rate that produces many nearly identical observations. You can repeat each selected action for several frames, such as four frames, before asking the agent to choose again.

Frame skipping speeds up training and helps the agent focus on meaningful movement. It also creates a trade-off: skip too many frames and the ball moves too far between decisions.

A common setup uses:

  • Grayscale observations.
  • 84 × 84 frame size.
  • Four stacked frames.
  • Four-frame action repeats.

Build a Convolutional DQN

A fully connected network struggles with images because it ignores spatial structure. A convolutional neural network (CNN) handles pixels much better.

The CNN scans local parts of each frame, detects patterns, and gradually combines them into higher-level features. In Pong, early layers may notice edges or moving dots, while later layers may represent paddles, ball trajectories, and useful game situations.

A classic DQN-style network might include:

Layer Example configuration Purpose
Input 4 × 84 × 84 frames Captures recent visual history
Conv 1 32 filters, 8 × 8 kernel, stride 4 Detects broad visual patterns
Conv 2 64 filters, 4 × 4 kernel, stride 2 Refines motion and object features
Conv 3 64 filters, 3 × 3 kernel, stride 1 Learns detailed spatial features
Dense 512 units Combines learned features
Output One value per action Estimates Q-values

The output layer contains one Q-value for each legal Pong action. If the environment allows three choices — up, down, and no-op — the network outputs three values.

Use Experience Replay

DQN learns from stored gameplay transitions:

(s_t, a_t, r_t, s_{t+1}, d_t)

A transition records:

  • Current stack of frames.
  • Chosen action.
  • Reward.
  • Next stack of frames.
  • Episode-end flag.

The replay buffer stores many of these transitions. During training, the agent samples random batches instead of learning only from the newest sequence of frames.

Why does that help? Consecutive game frames look extremely similar. If the agent trained only on the latest sequence, it could overfit to a short run and make unstable updates.

Replay sampling mixes wins, losses, near-misses, and ordinary paddle movements. It gives the network a richer training diet.

Add a Target Network

DQN uses two networks:

  • The online network chooses actions and learns from batches.
  • The target network supplies stable Q-learning targets.

The target network copies the online network's weights only at scheduled intervals. That delay prevents the Q-value target from changing at every update.

The DQN target looks like this:

y = r + γ·(1 − d)·max_a' Q_target(s', a')

The online network then learns to make its current Q-value match that target.

Without the target network, your agent tries to predict values using another set of values that it changes constantly. Training can wobble, collapse, or produce Q-values with the emotional stability of a shopping cart on a hill.

Balance Exploration and Exploitation

Your Pong agent needs to explore before it can play well. DQN typically uses epsilon-greedy exploration.

With probability ε, the agent chooses a random action. With probability 1 - ε, it selects the action with the largest Q-value.

A useful schedule might look like this:

Training phase Epsilon value Agent behavior
Early training 1.0 Mostly random exploration
Mid-training 0.3 to 0.1 Mixes learned actions with exploration
Late training 0.05 to 0.01 Mostly uses learned policy
Evaluation 0.0 Always chooses the best predicted action

Start high, lower epsilon gradually, and evaluate with no random actions. Otherwise, you may blame the model for an avoidable miss when the exploration setting literally told it to press a random button. FYI, RL logs can save you from plenty of false conclusions.

Design Rewards Without Overthinking Them

Pong already offers a simple reward structure:

  • Positive reward when the agent scores.
  • Negative reward when the opponent scores.
  • Zero reward during most rallies.

This sparse reward can make learning slow, but it also keeps the objective honest. The agent learns to win points rather than merely move in visually convincing ways.

You can experiment with small shaping rewards, such as rewarding paddle-ball contact. Be careful, though. If you reward contact too heavily, the agent may learn to return the ball without developing a scoring strategy. The reward function design guide walks through more shaping traps.

I prefer starting with the environment's original rewards. Once the agent learns basic play, you can add shaping only if training stalls.

Train the Agent Step by Step

Here is a practical sequence for training a Pong agent from pixels:

  1. Create a Pong environment with discrete actions.
  2. Apply grayscale conversion, resizing, frame stacking, and frame skipping.
  3. Initialize a CNN-based Q-network and matching target network.
  4. Create a replay buffer with room for hundreds of thousands of transitions.
  5. Collect random actions during an initial warm-up period.
  6. Store every transition in the replay buffer.
  7. Sample mini-batches after the warm-up phase.
  8. Train the online network against target Q-values.
  9. Update the target network at fixed intervals.
  10. Reduce epsilon gradually as training progresses.
  11. Evaluate the agent with epsilon set to zero.
  12. Save checkpoints, videos, and training logs.

The random warm-up matters. A replay buffer filled with a little variety gives DQN a better start than a buffer containing one repeated action and several immediate losses.

Useful Hyperparameters for Pong DQN

Start with sensible defaults before you start tweaking everything in sight.

Hyperparameter Starting value Why it matters
Learning rate 0.0001 Controls how quickly network weights update
Discount factor (gamma) 0.99 Values future points and rallies
Replay buffer size 100,000 to 1,000,000 Stores diverse gameplay
Batch size 32 Balances memory use and stable updates
Initial replay warm-up 10,000 to 50,000 steps Builds varied experience first
Target update interval 5,000 to 10,000 steps Stabilizes Q-learning targets
Initial epsilon 1.0 Encourages early exploration
Final epsilon 0.01 Retains a small amount of exploration
Frame stack 4 Reveals motion direction
Frame skip 4 Reduces redundant decisions

Do not treat these values as sacred. They give you a baseline, not a prophecy.

Change one element at a time. If you alter the network, reward scheme, learning rate, exploration schedule, and preprocessing in one run, you will learn almost nothing from the result.

Common Problems When Training Pong From Pixels

The Agent Never Scores

Check whether the agent actually sees a useful state. Incorrect preprocessing can crop out the paddle or make the ball nearly invisible.

Also inspect rewards. If your wrapper accidentally drops rewards, the agent has no reason to improve. It will happily wander forever because you accidentally removed the entire point of the game.

The Agent Moves but Misses the Ball

The agent may lack motion information. Confirm that frame stacking works and that the frame order stays correct.

Also reduce overly aggressive frame skipping. If your agent acts only after the ball crosses half the screen, it cannot make subtle paddle adjustments in time.

Q-Values Explode

Large or unstable Q-values often point to a learning rate that is too high, incorrect terminal-state handling, or a missing target network.

Try a lower learning rate, gradient clipping, reward clipping, or Huber loss. You can also use Double DQN to reduce overly optimistic action-value estimates.

Training Improves Then Collapses

This problem often appears when the policy overfits to a narrow strategy or the Q-network becomes unstable.

Try these fixes:

  • Increase replay-buffer diversity.
  • Slow down epsilon decay.
  • Update the target network more regularly.
  • Train with multiple random seeds.
  • Use Double DQN or Dueling DQN.
  • Evaluate with a fixed, deterministic policy.

Do not trust one spectacular game. A Pong agent can score a few lucky points and still have no clue how to play consistently. When you compare runs, the evaluation method here keeps you honest.

Improve Your Pong Agent

Once the basic DQN agent works, you can upgrade it.

Double DQN

Double DQN reduces Q-value overestimation. The online network selects the best next action, while the target network evaluates it.

That small change often makes training more stable:

y = r + γ·Q_target(s', argmax_a' Q_online(s', a'))

Dueling DQN

Dueling DQN separates state value from action advantage. It helps the network learn which states matter even when several actions produce similar results.

This can help in Pong when the ball sits far away and multiple paddle movements seem similarly reasonable.

Prioritized Experience Replay

Prioritized replay samples more informative transitions more often. Important moments, such as scoring, losing, or nearly missing the ball, can teach the agent more than another ordinary frame in the middle of a rally.

Use it carefully. You still need diverse replay data, not just a collection of dramatic disasters.

Want your agent to return aces instead of air? Grab the GPTAstra full course at https://cutt.ly/5yviN6qd and follow the full Atari pipeline with runnable preprocessing and training code.

Frequently Asked Questions

Why is Pong a good reinforcement learning project?

Pong pairs a tiny discrete action space with a hard visual learning problem: the agent sees only raw frames and must infer ball position, direction, and timing. You get a complete deep RL pipeline in one compact game.

Should I use DQN or PPO for pixel-based Pong?

Start with DQN. Its replay buffer reuses old frames, and the small fixed action set matches value-based learning well. PPO also works and stabilizes training with parallel environments, but adds policy-gradient machinery you do not need for a first pass.

Why stack four frames instead of using one?

A single frame shows position but not motion. Stacking recent frames gives the network a short motion history so it can infer the ball's direction and speed — without it, the agent plays with no sense of momentum.

Why grayscale and 84 x 84 resizing?

Pong does not need color to track the ball and paddles. Grayscale and an 84 x 84 frame cut input size and visual clutter while keeping the structure of the game — a proven starting point from the original DQN work.

Why does my Pong agent never score?

Check preprocessing first: a bad crop can hide the paddle or ball. Then check rewards — if a wrapper drops the reward signal, the agent has no reason to improve and will simply wander.

What should I try after vanilla DQN works?

Double DQN to reduce Q-value overestimation, Dueling DQN to separate state value from action advantage, and prioritized experience replay to sample the moments that teach the most — in that order, one change at a time.

Final Thoughts

Training an AI to play Pong from pixels teaches you the full deep reinforcement learning pipeline: image preprocessing, frame stacking, convolutional networks, replay buffers, target networks, exploration, and evaluation.

Start with a DQN agent, grayscale 84 × 84 frames, four-frame stacks, and a clean reward signal. Train patiently, log everything, and do not panic when your first agent behaves like it has never seen a paddle before. If pixel projects feel heavier than expected, the numerical-state Flappy Bird tutorial shows the lighter alternative.

Pong may look simple, but it forces your AI to learn perception, motion prediction, timing, and strategy from scratch. When your agent finally returns a fast ball and scores a point from raw pixels alone, that moment feels genuinely brilliant. Then you will probably spend another hour trying to figure out why it suddenly forgot how to move down.

Share this article X Facebook LinkedIn Reddit WhatsApp

Related Articles