Sam Austin AI

Train an AI to Play Flappy Bird with Reinforcement Learning

September 25, 2026 13 min read Sam Austin
Contents

A hummingbird hovering in mid-air, echoing the flap-or-glide choice of a Flappy Bird agent

Figure 1: One tiny action space, thousands of doomed flights — Flappy Bird is RL in its purest form

Flappy Bird looks simple until you hand the controls to an AI agent. The bird only flaps or does nothing, yet one bad tap sends it directly into a pipe, the ground, or a fresh batch of disappointment.

That tiny action space makes Flappy Bird an excellent reinforcement learning project. You can build an agent, define rewards, train it through thousands of failed flights, and eventually watch it navigate pipes without a human touching the keyboard. Train an AI to play Flappy Bird with reinforcement learning and you will learn more than just one game trick — you will understand the core loop behind many game-playing agents.

Why Flappy Bird Makes a Great RL Project

Flappy Bird gives your AI a simple goal: keep flying and pass through as many pipe gaps as possible. The agent sees the game state, chooses an action, earns a reward, and repeats the process.

That sounds easy, right? Then the bird flaps twice, rockets into the ceiling, and reminds you that reinforcement learning loves humble beginnings.

The game works so well for RL because it includes several useful ingredients:

  • Simple actions: Flap or do nothing.
  • Clear objective: Stay alive and pass pipes.
  • Fast feedback: the game quickly signals whether an action helped.
  • Repeatable episodes: each crash ends an episode and starts a new attempt.
  • Room for improvement: a random agent fails instantly, while a trained agent can learn timing and positioning.

You do not need a giant 3D world or a robot arm to learn reinforcement learning. Sometimes a stubborn pixel bird does the job nicely. The same loop scales up to the bigger builds in our reinforcement learning for games guide.

What Reinforcement Learning Means Here

Reinforcement learning teaches an agent through trial and error. Instead of giving the AI a fixed list of "correct" moves, you let it interact with the environment and learn which actions bring better outcomes.

In Flappy Bird, the agent repeatedly follows this loop:

  1. It observes the current game state.
  2. It chooses whether to flap.
  3. The environment updates the bird and pipes.
  4. The agent receives a reward or penalty.
  5. It uses that experience to improve future decisions.

The agent does not begin with an understanding of gravity, pipe gaps, or timing. It starts clueless, much like the rest of us during our first Flappy Bird session.

Over many episodes, the agent learns that flapping at the right moment keeps it near the next pipe gap. It also learns that flying directly into a pipe rarely produces career growth.

Choose the Right Reinforcement Learning Algorithm

For a basic Flappy Bird environment, Deep Q-Networks (DQN) make an excellent starting point.

Why? The game has a discrete action space. At any point, the agent usually chooses between:

  • 0: Do nothing.
  • 1: Flap.

DQN handles this setup well because it estimates a value for each possible action. It asks a useful question: "Given the current game state, which action should I choose to earn the most future reward?"

Why DQN Fits Flappy Bird

DQN works best when an environment offers a small, fixed menu of actions. Flappy Bird gives the agent exactly that.

The neural network takes the game state as input and outputs two Q-values:

Q(s, do nothing)
Q(s, flap)

The agent chooses the action with the higher Q-value most of the time. During training, it sometimes chooses randomly so it can explore unexpected actions.

That random exploration matters. Without it, the agent might decide that flapping immediately always looks promising, then repeatedly slam into the ceiling forever. Very determined. Not very productive.

Could You Use PPO Instead?

Yes, you can train a Flappy Bird agent with Proximal Policy Optimization (PPO). PPO works with discrete action spaces and often provides stable training.

Still, I would start with DQN for this project. DQN fits the tiny two-action setup naturally, uses a replay buffer to reuse old experiences, and teaches value-based reinforcement learning clearly. Our DQN vs PPO comparison explains the trade-off in depth.

Use PPO if you already built a PPO training pipeline — a Stable-Baselines3 setup covers that path — want to compare algorithms, or plan to add a more complex controller later. For a first pass, DQN keeps things focused. If you want the mechanics behind PPO, read our PPO explained guide.

Build the Flappy Bird Environment

Before training anything, you need an environment that follows a predictable interface. Gymnasium-style environments work especially well because they standardize the interaction loop.

Your environment should define:

  • Observation space: what the agent sees.
  • Action space: the choices it can make.
  • Reset method: how the game starts a fresh episode.
  • Step method: how the environment processes an action.
  • Reward function: how the agent receives feedback.
  • Termination rules: when the bird crashes or finishes.

The environment does not need fancy graphics during training. In fact, headless training often runs faster because your computer does not waste time drawing thousands of doomed birds.

Pick a State Representation

You have two common choices for the Flappy Bird state.

Option 1: Use Numerical Features

This option gives the agent a compact set of game values, such as:

  • Bird's vertical position.
  • Bird's vertical velocity.
  • Horizontal distance to the next pipe.
  • Vertical position of the next pipe gap.
  • Vertical distance between the bird and the gap center.

A compact state might look like:

s = [y_bird, v_y, x_pipe, y_gap - y_bird]

This approach trains much faster than raw-image input. It also helps you focus on reinforcement learning instead of computer vision.

For a first project, I strongly recommend numerical features. You will spend less time debugging pixel preprocessing and more time learning why your bird keeps performing aerial acrobatics into pipes.

Option 2: Use Raw Game Frames

This option gives the AI screenshots from the game. A convolutional neural network processes those frames and learns useful visual features. The Atari deep RL walkthrough shows the full pixel pipeline.

Raw pixels make the project more realistic, but they also make it harder. The agent needs more data, more compute, frame preprocessing, frame stacking, and careful tuning.

Use image input if you want the full "AI learns directly from the screen" experience. Use numerical state features if you want to build a working project before your coffee gets cold.

Design a Reward Function That Helps

Your reward function determines what the agent actually learns. If you reward the wrong behavior, your bird will find a strange shortcut faster than you can say "reward hacking."

A practical Flappy Bird reward setup might use:

Event Example reward
Survive one frame or time step +0.01
Pass through a pipe +1.0
Move closer to the next gap +0.05
Hit a pipe or the ground -1.0
Fly too high or too low Small negative penalty

The exact numbers can change, but the priorities should stay clear:

  • Reward survival and pipe progress.
  • Give a meaningful reward for passing a pipe.
  • Penalize crashes.
  • Avoid rewards that encourage pointless flapping.

Keep Reward Shaping Sensible

Reward shaping can speed up learning, but it can also produce weird behavior. For example, if you reward the bird only for staying near the gap center, it may hover around one safe-looking area instead of advancing through pipes.

I prefer a simple structure:

  • Small reward for surviving.
  • Large reward for passing a pipe.
  • Clear penalty for crashing.

That reward design keeps the objective honest. The bird should survive because survival helps it pass more pipes, not because it discovered a loophole in your scoring system. The reward function design guide covers more shaping pitfalls.

Set Up the DQN Agent

A DQN agent needs several moving pieces, but you can understand each one without summoning a reinforcement learning wizard.

The Q-Network

The Q-network maps a game state to one Q-value per action.

For a numerical state vector, a simple fully connected neural network can work well:

  • Input layer: state features.
  • Hidden layer: 64 or 128 neurons.
  • Hidden layer: 64 or 128 neurons.
  • Output layer: two Q-values.

The output might look like this:

[2.14, 2.37]

That result means the model currently believes flap offers a slightly better expected future reward than do nothing.

The Replay Buffer

The replay buffer stores past interactions as tuples:

(s, a, r, s', d)

Each entry includes:

  • Current state s.
  • Selected action a.
  • Reward r.
  • Next state s'.
  • Done flag d, which marks a crash or episode end.

During training, DQN samples random batches from the replay buffer. Random sampling reduces the risk that the model overfits to one recent sequence of events.

It also lets the agent learn from a useful pipe pass more than once. That is a big deal when successful flights stay rare early in training.

The Target Network

DQN uses a second network called the target network. This network provides stable targets while the main Q-network learns.

You periodically copy the main network weights into the target network. That may sound like unnecessary duplication, but it prevents the target values from changing too wildly during training.

Without a target network, DQN can chase its own changing predictions in circles. The agent already crashes enough; it does not need philosophical instability too.

The Flappy Bird Training Loop

The reinforcement learning loop looks like this:

state, info = env.reset()

while not done:
    action = agent.select_action(state)
    next_state, reward, terminated, truncated, info = env.step(action)
    done = terminated or truncated

    replay_buffer.add(state, action, reward, next_state, done)

    agent.train_on_batch(replay_buffer.sample(batch_size))

    state = next_state

This short loop hides a lot of learning. Each time the agent acts, it gathers evidence about whether flapping or waiting helped.

At first, the AI will make terrible choices. Expect crashes. Expect bizarre flapping patterns. Expect the bird to discover that gravity exists with great enthusiasm.

Then the agent starts to recognize patterns. It learns how fast it falls, how much a flap lifts it, and when to move toward the next pipe gap.

Handle Exploration with Epsilon-Greedy Actions

DQN often uses an epsilon-greedy strategy during training.

With probability ε, the agent picks a random action. With probability 1 - ε, it selects the action with the highest predicted Q-value.

A typical schedule might work like this:

  • Start with ε = 1.0: choose random actions often.
  • Gradually reduce epsilon over training.
  • End near ε = 0.05 or ε = 0.01: mostly exploit learned behavior.

Early exploration helps the agent discover pipe gaps and useful timing patterns. Later, lower exploration lets it use what it learned.

During evaluation, set epsilon to zero or choose the highest-value action every time. You want to measure skill, not watch the agent randomly flap itself into retirement. FYI, that evaluation step catches many "it looked good once" mistakes.

Tune the Important Hyperparameters

You do not need to tune everything at once. Start with reasonable DQN settings and adjust based on training curves.

Hyperparameter Good starting range What it controls
Learning rate 0.0001 to 0.001 How quickly the network changes
Discount factor (gamma) 0.95 to 0.99 How much the agent values future rewards
Replay buffer size 50,000 to 100,000 How much experience the agent remembers
Batch size 32 to 128 How many transitions each update uses
Target update interval 500 to 2,000 steps How often target weights refresh
Initial epsilon 1.0 Early exploration level
Final epsilon 0.01 to 0.05 Late-training exploration level

Start with a discount factor of 0.99, a batch size of 64, and a learning rate of 0.0005. Those values offer a useful baseline for many small DQN projects.

Change one setting at a time. If you alter the reward function, network size, learning rate, and epsilon schedule together, you will not know which change improved training. You will only know that something happened, which is technically true but not very helpful.

Common Flappy Bird RL Problems

The Bird Never Passes a Pipe

Check the reward structure first. If the bird only receives a reward after passing a pipe, it may struggle to learn because successful events happen too rarely.

Add a small survival reward or a carefully designed progress reward. Keep the pipe-passing reward much larger so the agent still focuses on the real goal.

The Bird Flaps Constantly

Your exploration rate may stay too high, or your reward function may accidentally reward upward movement more than useful positioning.

Also check your physics. If the flap force feels too strong or gravity feels too weak, the agent may find nonstop flapping strangely effective.

The Agent Improves, Then Gets Worse

DQN can become unstable when Q-values drift too high or training updates happen too aggressively.

Try these fixes:

  • Reduce the learning rate.
  • Update the target network more frequently.
  • Increase replay-buffer size.
  • Use Double DQN to reduce overestimated Q-values.
  • Clip rewards if they vary wildly.
  • Train with multiple random seeds.

Do not judge progress from one run. RL includes randomness, and one lucky bird does not prove that your setup works.

Evaluate Your Trained Flappy Bird Agent

After training, run several evaluation episodes with exploration turned off. Track more than just the single best score. Our reinforcement learning evaluation walkthrough shows the comparison method in detail.

Measure:

  • Average score across evaluation runs.
  • Maximum score.
  • Median score.
  • Pipe passes per episode.
  • Crash location and crash frequency.
  • Consistency across random seeds.

A trained agent that scores 20 once but averages 2 across 50 episodes has not mastered Flappy Bird. It has experienced a fortunate accident.

You can also record gameplay videos. Watching the agent fly gives you fast visual feedback about whether it behaves intelligently or merely survives through weird timing quirks.

Want your bird to actually clear the pipes? Grab the GPTAstra full course at https://cutt.ly/5yviN6qd and turn this tutorial into a trained agent, with ready-to-run training loops and hyperparameter tables.

Frequently Asked Questions

Which reinforcement learning algorithm is best for Flappy Bird?

Deep Q-Networks (DQN) are the natural first choice because Flappy Bird offers a discrete two-action space: do nothing or flap. PPO also works if you already have a PPO pipeline or plan a more complex controller later.

Why use numerical state features instead of raw pixels?

Numerical features such as bird height, velocity, and distance to the next gap train much faster, need less compute, and let you focus on reinforcement learning rather than computer vision. Raw frames require a convolutional network, frame stacking, and more data.

Why does my bird flap constantly?

Usually the exploration rate stays too high, the reward function accidentally rewards upward movement, or the physics are off. A flap force that is too strong or gravity too weak makes nonstop flapping strangely effective.

How should I design the Flappy Bird reward function?

Give a small reward for surviving each frame, a large reward for passing a pipe, and a clear penalty for crashing. Avoid any shaping that lets the bird farm reward by hovering instead of advancing.

How long does Flappy Bird training take?

Expect thousands of episodes. Early runs crash immediately; the agent gradually learns flap timing. Changes take effect slowly, so alter one hyperparameter at a time and watch training curves across several seeds.

How do I evaluate a trained Flappy Bird agent?

Turn exploration off and run multiple evaluation episodes. Track average score, median, maximum, pipe passes per episode, crash locations, and consistency across random seeds rather than the single best run.

Final Thoughts

Training an AI to play Flappy Bird with reinforcement learning offers one of the best entry points into practical game AI. The game gives you a simple action space, fast episodes, clear rewards, and enough challenge to teach real reinforcement learning lessons.

Start with a numerical state representation and a DQN agent. Build a clean environment, reward survival and pipe passes, use a replay buffer, and lower exploration gradually. Once that version works, try raw pixel input, Double DQN, prioritized replay, or PPO comparisons.

The first version of your agent will crash. The hundredth version may still crash in spectacular fashion. But when it finally threads a pipe gap with a perfectly timed flap, you will understand why reinforcement learning projects feel so addictive. It is not just a bird anymore — it is your tiny, stubborn pilot.

Share this article X Facebook LinkedIn Reddit WhatsApp

Related Articles