Sam Austin AI

Isaac Gym Tutorial: NVIDIA's GPU-Accelerated Robot Simulator

September 25, 2026 15 min read Sam Austin
Contents

A humanoid robot, the kind of machine Isaac Gym trains by the thousand

Figure 1: One robot on CPU, thousands on GPU — Isaac Gym's core promise

Training a robot one episode at a time can feel painfully slow. Your policy takes a step, the simulator thinks about gravity, your GPU waits around, and suddenly lunch arrives before the robot learns not to fall over.

Isaac Gym changed that workflow by running physics simulation, observations, reward calculations, and policy training on the GPU. Instead of training one robot in one virtual world, you can train thousands of robots across thousands of parallel environments. That speed matters when your robot needs millions — or billions — of simulation steps before it discovers the obvious, like keeping its feet underneath it.

One important note before we start: NVIDIA now positions Isaac Lab as the successor to the Isaac Gym Preview Release and recommends that new users build robot-learning work in Isaac Lab. Isaac Gym remains important because it introduced the GPU-native robot-learning workflow that many projects still reference.

What Is Isaac Gym?

Isaac Gym is NVIDIA's prototype physics simulation environment for reinforcement learning and robot-learning research. It uses NVIDIA PhysX for GPU-accelerated rigid-body simulation and exposes simulation data through PyTorch tensors.

That tensor-first approach creates the central benefit: the simulator can keep physics state on the GPU, and your neural network can consume it directly. You avoid repeatedly copying observations from GPU to CPU and back again, which can slow conventional robotics pipelines.

Isaac Gym supports:

  • GPU-accelerated physics simulation.
  • Massive parallel environments.
  • PyTorch tensor APIs.
  • Robot models in URDF and MJCF formats.
  • Joint position, velocity, force, and torque data.
  • Camera and sensor data.
  • Domain randomization for sim-to-real transfer.
  • Reinforcement learning tasks such as locomotion and manipulation.

NVIDIA designed Isaac Gym for high-throughput robot learning. The platform can simulate tens of thousands of environments on a single GPU, depending on your robot model, scene complexity, GPU memory, and physics configuration.

Why GPU Acceleration Changes Robot Learning

Traditional robot-learning pipelines often split the work across CPUs and GPUs:

  1. A CPU physics engine simulates the robot.
  2. The simulator sends observations to a GPU.
  3. The GPU runs the neural-network policy.
  4. The GPU sends actions back to the CPU.
  5. The cycle repeats.

That back-and-forth transfer adds overhead. It also limits how many environments you can run efficiently at once.

Isaac Gym keeps the full loop closer together:

  1. GPU physics simulates robot dynamics.
  2. GPU tensors store state, contacts, and actions.
  3. PyTorch computes observations and rewards on the GPU.
  4. The policy generates batched actions on the GPU.
  5. The simulator advances every environment in parallel.

This design lets you collect experience much faster than a single-environment CPU workflow. NVIDIA's Isaac Gym material specifically highlights that GPU simulation and tensor-based APIs eliminate costly CPU-GPU transfers in the RL loop.

Parallel Environments Do the Heavy Lifting

Suppose you train a quadruped to walk. A single environment gives you one robot trajectory at a time. Isaac Gym can launch thousands of copies of that quadruped, each with different initial positions, terrain conditions, or random disturbances.

Every simulation step produces a large batch of experience:

4,096 environments x 1 action step = 4,096 transitions

The policy then learns from a diverse batch instead of waiting for one robot to stumble through one episode. That massive parallelism gives PPO-style algorithms the data throughput they need.

The robot does not become intelligent because you created more copies of it. It simply fails much faster and learns from those failures in bulk. Efficiency at its finest.

Isaac Gym vs Isaac Sim vs Isaac Lab

NVIDIA's robotics ecosystem uses several similar names, so a quick comparison helps.

Tool Main role Current guidance
Isaac Gym GPU-first RL simulator and prototype platform Useful for legacy projects and learning the original GPU-native workflow
Isaac Sim High-fidelity Omniverse-based robotics simulation Use for photorealistic scenes, perception, sensors, and robotics workflows
Isaac Lab Robot-learning framework built on Isaac Sim Best starting point for new RL, imitation learning, and robot-learning projects

Isaac Lab provides APIs and examples for reinforcement learning, imitation learning, and related robot-learning methods. NVIDIA states that Isaac Lab replaces older robot-learning frameworks, including Isaac Gym Preview Release and IsaacGymEnvs, and encourages users to migrate.

So why learn Isaac Gym at all? Because the core ideas still matter:

  • Parallel simulation.
  • Batched observations and actions.
  • GPU-resident tensors.
  • Vectorized rewards.
  • Domain randomization.
  • Reinforcement learning at scale.

If you understand those concepts, Isaac Lab feels much less mysterious.

What You Need Before You Start

Isaac Gym targets NVIDIA GPU workflows, so check your setup before downloading anything.

You generally need:

  • A supported NVIDIA GPU.
  • Recent NVIDIA graphics drivers.
  • CUDA-compatible software components.
  • Python and PyTorch versions compatible with your Isaac Gym release.
  • Enough GPU memory for your number of environments.
  • A Linux or Windows setup supported by the release you use.

Isaac Gym's performance depends heavily on GPU resources. A simple reaching task may run well with hundreds or thousands of environments, while a complex humanoid task with detailed contacts can consume memory much faster.

Do not begin with 16,384 humanoids, textured scenes, cameras, and every sensor enabled. Start smaller, verify correctness, then scale. Your GPU does not need an immediate stress test disguised as a tutorial.

Install Isaac Gym and IsaacGymEnvs

Isaac Gym historically shipped as a preview release through NVIDIA's developer download page rather than as a normal pip install package. NVIDIA still labels it as a preview/prototype physics simulation environment and directs users toward Isaac Lab for new work.

A typical legacy setup looks like this:

python -m venv isaacgym-env  # create the virtual environment
source isaacgym-env/bin/activate  # activate it
pip install torch torchvision  # install compatible PyTorch first
pip install -e python  # from the extracted Isaac Gym package
git clone https://github.com/isaac-sim/IsaacGymEnvs.git  # NVIDIA example RL environments
cd IsaacGymEnvs
pip install -e .

The IsaacGymEnvs repository contains example reinforcement learning environments associated with the high-performance tasks described in NVIDIA's Isaac Gym work.

Version compatibility matters here. Match the Isaac Gym Preview Release, Python version, PyTorch build, CUDA support, GPU driver, and IsaacGymEnvs requirements. Treat that combination as a set, not a buffet.

Understand the Isaac Gym Workflow

An Isaac Gym RL task follows a familiar environment loop, but it processes almost everything in large GPU batches.

  1. Create the simulator.
  2. Add a ground plane or terrain.
  3. Load a robot asset from URDF or MJCF.
  4. Create many copies of the environment.
  5. Acquire GPU tensors for robot and object state.
  6. Build batched observations.
  7. Run the policy to produce batched actions.
  8. Apply actions to all robots.
  9. Advance physics.
  10. Compute rewards and reset completed environments.
  11. Repeat.

The big difference lies in scale. You do not loop through environments one by one in Python. You operate on tensors containing the state of all environments together.

For example, a tensor might have this shape:

joint_positions.shape == (4096, 12)

That tensor stores the joint positions for 4,096 robots, each with 12 controllable joints. A vectorized reward function can calculate all 4,096 rewards in one GPU operation.

Create a Basic Isaac Gym Simulation

The following outline shows the key setup stages. Exact APIs vary by Isaac Gym release, but the structure stays consistent.

from isaacgym import gymapi

gym = gymapi.acquire_gym()

sim_params = gymapi.SimParams()
sim_params.dt = 1.0 / 60.0
sim_params.substeps = 2
sim_params.up_axis = gymapi.UP_AXIS_Z
sim_params.gravity = gymapi.Vec3(0.0, 0.0, -9.81)

sim = gym.create_sim(
    compute_device_id=0,
    graphics_device_id=0,
    type=gymapi.SIM_PHYSX,
    params=sim_params,
)

plane_params = gymapi.PlaneParams()
plane_params.normal = gymapi.Vec3(0.0, 0.0, 1.0)

gym.add_ground(sim, plane_params)

This code acquires the Gym API, configures physics timing and gravity, creates a PhysX simulator, and adds a ground plane.

The compute_device_id identifies the GPU that handles simulation. The graphics device controls rendering. For headless RL training, you may disable or minimize rendering work because the policy does not need a cinematic view of 8,000 robots falling over.

Load a Robot Asset

Isaac Gym can import robot descriptions from URDF and MJCF files. URDF commonly describes robot links, joints, collision geometry, visual geometry, inertial properties, and joint limits.

Here is a simplified asset-loading example:

asset_options = gymapi.AssetOptions()
asset_options.fix_base_link = False
asset_options.disable_gravity = False

robot_asset = gym.load_asset(
    sim,
    asset_root="assets",
    asset_file="urdf/my_robot.urdf",
    options=asset_options,
)

Then define where each robot starts:

initial_pose = gymapi.Transform()
initial_pose.p = gymapi.Vec3(0.0, 0.0, 0.5)
initial_pose.r = gymapi.Quat(
    0.0,
    0.0,
    0.0,
    1.0,
)

For a fixed robot arm, set fix_base_link=True. For a quadruped or humanoid, let the base move freely so the physics engine can simulate balance and locomotion.

Create Thousands of Environments

This is where Isaac Gym earns its reputation. Instead of creating one robot scene, you create a grid of independent environments.

num_envs = 4096
num_per_row = 64
env_spacing = 2.0

lower = gymapi.Vec3(
    -env_spacing,
    -env_spacing,
    0.0,
)

upper = gymapi.Vec3(
    env_spacing,
    env_spacing,
    env_spacing,
)

envs = []
actors = []

for env_index in range(num_envs):
    env = gym.create_env(
        sim,
        lower,
        upper,
        num_per_row,
    )

    actor = gym.create_actor(
        env,
        robot_asset,
        initial_pose,
        "robot",
        env_index,
        1,
    )

    envs.append(env)
    actors.append(actor)

Every environment contains an independent copy of the robot. You can randomize initial joint states, terrain, target positions, object masses, and friction coefficients per environment.

That variation improves exploration and helps your policy avoid memorizing a single perfect simulation arrangement.

Work with GPU Tensors

Isaac Gym exposes simulation state through tensors. You can wrap those tensors as PyTorch tensors and run calculations directly on the GPU.

A typical pattern looks like this:

import torch
from isaacgym import gymtorch

root_state_tensor = gym.acquire_actor_root_state_tensor(sim)

gym.refresh_actor_root_state_tensor(sim)

root_states = gymtorch.wrap_tensor(
    root_state_tensor
)

The actor root state usually includes position, rotation, linear velocity, and angular velocity for each actor. You can similarly acquire degrees-of-freedom state, rigid-body state, contact-force, and Jacobian-related tensors.

Isaac Gym's tensor API lets developers evaluate environment state and apply actions on the GPU, while its documentation also highlights available sensor data such as position, velocity, force, and torque.

Why Tensor Operations Matter

Avoid code like this:

for env_id in range(num_envs):
    reward[env_id] = slow_python_reward(
        observation[env_id]
    )

That pattern forces Python to handle environments one by one. It defeats much of the speed advantage.

Instead, calculate rewards across all environments at once:

distance = torch.linalg.vector_norm(
    end_effector_positions - target_positions,
    dim=1,
)

rewards = -distance

That one expression calculates a reward for every environment in the batch. It looks almost too simple, which is exactly the point.

Build Observations and Actions

For a robot locomotion task, the policy may need:

  • Joint positions.
  • Joint velocities.
  • Base orientation.
  • Base angular velocity.
  • Previous actions.
  • Commanded velocity.
  • Contact forces.
  • Terrain height samples.

You might construct a batched observation tensor like this:

observations = torch.cat(
    [
        normalized_joint_positions,
        normalized_joint_velocities,
        base_angular_velocity,
        gravity_projection,
        velocity_commands,
        previous_actions,
    ],
    dim=-1,
)

Each row represents one environment. Each column block represents a part of the robot state.

Actions often represent joint targets, desired velocities, or torques:

actions = policy(observations)

joint_targets = default_joint_positions + (
    action_scale * actions
)

For a beginner locomotion task, position targets often make training easier. Torque control offers more realism but demands more careful reward design, tuning, and simulator configuration.

Design Rewards for Robot Learning

Reward functions teach the robot what you actually value. They also give it plenty of chances to surprise you.

For a quadruped velocity-tracking task, a reward might combine:

  • Forward velocity tracking.
  • Upright body posture.
  • Low energy use.
  • Smooth joint actions.
  • Foot-contact quality.
  • Survival time.
  • Penalties for falls.

A vectorized reward might look like this:

velocity_error = torch.sum(
    (commanded_velocity - base_velocity).square(),
    dim=1,
)

tracking_reward = torch.exp(
    -velocity_error / 0.25
)

upright_reward = torch.clamp(
    projected_gravity[:, 2],
    min=0.0,
)

action_penalty = torch.sum(
    actions.square(),
    dim=1,
)

rewards = (
    1.5 * tracking_reward
    + 0.5 * upright_reward
    - 0.01 * action_penalty
)

This reward encourages the robot to match a velocity command, remain upright, and avoid unnecessarily huge actions.

Keep reward terms interpretable. If you add 17 reward components before you train a baseline, you will not know why the robot learned to hop sideways while vibrating like a malfunctioning toothbrush. The reward function design guide covers more of this territory.

Reset Environments Efficiently

In Isaac Gym, different environments finish at different times. One robot falls, another reaches the target, and another somehow stays alive long enough to make you optimistic.

You should reset only the finished environments:

reset_ids = torch.nonzero(
    terminated,
    as_tuple=False,
).squeeze(-1)

if reset_ids.numel() > 0:
    reset_robot_states(reset_ids)
    reset_target_positions(reset_ids)
    episode_lengths[reset_ids] = 0

This selective reset keeps all other environments running. It avoids wasting simulation throughput while preserving diverse trajectories.

Common termination conditions include:

  • Robot base height drops below a threshold.
  • Robot tilts too far.
  • End effector reaches a target.
  • Object leaves a valid area.
  • Episode reaches a maximum number of steps.

Use a time-limit timeout as a separate condition from true failure or success. That distinction matters during RL return calculations.

Train with PPO

PPO became a common Isaac Gym choice because it works well with large batches of parallel simulation data. Isaac Gym includes a basic PPO implementation, but you can also use external RL libraries such as Stable-Baselines3 or frameworks.

At a high level, PPO training works like this:

  1. Run the current policy across thousands of environments.
  2. Collect a rollout of observations, actions, rewards, values, and episode endings.
  3. Estimate advantages.
  4. Update the policy and value network for several epochs.
  5. Collect a fresh rollout.
  6. Repeat.

The speed comes from parallel rollout collection. If you run 4,096 environments for 32 steps, one rollout contains 4,096 x 32 = 131,072 transitions.

That volume gives PPO a rich batch of recent experience. It also means reward bugs can generate tens of thousands of bad samples extremely quickly. Fast simulation speeds up learning and mistakes alike. Fair warning.

Add Domain Randomization

A policy that succeeds in one perfect simulated world may fail instantly on a real robot. Real motors differ. Surfaces vary. Sensors contain noise. Objects weigh slightly more or less than their virtual twins.

Domain randomization fights this simulation-to-reality gap by varying conditions during training. Isaac Gym supports runtime randomization of physics parameters for sim-to-real workflows, as covered in the sim-to-real transfer guide.

Randomize useful properties such as:

  • Mass and center of mass.
  • Friction and restitution.
  • Joint damping and motor strength.
  • Sensor noise.
  • Initial pose.
  • External pushes.
  • Target location.
  • Gravity perturbations.
  • Control latency.

For example, you can train a quadruped under different friction values so it does not panic when a real floor offers less grip than your default simulation surface.

Start with small randomization ranges. Increase them once the policy solves the nominal task. If you randomize every property wildly from step one, the robot may never learn a stable behavior in any environment.

Common Isaac Gym Problems

You Run Out of GPU Memory

Reduce the number of environments, simplify assets, lower camera resolution, reduce terrain complexity, or shorten rollout buffers.

Start with 256 or 512 environments. Scale gradually after you confirm that observations, actions, and rewards behave correctly.

Your Robot Does Not Move

Check your degree-of-freedom properties, action scaling, motor-control settings, joint limits, and actor indices. Also confirm that the policy actually outputs nonzero actions.

A zero-action policy can look calm and sophisticated while accomplishing absolutely nothing.

The Robot Falls Immediately

Inspect initial pose, gravity direction, collision settings, terrain location, joint defaults, and controller gains. Render a small number of environments while debugging.

A visual check often reveals obvious issues, such as a robot spawning halfway inside the ground. Physics engines have opinions about that.

Reward Never Improves

Verify reward signs, action normalization, observation normalization, terminal conditions, and target reachability. Log each reward component separately.

If the energy penalty outweighs movement reward, the optimal policy may simply stand still. Your agent does not lack ambition; it follows the incentives you gave it.

Training Is Slow

Make sure physics, observation construction, reward computation, and action application stay on the GPU. Avoid Python loops over environments, unnecessary CPU tensor conversions, and rendering during headless training.

Should You Start with Isaac Gym Today?

For a new NVIDIA robot-learning project, start with Isaac Lab, not the legacy Isaac Gym Preview Release. NVIDIA explicitly recommends Isaac Lab, which runs on Isaac Sim and replaces prior frameworks such as Isaac Gym and IsaacGymEnvs.

Still, Isaac Gym remains worth studying if you:

  • Maintain an existing Isaac Gym project.
  • Use older research code or benchmarks.
  • Want to understand GPU-native RL design.
  • Need to read IsaacGymEnvs examples.
  • Learn how large-scale parallel robotics training evolved.

The principles transfer directly: vectorize environments, keep data on the GPU, calculate rewards with tensors, reset selectively, randomize physics thoughtfully, and validate each task before launching huge runs.

Want to see 4,096 robots learn in one afternoon? Grab the GPTAstra full course at https://cutt.ly/5yviN6qd — it builds the full GPU-parallel training pipeline from simulator to policy.

Frequently Asked Questions

What is Isaac Gym?

Isaac Gym is NVIDIA's prototype physics simulation environment for reinforcement learning. It uses NVIDIA PhysX for GPU-accelerated rigid-body simulation and exposes simulation data as PyTorch tensors, so physics, observations, rewards, and policy training all run on the GPU.

Should I use Isaac Gym or Isaac Lab for a new project?

Start with Isaac Lab. NVIDIA positions Isaac Lab as the successor to the Isaac Gym Preview Release and recommends it for new robot-learning work. Isaac Gym remains valuable for legacy projects and for learning the original GPU-native workflow.

Why do parallel environments matter?

Thousands of environments produce huge batches of experience every step — 4,096 environments times 32 steps gives 131,072 transitions per rollout. PPO-style algorithms thrive on that throughput, and the full loop stays on the GPU with no CPU-GPU copy overhead.

What do I need to run Isaac Gym?

A supported NVIDIA GPU, recent drivers, CUDA-compatible components, Python and PyTorch versions matching your Isaac Gym release, and enough GPU memory for your environment count. Legacy installs come from NVIDIA's preview download rather than pip.

Why does my run run out of GPU memory?

Start with 256 or 512 environments, simplify assets, lower camera resolution, and shorten rollout buffers. Scale only after observations, actions, and rewards behave correctly. Complex humanoids with cameras consume memory much faster than simple reaching tasks.

How does Isaac Gym help sim-to-real transfer?

Runtime domain randomization varies mass, friction, motor strength, sensor noise, pushes, and latency during training, so the policy learns behavior that survives real-world variation instead of memorizing one perfect simulated world.

Final Thoughts

Isaac Gym made large-scale robot reinforcement learning dramatically faster by putting simulation and learning on the same GPU. Its PhysX-based simulation, PyTorch tensor interface, and massive environment parallelism let researchers train locomotion and manipulation policies at a scale that traditional one-environment workflows struggle to match.

Start with a small number of parallel environments, simple observations, position-based actions, and one clear task. Build a reaching or balance task first, inspect it visually, log reward terms, and scale only after the basics work.

Then move toward thousands of robots, domain randomization, complex terrains, vision, and sim-to-real experiments. Just remember: when you launch 4,096 robots at once, you do not eliminate failure — you simply industrialize it.

Share this article X Facebook LinkedIn Reddit WhatsApp

Related Articles