Isaac Lab on Linux (HPC recipe)¶
This page is the Linux counterpart of isaac-lab.md, which covers the Windows-native install. It was worked out on 2026-09-29 on the NCSU Hazel cluster (RHEL 9.8, glibc 2.34, Slurm, Apptainer, no internet on compute nodes), which is the hard case for Isaac: an older glibc than NVIDIA's wheels require, no root, no Docker, and offline GPU nodes. On an Ubuntu 22.04/24.04 workstation the same steps work without the container (see Ubuntu without a container).
Result. The core package and its GPU tests work natively. Isaac Lab 3.0 works in its kit-less mode (Newton or OV PhysX physics, no Isaac Sim), inside an Apptainer container that supplies a newer glibc. There, the official cartpole smoke run trains, our fleet environment runs on both physics backends, and all 9 Isaac-marked tests pass on both. The Isaac Sim (Kit) path does not install natively on glibc 2.34. The official NGC container, which needs no API key, runs Kit and NVIDIA's own cartpole smoke run, but our fleet env fails on it, because the only 3.0 image is a release candidate that predates release/3.0.0 (see Isaac Sim with Kit).
| Route | Status on Hazel | Why |
|---|---|---|
Core package (pip install -e .[dev]), torch cu126 + Triton |
works, native | plain manylinux wheels |
| Isaac Lab 3.0 kit-less, native venv | fails at install | omniverseclient 2.74.0 (a core Isaac Lab dependency) ships only manylinux_2_35 wheels, and its libomniverse_connection.so really imports GLIBC_2.35 symbols; usd-exchange and ovstage are also 2.35-only |
| Isaac Lab 3.0 kit-less in Apptainer (Debian bookworm, glibc 2.36) | works | the container supplies glibc; the GPU driver comes from the host through --nv |
Isaac Sim 6.1 pip (--extra isaacsim), native |
fails at install | all 25 isaacsim 6.1.0.0 wheels are manylinux_2_35 only |
Isaac Sim through the NGC isaac-lab:3.0.0-rc1 container |
Kit and its cartpole smoke work, our fleet env fails | the only published 3.0 image is rc1 (Isaac Lab 17.0.2); on it our env module loads pxr before SimulationApp starts (see below) |
Versions¶
| Component | Core env (envs/core) |
Isaac Lab env (envs/lab_ctr) |
|---|---|---|
| Host | RHEL 9.8, glibc 2.34, NVIDIA driver 595.58.03 (CUDA 13.2) | same host, container userland python:3.12-bookworm (glibc 2.36) under Apptainer 1.4.2 |
| GPU | L40 46 GB (gpu14), one GPU per job | same |
| Python | 3.11.16 (uv-managed) | 3.12.14 (uv-managed) |
| PyTorch | 2.14.0+cu126 | 2.12.0+cu130 (pinned by Isaac Lab) |
| Triton | 3.8.0 (bundled with the torch wheel) | 3.7.0 |
| Isaac Lab | not installed | release/3.0.0 at 60e28c1 (package isaaclab 25.0.0, isaaclab_tasks 20.0.0) |
| Physics | newton 1.6.0 (MuJoCo-Warp 3.12.0), warp-lang 1.17.0, ovphysx 0.6.3, isaaclab_physx 7.3.0 |
|
| RL | rsl-rl-lib 5.5.1 | |
| Tools | uv 0.12.21 standalone binary | same |
The cu130 torch build that Isaac Lab pins needs driver 580.65.06 or newer on Linux; Hazel's 595.58 is fine. Older clusters with a 55x/57x driver would need the cu126/cu128 build instead, which Isaac Lab 3.0 does not pin.
Cluster rules this recipe follows¶
Everything lives under one directory B on the shared file system (on Hazel /share/hpcproject/<user>/isaacnet), because /home is small. Installs run on the internet-connected transfer partition (sbatch -p xfer), and every test runs in a GPU batch job. Nothing heavy runs on the login node, and no job ssh-es to a compute node. The templates set PYTHONNOUSERSITE=1, and keep the uv, pip, Triton, Warp and XDG caches under B.
The templates are in scripts/hazel/ in the repository. They hard-code the Hazel paths in a B= line and in the #SBATCH -o line, and the account, QOS and GPU type in the #SBATCH header; change those for another cluster.
| Step | Template | Partition | What it does | Time |
|---|---|---|---|---|
| 0 | (by hand) | login | copy the code: git archive main \| ssh hazel "tar x -C $B/repo" (or git clone https://github.com/ZzZTripleZzZ/isaac-net on the login node, since the repository is public) |
|
| 1 | install_core.sbatch |
xfer | fetches the uv binary, creates a Python 3.11 venv, installs torch cu126 and isaac-net[dev] |
1.7 min |
| 2 | core_gpu.sbatch |
gpu | pytest -m gpu, pytest -m "not gpu", benchmarks/bench.py fast and reference backends, bench_nr.py |
10 min |
| 3 | install_isaaclab_kitless.sbatch |
xfer | clones Isaac Lab release/3.0.0, pulls python:3.12-bookworm as a SIF, runs uv sync --extra rsl-rl --extra ovphysx inside it, adds isaac-net (editable, --no-deps) and pytest |
6 min |
| 4 | prefetch_assets.sbatch |
xfer | downloads the cloud USD assets the runs need (ground plane, cartpole) into a local mirror | 1 min |
| 5 | isaac_gpu.sbatch |
gpu | official cartpole smoke on Newton and OV PhysX, the fleet env on every physics backend, pytest -m isaac on Newton and on OV PhysX, the GPU suite in the Isaac env |
15 min |
| 6 | isaac_scale.sbatch (optional) |
gpu | fleet-env throughput grid on both kit-less backends, and bench_nr.py without the GPL presets |
25 min |
| 7 | pull_isaaclab_container.sbatch |
xfer | pulls the official nvcr.io/nvidia/isaac-lab:3.0.0-rc1 image (Isaac Sim included) as a 12 GB SIF |
50 min |
| 8 | isaacsim_ctr_gpu.sbatch |
gpu | Kit cartpole smoke, then the fleet env check, the bench smoke and pytest -m isaac on Kit PhysX (LAB=release binds the release/3.0.0 checkout over the image's sources) |
8 min |
Step 4 copies the mirrored tree to $B/assets by hand once (cp -r $B/tmp/https/omniverse-content-production.s3-us-west-2.amazonaws.com/Assets $B/assets/), and the GPU templates set ISAACSIM_ASSET_ROOT=$B/assets/Assets/Isaac/6.1.
The kit-less recipe¶
Isaac Lab 3.0's recommended install is uv run / uv sync from a source checkout, which creates the environment from the checkout's uv.lock without Isaac Sim. The legacy ./isaaclab.sh --install path from the kit-less page also works in principle, but on a machine without cmake in PATH it runs sudo apt-get update, which fails without a terminal (and should not be attempted on a shared cluster). uv sync does not do that.
The commands inside the container are:
cd $B/IsaacLab # git clone --branch release/3.0.0 https://github.com/isaac-sim/IsaacLab.git (GIT_LFS_SKIP_SMUDGE=1)
UV_PROJECT_ENVIRONMENT=$B/envs/lab_ctr uv sync --extra rsl-rl --extra ovphysx
uv pip install --python $B/envs/lab_ctr/bin/python --no-deps -e $B/repo # isaac-net, keeps Isaac Lab's torch
uv pip install --python $B/envs/lab_ctr/bin/python pytest
The container is started with apptainer exec --nv --bind /gpfs_common,/gpfs_common/share:/share $B/sif/py312-bookworm.sif .... --nv binds the host driver. Both binds are needed on Hazel because /share is a symlink into /gpfs_common, and without them the container cannot see the venv or the checkout. The venv's interpreter is uv's standalone CPython under $B, so the image only has to supply glibc, git (for the two git-sourced Isaac Lab dependencies) and a shell; python:3.12-bookworm (350 MB as a SIF) does. The env is only valid inside the container.
A GPU job then runs, from the isaac-net checkout:
X="apptainer exec --nv --bind /gpfs_common,/gpfs_common/share:/share --pwd $B/repo \
--env TMPDIR=$B/tmp,ISAACSIM_ASSET_ROOT=$B/assets/Assets/Isaac/6.1,PATH=$B/envs/lab_ctr/bin:/usr/local/bin:/usr/bin:/bin \
$B/sif/py312-bookworm.sif"
$X isaaclab train --rl_library rsl_rl --task Isaac-Cartpole-Direct --num_envs 16 --max_iterations 10 physics=newton_mjwarp
$X env ISAAC_NET_PHYSICS=newton python -m pytest -m isaac tests/test_isaac_env.py
$X env ISAAC_NET_PHYSICS=newton python benchmarks/isaac/bench.py --num_envs 256 --num_robots 16 --level L2-legacy --backend triton
The templates also set OMNI_KIT_ACCEPT_EULA=YES for these processes only. This records the project's acceptance of the NVIDIA Omniverse EULA, given for the Windows install and confirmed for these templates. Whether kit-less runs need the variable was not tested.
Choosing the physics backend of the fleet env¶
The fleet env (isaac_net/examples/isaac_fleet_env.py) used to hard-code PhysxCfg, which is PhysX through Isaac Sim, so kit-less Isaac Lab refused it with "Isaac Sim is not installed or not found on PYTHONPATH. ... PhysX backend and Kit visualizer currently requires Isaac Sim." make_cfg(..., physics=...), or the environment variable ISAAC_NET_PHYSICS when it is not passed, now selects one of:
physics |
Isaac Lab config | Needs Isaac Sim |
|---|---|---|
isaacsim_physx (default, unchanged) |
isaaclab_physx.physics.PhysxCfg() |
yes |
ovphysx |
isaaclab_ov.physics.OvPhysxCfg() |
no |
newton |
NewtonCfg(solver_cfg=MJWarpSolverCfg()) |
no |
The sphere assets keep PhysxRigidBodyCfg(disable_gravity=True, ...), which Isaac Lab's own tasks also use under Newton and OV PhysX. For the two kit-less backends make_cfg also sets world gravity to zero, because Newton ignores the per-body disable_gravity flag (a PhysX schema attribute): without the fix the spheres fell from 0.5 m to the ground (height 0.30–0.32 m after 120 steps) and slid against friction, while with it they stay at 0.5 m. OV PhysX honors the flag either way. With a static ground and only gravity-free spheres, zero world gravity is the same scene. The 9 Isaac tests pass with and without the fix, because they check the network, not the trajectories. Because the test scripts and benchmarks/isaac/bench.py build the env through make_cfg, the environment variable reaches them (and the subprocesses of tests/test_isaac_env.py) without new command-line flags. bench.py now also records physics and the min and max sphere height in its RESULT line.
Offline GPU nodes and cloud assets¶
Isaac Lab loads task assets (the default ground plane, the cartpole USD) from NVIDIA's S3 bucket. With no route to the internet, the first run on a GPU node hung in environment creation for over 9 minutes after "Successfully created S3 provider", in omni.client.stat (Isaac Lab checks every locally cached copy against the server before using it), and was cancelled. Isaac Lab's cache under $TMPDIR alone therefore does not make GPU jobs work offline. What works is to download the assets on the xfer node with isaaclab.utils.assets.retrieve_file_path(url) (prefetch_assets.sbatch), which mirrors the S3 layout under $TMPDIR/https/<host>/..., copy that tree to $B/assets, and point ISAACSIM_ASSET_ROOT at $B/assets/Assets/Isaac/6.1. Asset paths are then local files, relative references inside the USD files resolve, and no network call is made. A task that needs other assets needs them added to the prefetch list.
Results (Hazel L40, dedicated)¶
All numbers below are from one L40 allocated to the job by Slurm, with no other process on it (the fleet runs recorded 0% utilization and 3 MiB in use before starting). Unlike the Windows numbers in performance.md and isaac-lab.md, they are not affected by contention. They are single runs.
Core package¶
| Suite | Env | Result | Time |
|---|---|---|---|
pytest -m gpu |
core (torch 2.14 cu126, Triton 3.8) | 26 passed | 95 s |
pytest -m "not gpu" |
core | 231 passed, 11 skipped | 315 s |
pytest -m gpu |
Isaac Lab env (torch 2.12 cu130, Triton 3.7) | 26 passed | 94 s |
The 11 CPU skips are the 9 Isaac tests (no Isaac Lab in the core env), a Sionna comparison (Sionna not installed) and a 5G-LENA table test (tables not generated, by design). benchmarks/bench_nr.py stops at the lena_like presets for the same reason; isaac_scale.sbatch runs it with those two presets removed.
benchmarks/bench.py, ms per control step (submit + step), R = 16 unless stated. The level names are the prototype's (L2 here is the slot-level engine that make_engine calls L2-legacy).
| E × R | L0 graph |
L1 graph |
L1 triton |
L2 graph |
L2 triton |
L2 reference |
L2 compile |
|---|---|---|---|---|---|---|---|
| 16 × 16 | 118.6 | 0.66 | |||||
| 256 × 16 | 0.41 | 2.12 | 0.43 | 11.0 | 0.53 | 122.0 | 0.74 |
| 1024 × 16 | 0.55 | 5.84 | 0.57 | 15.1 | 0.71 | ||
| 4096 × 16 | 1.64 | 17.3 | 1.77 | 28.3 | 2.21 | ||
| 1024 × 64 | 1.64 | 17.3 | 1.69 | 28.3 | 2.36 |
The triton backend costs 0.4–2.4 ms per 100 ms control step up to 65,536 robots, 5–21× less than graph, which replays thousands of small kernels per step. The eager reference of the slot-level engine takes 119–122 ms per step at 16 × 16 and 256 × 16. Under the contention of the Windows measurements the same step took about 1,400 ms, so the "up to about 20×" pessimism stated in performance.md is borne out. compile reaches 0.4–0.7 ms but pays up to 37 s of compilation per configuration. bench_nr.py (reference NR engine, R = 16): legacy NetSlot 90 ms, NR compat 250 ms and NR default (13 subbands, EESM) 318 ms per step at E = 64–256, flat in E because the reference engine is launch-bound.
Isaac Lab smoke runs (kit-less, in the container)¶
| Run | Result | Wall time of the job step |
|---|---|---|
isaaclab train ... Isaac-Cartpole-Direct --num_envs 16 --max_iterations 10 physics=newton_mjwarp |
trains, 886 steps/s, 47.5 s training | 163 s (first run, includes Warp kernel compilation) |
same, physics=ovphysx |
trains, about 540 steps/s, 7.2 s training | 74 s |
same, physics=newton_mjwarp, 4,096 envs × 30 iterations |
trains, about 197,000 steps/s, 15.8 s training | 62 s |
fleet env, ISAAC_NET_PHYSICS=isaacsim_physx |
fails as expected: "Isaac Sim is not installed" | 13 s |
fleet env, newton, 16 × 4, L2-legacy triton |
runs, finite observations | 37 s |
| fleet env, 64 × 16, sphere height after 120 steps | Newton 0.30–0.32 m before the gravity fix, 0.5 m after; OV PhysX 0.5 m both ways | 17–33 s |
fleet env, ovphysx, 16 × 4, L2-legacy triton |
runs, finite observations | 18 s |
pytest -m isaac tests/test_isaac_env.py, ISAAC_NET_PHYSICS=newton |
9 passed | 179 s |
same, ISAAC_NET_PHYSICS=ovphysx |
9 passed | 107 s |
The 9 tests include the three bitwise checks: the fleet env's network, fed by live physics poses, equals a reference-engine replay of the recorded poses, sends, tags and RNG stream through partial resets made by DirectRLEnv, at L2-legacy, L1 and L0DR. The network side is therefore validated on Linux with both kit-less physics backends. The physics side is a different simulator from the Windows validation (PhysX through Isaac Sim), so robot trajectories are not comparable across the two.
Fleet env throughput (kit-less)¶
benchmarks/bench.py in the container, random actions, 100 timed control steps after 20 warm-up steps, network L2-legacy on triton. "Off" is the same env without a network. Steps/s are control steps (0.1 s of simulated time each) per second of wall time.
| Physics | E × R | Robots | Startup (s) | Off (steps/s) | Network on (steps/s) | Network-only (ms/step) | Robot-steps/s, network on |
|---|---|---|---|---|---|---|---|
| Newton | 256 × 16 | 4,096 | 6–7 | 243 | 157 | 2.0 | 642,122 |
| Newton | 1024 × 16 | 16,384 | 7 | 221 | 156 | 2.0 | 2,553,802 |
| Newton | 1024 × 64 | 65,536 | 11–12 | 16.7 | 16.1 | 2.8 | 1,054,808 |
| Newton | 4096 × 32 | 131,072 | 13–14 | 29.4 | 25.2 | 5.6 | 3,308,432 |
| OV PhysX | 256 × 16 | 4,096 | 4–7 | 27.3 | 24.2 | 2.6 | 99,012 |
| OV PhysX | 1024 × 16 | 16,384 | 14 | 19.8 | 18.8 | 2.6 | 307,792 |
| OV PhysX | 1024 × 64 | 65,536 | 126–135 | 10.1 | 10.5 | 2.9 | 685,479 |
| OV PhysX | 4096 × 32 | 131,072 | 365 | 6.2 | 6.3 | 5.7 | 828,455 |
On a dedicated GPU the triton network costs 2.0–5.7 ms per control step up to 131,072 robots. Where physics is slow, which is OV PhysX at every size and Newton from 65k robots up, that is 3–14% of the step. Newton steps 4k–16k robots in about 4 ms, so there the network is about 30% of the step and lowers throughput by 29–36%. Newton starts in seconds at every size, while OV PhysX startup grows with the robot count (365 s at 131k robots), much like the PhysX-through-Kit startup on Windows. Newton with R = 64 is slower than with R = 32 at twice the robots, probably because robot-robot contact pairs grow with R per env. The Newton rows were taken after the zero-gravity fix; a first grid with gravity on (spheres resting on the ground) was 3–40% slower and is superseded. These throughputs are not comparable with the Windows table in isaac-lab.md (different physics, GPU and contention).
Isaac Sim with Kit (PhysX through Isaac Sim)¶
Isaac Sim 6.1's pip wheels need glibc 2.35, so on RHEL 9 the only Kit route is a container. NVIDIA publishes nvcr.io/nvidia/isaac-lab with Isaac Sim included; tags on 2026-09-29 were 2.0.0–2.3.2, 3.0.0-beta1, 3.0.0-beta2, 3.0.0-beta2-post1, 3.0.0-rc1 and 3.0.0-rc1-kitless (no final 3.0.0). The registry issues an anonymous pull token, so no NGC API key is needed. The Isaac Sim image nvcr.io/nvidia/isaac-sim (tags up to 6.1.0) also lists anonymously; it was not pulled.
pull_isaaclab_container.sbatch pulls 3.0.0-rc1 into a 12.1 GB SIF in about 50 minutes on the xfer partition. Its first attempt failed because mksquashfs was killed at the partition's default memory; 22 GB (the xfer QOS allows 24 GB per user) was enough. isaacsim_ctr_gpu.sbatch runs it with apptainer exec --nv --writable-tmpfs, HOME and Kit's cache, data and logs directories bound to $B/kit/, the local asset root, and OMNI_KIT_ACCEPT_EULA=YES. The image contains Isaac Sim 6.1.0-rc.26, Isaac Lab 3.0.0 rc1 (package isaaclab 17.0.2), torch 2.11.0+cu128, Triton 3.6.0 and pytest, and its Python is /isaac-sim/python.sh. An NVIDIA Vulkan ICD file (/etc/vulkan/icd.d/nvidia_icd.json) is present inside the container. The findings were:
- Kit starts headless and the official smoke run trains:
isaaclab.sh train --rl_library rsl_rl --task Isaac-Cartpole-Direct --num_envs 16 --max_iterations 10 physics=isaacsim_physxreached about 600 steps/s, with 14.8 s of training and 116 s for the whole job step. There was no EGL, Vulkan, driver or NGC-authentication problem. Kit only warned that it could not open an X display. - Our fleet env does not start on this image.
SimulationAppfails withAttributeError: module 'omni.usd' has no attribute 'get_context', after the warning "Please check to make sure no extra omniverse or pxr modules are imported before the call to SimulationApp(...)". With rc1's Isaac Lab, importingisaac_net.examples.isaac_fleet_envloadspxr(an import probe showed thatisaaclab.sim,isaaclab_physxand our mixins alone do not), and our scripts import the env module before they launch the app, because they need its config to choose the launcher. Withrelease/3.0.0the same import does not loadpxr, which is why the Windows install and the kit-less runs are unaffected. Kit then exits with status 0, so only the missingCHECK/RESULTline shows the failure, andpytest -m isaacreports all 9 cases as failed. - Binding the
release/3.0.0checkout over the image's/workspace/isaaclabdoes not help. The newer Isaac Lab needs a newer Warp than the image ships, andimport isaaclabfails inwarp.structwithTypeError: issubclass() arg 1 must be a class.
So the Kit route on this cluster needs an image with Isaac Lab release/3.0.0 on Isaac Sim 6.1.0: either a final isaac-lab:3.0.0 tag when NVIDIA publishes one, or an image built with Isaac Lab's own docker/container.py on a machine with Docker and converted with apptainer build x.sif docker-archive://x.tar. The alternative is to make the fleet env module importable without pxr on older Isaac Lab versions. Neither was tried. Until then, use the kit-less route on RHEL-type clusters.
Ubuntu without a container¶
On Ubuntu 22.04 or 24.04 (glibc 2.35/2.39), skip Apptainer and run step 3's commands directly: install uv, clone Isaac Lab release/3.0.0, uv sync --extra rsl-rl --extra ovphysx (add --extra isaacsim for Kit), then uv pip install --no-deps -e <isaac-net>. This was not run here; the only difference from the tested path is the missing container.
Failure modes seen¶
./isaaclab.sh --installrunssudo apt-get updatewhencmakeis not onPATH, which fails in a batch job ("sudo: a terminal is required") and is inappropriate on a shared machine. Useuv sync.- Native
uv syncon glibc 2.34: "Distributionomniverseclient==2.74.0can't be installed because it doesn't have a source distribution or wheel for the current platform ... You're on Linux (manylinux_2_34_x86_64)". Forcing the wheel would not help: itslibomniverse_connection.soimportsGLIBC_2.35symbols (checked withobjdump -T).ovphysxandusd-exchangeonly import up toGLIBC_2.34, butovstagebundles the same library. - Apptainer on Hazel:
/shareis a symlink, so the container needs--bind /gpfs_common,/gpfs_common/share:/share, otherwise "cd: ... No such file or directory". - Offline GPU nodes: without a local asset root, environment creation hangs in
omni.client.stat(see above). apptainer pullof the Isaac Lab image:mksquashfswas killed (out of memory) at the xfer partition's default job memory. The xfer QOS allows at most 4 CPUs and 24 GB per user, and--mem=22Gworks.- Kit in the
3.0.0-rc1image:omni.usdhas no attributeget_contextwhenpxrwas imported beforeSimulationApp. Our fleet env triggers this on rc1's Isaac Lab, and Kit still exits with status 0 (see above). - A30 jobs stayed pending with
QOSGrpGRESfor over an hour (the gpu QOS caps the whole group at 4 A30s), so all GPU results here are on the L40. benchmarks/isaac/bench.pyprinted aSyntaxWarning: invalid escape sequence '\i'from a Windows path in its docstring, now fixed.