Skip to content

Project status

State of 2026-10-06: main is packaged as version 0.2.0 (290195b plus the release commit). Unreleased work sits on integrate5: a usability and visualization wave of six branches (feat/uxcore, feat/uxcli, feat/viztrace, feat/vizlib, feat/onboard, feat/isaacux), listed under Unreleased in CHANGELOG.md and in Recently merged. The earlier waves of integrate2 (obstacles, Rician fading, RLF, TPC / CQI / antenna, RACH / DRX, DL traffic / FDD / REM), integrate3 (frequency-selective fading, QoS scheduler, spec check, known-limits closure) and integrate4 (two triton mirrors, rank-2 MIMO, mini-slot grants) are in 0.2.0.

isaac-net is a research prototype packaged as isaac_net. Every feature branch of the third round of work is merged on main: fast backends for the NR engine, channel models, traffic models, the edge loop, background users, the energy model, multi-GPU sharding, the Wi-Fi level, adaptive fidelity, differentiable models, radio maps from USD scenes, the benchmark suite, the OAI rfsim bridge, the Linux / HPC recipe and the MuJoCo Playground / MJX backend. The speed and scale tables come from an uncontended benchmark campaign on an idle GPU (performance.md), and the 5G-LENA load-gap mechanisms found in fidelity-load-gap.md are NR engine switches with the lena_match_v2 preset. CHANGELOG.md lists what 0.1.0 contains, and RELEASE.md how a release is cut. The paper describing the engine is on arXiv (arXiv:2610.02370).

What is merged

Each row is on main. The commits are the merge or the main component commits, oldest first within a row.

Area Component Commits How it was verified
Engine Prototype levels L0, L0DR, L05, L05Q, L1 and the slot-level uplink L2-legacy, with eager, graph, compile and triton backends e1afdb9 graph bitwise equal to the reference at E×R up to 256×16 and 64×100 over 300 steps (performance.md)
Engine Engine API: partial reset(env_ids), per-env clocks, submit / step dict outputs, fast backends for every level fbfcb44 New reference bitwise equal to the frozen original at all 6 levels; graph bitwise equal with random partial resets; untouched envs bitwise unaffected
Engine Package layout, configurable NR engine as L2, one NRConfig with presets, make_engine, multi-cell NetSlotMC, Sionna BLER tables 971fc12 test_nr_phy (TS 38.214 examples, 564,300-case TBS cross-check against Sionna), test_nr_harq, test_multicell, test_engine_api
Engine Surrogate levels TR, GE, QA, NN and the bounds ORACLE, NOCOMM, fit tool 4e73df5 test_levels; graph bitwise equal to the reference
Engine Multi-cell NR engine: per-cell PF and HARQ, UL and DL same-slot interference, fractional power control, A3 handover 096e86a, 661f780 test_nr_multicell; at one cell bitwise equal to a frozen single-cell engine
Engine Engine-owned random streams and configurable frame buffer, timeout and control step for the prototype, surrogate and bound levels 308fdfe test_rng_levels
Engine NR engine: scheduler variants (pf, pf_wideband, maxci, rr), engine RNG, graph (one or several cells) and triton (one cell, UL and DL) backends e5ffc42 … b79578f test_nr_sched, test_nr_fast (300-step graph == reference bitwise, triton to rounding, traffic sub-steps)
Radio Selectable channel models: log-distance variants, TR 38.901 (8 scenarios, spatially consistent LOS, O2I), radio maps, blockage, per-robot Doppler 388a7fb, 753619f test_channels (hand-computed path loss, LOS share, spatial consistency, default bitwise equal to the old radio, CUDA-graph capture)
Radio Radio maps from USD scenes: exporter with ITU-R P.2040 materials, Sionna RT bake CLI, Isaac hook with a scene-hash cache, warehouse demo a7ba517 test_scene_radio_map, test_scene_bake (slow, needs Sionna RT); synthetic scenes within about 1 dB of closed forms (scene-radio-map.md)
Traffic and stages Traffic models (periodic, bursty, video, event) inside the L2 step 1d533e9 test_traffic, including the graph test
Traffic and stages Edge-computing loop EdgeLoop over any level 0f8aaf9 test_edge (conservation, eager == graph bitwise with partial resets)
Traffic and stages Background users, radio energy model, multi-GPU ShardedEngine 5b64ca4 test_background, test_energy, test_sharded (two shards bitwise equal to one engine, background-energy-sharding.md)
Levels Level WIFI: mean-field 802.11 DCF / EDCA, event simulator, validation tables 2134085 test_wifi; ns-3 802.11ax saturation within 3.1% (wifi.md)
Levels Differentiable fluid models L1D, QAD, neural-proxy recipe 512f4d1 test_diff (exact at temperature 0, gradcheck, finite differences)
Levels Adaptive and mixed fidelity (AdaptiveEngine) 2d96580 … 2400ed9 test_adaptive (threshold 0 / ∞ bitwise equal to the plain levels, handoffs exact, graph == eager)
Validation ns-3 bridges (lockstep, pool, offline replay) and the 5G-LENA replay 651a93e Import tests; lockstep and pool smoke-tested against the lab ns-3 builds
Validation Formal comparison against 5G-LENA, and the load-gap diagnosis with the nr_loadfix.py prototype aac4e17, e4726f1, 5ff3eb9 CSVs and report scripts in benchmarks/fidelity/ (fidelity-vs-lena.md, fidelity-load-gap.md); test_nr_loadfix (all switches off bitwise equal to the engine)
Validation OAI 5G rfsim bridge, campaign, oai_rfsim preset 432c46d … 41b31e7 tests/bridges/oai (fake stack); campaign results in benchmarks/oai/results/ (bridges-oai.md)
Validation Measurement protocol and tools for a lab gNB and POWDER ec0037f, 5738520 test_measure_parsers, test_measure_calibrate on fixture logs
Simulators Isaac Lab layer on make_engine and NRConfig, IsaacNetCfg, observation selection, network domain randomization 2a98bc3, 848bcfa test_isaac_layer; test_isaac_env on the Windows Isaac Lab 3.0 install (in-env bitwise replay through partial resets)
Simulators Linux / HPC recipe: kit-less Isaac Lab on Newton or OV PhysX in Apptainer, Slurm templates ddfa6ed 9/9 isaac tests on both physics backends (isaac-lab-linux.md)
Simulators MuJoCo Playground / MJX backend, MJX fleet env, Brax PPO b871ef3, 9075a09 tests/mjx (in-env == direct torch replay bitwise, zero-copy check)
Benchmark suite isaac_net.bench: four tasks, metrics, baselines, runner, result format, load calibration 4ecb4f5, 32eb045 test_bench (reproducibility, partial-reset isolation, GPU graph == reference episodes)
Docs and packaging Docs site, tutorials, API reference, licensing page; 0.1.0 packaging with extras and console scripts 8cbd1a9, eaa9185, release branch mkdocs build --strict; wheel installed in a fresh CPU venv, suite run from outside the source tree

The CPU suite runs in CI (ruff, pytest -m "not gpu", strict docs build). Installed from the 0.1.0 wheel into a fresh venv with the CPU build of torch and run from outside the source tree, pytest -m "not gpu" gives 690 passed and 21 skipped (Isaac Lab, JAX, pxr, Sionna and the 5G-LENA tables absent). The one other test, the check of the prototype/ shims, needs the source tree and fails there by design. The last full GPU run on the lab box gave 75 GPU tests passed. Whether CI has run green on GitHub since the round-3 merges was not checked for this page.

The following results were produced outside the package code, with scripts in the project's research workspace and on the lab box, and are documented here: the ns-3.48 + 5G-LENA v5.1 reference sweep of 186 runs (validation-5g-lena.md), the public-data calibration that produced the srsran_like and oai_like presets (calibration-public-data.md), and the Isaac Lab 3.0 scale sweep up to 8,192 envs × 128 robots on Windows (isaac-lab.md).

Recently merged

Area What Evidence
Configuration Scenario presets warehouse_private_5g, factory_inf, outdoor_campus, urllc_control (representative, not calibrated; lena_validation_v2 MAC with fading on and the shipped Sionna PDSCH tables), NRConfig.describe() / diff(), make_engine(backend="auto") and resolve_backend (default still "reference"), output_schema() on every engine from core/schema.py, the UnusedFieldsWarning under the default strict=False (feat/uxcore, 622d43d..bd2f3b1, integrate5) test_ux_core (presets build and step on a CPU, triton refusals as the docstrings state, describe / diff, the auto choice with CUDA mocked, schema against real step dicts, the warning once per config, silent with strict=None, raising with strict=True) (configurability.md)
Diagnostics isaac-net-doctor: environment report, CPU / CUDA self-test, --config, --json, --quick (merge 8d9170b, integrate5) test_doctor on a CPU; the CUDA branch of the self-test is to be validated on the lab box (doctor.md)
Traces SlotTrace: slot-level MAC events of selected (env, robot) pairs on L2 reference, export to DataFrame / CSV / Parquet, per-robot summary; MacLink.slot_hook, SlotTap.add_observer; viz.trace.plot_timeline / plot_slot_heatmap (merge 0216c38, integrate5) test_trace (bitwise invisible, also with traffic, access and TPC; events against the step outputs and counters; fast backends refused; CSV round trip) (trace.md)
Recording and plots RecorderLoop / isaac_net.record: per-step KPIs accumulated on the device, Parquet or CSV tables, TensorBoard and W&B logging; the isaac_net.viz package and the viz extra; isaac-net-bench report --html / --figures, run --keep_rows; isaac-net-rem --panels (merge 0df0e3d, integrate5) test_record (bitwise invisible, no host sync in step, Parquet / CSV round trip; the CUDA no-sync test is to be run on the lab box), test_viz (every plot on synthetic records, HTML report); gallery from docs/img/viz/make_gallery.py (viz.md)
Onboarding Colab quick start (tutorial 00), choosing / cookbook / FAQ pages, a shorter landing page, docker/Dockerfile for the kit-less Linux path (hadolint-clean, not built), CI job docs-examples (merge e0bb485, integrate5) the docs-examples CI job runs the notebook and the ten cookbook recipes on a CPU (cookbook.md)
Isaac Lab Registered tasks Isaac-NetFleet-Direct-v0, -Direct-L0-v0, -Direct-Warehouse-v0, Isaac-NetFleet-Manager-v0 with rsl_rl and skrl configs, the manager-based workflow (NetManagerCfg, NetRuntime, net_* terms), viewport overlays NetMarkers (merge 1001121, integrate5) test_isaac_ux on a CPU (registration metadata, manager terms on a fake env, overlay geometry, inert when headless); the Isaac-marked task runs and the overlay screenshot are to be validated on the lab box (isaac-lab.md)
Performance Uncontended benchmark campaign on an idle RTX 4090 and a dedicated L40: every level and backend, the fleet task, Isaac Lab scale up to 1,048,576 robots with the NR engine, the ns-3 cost recomputation CSVs and raw JSONL in benchmarks/results/uncontended/ (performance.md)
Radio Obstacles and NLOS: LOS state from a baked map, a 2.5-D ray march or an Isaac callback (los_source), knife-edge diffraction (los_diffraction), soft LOS, TR 38.901 blockage models A and B (blockage_model), blockers= per step (merge 48d2f03, integrate2) test_obstacles on the synthetic rack hall, defaults bitwise unchanged (obstacles.md)
Radio Rician fast fading with K fixed or from the LOS state (fading_rician), reference, graph and triton (merge 4bb9be6, integrate2) test_rician (K-S distance below 0.002 against the Rician CDF), GPU equivalence lists (channels.md)
Multi-cell Radio link failure and re-establishment (rlf), A3 target admission (a3_min_target_rsrp_dbm) (merge 4ea800f, integrate2) test_rlf, defaults bitwise unchanged (multicell.md)
PHY and MAC Closed-loop UL power control (ul_tpc), the TS 38.214 CQI table (cqi_table="38214"), the TR 38.901 gNB sector antenna (gnb_antenna="sector") (merge 2142c00, integrate2) test_tpc_cqi_antenna (configurability.md, channels.md)
Access Contention-based RACH and connected-mode DRX (rach, drx), DRX sleep power in the energy model (merge da0856c, integrate2) test_access, defaults bitwise unchanged (access.md)
Downlink and tools DL traffic models, DL background, FDD (duplex="fdd"), radio environment map export (isaac-net-rem) (merge 8a665be, integrate2) test_dl_traffic_fdd_rem (configurability.md, rem.md)
Radio Frequency-selective fading: subband correlation from an exponential power-delay profile with a fixed or per-link TR 38.901 Table 7.5-6 delay spread (fading_freq_corr), reference, graph and triton (merge 052bcc7, integrate3) test_freqfade (cross-subband correlation within 0.01 of the model for 10 to 300 ns, defaults bitwise unchanged), GPU equivalence configs ul_fcorr, ul_fcorr_rician (channels.md)
PHY Rank-1/2 SU-MIMO model: rank per new TB from the wideband SINR and the Rician K (n_layers_max, rank_rule, rank_sinr_min_db, rank_k_max_db, rank_layer_penalty_db, ul_mimo, dl_mimo), TBS over the layers, per-layer SINR for MCS and decoding, rank kept per HARQ process, reference and graph (triton refuses it) (merge dbceecf, integrate4) test_mimo (off bitwise equal to the frozen engine, rank-2 TBS within 1.9–2.15× of rank 1 at the same MCS, saturated DL throughput 1.2–2× at high SINR, CPU graph capture path bitwise), GPU equivalence config ul_dl_mimo2 in G1 / G7 (configurability.md)
Backends triton mirror A: closed-loop UL power control (ul_tpc, accumulate and absolute) and the 38.214 CQI table (cqi_table="38214") in the fused kernel (constexprs TPC, CQI38214, per-robot TPC state carried through the step) (feat/tritonA, integrate4) test_triton_tpc_cqi, GPU equivalence configs ul_tpc, ul_tpc_abs, ul_dl_cqi38214, ul_tpc_cqi in G1 / G2 / G7 (configurability.md)
Backends triton mirror B: the RACH / DRX access gate (ACCESS, UL and DL slots in one kernel), FDD with a DL carrier of its own width (FDD, NPRB_D) and the DL traffic arrival gate (DLGATE) in the fused kernel (merge 2989ded, integrate4) test_triton_access_fdd_dl, GPU equivalence configs ul_access, ul_fdd, ul_dl_fdd, ul_dl_traffic in G1 / G2 / G7 (configurability.md, access.md)
MAC Mini-slot (type B) grants: ul_mini_slot_symbols (2, 4 or 7) splits every UL data slot into scheduling occasions, mini_slot_dl the DL slots too; grant, PF, MCS, TBS, decode, HARQ and RLC per occasion, timers in slots, reference and graph (triton refuses it) (merge c88bd3f, integrate4) test_minislot (off bitwise equal to the frozen slot loops, occasion timing, capacity, conservation, E-independence), GPU equivalence configs ul_minislot2, ul_minislot4 in G1 / G7 (configurability.md)
MAC QoS scheduler after 5G-LENA NrMacSchedulerOfdmaQos (scheduler="qos", qos_classes, qos_priority, qos_pdb_ms, qos_gamma), class-ordered byte assignment in one queue per robot, reference, graph and triton (merge c32c5eb, integrate3) test_qos (Q1 bitwise equal to the frozen engine when off), GPU equivalence configs qos, qos_rbg (configurability.md)
Validation Spec check against ETSI TR 138 901 V17.0.0 and TS 138 214 V17.1.0: blockage model A now applies its loss only inside the angular window of §7.6.4.1, every [verify] marker resolved (commits d3c04f3, 135d90e, integrate3) test_obstacles; the check is recorded in obstacles.md and channels.md
Engine Known limits closed: DL traffic models on graph, the EdgeLoop delay return path on the FDD DL carrier, RLF re-establishment through RACH, the sector pattern at bake time, all triton refusals in NRTritonEngine.__init__ (merge 63941bb, integrate3) test_limits_closed, tests/limits_off_scenarios.py with ten switch-off goldens; items 18, 19, 20 and 22 below
Validation 5G-LENA load-gap mechanisms as NRConfig switches (pf_update, pf_avg_idle, ul_retx_sched, ul_amc_alloc, ul_grant_model), defaults bitwise unchanged, presets lena_match_v2 and lena_validation_v2 with no fitted parameter; graph equals the reference exactly, triton covers all but the BSR grant pipeline 153-run replay in benchmarks/fidelity/results/loadfix_v2/ (fidelity-vs-lena.md, fidelity-load-gap.md); test_nr_loadfix matches a frozen copy of the prototype exactly

Open items, in priority order

# Item Why it matters Pointer
1 NR engine triton for several cells, and at large R graph covers one or several cells and triton one cell. The fused kernel holds one env's robots in one program, so large R needs a tiled kernel, and several cells need per-cell schedulers and interference inside it performance.md, ARCHITECTURE follow-up 1
2 Remaining differences to 5G-LENA after the load-gap fixes The engine's UL MCS is one step above 5G-LENA's for about half the UEs at 3–15 dB SNR, and the first-transmission BLER differs by a few tenths of a point, causes not isolated. The TBS and overhead check against NrUlMacStats.txt is open fidelity-load-gap.md
3 Lab measurement campaign Public data cannot validate multi-UE uplink contention, per-TB BLER against SINR, the configured SR period and proactive grants, latency with large frames under load, an indoor channel, or fading correlation at robot speeds. The protocol and tools are ready for the lab gNB (up to 4 UEs) and POWDER (up to 2 UEs) measurement-protocol.md
4 Multi-UE tails against OAI No preset reproduces OAI's multi-UE delay tails or its saturation at 4.8 Mbit/s of cell load bridges-oai.md
5 Multi-cell against a reference P0 and α of the uplink power control come from a sanity sweep, and multi-cell has not been compared with any reference simulator. The downlink has no power control or interference coordination. Radio link failure is done (multicell.md), but its defaults have not been compared with 5G-LENA RLF traces multicell.md
6 16 HARQ processes at overload in the netslot_compat geometry With 16 processes, loss is lower at light load (0.353 vs 0.390) but higher at overload (0.797 vs 0.724), likely because UEs in trouble keep taking PRBs for new TBs while retransmissions block in-order delivery validation-5g-lena.md
7 Isaac Sim (Kit) on Linux clusters Kit-less Isaac Lab works in Apptainer, but the only Isaac Lab 3.0 Kit container is a release candidate on which the fleet env fails. It needs the final isaac-lab:3.0.0 image or a custom image isaac-lab-linux.md
8 Isaac startup and memory at scale Scene startup grows by about 1.1 ms per robot (19 minutes at 1M robots). Options: clone_in_fabric=True, one multi-instance asset per env, or a kinematic pose integrator isaac-lab.md
9 Run Isaac jobs on Windows as a console user CUDA is unavailable from the ssh session on the Windows lab box, so every GPU run there was a one-shot SYSTEM task isaac-lab.md
10 MJX host sync and NetModule capture A masked reset in the core would remove the one host sync per step, and one CUDA graph for the whole NetModule step would cut its host-side launches backends-mjx.md
11 Differentiable models at full scale The full benchmark runs (benchmarks/diff/run_all.sh), the sign and magnitude comparison against L2-legacy finite differences, and the fit quality of the neural proxy are pending differentiable.md
12 Shard invariance of L2 with traffic models ShardedEngine is bitwise one engine on L2 (one or several cells), multi-cell L2-legacy and WIFI under rng="engine". L2 with traffic models or background users (one TrafficGen generator per engine), and any level with an edge loop that draws from the global RNG, still get derived per-shard seeds, so its results depend on the split background-energy-sharding.md
13 Wi-Fi scope The mean-field model loses accuracy for small contention windows (VO-only saturation, −41% at 10 stations) and has no downlink, OFDMA or MU-MIMO. Validation tables B, D and E predate the Poisson-cap and FER channel-time fixes and are to be regenerated with validate.py (which now has FER = 0.2 rows) wifi.md
14 Bridge follow-ups Stale frames are not purged from the ns-3 RLC queue, partial reset works only in the process-per-env mode, and the Windows launcher of the pool is untested. The pose-interpolation fix in the lockstep bridge program needs a rebuild on the lab box bridges.md
15 Smaller PHY and MAC gaps MIESM, FR2 numerology, MIMO layers and DL power control are not modelled, and the UL DCI-slot constraint for K2 is missing. Done: FDD (Duplexing), closed-loop UL power control, the 38.214 CQI table and the gNB sector antenna (configurability.md) configurability.md
16 Reference scenario in the repository The standalone 5G-LENA scenario netslot-ref.cc, its crash-guard patch and the sweep scripts live outside this repository validation-5g-lena.md
17 Public release Done: the repository is public, 0.1.0 and 0.2.0 are on PyPI (pip install isaac-net), and the README carries the citation; remaining: a trusted-publisher release workflow RELEASE.md
18 DL traffic models on the graph backend Done (feat/limits): the DL arrival gate reads static buffers refilled before each replay, as the UL one does, so graph is bitwise equal to the reference with DL models (CPU code-path test; the CUDA test test_dl_traffic_graph_cuda_bitwise is to be run on the lab box); triton runs them too since feat/tritonB configurability.md
19 EdgeLoop delay return path under FDD Done (feat/limits): EdgeConfig(return_path="delay") sizes the command rate on the DL carrier (dl_nprb) configurability.md
20 RLF re-establishment through RACH Done (feat/limits): with rach=True the robot re-establishes through contention-based RACH toward the selected cell, so the outage includes the access delay, collisions and backoff; with rach=False the fixed reest_delay_ms is bitwise unchanged. Contention-free access and T301 are not modelled multicell.md, access.md
21 triton mirrors The fused kernel still refuses several cells, the SR / BSR grant pipeline, SINR hooks, rank-2 MIMO (n_layers_max=2) and mini-slot grants (ul_mini_slot_symbols) (in NRTritonEngine.__init__, SINR hooks at the first step), so these run on reference and graph only. Done: closed-loop UL TPC, the 38.214 CQI table, RACH, DRX, FDD and DL traffic models run in the kernel configurability.md, access.md
22 Sector pattern in the Sionna bake Done (feat/limits): bake.py --gnb-antenna sector --cell-azimuth ... applies the TR 38.901 pattern to the baked map along the direction of each grid point (not per traced ray) and records it in the metadata; RadioMC refuses gnb_antenna="sector" with such a map. Tested on the synthetic map tool; a bake with Sionna RT installed is still to be run scene-radio-map.md, channels.md
23 5G-LENA re-validation after the MAC and radio-RNG fixes Done (2026-10-05): the engine side of the whole 5G-LENA sweep was replayed on the fixed engine; lena_validation_v2(), "v2 minus BSR" and the held-out carrier are bitwise identical, so Table II and the held-out claim stand; v1, NRConfig() and the legacy rows moved by small amounts (6 of 153 v1 runs touched by the retransmission-priority fix). Report in benchmarks/fidelity/results/rerun_2026-10-05/ fidelity-vs-lena.md
24 Lab-box validation of the usability wave The branches were tested on the CPU paths. To be validated on the lab box: the CUDA branch of the isaac-net-doctor self-test, the recorder's CUDA no-sync test (test_recorder_no_sync_cuda), the Isaac-marked cases of test_isaac_ux (tests/scripts/isaac_tasks_check.py builds and steps the registered tasks with overlays on), a short training run of each registered task, and the NetMarkers screenshot docs/img/isaac_markers.png that isaac-lab.md still waits for doctor.md, isaac-lab.md
25 Docker image docker/Dockerfile passes hadolint but has not been built or run; the core stage and the isaaclab stage (pinned to Isaac Lab 60e28c1) are to be built on a Linux machine with Docker and a GPU docker/README.md
26 Scenario presets are not calibrated The four presets are representative starting points: only their lena_validation_v2 MAC was compared with 5G-LENA, and their channel, fading, multi-cell, QoS and mini-slot settings have no reference measurement configurability.md

Suggested starter tasks

A good first contribution touches one module, comes with a test, and does not need the shared lab GPU for long. The tasks below are ordered roughly from least to most context needed.

  1. Run the suite and read one engine end to end. Install the package on a CPU machine, run python -m pytest -m "not gpu and not slow", and read isaac_net/core/proto/netsim.py (the eager reference of every prototype level) alongside tests/test_mac_l2.py, which pins the HARQ and RLC timeline slot by slot.
  2. A benchmark-suite baseline. Train a new baseline on one task of isaac_net.bench and submit its result files as described in benchmark-suite.md. No engine change is needed.
  3. Bring the reference scenario into the repository. Move netslot-ref.cc, parse_run.py and the crash-guard patch script into isaac_net/bridges/ns3/ref/ next to the bridge programs, with a build script that follows the recipe in validation-5g-lena.md. No ns-3 or 5G-LENA source may be copied in.
  4. A scheduler or PHY gap. Pick one of the gaps in item 15, for example DL power control, add it behind an NRConfig field whose default keeps the engine bitwise, and test it next to tests/test_nr_sched.py.
  5. Masked reset. Add reset_mask(mask [E]) to one level next to reset(env_ids), drawing for all envs and keeping the masked rows, with a test that the kept rows match reset(env_ids) in distribution and that untouched envs are bitwise unaffected (item 10).
  6. Tiled NR triton kernel. Split the robots of one env over several programs in nr_triton.py and check the result against the reference with the harness in tests/nr_equiv.py. This is the largest item and the one with the highest payoff at scale.