Fidelity levels¶
Every level comes from make_engine(level, E, R, device, config, backend) and honors the engine contract. A task switches fidelity by changing the level string.
| Level | Model | Parameters | Typical use |
|---|---|---|---|
L0 |
i.i.d. lognormal delay and loss per message, no queue interaction | l0_delay_median_steps, l0_delay_log_sigma, l0_loss from the config, or params={"mu", "sig", "p"} |
the usual randomized-delay baseline |
L0DR |
L0 whose median delay, spread and loss are redrawn per environment at every reset |
dr_* ranges from the config |
domain randomization over delay |
L05, L05Q |
delay and drop looked up in tables fitted offline from L2 rollouts, per cell of backlogged robots × SNR × class (L05Q also by the robot's own queue) |
params={"q", "pdrop"} (required) |
cheap state-conditioned delay |
L1 |
fluid slot model: robots with queued data share each uplink slot equally, FIFO queues | l1_eta (goodput factor) |
contention without MAC detail |
L2 |
configurable NR MAC and PHY: 3GPP MCS/TBS and BLER tables, multiple HARQ processes, optional downlink, 1 to 7 cells with interference, power control and handover | the whole NRConfig |
the fidelity model |
L2-legacy |
the prototype slot-level MAC and PHY, frozen; also runs 1 to 7 cells with a multi-cell config | application fields; the cell block for multi-cell | reproducing earlier prototype runs, and speed at scale |
TR |
trace replay: each environment replays one recorded L2 or L2-legacy env-episode, open loop |
fit file (required) | the replayed-trace baseline |
GE |
three-state Markov-modulated delay and loss, one chain per environment | fit file (required) | a Gilbert–Elliott-style baseline |
QA |
analytic processor-sharing queue per control step, FIFO service, scheduling-request delay | fit file, or an uncalibrated default | contention without slot simulation |
NN |
learned stateful surrogate: an MLP predicts a drop probability and delay quantiles from features at submission | fit file (required) | the learned-surrogate baseline |
ORACLE |
every message delivered at capture, delay 0, never lost | none | upper bound on what any network gives a task |
NOCOMM |
no message is ever delivered | none | lower bound: the task without communication |
Semantics shared by the levels¶
- Message model. Every level uses the same per-robot FIFO of
F = 16message slots, the same 20-step (2 s) application deadline and the same output dict. The prototype levels, the surrogates and the bounds are compiled around these constants and a 100 ms control step, somake_enginerefuses a config that asks for other values. The NR engineL2readsframe_buffer,timeout_stepsandcontrol_step_msfrom the config. - Radio.
L0toL1, the surrogates and the bounds use the fixed legacy radio (one gNB at the origin, log-distance path loss with correlated shadowing, a −90 dBm noise floor).L1andQArefuse any other cell setting.L2and multi-cellL2-legacybuild their radio from the config, and only they acceptn_cells > 1(up to 7, with uplink power control on by default). - Config fields. Each level reads only some
NRConfigfields.NRConfig.unused_fields(level)lists the non-default fields a level ignores, andmake_engine(..., strict=True)raises on them (see Configuration). - Resets. Every level keeps per-env clocks and exact partial resets. The surrogates draw their reset state (the replayed trace of
TR, the initial state ofGE) from the engine generator.
Backends per level¶
| Level | reference |
eager |
graph |
compile |
triton |
|---|---|---|---|---|---|
L0, L0DR, L05, L05Q |
✓ | ✓ | ✓ | ✓ | |
L1, L2-legacy (one cell) |
✓ | ✓ | ✓ | ✓ | ✓ |
L2-legacy (multi-cell) |
✓ | ||||
L2 (one or several cells) |
✓ | ||||
TR, GE, QA, NN, ORACLE, NOCOMM |
✓ | ✓ | ✓ |
For the surrogates and bounds, reference and eager run the same graph-safe code, and graph captures it. The fast backends of the prototype levels need a CUDA GPU.
Fitting the surrogates¶
TR, GE and NN exist only as fits, and QA is calibrated by the same tool. The fit rolls out L2 or L2-legacy on the example fleet task under a behavior policy, logs every message through the public API, and writes all four surrogates into one file outside the repository:
python -m isaac_net.tools.fit_levels --source L2-legacy --task T1 --device cuda --backend graph
python -m isaac_net.tools.fit_levels --source L2 --preset netslot_compat --task T1 --out ~/fits/T1_nr.pt
The default location is ~/.cache/isaac_net/levels/<source>_<task>.pt, or $ISAAC_NET_LEVELS_DIR. The tool refuses a path inside the source tree, and fitted files are never committed. A JSON summary of the fit is written next to the file. Load the file with params:
net = make_engine("NN", E, R, device, params="~/.cache/isaac_net/levels/L2-legacy_T1.pt", backend="graph")
The file records the message sizes it was fitted with, and make_engine refuses an engine whose msg_sizes differ. The surrogates share the prototype constants, so the fit tool refuses a source config with a different frame buffer or timeout.
fit_levels
¶
fit_levels(source='L2-legacy', E=64, R=16, episodes=12, test_episodes=4, T=300, sizes=TASK_SIZES['T1'], device='cuda', preset=None, backend='reference', seed=31337, nn_steps=6000, qa_envs=32, qa_etas=(0.4, 0.55, 0.7, 0.85, 1.0, 1.15), log=print, frame_buffer=None, timeout_steps=None, control_step_ms=None, slots_per_step=None)
Roll out source, fit TR / GE / NN, calibrate QA; returns (fit dict for torch.save, info dict).
frame_buffer, timeout_steps, control_step_ms, slots_per_step: overrides of the preset's application fields
(None = the preset's); the fit records the values it used.
save_fit
¶
torch.save the fit (plus a JSON summary next to it). Refuses paths inside the source tree.
load_level_params
¶
Parameters of one level from params, which is one of
None QA falls back to QA_DEFAULT (uncalibrated); ORACLE / NOCOMM need nothing;
TR, GE and NN raise, because they only exist as fits
a path (str / PathLike) to a fit file written by isaac_net.tools.fit_levels (or the legacy
baseline fitter), loaded with torch.load(weights_only=True)
a fit-file dict {"TR": ..., "GE": ..., "QA": ..., "NN": ..., "meta": {...}}: the entry of level
the level's own dict used as is
A fit file records the message sizes it was fitted with (meta["sizes"]); they must equal sizes. It also records
the frame buffer, timeout, control step and UL slots per step (FIT_APP_KEYS); they must equal app (a dict as
returned by fit_app), because the fitted delays are in control steps of that configuration.
Level classes¶
These classes are what make_engine returns for the surrogate and bound levels. Build them through make_engine, which checks the config first.
NetTR
¶
NetTR(E, R, device, sizes, params=None, backend='reference', inject=False, seed=None, fb=F, timeout=TIMEOUT, ul_per_step=_ns.UL_PER_STEP, rng='global')
Bases: LevelNet
Trace replay. At every reset, env e picks one recorded trace j uniformly (engine generator).
A frame of class c captured at env step t gets an outcome from trace j: the outcomes of the class-c frames
captured at t* are indexed by start[j, c, t] and cnt[j, c, t] (the fit resolves t* = the nearest step with
at least one class-c frame, ties to the earlier step), and one of them is drawn uniformly. If the trace has
no class-c frame at all (cnt = 0), the draw comes from the pooled class-c outcomes. The frame is delivered
at t + delay, or lost (delay = inf). Steps beyond the recorded horizon use its last step. Position, SNR,
queue state and the policy's sends are ignored.
NetGE
¶
NetGE(E, R, device, sizes, params=None, backend='reference', inject=False, seed=None, fb=F, timeout=TIMEOUT, ul_per_step=_ns.UL_PER_STEP, rng='global')
Bases: LevelNet
K-state Markov chain per env (Gilbert-Elliott style), one transition per control step. The state s_t
applies to the frames captured at step t. Each (state, class) has a loss probability p [K,C] and a
delivered-delay distribution: the 101-point empirical quantile function q [K,C,101] or, if the fit has
no q, a lognormal (mu, sig [K,C]). The initial state of a reset env is drawn from pi0 [K] (engine
generator); the chain is exogenous, so the policy's sends never change it.
NetQA
¶
NetQA(E, R, device, sizes, params=None, backend='reference', inject=False, seed=None, fb=F, timeout=TIMEOUT, ul_per_step=_ns.UL_PER_STEP, rng='global')
Bases: LevelNet
Analytic per-step processor-sharing queue, once per control step (not per slot).
n = backlogged robots in the env after arrivals. Each gets a time share of min(S/n, n_max) subbands (n_max =
the L2-legacy power-headroom cap), with its power split over more than one subband as in L1. With pf, the
per-subband SNR gains the analytic PF multi-user-diversity term 10 log10(H_n) + RAYLEIGH_DB (H_n = the n-th
harmonic number). The byte rate per step is b = share * eta * 0.75 log2(1 + SNR) (SE capped) * BYTES_PER_SE
* K (UL slots per control step). A robot whose queue was empty at the end of the previous step starts SR_DELAY slots late. Frames are
served FIFO in continuous time: a frame finishes at t + off + (bytes through it) / b if that is within the
step. Params: eta (default 0.9) and pf (default True), calibrated by the fit tool.
NetNN
¶
NetNN(E, R, device, sizes, params=None, backend='reference', inject=False, seed=None, fb=F, timeout=TIMEOUT, ul_per_step=_ns.UL_PER_STEP, rng='global')
Bases: LevelNet
Learned stateful surrogate (MimicNet / DeepQueueNet style). At submit every enqueued frame is dropped with
probability sigmoid(logit), otherwise its delay is sampled by inverse-CDF interpolation through the Q
quantiles. The history features (hist, to) come from this engine's own sampled outcomes, so the model is
autoregressive. With fifo (default), a delivered frame cannot arrive before the frames queued ahead of it,
and a lost frame ahead blocks until its own timeout.
Params (the fit file's "NN" entry): state (DelayNet state dict), xm / xs (input normalizer), din, Q, h, and
optionally fifo. The MLP runs on every robot at every submit (fixed shape); only enqueued frames use it.
NetOracle
¶
NetOracle(E, R, device, sizes, params=None, backend='reference', inject=False, seed=None, fb=F, timeout=TIMEOUT, ul_per_step=_ns.UL_PER_STEP, rng='global')
Bases: LevelNet
Every accepted message is delivered at its capture step t, delay 0, never lost.
NetNoComm
¶
NetNoComm(E, R, device, sizes, params=None, backend='reference', inject=False, seed=None, fb=F, timeout=TIMEOUT, ul_per_step=_ns.UL_PER_STEP, rng='global')
Bases: LevelNet
No accepted message is ever delivered; each one times out after the application deadline.