The most dangerous failure in production reinforcement learning (RL) is often not a crashed job. It is a job that keeps running while it optimizes a trajectory the system never actually generated.
Here is how that happens. The rollout engine samples tokens from the model, and the trainer later recomputes the probability of those same tokens to compute the gradient. If the trainer re-tokenizes a tool call and gets different token IDs, or a mixture-of-experts (MoE) layer picks different experts on the trainer side, the two sides now describe different policies, even with identical weights. Kernels, numerical precision, and batch shapes cause a third kind of drift, in the numerics. Fully asynchronous execution, where generation keeps running while the trainer updates the weights, adds a fourth gap: samples that come from an older weight version.
Miles v0.1 does not introduce a new loss. Its central contribution is to treat this train–rollout mismatch as a core systems problem. Miles breaks “the rollout engine and the trainer run the same policy” into a few concrete guarantees that it can check one by one. On top of that foundation it builds fully asynchronous scheduling, low-precision execution, memory offload, and weight synchronization for trillion-parameter models.
This post combines the paper, the embedded eight-minute talk, the accompanying 38-slide deck, and the Miles source code. It assumes you know LLMs, GPUs, and basic RL: policies, rewards, and algorithms such as Proximal Policy Optimization (PPO) and Group Relative Policy Optimization (GRPO). It does not assume you have looked inside an RL training system. The sections build on each other:
- §1–2: the training loop and the four ways the two sides can disagree.
- §3–4: the fixes for tokens, MoE routing, and numerics.
- §5–7: the machinery for speed and scale: fully asynchronous training, low precision, memory, and weight transport.
- §8–9: the two training backends, recipes beyond RL, and an end-to-end GLM-5.2 case study.
- §10–11: what the source code reveals and a staged path for adopting Miles.
Companion Materials
Video
Open the video on YouTube. Quick links: 00:00 Overview · 01:28 Train–rollout mismatch · 03:07 Low precision · 03:49 Fully async · 05:15 Weight updates · 07:17 GLM-5.2 case study.
Slides
Open the deck in Google Slides.
1. The System at a Glance: Three Stages, Four Core Objects
Classic RL for LLMs alternates two steps: generate a batch of short answers, then update the model. Agentic RL stretches each answer into a long-running chain. The model holds a multi-turn conversation, calls tools, waits for a sandbox, and resumes generating. Meanwhile rollout needs low latency, the trainer needs throughput, and a very large MoE must span many GPUs.
Miles keeps this manageable by splitting the job into three stages with clear boundaries:
- Rollout. SGLang inference engines generate trajectories. In agentic RL, each rollout session talks to its own isolated environment, which executes the model’s actions and produces the reward.
- Training. The trainer, built on NVIDIA Megatron-LM or PyTorch FSDP (Fully Sharded Data Parallel), consumes completed trajectory groups, computes the RL loss, and updates the policy.
- Weight update. After each training step, Miles sends the new weights to the rollout engines while interrupting in-flight rollouts as little as possible.
The stages do not have to take turns. In lockstep, the trainer waits for the slowest trajectory, and then the engines wait for the optimizer. The fully asynchronous mode in Section 5 lets the engines keep generating while the trainer works.
Four objects in this loop are easy to conflate:
- Prompt: one task drawn from the dataset.
- Trajectory: one complete attempt by the policy, potentially spanning many turns.
- Group: multiple trajectories for the same prompt. Algorithms such as GRPO compare rewards within a group, so an attempt counts as good when its reward is above the group’s average.
- Session: one multi-turn conversation as seen by the serving layer. A session can produce one or more trajectories (Section 3 explains when).
A concrete picture: the prompt is “make this failing test pass.” The policy makes several attempts, which form one group. In the simplest case each attempt is one session: the model reads files, runs the test in its sandbox, and edits code over many turns. That session then yields one trajectory.
The synchronous driver that runs these stages is short. For each rollout step it:
- collects a batch of rollout data;
- trains on it;
- saves a checkpoint when one is due;
- sends the new weights back to SGLang;
- runs evaluation when it is due.
The source deliberately keeps this layer close to pseudocode and pushes complexity into components with well-defined boundaries.
Implementation notes
- This loop lives in
train.py, which refuses--fully-async; the asynchronous driver istrain_async.py(Section 5). - When the trainer and the rollout engines share GPUs, the loop also moves model state and caches on and off the GPU around each phase (
--offload-train,--offload-rollout; Section 7.1). - After the final rollout step, the loop skips the weight update unless an evaluation is still due.
One detail matters for correctness. Before the first optimizer step, Miles runs one weight update, pushing the trainer’s freshly loaded weights to the rollout engines. SGLang therefore starts from weights that went through the trainer’s own loading and conversion path, not from a checkpoint it loaded separately, so both sides begin with exactly the same weights.
Disk-delta transport is the one exception: its first update publishes nothing, and both sides start from the same checkpoint on disk instead (Section 7.2.4). An optional preflight check, --check-weight-update-equal, can compare the engine weights right after this initial sync (Section 7.2.5).
2. Four Kinds of Mismatch
Policy-gradient math assumes that the trainer scores exactly the tokens the rollout engine sampled, under the same policy that sampled them. In a production system that assumption can fail in at least four separate places, and Miles treats each one as its own contract with its own mechanism.
The simplest health check is a per-token ratio. For every sampled token, the rollout engine records the log-prob it assigned, $\ell_t^{\mathrm{roll}}$. Before the update, the trainer recomputes the log-prob of that same token with its pre-update weights, $\ell_t^{\mathrm{train,old}}$. Ideally the two agree:
\[m_t = \exp\left(\ell_t^{\mathrm{train,old}}-\ell_t^{\mathrm{roll}}\right)=1\]In plain terms: if rollout gave a token probability 0.40 and the trainer also computes 0.40, then $m_t=1$. If the trainer computes 0.44 instead, $m_t=1.1$, and that token’s gradient is computed under a slightly different policy from the one that actually produced it. Section 4.2 returns to this ratio in detail.
The table shows what breaks at each of the four places and how Miles responds:
| Layer | What breaks | Miles’s response | Cost or limitation |
|---|---|---|---|
| Token | Text parsing, tool-call JSON reordering, or reapplying the chat template changes the token sequence | The TITO (Token-In-Token-Out) session server preserves the original token IDs (Section 3) | The current session server does not support image or video input |
| Routing | MoE top-k expert selection is sensitive to tiny numerical errors | R3 (Rollout Routing Replay) replays the rollout’s expert assignments (Section 4.1) | Long sequences require additional storage and transfer |
| Numerics | Different kernels, precision, batch shapes, or parallel layouts | TIS (Truncated Importance Sampling) or clip-or-pop, which drops out-of-range tokens (Section 4.2), or true-on-policy mode (Section 4.3) | Correction does not eliminate the mismatch; strict alignment costs throughput and is registered only for dense Qwen3 0.6B/4B |
| Time | In fully async mode, trajectories come from older weights | Tag every token span with its weight version, measure staleness at consumption, and optionally cap it (Section 5) | Training remains off-policy; the lag is bounded only if you set a staleness limit |
These mechanisms are not interchangeable, because each one guards a different equality. Written per token $t$, the four contracts are:
\[z^{\mathrm{roll}}_{<t}=z^{\mathrm{train}}_{<t},\qquad A^{\mathrm{roll}}_{j,t}=A^{\mathrm{train}}_{j,t},\qquad \ell^{\mathrm{roll}}_t=\ell^{\mathrm{train,old}}_t,\qquad v_t=v_{\mathrm{current}}.\]- Same tokens. $z_{<t}$ is the token prefix before position $t$. Both sides must condition on identical token IDs, not merely on the same text. TITO enforces this.
- Same routes. $A_{j,t}$ is the set of experts selected at MoE layer $j$ for token $t$, such as experts
{2, 7}. R3 enforces this. Dense models have no routes, so this contract does not apply to them. - Same log-probs. $\ell_t$ is the log-prob of the sampled token. True-on-policy alignment enforces this for the configurations it covers.
- Same weight version. $v_t$ is the weight version that generated token $t$. Fully async execution breaks this contract on purpose; the asynchronous buffer measures how far it is violated and, if you set a limit, rejects groups that exceed it.
The first three contracts build on each other. A routing record is only meaningful if it belongs to the same token, and log-probs can only match once the tokens (and, in an MoE model, the routes) match. Each step is still a separate guarantee: matching tokens do not imply matching routes, and matching routes do not imply matching numerics. The fourth contract is independent of the other three. Exact tokens, routes, and numerics do not make an old trajectory current.
TIS and clip-or-pop do not establish any of these equalities. They limit how much a remaining policy-ratio mismatch can move the gradient. That remainder includes numerical drift and, under fully async execution, weight staleness.
Two other mechanisms in this post are easy to mistake for mismatch fixes. Faster P2P (peer-to-peer) or disk-delta weight delivery (Section 7) can shorten version lag, but it cannot correct a sample that is already stale. OPD (On-Policy Distillation, Section 8) changes the learning signal; it does not make rollout and trainer the same policy again.
3. TITO: Do Not Turn Tokens Back into Text and Guess Again
A multi-turn agent turns the model’s tokens into text, hands that text to its harness (the agent program that parses replies, runs tools, and builds the next request), and turns the text back into tokens on the next turn. That round trip can be lossy. The next turn, and later the trainer, can end up with a different token sequence from the one the model actually sampled. TITO keeps the sampled token IDs on the server and never rebuilds them from text.
What goes wrong. Agents usually talk to the model through a message API. The model generates tokens, the harness receives parsed text or a tool call, the tool runs, and the harness sends the full message history again. Decoding, parsing, reserializing, and retokenizing is not an identity transform. Some concrete ways it drifts:
- the harness rewrites tool-call JSON with different key order or whitespace, or substitutes
{}for absent arguments; - the harness drops the reasoning text from the previous turn;
- the chat template renders a turn boundary differently from what the model emitted, such as a trailing newline after a stop token.
Any one of these changes the token IDs, and the damage shows up in two places:
- At the next serving turn. Suppose turn 1 used prompt IDs $P_1$ and sampled output IDs $O_1$. If turn 2 retokenizes the whole transcript, its reconstructed prompt $\widetilde P_2$ need not start with $P_1\Vert O_1$ (the two ID lists concatenated). Turn 2 is then conditioned on a history the engine never generated.
- At training. If the data pipeline renders the transcript yet again, the trainer scores a reconstructed sequence instead of the token lineage that rollout produced. Its log-probs describe a trajectory that never happened.
TITO fixes the turn boundary first and then carries the same token lineage into training. Asking the trainer to use the same tokenizer would not be enough: the text it would tokenize has already drifted.
How TITO works. TITO makes the session server, which sits between the agent and SGLang (the rollout engine), the single authority on tokens. The agent still sends ordinary messages with its full history every turn. Those messages only describe the conversation and help the server find what it can reuse. A two-turn tool call shows the flow:
- Turn 1. From the first messages $M_1$, the server renders the chat template once to get $P_1$ and sends those IDs to SGLang. SGLang samples $O_1$ and returns each output ID together with its log-prob.
- Commit. The server stores a token checkpoint $C_1=P_1\Vert O_1$, the full snapshot the next turn will extend. It also keeps a per-turn record of the request and response, which training later reads.
- Turn 2. The agent sends the full history $M_2$: everything from turn 1 plus a tool result. The server matches that history against its stored checkpoints, picks the deepest one that applies ($C_1$ here), and treats its token prefix as authoritative. It tokenizes only the newly appended messages (here, the tool result) plus the opener for the next assistant turn. The assistant text the agent replays helps locate $C_1$; it is never used to recreate $O_1$.
- Collect. After SGLang samples $O_2$, the server commits $C_2=P_2\Vert O_2$. Training reads the server-owned prompt IDs, exact output IDs, and rollout log-probs from the stored records instead of tokenizing the final transcript again.
Let $S_2$ be the ordered list of messages appended after the matched turn-1 history. The second prompt is
\[P_2=\operatorname{merge}\!\left(C_1,\Delta_\tau(S_2)\right),\]where $\Delta_\tau$ renders and tokenizes only the appended messages, and merge applies model-family-specific boundary repair such as a missing newline or an overlapping stop delimiter.
In plain terms: $P_2$ is the exact turn-1 token row with a freshly tokenized suffix glued on, and merge fixes the seam. Two real seams:
- Qwen3: the template writes
<|im_end|>followed by a newline after every message, but the model stops at<|im_end|>without sampling the newline.mergeinserts it. - GLM-4.7:
<|observation|>and<|user|>are stop tokens for the assistant and also openers for the next message.mergedrops that stop token from the end of the stored row, and the next message’s own opener marks the boundary, so it appears exactly once.
Implementation notes
The paper describes the checkpoint and the per-turn evidence together. The source keeps them in two structures:
- Token checkpoint: the full snapshot $C_t=P_t\Vert O_t$ that the next turn extends.
SessionRecord: the resolved request, including $P_t$ ininput_ids, and the raw SGLang response. The response carries output-ID/log-prob pairs (output_token_logprobs) and, when enabled, routed-expert data. Training samples are assembled from these records.
The boundary rules live in per-family subclasses of the TITO tokenizer (merge_tokens in miles/utils/chat_template_utils/tito_tokenizer.py); the base class simply concatenates.
Building the training sample. Training consumes the same token row that serving built. For the two-turn example, the assembled sample looks like this:
| Span | Comes from | Loss mask |
|---|---|---|
| Original prompt $P_1$ | agent’s first messages $M_1$, rendered by the template | context, not a target |
| $O_1$ | sampled by the policy | 1 |
| Tool result and next assistant opener | environment and template | 0 |
| $O_2$ | sampled by the policy | 1 |
To build it, Miles aligns each turn’s output IDs with the accumulated checkpoint and removes only registered boundary overlap, such as the GLM stop token above. Later user messages and other environment text get mask 0 just like the tool result. The trainer therefore gets one contiguous, token-exact sequence: the rollout log-probs and the loss mask index exactly the same tokens, and the trainer never treats environment text as if the policy had produced it.
Implementation notes
Assembly lives in miles/rollout/session/samples/merge.py. For each record, it reads input_ids from the request and output IDs from output_token_logprobs, then greedily matches the output IDs against the accumulated checkpoint. Unmatched trailing tokens are the ones the next turn’s template absorbed. Their count must not exceed the family’s registered allowance (one token for GLM-4.7, zero for the final turn), and assembly fails loudly if it does.
Ownership of shared completions is computed after the sample picker drops leaves, so a dropped leaf cannot own a span.
The agent cannot override tokens. Because the server owns the token row, the agent may not supply token IDs or the offsets that index into them. A request that tries is rejected with an error rather than silently corrected. The server also sets the capture options itself (log-probs, metadata, and, when the run enables replay, routing data), whatever the agent asked for.
Implementation notes
The current server distinguishes rejection from forced capture (miles/rollout/session/request_args.py):
- Rejected with HTTP 400: a non-null client value for
input_idsorlogprob_start_len, and, on the training session path,routed_experts_start_len. - Forced by the server:
logprobs=true,return_meta_info=true, andno_stop_trim=false. The routing and indexer capture flags are derived from the run configuration.
The paper summarizes both cases as “agent-supplied token fields are overridden”; in the source, only the capture settings are overridden.
A built-in diagnostic. Miles can still render the complete messages through the chat template the canonical way and compare the result with the server-owned row. The result is reported as tito_session_mismatch. These canonical IDs never replace the sequence used for serving or training. Reading the diagnostic:
- Assistant-text differences are expected. They are evidence that TITO kept the sampled IDs instead of the template’s preferred tokenization.
- Special-token or non-assistant-text differences point to a broken chat template or boundary merge.
Why this matters beyond tokenization. Three later mechanisms index per-token data and are only correct if each position holds the same token on the rollout side and the training side. R3 (Section 4.1) needs each token’s expert route. OPD (Section 8.1) needs the student’s per-token log-prob. True-on-policy mode (Section 4.3) requires the rollout engine and the trainer to compare the same tokens. All three depend on the trajectory TITO preserves.
Linear and branching sessions. A session’s stored history can grow in two ways, and the choice decides how many training sequences one session yields:
- Linear: each request may only extend the tail. The agent may retry its latest turn by rolling back one assistant checkpoint; anything that diverges earlier is rejected. A linear session produces exactly one training sequence.
- Branching: the history is an append-only tree, and every retained leaf can become a trajectory. This fits coding harnesses, such as Claude Code, that fork sub-agents or compact their context mid-task. Such a harness cannot know in advance how many trajectories a session will produce, so the linear rule cannot serve it. When several leaves share an earlier completion, that completion keeps loss mask 1 only in the earliest retained leaf, so it contributes one gradient instead of several.
Implementation notes
In branching mode, the server attaches each request to the deepest checkpoint whose message path is a prefix of the request; any unmatched remainder opens a new branch, and branches are never deleted. A branch whose last generation stopped at the length limit cannot be extended. The linear rollback depth is fixed at one assistant turn (MAX_ASSISTANT_ROLLBACK_STEPS = 1 in linear_trajectory.py).
Deciding when history matches. Before reusing a stored prefix, the server compares each replayed message with the message stored at that position. The message matcher sets how strict that comparison is:
strictis the safe default.loose_tool_callrelaxes only JSON serialization differences in tool-call arguments.role_content_onlyignores tool calls, which can incorrectly merge distinct execution paths.
The last setting is dangerous for a concrete reason. A tool-calling turn often has empty visible text, so two turns that called different tools look identical under role_content_only. The stored prefix wins, and the trainer keeps a history showing a call the agent never made. Reuse should never come at the cost of a weaker correctness check.
Implementation notes
The matcher is selected with --session-message-matcher (default strict); a trusted dotted import path to a custom matcher is also accepted. strict compares only the fields a chat template reads: role, content, reasoning content, and tool calls. loose_tool_call still requires call IDs, function names, and ordering to agree. When role_content_only collapses two histories, Miles does not reconcile the tool-call IDs, so the mismatch stays silent.
Current limits. TITO is still a model-specific contract. Each registered model family has CPU tests that check the template is append-only (rendering a longer history only adds tokens after the shorter history’s tokens, never changes them) plus checks against live GPU inference; the generic handler is best-effort. The session path also cannot carry image or video inputs yet.
Implementation notes
A model family is a set of checkpoints that share one chat template and one pair of reasoning and tool-call parsers. Miles never detects the family from a checkpoint; the user names it at launch (--tito-model). Registrations span the Qwen3, GLM, Nemotron, Kimi, MiniMax, DeepSeek, and Inkling lines, and other checkpoints fall back to the generic handler with a warning. The CPU test alone is not enough: a template can be append-only in isolation and still break once a real parser consumes model output. Vision-language models bypass the session server and drive SGLang through its lower-level token-in, token-out interface.
A throughput bonus. Once a session is valid, the same session key keeps all its turns on one SGLang engine, or on one data-parallel (DP) rank when DP attention is enabled. That engine already holds the conversation prefix in its KV cache (the stored attention keys and values of earlier tokens), so each turn prefills only the new suffix instead of the whole history.
4. R3, TIS, and True-on-Policy: Answers at Three Different Layers
TITO guarantees that the trainer scores the same tokens that rollout sampled. That is necessary but not sufficient: the trainer can still assign those tokens a different probability. Miles attacks the remaining gap at three layers:
- R3 makes the trainer send each token through the same MoE experts that rollout used.
- TIS and clip-or-pop accept that some numerical gap remains and limit how far it can push the gradient.
- True-on-policy mode makes sampled-token log-probs match exactly, but only for a narrow set of models.
4.1 R3 Pins Down Discrete MoE Routing
In an MoE model, a router in every MoE layer scores all experts for each token and keeps only the top $k$. That choice is discrete, so a tiny difference in the scores can change which experts win. R3 records the experts rollout chose and makes the trainer reuse them instead of choosing again.
A concrete case shows why this matters. Suppose rollout selects experts {2, 7} for a token, while a few bits of floating-point error cause the trainer to select {2, 8}. Expert 8 then receives a gradient for a token it never helped generate, while expert 7, the one that did, receives none.
This error compounds across layers, tokens, and steps. Each update produces the policy that samples the next batch, so the drift feeds back into training. The R3 paper reports that such routing discrepancies can destabilize RL in MoE models and even end in training collapse.
Formally, for MoE layer $j$, let $g_{j,t}$ denote the router scores for token $t$. Rollout computes an integer expert set
\[A^{\mathrm{roll}}_{j,t} =\operatorname{TopK}\!\left(g^{\mathrm{roll}}_{j,t},k\right).\]R3 serializes those IDs and substitutes them for the trainer’s newly computed top-k result:
\[A^{\mathrm{train}}_{j,t}\leftarrow A^{\mathrm{roll}}_{j,t}.\]In plain terms: rollout writes down “token $t$, layer $j$ → experts 2 and 7”, and the trainer’s forward pass uses that list instead of its own top-k.
R3 pins expert membership, not the entire router computation. The trainer still gathers its own router scores at the replayed indices, and the arithmetic inside each expert can still differ. R3 therefore does not guarantee equal gate weights or equal final log-probabilities, and TIS or true-on-policy mode may still be needed for the numerical difference that remains.
Implementation notes
- R3 is enabled with
--use-rollout-routing-replay. SGLang then returns the routed experts alongside the generated tokens, and the trainer replays them during its forward pass. - Because the TITO session server records routed experts together with token IDs and log-probs, replay covers whole multi-turn episodes as well as single completions.
- Because only the final top-k selection is replaced, the router stays differentiable and still receives gradients.
Replay is cheap to apply but expensive to carry. The raw payload is
B_R3 = (tokens - 1) × layers × top-k × sizeof(int32)
At 32K tokens, 60 layers, and top-k=8, that is roughly 60 MiB per trajectory, before surrounding serialization overhead. The payload stays in memory and travels with the trajectory, and it grows with sequence length, which is exactly what agentic RL makes long.
Whether to enable R3 is therefore a deliberate per-recipe decision:
- Dense models have no expert routing, so R3 does nothing for them.
- Fully asynchronous training adds other discrepancies, such as weight staleness, which can limit R3’s marginal value.
- Several shipped MoE recipes enable R3. The GLM-5.2 case study later in this post disables it and uses TIS to control the remaining mismatch.
4.2 TIS and Clip-or-Pop Bound the Remaining Policy-Ratio Mismatch
Even with the same tokens and the same experts, the two engines disagree slightly on probabilities. Miles does not try to remove that gap here. It measures the gap for every sampled token and limits how much an outlier token can sway the update.
The gap has two sources. Even with identical weights, SGLang and Megatron can produce different log-probs because their kernels, precision, and batching differ. Under same-version synchronous execution, that numerical difference is most of what the ratio below measures. Under fully async execution, it can also include temporal drift between stale rollout weights and the trainer’s current pre-update weights.
It helps to separate this systems ratio from the ordinary optimizer ratio. For sampled token $x_t$ with history $h_t$, there are three probabilities:
- $\mu$: the behavior probability that rollout recorded when it sampled the token;
- $\bar\pi$: the trainer’s pre-update policy, recomputed on the same tokens;
- $\pi_\theta$: the policy being optimized in this step.
In plain terms, $\rho_t$ is the familiar PPO/GRPO ratio, and $m_t$ is the systems correction on top of it. Before objective-specific clipping, $\rho_t m_t=\pi_\theta/\mu$: the ratio against the policy that actually generated the token.
Miles forms the policy loss first and then multiplies each token’s loss by one of two mismatch weights, with interval bounds $l$ and $u$:
\[w_t^{\mathrm{TIS}}=\operatorname{clip}(m_t,l,u), \qquad w_t^{\mathrm{pop}}=m_t\,\mathbf 1[l\le m_t\le u].\]- TIS caps an outlier’s leverage but keeps the token in the gradient.
- Clip-or-pop passes an in-range ratio unchanged and removes an out-of-range token from the gradient by giving it weight 0.
A small example with the default interval [0, 2]. Suppose rollout sampled a token with probability 0.20:
| Trainer’s probability $\bar\pi$ | $m_t$ | TIS weight | Clip-or-pop weight |
|---|---|---|---|
| 0.30 | 1.5 | 1.5 | 1.5 |
| 0.60 | 3.0 | 2.0 (capped) | 0 (dropped) |
| 0.02 | 0.1 | 0.1 | 0.1 |
Ratios are always positive, so with a lower bound of 0 only the upper tail is affected. Both methods trade bias for lower variance; neither makes the two engines identical.
Miles reports the unclipped ratio, the clipped fraction, and the mean |m_t-1|. Inspecting that distribution before deciding how to handle outliers is more informative than watching only the mean loss.
Implementation notes
--use-tisturns the correction on. The interval comes from--tis-clip-low(default 0) and--tis-clip(default 2.0).- The built-in correction is TIS. Clip-or-pop is selected by pointing
--custom-tis-function-pathaticepop_function(a custom correction, not a separate top-level mode). - The weight multiplies the per-token policy loss after the objective’s own clipping has been applied.
- Both functions log the same three metrics,
tis,tis_clipfrac, andtis_abs, so runs using either correction can be compared directly.
4.3 True-on-Policy Tries to Eliminate the Difference
TIS reweights the effect of the gap. True-on-policy mode goes after its cause: it makes rollout and training run the same computation, so that, in supported configurations, the two engines produce identical log-probs for every sampled token. The public launcher exposes this as a single --true-on-policy flag, which expands into a cross-engine kernel contract:
| Source of numerical drift | Alignment rule |
|---|---|
| Attention | Use the same FlashAttention-3 path for rollout and training |
| GEMM (general matrix multiplication) and batch shape | Use batch-invariant kernels |
| Fused operations | Disable or replace fused RoPE and bias-SwiGLU paths that do not match |
| Runtime nondeterminism | Pin deterministic cuBLAS, Transformer Engine, NCCL (NVIDIA’s GPU communication library), and SGLang settings |
| Tensor parallelism (TP) | Use TP-invariant row-parallel layers and reductions where required |
| Decode vs prefill scoring | Re-score completed sequences with a prefill pass |
Batch invariance deserves a word: the rollout engine batches requests very differently from the trainer, so the matrix-multiply result must not depend on how many requests share a batch.
With all six rules in place, the absolute log-prob difference that Miles reports between the two engines is exactly zero in supported configurations.
That is a strong result, but its scope is narrow, and the limits change how you should use it:
- Model coverage. The paper and current source register profiles only for the dense Qwen3 0.6B/4B family. For any other model, Miles refuses to start a true-on-policy run, so it never runs with only a partial guarantee.
- What is equal. The guarantee covers sampled-token log-probs, not the full vocabulary distribution.
- What it ignores. It does not address staleness from older weights; the asynchronous buffer in Section 5 can bound that separately (
--max-weight-staleness). - Cost. Determinism and batch invariance forgo some throughput optimizations. In the documented Qwen3-4B-Base run, the reward curve matches the baseline while rollout takes longer.
- Backend. The paper describes both Megatron and FSDP integration, but at the pinned source revision the bundled Megatron argument path is gated off pending follow-up work. FSDP is the runnable path.
True-on-policy is therefore best understood as a narrow, strict kernel contract and a diagnostic tool. Turn it on deliberately; production workloads should not enable it by default.
Implementation notes
- The launcher flag maps to the trainer’s
--true-on-policy-mode. With--train-backend megatron, the pinned revision raisesNotImplementedErrorand points the user to--train-backend fsdp. - The paper says the registered Qwen3 profile covers data, tensor, pipeline, and context parallelism on the training side. The tensor, pipeline, and context layouts belong to the gated Megatron path, though; at this revision the runnable FSDP path is data-parallel only.
5. Fully Async: Throughput from Decoupling, Correctness from Observability
In a synchronous loop, generation and training take turns, so one side of the cluster is always idle while the other works. Miles runs the two sides at the same time on separate GPUs. The price is that the trainer now learns from slightly older weights, so Miles measures that age, reports it on every step, and lets you set a cap on it.
Why turn-taking wastes GPUs. Suppose most trajectories in a batch finish in two minutes but one agentic episode runs for twenty. The rollout GPUs that finished early sit idle until the straggler ends, and the trainer cannot start until the batch is complete. When the trainer then runs the optimizer, the rollout engines wait in turn. Long-context and tool-using tasks spend much of each batch in these bubbles.
How Miles decouples the two sides. With train_async.py and --fully-async, rollout and training run on separate GPU pools connected by a bounded buffer. Miles refuses to start a fully async run in colocated mode, where the trainer and rollout engines would share GPUs. A background worker keeps a fixed budget of trajectories generating, and the trainer pulls finished work from the buffer whenever it needs a batch.
The default refill rule works at sample granularity:
- Each completed trajectory frees one slot (a “sample credit”).
- Once
n_samples_per_promptcredits have accumulated, the worker submits another complete prompt group. - The trainer always consumes complete groups, which group-relative objectives such as GRPO need to compute advantages.
As a result, the number of trajectories in flight stays near the budget even when trajectory lengths differ by an order of magnitude. Under group granularity, one slow trajectory would keep its whole group’s slots occupied until it ends.
Implementation notes
--rollout-submission-granularityselectssampleorgroup; fully async defaults tosample.- The in-flight budget is
rollout_batch_sizegroups by default, or floor(--async-max-concurrent-samples/n_samples_per_prompt) groups when that flag is set (the flag must be at leastn_samples_per_prompt). - The buffer holds at most
floor(--async-data-buffer-capacity-factor × rollout_batch_size)groups (factor 2.0 by default). When it is full, adding a group blocks until the trainer takes one out. --async-unused-samples-handlerdecides what happens to aborted and stale groups:drop(the default) discards them,retryreturns their prompts for regeneration. Groups with a missing reward or rejected by a user filter are always discarded.- In-flight handling during a weight update is
--pause-generation-mode(retractby default, orin_place). Fully async refusesabort, because generation is always in flight and every weight update would kill it. - The buffer exposes only
put,get, and a metrics call, so a run can swap in its own selector by import path (--custom-async-data-buffer-path).
What still waits. Fully async does not mean nothing ever waits. Installing new weights still pauses rollout, and evaluation pauses new generation when it shares the rollout engines. By default, requests that are in flight during a weight update are retracted and then resume under the new weights, so one long trajectory can contain tokens from two or more policy versions. What asynchronous execution buys is overlap of the major compute phases. Its cost shows up as staleness:
staleness = current published rollout/engine version - oldest token version in the group
Using the oldest token in the group is deliberately conservative: a group is never treated as fresher than its oldest token. The buffer checks failures, timeouts, and user-defined filters on put, because those verdicts are fixed once the group is generated. It checks age on get, because data keeps growing stale while it waits in the queue.
Three ways to measure age. A long trajectory can itself span several weight versions published to rollout, so a single number hides useful detail. Let $v_{\min}$ and $v_{\max}$ be the oldest and newest weight versions attached to generated-token spans in group $G$, and let $v_c$ be the current published rollout version when the trainer consumes it. Miles can distinguish
\[S_{\mathrm{old}}=v_c-v_{\min},\qquad S_{\mathrm{post}}=v_c-v_{\max},\qquad \Delta_{\mathrm{gen}}=v_{\max}-v_{\min}.\]- $S_{\mathrm{old}}$ is the conservative value used for filtering.
- $S_{\mathrm{post}}$ measures how much the group aged after its newest span was generated, for example while it waited in the queue.
- $\Delta_{\mathrm{gen}}$ reveals weight updates that landed during generation.
For example, take a trajectory whose first tokens were decoded under version 12 and whose last tokens were decoded under version 13, consumed when the current version is 14. Then $S_{\mathrm{old}}=2$, $S_{\mathrm{post}}=1$, and $\Delta_{\mathrm{gen}}=1$. In plain terms: the group is two versions behind at worst, one version of that lag came from waiting, and one update arrived mid-generation.
A token-weighted staleness metric describes the typical token instead of only the worst one. It is $v_c$ minus the token-weighted mean version. If the trajectory above had 300 tokens from version 12 and 100 from version 13, its mean version would be 12.25 and its token-weighted staleness 1.75.
Two easy-to-miss details. First, --max-weight-staleness is unset by default. A bounded queue provides backpressure (a full queue makes rollout wait) but no freshness guarantee. When a limit is configured, a group is rejected only if $S_{\mathrm{old}}$ is strictly greater than the limit at get() time. In the example above, --max-weight-staleness 1 would reject the group, while a limit of 2 would let it through.
Second, the unit is a published rollout weight version, which is not necessarily an optimizer step. The two normally coincide only when --update-weights-interval=1. Spans with no version provenance (no record of which weight version produced them) are excluded from lag statistics, and a group with no versioned tokens at all is never rejected as stale. That is why weight_version_sample_coverage matters too.
Reading the queue. The most useful diagnostic is often simply queue_size:
- Persistently near zero: the run is rollout-bound, and the trainer is waiting for data.
- Persistently at capacity: the run is trainer-bound, and queued data is growing stale.
- A spike in
stale_groups_filtered: generation compute is producing data that will never be trained on.
Together with the staleness measures above, these queue counters turn “async is faster” into an operating state you can check: you can see which stage is the bottleneck and how fresh the training data is.
Implementation notes
All buffer metrics are reported on every training step under the rollout/fully_async/ prefix.
| Metric | What it reports |
|---|---|
avg_staleness, max_staleness |
$S_{\mathrm{old}}$ over the groups this step consumed |
avg_post_generation_staleness, max_post_generation_staleness |
$S_{\mathrm{post}}$ over the groups this step consumed |
avg_generation_version_span, max_generation_version_span |
$\Delta_{\mathrm{gen}}$ over the groups this step consumed |
token_weighted_staleness |
Token-weighted lag over the groups this step consumed |
weight_version_sample_coverage |
Fraction of consumed samples that carry any version provenance |
buffer_avg_staleness, buffer_max_staleness |
$S_{\mathrm{old}}$ over the groups still waiting |
queue_size |
Groups waiting in the buffer |
aborted_groups_filtered, stale_groups_filtered |
Groups dropped on put because generation gave up, and on get for exceeding the limit |
The paper’s metrics table covers queue_size, the $S_{\mathrm{old}}$ metrics, and the two drop counters. The post-generation, span, token-weighted, and coverage metrics appear only in the source.
Asynchronous evaluation. Evaluation follows the same principle. Miles can reuse the rollout engines, use a dedicated GPU fleet, or hand work to an external backend. Each returned score is attributed to the step of the evaluated weights, and its latency is recorded separately, so a score that arrives several steps late still lands on the right step. A dedicated fleet also verifies every engine’s weight version, so one score cannot silently combine multiple checkpoints.
6. Low Precision Is Not a Dtype; It Is a Four-Stage Contract
Lower-precision number formats make matrix multiplication faster: a GPU’s tensor cores roughly double their peak rate each time the precision halves. The risk in RL is that the trainer and the rollout engine quantize the same weights under different rules. They then compute different policies from the same checkpoint, the disagreement grows layer by layer, and a change made for memory and speed turns into train–rollout mismatch in the gradient. Miles therefore treats precision as one contract that every stage touching the weights must obey.
A quick vocabulary note. A dtype (data type) is the numeric format a tensor is stored in. BF16 (bfloat16) is the standard 16-bit floating-point training format and serves as the baseline here. FP8 and FP4 are 8-bit and 4-bit floating-point formats, available in hardware on NVIDIA Hopper and Blackwell GPUs respectively. Block formats store one shared scale per block of values, which widens the range a narrow format can represent.
The four stages. The contract covers every place where weights are converted or used:
- Checkpoint conversion.
- Trainer forward pass.
- Live weight export, which converts the trainer’s weights into the rollout engine’s format at each weight update.
- SGLang rollout.
The end-to-end MXFP8 and NVFP4 recipes share a bit-exact quantizer, so the training and rollout kernels see identical quantized values. A few tensors stay in BF16 because their contraction axis does not line up with a one-dimensional scaling block. The paper names the final transformer layers, the shared experts, and the projections in multi-head latent attention. In those recipes, configured BF16 tensor and layer exceptions must agree across all four stages; an override applied at only one stage breaks the contract.
| Format | Block size / scale format | Maturity | Hardware scope in the paper | Models tested in the paper |
|---|---|---|---|---|
| BF16 | — | Baseline | A100 and later NVIDIA GPUs; supported AMD GPUs | All |
| FP8 blockwise | 128×128 / FP32 | GA (generally available) | Hopper, Blackwell, MI350X/MI355X | Qwen3-4B, Qwen3-30B-A3B, DeepSeek-V4 |
| MXFP8 | 1×32 / UE8M0 | Beta | Blackwell | Qwen3-30B-A3B, DeepSeek-V3.2 |
| NVFP4 (E2M1) | 1×16 / E4M3 + per-tensor FP32 | Beta | Blackwell | Qwen3-30B-A3B |
To read the block column: MXFP8 shares one UE8M0 scale (an 8-bit exponent, so a power of two) across 32 consecutive FP8 (E4M3) values, where E4M3 means 4 exponent bits and 3 mantissa bits. NVFP4 gives each block of 16 four-bit values (E2M1 = 2 exponent bits, 1 mantissa bit) its own E4M3 scale and nests those inside one FP32 scale per tensor.
Implementation notes
- MXFP8 keeps its format through rollout, the forward pass, and both the weight-gradient and data-gradient GEMMs, holding the configured exceptions in BF16.
- NVFP4 scales activations per token, so quantization artifacts do not depend on how a batch is composed. Miles quantizes the gate and up projections together, so the fused rollout GEMM uses one outer weight scale. Tensors the recipe does not cover stay in BF16, and rollout keeps a BF16 KV cache.
- NVFP4 has two optional refinements, selected through environment variables rather than flags. Dequantized backward runs the backward GEMMs in BF16 on operands dequantized from the forward pass’s NVFP4 values; it touches only training. Four-over-six tries both 4 and 6 as the FP4 magnitude for each block’s largest value and keeps whichever gives less error. Because it changes the quantized values, it is enabled in both the trainer’s Transformer Engine kernels and SGLang’s FlashInfer kernels.
Which combinations work. Valid combinations are generally either “trainer and rollout use the same format” or “BF16 training with quantized rollout.” NVFP4 is the exception: every stage that touches the weights must quantize them, so BF16 trainer + NVFP4 rollout is unsupported.
Beyond the table, the paper also mentions INT4 quantization-aware training and a “BF16 train, FP8 serve” mode that is easier to stand up when a new architecture first comes online. The paper’s GLM-5.2 case study uses that more practical combination: BF16 training with FP8 serving weights and KV cache.
How to validate a recipe. The reusable lesson is the validation sequence more than any particular configuration line:
- Share the quantizer.
- Align exception tensors.
- Verify weight updates tensor by tensor.
- Compare log-probs on the same sampled tokens.
- Monitor importance ratios online.
The paper validates its recipes with step 4, using SGLang and Megatron-LM as the two engines. On the configurations measured so far, reward curves track the BF16 baseline closely and rollout time drops significantly. The same quantization may behave differently on other models, and MXFP8 and NVFP4 remain Beta.
7. Memory and Weight Synchronization: Making a 1T Model Actually Run
At trillion-parameter scale, two engineering problems decide whether a run is practical at all. First, the training state is larger than GPU memory. Second, after every training step, the new weights have to reach a fleet of rollout GPUs quickly enough that generation is not left waiting. Section 7.1 covers memory; Section 7.2 covers weight transport.
7.1 Two Forms of Offload for Two Time Scales
A GPU’s high-bandwidth memory (HBM) has to hold the model’s weights, gradients, and optimizer state. When they do not fit, some of that state must live elsewhere. Miles moves state off the GPU at two time scales: between training steps, and within a single step.
- Between steps:
--offload-train. While the training actor (the trainer process that holds the model, gradients, and optimizer state) is paused, its weights, gradient buffers, and optimizer state leave the GPU together, freeing the memory for another process such as a rollout engine. Megatron can offload to CPU memory or local disk; FSDP currently supports host RAM only. - Within a step:
--stream-optimizer-state-to-disk. The forward and backward passes never read the optimizer state, so Miles keeps it on disk and loads and updates it bucket by bucket only when the optimizer step needs it.
The optimizer state is usually the largest item. FP32 master weights take 4 bytes per parameter and each of Adam’s two moment estimates takes another 4, so the total is roughly 12 bytes per parameter. For a 1T-parameter model, that is about 12 TB for the optimizer state alone. For very large models it may not fit in HBM even after sharding; the GLM-5.2 run streams roughly 279 GB per rank from local disk.
The two mechanisms also compose. On Qwen3-30B-A3B, the paper reports that enabling the streaming optimizer cuts actor offload time from 24 s to 5.2 s and reload time from 8.9 s to 1.3 s, because state that already lives on disk does not have to move again when the actor is evicted. The costs are slower checkpoint saves and a requirement to keep the same parallel layout when restoring.
Implementation notes
- Offload defaults follow placement. In a colocated run, where rollout engines share the trainer’s GPUs,
--offload-traindefaults to on; in a disaggregated run, where they have separate GPUs, it defaults to off. PPO with a critic also evicts the actor by default, because Miles always colocates actor and critic. - Megatron offloads at the allocator level, so weights, gradients, and optimizer state move as one block. Its disk path streams through a fixed-size pinned staging buffer, which keeps host memory bounded.
- A streamed run cannot resume from a checkpoint written without streaming; Miles refuses the resume rather than silently restarting Adam from step zero. Streamed state is copied into the checkpoint directory synchronously, which is why saves are slower.
7.2 Weight Transport: Prepare Once, Move According to Topology
After each training step, the trainer must hand one complete new policy version to every rollout engine before they generate with it. When trainer and rollout run on different GPUs, this handoff can dominate the step: broadcasting the 1T-parameter Kimi K2 takes almost a minute in the measurements below.
A weight update has two distinct jobs:
- Prepare: reassemble serving-ready tensors from the trainer’s sharded layout.
- Move: deliver them so that every rollout engine sees the same complete version.
Miles keeps those jobs separate. On Megatron, broadcast, P2P, and disk-delta all consume the same prepared buckets; disk-delta then follows its own storage-oriented publish lifecycle. FSDP has a separate, simpler broadcast updater. When trainer and rollout are colocated on the same GPUs, the handoff stays local through CUDA IPC (inter-process memory sharing on one GPU) and no network transport is involved.
A few terms recur below:
- Parallelism. TP splits each weight matrix across GPUs; PP (pipeline parallelism) splits the layers into stages; EP (expert parallelism) gives each GPU a different subset of an MoE layer’s experts; ETP (expert tensor parallelism) additionally splits individual experts.
- Fabrics. Broadcast runs over NCCL. RDMA (remote direct memory access) lets one machine write directly into another machine’s registered memory without involving the remote CPU.
- Formats. “HF” means the Hugging Face checkpoint names and layouts that SGLang expects.
7.2.1 The Shared Bucketed Pipeline
A large model has many thousands of tensors. Instead of calling the rollout engines once per tensor, Miles packs converted tensors into size-bounded buckets, 512 MiB by default, and hands each full bucket to the selected transport.
For ordinary parameters, each Megatron pipeline stage:
- all-gathers its TP shards into full tensors;
- converts Megatron names and layouts to the HF/SGLang representation;
- appends the result to the current bucket and flushes it when full.
Some tensors must travel together. A weight stays with its quantization scales, and model-specific parameter groups that SGLang must load atomically stay in the same bucket.
Routed experts take a second pass with their own ETP/EP gather, because under EP each rank holds different experts rather than a slice of one shared tensor. Miles fills the expert bucket before the EP gather, so the gather multiplies its size by the EP degree. With EP = 8, a bucket that looks like 512 MiB before the gather would become 4 GiB after it. Miles therefore multiplies the pre-gather byte count by the EP degree when deciding when to flush, which keeps the gathered bucket near its nominal size.
Pipeline stages own disjoint parameters and prepare them independently, which is what lets the pipeline scale to many stages. Two mechanisms still synchronize them: lifecycle barriers at phase boundaries, and a lock that broadcast takes during each bucket flush to keep collectives in order. A slow stage stalls the others at those points.
7.2.2 Broadcast: Simple, General, and Redundant at Scale
Broadcast is the default transport when trainer and rollout use separate GPUs. One trainer rank per pipeline stage sends every bucket to every rollout GPU.
For each bucket, Miles first sends the tensor names, dtypes, and shapes over a control path, then broadcasts the payload through an NCCL group that contains that stage’s sender and all rollout ranks, and waits for the collective to finish.
Broadcast needs no model-specific logic for re-sharding on the target side, and it works well when both sides already share an NCCL fabric. Its weakness is fan-out. A rollout GPU that holds only one-eighth of the experts still receives all of them. Adding rollout replicas adds traffic, but each pipeline stage still has only one sender.
Implementation notes
Each pipeline stage gets its own NCCL group, named miles-pp_{pp_rank}, and only the rank with DP = TP = 0 in that stage broadcasts. Metadata travels over HTTP. Non-expert and expert parameters use separate buckets. When there is more than one pipeline stage, a world-wide ticket lock serializes bucket flushes so that concurrent broadcasts cannot deadlock NCCL.
7.2.3 P2P RDMA: Send Only the Shard a Target Needs
P2P transfer replaces “one sender, everyone receives” with “many senders, each target receives only its own shard.” Several trainer ranks write directly into rollout-rank memory over RDMA at the same time. Gathering, conversion, and bucketing stay the same as for broadcast.
The topology-specific work happens once, at initialization:
- Transfer plan. A plan maps sources to targets. The first
min(sources, targets)assignments are one-to-one; the remaining targets are spread evenly across sources to limit the number of RDMA sessions per sender. - Target metadata. Each source queries its targets’ memory registrations and parallel-layout metadata.
- CPU replica. The source builds an SGLang model replica on CPU under the target rank’s layout. The replica never serves tokens. Its
weight_loaderis the executable definition of SGLang’s sharding and fusion rules, so the sender never has to reimplement them. - One staging buffer. A single reusable pinned (page-locked) host buffer is registered with the RDMA transfer engine. All target replicas reuse it, so staging memory stays constant instead of growing with the number of engines.
During an update, converted HF tensors pass through the CPU replica, which re-shards them exactly as the target expects. Fused parameters such as the Q/K/V projections wait until every constituent shard has arrived. The source then issues batched RDMA writes; the writes for the final target group run in a background pool and overlap preparation of the next bucket.
The scaling argument explains when P2P wins. With $M$ source ranks, trainer pipeline depth $p$, and rollout expert-parallel degree $e$, P2P exposes roughly $M/p$ times the aggregate sender bandwidth of broadcast, while each MoE target receives roughly $e$ times less data because it gets only its local experts. For example, 64 source ranks at $p=4$ give about 16 times the sender bandwidth, and a rollout GPU at $e=8$ receives about one-eighth of the expert bytes. Fleet width and expert sharding matter more than parameter count alone.
The pinned source revision documents a broader H100-80GB sweep than the three rows in the paper. All runs use 1 GiB buckets, and the reported times are averages over steady-state steps 3–12; a positive change means P2P is slower.
| Model | Nodes/side | Broadcast | P2P RDMA | Change |
|---|---|---|---|---|
| GLM-4.7-9B-Flash | 1 | 2.51 s | 4.23 s | +68.6% |
| GLM-5, 4-layer test | 1 | 0.73 s | 1.26 s | +72.2% |
| Qwen3-30B-A3B | 2 | 2.67 s | 2.16 s | −19.1% |
| GLM-4.5-Air | 4 | 5.00 s | 2.64 s | −47.3% |
| Qwen3-235B-A22B | 8 | 10.75 s | 3.16 s | −70.6% |
| GLM-5 744B-A40B | 16 | 58.30 s | 8.48 s | −85.5% |
| Kimi K2 1T | 32 | 53.28 s | 7.23 s | −86.4% |
On one node, P2P is slower: a single node offers no extra aggregate bandwidth, yet P2P still pays for CPU re-sharding and pinned-memory staging on every update. From two nodes per side it starts to win, and at 16–32 nodes per side it cuts an almost one-minute broadcast to 7–8 seconds. These are transfer-window measurements, not end-to-end training speedups, so P2P remains a benchmark-first optimization.
Implementation notes
- The paper’s table contains the 2-, 16-, and 32-node rows. Trainer and rollout fleets have equal size in every row.
- Timing starts when the generation-pause call returns and ends when the update call exits, so request-abort time is excluded.
- The Kimi K2 P2P time includes about 884 ms of on-GPU requantization that its checkpoint needs after every transfer.
P2P also covers fewer configurations. At commit 3439ec751, it needs a supported Megatron-to-SGLang weight mapping and direct rank-to-rank reachability. Startup rejects colocation, LoRA (low-rank adaptation, Section 8), Megatron Bridge (Megatron’s direct HF loader, Section 8), and prefill–decode (PD) disaggregation, where prompt prefill and token decoding run on separate servers. Rollout pipeline parallelism other than one is untested and also raises an error.
In this revision, some background RDMA writes that fail are only logged, not raised as errors, so an update can finish with tensors that never arrived. That makes the preflight weight-update check --check-weight-update-equal (Section 7.2.5) especially important before relying on P2P.
Implementation notes
- Weight mappings exist for the Qwen2 and Qwen3 dense families, Qwen3-MoE, the GLM4-MoE families, and DeepSeek-V3/V3.2 derivatives (which include GLM-5 and Kimi K2).
- The background writes for the last target group are collected by a wait step that catches each failed future and logs it instead of re-raising.
7.2.4 Disk-Delta: Publish a Version Instead of Sending a Model
Disk-delta serves a different boundary: trainer and rollout cannot share an NCCL or RDMA fabric, but every host can reach shared storage and keep a full checkpoint on local NVMe. Instead of sending the model, the trainer publishes only the bytes that changed since the previous version, and each rollout host patches its own local copy.
The diff is taken over canonical checkpoint bytes, not over tensor values such as $W_v-W_{v-1}$. With $B_v$ the checkpoint bytes at version $v$ and $i$ a byte position:
\[D_v^{\mathrm{xor}}=B_v\oplus B_{v-1}, \qquad D_v^{\mathrm{overwrite}}=\{(i,B_v[i])\mid B_v[i]\ne B_{v-1}[i]\}.\]In plain terms: if a byte changes from 0x3C to 0x3D, the XOR delta stores 0x01, and every unchanged byte becomes 0x00, which compresses to almost nothing. The overwrite delta instead lists only the changed positions together with their new values.
The lifecycle is explicitly versioned:
- Snapshot. The first update publishes no delta, so on this path SGLang starts from the checkpoint on disk rather than from a trainer push (the exception noted in Section 1). Trainer ranks read their byte baseline from the same canonical local HF checkpoint that every rollout host materializes as version zero, so the two baselines are byte-identical. The trainer also checks every tensor it exports against that checkpoint’s names, dtypes, shapes, and byte sizes, and any mismatch fails the run before training starts. Deltas begin with the second update.
- Diff and compress. Later updates gather tensors under canonical HF names. CPU workers compare each tensor byte for byte with the previous snapshot and compress the changed data with Zstandard; unchanged tensors are left out of the artifact.
- Publish. Each
weight_vNNNNNN/version records its required base, encoding, compression and checksum formats, and a tensor-to-shard map. Data shards carry the compressed changes and digests of the target state. - Pull and patch. Every rollout host patches its local full checkpoint and verifies the version lineage plus the checksum of the resulting tensor, not merely of the transferred delta.
- Reload and activate. Only after the patch succeeds does Miles pause generation, reload SGLang, advance the engine weight version, and resume.
Only step 5 pauses generation. Steps 1–4 are CPU and storage work and can overlap generation, so the final reload is the activation boundary.
Two safeguards keep a partial version from becoming visible. The index is written last, after the shards are flushed, fsynced, and atomically renamed, so its presence marks a complete version. And if two pipeline stages publish the same replicated tensor with conflicting checksums, publication fails before any engine reloads.
Implementation notes
- Disk-delta is selected with
--update-weight-transfer-mode=disk-delta. Required flags:--update-weight-disk-dirmust point at storage shared by trainer and rollout,--update-weight-local-checkpoint-dirat a rollout-host-local directory such as NVMe, and--hf-checkpointmust be a local directory because the baseline is seeded from its safetensors bytes. - Each rank that has changes writes them as one
model-XXXXX-of-YYYYY.safetensorsfile; rank 0 writesmodel.safetensors.index.jsonwith the version, base version, delta encoding, compression format, and checksum format.
The two encodings trade size against retry safety:
| Encoding | Stored data | Retry property | Consequence |
|---|---|---|---|
xor (default) |
new_bytes XOR old_bytes, then compressed |
Not idempotent | Must be applied exactly once to the declared base; applying twice restores the old bytes |
overwrite |
Changed positions and their new absolute values | Idempotent | Larger artifact, but safe to retry |
Both are byte operations that ignore dtype, so tensor name, dtype, shape, and byte layout must match the base exactly.
Disk-delta is currently Megatron-only and requires a local HF safetensors baseline, shared publication storage, and rollout-local checkpoint directories. Colocation, LoRA, PD disaggregation, and hot restart are unsupported. Its benefit depends on how many bytes actually change and how well they compress, and not every optimizer step is guaranteed to produce a small delta. Watch perf/update_weights_density (fraction of bytes changed) and perf/update_weights_wire_bytes (bytes written) when deploying it.
7.2.5 Choosing a Transport
No transport is best everywhere; the choice depends on whether trainer and rollout share GPUs, a fast GPU fabric, or only storage.
| Deployment | Recommended path | Why |
|---|---|---|
| Trainer and rollout share GPUs | CUDA IPC in colocated mode | No network transport is needed |
| Separate GPUs on one node or a modest shared NCCL fabric | Broadcast | Lowest setup and host-staging overhead; widest coverage |
| Wide multi-node fleets with direct RDMA and supported mappings | Benchmark P2P | Multiple senders and target-specific shards can outweigh staging overhead |
| No common GPU fabric, but shared storage and rollout-local NVMe | Disk-delta | Publish compressed changed bytes and prepare them while generation continues |
| FSDP, remote LoRA, or PD-disaggregated rollout | Broadcast in the current implementation | P2P and disk-delta are not general paths for these configurations |
Whichever transport you pick, it changes the cost of synchronization and leaves policy semantics untouched. A faster update can shrink the version gap between trainer and rollout, but correctness still requires activating one complete version and recording that version on every generated-token span. One long trajectory can legitimately span several publications.
To catch a transport that silently skips tensors, Miles ships an opt-in preflight check, --check-weight-update-equal. Before the first update, it fills the rollout tensors with random data; after synchronizing, it compares every expected tensor with the trainer’s copy. A write that never happened leaves noise behind, so it cannot pass just because an old value happened to match.
Implementation notes
The check runs once, at the start of training, and is enabled automatically only in Miles’s own continuous-integration (CI) tests. It requires bit-exact equality by default; a run may instead allow the rounding error of a quantized round trip, with the tolerance derived from the quantized format. Tensors that exist only on the rollout side, such as inference-only KV-cache scales, are skipped.
8. Two Trainers and Recipes Beyond RL
Miles treats the rollout engines, the trainer, and the weight-update paths as shared parts that other jobs can reuse. Two consequences follow. You can swap the training backend, the component that holds and partitions the model on the GPUs, without touching the rest of the loop. And you can recombine the same parts into post-training jobs other than full-parameter RL.
Miles ships two backends: NVIDIA Megatron-LM and PyTorch FSDP. Both can start from an HF checkpoint. The rollout, loss, and weight-update interfaces are the same for both, so the choice comes down to the model and the scale:
| Megatron-LM | FSDP | |
|---|---|---|
| Parallelism | TP × PP × CP × EP × ETP, plus DP | replicate × shard (HSDP) |
| Model input | Megatron distributed checkpoint, or direct HF loading through Bridge | Directly reads an HF directory |
| Disk offload | Supported | Unsupported |
| LoRA | Supported, with validated recipes | Currently unsupported |
| Best fit | Very large MoEs, complex parallelism, multi-node production workloads | Fast bring-up of new HF architectures; small- and medium-scale experiments |
Megatron-LM is the default; its five model-parallel axes (CP is context parallelism) are what large multi-node MoE runs need. FSDP gives up those axes for data-parallel sharding, and in return it loads an HF directory as-is, with no conversion step and no architecture flags.
Implementation notes
- One launch flag,
--train-backend, selectsmegatron(the default) orfsdp. - HSDP (hybrid sharded data parallel) replicates the model across groups and shards it within each group; that is the “replicate × shard” entry in the table.
- Megatron-LM does not need offline conversion either: Megatron Bridge reads an HF directory directly. A written Megatron checkpoint is parallelism-agnostic, so a run can change its parallel layout later without reconverting.
- “Unsupported” in the FSDP column means the current FSDP backend does not implement the feature yet; it is not a limit of FSDP itself.
The same components also compose into other post-training workloads, each reusing whichever parts it needs:
- SFT (supervised fine-tuning): has no online rollout engine. The data pipeline directly produces token IDs and an assistant-only loss mask while reusing the trainer, parallelism, and checkpoint machinery.
- LoRA RL: LoRA freezes the base model and learns a small adapter made of low-rank matrices; Miles trains and synchronizes only that adapter. Colocated deployments pass it through CUDA IPC; disaggregated deployments use NCCL. Experimental multi-LoRA support lets multiple adapters share one base model; it currently runs with disaggregated rollout only.
- Miles-Diffusion: extends the same generate → train → update structure to image and video diffusion, where a trajectory becomes a complete denoising path.
8.1 OPD: Dense Teacher Feedback on States the Student Actually Visits
Conventional distillation trains a student on text the teacher wrote. The student would rarely produce that text itself, so at inference time it drifts into states the teacher never demonstrated. On-policy distillation swaps the roles: the student writes the response, and a stronger teacher grades every token of it. The student learns on its own state distribution and receives a signal at every response token instead of a single end-of-sequence reward.
Concretely, the teacher only scores; it never generates. It reads the exact prefixes and token IDs the student produced and reports how likely it would have been to pick each of those tokens.
Let $h_t=(x,y_{<t})$, $p_t=\pi_{\mathrm{old}}(\cdot\mid h_t)$, and $q_t=\pi_T(\cdot\mid h_t)$. If the student scorer matches the behavior policy, so $y_t\sim p_t$, sampled-token OPD records
\[\hat d_t=\log p_t(y_t)-\log q_t(y_t), \qquad \mathbb E_{y_t\sim p_t}[\hat d_t]=D_{\mathrm{KL}}(p_t\Vert q_t).\]In plain terms: for the token the student actually chose, compare how surprised the old student and the teacher were. For example, if the student gave its sampled token probability 0.6 and the teacher gave it 0.2, then $\hat d_t=\log 0.6-\log 0.2\approx 1.10$. Swap the two numbers and $\hat d_t\approx -1.10$.
The sign says who favored the token more (positive: the student; negative: the teacher). A single $\hat d_t$ can therefore be negative, even though its expectation is a KL divergence and never negative. This is the reverse KL, because the student distribution is the first argument.
The expectation identity holds only if the tokens really came from $p_t$, that is, if the rollout distribution $\mu$ equals $p_t$. Numerical mismatch or asynchronous staleness can break that condition. TIS can bound the behavior-policy ratio, but it does not restore equality.
Miles does not add a separate distillation loss. It folds the teacher’s signal into the advantage. First it computes ordinary token advantages with GRPO, PPO, GSPO (Group Sequence Policy Optimization), or another estimator, then it applies
\[A_t^{\mathrm{train}}=A_t^{\mathrm{task}}-\beta_{\mathrm{OPD}}\hat d_t.\]A positive log-ratio lowers that token’s advantage; a negative one raises it. In the example above, with $\beta_{\mathrm{OPD}}=1$, the token the student overrated loses about 1.10 from its advantage, and the token the teacher preferred gains about 1.10.
The old-student and teacher scores are detached inputs, so gradients still flow only through the ordinary policy loss; there is no second, differentiable teacher loss. Task reward is optional: set it to zero for pure distillation, or keep it for reward-grounded OPD.
Implementation notes
--use-opdturns OPD on and requires--opd-type(sglangormegatron, see the next table).--opd-kl-coefis $\beta_{\mathrm{OPD}}$ and defaults to 1.0.- The shaping runs after the configured advantage estimator and before advantage normalization, so with
--normalize-advantagesthe OPD term is normalized together with the task advantage. - On the sampled-token path, the “old student” log-probs are the trainer’s recomputed log-probs by default, or the rollout engine’s under
--use-rollout-logprobs; top-K OPD computes its term during rollout. - OPD expects per-token advantages and raises an error on length or shape mismatches between advantages, student log-probs, and teacher log-probs.
Where the teacher runs decides when its scores arrive. An external SGLang teacher scores each sample while rollout samples are being processed, and the scores travel to the trainer with the trajectory. An in-process Megatron teacher instead runs an additional forward pass inside the training step.
| Teacher placement | Scoring time | Strength | Constraint |
|---|---|---|---|
| External SGLang teacher | During rollout processing | Can use a different or much larger architecture; supports top-K and metadata-routed teachers | Must share the student’s exact tokenizer/token-ID mapping; adds a scoring service |
| In-process Megatron teacher | Extra forward pass during training | No external teacher service | Compatible Megatron architecture/checkpoint; extra trainer memory; sampled-token path only |
The tokenizer constraint follows from the design: the teacher scores the student’s token IDs, so both models must map IDs to the same strings.
Implementation notes
--opd-type=sglangsends samples to the teacher at--rm-url.--opd-type=megatronloads the teacher checkpoint from--opd-teacher-load, which is required in that mode and rejected in SGLang mode.- Top-K scoring (
--opd-log-prob-top-k> 0) and multi-teacher routing (--opd-teacher-urls) are accepted only with--opd-type=sglang.
When $\mu=p_t$, sampled-token OPD is an unbiased but noisy one-sample reverse-KL estimate at each position: it looks at one token out of the whole vocabulary. The served-teacher path (the external SGLang teacher) can look at more. It chooses a candidate set $C_t$ (student top-K, teacher top-K, intersection, union, or symmetric difference) and computes
\[\widetilde d_t^{(K)}= \sum_{v\in C_t}w_t(v) \left[\log p_t(v)-\log q_t(v)\right].\]Weights can follow student probabilities, teacher probabilities, or be uniform. For example, with student top-5 and student-probability weights, each position compares the two models on the student’s five most likely tokens, weighted by how much the student favors each one.
This buys denser supervision with more scoring and transport. It still omits probability mass outside $C_t$, so it is a selected-support approximation, not exact full-vocabulary KL. At the pinned source revision, the default class-based rollout path supports only the teacher-top-K strategy; the others require the deprecated v1 rollout path (see the notes below).
Implementation notes
--opd-log-prob-top-ksets $K$; 0 (the default) means sampled-token OPD.--opd-top-k-strategypicks $C_t$:only-student,only-teacher,intersection,union, orxor(symmetric difference). The default isonly-student.--opd-reward-weight-modepicks $w_t$:student_p(default),teacher_p, ornone(uniform). The weights are renormalized over $C_t$, except underxor, which leaves them unnormalized (raw probabilities, or 1 per token undernone).- The class-based rollout path supports only
only-teacher. Strategies that need student top-logprobs, including the defaultonly-student, requireMILES_USE_LEGACY_ROLLOUT_V1=1; otherwise argument validation fails at startup.
Multi-teacher OPD is routing, not ensembling. A tag in each sample’s metadata selects exactly one teacher, so a math specialist can score math prompts while a coding specialist scores coding prompts. If a sample has no matching route and no default teacher is configured, Miles fails rather than silently choosing the wrong teacher.
Implementation notes
--opd-teacher-urls takes NAME=URL entries; --opd-teacher-key (default opd_teacher) names the metadata field that holds the teacher name. The reserved name default catches samples whose name is missing or unknown. Malformed or duplicate entries fail at startup.
The paper’s self-distillation result is useful but narrow. A Qwen3.5-35B-A3B teacher was produced by five RLVR (RL with verifiable rewards) steps; the student restarted from the base checkpoint and trained for five steps with zero task reward, so the teacher’s signal was the only training signal.
On held-out prompts from the DAPO math dataset, mean response length fell from 14,070 to 6,132 tokens, a 56.4% reduction, while accuracy moved from 84.0% to 85.2%. The roughly 1.6-point standard error exceeds the 1.2-point movement. The supported conclusion is therefore “substantially shorter responses with no reliable accuracy change,” not an accuracy improvement. It also comes from a single documented run.
Finally, OPD and true-on-policy (Section 4.3) use “on-policy” in different senses:
- OPD is algorithmically on-policy: the student generates the states it trains on.
- True-on-policy is a numerical systems contract: rollout and trainer must agree on sampled-token log-probabilities. OPD does not provide that equality.
9. GLM-5.2 744B: An End-to-End Case Study on 64 GB300s
The earlier sections test one mechanism at a time. A production run needs all of them at once: exact multi-turn tokens, asynchronous scheduling, low-precision rollout, mismatch correction, and weight updates for a very large model. The case study asks whether these pieces still work together at 744B scale. The paper answers with one run: a coding agent that learns to solve tasks at a terminal.
Fully async≤128 in flightTIS onR3 offBF16 trainFP8 rollout + KV cache
GLM-5.2 is an MoE model. “744B total, A40B active” means it stores 744 billion parameters but routes each token through only about 40 billion of them.
How the Run Is Set Up
The task. Each trajectory is one agent solving one terminal-bench-2 task at a command line. The task runs inside its own Daytona sandbox, built from the task’s official image and deleted when the episode ends; Miles reaches the sandboxes through its OpenEnv connector. An episode ends after 30 turns or one hour, whichever comes first. The task’s own test script grades the finished session.
The budget. Each model reply may use at most 8,192 tokens. The 65,536-token limit covers the agent’s whole multi-turn session, not a single reply. A training batch holds 64 trajectories: 8 tasks with 8 attempts each, so GRPO can compare every attempt with the other seven attempts at the same task.
The loop. The 64 GPUs sit in 16 GB300 nodes of four GPUs each. Eight nodes generate and eight nodes train. The two halves overlap: generation pauses only to install new weights or to run an evaluation.
- Generate. Eight rollout engines, one per generation node, serve an FP8 copy of the model with an FP8 KV cache. The session server keeps each agent’s exact tokens across turns (TITO, Section 3). Up to 128 trajectories are in flight at once.
- Buffer. Finished trajectory groups wait in the fully async buffer (Section 5) until the trainer takes a batch of 64. The 128-trajectory in-flight limit from step 1 is set separately from that batch size.
- Learn. The trainer keeps its weights in BF16 (Section 6), computes GRPO advantages, and applies TIS to the remaining train–rollout mismatch (Section 4.2). R3 is off.
- Refresh. The trainer broadcasts the new weights to the rollout engines (Section 7.2). Every ten steps, the run also pauses generation to evaluate on a disjoint held-out set of tasks, using the same engines.
How 32 GPUs hold a 744B trainer. Memory, not speed, decides the training layout. Tensor, pipeline, and context parallelism (TP2 × PP4 × CP4) use all 32 training GPUs, which leaves a single data-parallel replica (DP=1). EP8 does not add GPUs: it divides the experts among the same 32 ranks, so multiplying by 8 again would wrongly suggest 256 GPUs.
DP=1 has a cost. With no second replica to share it, each rank holds its full share of the optimizer state, unpartitioned: about 279 GB, more than a GB300’s memory. The run therefore streams the optimizer state to node-local disk (Section 7.1). Here streaming is what makes the run fit at all, not a speed-up.
The rollout side. Each generation node runs one four-GPU engine with DP-attention (DP4): the attention layers run data-parallel, so each of the four GPUs serves its own requests, while the experts are split across the four GPUs. Eight such engines give 32 inference DP ranks. Each engine also uses multi-token prediction (MTP): a small draft model proposes more than one token per forward pass.
Full run configuration
| Item | Configuration |
|---|---|
| Model | GLM-5.2, 744B total / A40B active parameters |
| Hardware | 64 × NVIDIA GB300; 32 for rollout / 32 for training |
| Environment | terminal-bench-2, OpenEnv, and one Daytona sandbox per episode |
| Training parallelism | TP2 / PP4 / CP4 / EP8 |
| Rollout | Eight DP-attention engines (DP4), with MTP enabled |
| Precision | BF16 training; FP8 rollout weights + FP8 KV cache |
| Length and batch | Up to 65,536 tokens; 64 trajectories = 8 prompts × 8 attempts |
| Episode limits | 30 turns or one hour; at most 8,192 tokens per reply |
| Scheduling | Fully async; up to 128 trajectories in flight; TIS on, R3 off |
| Optimizer state | Streamed to node-local disk (--stream-optimizer-state-to-disk) |
| Weight update | Broadcast (--update-weight-transfer-mode broadcast) |
Why the other degrees are what they are:
- TP2, not TP1. At TP1, a rank’s non-expert weights alone overflow its GPU while the checkpoint loads.
- CP4. Context parallelism splits each training sequence across four ranks, so each rank holds roughly a quarter of the activations.
- PP4 with an uneven split. The model’s 78 layers are divided 18 / 20 / 20 / 20. GLM-5.2’s sparse attention shares indices across layers, so every pipeline stage must start on a layer that computes its own indices.
The launch script that reproduces the run is examples/experimental/openenv/glm52_tbench2/. It plugs in the terminal agent and the reward through import-path flags (--custom-agent-function-path, --custom-rm-path), the extension mechanism described in Section 10.
What the Run Showed
One run100 stepsOne task distribution
The run answers three systems questions and gives one hint about learning.
- Does it fit? Yes. The parallel layout plus the memory mechanisms of Section 7 fit the 744B trainer on half of the fleet and leave the other half to generation.
- How fast is a step? About 263 seconds at the median (first 30 measured steps), after a 1,042-second warm-up at step 0.
- Do generation and training overlap? Yes. Sample-granularity refill (Section 5) starts a new trajectory as soon as one finishes, so about 90–100 requests generate at once. The count stays below the 128 limit because a trajectory waiting on a tool call keeps its slot without generating. Each later turn returns to the data-parallel rank that already holds its prefix (Section 3), which keeps the prefix-cache hit rate at 96%.
Numerical health. The divergence metric compares two log-probs for each sampled token: the one the rollout engine reported and the one the trainer recomputes. It averages 0.0369 over the 100 steps and ends near its starting value, so the gap does not grow as training proceeds. TIS corrects for the remaining gap in each update.
Reward. The raw reward is the score each task’s test script returns, before GRPO centers it within the group. Its nine-step moving average rises from 0.438 to 0.556 over the 100 steps, which the paper reports as an observation, not a measured improvement.
10. What the Source Reveals About Miles’s Engineering Philosophy
A paper explains what a system is meant to do; the code shows what it actually does. This section gives a reading map from common questions to the files that answer them, then names three habits that recur across the code base.
Instead of reading the repository directory by directory, start with the question you want to answer:
- Where do sync and async loops diverge?
train.py·train_async.py - Who owns exact multi-turn tokens?
rollout/session/ - Where are queue age and filtering measured?
fully_async_data_buffer.py
- How do TIS and clip-or-pop differ?
corrections.py - Where does OPD score and apply reverse KL?
on_policy_distillation.py·opd.py - What does true-on-policy constrain?
true_on_policy/
- Where are broadcast, P2P, and disk-delta implemented?
weight_update/ - Where are low-precision weights converted?
megatron_to_hf/processors/
Reading these files side by side, three engineering habits stand out:
Rollout, generation, agents, rewards, losses, corrections, data sources, and buffers are injectable. Application logic can change without forking the framework.
Invalid combinations—such as fully async with colocation—are rejected at startup instead of silently changing semantics.
A training curve, a determinism test, a reduced-scale proxy, and a smoke test are not equivalent forms of support.
Replaceable boundaries in practice. Each extension point is a flag that takes a Python import path, so your code loads only when a run asks for it and nothing inside Miles changes. The GLM-5.2 run in Section 9 works this way: its terminal agent and its reward are plain functions named on the command line. The rollout stack is also split into agent, generation, and rollout layers, so an environment can replace exactly the layer it needs and reuse the rest.
Implementation notes
Some of the import-path flags at the pinned revision, by boundary:
| Boundary | Flag |
|---|---|
| Rollout | --rollout-function-path |
| Generation | --custom-generate-function-path |
| Agent | --custom-agent-function-path |
| Reward | --custom-rm-path |
| Loss | --custom-loss-function-path |
| Importance-ratio correction | --custom-tis-function-path |
| Data source | --data-source-path |
| Fully async buffer | --custom-async-data-buffer-path |
Failing early in practice. A configuration that cannot keep its promise stops the run at startup. Fully async generation must keep running while the trainer trains, so --fully-async together with --colocate is rejected with an error. The same habit appears earlier in this post: true-on-policy refuses unregistered models (Section 4.3), and multi-teacher OPD raises an error as soon as a sample matches no teacher route and no default exists, instead of silently picking a teacher (Section 8.1).
Evidence levels in practice. Miles-Diffusion labels every recipe with one of four levels, and the label applies to the exact script and topology, not to the model family:
- Fully gated: a complete training curve, plus a deterministic end-to-end test of the recipe itself, run at least nightly, whose every registered metric matches a committed standard exactly.
- Proxy gated: the same, except that the nightly test runs a documented, scaled-down single-node proxy.
- Verified: a complete training curve, but no deterministic test that meets either standard above.
- Not verified: no complete training curve; smoke tests and short debugging runs do not count.
Section 9 applies the same discipline to the GLM-5.2 result: one run is reported as one run.
Implementation notes
The paper also describes how the readability rules are enforced. Formatting and import-order checks run as pre-commit hooks that CI reruns on every pull request. Further hooks ban Miles-specific anti-patterns and point to the API to use instead; for example, ban-mpu-get rejects direct mpu.get_* calls because the two training backends share one ParallelState object. CI also keeps each test’s training metrics across runs and checks every new value against that history, which catches slow drift in reward or divergence that no single run would reveal.
11. From a Correct Baseline to Scale: A Four-Stage Adoption Path
Miles has many switches, and each one changes either what the trainer learns from or how fast it runs. If you turn several on at once and the reward curve goes wrong, you cannot tell which switch caused it. The safer path has four stages. Each stage ends with a check you pass before moving on, so every new problem has one obvious suspect.
Skip any part of stage 2 that does not apply to your model.
- Start with a small model, BF16, and synchronous scheduling.
- Validate rewards, samples, per-token log-probs, and the first weight synchronization.
- Add TITO with the
strictmatcher. - A reasonable bar before moving on: rollout and trainer log-probs agree closely on the same tokens and the synchronized weights match the trainer's copy (
--check-weight-update-equal, Section 7.2.5).
- For an MoE, measure whether R3 is worth its routing payload.
- For low precision, validate conversion → trainer → export → rollout as one chain.
- A reasonable bar before moving on: the train–rollout log-prob mismatch (the ratio metrics of Section 4.2) is small and stays flat over training.
- Enable fully async only after the synchronous path is stable.
- Treat queue size, staleness, discard rate, and version coverage as launch criteria.
- A reasonable bar before moving on: those metrics hold steady and the queue is not pinned at capacity (Section 5).
- Prefer broadcast when rollout runs on separate GPUs within one node (colocated runs use CUDA IPC).
- Benchmark P2P on wide multi-node fleets.
- Use disk-delta for cross-cluster storage-mediated delivery.
Miles v0.1 does not claim that one transport, dtype, or optimizer is always best. Its most valuable habit is to document, for every choice, the topology it fits, the conditions under which it fails, what to observe, and how to validate it. In RL, a mistake may surface only hundreds of steps later. Measuring, bounding, and correcting how rollout and training differ is therefore worth more than any single throughput number.