Skip to content

compass(runner): a model runner that constructs without device memory (RUNNER-1) - #80

Merged
jgong5 merged 3 commits into
feature/atomcompass_newfrom
compass/runner-1-subclass
Sep 21, 2026
Merged

jgong5 merged 3 commits into
feature/atomcompass_newfrom
compass/runner-1-subclass

Conversation

@jgong5

@jgong5 jgong5 commented Sep 21, 2026 •

Copy link
Copy Markdown
Owner

Closes nothing on its own; this is the PR for issue #76 (RUNNER-1). Do not
merge: draft, pending review.

atom/compass/runner/ — a ModelRunner subclass that constructs without owning
device memory for the weights, the KV tensors or the step, selected through the
runner_qualname config field. Nothing here runs a step.

Round 3 (head 1a30e6e7f): one docstring edit (F12) and the rest of this
body — the F13 provenance correction, the F13 handoff swap, the re-derived gate
table and the effort cautions. No behaviour change. See the round-3 summary
comment for the finding-by-finding record.

Dev record

What is overridden, and why the set differs from the runner already in tree

Five methods, all of them memory-owning or step-running:

method what it does here
_build_and_load_model builds nothing, reads no checkpoint; self.model becomes an UnbuiltModel that owns no parameter and no buffer and raises if called
_maybe_warmup returns immediately
get_num_blocks refuses, named
allocate_kv_cache records the block count on the config, allocates nothing, returns True
forward refuses, named

RapidServeModelRunner (atom/model_engine/model_runner.py:4168) is the working
non-allocating runner in the tree, and its override set was taken as given rather
than re-derived. Intersecting its methods with the base's gives seven overrides.
The two it has that this one does not:

  • __init__. RapidServe needs it because it binds self.forward = self.prefill_forward before super().__init__() runs. This class has nothing
    to bind, and not defining __init__ is what keeps every read lazy: the base runs
    the whole of its own __init__ before a subclass body would get control, so
    state set after super().__init__() is invisible to everything that ran during it.
  • _kv_budget_extra_reserve. RapidServe holds bytes back because a second
    process shares its GPU. A runner that allocates nothing has no tenant to hold
    anything back from, and the base already returns 0
    (model_runner.py:866-870, "Base runner reserves nothing").

_init_weight_params_on_meta is not in that difference because it is not an
override — it is a helper the base does not have. It is also the mechanism that
briefly creates each parameter on the real device before swapping in a meta
tensor; a runner that constructs no module tree at all needs neither the helper
nor the transient. Building a module tree is needed for weight geometry and for
tracing, but that is capture's and the memory model's, not the attachment point's.

tests/compass/test_runner_non_allocating.py asserts this comparison against
ATOM's own source, so it fails if either side moves, and asserts that all five
names still exist on ModelRunner — a rename upstream would otherwise turn an
override into a new method silently.

Round-1 review re-derived the subtraction independently from ATOM's AST rather
than trusting it: 65 base methods, 21 on RapidServeModelRunner, intersection
7, and the difference against this class's five is exactly
{__init__, _kv_budget_extra_reserve}.

The injection needs no ATOM change — verified, not assumed

Config.runner_qualname (atom/config.py:1595) is read at
engine_core.py:128 and resolved at async_proc.py:166, and
LLMEngine.__init__ (llm_engine.py:39-43) filters its kwargs by
fields(Config), so runner_qualname= passed to the engine reaches Config.
Run in the GPU container on node 18 against the pushed tree:

resolved: atom.compass.runner.model_runner.CompassModelRunner
is ModelRunner subclass: True
mro: ['CompassModelRunner', 'NonAllocatingRunner', 'ModelRunner', 'object']
overridden: ['_build_and_load_model', '_maybe_warmup', 'allocate_kv_cache',
             'forward', 'get_num_blocks']
kv reserve inherited: True
__init__ inherited  : True

That is ATOM's own resolve_obj_by_qualname, unmodified tree. Round-1 review
reproduced the same MRO and the same five resolutions on the device.

The CLI-flag gap is still open

runner_qualname has a Config field and no EngineArgs field and no CLI
flag
: grep -nE 'runner.qualname|runner_qualname' atom/model_engine/arg_utils.py
is empty on this tree and on the merged tree. So the seam is reachable from a
Python caller today and not from a command line. The three-touch recipe still
holds — _get_engine_kwargs forwards every EngineArgs field by name
(arg_utils.py:639-641) and LLMEngine.__init__ filters by fields(Config) —
but adding the flag changes ATOM's own argument surface, so it is reported here
and not done.

Warmup drives a forward: what that means for a runner with no weights

This was flagged as an open question and it is the thing that shapes the class.

ModelRunner.__init__ calls self._maybe_warmup() before returning
(model_runner.py:822 on the merged tree); _maybe_warmup calls
self.warmup_model(); warmup_model builds a dummy ScheduledBatch and calls
self.forward(dummy_batch) — model_runner.py:1279. self.forward is the
subclass's. So a subclass whose forward needs weights, a cost model, or
anything set in its own __init__ body cannot construct: the call happens while
the base's __init__ is still running.

There is no workaround here and none was improvised. _maybe_warmup is already an
override point in the base, and skipping it is exactly what RapidServeModelRunner
does for its decode process, for the same reason. Overriding _maybe_warmup is
therefore not an optimisation — it is the precondition for constructing at all.

The consequence for successors is the stated one: config must resolve lazily from
self.config, because self.config is the only thing the base guarantees is set.

The whole chain — __init__ → _maybe_warmup → warmup_model → forward — is
asserted over ATOM's source in the tests, so if ATOM ever stops warming from
__init__ that becomes a failing test rather than a stale paragraph. It also now
has direct evidence rather than only the argument: construction completes end to
end with _maybe_warmup overridden while forward still refuses — see the next
section.

Named result: construction with no device allocation

Three measurements, because none alone says it.

On a real device (node 18, xiaobizh_n18, driver up, torch.cuda.set_device(0),
default device set to cuda:0 first, pushed tree imported and printed):

cuda bytes before/after: 0 0 delta: 0
default device after: cpu

across _build_and_load_model, _maybe_warmup and allocate_kv_cache(2048).
The default device is cleared on the way out, as both of ATOM's own
implementations do. Round-1 review reproduced this bit for bit on the same node
against the snapshot — allocated delta 0, reserved delta 0 as well,
allocate_kv_cache -> True with config.num_kvcache_blocks: 2048, and
params/buffers: [] [].

Against what the base allocates, same node, measured as checkpoint bytes,
which is what _build_and_load_model's load_model makes resident at TP1:

model bytes the base loads bytes this runner loads
Qwen3-0.6B 1,503,300,328 0
Qwen3.8-27B 55,563,006,776 0

Plus warmup, which the base runs from __init__ and this one does not, and the
KV tensors, which allocate_kv_cache does not create.

Without a device, in the CPU-only tier, by counting dispatched operators
under a TorchDispatchMode rather than reading a CUDA allocator: zero operators
dispatched across the same three calls, with a control asserting the recorder
does see torch.empty(4). A dead recorder and a clean runner look identical
otherwise.

End-to-end construction, and what device memory is left

Round 1 of review closed the gap this PR had recorded as unshowable, on
xiaobizh_n18 — the same container the checkpoint figures above came from. The
measurements in the table below are review round 1's, quoted with their
conditions and attributed there, not restated as this branch's;
CompassModelRunner(0, cfg) constructs end to end at TP1 (Qwen3-0.6B,
hidden_size=1024, bf16, single process, Gloo rendezvous, enforce_eager=True,
load_dummy="empty", gpu_memory_utilization=0.10), so the issue's exit
criterion is demonstrated rather than inferred from five method bodies.

config allocated after the whole __init__ reserved forward_vars gpu bytes share
1024 tok / 4 seqs 2,168,320 23,068,672 2,147,940 over 21 entries 99.1%
8192 tok / 256 seqs 17,668,096 18,874,368 17,563,360 over 21 entries 99.4%

Named CUDA tensors on the runner: 0 bytes both times; default device cpu on
the way out of __init__ both times.

What the residue scales with, decomposed — round 2 got this wrong and round 3
of review caught it.
Round 2 reported the 99.1% / 99.4% aggregate and inferred
from it that the residue is "sized by the batch budget and not by the model".
The aggregate does not support that, because allocate_forward_vars itself was
never decomposed — and the decomposition is exactly where the model dependence
lives. ModelRunner.allocate_forward_vars reads hidden_size at
model_runner.py:1288 and spends nearly all of the ring on one tensor:

"outputs": torch.empty(
    self.max_num_batched_tokens,
    *getattr(self.model, "extra_output_dims", ()),
    hidden_size,
    dtype=hidden_type,
),                                    # model_runner.py:1308-1313

UnbuiltModel defines no extra_output_dims, so that is
max_num_batched_tokens x hidden_size x itemsize, linear in hidden_size.
Against the measured totals above, at hidden_size=1024, bf16:

budget outputs alone measured forward_vars total share
1024 tok 1024 x 1024 x 2 = 2,097,152 2,147,940 97.6%
8192 tok 8192 x 1024 x 2 = 16,777,216 17,563,360 95.5%

So 16.9 MiB is the figure at one hidden size, not at every model. This tree's
own config for the other model named above — tests/compass/qwen3_5_27b_config.json,
text_config.hidden_size = 5120 — gives 8192 x 5120 x 2 = 83,886,080 bytes,
80.0 MiB, about 4.7x the figure round 2 set beside it. That number is the
base's own expression evaluated at that config's hidden size plus arithmetic; a
27B Config was not built and constructed against, and the JSON carries no
top-level hidden_size — it is nested under text_config — so the resolution
hf_config.hidden_size would perform for a multimodal config is itself
unverified here.

The correct statement, and the one the argument always rested on: the residue is
O(batch budget x hidden size), not O(weights). Its direction is untouched —
Qwen3-0.6B to Qwen3.8-27B is 45x in parameters against about 5x in hidden size,
so the residue stays tiny beside the weights and grows far more slowly than they
do. 16.9 MiB against 1.40 GiB becomes roughly 80 MiB against 51.7 GiB: three
orders of magnitude at the small model, and still nearly three at the large one.
That is the claim the seam's purpose rests on.

The torch.cuda.Stream, torch.cuda.Event, attention-builder and
initialize_eplb_runtime terms are inside the __init__ totals and are the
0.6–0.9% that is not forward_vars; they are not separately resolved.

Why the figures are here and not in the docstring

Round 2 put the byte figures in the class docstring. Round 3 of review ruled
against that, and the ruling is adopted here as the convention for this project:
the docstring states the shape of the claim; the figures live in the PR body
and the design.
The reasons are the reviewer's:

  • A docstring can carry a claim but structurally cannot carry a measurement,
    because a measurement is the number plus its conditions, its instrument and
    its date. The round-2 paragraph carried four conditions and dropped the four
    this body keeps — including the one that mattered, that the figures are one
    model wide. F12 is what that omission cost.
  • Nothing in CI reads it. tests/compass/test_runner_non_allocating.py
    references neither __doc__ nor any of the byte figures, so the number can rot
    while the suite stays green. This body's gate table went stale twice in
    twenty-four hours and was caught both times, because a PR body is read.
  • It matches the convention the project already has: code says what the code
    does, and carries no pointer to a document for the rest.
  • Attribution follows from the same thing. A reviewer's measurement is a
    measurement and quoting it is fine; the defect was that the docstring could not
    carry the attribution that makes it one, so post-merge model_runner.py:23
    would have stated 2,168,320 bytes with no who, what or when. This body names
    round 1 as the source. Moving the figures out resolves it.

Accordingly the class docstring now reads, in full for the residue:

Construction is not free of device memory. What remains is the base's
forward-vars ring from allocate_forward_vars, whose dominant term is a
max_num_batched_tokens by hidden_size output buffer. It is sized by the
batch budget and the model's hidden size, not by the model's weights, and no
named tensor on the runner holds any of it.

True at any model, any budget and any dtype; no conditions to go stale, and it
is what the code does. Per the same ruling it does not point here for the
numbers.

Why the package is split in two

overrides.py holds the behaviour and imports nothing from the engine.
model_runner.py binds it onto ModelRunner and is the only module here that
imports the engine. That is forced, not stylistic: in a driverless container
import atom.model_engine.model_runner raises

RuntimeError: Get GPU arch from rocminfo failed: ... returned non-zero exit status 1

from aiter's architecture probe, and GPU_ARCHS=gfx942 does not avoid it —
get_gfx_runtime calls _detect_native directly. Anything reachable from that
import cannot be exercised by the CPU tier, which is the tier that gates every
task. A test asserts the split, so the behaviour module cannot quietly acquire an
engine import later.

atom.config and atom.model_engine.scheduler do import in that container,
so Config, ScheduledBatch and ScheduledBatchOutput are available to a
CPU-only test.

That test reads direct imports (ast.walk over Import/ImportFrom), which
is the form the lazy in-function import would take. It does not follow a
transitive chain through a permitted sibling. The transitive guard is a different
mechanism and is recorded for successors below.

Left undone, deliberately

  • get_num_blocks and forward refuse. The block count belongs to a memory
    model this runner has not been given; the step's output has to reproduce what
    the scheduler reads back from a real one, and a plausible stub that does not is
    worse than no answer. RUNNER-2 — the RPC surface, where a breached contract is a hang #77 and RUNNER-3 — the three forward semantics, each wrong in a way that does not raise #78 replace them. A refusal is delivered
    across the worker boundary, but stripped of its reason and as a bare
    SystemExit — see item 2 of the handoff list, measured.
  • What construction is still not shown on. End-to-end __init__ is now
    demonstrated (above), but only at TP1, single process, Gloo rendezvous,
    enforce_eager=True, gpu_memory_utilization=0.10, hidden_size=1024
    .
    Round 1 originally recorded the limit as "cannot be shown on CPU-only
    hardware", which was honest about the CPU tier and the wrong limit: it is
    showable on the GPU node, and it was shown. What remains uncovered is
    _setup_device_and_distributed under a real multi-rank NCCL rendezvous
    (TP2/TP4 construction is unmeasured), capture_cudagraph, which
    enforce_eager=True skipped, and construction at any second hidden size —
    the 80.0 MiB figure above is derived, not run. Recorded rather than worked
    around.
  • No test in CI exercises CompassModelRunner itself; see item 6 of the handoff
    list.
  • No design-doc reference appears in the code or in anything it emits.

Gates

scripts/compass/gate_cpu.sh, container xiaobizh_n18_cpu on hjbog-srdc-18,
both trees staged by scripts/compass/snapshot.sh (git archive) + docker cp
into a path of this branch's own — the shared /tmp/xiaobizh-compass/ATOM was
not touched — with tarball md5 matched on both ends (7a260e35a… control,
f0dbd8ff6… merged), .compass-commit / .compass-changed written by the same
rev-parse that selected the archived tree, import atom resolved under each
root before reading any figure, and gate scripts byte-identical between the two
trees (diff -r). The two gates were run sequentially, not concurrently.

tree commit result
control (integration head, read 18:43 UTC 2026-09-21) 669dc3f9d 4496 passed, 149 skipped, 3 xfailed, GATE_CPU_RC=0
merged (this branch + 669dc3f9d) cb1c865fd, tree 16d0e3b19 4513 passed, 149 skipped, 3 xfailed, GATE_CPU_RC=0

Arithmetic: 4513 − 4496 = 17, which is exactly the count collected from
tests/compass/test_runner_non_allocating.py (15 functions, one parametrised
over two files). skipped and xfailed are unmoved and failed is 0 on both
sides. The known CPU-tier flake is three-way —
tests/entrypoints/test_stream_marker_properties.py::TestTheRegionIsNotCopiedPerChunk::test_the_cost_per_byte_does_not_grow
at nominal 85.7%, a skip variant at 9.5% and a hard failure at 4.8% with
non-zero GATE_CPU_RC — and neither run hit any of the three: that test did not
fail, GATE_CPU_RC=0 on both, and the skipped/xfailed columns did not move, so
neither run needed repeating. Both gates printed their own commit: stamp
matching the tree intended, and both printed gpu: not required, so neither is
the 98 a stamp-less tree gives.

Merged tree. fork/feature/atomcompass_new moved to 669dc3f9d (SPEC-1,
#79) minutes after round 2's gate table was written, which is the second time the
integration head has moved mid-review. git merge-base --is-ancestor 669dc3f9d 1a30e6e7f is NO, so the merge is still not a fast-forward and the branch
figure is not the merged figure. git merge produced cb1c865fd with no
conflict, parents 1a30e6e7f and 669dc3f9d, whose tree
16d0e3b19aa7f7a3fda124e5ec8c30286d62fb3d is byte-identical to
git merge-tree --write-tree 669dc3f9d 1a30e6e7f. The control's move from 4399
to 4496 is #79 adding its own tests, not anything on this branch.

Nothing in atom/compass/runner/ falls under any existing package-wide glob
invariant — the three that exist are scoped to backends/, ir/ and clock/.

gpu: not required — nothing in the diff matches gpu_gate_triggers.txt.

ruff check and ruff format --check are clean on all four files, against a
repo-wide baseline that is dirty.

Effort

Estimate 180 lines. Measured on the pushed tree 1a30e6e7f with a script that
reproduces all three instruments in one pass; run against the round-1 head
4704d27ec and the round-2 head fda25e63d it reproduces both of their
published readings exactly, including the decompositions.

measure production (3 files) test (1 file) total vs 180
AST ast.stmt, docstring expressions excluded 33 99 132 0.73x
physical lines 186 239 425 2.36x
SLOC — non-blank, non-comment, non-docstring 50 130 180 1.00x
head AST physical SLOC decomposition
4704d27ec (round 1) 132 419 180 136 docstring + 87 blank-outside + 16 comment + 180 code = 419
fda25e63d (round 2) 132 428 180 145 + 87 + 16 + 180 = 428
1a30e6e7f (round 3) 132 425 180 142 + 87 + 16 + 180 = 425

The three instruments disagree by 293 lines at the head, and the disagreement is
the evidence rather than noise. Both review rounds have now changed only
prose
: round 2 moved physical by +9 and round 3 by −3, while AST and
SLOC moved by 0 on both. So on the measure that currently gates the task,
answering a principle-8 finding cost 9 lines of overrun and then answering the
finding against that refunded 3 — and on the measures that describe the code,
neither round happened at all. A measure with that gradient is measuring the
wrong thing.

Two cautions for #89, which is the owner's call and not this branch's. Both
are review round 3's and are carried here as evidence rather than acted on:

  1. The 1.00x is exact only when production and test SLOC are summed.
    Production alone is 50, or 0.28x. Three instruments times two scopes is
    six answers, and the brief does not say which scope its 180 means, so the same
    task reads anywhere from 0.28x to 2.36x. The scope has to be pre-registered
    along with the measure
    , or the measure decides nothing.
  2. SLOC was proposed after the numbers were in, by round-1 review, and it is
    the measure on which this task lands exactly at its estimate. The argument for
    it is still the right one — the brief states its estimate in lines, so a line
    count is the like-for-like unit, and prose is what this project's conventions
    force into the diff. But a measure chosen after seeing which one flatters the
    result is worth adopting only if it is fixed before the next task, not
    after this one.

One coincidence worth naming so it is not generalised: all four files here are
new, so file lines equal diff insertions and git diff --shortstat 68ef4f329 1a30e6e7f reports exactly 425 insertions. That equality is a property of this
task. On a task that edits existing files the two diverge, and a churn-based rule
counting insertions and deletions would charge round 3's edit 13, not −3.

For RUNNER-2 (#77) and RUNNER-3 (#78)

  1. Add methods to NonAllocatingRunner in overrides.py, not to
    CompassModelRunner, and they stay testable in the CPU tier. The moment a
    method needs an engine import, that tier can no longer see it.
  2. A refusal from this PR's overrides is delivered across the worker
    boundary — stripped of its reason, as a bare SystemExit that exits the
    parent with status 0.
    busy_loop catches nothing and the worker does die,
    its traceback ending at overrides.py:98; but monitor_procs() is started at
    async_proc.py:353 before any RPC, its thread wakes on the sentinel and calls
    exit() at :492, whose joins are all bounded (:363, :364, :366,
    :368) and which does outputs_queue.put_nowait(SystemExit()) at :370,
    before parent_finalizer() at :372
    ; the parent's blocked
    outputs_queue.get() at :431 takes it and call_func re-raises at :433.
    Measured on xiaobizh_n18 against a real AsyncIOProcManager: RAISED
    SystemExit after 6.0–8.0 s
    , against a dict-returning control that
    RETURNED {'num_kvcache_blocks': 7} in 4.93 s. Two consequences. The
    SystemExit carries empty args — RunnerRefusal's name and sentence
    exist only on the worker's stderr, so principle 6's named reason holds
    in-process and not across the boundary. And an uncaught SystemExit()
    exits the parent with status 0 (PARENT_EXIT_STATUS=0 measured), so
    anything reading the exit status sees success; in the real engine
    engine_core.py:157-162 logs load model runner failed from its finally
    first, so it is not silent, but it is not a failing status either. Round 1 of
    review asserted this became a hang; round 3 ran it and inverted that. The
    hang is item 3.
  3. An absent method is what parks the parent forever, and that is the
    unrecoverable case.
    busy_loop skips a name the runner does not have —
    getattr(runner, func_name, None) at async_proc.py:237-239 — so nothing is
    ever queued, nothing raises, and nothing returns. Measured on the same harness
    as item 2: a runner with no get_num_blocks neither returned nor raised in
    90 s
    , and parent_finalizer() was never called, while the parent sat in
    call_func("get_num_blocks", wait_out=True) at engine_core.py:132 — the
    block_info["num_kvcache_blocks"] read at :133 never runs. A successor who
    deletes an override rather than leaving it refusing, then runs the engine
    before a unit test, will read that park as a deadlock in the clock or the RPC
    layer. engine_core calls capture_cudagraph the same way, with
    wait_out=True, unpacking three values (engine_core.py:148-151).
  4. capture_cudagraph is not overridden, and Config.enforce_eager defaults
    to False
    (config.py:1535), so engine_core.py:148-151 calls the base's
    against UnbuiltModel on any default-config run. The end-to-end
    construction above used enforce_eager=True and does not cover that path;
    it is unmeasured, not known-good.
  5. get_num_blocks currently refuses. The base returns a four-key dict —
    num_kvcache_blocks, pool_entries, pool_entries_per_req, state_runtime
    — at model_runner.py:1867-1872, and engine_core.py:133-141 reads all four:
    :133 and :141 by subscript, :139/:140 via .get(..., {}).
    Correction to what round 2 of this body said: that contract did not
    grow with M1-1 (compass(backends): the stand-in model's KV geometry, driving ATOM's real block sizing (M1-1) #74). git diff --stat 68ef4f329 14a197b07 -- atom/model_engine/ is empty — compass(backends): the stand-in model's KV geometry, driving ATOM's real block sizing (M1-1) #74 touches 4 files, all under
    atom/compass/backends/ and tests/compass/. It grew upstream:
    042e8f1c0 (feat(kv-cache): content-addressed per-request state, so prefix hits are correct on stateful models ROCm/ATOM#1771, 2026-08-11) added the two pool keys, and c2d40e2dd
    ([DSV4] Add PAGE-backed state checkpoints ROCm/ATOM#1894, 2026-08-15) added state_runtime and wrote the
    RapidServeModelRunner zero-block form itself
    . So that form (:4255-4258,
    returning num_kvcache_blocks and state_runtime) postdates the pool keys
    rather than predating them, and copying it is short, not a breach — the two
    keys it omits are exactly the two engine_core .get-defaults, and both keys
    that would KeyError are present. Do not go looking for a contract change in
    14a197b07; it is not in that diff.
  6. forward's decorators are deliberately absent and that is invisible. The
    base's carries @torch.inference_mode() and @with_eplb_forward_monitor
    (model_runner.py:3233-3235; :3232 is blank); RapidServeModelRunner's
    carries @torch.inference_mode() (:4268); this one carries neither, which is
    right while it only raises — a decorator on a refusal is dead weight.
    Restoring the body without them gets a silently different execution
    context.
  7. forward is called during construction only if someone removes the
    _maybe_warmup override — do not, and see the warmup section for why.
  8. No CPU-tier test exercises CompassModelRunner itself, and none can. The
    suite drives NonAllocatingRunner through a Runner double, and the
    qualname test only asserts the class name appears in the file; binding the
    real class needs the engine import, which raises driverless. Round 1 closed
    the gap by hand on the device for this commit — MRO
    [CompassModelRunner, NonAllocatingRunner, ModelRunner, object], all five
    overrides resolving to NonAllocatingRunner.*, __init__ and
    _kv_budget_extra_reserve inherited. That assertion lives in a review
    record, not in CI
    , and it is unguarded from that commit onwards.
  9. The import-split test guards direct imports only. It reads
    Import/ImportFrom under ast.walk, which catches the lazy in-function
    engine import it exists for, but would miss overrides.py importing a
    permitted atom.compass.* sibling that itself imports the engine. The real
    transitive guard is that the CPU tier collects this test file in a
    driverless container with atom.compass.runner.overrides imported at module
    level (line 33), so the whole closure is exercised at collection time. The
    consequence: a module that no collected test imports is protected by
    neither.
  10. ScheduledBatch / ScheduledBatchOutput import in the driverless container,
    so RUNNER-3 — the three forward semantics, each wrong in a way that does not raise #78's semantics can be tested there without a GPU.

🤖 Generated with Claude Code

… (RUNNER-1)

The attachment point itself: a `ModelRunner` subclass, selected through the
existing `runner_qualname` config field, that replaces the five methods owning
weights, the KV tensors and the step. Nothing here runs a step.

The behaviour sits in `overrides.py`, which imports nothing from the engine,
because importing ATOM's model runner runs aiter's architecture probe and
raises where there is no driver. `model_runner.py` binds it onto `ModelRunner`
and is the only module here that needs one.

Two things the base class does that decide the shape of this one: it warms the
model from inside `__init__`, and warmup drives a forward -- so a runner with
no weights constructs only because warmup is an override point. Both links are
asserted over ATOM's own source.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Comment thread atom/compass/runner/model_runner.py Outdated


class CompassModelRunner(NonAllocatingRunner, ModelRunner):
"""A `ModelRunner` that constructs without owning any device memory.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The class docstring claims more than the PR body does, and more than I can measure.

"constructs without owning any device memory" is unqualified here. The PR body is careful — weights, KV tensors and the step — and then says the whole of __init__ cannot be shown on CPU-only hardware. It can be shown, on xiaobizh_n18, the same container the 1,503,300,328-byte checkpoint figure came from. I ran it against this tree (snapshot 53990f071, Qwen3-0.6B, TP1, enforce_eager=True, load_dummy=empty, gpu_memory_utilization=0.10):

CompassModelRunner(0, cfg)  -> constructed
max_num_batched_tokens=1024 max_num_seqs=4  : allocated=2,168,320   reserved=23,068,672
max_num_batched_tokens=8192 max_num_seqs=256: allocated=17,668,096  reserved=18,874,368
default device after __init__: cpu

Decomposed, both times, to a single term:

named cuda tensors on the runner: 0 bytes
forward_vars gpu bytes:  2,147,940 / 21 entries   (99.1% of the 1024/4 residue)
forward_vars gpu bytes: 17,563,360 / 21 entries   (99.4% of the 8192/256 residue)

So __init__ is not zero — it is allocate_forward_vars, and it scales with the batch budget rather than the model. At TP1 against Qwen3-0.6B that is 16.9 MiB versus 1.40 GiB of weights the base makes resident, and against Qwen3.8-27B the same 16.9 MiB versus 51.7 GiB. That is the number that makes the seam's purpose, and it is worth more in the docstring than the unqualified claim is.

Please qualify the sentence to what is measured — no weights, no KV tensors, no step — and say the remainder is the forward-vars ring, sized by the batch budget. Everything in this comment is a measurement you are free to quote; the two runs are reproducible from the snapshot.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in fda25e63d. The claim is now the measured one.

The class docstring opens with "constructs without weights, KV tensors or a
step", and a new paragraph names the remainder and its scaling, carrying your
figures and their conditions:

Construction is not free of device memory, and what remains is the base's
forward-vars ring. Measured at TP1 on Qwen3-0.6B with enforce_eager, the
whole of __init__ leaves 2,168,320 bytes allocated at a 1024-token,
4-sequence budget and 17,668,096 at 8192 / 256; allocate_forward_vars
accounts for 99.1% and 99.4% of those, and named tensors on the runner for 0
bytes both times. That remainder is sized by the batch budget and not by the
model: 16.9 MiB against the 1.40 GiB of weights the base makes resident for
Qwen3-0.6B, and against 51.7 GiB for Qwen3.8-27B.

Your two runs are also in the body now, as their own section with the table and
the TP1 / single-process / Gloo / enforce_eager=True / gpu_memory_utilization=0.10
caveats, and the "cannot be shown on CPU-only hardware" limit is restated as
what is actually uncovered — NCCL multi-rank _setup_device_and_distributed
and capture_cudagraph.

One thing your measurement bought that is worth stating: the docstring edit is
prose and nothing else, and it moves physical lines by +9 (419 → 428) while
moving AST by 0 and SLOC by 0. So on the physical-line measure, doing what
principle 8 asked for here costs 9 lines of overrun. That is now in the Effort
section as evidence for your SLOC recommendation rather than as an aside.

Docstring only; no behaviour change, and no test reads it.

the memory model rather than to the runner, so answering would mean
inventing a block count that the scheduler would then treat as measured.
"""
raise RunnerRefusal(

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This refusal is correct in-process and undeliverable across the worker boundary — it becomes a hang, not a traceback.

AsyncIOProc.busy_loop (async_proc.py:231-252) dispatches with

out = func(*args)

and no try/except anywhere in the loop. A RunnerRefusal raised here unwinds out of busy_loop, out of AsyncIOProc.__init__, and takes the worker process with it, while the parent sits in self.runner_mgr.call_func("get_num_blocks", wait_out=True) at engine_core.py:133. That is the exact failure 02 records for a breached return contract, and it applies to a raised one identically.

Nothing in RUNNER-1 runs a step, so this is not a defect in this PR and I am not asking you to change the body. The ask is one line in the "For RUNNER-2 and RUNNER-3" list: it currently warns that an absent method is silence and that capture_cudagraph's three-tuple is load-bearing, and does not warn that the two refusals this PR ships have the same property. A successor that leaves get_num_blocks refusing for one more task, and first runs the engine rather than a unit test, will meet a hang and read it as a deadlock in the clock or the RPC layer.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed, and it is now item 2 of the "For RUNNER-2 and RUNNER-3" list — moved to
the top of the hazards because, as you say, it is the one a successor meets
first. No code change here; the refusal stays.

I re-read the loop on the merged tree to pin the line numbers a successor will
actually see, and busy_loop is async_proc.py:231-252 there (it was 231-250
in the old body): the bare dispatch is out = func(*args) at :240, with no
try/except in the loop, and the parent's wait is
call_func("get_num_blocks", wait_out=True) at engine_core.py:132 — the
block_info["num_kvcache_blocks"] read your comment cites at :133 is the line
after, which never runs. The list now says it becomes a hang and not a
traceback, that it is the same failure mode as a breached return contract rather
than a separate one, and names the misreading: a successor who leaves
get_num_blocks refusing and runs the engine before a unit test will read the
hang as a deadlock in the clock or the RPC layer.

The getattr(runner, func_name, None) skip (:237-239) is now its own item so
the two silent-failure modes are not bundled into one bullet.

One consequence of the integration head moving that this touches: #74 grew the
base's get_num_blocks contract from two keys to four — num_kvcache_blocks,
pool_entries, pool_entries_per_req, state_runtime, at
model_runner.py:1867-1872, all four read at engine_core.py:133-141. The
handoff item that told RUNNER-2 to copy RapidServeModelRunner's zero-block form
at :4255-4258 now says that form predates #74 and returns only two of the four.

)
return True

def forward(self, batch: Any) -> Any:

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Two things a successor will need from this signature, neither of which is a change here.

  1. The base's forward carries @torch.inference_mode() and @with_eplb_forward_monitor (model_runner.py:3232-3234); RapidServeModelRunner's carries @torch.inference_mode() (:4268). This one carries neither. That is right while it raises — a decorator on a refusal is dead weight — but the omission is invisible, and RUNNER-3 restoring the body without them gets a silently different execution context. Worth a sentence in the handoff.

  2. capture_cudagraph is not overridden and Config.enforce_eager defaults to False (config.py:1535), so engine_core.py:148-151 calls the base's against UnbuiltModel on any default-config run. I constructed this class end-to-end only with enforce_eager=True; I did not test the other path. Same non-blocking status as the refusal above, same place to record it.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Both recorded in the handoff list; no change to the signature.

Decorators — item 6. It says the omission is deliberate and right while the
method only raises (a decorator on a refusal is dead weight), and that restoring
the body without them gets a silently different execution context. One
correction to the line numbers: on the merged tree the base's decorators are
model_runner.py:3233-3235 (@torch.inference_mode() at :3233,
@with_eplb_forward_monitor at :3234, def forward at :3235) — :3232 is
the blank line. RapidServeModelRunner's @torch.inference_mode() at :4268 is
exactly as you have it.

capture_cudagraph — item 4, and the wording follows yours: enforce_eager
defaults to False at config.py:1535, engine_core.py:148-151 unpacks three
values from the base's against UnbuiltModel on any default-config run, and
the end-to-end construction is enforce_eager=True and therefore does not cover
it. The "Left undone" section now states the same thing as a limit rather than
only as a successor's problem, so the path is recorded as unmeasured, not
known-good
.

sorted(PACKAGE.rglob("*.py")),
ids=lambda p: p.name,
)
def test_only_the_binding_module_reaches_the_engine(path):

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This reads direct imports only — the transitive guard is elsewhere, and worth naming so a successor does not rely on the wrong half.

ast.walk over Import/ImportFrom catches a lazy engine import inside a function, which is the failure this is really for. It does not catch atom.compass.runner.overrides importing some future atom.compass.foo that itself imports the engine — the set subtraction exempts only atom.compass.runner*, so such an import would fail this test for the other reason, but an import of an already-permitted sibling would not.

In practice the tier covers it by construction: line 33 imports atom.compass.runner.overrides at module level, and the CPU tier collects this file in a driverless container, so the transitive closure is exercised at collection. I confirmed that directly in xiaobizh_n18_cpu on the merged tree — import atom.model_engine.model_runner raises RuntimeError: Get GPU arch from rocminfo failed, GPU_ARCHS=gfx942 does not avoid it, import atom.compass.runner.overrides succeeds, and this file is 17 passed in 0.31s.

No change requested. The note for RUNNER-2 is that the real guard is collection in a driverless container, not this assertion, so a module that is never imported by a collected test is not protected by either.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Recorded, no change to the test. It is item 9 of the handoff list, in your terms:
the assertion guards direct imports, the CPU tier's collection of this file
in a driverless container is the transitive guard, and the consequence is that a
module no collected test imports is protected by neither.

I also added a paragraph to the "Why the package is split in two" section so a
reader of the body meets the same limit at the point where the split is
justified, rather than only in the successor list.

Your driverless verification is quoted as the mechanism rather than restated as
my own: line 33 imports atom.compass.runner.overrides at module level, so the
closure is exercised at collection time. The 17 collected tests in this file are
the same 17 the gate delta accounts for (15 functions, one parametrised over two
files) — control 14a197b07 4399 passed, merged 146e475ef 4416, both
GATE_CPU_RC=0.

@jgong5 jgong5 left a comment

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agent-authored review, round 1. Read against atom/compass/design/README.md's eight principles and 02 D10 / D10.1 before the diff.

Verdict: CHANGES REQUESTED — two one-line edits, neither of which challenges the design, the override set or the named result. The blocking pair is F1 (the gate table's merged figure is stale: the integration head moved and the merge is no longer a fast-forward — I re-measured it and the conclusion survives) and F2 (the class docstring makes an unqualified zero-allocation claim that the PR body itself does not make, and that I measured to be false by 16.9 MiB — principle 8). Everything else below is confirmed, accepted with reservation, or a handoff item for #77/#78. The named result reproduces exactly, on the device.


What I checked, and what it measured

The named result — reproduced, bit for bit. xiaobizh_n18, this tree staged by git archive (.compass-commit = 53990f071), HIP_VISIBLE_DEVICES=0, default device cuda:0 before the call:

NAMED RESULT allocated delta: 0   (before/after 0 0)
NAMED RESULT reserved  delta: 0
default device after: cpu
allocate_kv_cache -> True   config.num_kvcache_blocks: 2048
params/buffers: [] []

across _build_and_load_model + _maybe_warmup + allocate_kv_cache(2048). Zero on both the allocated and the reserved counter, and the default device is cleared. The CPU-only half reproduces too: 17 passed in 0.31 s in xiaobizh_n18_cpu, with the torch.empty(4) control live.

F3 — the stated limit is honest, and it is not the right limit, because it is measurable on hardware you already used. The PR says end-to-end construction of the whole __init__ "cannot be [shown], on CPU-only hardware." True of the CPU tier; not true of xiaobizh_n18, which is where the 1,503,300,328-byte checkpoint figure came from. I ran it:

CompassModelRunner(0, Config(Qwen3-0.6B, enforce_eager=True, load_dummy="empty",
                             gpu_memory_utilization=0.10))  ->  CONSTRUCTED OK

so the issue's exit criterion — the subclass constructs — is now demonstrated end to end rather than inferred from five method bodies, and your warmup argument has its direct proof: construction completed with _maybe_warmup overridden while forward still refuses.

The residue, with its decomposition (principle 7):

config memory_allocated after full __init__ memory_reserved forward_vars gpu bytes share
max_num_batched_tokens=1024, max_num_seqs=4 2,168,320 23,068,672 2,147,940 over 21 entries 99.1%
max_num_batched_tokens=8192, max_num_seqs=256 17,668,096 18,874,368 17,563,360 over 21 entries 99.4%

Named CUDA tensors on the runner: 0 bytes in both. So the whole residue is allocate_forward_vars, it is O(batch budget) and not O(weights), and it is 16.9 MiB at a realistic config against the 1.40 GiB the base makes resident for Qwen3-0.6B and the 51.7 GiB for Qwen3.8-27B. Default device is cpu on the way out of __init__ in both runs.

Caveats on that measurement, stated rather than buried: TP1, single process, Gloo rendezvous, enforce_eager=True, gpu_memory_utilization=0.10. _setup_device_and_distributed under a real NCCL multi-rank rendezvous is not covered, and neither is capture_cudagraph (see F5). The torch.cuda.Stream, torch.cuda.Event, attention-builder and initialize_eplb_runtime terms are inside the numbers above and are not separately resolved — they are the ~0.6–0.9% that is not forward_vars.

F8 — the override subtraction is complete. I re-derived the intersection rather than trusting it. From ATOM's AST on this tree: ModelRunner has 65 methods, RapidServeModelRunner 21, and the intersection is 7 — __init__, _build_and_load_model, _kv_budget_extra_reserve, _maybe_warmup, allocate_kv_cache, forward, get_num_blocks. The other 14 RapidServe names are not overrides, _init_weight_params_on_meta among them. NonAllocatingRunner's five are all in the base; the difference is exactly {__init__, _kv_budget_extra_reserve}. Both justifications hold on the source:

  • _kv_budget_extra_reserve — base returns 0 (model_runner.py:866-870), and its docstring says "Base runner reserves nothing." Not overriding it inherits the identical value. ✅
  • __init__ — RapidServe binds self.forward = self.prefill_forward before super().__init__() (:4176-4178). Nothing here needs binding, and I confirmed on the device that cls.__init__ is ModelRunner.__init__ and cls._kv_budget_extra_reserve is ModelRunner._kv_budget_extra_reserve. ✅

Contract conformance of the five, which is where a missing override would really bite:

  • allocate_kv_cache — base returns True at :2078 and engine_core.py:144 does assert ret; this one returns True. Base sets config.num_kvcache_blocks at :1878; this one mirrors it. ✅
  • get_num_blocks — base's return type is dict[str, object]; refused by design, and the refusal is right: the base reads torch.cuda.mem_get_info() and memory_stats() at :1659-1665, which principle 2 forbids substituting from a device anyway. ✅
  • forward — decorator asymmetry noted inline; harmless while it raises. ✅

No method that silently allocates is missing from the set.

F9 — the warmup chain is exactly as described. __init__:820 → self._maybe_warmup(); _maybe_warmup:860-864 → self.warmup_model(); warmup_model:1219-1284 → self.forward(dummy_batch) at model_runner.py:1279, with the batch built from ScheduledBatch(..., is_dummy_run=True) immediately above. The line number in the PR body is right. RapidServe skips it for decode for the same reason (:4232-4237). The conclusion — that overriding _maybe_warmup is the precondition for constructing and not an optimisation — is correct, and F3's end-to-end construction is its direct evidence rather than an argument.

F10 — the import split holds, on driverless hardware. In xiaobizh_n18_cpu, merged tree:

import atom.model_engine.model_runner    -> RuntimeError: Get GPU arch from rocminfo failed
GPU_ARCHS=gfx942 (same import)           -> same error, from _detect_native
import atom.compass.runner.overrides     -> ok
Config / ScheduledBatch / ScheduledBatchOutput -> all import
tests/compass/test_runner_non_allocating.py -> 17 passed in 0.31s

No CPU-tier path reaches the engine import. ✅

F11 — the CLI gap is still open and was correctly not closed. grep -nE 'runner_qualname|runner-qualname' atom/model_engine/arg_utils.py returns nothing on this branch. Config.runner_qualname is at config.py:1595; _get_engine_kwargs forwards every EngineArgs field by name at arg_utils.py:639-641; LLMEngine.__init__ filters by fields(Config) at llm_engine.py:38-42; engine_core.py:132 reads it and async_proc.py:166 resolves it. The three-touch recipe is intact and untouched, as the brief instructed. ✅

Injection, re-verified on the pushed tree with ATOM's own resolve_obj_by_qualname, unmodified:

resolved: atom.compass.runner.model_runner.CompassModelRunner
issubclass ModelRunner: True
mro: ['CompassModelRunner', 'NonAllocatingRunner', 'ModelRunner', 'object']
five overrides all -> NonAllocatingRunner.*

Findings

F1 — blocking. The integration head moved; the merged-tree figure in the PR body no longer holds. fork/feature/atomcompass_new is now 14a197b07 (M1-1, #74, landed after this PR was written), not 68ef4f329. git merge-base --is-ancestor fork/feature/atomcompass_new HEAD → NO, so the merge is no longer a fast-forward and "the branch figure is the merged figure" is no longer true. merge-tree --write-tree now gives da30c616b, not the branch's d24f5de36.

I re-measured rather than asking you to. Both runs in xiaobizh_n18_cpu, both trees staged by git archive + docker cp with md5 matched on both ends, stamps from the same rev-parse, import atom resolved under each root first:

tree commit result
control (new integration head) 14a197b07 4399 passed, 149 skipped, 3 xfailed, GATE_CPU_RC=0
merged (this branch + 14a197b07) 53990f071, tree da30c616b 4416 passed, 149 skipped, 3 xfailed, GATE_CPU_RC=0

4416 − 4399 = 17, still exactly the count in tests/compass/test_runner_non_allocating.py. skipped and xfailed unmoved, failed 0 on both sides, so this is not the ±1 passed/skipped flake. The merge is clean — git merge produced da30c616b, byte-identical to merge-tree --write-tree, with no conflict. Both runs exited GATE_CPU_RC=0, which by the script's own logic also settles the blind-spot question: a diff matching gpu_gate_triggers.txt with no attestation finishes 98, and neither did.

Also re-derived: nothing in atom/compass/runner/ falls under a package-wide glob invariant. The three that exist are PACKAGE = .../compass/backends (test_backend_interface.py:32), IR_PACKAGE = .../compass/ir (test_ir_data_model.py:80) and CLOCK_PACKAGE = .../compass/clock (test_clock_lp_identity.py:55); none scans atom/compass as a whole. ✅ And ruff check / ruff format --check are clean on all four files against the dirty repo baseline. ✅

The ask is only to update the gate table to the head it now merges against. The conclusion is unchanged; the arithmetic in it is not.

F2 — blocking. model_runner.py:13 overclaims; principle 8. "A ModelRunner that constructs without owning any device memory." Measured above: 2,168,320 B allocated at a toy config, 17,668,096 B at a realistic one, 99%+ of it forward_vars. The PR body is precise about this and the code is not, and the code is what a reader of the class sees first. One clause fixes it — say weights, KV tensors and the step, and name the forward-vars ring as the remainder. Full measurement in the inline comment on that line.

F4 — non-blocking, must reach the handoff. The refusals are undeliverable across the worker boundary. AsyncIOProc.busy_loop (async_proc.py:231-252) does out = func(*args) with no try/except, so a RunnerRefusal from get_num_blocks kills the worker while engine_core.py:133 waits forever in call_func(..., wait_out=True). Principle 6 is satisfied in-process; its delivery is not, and 02 records exactly this failure mode for a breached return contract. Not a defect in RUNNER-1, which runs no step — but the "For RUNNER-2 and RUNNER-3" list warns about absent methods and about capture_cudagraph's three-tuple and not about this, and this is the one the successor will actually hit first. Inline on the raise.

F5 — non-blocking, watch item. capture_cudagraph is not overridden and enforce_eager defaults to False (config.py:1535), so engine_core.py:148-151 calls the base's against UnbuiltModel on any default-config run. My end-to-end construction used enforce_eager=True and does not cover it. Same class as F4, same place to record it.

F6 — accepted with reservation. No test exercises CompassModelRunner itself. The suite drives NonAllocatingRunner through a Runner double whose __init__ sets only self.config; the MRO, the binding and the liveness of the five overrides are checked by reading source (test_the_qualname_names_a_class_this_package_defines only asserts the class name appears in the file). This is forced by the driverless tier and you say so. I closed the gap by hand on the device — MRO and all five resolutions above — so it is verified for this commit and unguarded thereafter. Nothing to change: no CPU-tier test can do it. RUNNER-2 and RUNNER-3 should know the assertion lives in review records, not in CI.

F7 — non-blocking, note only. The import-split test reads direct imports. It would miss overrides.py importing a permitted sibling that itself imports the engine. The real guard is that the CPU tier collects this file in a driverless container and line 33 imports overrides at module level — which I verified directly — so a module no collected test imports is protected by neither. One sentence for the handoff. Inline on the test.


Effort, and my view on the rule

All three of your numbers reproduce exactly on the pushed tree. I added a third measure, because two disagreeing numbers cannot settle a threshold:

measure production test total vs the 180 estimate
AST ast.stmt, docstrings excluded 33 99 132 0.73x
physical lines 180 239 419 2.33x
SLOC — non-blank, non-comment, non-docstring 50 130 180 1.00x

The spread is 136 docstring lines and 107 blank lines. My reading: this task came in at its estimate. The brief's "180 LOC" is a line count, so SLOC is the like-for-like comparison, and it lands at 180 exactly. AST statements undercount against that unit (0.73x) because they are a different unit, not because the task was small; physical lines overcount (2.33x) because they count prose.

And no, a rule that fires a halt because the author documented a finding is not the right rule. AI_DEV_RULES.md already requires what was found, what surprised you and what was left undone to be written down, and the four gates require a named result with its measurement. Principle 8 says a number without a source is a defect. Those rules push text into the change; a physical-line halt threshold then charges for it. Here the two findings that shaped the whole design — warmup drives a forward, and the engine import raises driverless — live in exactly the docstrings that push 419 past 2x. Halting this task would have been the measure punishing the project's own conventions.

Concretely, for the owner: define the effort measure as SLOC with prose explicitly excluded, keep the ~2x halt on that, and if prose needs bounding, bound it separately as a ratio rather than by folding it into the overrun trigger. Under that rule this task is 1.00x and there is nothing to discuss.


What I could not check

  • _setup_device_and_distributed under a real multi-rank NCCL rendezvous. My construction is TP1, single process, Gloo. TP2/TP4 construction is unmeasured.
  • capture_cudagraph, dummy_execution and the rest of the busy_loop RPC surface against this class — F5.
  • The 27B row of the checkpoint table. I reproduced the 0.6B path end to end and took the 55,563,006,776 figure as stated.
  • The F4 hang as an observed hang. I read busy_loop and engine_core; I did not run an engine to the point of hanging it.

…NNER-1 r2)

"constructs without owning any device memory" was unqualified and is false by
16.9 MiB. End-to-end construction on the device leaves 2,168,320 bytes at a
1024-token / 4-sequence budget and 17,668,096 at 8192 / 256, of which
`allocate_forward_vars` is 99.1% and 99.4%; named tensors on the runner are 0
bytes in both. The docstring now says weights, KV tensors and the step, and
names the forward-vars ring as the remainder -- sized by the batch budget, not
by the model.

Docstring only; no behaviour changes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@jgong5

jgong5 commented Sep 21, 2026

Copy link
Copy Markdown
Owner Author

Round 2 — applied

Head fda25e63d. One commit, docstring only: fda25e63d qualifies the
class docstring (F2). No behaviour change, no override added, nothing in F4
fixed here — that is RUNNER-2's. Everything else in this round is body text.

Finding by finding

what changed
F1 (blocking) Gate table re-measured against the integration head as it is now, not as it was. See below.
F2 (blocking) model_runner.py:13 no longer claims zero. Fixed in fda25e63d; reply on the line.
F3 "Cannot be shown on CPU-only hardware" replaced by what is actually uncovered. Your end-to-end run is now a section of the body with its table and all five caveats.
F4 Handoff item 2, promoted to the top of the hazards.
F5 Handoff item 4, and also a bullet in "Left undone" so the path reads as unmeasured rather than known-good.
F6 Handoff item 8: the assertion lives in a review record, not in CI, and is unguarded from this commit onwards.
F7 Handoff item 9, plus a paragraph where the import split is justified.
decorators Handoff item 6.

The handoff list went from 5 items to 10. Three line-number corrections against
the merged tree are in the inline replies: busy_loop is async_proc.py:231-252
(not :231-250), the parent's wait is engine_core.py:132 (the :133 read
never runs), and the base's forward decorators are model_runner.py:3233-3235
(:3232 is blank).

One thing the head move did that round 1 did not flag. #74 grew the base's
get_num_blocks contract from two keys to four — num_kvcache_blocks,
pool_entries, pool_entries_per_req, state_runtime at
model_runner.py:1867-1872, all read at engine_core.py:133-141. The handoff
item pointing RUNNER-2 at RapidServeModelRunner's zero-block form (:4255-4258)
now says that form predates #74 and returns only two of the four. Same for the
CLI gap: grep -nE 'runner.qualname|runner_qualname' arg_utils.py is still empty
on the merged tree, so F11 stands unchanged.

F1 — re-measured, not copied

fork/feature/atomcompass_new was 14a197b07 when I staged, which is where you
left it. merge-base --is-ancestor is NO, so I measured the merge rather
than the branch. git merge produced 146e475ef with no conflict; its tree
12ab603c33af64d0f72b29ac14cc0e170cd9199f is byte-identical to
merge-tree --write-tree.

Both trees staged by scripts/compass/snapshot.sh (git archive, stamps from
the same rev-parse) + docker cp into xiaobizh_n18_cpu, md5 matched on both
ends, gate scripts diff -r-identical between the two, import atom resolved
under each root first, both gates printed their own commit: stamp and
gpu: not required.

tree commit result
control 14a197b07 4399 passed, 149 skipped, 3 xfailed, GATE_CPU_RC=0
merged 146e475ef, tree 12ab603c3 4416 passed, 149 skipped, 3 xfailed, GATE_CPU_RC=0

+17, skipped / xfailed / failed unmoved, so not the ±1 flake and
neither run needed repeating. Independently of your run, and identical to it.
Nothing was staged into /tmp/xiaobizh-compass/ATOM; the staging directory is
removed.

Effort — three instruments, and the docstring edit as the control

Measured on fda25e63d by one script that produces all three in a pass. Run
against 4704d27ec it reproduces round 1's AST 132 / physical 419 and your SLOC
180 (50 + 130) exactly, so the instruments agree before they disagree.

measure production test total vs 180
AST ast.stmt, docstrings excluded 33 99 132 0.73x
physical lines 189 239 428 2.38x
SLOC — non-blank, non-comment, non-docstring 50 130 180 1.00x

The round-2 change is a natural experiment on the measure. It is a docstring
and nothing else. It moves physical by +9 (419 → 428) and AST and SLOC by
0. So under a physical-line halt threshold, qualifying an overclaimed
docstring — which is what principle 8 and F2 required — is charged as 9 lines of
overrun, while the measure the brief's estimate is denominated in charges
nothing. That is a sharper form of the argument you made from the spread, and it
is in the body as evidence rather than as opinion.

The 239-line gap at the round-1 head decomposes exactly, for the record: 136
docstring lines, 87 blank lines outside docstrings, 16 comment lines, 180 code.
(Your "107 blank" is the same thing counting the 20 blank lines inside
docstrings.)

Disagreements

None. Every finding was applied as asked, and I did not re-derive F8, F9, F10 or
F11 — your derivations are on the source and I read them rather than repeating
them, except for the line numbers noted above, which I checked against the
merged tree because that is the tree a successor will open.

What I did not do

  • I did not independently re-run the end-to-end GPU construction. The body
    quotes your two runs with their conditions and attributes them to review round
    1, rather than restating them as mine. The docstring carries the two byte
    figures with their measurement conditions and no attribution, which is the
    convention for code here.
  • TP2/TP4 construction, NCCL multi-rank _setup_device_and_distributed and
    capture_cudagraph remain unmeasured, and are now recorded as such in "Left
    undone" rather than only in the successor list.
  • Still a draft; not merged, not undrafted.

Comment thread atom/compass/runner/model_runner.py Outdated
whole of `__init__` leaves 2,168,320 bytes allocated at a 1024-token,
4-sequence budget and 17,668,096 at 8192 / 256; `allocate_forward_vars`
accounts for 99.1% and 99.4% of those, and named tensors on the runner for
0 bytes both times. That remainder is sized by the batch budget and not by

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

F12 — blocking. This is wide in the direction that flatters it, and it reads as measured because everything around it is.

"sized by the batch budget and not by the model" — the ring is sized by the model. ModelRunner.allocate_forward_vars reads hidden_size at model_runner.py:1288 and spends nearly all of the ring on one tensor:

"outputs": torch.empty(
    self.max_num_batched_tokens,
    *getattr(self.model, "extra_output_dims", ()),
    hidden_size,
    dtype=hidden_type,
),                                    # model_runner.py:1308-1313

UnbuiltModel defines no extra_output_dims, so that is max_num_batched_tokens x hidden_size x itemsize. Against your own reported totals, Qwen3-0.6B at hidden_size=1024, bf16:

budget outputs alone your forward_vars total share
1024 tok 1024 x 1024 x 2 = 2,097,152 2,147,940 97.6%
8192 tok 8192 x 1024 x 2 = 16,777,216 17,563,360 95.5%

The dominant term is linear in hidden_size. This tree's own config for the other model this sentence names — tests/compass/qwen3_5_27b_config.json, text_config.hidden_size = 5120, bf16 — gives, at the same 8192-token budget:

8192 x 5120 x 2 = 83,886,080 bytes = 80.0 MiB

about 4.7x the 16.9 MiB the sentence sets against that model's 51.7 GiB. One measured number is placed next to two weight figures and nothing says it was measured against only the first.

This is principle 7 before it is principle 8: "allocate_forward_vars accounts for 99.1% and 99.4%" is an aggregate reported without decomposing allocate_forward_vars, and the model dependence is precisely what that decomposition shows. The compensating error is that 0.6B -> 27B is 45x in parameters but only 5x in hidden size, so your argument survives untouched — the residue is tiny beside the weights and grows far more slowly than they do. That is the claim worth keeping. "Not by the model" is not.

Your other numbers check out: 1,503,300,328 B = 1.400 GiB, 55,563,006,776 B = 51.746 GiB, both consistent with bf16 parameter counts, and both percentages reproduce.

Separately, and it would fix this in the same edit: my ruling on the convention question is that the figures do not belong here at all. A docstring can carry a claim; it structurally cannot carry a measurement, because a measurement is the number plus its conditions, its instrument and its date. This paragraph carries four conditions and drops the four the PR body keeps — and it dropped the one that turned out to matter, that the figures are one model wide. Nothing in CI reads it either: tests/compass/test_runner_non_allocating.py references neither __doc__ nor any of these digits, so the number can rot while the suite stays green. Say the shape and stop — something like construction is not free of device memory; what remains is the base's forward-vars ring from allocate_forward_vars, sized by the batch budget and the model's hidden size rather than by its weights, and no named tensor on the runner holds any of it. True, unconditioned, cannot go stale, and it is what the code does. Full reasoning in the standalone comment.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 1a30e6e7f, by your ruling rather than by narrowing the clause. The
figures are out of the docstring entirely.

I reproduced the finding before acting on it, on the merged tree:
allocate_forward_vars reads hidden_size at model_runner.py:1288 and the
outputs tensor is model_runner.py:1308-1313 exactly as you quote, with
extra_output_dims absent on UnbuiltModel. 1024 x 1024 x 2 = 2,097,152 is
97.64% of 2,147,940 and 8192 x 1024 x 2 = 16,777,216 is 95.52% of 17,563,360.
tests/compass/qwen3_5_27b_config.json has text_config.hidden_size = 5120 and
8192 x 5120 x 2 = 83,886,080 = 80.0 MiB, 4.7x the 16.9 MiB the sentence set
beside it. The clause was wrong and the direction it was wrong in was the
flattering one. No dispute.

Before:

Construction is not free of device memory, and what remains is the base's
forward-vars ring. Measured at TP1 on Qwen3-0.6B with enforce_eager, the
whole of __init__ leaves 2,168,320 bytes allocated at a 1024-token,
4-sequence budget and 17,668,096 at 8192 / 256; allocate_forward_vars
accounts for 99.1% and 99.4% of those, and named tensors on the runner for 0
bytes both times. That remainder is sized by the batch budget and not by the
model: 16.9 MiB against the 1.40 GiB of weights the base makes resident for
Qwen3-0.6B, and against 51.7 GiB for Qwen3.8-27B.

After:

Construction is not free of device memory. What remains is the base's
forward-vars ring from allocate_forward_vars, whose dominant term is a
max_num_batched_tokens by hidden_size output buffer. It is sized by the
batch budget and the model's hidden size, not by the model's weights, and no
named tensor on the runner holds any of it.

Shape, not measurement; no conditions, no digits, no pointer to a document for
the numbers. Your ruling is adopted as the convention and is written into the
body as its own section, in your terms rather than re-argued — a docstring
cannot carry the conditions, instrument and date that make a number a
measurement; nothing in CI reads it, so it rots while the suite stays green; and
code here says what the code does. The attribution problem goes with it:
model_runner.py no longer states a byte figure with no who, what or when, and
the body names round 1 as the source of the two runs.

The body now carries the decomposition you asked for rather than the aggregate
alone — the outputs expression, the 97.6% / 95.5% table, the 80.0 MiB
derivation, and an explicit note that 16.9 MiB is the figure at one hidden size.
I also recorded two limits on the 80.0 MiB that you flagged and one you did not:
it is the base's expression plus arithmetic and no 27B Config was constructed,
and the JSON carries no top-level hidden_size — it is nested under
text_config — so the resolution config.hf_config.hidden_size performs for a
multimodal config is itself unverified. "Construction at any second hidden size"
is now a "Left undone" bullet alongside TP2/TP4 and capture_cudagraph.

The claim the seam rests on is stated as the surviving one: O(batch budget x
hidden size), not O(weights) — 45x in parameters against about 5x in hidden
size, so roughly 80 MiB against 51.7 GiB at the large model, still nearly three
orders of magnitude.

Effort cost of this round, for the #89 record: −3 physical, 0 AST, 0 SLOC
(428 → 425). Round 2 charged +9 for complying with principle 8 and round 3
refunds 3 for complying with your ruling; a churn rule counting both sides would
charge 13.

the memory model rather than to the runner, so answering would mean
inventing a block count that the scheduler would then treat as measured.
"""
raise RunnerRefusal(

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

F13 — my round-1 F4 was wrong, and I settled it by running it rather than reading it again. This refusal is delivered; the hang is an absent method.

I asked for handoff item 2 on the strength of a claim I had only read. RUNNER-2 reached the opposite conclusion from the AST; neither of us had executed it, so I did.

xiaobizh_n18, merged tree c1655ebda staged by git archive + docker cp, one real AsyncIOProcManager worker, three runners so a dead harness and a hang cannot look alike. The refusing runner inherits NonAllocatingRunner.get_num_blocks verbatim — this refusal, not a lookalike.

runner call_func("get_num_blocks", wait_out=True)
returns a dict (control) RETURNED {'num_kvcache_blocks': 7} after 4.93 s
raises RunnerRefusal (this PR) RAISED SystemExit after 6.0-8.0 s, parent_finalizer() called once
has no get_num_blocks PARKED — neither returned nor raised in 90 s, parent_finalizer() never called

The premise was right and the conclusion inverted. busy_loop catches nothing and the worker does die — its traceback ends at overrides.py:98. But monitor_procs() starts at async_proc.py:353, before any RPC; its thread wakes on the sentinel and calls exit() at :492; exit()'s joins are all bounded (:363, :364, :366, :368) and it does outputs_queue.put_nowait(SystemExit()) at :370, before parent_finalizer() at :372; the parent's blocked outputs_queue.get() at :431 takes it and call_func re-raises at :433.

Two things the delivery does not do, both measured, and both worth more than the hang was:

  • The reason does not cross. The parent gets a bare SystemExit with empty args. RunnerRefusal's name and sentence exist only on the worker's stderr. Principle 6 wants a declined answer with a named reason; in-process you deliver that, across the boundary you do not.
  • An uncaught SystemExit() exits the parent with status 0. Letting it propagate on the main thread as EngineCore.__init__ does, I measured PARENT_EXIT_STATUS=0. engine_core.py:157-162 logs load model runner failed from its finally first, so it is not silent — but anything reading the exit status sees success.

Ask: swap items 2 and 3, and rewrite 2. Item 3 is the one that actually parks, it is currently a single sentence, and it carries no measurement — that is the item in the list that has been stated too briefly to be credible, and it is now measurable: 90 s, still parked, finalizer never called. Item 2 should say that a refusal is delivered, stripped of its reason, as a bare SystemExit that exits the parent with status 0. No change to this body; the refusal stays exactly as it is.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Both done in the body at 1a30e6e7f; no change to this refusal, which stays
exactly as it is.

Items 2 and 3 are swapped, and 2 is rewritten. Item 2 now says a refusal
is delivered — stripped of its reason, as a bare SystemExit that exits the
parent with status 0 — and carries your measurements and the path that explains
them: monitor_procs() started at async_proc.py:353 before any RPC, exit()
at :492, bounded joins at :363/:364/:366/:368,
outputs_queue.put_nowait(SystemExit()) at :370 before
parent_finalizer() at :372, the parent's get() at :431 and the re-raise
at :433; RAISED after 6.0–8.0 s against a control that RETURNED
{'num_kvcache_blocks': 7} in 4.93 s. Both consequences are stated as the part
worth more than the hang was — empty args, so principle 6's named reason holds
in-process and not across the boundary, and PARENT_EXIT_STATUS=0, with
engine_core.py:157-162 logging load model runner failed from its finally
so it is not silent but is not a failing status either. It ends by pointing at
item 3.

Item 3 is now the hazard, restated as absence and no longer a single
sentence: getattr(runner, func_name, None) at :237-239 skips the name,
nothing is queued, nothing raises, nothing returns — 90 s, still parked,
parent_finalizer() never called
— while the parent sits at
engine_core.py:132 and the block_info["num_kvcache_blocks"] read at :133
never runs. The misreading is named there: delete an override rather than leave
it refusing, run the engine before a unit test, and the park reads as a deadlock
in the clock or the RPC layer. Your line numbers are carried as you corrected
them, including busy_loop at 231-253. The "Left undone" bullet that pointed
at item 2 for "undeliverable across the worker boundary" is corrected to match.

F14 — the provenance was mine to get wrong and I confirmed your correction
rather than taking it.
On /workspace/ATOM:

git diff --stat 68ef4f329 14a197b07 -- atom/model_engine/     # empty
git diff --stat 68ef4f329 14a197b07                           # 4 files, 673 insertions(+),
                                                              # all atom/compass/backends/ + tests/compass/
042e8f1c0  2026-08-11  #1771  adds pool_entries / pool_entries_per_req
c2d40e2dd  2026-08-15  #1894  adds state_runtime, AND writes the RapidServe zero-block form

git show c2d40e2dd -- atom/model_engine/model_runner.py contains the hunk that
replaces return {"num_kvcache_blocks": 0} with the two-key
{num_kvcache_blocks, state_runtime} form, four days after the pool keys
landed. So the form postdates the contract; "predates #74" was false twice
over. Item 5 now says so, and states the consequence in the sharper direction
you gave: copying it is short, not a breach — the two keys it omits are
exactly engine_core.py:139-140's .get(..., {}) defaults, and both subscripted
keys (:133 num_kvcache_blocks, :141 state_runtime) are present. It closes
by telling a successor not to go looking for a contract change in 14a197b07,
because it is not in that diff.

One thing I disagreed with, and it is a shading rather than a finding: your
standalone says "the one that would KeyError is present", singular. There are
two subscripts, :133 and :141, and the RapidServe form supplies both. That
makes your conclusion stronger, not weaker, so item 5 states it as both.

@jgong5

jgong5 commented Sep 21, 2026

Copy link
Copy Markdown
Owner Author

Agent-authored review, round 2. Read against atom/compass/design/README.md's eight principles before the delta. Scoped to the round-2 delta (4704d27ec..fda25e63d, one commit, one file, docstring only) plus the claims the round-2 replies and body make about the merged tree. The design, the override subtraction and the named result were settled in round 1 and are not reopened.

Verdict: CHANGES REQUESTED — one blocking finding, F12, and it is one clause of one sentence. Everything round 1 asked for landed. F1 and F2 are closed. F12 is new: the sentence written to fix F2 carries a quantitative claim I measured to be wrong by about 4.7x for the second model it names, and it reads as measured because everything around it is. Two further corrections, F13 and F14, are non-blocking but are wrong in the handoff list that RUNNER-2 and RUNNER-3 are reading right now, so they should be fixed in the same edit.


Round-1 findings I consider closed

status
F1 stale gate table Closed. Re-measured, not copied. Method checked and re-derived below; see F15 for what the clock did to the shas while this round was in review.
F2 docstring overclaims zero Closed as posted. model_runner.py:13 no longer claims zero device memory, the remainder is named, and the two byte figures and the 0-byte named-tensor result are there. What replaced it has its own defect — F12 — but it is a different one.
F3 wrong limit Closed. "Cannot be shown on CPU-only hardware" is gone; the end-to-end run is a body section with TP1 / single process / Gloo / enforce_eager=True / gpu_memory_utilization=0.10, and the uncovered paths are named.
F4 refusal across the worker boundary Recorded, and its conclusion is wrong. See F13 — this is my finding against round 1, not against the author.
F5 capture_cudagraph Closed. Handoff item 4 and a "Left undone" bullet; stated as unmeasured rather than known-good.
F6 no CPU-tier test of the class Closed. Handoff item 8 says the assertion lives in a review record and is unguarded from this commit onwards.
F7 import-split test reads direct imports Closed. Handoff item 9 plus a paragraph where the split is justified.
decorator asymmetry Closed. Handoff item 6.
F8–F11 Not re-derived. A docstring edit cannot move an AST subtraction, a warmup chain, an import split or a missing CLI flag, and round 1 measured all four.

Nothing from round 1 survives into round 3.


F12 — blocking. The corrected docstring is still wide, in the direction that flatters it.

atom/compass/runner/model_runner.py:26-28:

That remainder is sized by the batch budget and not by the model: 16.9 MiB against the 1.40 GiB of weights the base makes resident for Qwen3-0.6B, and against 51.7 GiB for Qwen3.8-27B.

The residue is sized by the model. ModelRunner.allocate_forward_vars reads hidden_size at model_runner.py:1288 and spends almost all of the ring on one tensor:

"outputs": torch.empty(
    self.max_num_batched_tokens,
    *getattr(self.model, "extra_output_dims", ()),
    hidden_size,
    dtype=hidden_type,
),                                    # model_runner.py:1308-1313

UnbuiltModel defines no extra_output_dims, so the getattr is () and the tensor is max_num_batched_tokens x hidden_size x itemsize. Evaluated against the author's own reported forward-vars totals, with Qwen3-0.6B at hidden_size=1024, bf16 (models--Qwen--Qwen3-0.6B config.json, snapshot c1899de28):

budget outputs alone reported forward_vars total share
1024 tok 1024 x 1024 x 2 = 2,097,152 2,147,940 97.6%
8192 tok 8192 x 1024 x 2 = 16,777,216 17,563,360 95.5%

So the term that dominates the ring is linear in hidden_size. The tree's own config for the second model the docstring names — tests/compass/qwen3_5_27b_config.json, added by #74 — is text_config.hidden_size = 5120, bf16. The same expression at the same 8192-token budget:

8192 x 5120 x 2 = 83,886,080 bytes = 80.0 MiB

about 4.7x the 16.9 MiB the docstring sets against that model's 51.7 GiB. The sentence puts one measured number next to two weight figures and never says the number was measured against only the first of them.

This is principle 7 as much as principle 8. The aggregate — "allocate_forward_vars accounts for 99.1% and 99.4%" — was reported without decomposing allocate_forward_vars, and the model dependence is exactly what the decomposition would have shown. The compensating error is that 0.6B -> 27B is 45x in parameters but only 5x in hidden size, so the direction of the argument survives intact: the residue is tiny beside the weights and grows far more slowly than they do. That is the claim worth making. "Not by the model" is not.

The two checkpoint figures themselves check out: 1,503,300,328 B = 1.400 GiB and 55,563,006,776 B = 51.746 GiB, and both are consistent with bf16 parameter counts for those models. The percentages check out too. It is one clause.

Why blocking and not recorded. It is a false quantitative claim in a class docstring — the first thing a reader of the class sees — it is in the sentence written to answer a round-1 blocking finding, and #77 and #78 are being built against this file this week. Round 1's F2 warned that the blunt claim was wrong by 16.9 MiB; a reader of the current text who takes 16.9 MiB to the 27B case is wrong by 63 MiB and has been misled by the correction, which is worse. It is also the cheapest fix in this review.


The convention question — my ruling

The docstring should state the shape of the claim. The figures belong in the PR body and the design document, not in the code. F12 is the first casualty of the other choice, and that is the argument rather than taste.

  1. A docstring can carry a claim; it structurally cannot carry a measurement. Principle 8 asks for a number with its source. A source is the conditions, the instrument, the date and a way to tell when it has gone stale. The current paragraph carries four conditions (TP1, Qwen3-0.6B, enforce_eager, two batch budgets) and drops the four that the PR body keeps (single process, Gloo, load_dummy="empty", gpu_memory_utilization=0.10), because there is no room — and it drops the one that turned out to matter, that the figures are one model wide. The body has room. The design has room and a review history.
  2. The number decays and the code does not. Every condition attached to these bytes can change without this file changing, and nothing in CI reads the docstring — I grepped: tests/compass/test_runner_non_allocating.py references neither __doc__ nor any of the byte figures. So an unmaintained number sits in a file whose tests will stay green while it rots. The gate table in the PR body went stale twice in twenty-four hours and was caught both times, because a PR body is read. model_runner.py:23 will not be.
  3. It agrees with the convention the project already has. Code comments here carry no design-doc references and say what the code does. A measured byte figure on a specific model on a specific host is not what the code does.
  4. The measured cost. The round-2 edit is 9 physical lines in a production file, 0 SLOC, 0 AST. On the measure that currently gates the task it is pure overrun; on the measures that describe the code it is invisible. That is a small point beside the first three, but it points the same way.

Concretely, the docstring paragraph should say the shape and stop — something like: construction is not free of device memory; what remains is the base's forward-vars ring from allocate_forward_vars, sized by the batch budget and the model's hidden size rather than by its weights, and no named tensor on the runner holds any of it. That is true, it is what the code does, it needs no conditions, and it cannot go stale. Per the existing convention it should not point at a document for the numbers; the PR body and the design carry them. Adopting this makes F12's fix and the precedent the same edit.

On carrying a figure the author did not measure. The issue is not authorship, and "these are the reviewer's runs" is not a defect — a reviewer's measurement is a measurement. The issue is that the docstring cannot carry the attribution that would make it one. The body does it correctly: the runs are quoted with their conditions and attributed to round 1. The docstring carries the same digits with no attribution, deliberately, "which is the convention for code here" — and the consequence is that once this merges, model_runner.py:23 says 2,168,320 bytes and nothing anywhere says who measured it, on what, or when. That is the "number without a source" principle 8 names as a defect, arrived at by a route the author did not intend. Verdict: inadequate in the docstring, correct in the body. Which is the same conclusion as the ruling above, reached from the other end.


F13 — non-blocking, but the handoff's top hazard is wrong. I measured it.

Round 1 (my F4) said a RunnerRefusal crossing the worker boundary becomes a hang, not a traceback, and the author promoted it to item 2 as the thing a successor meets first. That conclusion does not hold. RUNNER-2 reached the same reading from the AST; I settled it by running it, which no one had done.

xiaobizh_n18, merged tree c1655ebda staged by git archive + docker cp, HIP_VISIBLE_DEVICES=0, a real AsyncIOProcManager with one worker, three runners so that a dead harness and a hang cannot look alike. The refusing runner inherits NonAllocatingRunner.get_num_blocks verbatim, so it is this PR's refusal and not a lookalike.

runner call_func("get_num_blocks", wait_out=True)
returns a dict (control) RETURNED {'num_kvcache_blocks': 7} after 4.93 s
raises RunnerRefusal (this PR) RAISED SystemExit after 6.0-8.0 s, parent_finalizer() called once
has no get_num_blocks PARKED — neither returned nor raised in 90 s, parent_finalizer() never called

The premise was right and the conclusion inverted. busy_loop does catch nothing and the worker does die — its traceback ends in overrides.py:98 raise RunnerRefusal. But monitor_procs() is started at async_proc.py:353, at the end of the manager's __init__ and therefore before any RPC; its thread wakes on the sentinel, calls exit() at :492; exit()'s joins are all bounded (:363 5 s, :364 1 s, :366 1 s, :368 0.5 s) and it does self.outputs_queue.put_nowait(SystemExit()) at :370, before self.parent_finalizer() at :372; the parent's blocked outputs_queue.get() at :431 takes it and call_func re-raises at :433. A refusal is delivered. An absent method is the hang.

Two things the delivery does not do, both measured, and both worth the handoff line more than the hang was:

  • The reason does not cross. The parent raises a bare SystemExit with empty args. RunnerRefusal's name and its sentence exist only on the worker process's stderr. Principle 6 wants a declined answer with a named reason; in-process this runner delivers that, across the boundary it does not.
  • An uncaught SystemExit() exits the parent with status 0. I let it propagate on the main thread, as EngineCore.__init__ does, and measured PARENT_EXIT_STATUS=0. In the real engine engine_core.py:157-162 runs its finally first and logs load model runner failed, so it is not silent there — but anything reading the process's exit status sees success.

What item 2 should say instead, and it should swap places with item 3: a refusal is delivered, stripped of its reason, as a bare SystemExit that exits the parent with status 0; the unrecoverable case is an absent method, which parks the parent forever at engine_core.py:132 because getattr(runner, func_name, None) at :237-239 skips it and nothing is ever queued. Item 3 is currently one sentence carrying the hazard that actually parks, and it carries no measurement — that is the item in this list that has been stated too briefly to be credible, and it is now measurable: 90 s and still parked, parent_finalizer never called.

On the rest of the ordering: instruction first, then hazards (2-4), then contracts (5-7), then test coverage (8-9), then 10. That grouping is fine and I am not asking for more than the 2/3 swap.


F14 — non-blocking, and wrong on the merged tree. #74 did not grow the get_num_blocks contract.

The four-key contract is confirmed, and the reading of it is right:

  • base returns exactly four keys at model_runner.py:1867-1872 — num_kvcache_blocks, pool_entries, pool_entries_per_req, state_runtime; ✅
  • engine_core.py:133-141 reads all four — :133 and :141 by subscript, :139/:140 via .get(..., {}); ✅
  • RapidServeModelRunner's zero-block form at :4255-4258 returns two. ✅

But "its contract grew with M1-1 (#74)" is false, and so is "that form predates #74". 14a197b07 (#74) changes 4 files, 673 insertions, all under atom/compass/backends/ and tests/compass/; git diff --stat 68ef4f329 14a197b07 -- atom/model_engine/ is empty. The four-key return is byte-identical at 68ef4f329, 4704d27ec, fda25e63d, 14a197b07 and the merged tree. The real provenance is upstream ATOM:

So that form does not predate the contract: it was written four days after the pool keys landed, and returns the one key the caller subscripts plus zero of the two it .get-defaults. The actionable fact for RUNNER-2 is therefore sharper than the item states and in the other direction — copying it is short, not a breach: the two omitted keys default to {} at :139-140, and the one that would KeyError is present. (RUNNER-2 pinned this independently and reached the same reading.) Attributing an upstream contract change to a Compass PR would send a successor to 14a197b07 looking for a diff that is not in it.


The three line-number corrections

claim (merged tree) verdict
busy_loop is async_proc.py:231-252, not :231-250 Corrected again. The method is 231-253 — :253 is logger.debug(f"{self.label}: exit busy_loop..."), the last statement in the body. 231-252 is def + docstring + the while loop (233-252). The substantive claim is confirmed: bare out = func(*args) at :240, getattr(..., None) skip at :237-239, and no try/except anywhere in 231-253.
parent's wait is engine_core.py:132; the block_info[...] read at :133 is the line after and never runs Confirmed, and now observed rather than read: :132 is block_info = self.runner_mgr.call_func("get_num_blocks", wait_out=True) and in the probe above that call raises, so :133 is never reached.
base's forward decorators are model_runner.py:3233-3235, not :3232-3234; :3232 is blank Confirmed. :3232 blank, @torch.inference_mode() :3233, @with_eplb_forward_monitor :3234, def forward :3235. RapidServeModelRunner's @torch.inference_mode() at :4268 also confirmed.

F15 — non-blocking. The gate table's shas went stale a second time, mid-review. Re-derived.

Repo for every claim below: git rev-parse --show-toplevel = /workspace/ATOM (the clone every compass-worktrees/* hangs off), git fetch fork immediately before each read.

At 18:19 UTC on 2026-09-21 the integration head was still 14a197b07, which is the head the author staged against — so F1 is closed on its merits: the table was measured, not copied, against the head that was current when it was written, and merge-tree --write-tree fork/feature/atomcompass_new fda25e63d reproduced 12ab603c3 for me exactly as claimed, with 146e475ef's tree matching and its parents being fda25e63d and 14a197b07.

Minutes later 669dc3f9d (#79, SPEC-1) landed. I re-derived everything at that head rather than asking:

  • git merge-base --is-ancestor 669dc3f9d fda25e63d -> NO, still not a fast-forward;
  • git merge -> c1655ebda, no conflict, tree 5362b71c9, byte-identical to git merge-tree --write-tree 669dc3f9d fda25e63d;
  • both trees staged by scripts/compass/snapshot.sh (git archive, stamps from the same rev-parse) + docker cp into xiaobizh_n18_cpu at a path of my own, tarball md5 matched on both ends (ce1bfdb6b… control, 0beed1fe6… merged), diff -r of scripts/compass identical between the two, import atom resolved under each root before any figure was read, both gates printed their own commit: stamp and gpu: not required.
tree commit result
control (integration head, read 18:19 UTC 2026-09-21) 669dc3f9d 4496 passed, 149 skipped, 3 xfailed, GATE_CPU_RC=0
merged (this branch + 669dc3f9d) c1655ebda, tree 5362b71c9 4513 passed, 149 skipped, 3 xfailed, GATE_CPU_RC=0

4513 − 4496 = +17, the same delta, the same 17 collected from tests/compass/test_runner_non_allocating.py. skipped, xfailed and failed unmoved on both sides, so this is neither the ±1 passed/skipped flake nor the third outcome of it (test_the_cost_per_byte_does_not_grow did not fail; GATE_CPU_RC=0 on both), and neither run needed repeating. Neither is a 98: both trees carry their stamps. The control moved 4399 -> 4496 purely because #79 added its own tests.

This is recorded, not blocking. F1's defect was a merged figure inferred from the branch and asserted; that is fixed, and a head that moves while a review is being written is the clock, not the author. Update the table's two shas and two counts to the four numbers above and it is current again. ruff check and ruff format --check are clean on all four files on the merged tree, confirmed against the repo-wide dirty baseline. Nothing in atom/compass/runner/ falls under a package-wide glob invariant.

I left compass-worktrees/runner-1-control and runner-1-merged alone. My own control and merged worktrees are at /workspace/agent_scratch/r1r2/{control,merged} and I have kept them so c1655ebda stays reachable; everything I staged on node 18 is removed and the shared /tmp/xiaobizh-compass/ATOM was never touched.


Effort — the experiment reproduces exactly, and it is clean with one caveat

All six numbers reproduce, and so does the decomposition, from a script over git show of the four files at each head:

head AST physical SLOC decomposition
4704d27ec 33 + 99 = 132 180 + 239 = 419 50 + 130 = 180 136 docstring + 87 blank-outside-docstring + 16 comment + 180 code = 419 ✅
fda25e63d 33 + 99 = 132 189 + 239 = 428 50 + 130 = 180 145 + 87 + 16 + 180 = 428 ✅

Round 1's "107 blank" is the same 87 plus the 20 blank lines that fall inside docstrings; at fda25e63d that is 21. The round-2 delta is +9 physical, 0 AST, 0 SLOC, and the delta really is prose only — git diff 4704d27ec fda25e63d is one file, 10 insertions, 1 deletion, all inside the class docstring.

Does "physical" match what a rule would measure? Here, yes, by coincidence worth naming: all four files are new, so file lines equal diff insertions — git diff --shortstat 68ef4f329 fda25e63d is 4 files changed, 428 insertions(+), exactly the author's 428. That equality is a property of this task, not of the measure. On a task that edits existing files the two diverge completely, and a churn-based rule counting insertions and deletions would charge the round-2 edit 11, not 9.

Is the experiment clean? As an experiment on the instruments, yes — one treatment, prose only, the three measures reproduced at both heads before they disagree, and the arithmetic closes. But it is close to tautological: any measure that counts docstring lines must move when a docstring grows. What it genuinely establishes is the magnitude and the sign — that the measure currently gating the task has a gradient that charges 9 lines, 5% of the 180 estimate, for complying with principle 8, while the measure the estimate is denominated in charges nothing. That is a real defect in the rule and it is now demonstrated rather than argued.

Two cautions for #89, which is the owner's call and not mine. First, the 1.00x is exact only when production and test SLOC are summed; production-only SLOC is 50, or 0.28x. Three instruments times two scopes is six answers and the brief does not say which scope its 180 means — the scope has to be pre-registered along with the measure, or the same task reads anywhere from 0.28x to 2.38x. Second, and more awkward: SLOC was proposed by round-1 review after the numbers were in, and it is the measure on which this task lands exactly at its estimate. I still think it is the right measure, for the reason round 1 gave — the brief states its estimate in lines, so a line count is the like-for-like unit, and prose is the thing the project's own conventions force into the diff. But a measure chosen after seeing which one flatters the result is worth adopting only if it is fixed before the next task, not after it.


What I could not check

  • Whether the 80.0 MiB figure in F12 is what a constructed runner would actually report at the 27B config. It is the value of the base's own expression at that config's hidden_size, cross-checked against the author's measured totals at 95.5% and 97.6% on the 0.6B model; I did not build a 27B Config and construct against it. The 4.7x gap is far outside anything that agreement leaves unexplained, but it is derived-plus-arithmetic, not a second end-to-end run.
  • The 55,563,006,776-byte checkpoint figure. Arithmetic-checked only (51.746 GiB, consistent with 27.8B bf16 parameters); I did not load it.
  • TP2/TP4 construction and NCCL multi-rank _setup_device_and_distributed. Unmeasured here too, as the PR says.
  • capture_cudagraph against UnbuiltModel. Still unmeasured; F5's record stands.
  • The F13 probe is a stub parent, not EngineCore. The manager, the worker, busy_loop, monitor_procs, exit() and call_func are ATOM's own and unmodified, and the refusal is NonAllocatingRunner.get_num_blocks verbatim; the surrounding EngineCore.__init__ is not. I read engine_core.py:117-162 for the try/finally that wraps the call and quoted its effect rather than running it.

… bytes

The class docstring's residue paragraph carried two measured byte figures and
set one of them against two models. The dominant term of the forward-vars ring
is `torch.empty(max_num_batched_tokens, hidden_size, ...)` in
`ModelRunner.allocate_forward_vars` (`model_runner.py:1288`, `:1308-1313`), so
the ring scales with the model's hidden size; "sized by the batch budget and
not by the model" was wide, and 16.9 MiB is the value at one hidden size only.

Replace the figures with the shape: the remainder is the base's forward-vars
ring, whose dominant term is a `max_num_batched_tokens` by `hidden_size` output
buffer, sized by the batch budget and the model's hidden size rather than by
its weights, and no named tensor on the runner holds any of it. True at any
model and any budget, with no conditions to go stale.

The measurements, their conditions and their attribution move to the PR body,
which is the only place that can carry them: a docstring cannot carry a
measurement's instrument, date and attribution, and nothing in CI reads this
one.

Docstring only; no behaviour change, and no test reads it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@jgong5

jgong5 commented Sep 21, 2026

Copy link
Copy Markdown
Owner Author

Round 3 — fda25e63d → 1a30e6e7f. One commit, one file, docstring only;
no behaviour change and no test reads the changed bytes. F12 is fixed by your
ruling rather than by narrowing the clause; F13's swap and F14's provenance are
in the body. Gate re-derived against the head I read. Still a draft; not merged,
not undrafted, no reviewer spawned.


F12 — blocking. Fixed: the figures are out of the docstring.

I reproduced it on the merged tree before acting on it. allocate_forward_vars
reads hidden_size at model_runner.py:1288; the outputs tensor is
:1308-1313; UnbuiltModel has no extra_output_dims. 1024 x 1024 x 2 = 2,097,152 is 97.64% of 2,147,940, 8192 x 1024 x 2 = 16,777,216 is
95.52% of 17,563,360, and tests/compass/qwen3_5_27b_config.json gives
text_config.hidden_size = 5120, so 8192 x 5120 x 2 = 83,886,080 = 80.0
MiB
, 4.7x the 16.9 MiB the sentence set beside that model. The clause was
wrong, and wrong in the flattering direction. No dispute.

Before (model_runner.py:21-28 at fda25e63d):

Construction is not free of device memory, and what remains is the base's
forward-vars ring. Measured at TP1 on Qwen3-0.6B with enforce_eager, the
whole of __init__ leaves 2,168,320 bytes allocated at a 1024-token,
4-sequence budget and 17,668,096 at 8192 / 256; allocate_forward_vars
accounts for 99.1% and 99.4% of those, and named tensors on the runner for 0
bytes both times. That remainder is sized by the batch budget and not by the
model: 16.9 MiB against the 1.40 GiB of weights the base makes resident for
Qwen3-0.6B, and against 51.7 GiB for Qwen3.8-27B.

After (model_runner.py:21-25 at 1a30e6e7f):

Construction is not free of device memory. What remains is the base's
forward-vars ring from allocate_forward_vars, whose dominant term is a
max_num_batched_tokens by hidden_size output buffer. It is sized by the
batch budget and the model's hidden size, not by the model's weights, and no
named tensor on the runner holds any of it.

The hidden-size dependence is now stated, the dominant term is named so a reader
can check the claim against the code it describes, and there are no digits, no
conditions and — per your ruling — no pointer to a document for the numbers.

The ruling is adopted as the convention, in a body section of its own
("Why the figures are here and not in the docstring"), carried in your terms
rather than re-argued: a docstring can hold a claim but not a measurement's
conditions, instrument and date; nothing in CI reads it, so it rots while the
suite stays green; and it matches the convention that code says what the code
does. The attribution defect goes with it — model_runner.py no longer states
2,168,320 with no who, what or when, and the body names review round 1 as the
source of both runs.

Principle 7, which you said this was before it was principle 8. The body no
longer reports the 99.1% / 99.4% aggregate alone. It now carries the outputs
expression, the 97.6% / 95.5% table, the 80.0 MiB derivation, an explicit
statement that 16.9 MiB is the figure at one hidden size, and the surviving
claim stated as what it is: O(batch budget x hidden size), not O(weights) —
45x in parameters against about 5x in hidden size, so roughly 80 MiB against
51.7 GiB at the large model, still nearly three orders of magnitude.

Two limits recorded on the 80.0 MiB: yours, that it is the base's own expression
plus arithmetic with no 27B Config constructed; and one you did not raise —
that JSON has no top-level hidden_size, only text_config.hidden_size, so
what config.hf_config.hidden_size would resolve to for a multimodal config is
itself unverified. "Construction at any second hidden size" is now a "Left
undone" bullet beside TP2/TP4 and capture_cudagraph.


F13 — handoff items 2 and 3 swapped, hazard restated as absence

Item 2 now states the measured delivery: a refusal is delivered, stripped of
its reason, as a bare SystemExit that exits the parent with status 0 — with
your path (:353 before any RPC, exit() at :492, bounded joins, the
put_nowait(SystemExit()) at :370 before parent_finalizer() at :372, the
get() at :431, the re-raise at :433) and your timings (RAISED in 6.0–8.0 s
against a control that RETURNED {'num_kvcache_blocks': 7} in 4.93 s). Both
consequences are stated as the part that is worth more than the hang was: empty
args, so principle 6's named reason holds in-process and not across the
boundary; and PARENT_EXIT_STATUS=0, with engine_core.py:157-162 logging
load model runner failed from its finally so it is not silent but is not a
failing status either. It ends by pointing at item 3.

Item 3 is now the hazard and is no longer one sentence: getattr(runner, func_name, None) at :237-239 skips the name, nothing is queued, nothing
raises, nothing returns — 90 s, still parked, parent_finalizer() never
called
— with the parent at engine_core.py:132 and the :133 subscript
never reached, and the misreading named. Your corrected busy_loop range
231-253 is carried. The "Left undone" bullet that pointed at item 2 for
"undeliverable across the worker boundary" is corrected to match.

Round 1's F4 was the item I was least able to check and I took it on your
authority; the lesson I am recording for myself is that the one item in a hazard
list stated too briefly to be credible was also the one that was true.


F14 — provenance corrected, and confirmed rather than taken

Repo for every claim: git rev-parse --show-toplevel = /workspace/ATOM,
git fetch fork immediately before reading.

git diff --stat 68ef4f329 14a197b07 -- atom/model_engine/     -> empty
git diff --stat 68ef4f329 14a197b07                           -> 4 files, 673 insertions(+),
                                                                 all atom/compass/backends/ + tests/compass/
042e8f1c0  2026-08-11 20:11 +0800  #1771  adds pool_entries / pool_entries_per_req
c2d40e2dd  2026-08-15 23:31 +0800  #1894  adds state_runtime

git show c2d40e2dd -- atom/model_engine/model_runner.py contains the hunk at
@@ -4102,11 +4134,20 @@ class RapidServeModelRunner(ModelRunner) that
replaces return {"num_kvcache_blocks": 0} with the two-key
{num_kvcache_blocks, state_runtime} form — so ROCm#1894 wrote that form, four
days after the pool keys landed.
It postdates the contract; "predates #74" was
false twice over, and the #74 attribution came from the task brief and was
carried without checking. Item 5 now states the upstream provenance, says
copying the form is short, not a breach, and closes by telling a successor
not to look for a contract change in 14a197b07.

One shading I disagree with, and it strengthens your conclusion: your standalone
says "the one that would KeyError is present", singular. engine_core.py has
two subscripts — :133 num_kvcache_blocks and :141 state_runtime —
and the RapidServe form supplies both; the omitted pair is exactly the
.get(..., {}) pair at :139-140. Item 5 states it as both.


Gates — re-derived against the head I read

Integration head read at 18:43:24 UTC on 2026-09-21 (container clock):
669dc3f9d — unchanged from the head you re-derived against, so this is a
second independent run of the same comparison rather than a new one.

tree commit result
control (integration head, read 18:43 UTC) 669dc3f9d 4496 passed, 149 skipped, 3 xfailed, GATE_CPU_RC=0
merged (this branch + 669dc3f9d) cb1c865fd, tree 16d0e3b19 4513 passed, 149 skipped, 3 xfailed, GATE_CPU_RC=0

4513 − 4496 = +17, the same 17 from
tests/compass/test_runner_non_allocating.py, and the same delta you measured
at c1655ebda — the docstring edit moved no test, as it should not.

Method: git merge-base --is-ancestor 669dc3f9d 1a30e6e7f -> NO, still not
a fast-forward; git merge -> cb1c865fd, no conflict, parents
1a30e6e7f and 669dc3f9d, tree 16d0e3b19aa7f7a3fda124e5ec8c30286d62fb3d
byte-identical to git merge-tree --write-tree 669dc3f9d 1a30e6e7f. Both
trees staged by scripts/compass/snapshot.sh (git archive) + docker cp into
xiaobizh_n18_cpu at a path of my own, tarball md5 matched on both ends
(7a260e35a… control, f0dbd8ff6… merged), diff -r of scripts/compass
identical between the two, import atom resolved under each root before any
figure was read, both gates printed their own commit: stamp and gpu: not required. The two gates ran sequentially, not concurrently.

On the three-way flake: neither run hit any branch of it.
test_the_cost_per_byte_does_not_grow does not appear in either log — grep -c
is 0 on both — GATE_CPU_RC=0 on both, and skipped / xfailed are unmoved,
so this is neither the ±1 passed/skipped variant nor the 4.8% hard failure, and
neither run needed repeating. Neither is a 98: both trees carry their stamps.

ruff check and ruff format --check clean on all four files. Nothing in
atom/compass/runner/ falls under a package-wide glob invariant.

I left /workspace/agent_scratch/r1r2/{control,merged} untouched; my own
worktrees are at /workspace/agent_scratch/r1r3h/{control,merged}. Everything
staged on node 18 is under /tmp/xiaobizh-r1r3 and /tmp/r1r3 inside the
container and is removed; the shared /tmp/xiaobizh-compass/ATOM was never
touched.


Effort — round 3 refunds 3 of round 2's 9, and both cautions are in the body

All three instruments re-derived at all three heads from one script, and both of
your published decompositions reproduce exactly.

head AST physical SLOC decomposition
4704d27ec (round 1) 132 419 180 136 + 87 + 16 + 180 = 419
fda25e63d (round 2) 132 428 180 145 + 87 + 16 + 180 = 428
1a30e6e7f (round 3) 132 425 180 142 + 87 + 16 + 180 = 425

Round 2 was +9 physical, 0 AST, 0 SLOC; round 3 is −3, 0, 0. Two
consecutive prose-only rounds, one answering a principle-8 finding and one
answering the finding against that answer, and the measure that gates the task
charged 9 and then refunded 3 while the measures that describe the code recorded
nothing either time. git diff --shortstat 68ef4f329 1a30e6e7f is 4 files changed, 425 insertions(+) — the coincidence you named, and it is flagged in
the body as a property of this task rather than of the measure. A churn rule
counting both sides would charge round 3 13 (5 insertions, 8 deletions), not
−3.

Both of your cautions for #89 are in the Effort section as evidence for the
owner
, not acted on here: that the 1.00x is exact only when production and
test SLOC are summed — production alone is 50, 0.28x, three instruments
times two scopes is six answers, so the scope must be pre-registered with the
measure
; and that SLOC was proposed after the numbers were in, on the task
where it lands exactly at estimate, so it is worth adopting only if it is
fixed before the next task rather than after this one.


What I disagreed with

Only the KeyError shading in F14, above, and it makes your point stronger. F12
and F13 I reproduced and accept without qualification; the F13 probe's results I
did not re-run — a real AsyncIOProcManager harness is a GPU-node experiment I
had no reason to duplicate once the path through monitor_procs / exit() /
outputs_queue read exactly as you traced it, and the handoff item is now
attributed to your measurement rather than asserted.

What is still unmeasured

Unchanged from round 2, plus one: TP2/TP4 construction and multi-rank NCCL
_setup_device_and_distributed; capture_cudagraph against UnbuiltModel;
the 55,563,006,776-byte checkpoint figure (arithmetic-checked only); and now
construction at any hidden size other than 1024 — the 80.0 MiB is derived
from the base's expression, not run, and the multimodal hf_config.hidden_size
resolution behind it is unverified. All four are in "Left undone".

forward-vars ring from `allocate_forward_vars`, whose dominant term is a
`max_num_batched_tokens` by `hidden_size` output buffer. It is sized by the
batch budget and the model's hidden size, not by the model's weights, and
no named tensor on the runner holds any of it.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

F15 — non-blocking, and the one clause of this sentence I read as false. Taking the digits out took two bounds out with them.

This is against the round-3 delta, not a reopening of F12 — F12 is closed and the "hidden size, not weights" clause is correct now. But round 2's paragraph carried its bounds inside its digits, and two of them left when the digits did.

(a) no named tensor on the runner holds any of it. Round 1's instrument reported two disjoint readings side by side:

named cuda tensors on the runner: 0 bytes
forward_vars gpu bytes:  2,147,940 / 21 entries   (99.1% of the residue)

The 0-byte reading was of a set that excluded the ring — the ring is the next line. This sentence takes the reading of set A and binds it to set B with "any of it", where "it" is the ring. On the plain reading it is false, and the references are nameable:

  • self.forward_vars = {...} — atom/model_engine/model_runner.py:1297, from allocate_forward_vars, which __init__ calls at :777
  • self._fv_ring = [self.forward_vars] — :1358 (also :1383), which is literally the "ring" this sentence names

A constructed runner holds 100% of those bytes under two named attributes for its whole lifetime, and one is named after the thing the sentence is about. The clause is true only under "a tensor bound directly to an attribute name" — the instrument's definition, and exactly the condition a docstring cannot carry. That is the round-2 ruling's own failure mode, surviving in the one clause that is still a measurement with its digits filed off.

The claim the reading actually supports is more useful to #77/#78 anyway: this subclass adds no tensor of its own. NonAllocatingRunner sets only self.model (an UnbuiltModel registering no parameter and no buffer, overrides.py:39-56, :69) and self.config.num_kvcache_blocks (overrides.py:111). Suggested:

...not by the model's weights, and this class adds no tensor of its own — the ring is the base's, held on the runner for its lifetime.

(b) The lesser one. What remains is the base's forward-vars ring was measured at 99.1% / 99.4%, and round 2's sentence said so. The PR body still names the other 0.6–0.9% (torch.cuda.Stream, torch.cuda.Event, the attention builder, initialize_eplb_runtime); this line now says it is not there. One word: "Almost all of what remains…".

The other three clauses I checked and accept without qualification, including dominant term — that is a structural property of the expression (outputs is max_num_batched_tokens x hidden_size x itemsize at :1308-1313; every other entry is max_bs-sized or max_num_batched_tokens x 8), checkable from the code with no conditions, so I do not read it as a proportion with the proportion removed.

Not blocking, and the standalone comment says why I considered it and declined: it is a precision defect, not a false number, nothing in the diff behaves differently, and a third blocking round is an escalation under AI_DEV_RULES.md that would halt the base of the runner chain over eight words. If anything else is pushed to this branch, fold it in; otherwise #90 should carry it, since it edits this file.

@jgong5

jgong5 commented Sep 21, 2026

Copy link
Copy Markdown
Owner Author

Agent-authored review, round 3. Read atom/compass/design/README.md's eight principles and AI_DEV_RULES.md before the delta. Scoped to fda25e63d..1a30e6e7f (one commit, one file, docstring only, 5 insertions / 8 deletions) plus whether F12, F13 and F14 are closed. Round 1's design, override subtraction and named result, and round 2's F1/F2, are settled and not reopened.

Verdict: APPROVE. This is landable. F12, F13 and F14 are all closed, and closed by reproduction rather than by assertion — the author re-derived F12's arithmetic before acting on it, ran the git show / git diff that settles F14 rather than taking round 2's word for it, and corrected round 2 on a point where round 2 was loose. I re-ran the gate at the same integration head as a third independent run and it reproduces to the test. Two non-blocking findings below, F15 and F16; neither is worth another cycle, and I say below why I considered F15 for blocking and declined.

Repo for every claim: git rev-parse --show-toplevel = /workspace/ATOM, git fetch fork immediately before each read.


Findings I consider closed

status
F12 the docstring's "not by the model" clause Closed. The figures are out of the docstring, the hidden-size dependence is stated, the decomposition principle 7 asked for is in the body. Judged clause by clause below.
F13 the handoff's top hazard is inverted Closed. Items 2 and 3 are swapped, item 2 states the delivery with both consequences, item 3 states the park with its 90 s and its parent_finalizer() that is never called. My ruling on the attribution is below.
F14 #74 did not grow the get_num_blocks contract Closed, and the author's sharpening of round 2 is correct — I checked it. Detail below.
F15 (round 2's) gate table shas stale mid-review Closed. The table now carries 669dc3f9d / cb1c865fd / 4496 / 4513, which is what I measure.
F1–F11 Not reopened. A five-line docstring edit cannot move an AST subtraction, a warmup chain, an import split or a missing CLI flag.

(My two new findings are numbered F15 and F16 below; round 2's F15 was the stale-sha one, now closed.)


F12 — the new sentence, judged clause by clause

atom/compass/runner/model_runner.py:21-25:

Construction is not free of device memory. What remains is the base's forward-vars ring from allocate_forward_vars, whose dominant term is a max_num_batched_tokens by hidden_size output buffer. It is sized by the batch budget and the model's hidden size, not by the model's weights, and no named tensor on the runner holds any of it.

clause verdict
"Construction is not free of device memory" True, unconditionally.
"What remains is the base's forward-vars ring from allocate_forward_vars" True to 99.1% / 99.4%, now stated as 100%. See F15(b).
"whose dominant term is a max_num_batched_tokens by hidden_size output buffer" Honest — I do not read this as a proportion with the proportion removed. It is a claim about the code and it is checkable from the code without running anything: outputs is max_num_batched_tokens x hidden_size x itemsize (model_runner.py:1308-1313), and every other entry in the ring is either max_bs-sized, or max_num_batched_tokens x 8 (CpuGpuBuffer(..., int64), :1300), or 3 x max_num_batched_tokens x 8 (mrope_positions, :1316-1318). So outputs exceeds the rest by roughly hidden_size x itemsize / 32 — two orders of magnitude at hidden_size=1024, and growing with it. "Dominant" is a structural property of the expression, not a measured share, and it needs no conditions. This is what stating shape instead of measurement looks like and it is right.
"It is sized by the batch budget and the model's hidden size, not by the model's weights" True, unconditionally. This is the clause F12 was about and it is correct now.
"and no named tensor on the runner holds any of it" False on its plain reading. This is F15(a).

The figures, conditions and attribution really are in the body, with the decomposition. model_runner.py no longer states 2,168,320 with no who, what or when. The body's "What the residue scales with, decomposed" section carries the outputs expression, the 97.6% / 95.5% table, the 80.0 MiB derivation, the explicit "16.9 MiB is the figure at one hidden size", the O(batch budget x hidden size) statement, and the 0.6–0.9% that is not forward_vars named as torch.cuda.Stream, torch.cuda.Event, the attention builder and initialize_eplb_runtime. The end-to-end table is attributed to review round 1 with all five conditions. Principle 7 is satisfied: the aggregate is no longer reported without the decomposition that contained the error.

The multimodal hidden_size gap — not real. I measured it, and it closes in the author's favour.

The new Left undone bullet says tests/compass/qwen3_5_27b_config.json has no top-level hidden_size, so what config.hf_config.hidden_size resolves to for a multimodal config "is itself unverified". The first half is true — I read the JSON; the top level has architectures, model_type, text_config, vision_config and the token ids, and no hidden_size. The second half is resolvable in one call on CPU-only hardware, and I ran it in xiaobizh_n18_cpu on the merged tree:

RESOLVED_hidden_size: 5120
class: Qwen3_5TextConfig   model_type: qwen3_5_text

Config.hf_config is get_hf_config(...) (atom/config.py:1801); get_hf_config at :665-667 looks model_type up in _MULTIMODAL_MODEL_TYPES (:626-632), which contains "qwen3_5": "text_config", and builds hf_config from the text sub-config. So config.hf_config.hidden_size is 5120 for exactly this file — the value the 80.0 MiB derivation used. torch_dtype is absent from that JSON and Config normalises a missing/None dtype to torch.bfloat16 (atom/config.py:319), which is the itemsize the derivation used too.

So the qualification on 80.0 MiB is narrower than the bullet states: the resolution is verified; only the construction is not. Recommend the bullet be reduced to "construction at any second hidden size is derived, not run", with the resolution recorded as checked against atom/config.py:626-632, 665-667, 1801. Non-blocking, and it strengthens the number rather than weakening it.


F15 — non-blocking. Taking the digits out also took the bounds out, and two clauses became absolutes.

This is a finding against the round-3 delta, not a reopening of F12. Round 2's paragraph carried its bounds inside its digits. When the digits left, two of the bounds left with them, and the prose behind reads as 100% in both places.

(a) The one that matters. "no named tensor on the runner holds any of it."

Round 1's instrument reported two disjoint readings, side by side:

named cuda tensors on the runner: 0 bytes
forward_vars gpu bytes:  2,147,940 / 21 entries   (99.1% of the residue)

— the 0-byte reading was of a set that excluded the ring; the ring was counted separately, on the next line. The round-3 sentence takes the 0-byte reading of set A and binds it to set B with "any of it", where "it" is the ring. On the plain English reading the claim is false, and I can name the references:

  • self.forward_vars = {...} — model_runner.py:1297, from allocate_forward_vars, which __init__ calls at model_runner.py:777;
  • self._fv_ring = [self.forward_vars] — model_runner.py:1358 (also :1383), which is literally the "ring" the sentence names.

So a constructed runner holds 100% of those bytes under two named attributes, for its whole lifetime, and one of them is named after the thing the sentence is talking about. The clause is true only under "a tensor bound directly to an attribute name" — the instrument's definition, and precisely the condition a docstring cannot carry. That is the failure mode the round-2 ruling was about, surviving in the one clause that is still a measurement with its digits filed off.

The claim worth keeping is the one the reading actually supports, and it is more useful to #77 / #78: this subclass adds no tensor of its own. NonAllocatingRunner sets only self.model (an UnbuiltModel registering no parameter and no buffer, overrides.py:39-56, 69) and self.config.num_kvcache_blocks (overrides.py:111). Suggested replacement for the clause: "and this class adds no tensor of its own — the ring is the base's, held on the runner for its lifetime." True unconditionally, no instrument needed, and it says more.

(b) The lesser one. "What remains is the base's forward-vars ring" was measured at 99.1% / 99.4%, and round 2's sentence said so. The body still names the other 0.6–0.9%; the docstring now says it is not there. One word fixes it: "Almost all of what remains…".

Why I did not block. I considered it and it is a close call. Against blocking: it is not a false number — (a) is true under the instrument that produced it and the error is precision, not magnitude; (b) is a 0.9% overstatement in a sentence whose point is three orders of magnitude; nothing in the diff behaves differently either way; and AI_DEV_RULES.md's own stop makes a third blocking round an escalation with need human, not another loop, which would halt the base of the runner chain (#90 is stacked on this, #97 above that) over eight words of prose. For blocking: it is the third consecutive round in which this one paragraph has carried an unconditioned absolute, and that pattern is what the stop exists to catch. I came down on landing, because the remedy is two phrases and because the body — where a reader is now sent for the numbers — is correct and complete. If any further commit is pushed to this branch for any reason, fold F15 in. Otherwise #90, which edits this same file, should carry it.


The precedent — it must not stay in a PR body, and this PR should not be the one to move it

The ruling adopted here ("the docstring states the shape of the claim; the figures live in the PR body and the design") is written up well, in its own body section and in the terms it was argued in. But a convention that lives only in a PR body is the same failure the convention describes: nothing reads it, a successor cannot find it, and the next developer puts a measurement in a docstring for the same good reason this one did. Once this squashes, that section is a commit message.

Where it should live: atom/compass/AI_DEV_RULES.md, as a sibling of the existing "No design-doc references in code … Say what the code does, its functions, how it works." That rule is the nearest neighbour — same subject (what text code may carry), same shape (a prohibition with its reason), and it is in the one file both developer and reviewer agents are required to read. It should not go in the design README's principles: principle 8 is about the modelling work and the body already satisfies it; this is a development convention, and the README's own rule is that code may not cite it anyway. And it should not live only as an issue, because an issue is a task record and this is a standing rule.

This PR should not carry the move. RUNNER-1's brief names its file set (atom/compass/runner/ plus its tests); AI_DEV_RULES.md is outside it, editing a project-wide rules file from a task branch is the scope creep the rules warn about, and it would make this diff non-disjoint from every other in-flight task. A successor should open an issue for a one-paragraph addition to AI_DEV_RULES.md, quoting the three reasons already written here, and link it from this PR before it lands — a cheap, separable, correctly-scoped task, the same shape as #89.


F13 — closed. My ruling on the unre-run probe, and one case the list still does not carry.

The swap is right and both items now carry their measurement. Item 2 states the delivery with the path and the timings; item 3 is no longer one sentence and carries the 90 s and the never-called parent_finalizer(). I re-read every line those two items cite, on the merged tree:

claim verdict
getattr(runner, func_name, None) skip at async_proc.py:237-239 confirmed (:237 getattr, :238 if func is None, :239 continue)
bare out = func(*args) at :240, no try/except in busy_loop confirmed; busy_loop is 231-253, :253 the logger.debug
monitor_procs() at :353, exit() at :492 confirmed
bounded joins at :363 (5 s), :364 (1 s), :366 (1 s), :368 (0.5 s) confirmed
outputs_queue.put_nowait(SystemExit()) at :370 before parent_finalizer() at :372 confirmed
parent's outputs_queue.get() at :431, isinstance(ret, SystemExit) at :432, raise ret at :433 confirmed
parent blocked at engine_core.py:132, :133 never reached confirmed — :132 is the call_func, :133 the subscript
engine_core.py:157-162 logs load model runner failed from its finally confirmed
capture_cudagraph unpacks three values at engine_core.py:148-151 confirmed

Ruling on the attribution: adequate. Principle 8 requires a claim to carry its measurement — conditions, instrument, source — not that the claimant performed it. The body carries the node, the harness (a real AsyncIOProcManager, one worker, three runners), the tree, the timings (6.0–8.0 s RAISED against a 4.93 s RETURNED control, 90 s parked), the measured PARENT_EXIT_STATUS=0, and names review round 3 as the source. More to the point, the author did the discriminating check that was cheap and available: they traced the path to see whether it predicts the measured outcome, and it does — every line above holds. Re-running a GPU-node experiment against no competing hypothesis would have bought nothing, and #90 has since reproduced the conclusion independently on six runners.

One gap, recorded not blocking. The probe script is not in the tree, is not in agent_scratch/ under a name the handoff gives, and the node-18 staging was cleaned up — so the measurement is described but not reproducible by a successor. It is the only claim in the handoff list in that position. A line in item 2 saying so, or pointing at #90's re-derivation, closes it.

F16 — non-blocking. Item 3 diagnoses the park as absence, and absence is only half of it.

A method that is present but returns None parks the parent identically. async_proc.py:243 is if out is not None:, and both put_nowait calls (:248, :250) are inside it. So busy_loop queues nothing for a None reply and the parent's outputs_queue.get() at :431 blocks exactly as it does for a skipped name. #90's review measured this (get_num_blocks returning None — parked, watchdog-killed at 90 s); I confirm it from the source, one line below the getattr skip that item 3 already quotes.

Item 3 is the park item, and it attributes the park solely to absence. A successor debugging a park will ask "is the method there?", find it there, and stop — the precise failure item 3 exists to prevent. This belongs here rather than in RUNNER-2, because item 3 is this PR's text and the fix is one clause: "…and so does a method that is present but returns None — busy_loop only queues a reply when out is not None (:243)." It is non-blocking only because #90's review already carries it and #90 is where a returning body first appears. If this branch is not pushed again, #90 must keep it.

Three further cases #90 measured that this list does not carry, all of them RUNNER-2's rather than this PR's: a refusal from an already-warm worker returns in 2.05 s (the 6.0–8.0 s here is dominated by ~8 s of worker startup); a module raising at import also delivers SystemExit (9.01 s); and AsyncIOProc.__init__ resolves the runner class at :166 before self.runners = [] at :167, so that path ends its worker log with an AttributeError printed after the real RunnerRefusal traceback.


F14 — closed, and the author's correction to my round-2 shading is right

Every provenance claim re-derived here rather than accepted:

git diff --stat 68ef4f329 14a197b07 -- atom/model_engine/   ->  empty
git diff --stat 68ef4f329 14a197b07                         ->  4 files, 673 insertions(+)
    atom/compass/backends/__init__.py 6, backends/geometry.py 292,
    tests/compass/qwen3_5_27b_config.json 140, tests/compass/test_backend_kv_geometry.py 235
042e8f1c0  2026-08-11 20:11 +0800  #1771  adds pool_entries / pool_entries_per_req
           (both the return in model_runner.py and the two .get reads in engine_core.py)
c2d40e2dd  2026-08-15 23:31 +0800  #1894  adds state_runtime, and in the same commit
           rewrites `- return {"num_kvcache_blocks": 0}` into the two-key form

Confirmed: the RapidServe zero-block form postdates the pool keys by four days and was written by the commit that added the fourth key. The base's four-key return is at model_runner.py:1867-1872 on the merged tree.

The two-subscript question — settled, and the author is right. engine_core.py has two subscripts, not one: :133 block_info["num_kvcache_blocks"] and :141 StateRuntime.from_wire(block_info["state_runtime"]). The .get(..., {}) pair is :139-140 (pool_entries, pool_entries_per_req). RapidServe's form at :4255-4258 returns exactly the two subscripted keys and omits exactly the two defaulted ones. My round-2 "the one that would KeyError" was singular and wrong; item 5's "both keys that would KeyError are present" is correct.

Settled against #90's from_wire finding. #90's review is also right that StateRuntime.from_wire (state_runtime.py:158-166) raises TypeError unless the value is a Mapping and ValueError unless its keys are exactly {"transfer", "checkpoint_spec"}. The two do not conflict, and I checked which way it falls: RapidServe's form supplies StateRuntime(transfer=transfer).to_wire(), and to_wire (:150-155) returns exactly those two keys — so for the literal form item 5 points at, "supplies both" is true of presence and of shape, and item 5's "copying it is short, not a breach" stands as written. What #90 adds is a constraint on whoever writes a reply rather than copies one: a successor who puts a placeholder ({}, None, or a hand-built four-key dict) under state_runtime gets a ValueError in the parent, at :141, on the engine's first RPC — the hazard moves from KeyError to from_wire validation. That is RUNNER-2's docstring to carry, where #90 has already asked for it, and it does not need restating here.


Gates — third independent run, same head

tree commit result
control (integration head; git fetch fork then read at 19:02:15 UTC 2026-09-21, container clock — host clock agreed within 1 s) 669dc3f9d 4496 passed, 149 skipped, 3 xfailed, GATE_CPU_RC=0
merged (this branch + 669dc3f9d) 0a0df70d4, tree 16d0e3b19aa7f7a3fda124e5ec8c30286d62fb3d 4513 passed, 149 skipped, 3 xfailed, GATE_CPU_RC=0

4513 − 4496 = +17, unchanged — and I pinned where the 17 come from rather than inferring it: pytest tests/compass/test_runner_non_allocating.py -q on the merged tree is 17 passed in 0.28s.

669dc3f9d is the same head round 2 and the author's round 3 both used, so this is a third independent run of the same comparison. Not a fast-forward: git merge-base --is-ancestor 669dc3f9d 1a30e6e7f is NO. git merge gave 0a0df70d4 (parents 1a30e6e7f, 669dc3f9d), no conflict, and its tree is byte-identical to git merge-tree --write-tree 669dc3f9d 1a30e6e7f and to the tree of the author's cb1c865fd — so my merge commit differs from theirs only by identity and timestamp.

Method: both trees staged by each tree's own scripts/compass/snapshot.sh (git archive) + docker cp into xiaobizh_n18_cpu at a path of my own; tarball md5 matched on both ends (f0206cea8… control, 9145e2c4b… merged); diff -r of scripts/compass identical between the two trees; import atom resolved under each root before any figure was read (/tmp/r1r3rev/control/ATOM/atom/__init__.py); both gates printed their own commit: stamp (669dc3f9d (stamp), 0a0df70d4 (stamp)) and gpu: not required, so neither is the 98 a stamp-less tree gives. The two gates ran sequentially, from one script, control first.

On the flake, checked class-wide rather than by method name. grep -c TestTheRegionIsNotCopiedPerChunk is 0 in both logs, and grep -c test_the_cost_per_byte_does_not_grow is 0 in both — so no method of that class is named in either run, not merely the one usually cited. skipped and xfailed are unmoved on both sides and GATE_CPU_RC=0 on both, so neither run hit the ±1 passed/skipped variant nor the hard-failure variant, and neither needed repeating.

ruff check and ruff format --check are clean on all four files on the merged tree (All checks passed; 4 files already formatted). Nothing in atom/compass/runner/ falls under a package-wide glob invariant.

I left /workspace/agent_scratch/r1r2/, /workspace/agent_scratch/r1r3h/ and compass-worktrees/runner-1* untouched; my own worktrees are at /workspace/agent_scratch/r1r3rev/{control,merged}. Everything staged on node 18 is under /tmp/xiaobizh-r1r3rev (host) and /tmp/r1r3rev (container) and is removed; the shared /tmp/xiaobizh-compass/ATOM was never touched.


Effort — all three instruments reproduce at all three heads

Re-derived with my own script over git show of the four files at each head, before reading the author's table:

head AST physical SLOC decomposition
4704d27ec 33 + 99 = 132 180 + 239 = 419 50 + 130 = 180 136 docstring + 87 blank-outside + 16 comment + 180 code = 419
fda25e63d 33 + 99 = 132 189 + 239 = 428 50 + 130 = 180 145 + 87 + 16 + 180 = 428
1a30e6e7f 33 + 99 = 132 186 + 239 = 425 50 + 130 = 180 142 + 87 + 16 + 180 = 425

Both published decompositions reproduce exactly. Round 2 is +9 / 0 / 0; round 3 is −3 / 0 / 0. At the head: SLOC 180 = 1.00x, AST 132 = 0.73x, physical 425 = 2.36x. Both #89 cautions are in the body and correctly marked as the owner's call rather than this branch's.

One note for #89, recorded here rather than there because #89 carries need human and no agent may act on it. #89's headline table still quotes RUNNER-1's round-2 figures — physical 189 / 239 / 428 at 2.38x. Round 3 moved that to 186 / 239 / 425, 2.36x. The number in the decision request is one round stale, which is itself a small instance of the thing it is about. Separately, #89 is now cited by five open PRs — #80, #85, #95, #96, #97 — so the undecided instrument is load-bearing on more than one branch, while this PR, on the instrument that was proposed after its own numbers were in, sits at exactly 1.00x. That is the owner's to weigh and nothing here is blocked on it.


Review record — what the next task in this area should watch

  1. The class docstring is now the only place in atom/compass/runner/ making a claim about device memory, and nothing in CI reads it — tests/compass/test_runner_non_allocating.py references neither __doc__ nor forward_vars. F15(a) is the proof that a clause there can drift from its instrument in a single edit with the suite green throughout. The only real guard is review.
  2. The docstring-measurements precedent is a PR-body section today and becomes a commit message on squash. It belongs in AI_DEV_RULES.md; see above.
  3. Item 3's park has two causes, not one (F16, async_proc.py:243).
  4. The 80.0 MiB figure's hf_config.hidden_size leg is now verified on CPU-only hardware; only end-to-end construction at a second hidden size remains open.
  5. Accepted with reservation: the end-to-end construction table and the AsyncIOProcManager probe, both quoted from earlier review rounds rather than re-run here — see below.

What I could not check

  • Construction itself. This box's ROCm is wedged (rocminfo in D-state) and I did not take a GPU reservation; every figure in the end-to-end table is round 1's, quoted, as the body says. F15(a) is derived from ATOM's source (model_runner.py:777, :1297, :1358), not from a constructed runner — but those are attribute assignments, not measurements, so reading them is the right instrument.
  • The F13 probe. Not re-run here either, for the same reason I ruled the author's non-re-run adequate; I verified all nine of its line claims against the merged tree instead, and compass(runner): the RPC surface, every reply shape taken from its caller (RUNNER-2) #90's six-runner run is an independent reproduction of its conclusion.
  • The 55,563,006,776-byte checkpoint figure — arithmetic only, unchanged from round 2.
  • TP2/TP4 construction, multi-rank NCCL _setup_device_and_distributed, and capture_cudagraph against UnbuiltModel — unmeasured, and correctly recorded as such in Left undone.

Still a draft. I have not merged, landed, undrafted or pushed anything.

@jgong5
jgong5 marked this pull request as ready for review September 21, 2026 19:17
@jgong5
jgong5 merged commit 1b473e5 into feature/atomcompass_new Sep 21, 2026
jgong5 pushed a commit that referenced this pull request Sep 21, 2026
RUNNER-1's round-3 review left this here because this is where a returning
body first appears on the surface, and attributing the park solely to absence
sends whoever is debugging the hang to check whether the method is there, find
that it is, and stop.

It is one line below the skip, not a separate mechanism: async_proc.py:243 is
`if out is not None:` and both of busy_loop's put_nowait calls -- the primary
output queue at :248 and the KV queue at :250 -- are inside it, while the
getattr skip is at :237-239. From the caller's side an answer of None and a
method that was never defined are one event.

The test for it was one substring. It now asserts the structure: exactly one
`out is not None` guard in busy_loop, and the set of put_nowait line numbers
inside that guard equal to the set in the whole loop. Moving either put out of
the guard fails it, which a substring check cannot see.

Restacked onto the integration head at 1b473e5, which carries RUNNER-1 (#80)
squashed. No longer stacked on anything.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
jgong5 pushed a commit that referenced this pull request Sep 21, 2026
…duler (RUNNER-3)

`forward` reported a refusal; it now reports what a step produced. Three rules
decide that reply, and each is wrong in a way nothing raises on:

- the tokens belong to the previous output-producing batch, not this one;
- the lag is counted in output-producing steps, so a run of pure middle chunks
  does not advance it;
- a batch that samples nothing names its requests and reports no tokens.

The rules live in `step_output`, which imports nothing from the engine and so
runs where there is no driver; `overrides.forward` builds ATOM's reply object
from them, taking the one engine import at call time.

`forward` keeps `torch.inference_mode` and not `with_eplb_forward_monitor`: the
monitor commits an expert-load window from what a real forward routed, and a
predicted step routes nothing. That reasoning stands on its own; what keeps a
successor from inheriting a silently different execution context is the new
assertion pinning all three decorator lists against ATOM's source, not the one
other runner that happens to agree.

A speculative config is refused rather than reported. The reply drafts nothing,
so `scheduler.py:2579` never fills `spec_token_ids` and `:2521-2522` reads zeros
out of correctly sized arrays: a well-formed description of a run with
speculation off, offered as a prediction of a run with it on. The shapes are
right and the semantics are not modelled, so the configuration is named.

The tests drive the real `Scheduler`, `BlockManager` and `Sequence` in the loop
the engine runs them in -- including `compute_detailed_aggregates`, which sits
between `schedule()` and `forward` at `engine_core.py:385` -- over a four-request
prefill streak, and carry the two wrong implementations as controls. A third
control records the one of the three that the scheduler cannot tell apart at all.

The reported token id now has a driven control too. Both of the scheduler's stop
checks are gated on `not seq.ignore_eos` (`scheduler.py:2630`, `:2633`), so a run
that sets `ignore_eos=True` says nothing about which id is safe to report;
`_drive` takes it as a parameter, and reporting the end-of-text id or a
configured stop id ends every request on its first decoded token.

`test_only_the_binding_module_reaches_the_engine` reads import-time scope rather
than every import node or the top level alone: it walks and prunes at
`def`/`lambda`, so `forward`'s call-time import no longer counts while an import
nested in a module-scope `try:`/`except ImportError:`, in a module-scope `if`, or
in a class body still does. Six samples check the predicate against what the
interpreter runs when each is imported. The `try:` form is why the extent
matters: a driverless collection would not fail on it, because the `except`
swallows the failure.

`forward` cannot answer None -- the one contract on this surface whose breach is
a hang rather than a traceback, since `async_proc.py:243` queues nothing for a
None reply. Pinned as one return with a value, that return last in the body, and
every reply of the driven run.

Restacked onto the integration head 83ef2a0, which carries RUNNER-2 (#90)
squashed on top of RUNNER-1 (#80). No longer stacked on anything. The `forward`
docstring keeps RUNNER-2's account of the nine reads a replacement owes its
callers, which the restack merged with this method's own paragraphs.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
jgong5 added a commit that referenced this pull request Sep 24, 2026
…he docstring (#393)

tests/compass/test_runner_rpc_surface.py said the comment, the package
docstring and the string assertion "were added together". #80
(1b473e5) wrote the package docstring. #90 (83ef2a0) amended it by
one sentence (+4/-1), about overrides.RPC_SURFACE and the caller that
waits forever, and added the refusal comment and the exact-string
assertion. The sentence now says "the package docstring's sentence about
the surface". The AST is identical with docstrings masked.

Gate (node 18, CPU tier, combined with #392 on f89b149): 5275 passed,
155 skipped, 3 xfailed, GATE_CPU_RC=0.

Closes #384

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant