Skip to content

compass(runner): the exit bullet points at the refusal comment instead of restating its losses - #402

Merged
jgong5 merged 1 commit into
feature/atomcompass_newfrom
compass/issue-398
Sep 24, 2026
Merged

jgong5 merged 1 commit into
feature/atomcompass_newfrom
compass/issue-398

Conversation

@jgong5

@jgong5 jgong5 commented Sep 24, 2026

Copy link
Copy Markdown
Owner

Closes #398

Decision: point at the comment, not pin

The brief left one choice: pin the bullet's loss list with about 10 test lines, or point the bullet at the refusal comment. I pointed. The comment's loss list is already pinned: test_what_the_comment_says_a_hole_at_exit_loses_is_what_exit_does holds it against ModelRunner.exit's body. The bullet was a second copy that nothing pinned, and it had drifted into claiming the opposite of the comment. Deleting the copy removes the drift at its source. Pinning would have added a second pin for the same facts, plus about 10 test lines to maintain. Under /ponytail full, deleting beats adding.

The change: 3 docstring lines in atom/compass/runner/__init__.py

Before (:32-34 at b7cd11d48; the bullet runs to :36):

exit (engine_core.py:260) never reaches ModelRunner.exit, so the distributed environment is never destroyed and the graphs and five KV tensors it deletes stay held.

After (:32-34):

exit (engine_core.py:260) never reaches ModelRunner.exit; the comment on the RPC_SURFACE check in model_runner says what that loses here.

The rest of the bullet is unchanged: "The worker still leaves its loop -- busy_loop breaks on the dispatched name, [...] so the symptom is what shutdown failed to release, not a hang."

Evidence: every claim the bullet makes, quoted at the tip b7cd11d48

The new bullet makes three claims, and each one is quoted below. It names no loss of its own: the losses are now named only in the comment.

claim in the bullet quote
exit is dispatched at engine_core.py:260 atom/model_engine/engine_core.py:260: self.runner_mgr.call_func("exit")
a hole "never reaches ModelRunner.exit" atom/model_engine/model_runner.py:1052: def exit(self):. The comment says the same thing: "ModelRunner.exit never runs" (atom/compass/runner/model_runner.py:67).
"the comment on the RPC_SURFACE check in model_runner" atom/compass/runner/model_runner.py:54-55: _UNANSWERED = unanswered_rpc_names(CompassModelRunner) / if _UNANSWERED:. The comment is at :56-85, inside that if. unanswered_rpc_names (overrides.py:147-148) is "The dispatched names getattr would answer with None for this runner", drawn from RPC_SURFACE. The model_runner module bullet in the same docstring already says it is "where the composed class is checked against overrides.RPC_SURFACE".

What the comment names (model_runner.py:67-71), against ModelRunner.exit's body (atom/model_engine/model_runner.py:1052-1080)

the comment says ModelRunner.exit Compass overrides
"the distributed environment is never destroyed" :1058 destroy_dist_env() not overridden
"self.model is never dropped" :1074-1075 if hasattr(self, "model"): / del self.model overrides.py:248 self.model = UnbuiltModel(model_class): the model exists, but no weight was built
"torch.cuda.empty_cache() never runs" :1078 torch.cuda.empty_cache() not overridden
"Its five KV-tensor deletions are hasattr-guarded and find nothing here, because this runner allocated none" :1065-1073 loop over kv_cache, kv_scale, index_cache, mamba_k_cache, mamba_v_cache, each if hasattr(self, attr): delattr(self, attr) overrides.py:396-397 allocate_kv_cache: "Record the block count and allocate nothing."

What the old bullet claimed, and why it was false for this runner

old claim ModelRunner.exit Compass overrides
"the graphs [...] it deletes stay held" :1060-1061 if not self.enforce_eager: / self.graphs = self.graph_pool = None. ATOM creates self.graphs only inside its own capture_cudagraph, at :3832. overrides.py:410-411 capture_cudagraph: "Capture nothing". There is no graph for exit to release.
"five KV tensors it deletes stay held" :1065-1073 overrides.py:396-397 "allocate nothing". There is no tensor for exit to release.

The #399 review measured the same thing with an AST walk (#399 (comment)). ModelRunner writes the KV names and the graph attributes in only three methods: allocate_kv_cache, capture_cudagraph and exit. NonAllocatingRunner replaces the first two.

Checks on the docstring's readers

Only tests/compass/test_runner_rpc_surface.py reads PACKAGE_DOC (:48). Its readers are at :1118, :1129, :1152, :1170, :1174 and :1214. I re-ran their text logic, with no torch import, on the tip docstring and on the head docstring. The results were identical:

  • Bullet keys, used by the name-partition test and the module-bullet test. Both sides give exit, model_runner, overrides, process_kvconnector_output, step_output. The edit keeps the - `exit` lead, and its new backticked names sit mid-line, where _bullets does not read them as keys.
  • CITATION matches, used by the citation check. Both sides give the same six sites, with engine_core.py:260 the only one in the exit bullet. The pointer names model_runner without a file.py:NNN, so it adds no citation.
  • Stated module counts. None on either side.

Masked AST: atom/compass/runner/__init__.py, with docstrings masked (1/1), is identical between b7cd11d48 and the head. The diff is 3 docstring lines and nothing else.

Gate 1: node 18, xiaobizh_n18_cpu

  • Tree gated. git merge-tree --write-tree b7cd11d48 cfb36f81f = 4529ebe10d78, which equals cfb36f81f^{tree}. That is the tree staged with git archive plus docker exec -i ... tar -x into /tmp/i398gate/ATOM. The tarball md5 6483d50d7667... matched on both ends. .compass-commit and .compass-changed were written from the same rev-parse.
  • Scripts and import. The gate ran from the tree's own scripts/compass/gate_cpu.sh, with timeout -k 10 3000 and no pipe. It printed commit: cfb36f81f (stamp). atom.__file__ = /tmp/i398gate/ATOM/atom/__init__.py.
  • Result. 5275 passed / 155 skipped / 3 xfailed, GATE_CPU_RC=0.
  • Control. The tip b7cd11d48 has tree 869545f94c51, which is the tree compass(tests): say the refusal comment is what the exit test holds against ModelRunner.exit #399 gated at 5275 / 155 / 3, rc=0. Delta: 0. No test was added.
  • Aborted first attempt. My first launch overlapped by 1 s with another agent's gate. I stopped mine (EXIT=143) and re-ran alone, and the result above is from that re-run.

Not covered

The comment names no graph release and no del self.drafter (:1076-1077). The graph release finds nothing here (see above). For the drafter, I did not measure whether a speculative config builds one in __init__ (model_runner.py:771-776) before overrides.py:516 refuses it. The comment is outside this PR's file set, so that question stays with the comment and its test.

🤖 Generated with Claude Code

…d of restating its losses

The package docstring's `exit` bullet said a hole at exit leaves "the graphs
and five KV tensors it deletes" held. This runner allocates neither:
`NonAllocatingRunner.allocate_kv_cache` records a block count and allocates
nothing, and `capture_cudagraph` captures nothing, while ATOM creates
`self.graphs` only inside its own `capture_cudagraph`. The comment on the
`RPC_SURFACE` check in `model_runner` already says the opposite, and a test
holds it against `ModelRunner.exit`'s body. The bullet was a second, unpinned
copy of that loss list, and it had drifted.

The bullet now names the dispatch site, says `ModelRunner.exit` is never
reached, and points at that comment for what is lost. The sentence about the
loop still breaking is unchanged.

Closes #398

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
distributed environment is never destroyed and the graphs and five KV tensors
it deletes stay held. The worker still leaves its loop -- `busy_loop` breaks
- `exit` (`engine_core.py:260`) never reaches `ModelRunner.exit`; the comment
on the `RPC_SURFACE` check in `model_runner` says what that loses here.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Non-blocking. Filed as #405, not this PR's to fix.

The pointer is right, and the comment is the only place the losses should live. But "says what that loses here" hands the comment a completeness claim that it does not meet under one config.

  • The comment's list (model_runner.py:66-71) names no del self.drafter (atom/model_engine/model_runner.py:1076-1077).
  • I measured that a drafter can exist on this runner at exit. ModelRunner.__init__ owns construction, and it calls build_drafter at :776, after the _build_and_load_model override at :757.
  • The only speculative refusal in the runner is in forward (overrides.py:517 at 37fba4df0), so a step is the only thing that reaches it.
  • With an eagle3 config at the default aux setting, Drafter.load_model(UnbuiltModel) returns None and arm_aux_capture returns None. So the runner constructs holding a drafter, and one shut down before its first step reaches exit with it.
  • The pin (test_what_the_comment_says_a_hole_at_exit_loses_is_what_exit_does) checks only comment ⊆ exit body, so it cannot see the omission.

The larger problem is in #405: the drafter is built on the device and its checkpoint is read. That breaks principle 2: "Simulated execution touches no GPU. No compute, no device allocation."

Rule: "A finding not fixed in the PR that found it gets an issue." (AI_DEV_RULES)

@jgong5

jgong5 commented Sep 24, 2026

Copy link
Copy Markdown
Owner Author

Review cycle 1: PR #402, head cfb36f81f0d13fc5ea542f61002d7a9c7af71746

Agent-authored review. I read the Design principles of atom/compass/design/README.md and AI_DEV_RULES.md at 37fba4df0 first.

Verdict: APPROVE. There are no blocking findings. One finding is non-blocking: it is outside this PR's file set and is filed as #405, with an inline comment at __init__.py:33. The gate on the merged tree is green: 5276 passed, GATE_CPU_RC=0.

1. The new bullet, clause by clause

Clause Check Result
"exit (engine_core.py:260)" atom/model_engine/engine_core.py:260 is self.runner_mgr.call_func("exit"). The file is unchanged between cfb36f81f and 37fba4df0. holds
"never reaches ModelRunner.exit" busy_loop resolves getattr(runner, func_name, None) and continues on None (async_proc.py:237-239). A hole therefore never calls the method. holds
"the comment on the RPC_SURFACE check in model_runner" model_runner.py has exactly one indented comment: lines 56-85, inside if _UNANSWERED: (:55). The only other # lines are the SPDX header at :1-2. _refusal_comment() in the test file relies on that same uniqueness. The check is unanswered_rpc_names(CompassModelRunner), drawn from RPC_SURFACE, and the docstring's own model_runner bullet (:21-22) calls it "checked against overrides.RPC_SURFACE". unambiguous
"says what that loses here" :66-71 names the lost destroy_dist_env, the lost del self.model and the lost torch.cuda.empty_cache(), and says the five KV deletions find nothing. That is a loss list, and it is scoped to this runner. It omits del self.drafter, which is reachable under eagle3. See §3 and #405. holds, but incomplete for one config (non-blocking)
The rest of the bullet (unchanged) The break is a sibling of the per-runner loop (async_proc.py:251-252). holds

The PACKAGE_DOC tests, run on node 18. I ran tests/compass/test_runner_rpc_surface.py whole in xiaobizh_n18_cpu on the staged merged tree: 69 passed, rc=0, with atom.__file__ = /tmp/r402gate/ATOM/atom/__init__.py. That run covers:

  • test_the_package_docstring_lists_every_module_beside_it (bullet partition);
  • test_a_module_count_the_package_docstring_states_is_the_packages;
  • test_the_package_docstring_partitions_the_surface_the_way_the_table_does;
  • test_every_site_the_package_docstring_cites_is_one_no_caller_waits_for (citation check);
  • test_a_module_bullet_cites_the_engine_imports_its_module_makes;
  • test_what_the_comment_says_a_hole_at_exit_loses_is_what_exit_does.

The full gate below includes them again.

Design-doc references: none in atom/compass/runner/__init__.py at the head. This is the PR's whole file set.

2. Point versus pin

Pointing is the right call, and it loses nothing a reader needs.

  • The losses are facts about ATOM's exit body. They drift whenever that body is edited.
  • The comment already holds the list, and a test already holds the comment against that body. A second copy in the docstring is the copy that drifted: it claimed graphs and KV tensors are held, when this runner allocates neither.
  • A reader now needs one hop, within the same package, to a location the bullet names without ambiguity.
  • Pinning the bullet would have pinned the same partial list twice. It would not have caught the drafter omission either, because the existing pin is one-directional (each loss named ⊆ exit's body).

Principle 3: "Prioritise simplicity. Add only what is necessary, and nothing more."

3. The drafter question: measured, reachable, filed as #405

I ran drafter_probe.py in xiaobizh_n18 (one device visible) on the same staged tree. The draft-checkpoint loader was stubbed. The results:

  • Method owners. __init__ and exit are ModelRunner's. _build_and_load_model, _maybe_warmup and forward are NonAllocatingRunner's.
  • Call order in ModelRunner.__init__ (AST). :757 _build_and_load_model, then :776 build_drafter, :784 drafter.load_model, :832 drafter.arm_aux_capture and :836 _maybe_warmup.
  • Refusals before a step. Neither override on that path raises. forward is the only NonAllocatingRunner method that reads speculative_config.
  • Drafter.load_model(UnbuiltModel). eagle3 returns None. mtp raises AttributeError: 'UnbuiltModel' object has no attribute 'model'.
  • EagleProposer.arm_aux_capture(UnbuiltModel). It returns None with aux off (the default) and with aux on and no ids. It raises with configured ids.

So self.drafter can exist when exit would run. An eagle3 config at default aux settings constructs a live runner holding a drafter, and the refusal fires only at the first step. The comment's loss list is therefore incomplete for that config.

The bigger finding is that the drafter is built on the device and its checkpoint is read. That breaks principle 2: "Simulated execution touches no GPU. No compute, no device allocation." The MTP path dies with an unnamed AttributeError instead of the refusal's reason, which breaks principle 6: "Refuse rather than fall back. A declined answer with a named reason is a result."

The fix belongs in overrides.py and the comment, both outside this PR's file set. Filed as #405 per AI_DEV_RULES: "A finding not fixed in the PR that found it gets an issue." Non-blocking for #402. This PR's pointer stays correct whatever #405 does to the comment.

4. ponytail-review

The diff is 3 docstring lines, a net-zero line count, and it replaces a second copy of the facts with a reference. Nothing to cut.

Lean already. Ship.

5. Gate 1 on the merged tree

Findings

# Finding Severity Where
1 The refusal comment's loss list omits del self.drafter, which is reachable under eagle3. The underlying defect is that the drafter is built on the device before any refusal. non-blocking; filed as #405 inline at atom/compass/runner/__init__.py:33

No blocking findings.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant