Skip to content

compass(gates): locate aiter for gate_gpu.sh without importing it - #271

Merged
jgong5 merged 1 commit into
feature/atomcompass_newfrom
compass/issue-166-gate-gpu-find-spec
Sep 23, 2026
Merged

jgong5 merged 1 commit into
feature/atomcompass_newfrom
compass/issue-166-gate-gpu-find-spec

Conversation

@jgong5

@jgong5 jgong5 commented Sep 23, 2026

Copy link
Copy Markdown
Owner

Closes #166

What changed

scripts/compass/gate_gpu.sh located aiter's checkout from aiter.__file__, so recording the aiter version imported aiter. That import runs aiter's architecture probe, which shells out to rocminfo and hangs uninterruptibly on a wedged driver.

Only the locating step changes. It now calls the artifact store's module_root("aiter"), which resolves the top-level name with importlib.util.find_spec and executes nothing. The git describe --tags --always --dirty call and the BASE_AITER comparison are unchanged.

file lines
scripts/compass/gate_gpu.sh (production) +6 / -1: one code line and five comment lines
tests/compass/test_gate_gpu_aiter_version.py (test) +65

The total is 71 lines against a 20-40 estimate. That is inside the 2x stop, and most of it is the test's stub-package setup.

Decision: reuse ART-1's helper instead of duplicating it

The gate calls atom.compass.artifacts.provenance.module_root and has no second copy of the lookup. Reuse is only safe if importing that module does not itself reach aiter or the driver, so I measured that in xiaobizh_n18_cpu, with PYTHONPATH set to the staged branch root:

  • from atom.compass.artifacts.provenance import module_root took 38.3 ms and added 73 modules. None of them is torch, aiter, numpy, zmq or triton. The atom modules it loads are atom, atom.sampling_params, atom.plugin{,.prepare,.sglang,.sglang.prepare} and atom.compass.artifacts.*, all standard-library only. atom/__init__.py already defers LLMEngine to first attribute access.
  • module_root("aiter") returned /app/aiter-test/aiter, and afterwards aiter was not in sys.modules.

The test enforces this condition rather than taking it for granted. Its stub torch also raises, so a future eager import on this chain fails the test.

Named result

No module is executed. The test cuts the gate's own lines out of gate_gpu.sh, from AITER_DIR= to the closing fi, and runs them with bash. PYTHONPATH is <stub site>:<repo>, where the stub site is a committed, tagged git checkout that holds an aiter/__init__.py and a torch/__init__.py. Each stub writes a marker file and raises when executed. The test asserts three things:

  • no marker exists;
  • AITER_DIR is the stub's package directory;
  • AITER equals git describe --tags --always --dirty run on that directory.

It compares against that command's output, never a version literal, because two aiter versions are in circulation.

The test catches a restored import. I ran a line-count-preserving mutation and a null control in xiaobizh_n18_cpu. Each ran against its own copy of the staged tree, gate_gpu.sh stayed at 326 lines in every copy, and atom.__file__ was printed under each copy's root:

variant edit to gate_gpu.sh result
branch none 2 passed
mutation line 158 restored to AITER_DIR=$(python -c 'import aiter, os; print(os.path.dirname(aiter.__file__))' 2>/dev/null) 2 failed, rc=1
null control line 154 comment: wedged driver changed to stuck driver 2 passed

The mutation fails these two tests:

  • tests/compass/test_gate_gpu_aiter_version.py::test_aiters_version_is_read_without_executing_aiter_or_torch, with AssertionError: a stub was executed / assert ['aiter'] == []. This is the behavioural failure: the old path executed the stub aiter and read UNKNOWN.
  • tests/compass/test_gate_gpu_aiter_version.py::test_the_located_name_is_top_level, where 'module_root("aiter")' no longer appears in the resolution lines.

The same string on a healthy box. I ran a read-only check in node 18's GPU container xiaobizh_n18. It ran only the resolution snippet from a staged copy of this branch; it did not run the GPU gate and took no GPU. The staged copy was removed afterwards. Results:

path directory version time
new (module_root) /app/aiter-test/aiter v0.1.21.dev0-49-gf4e7c7509 0.07 s
old (import aiter) /app/aiter-test/aiter v0.1.21.dev0-49-gf4e7c7509 8.33 s

The two versions match each other and BASE_AITER.

Gates

Gate 1: ATOM's suite, unmodified. Measured on node 18, xiaobizh_n18_cpu, using the tree's own scripts/compass/gate_cpu.sh. Both trees were git archive stages with .compass-commit and .compass-changed stamps, and atom.__file__ was asserted under each staged root.

control ff9617f30 (tip) branch 1b8d1e3cf
pytest 5120 passed, 149 skipped, 3 xfailed 5122 passed, 149 skipped, 3 xfailed
GATE_CPU_RC 0 0
gpu line not required not required (.compass-changed stamp)
  • Node-id delta, from junit XML: 2 node ids exist only on the branch, the two new tests. None exists only on the control, and no node id changed status.
  • GPU tier: GATE_CPU_RC is not 98. gate_gpu.sh is not a GPU-tier trigger, because the trigger list names atom/ modules, so the GPU tier did not fire and was not run.
  • Merge check: git merge-tree --write-tree ff9617f30 1b8d1e3cf gives 2114ab788302bc26e719cf499538211618ea0ff5 with rc 0. The branch is based on the current tip.

Gate 2: CPU-only tests. The new test file passes, and ruff check, ruff format --check and black --check are clean on it.

Gate 3: the named result, shown above.

Gate 4: independent review. Pending; the coordinator dispatches it.

Surprises

  • The premise of reuse has changed. scripts/compass/README.md still says that on a wedged node ATOM's import hangs. That is no longer true of import atom: LLMEngine is now lazy, and the chain used here loads neither torch nor aiter. I did not edit the README, because it is outside this file set.
  • The GPU tier does not cover this change. gpu_gate_triggers.txt names atom/ modules, not scripts/compass/, so the CPU gate says "gpu: not required" for this change. The only real-aiter evidence is the read-only node-18 check above.

Left undone

  • No check on the local wedged box. I did not run anything in the local gpu_docker container, as the brief directs, so there is no direct demonstration on a wedged driver. The raising stub is the stand-in: it proves nothing is executed, which is the property that decides whether a hang can happen.
  • The README line. The scripts/compass/README.md statement about import atom hanging is left for a follow-up.

🤖 Generated with Claude Code

gate_gpu.sh read aiter's checkout from aiter.__file__, so recording the
aiter version imported aiter, whose architecture probe shells out to
rocminfo and hangs uninterruptibly on a wedged driver. The checkout is
now located by the artifact store's module_root, which resolves the
top-level name with importlib.util.find_spec and executes nothing. The
git describe call and the BASE_AITER comparison are unchanged.

The new test runs the gate's own resolution lines against stub aiter
and torch packages that raise when executed, and fails if the import is
restored.

Closes #166

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
# version is most worth recording. module_root resolves a top-level name with
# importlib.util.find_spec, which executes nothing, and its import chain loads
# neither aiter nor torch.
AITER_DIR=$(python -c 'from atom.compass.artifacts.provenance import module_root; print(module_root("aiter"))' 2>/dev/null)

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reuse ruling (principles 3, 4 and 2): accepted. Keep the atom import. Do not inline it. Not blocking.

  • The gate already depends on import atom being cheap. Line 148 calls compass_require_tree "$ROOT", which runs python -c 'import atom; ...' under the same PYTHONPATH=$ROOT, ten lines before this one. This line only extends an existing dependency into the atom.compass.artifacts package; it adds no new kind of dependency.
  • One locator, not two (principle 4). The provenance stanza and this gate record the same aiter version. Because they share module_root, they cannot disagree about which checkout that version came from. An inlined find_spec one-liner would be a second definition that could drift from the first.
  • Measured on the merged tree 2114ab788, in xiaobizh_n18_cpu:
    • from atom.compass.artifacts.provenance import module_root took 37.7 ms, loaded 146 modules, and loaded none of torch, aiter, numpy, zmq or triton.
    • Bare import atom took 16.2 ms and loaded 104 modules, also none of them.
  • The torch stub does catch a regression on this chain. Each mutation below appends ; import torch # noqa to an existing line, so the line count is unchanged:
mutation lines test_aiters_version_is_read_without_executing_aiter_or_torch test_the_located_name_is_top_level
atom/compass/artifacts/rules.py:23 47→47 FAILED, assert ['torch'] == [] passed
atom/sampling_params.py:4, via atom/__init__ 30→30 FAILED, assert ['torch'] == [] passed
rules.py:23 gets an aiter import that is swallowed: exec('try: import aiter / except BaseException: pass') 47→47 FAILED, assert ['aiter'] == [] passed

The last row matters. The marker file catches an import that a try block hides, and a raise-only stub would miss it.

Residual, pre-existing and not introduced here (principle 6). 2>/dev/null still turns any failure of this line into AITER=UNKNOWN. That includes a future broken import in atom.compass.artifacts. UNKNOWN is treated as a mismatch, so the run warns loudly, but the warning reads "toolchain differs" rather than naming the lookup failure. This was the same before the PR, so I am not asking for a change here.

"""From the `AITER_DIR=` assignment through the `fi` that closes its branch."""
lines = GATE_GPU.read_text().splitlines()
start = next(i for i, line in enumerate(lines) if line.startswith("AITER_DIR="))
end = lines.index("fi", start)

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Extraction ruling (principles 6 and 5): robust, and it fails loudly, but the failures do not say why. Not blocking.

It is not a line-drift guard. The extraction is anchored on content, not position. Every run was on node 18, with atom.__file__ asserted under each copy's root:

edit to gate_gpu.sh lines result
comment word: wedged → stuck (N1) 326→326 2 passed
extra comment line inserted above AITER_DIR=, so every line after it shifts (N2) 326→327 2 passed
trailing comment on the closing fi: fi # end (N3, no effect on behaviour) 326→326 test 1 FAILED: ValueError: too many values to unpack (expected 2)
AITER_DIR renamed to AITER_ROOT (M7) 326→326 both FAILED: bare StopIteration
if/fi collapsed into one AITER=$(git -C ...) line (M8) 326→325 test 1 FAILED: ValueError: too many values to unpack

The test never passes silently when the markers move or vanish. But three of these failures name nothing:

  • N3 changes no behaviour and still reddens. lines.index("fi", start) needs a bare fi. When the closing fi carries a comment, the search runs on to the next bare fi, which closes the toolchain-WARNING block. The script then prints the torch:/baseline: lines, and the unpack fails.
  • M7 reports a bare StopIteration.

Suggestion: have _resolution_lines check its own boundaries, with messages that say what went wrong. For example, assert start is not None, "gate_gpu.sh has no top-level AITER_DIR= line". Then check that the extracted block contains exactly one AITER= assignment, so a run-on is refused by name rather than by the unpack.

Scope limit, measured (M6c). Adding AITER_PY=$(python -c 'import aiter' 2>/dev/null) on its own line after the fi leaves 2 passed. The test holds the resolution lines, which is what it claims; it does not hold the whole script. A cheap whole-file text check would close that gap, for example no import aiter anywhere in gate_gpu.sh. That is optional: it is outside the named result.

@jgong5

jgong5 commented Sep 23, 2026

Copy link
Copy Markdown
Owner Author

This review was written by an agent. Before reviewing, it read the eight Design principles in atom/compass/design/README.md and atom/compass/AI_DEV_RULES.md.

Verdict: APPROVE at head 1b8d1e3cfe22b86d3649eb8940efdc4ff503335a. Nothing blocks. The two inline comments and the stale-text follow-up below do not block either.

What I checked

1. The named result reproduces (principle 8). I ran the mutations in xiaobizh_n18_cpu, each on its own copy of the merged tree, with atom.__file__ asserted under that copy's root.

Node ids:

  • A = tests/compass/test_gate_gpu_aiter_version.py::test_aiters_version_is_read_without_executing_aiter_or_torch
  • B = tests/compass/test_gate_gpu_aiter_version.py::test_the_located_name_is_top_level
variant edit lines A B
B0 branch none 326 passed passed
M1 import restored line 158 set back to import aiter, os; ... aiter.__file__ 326→326 FAILED, assert ['aiter'] == [] FAILED, text absent
M3 dotted find_spec u.find_spec("aiter.ops").origin, which imports the parent aiter 326→326 FAILED, ['aiter'] FAILED
M3b dotted find_spec, with the module_root("aiter") text kept u.find_spec("aiter.ops"); ... module_root("aiter") 326→326 FAILED, ['aiter'] passed
M4 torch on the chain rules.py:23 + ; import torch 47→47 FAILED, ['torch'] passed
M4b torch via atom/__init__ sampling_params.py:4 + ; import torch 30→30 FAILED, ['torch'] passed
M5 aiter imported and the error swallowed rules.py:23 + exec('try: import aiter except BaseException: pass') 47→47 FAILED, ['aiter'] passed
N1 null control comment word wedged → stuck 326→326 passed passed
N2 null control, lines shifted extra comment line above AITER_DIR= 326→327 passed passed

M3b isolates the behavioural test. With the text pin (B) still satisfied, a parent import is caught by A alone, through the marker file. M5 shows why the marker matters: a stub that only raises would miss an import wrapped in try. The reuse precondition is held by A, by name.

2. Reuse via atom (principles 3, 4 and 2): accepted. The details and the torch-mutation table are inline on gate_gpu.sh:158.

The deciding fact is that gate_gpu.sh:148 (compass_require_tree) already runs import atom under the same PYTHONPATH, before this line. The gate's dependency on a cheap import atom predates this PR. Sharing module_root keeps the gate and the provenance stanza on one locator.

Measured on node 18:

import time modules loaded heavy modules loaded (torch, aiter, numpy, zmq, triton)
from atom.compass.artifacts.provenance import module_root 37.7 ms 146 none
bare import atom 16.2 ms 104 none

3. The extraction markers (principle 6): robust, and loud, but the failures are unnamed. The details are inline on the test's _resolution_lines.

  • It is not a line-drift guard: N2 shifts every later line and passes.
  • A comment on the closing fi changes no behaviour, yet it reddens A with ValueError: too many values to unpack.
  • Renaming AITER_DIR reddens both tests with a bare StopIteration.

In every case it fails; it never passes silently. My suggestion is named asserts. That is not blocking.

  • Scope limit (M6c): a new import aiter on its own line after the fi gives 2 passed. That is outside the named result.

4. No design-doc references (rules). The diff adds none. I grepped both files for D<n>, T<nn>, P0., W<n>. and principle, and found no hits. ruff check, ruff format --check and black --check are clean on the new test.

5. Size. 71 lines against a 20–40 estimate, which is inside the 2x stop. Most of it is the stub setup, which is what makes the mechanism testable.

Rulings on the two surprises

GPU-tier trigger for gate_gpu.sh: no follow-up.

  • gpu_gate_triggers.txt is generated from atom/ source paths that no running CPU-tier test names. Nothing in it is maintained by hand.
  • gate_gpu.sh is the GPU tier's harness, not code under its test. Running that tier's pytest superset would exercise this line only by printing it.
  • What does cover the change is the new CPU test, which runs the gate's actual lines, plus the developer's read-only check on node 18. That check found the same v0.1.21.dev0-49-gf4e7c7509 from both paths, 0.07 s against 8.33 s.

Watch item for the next GPU wave: on node 18, gate_gpu.sh must print aiter: v0.1.21.dev0-49-gf4e7c7509 with no toolchain WARNING.

Stale README line: yes, file one follow-up. Widen it, because this PR makes more text stale than the README. Measured on node 18: bare import atom loads no torch or aiter, as above. import atom.config takes 3312 ms and loads torch, numpy, zmq and triton, but not aiter. So "on a wedged node ATOM's import hangs, because aiter shells out to rocminfo" is no longer true as stated. It appears at:

  • scripts/compass/README.md:16
  • scripts/compass/preflight.sh:17-19

This PR also falsifies text that is outside its diff, so it cannot be commented inline:

  • atom/compass/artifacts/provenance.py:31-34 says that gate_gpu.sh "imports aiter to read aiter.__file__". After this lands, that is no longer true.
  • provenance.py:31, provenance.py:300 and atom/compass/artifacts/__init__.py:17 cite gate_gpu.sh:153-159. That range now covers mostly the new comment block, and the describe call is at line 160.

This is one small documentation issue covering all five sites (principle 8). It is not a reason to widen this PR's file set.

One piece of context for the named result, not a finding. gate_gpu.sh runs preflight.sh first, and preflight's 25-second rocminfo check aborts the gate with 91 on a node that is fully wedged. So line 158 matters on a node that passes preflight and then stalls, and on every healthy node, where it saves 8.3 s.

Gate: the tree that will land

  • The tip is fork/feature/atomcompass_new = ff9617f30, unchanged; I read it again after git fetch.
  • git merge-tree --write-tree ff9617f30 1b8d1e3cf gives 2114ab788302bc26e719cf499538211618ea0ff5, rc 0. That is identical to the head's own tree.
  • Staging: git archive 2114ab788 went into a path of my own in xiaobizh_n18_cpu, with matching md5 on both ends and the .compass-commit (1b8d1e3cf) and .compass-changed (2 files) stamps. atom.__file__ was /tmp/pr271rev/merged/ATOM/atom/__init__.py. I used the tree's own scripts/compass/gate_cpu.sh, bounded by timeout -k 10 1500.

The gate printed:

line value
commit: 1b8d1e3cf (stamp)
gpu: not required (.compass-changed stamp)
pytest 5122 passed, 149 skipped, 3 xfailed in 167 s
GATE_CPU_RC 0

This matches the developer's 5122 for the branch. Their control at ff9617f30 was 5120, and the +2 are A and B, both of which appear in the junit. Nothing failed, so there was no flake to re-run. I removed the staging afterwards.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant