Skip to content

compass(memory): rename ModelTerms.declared_for_m1 to declared, and let the guard enforce milestone labels - #312

Merged
jgong5 merged 2 commits into
feature/atomcompass_newfrom
compass/issue-307-declared-name
Sep 23, 2026
Merged

jgong5 merged 2 commits into
feature/atomcompass_newfrom
compass/issue-307-declared-name

Conversation

@jgong5

@jgong5 jgong5 commented Sep 23, 2026 •

Copy link
Copy Markdown
Owner

Closes #307

Dev record

Decision: rename ModelTerms.declared_for_m1 to ModelTerms.from_declared_config. It is not deleted. Round 1 used declared. The first review showed that name collides with DeviceReadings.declared and Reading.declared, both properties that return the names of terms on a declared basis. Round 2 (40ae3faad) renamed it; the round-2 comment has the re-verified gates.

Evidence:

  • It has no production caller. At 93be90841, grep -rn declared_for_m1 over the whole tree finds the definition and one prose mention in atom/compass/memory/readings.py, and 10 call sites, all in tests/compass/: 9 in test_memory_readings.py and 1 in test_memory_compare.py. Beyond those, only the guard test in test_memory_compare.py names it, to assert the name exists. No design doc names it. At the new tip f84b4d778, git grep -l declared_for_m1 outside those three files finds nothing.
  • Deleting it removes the only way those 10 call sites get the declared formulas. Without it, each test would build three Terms by hand. The tests would then duplicate the formulas instead of exercising them. The brief says to rename in that case.
  • from_declared_config says what it does. It is an alternate constructor, and every term it returns is Basis.DECLARED, built from config geometry and stated coefficients. It leaves declared to the two existing queries.

What the guard needed, which the brief did not say. _TAGS never matched M1 or declared_for_m1. KEPT only documented why, and the one test that read KEPT asserted the kept forms did not match. So taking them out of KEPT alone would have enforced nothing. The pattern gains two forms:

  • \bM\d+\b for a milestone label in prose;
  • \w*_m\d+\b for one used as an identifier suffix.

It still does not match a bare letter-and-digit, so TP1, w1_base_bytes and the dtype widths stay clear. That check is kept as its own test. The group (\.\d+)? became non-capturing, and the assertion message now prints found. Before, findall returned the group's contents (''), so a failure could not name what it found. Now it can: see the named result.

Changes:

  • readings.py: the method is renamed. Three M1 mentions in prose now say what they meant:
    • "fills all three with declared formulas";
    • "With fake models, ...";
    • "on a real model".
  • graph_pool.py: "Neither fires at M1 -- no drafter, one data-parallel rank --" becomes "Neither fires with no drafter and one data-parallel rank".
  • test_memory_readings.py: 9 call sites renamed.
  • test_memory_compare.py:
    • 1 call site renamed;
    • KEPT and its comment deleted;
    • _TAGS widened, as above;
    • the two swept forms added to the guard's driven removed list;
    • test_the_kept_forms_are_kept_on_purpose_and_stay_readable replaced by test_the_widths_and_dtypes_that_share_the_shape_are_not_swept. It keeps only the benign-forms half, because the other two assertions tested the exemption.

Lines: 27 insertions and 41 deletions over 4 files.

  • Production: 7+/7- (readings.py 5/5, graph_pool.py 2/2).
  • Tests: 20+/34-.

This is within the 20-40 line estimate.

Remaining M1 in atom/compass/memory/: none. grep -nE "\bM[0-9]+\b|_m[0-9]+\b" atom/compass/memory/*.py is empty at the head. The guard now enforces that over the package glob.

Gate 1: ATOM's CPU suite, unmodified, as a delta (round 1, head 306b9b7e8; round 2 is in the PR thread)

This ran on node 18 xiaobizh_n18_cpu, using each tree's own scripts/compass/gate_cpu.sh. Trees were staged by git archive plus docker exec -i tar -x into /tmp/i307gates/<label>/ATOM, with .compass-commit and .compass-changed stamps. The tarball md5 matched on both ends. atom.__file__ resolved under each staged root, for example /tmp/i307gates/head/ATOM/atom/__init__.py.

tree commit result GATE_CPU_RC
control (branch point) 93be90841 5209 passed, 155 skipped, 3 xfailed 0
head 306b9b7e8 5209 passed, 155 skipped, 3 xfailed 0
control (new tip) f84b4d778 5214 passed, 155 skipped, 3 xfailed 0
merged (new tip + head) 09567002d, tree cc28111b4 5214 passed, 155 skipped, 3 xfailed 0

Node-id delta, from junit XML: 5367 ids on each side against 93be90841, and 5372 ids on each side against f84b4d778. The delta is identical on both:

  • removed: tests/compass/test_memory_compare.py::test_the_kept_forms_are_kept_on_purpose_and_stay_readable (passed)
  • added: tests/compass/test_memory_compare.py::test_the_widths_and_dtypes_that_share_the_shape_are_not_swept (passed)
  • no state changes. Net 0.

git merge-tree --write-tree:

ruff and black on the four changed files, at control and head: RUFF_RC=0 and BLACK_RC=0 on both.

Gate 2: CPU-only tests

tests/compass/test_memory_compare.py and tests/compass/test_memory_readings.py pass 109 of 109 at the head. The renamed method is exercised by the 10 existing call sites, none of which this PR adds. The guard it widens runs over the package glob, which this PR does not add.

Gate 3: named result

The mutation trees are commit objects made with git commit-tree on the head. Each changes one file with its line count preserved. Each ran pytest tests/compass/test_memory_compare.py tests/compass/test_memory_readings.py on node 18.

tree edit result
head 306b9b7e8 none 109 passed
mutA c7b71bbdd readings.py:42: `declared` fills becomes `declared_for_m1` fills 1 failed, 108 passed: test_no_module_in_the_package_carries_a_design_reference[readings.py], with AssertionError: readings.py carries design references: ['declared_for_m1']
mutB 074a1053a readings.py:122: "With fake models," becomes "For M1, with fake models," 1 failed, 108 passed: ...[readings.py], with ['M1']
mutC 366b9d7ca graph_pool.py reverted to 93be90841 via git show 1 failed, 108 passed: ...[graph_pool.py], with ['M1']
mutD 80d9a2a23 readings.py reverted to 93be90841 via git show, which is the full pre-fix module 32 failed, 77 passed: ...[readings.py] with ['declared_for_m1', 'M1', 'declared_for_m1', 'M1', 'M1'], plus 31 AttributeErrors from callers of declared
null 215bbc73c readings.py:131: "on a real model" becomes "on any real model" 109 passed
tip 93be90841 none; it carries all of the above forms (declared_for_m1 twice in readings.py, M1 in graph_pool.py) 109 passed; test_no_module_in_the_package_carries_a_design_reference[readings.py] and [graph_pool.py] both pass. The forms are exempt there.

On the tip, the lines mutA, mutB and mutC reinstate are the tip's own lines, so "the same edit on the tip" is the tip itself.

AST comparison, tip against head, with declared_for_m1 normalised to declared in FunctionDef.name and Attribute.attr:

  • tests/compass/test_memory_readings.py: identical.
  • atom/compass/memory/readings.py and graph_pool.py: identical once docstrings are masked. The only non-name difference is the docstring prose that removes M1.
  • tests/compass/test_memory_compare.py: differs only in the intended guard changes:
    • _TAGS changed;
    • KEPT removed;
    • test_no_module_in_the_package_carries_a_design_reference changed (the message);
    • test_the_guard_catches_the_forms_that_were_actually_removed changed (two strings added);
    • test_the_kept_forms_... removed;
    • test_the_widths_and_dtypes_... added.

Gate 4

The coordinator dispatches the independent reviewer.

Left undone

Nothing from the brief.

🤖 Generated with Claude Code

The classmethod carried a milestone label in a public name. Its only
callers are in tests/compass/, and it is the only path to the declared
formulas those tests build ModelTerms from, so it is renamed, not deleted:
every term it returns is Basis.DECLARED, from geometry and stated
coefficients.

The memory package's design-reference guard now matches a milestone label,
as `M1` in prose and as an `_m1` identifier suffix, and the KEPT exemption
is gone. The pattern no longer captures a group, so a failure prints the
forms it found. The remaining M1 mentions in readings.py and graph_pool.py
say what they meant instead.

Closes #307

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Comment thread atom/compass/memory/readings.py Outdated

@classmethod
def declared_for_m1(
def declared(

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocking (principle 4): declared is already a query in this package, and this makes it a constructor too.

At 306b9b7e8, three classes in atom/compass/memory/ have a member called declared. Measured on node 18 with inspect.getattr_static against the merged tree:

member kind returns
readings.DeviceReadings.declared (this file, :219) property tuple[str, ...], the names of terms still on a declared basis
terms.Reading.declared (terms.py:159) property tuple[str, ...], the same question for one reading
readings.ModelTerms.declared (here) classmethod a new ModelTerms

So x.declared asks "which terms are declared?" on two classes and means "build me declared terms" on the third. On a ModelTerms instance, .declared is a bound method: tuple(model_terms.declared) raises TypeError: 'method' object is not iterable, where the same expression on either sibling returns names.

The two meanings already sit side by side in the same files:

  • readings.py:224 iterates reading.declared, 113 lines below this definition.
  • test_memory_readings.py calls ModelTerms.declared(...) nine times, and at :324 asserts on readings.peak_torch.declared.

The old name did not collide. This PR is a naming change, so the name is the deliverable.

Fix: use an alternate-constructor name that no query in the package uses. The brief's own example, from_declared_config, works, and so does declared_from_config. That is 11 mechanical sites: this definition, the prose at :42, nine in test_memory_readings.py and one in test_memory_compare.py. Neither name trips _TAGS.

Deleting the method stays wrong, for the reason the dev record gives: 10 test sites would each rebuild three Terms by hand. Keeping it with no production caller is justified. Only the name needs to change.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 40ae3faad. The method is renamed to ModelTerms.from_declared_config, the brief's own example, at all 11 sites: the definition, the module docstring at :42, nine calls in test_memory_readings.py and one in test_memory_compare.py.

DeviceReadings.declared (:219) and Reading.declared (terms.py:159) are untouched, so declared is now only ever the query.

AST check, tip 2565b5f57 against head 40ae3faad, with declared_for_m1 normalised to from_declared_config, run on node 18:

  • test_memory_readings.py is identical.
  • readings.py and graph_pool.py are identical once docstrings are masked.

Re-run of mutA with the new name's line (26f1d89a3): `from_declared_config` at :42 becomes `declared_for_m1`. test_no_module_in_the_package_carries_a_design_reference[readings.py] fails, with readings.py carries design references: ['declared_for_m1']. "def declared_for_m1(" stays in the driven list.

#: prints the forms it found.
_TAGS = re.compile(
r"\b[DTW]\d+(\.\d+)?\b|\bP\d+\.\d+\b|\b[A-Z]{2,4}-\d+\b"
r"\b[DTW]\d+(?:\.\d+)?\b|\bP\d+\.\d+\b|\b[A-Z]{2,4}-\d+\b|\bM\d+\b|\w*_m\d+\b"

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Non-blocking (principle 6): near variants of the removed form escape. They pass silently, not loudly.

The named result holds for the exact forms: declared_for_m1 and M1 both fail the guard by name at the head (details in the review comment). One step away, they pass. Measured two ways at 306b9b7e8:

  1. A line-count-preserving edit to readings.py, tree 1bd8572e3. Three prose lines became `declared_for_m1_terms` fills, At m1, with fake models, and past M1a or declared_for_M1. Result: 109 passed. test_no_module_in_the_package_carries_a_design_reference[readings.py] is green.
  2. Probing this compiled pattern directly (findall, node 18):
form caught?
declared_for_m1, M1, M1's, M1.2, M01 yes
declared_for_m1_terms, _m1_ no. \w*_m\d+\b needs a non-word character after the digits, and _ is a word character.
declared_for_M1, declaredForM1, m1_declared no. The suffix form is lowercase-only and suffix-only, and \bM has no boundary after _.
m1, at m1, M1a, Milestone1 no

Nothing else in the suite notices either. git grep for any other design-reference guard over this package finds only this file. The whole CPU gate on tree 1bd8572e3 gave 5209 passed, 155 skipped, 3 xfailed, GATE_CPU_RC=0: the same count as the head on its own base.

Ask: a regex cannot close all of these without taking more of ATOM's own names with it (see the next comment). So I am not asking for a wider pattern. Instead, state the escapes beside the benign list, the way the deleted KEPT comment stated its absence. The test file's own standard is that "an absence nobody explained is the same defect one level up."

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Recorded in 40ae3faad; the pattern is not widened. test_what_the_pattern_matches_beside_its_targets_is_on_record asserts that declared_for_m1_terms, At m1, M1a and declared_for_M1 are not matched. Its docstring says why: a pattern wide enough to take them would take more of ATOM's names.

If someone later widens the pattern to catch one of them, this assertion fails. That forces them to update the record, instead of the record going stale.

# share its shape, which is why it is not a bare letter-and-digit.
def test_the_widths_and_dtypes_that_share_the_shape_are_not_swept():
"""The pattern is not a bare letter-and-digit, and this is why."""
for benign in ("TP1", "w1_base_bytes", "fp8", "int8", "bf16"):

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Non-blocking (principle 8): this list claims to cover what shares the pattern's shape, but ATOM's own model names are not in it, and both new alternatives match them.

Measured over the merged tree (d45ecb5b5, tree 149965dee), with this file's compiled _TAGS:

  • In the guard's actual scope (atom/compass/memory/*.py, 6 files, no subdirectories): 0 hits. So there is nothing to fix today, and any future hit fails loudly.
  • In atom/ outside compass/:
    • \bM\d+\b hits 140 times, all M3, across 28 files. Almost all of it is ATOM's MiniMax-M3, for example atom/entrypoints/openai/chat_encoders.py:162.
    • \w*_m\d+\b hits 50 times: minimax_m3 45 times and minimax_m2 5 times. These are ATOM's module and model-type names, for example atom/config.py:766 and atom/model_engine/model_runner.py:124.
  • Probes:
    • Matched: MiniMax-M1 gives ['M1'], MiniMax-M2.5 gives ['M2'], M128 gives ['M128'], tile_m128, block_m16, seq_len_m1, n_m1, fp8_m3 and int4_m2.
    • Clear: float8_e4m3fn, fp8_e5m2, BLOCK_M128 and MI308X.
  • A line-count-preserving edit to readings.py (tree e7a5729db) put MiniMax-M1, an M128 tile and seq_len_m1 = seq_len - 1 into prose. It fails test_no_module_in_the_package_carries_a_design_reference[readings.py] with ['M1', 'M128', 'seq_len_m1'].

This package sizes model memory, so a sentence or a key naming one of ATOM's supported models is plausible here. The next developer to write one will be refused, and the obvious reaction is to add an exemption back.

Ask: say it where the next reader looks. Either:

  • add MiniMax-M3 and minimax_m3 to this test as known collisions, asserting that they are matched, so the trade-off is on record; or
  • narrow the prose form with a lookbehind ((?<![-\w])M\d+\b leaves MiniMax-M3 alone), then add the model name to the benign loop.

Either way the docstring's claim then matches its assertions.

I measured the lookbehind on node 18:

  • MiniMax-M3 and MiniMax-M1 give [].
  • at M1, (M1), M1-era, For M1, with and past M1. give ['M1'].
  • Its cost: a hyphenated the-M1 would escape.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Recorded in 40ae3faad, using your first option. The same test now asserts that MiniMax-M3, minimax_m3, M128, tile_m128, seq_len_m1 and fp8_m3 are matched. Its docstring says that none occurs in this package, and that a hit is fixed by rewording the line, not by adding an exemption. I left the lookbehind out, so _TAGS is unchanged this round.

The record also holds the pattern. On mutE, which puts back the tip's _TAGS, this test fails with AssertionError: MiniMax-M3, alongside test_the_guard_catches_the_forms_that_were_actually_removed.

@jgong5

jgong5 commented Sep 23, 2026

Copy link
Copy Markdown
Owner Author

Review by an AI agent: the independent reviewer for #307, first cycle.

Verdict: REQUEST_CHANGES at 306b9b7e810b4441f588d489f1c669bdfe7dec88.

One finding blocks: the new name declared collides with two existing declared members in the same package (principle 4, inline on readings.py:111). The fix is mechanical, 11 sites.

Everything else checks out, measured on node 18:

  • the named result;
  • the AST claim;
  • the rename-over-delete decision;
  • the guard widening;
  • the merged tree's gate.

The two guard comments are non-blocking.

I read the eight design principles (atom/compass/design/README.md) and atom/compass/AI_DEV_RULES.md at fbfc37ef8 before the diff.

1. Rename vs delete (principle 3): rename is right; the name is not

  • Keeping a test-only method is justified. At the merged tree, git grep ModelTerms.declared -- atom finds no production caller. The 10 test sites each need all three declared Terms, so deleting the method would copy the formulas into 10 places instead of exercising them once. The brief anticipated exactly this case.
  • declared is not a clear name here. Measured with inspect.getattr_static on node 18:
    • DeviceReadings.declared is a property in the same module, and Reading.declared is a property in terms.py. Both return the names of terms still on a declared basis.
    • ModelTerms.declared is a classmethod constructor. On an instance it is a bound method, so tuple(model_terms.declared) raises TypeError: 'method' object is not iterable.
    • test_memory_readings.py uses both meanings: nine ModelTerms.declared(...) calls, and the assertion at :324 on readings.peak_torch.declared.
    • Blocking. Use from_declared_config, the brief's own example, or declared_from_config.

2. The new guard patterns (principle 6): sound for the named forms; both edges measured

  • In scope, the guard is clean. atom/compass/memory/*.py is 6 files with no subdirectories, and it has 0 hits at the merged tree.
  • False positives exist just outside scope. In atom/ outside compass/:
    • \bM\d+\b hits 140 times, all M3, mostly ATOM's MiniMax-M3.
    • \w*_m\d+\b hits 50 times: minimax_m3 45 times and minimax_m2 5 times.
    • Probes: MiniMax-M1, M128, tile_m128, seq_len_m1 and fp8_m3 all match. float8_e4m3fn, fp8_e5m2, BLOCK_M128 and MI308X do not.
    • A prose edit to readings.py (tree e7a5729db) fails the guard with ['M1', 'M128', 'seq_len_m1'].
    • This fails loudly, so it is not a principle 6 breach. But the new benign-forms test does not state the collision. Non-blocking, inline on test_memory_compare.py:1108.
  • False negatives. Tree 1bd8572e3 has declared_for_m1_terms, At m1, M1a and declared_for_M1 in prose.
    • Result: 109 of 109 green on the two files, and the whole gate green at 5209 / 155 / 3, GATE_CPU_RC=0. Nothing notices.
    • \w*_m\d+\b misses an identifier that continues past the digits (_m1_), and misses any uppercase _M1.
    • Non-blocking. Record these escapes in the test rather than widen the pattern further. Inline on test_memory_compare.py:1043.

3. Named result: reproduced

tree edit result failing node id, assertion
head 306b9b7e8 none 109 passed none
mutA 88bb16a6f readings.py:42: `declared` fills becomes `declared_for_m1` fills 1 failed, 108 passed test_memory_compare.py::test_no_module_in_the_package_carries_a_design_reference[readings.py]: readings.py carries design references: ['declared_for_m1']
mutB 5fe202468 readings.py:122: "With fake models," becomes "For M1, with fake models," 1 failed, 108 passed same id: ['M1']
mutC 9f8602ab1 graph_pool.py taken whole from the tip 1 failed, 108 passed ...[graph_pool.py]: ['M1']
null 365f4eda1 readings.py:131: "on a real model" becomes "on any real model" 109 passed none
mutE (mine) d3598d53f test_memory_compare.py:1043: _TAGS's first line reverted to the tip's pattern, all else kept 1 failed, 108 passed test_memory_compare.py::test_the_guard_catches_the_forms_that_were_actually_removed: AssertionError: For M1, with fake models, a declared formula suffices
tip fbfc37ef8 none (carries declared_for_m1 twice, M1 three times in readings.py, and M1 once in graph_pool.py) 109 passed exempt
tip null 79ac3091c readings.py:131: "past M1" becomes "past M1" 109 passed exempt

mutE shows that the widening itself is pinned. A developer who reverts the pattern and keeps the prose clean is caught by the driven removed list, which names the exact string. Without that pin, the head would pass with a guard that enforces nothing new.

4. AST claim: confirmed

The check parsed tip and head on node 18, normalised declared_for_m1 to declared in FunctionDef.name, Attribute.attr and Name.id, and masked docstrings.

  • test_memory_readings.py is identical even without masking.
  • readings.py and graph_pool.py are identical with docstrings masked, so every M1 edit is docstring prose.
  • test_memory_compare.py differs only in the intended places:
    • _TAGS changed;
    • KEPT removed;
    • test_the_kept_forms_... removed;
    • test_the_widths_and_dtypes_... added;
    • two bodies changed: the assertion message, and the two driven strings.
  • The --collect-only node-id delta on the two files is exactly that 1-for-1 swap, at 109 ids each side.

5. ponytail-review

Over the 27+/41- diff:

  • declared is kept for 10 real callers.
  • The new benign-forms test is a 3-line self-check, which is the minimum.
  • The comment added above _TAGS replaces a longer deleted KEPT comment.

Nothing is left to delete.

Lean already. Ship.

Gate: the tree that will land

  • Merge tree. The tip moved to fbfc37ef8 (compass(tests): hold four refusals the artifact store states and nothing reached #306). git merge-tree --write-tree fbfc37ef8 306b9b7e8 gives tree 149965dee1e466d965ad82b02c22412cf7380a9f with no conflict. That is neither 306b9b7e8^{tree} (64d2ec9c4) nor the developer's cc28111b4, so the developer's merged-tree gate no longer covers what would land.
  • Stamp commit. git commit-tree 149965dee -p fbfc37ef8 -p 306b9b7e8 gives d45ecb5b5469e44825b5133bb87240085c75014c.
  • Staging. The tree was staged with .compass-commit and .compass-changed stamps; .compass-changed lists the four PR files. The gate printed commit: d45ecb5b5 (stamp), atom: /tmp/r312/merged/ATOM/atom/__init__.py, and gpu: not required.
  • Run. It used the tree's own scripts/compass/gate_cpu.sh under timeout -k 10 3000.
  • Result: 5222 passed, 155 skipped, 3 xfailed. GATE_CPU_RC=0 PASSED. The junit has 5380 ids. It includes test_the_widths_and_dtypes_that_share_the_shape_are_not_swept and no test_the_kept_forms_....
  • Against the developer's merged count at f84b4d778 (5214), the difference is +8. That is exactly the 8 unparametrised def test_ functions compass(tests): hold four refusals the artifact store states and nothing reached #306 adds to tests/compass/test_artifact_invalidation.py, so this PR's delta stays net 0.
  • No failures, so no timing-flake re-runs were needed.

I also ran the gate once on the false-negative tree 1bd8572e3, reported under item 2.

What the next cycle should check

  1. The rename to a non-colliding name, across all 11 sites. Re-run mutA with the new name's line, and keep "def declared_for_m1(" in the driven list.
  2. The two non-blocking guard notes, if taken.
  3. If the tip moves again, recompute merge-tree and gate the result once.

…d the guard's reach

`declared` collided with two queries in the same package:
DeviceReadings.declared and Reading.declared are properties that return
the names of terms still on a declared basis, so `x.declared` meant a
query on two classes and a constructor on the third. The constructor is
now `ModelTerms.from_declared_config`. The definition, the module
docstring and all ten test call sites are updated.

The benign-forms test now records the pattern's reach both ways. Some of
ATOM's own names match it (MiniMax-M3, minimax_m3, M128, tile_m128,
seq_len_m1, fp8_m3), and none occurs in this package. Some forms one step
from the removed ones escape it (declared_for_m1_terms, At m1, M1a,
declared_for_M1). The pattern itself is unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@jgong5

jgong5 commented Sep 23, 2026

Copy link
Copy Markdown
Owner Author

Developer round 2, head 40ae3faade7b93d28727e9d0ed7ba5a1ba893d0b. This adds one commit on 306b9b7e8, with no force-push. There are no blocking issues outstanding from my side.

Findings addressed:

  1. readings.py:111, blocking. The method is renamed to ModelTerms.from_declared_config at all 11 sites. DeviceReadings.declared and Reading.declared are untouched.

  2. test_memory_compare.py:1043 and :1108, non-blocking. The benign-forms test becomes test_what_the_pattern_matches_beside_its_targets_is_on_record. It asserts three things:

    • benign forms are not matched: TP1, w1_base_bytes, fp8, int8, bf16;
    • known collisions are matched: MiniMax-M3, minimax_m3, M128, tile_m128, seq_len_m1, fp8_m3;
    • known escapes are not matched: declared_for_m1_terms, At m1, M1a, declared_for_M1.

    _TAGS itself is unchanged.

Lines against the branch point 93be90841: 43 insertions and 39 deletions.

  • Production: 7+/7- (readings.py 5/5, graph_pool.py 2/2).
  • Tests: 36+/32-.

Named result, re-verified on node 18 (xiaobizh_n18_cpu):

  • Each tree is a git commit-tree object on 40ae3faad, with one file changed and its line count preserved.
  • Each was staged by git archive plus docker exec -i tar -x into /tmp/i307gates/<label>/ATOM. The md5 matched on both ends.
  • atom.__file__ resolved under each staged root.
  • Each ran pytest tests/compass/test_memory_compare.py tests/compass/test_memory_readings.py.
tree edit result
head 40ae3faad none 109 passed
mutA 26f1d89a3 readings.py:42: `from_declared_config` becomes `declared_for_m1` 1 failed: test_memory_compare.py::test_no_module_in_the_package_carries_a_design_reference[readings.py], with ['declared_for_m1']
mutB 4f9f7af6f readings.py:122: "With fake models," becomes "For M1, with fake models," 1 failed: same id, with ['M1']
mutC 532a76154 graph_pool.py taken whole from the tip 1 failed: ...[graph_pool.py], with ['M1']
mutE 58494a277 _TAGS first line reverted to the tip's pattern 2 failed: test_the_guard_catches_the_forms_that_were_actually_removed (AssertionError: For M1, with fake models, a declared formula suffices) and test_what_the_pattern_matches_beside_its_targets_is_on_record (AssertionError: MiniMax-M3)
null 9e40a1292 readings.py:131: "on a real model" becomes "on any real model" 109 passed
tip 2565b5f57 none; it carries declared_for_m1 and M1 109 passed, so the forms are exempt there

The four PR files at the tip 2565b5f57 are byte-identical to the branch point 93be90841.

AST check, tip against head, with declared_for_m1 normalised to from_declared_config:

  • test_memory_readings.py is identical.
  • readings.py and graph_pool.py are identical with docstrings masked.
  • test_memory_compare.py differs only in:
    • _TAGS;
    • KEPT removed;
    • the guard's assertion message;
    • the driven list;
    • test_the_kept_forms_... removed;
    • test_what_the_pattern_matches_... added.

Gate: the tree that will land.

  • git merge-tree --write-tree 2565b5f57 40ae3faad gives 4c05f3420988c2f70d8162537c59454cb8bdff1f, with no conflict. The stamp commit is 445075711 (parents 2565b5f57 and 40ae3faad), and the head tree is 18c4f8213.
  • The gate is the tree's own gate_cpu.sh, with stamps, under timeout -k 10 1500:
tree result GATE_CPU_RC
control 2565b5f57 5225 passed, 155 skipped, 3 xfailed 0
merged 445075711 5225 passed, 155 skipped, 3 xfailed 0
  • Node-id delta: 5383 ids on each side.
    • removed: tests/compass/test_memory_compare.py::test_the_kept_forms_are_kept_on_purpose_and_stay_readable
    • added: tests/compass/test_memory_compare.py::test_what_the_pattern_matches_beside_its_targets_is_on_record
    • Both passed; no other changes. Net 0.
  • Lint: ruff check and black --check on the four files give RUFF_RC=0 and BLACK_RC=0 at both control and head.

What the next review should check: the delta 306b9b7e8..40ae3faad.

# share its shape, which is why it is not a bare letter-and-digit.
for benign in ("TP1", "w1_base_bytes", "fp8", "int8", "bf16"):
assert not _TAGS.search(benign), benign
for collision in (

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Non-blocking (principle 8): four of these six strings are not ATOM's names. They occur nowhere in the repo except this test.

The docstring says the pattern "does match some of ATOM's own names". Round 2's comment and the cycle-2 brief go further and call all six "ATOM's colliding names". I counted each one at 40ae3faad with git grep -wF -c over the whole tree:

string whole repo atom/
MiniMax-M3 258 120
minimax_m3 47 44
M128 1 0
tile_m128 1 0
seq_len_m1 1 0
fp8_m3 1 0

For the last four, the single hit is this line. They are shapes that cycle 1's reviewer invented as probes, not names ATOM carries.

On node 18, the head's \bM\d+\b|\w*_m\d+\b finds exactly three distinct forms across atom/**/*.py (489 files): M3 140 times, minimax_m3 45 times and minimax_m2 5 times.

The asserts are correct either way; only the record's wording is off. One comment line would fix it, for example: "The first two are in ATOM (MiniMax-M3 120 times under atom/); the rest are shapes it could plausibly grow." Otherwise the next reader will believe the four probes were measured in ATOM.

@jgong5

jgong5 commented Sep 23, 2026

Copy link
Copy Markdown
Owner Author

Review by an AI agent: the independent reviewer for #307, second cycle.

Verdict: APPROVE at 40ae3faade7b93d28727e9d0ed7ba5a1ba893d0b.

Nothing blocks. Cycle 1's blocking finding is fixed. There is one non-blocking note, posted inline on test_memory_compare.py:1117: four of the six "collisions" are not ATOM names (principle 8).

Scope. This is a delta review of 306b9b7e8..40ae3faad. When I posted, 40ae3faad was the branch head, checked with git ls-remote. I read the eight design principles in atom/compass/design/README.md and atom/compass/AI_DEV_RULES.md (the current revision, from #137) at the head, before reading the diff.

1. The rename (principle 4): no collision left, and the name reads right

At 40ae3faad, a git grep over the whole repo, design docs included, finds:

  • from_declared_config: 12 hits. These are the def at readings.py:111, the module docstring at :42, test_memory_compare.py:453, and the nine test_memory_readings.py calls.
  • declared_for_m1: two hits, both deliberate driven strings in the guard tests (test_memory_compare.py:1090 and :1126).
  • ModelTerms.declared and .declared(: none. No .md file names ModelTerms.
  • Other members spelled exactly declared: only the two tuple-returning properties cycle 1 named, DeviceReadings.declared (readings.py:219) and Reading.declared (terms.py:159). The remaining declared* members are distinct names (declared_terms, declared_bytes, declared_fraction, declared_line in budget.py), and the other bare declared hits are local variables. So .declared now has one meaning in memory/: "the names still on a declared basis". The assertion at test_memory_readings.py:324 now reads unambiguously.
  • The name fits the package. from_* is how the package names its alternate constructors: from_hf_config (backends/geometry.py), TransferModel.from_spec, from_readings (memory/compare.py), and several from_json and from_mapping. from_declared_config says what goes in (a config) and what basis comes out (declared).

2. The record test (principle 8): an honest pin of the pattern's current reach

Ruling: accepted. It pins a limitation, and it does not block a future fix. A future fix that changes the pattern's reach does fail this test, by design. I measured both directions with line-count-preserving edits to _TAGS at :1043:

tree edit result
mutF 86ec2194f (narrowing "fix") \bM\d+\b|\w*_m\d+\b becomes (?<![-\w])M\d\b|\w*_m\d\b 1 failed, 108 passed: tests/compass/test_memory_compare.py::test_what_the_pattern_matches_beside_its_targets_is_on_record, AssertionError: MiniMax-M3. The package guard and the removed-forms test stay green.
mutG 6a2cd251c (widening "fix") the same span becomes \b[Mm]\d+|_[Mm]\d+ 1 failed, 108 passed: same node id, AssertionError: declared_for_m1_terms. The package guard stays green, with 0 hits in memory/.

Why I accept it:

  • The docstring makes a claim about reach in both directions, and these asserts are its measurement. Without them, a narrowed or widened pattern would leave the docstring stating something false, and nothing would notice. That is the drift principle 8 is meant to stop.
  • A fix still costs little. The failure names one test and the exact string. The fixer moves that string from one tuple to the other in the same commit. That is the record being updated, not the fix being blocked.
  • It is the only test that fails. In both mutants the package guard and the removed-forms test stay green, so the record does not hide a regression of the guard itself.

The reason the docstring gives for not widening, measured. On node 18 I ran the head pattern and the mutG widening over the head tree:

scope head pattern mutG widening
memory/*.py (6 files) 0 hits 0 hits
atom/compass/**/*.py (52 files) 0 hits 0 hits
atom/**/*.py (489 files) 190 hits, 3 distinct 373 hits, 4 distinct

So "would take more of ATOM's names with it" is true (+183 hits, mostly _m3). Inside the guard's own scope, though, the widening costs nothing today. Either choice is defensible, and cycle 1 asked for the record rather than the widening. I don't ask for a change.

Non-blocking, inline on :1117. M128, tile_m128, seq_len_m1 and fp8_m3 occur nowhere in the repo except this test. Only MiniMax-M3 (120 lines under atom/) and minimax_m3 (44) are ATOM's names. One comment line would make the record accurate.

3. Named result: reproduced on node 18

Method:

  • Every tree is a git commit-tree object on 40ae3faad, changes one file, and keeps its line count.
  • Each was staged by git archive (with the .compass-commit and .compass-changed stamps) into /tmp/r312c2/<label>/ATOM inside xiaobizh_n18_cpu, via docker exec -i. The tarball md5 matched on both ends for all 11 tarballs.
  • atom.__file__ resolved under each staged root, for example /tmp/r312c2/mutA/ATOM/atom/__init__.py.
  • Each ran pytest tests/compass/test_memory_compare.py tests/compass/test_memory_readings.py under timeout -k 10 900.
tree edit lines result failing node id, assertion
head 40ae3faad none 109 passed none
tip 2565b5f57 none. The four PR files are byte-identical to the branch point 93be90841, so the tip carries declared_for_m1 and M1 109 passed none (exempt)
null 97eaffca1 readings.py:131: "on a real model" becomes "on any real model" 366/366 109 passed none
mutA e9603e7ec readings.py:42: `from_declared_config` becomes `declared_for_m1` 366/366 1 failed, 108 passed tests/compass/test_memory_compare.py::test_no_module_in_the_package_carries_a_design_reference[readings.py]: readings.py carries design references: ['declared_for_m1']
mutA2 2e42dc257 readings.py:111: def from_declared_config( becomes def declared_for_m1(, callers left alone 366/366 32 failed, 77 passed the same guard id with ['declared_for_m1'], plus 31 AttributeError: ... no attribute 'from_declared_config' at the call sites
mutB 6497f9821 readings.py:122: "With fake models," becomes "For M1, with fake models," 366/366 1 failed, 108 passed the same guard id: ['M1']
mutE cf273f7d9 test_memory_compare.py:1043: _TAGS's first line is replaced by the tip's :1039, byte for byte 1127/1127 2 failed, 107 passed ::test_the_guard_catches_the_forms_that_were_actually_removed: AssertionError: For M1, with fake models, a declared formula suffices; and ::test_what_the_pattern_matches_beside_its_targets_is_on_record: AssertionError: MiniMax-M3

Each red names the test that holds the property, with the exact form found. None of them is a line-drift guard: the null control shifts nothing and stays green, and mutA and mutB fail on the reinstated text itself. This matches the developer's round-2 table row for row.

4. AST claim: confirmed

I parsed both sides on node 18. The only normalisation is the ModelTerms constructor's name: the method inside class ModelTerms, and ModelTerms.<name> attribute access, are renamed to from_declared_config. Other .declared attributes are left alone.

Tip 2565b5f57 against head (declared_for_m1 normalised, 11 renames on the tip side):

  • test_memory_readings.py: identical even without masking.
  • readings.py and graph_pool.py: identical once docstrings are masked, so every remaining edit is docstring prose.
  • test_memory_compare.py: the only differences are _TAGS, KEPT removed, test_the_kept_forms_... removed, the new record test added, and the bodies of the guard and removed-forms tests changed.

Delta 306b9b7e8 against head (declared normalised on ModelTerms only):

  • graph_pool.py and test_memory_readings.py: identical.
  • readings.py: identical with docstrings masked (the :42 wording).
  • test_memory_compare.py: the only difference is test_the_widths_and_dtypes_that_share_the_shape_are_not_swept giving way to test_what_the_pattern_matches_beside_its_targets_is_on_record.

5. ponytail-review

This covers the 33+/15- delta. The rename adds no lines. The record test is three assert-based loops, the ponytail minimum. The 8-line collision tuple looks like a shrink:, but it is not one: I checked on node 18 that black 26.5.1 re-explodes it to exactly this shape without the trailing comma, and black --check gives BLACK_RC=0 on the head file.

Lean already. Ship.

Gate: the tree that will land

  • Tip. Re-read just before the gate with git fetch and git ls-remote: feature/atomcompass_new is 2565b5f57 (compass(gates): join the README's exclude-list counts to the list #310), unchanged since round 2.
  • Merge tree. git merge-tree --write-tree 2565b5f57 40ae3faad gives 4c05f3420988c2f70d8162537c59454cb8bdff1f, with rc 0 and no conflict. This matches the developer's tree. It is not 40ae3faad^{tree} (18c4f8213), so the combined tree is what gets gated.
  • Stamp commit. git commit-tree 4c05f3420 -p 2565b5f57 -p 40ae3faad gives d910a176adde1b485b467f489ebe3d6e38a13e70. .compass-changed lists the four PR files.
  • Staging and run. The tree was staged at /tmp/r312c2/merged/ATOM and gated once with its own scripts/compass/gate_cpu.sh under timeout -k 10 3000, unpiped. The gate printed commit: d910a176a (stamp), atom: /tmp/r312c2/merged/ATOM/atom/__init__.py and gpu: not required.
  • Result: 5225 passed, 155 skipped, 3 xfailed. GATE_CPU_RC=0 PASSED, in 186 s.
    • The junit has 5383 ids, with 0 failures and 0 errors. It contains test_what_the_pattern_matches_beside_its_targets_is_on_record, and neither test_the_kept_forms_... nor test_the_widths_and_dtypes_....
    • The counts are identical to the developer's control and merged runs on the same tip (5225/155/3, 5383 ids). Net 0.
    • Nothing failed, so no timing-flake re-runs were needed.

What the next cycle should check

Nothing is outstanding. If the tip moves before landing, recompute merge-tree and gate the result once.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant