Skip to content

compass(backends): a step price that reads the batch shape, and the schedule that moves with it (M1-2) - #81

Draft
jgong5 wants to merge 4 commits into
feature/atomcompass_newfrom
compass/m1-2-cost-stub
Draft

jgong5 wants to merge 4 commits into
feature/atomcompass_newfrom
compass/m1-2-cost-stub

Conversation

@jgong5

@jgong5 jgong5 commented Sep 21, 2026 •

Copy link
Copy Markdown
Owner

Part of #70; closes nothing yet. Blocked on #89 (effort overrun): need human applied (comment). Draft; do not merge or undraft.
Head 58cb9dcce. APPROVE covers 0fc02f458 (review); the base-update merge's resolutions were reviewed sound (review). No blocking issues in the code.

A cost backend, atom/compass/backends/shape.py, whose price is a declared linear form over shapes the scheduled batch already carries, so the schedule moves with the batch shape. Constant pricing stays as a coefficient set, not the default.

prefill:  a  + b .tokens   + c .Sum N_Q^2 + d .Sum N_Q.(N_KV - N_Q)
decode:   a' + b'.requests + c'.Sum N_KV  + e'.(rung.max N_KV - Sum N_KV)
What the backend is

shape.py, four names on the package's exports, and one new assertion in the seam's own test file. The form above, plus one term per collective the widths admit. A mixed step is priced as both groups and pays both intercepts. Constant pricing is Coefficients.constant(...), every shape coefficient zero: no branch in estimate, and the zeros stay visible in the breakdown rather than implied by a flag.

Nothing here is an accuracy claim, and the output says so. Every term's provenance reads 5e-07 s x 128; declared, not measured -- a plumbing figure, not an accuracy claim, which is what StepCost.rows() hands an artifact; describe() repeats it. A collective's term adds a collective these widths admit, not one observed to run.

Dev record

Dev record, in full

Ungating. compute_detailed_aggregates is already called on every step that has sequences (EngineCore._process_engine_step_inner); the gate is inside the method and reads two things. ATOM_ENABLE_DETAILED_ANNOTATION is a real env flag (atom/utils/envs.py), cached into Scheduler._detailed_annotation_enabled. profile_active is a public attribute, False from Scheduler.__init__, flipped only by _handle_start_profile/_handle_stop_profile in atom/model_engine/engine_utility.py; it has no env flag, so ungating means setting it from outside. _handle_start_profile calls runner_mgr.call_func("start_profiler") before it flips the flag, and ModelRunner.start_profiler starts a real torch.profiler and installs a trace-export callback whenever profiler_dir is set. So assign the attribute; the RPC drags a profiler and its trace files into a run with no device. Recorded in the comment above PUBLISHING in test_backend_shape_stub.py, in exit criterion 2 of #124, and in #70's closing handoff when #70 closes.

The cross term. ATOM's sqsk is Sum N_Q.N_KV; the form wants Sum N_Q.N_KV_cached. On both branches the method uses N_KV = cached + N_Q (decode via seq.num_tokens, the same quantity), so Sum N_Q.N_KV_cached == sqsk - sqsq exactly: a difference of two sums ATOM already publishes, no fourth reading. Asserted.

Every shape sum is a function of the rows. BatchView holds requests and capture_rung and nothing else. No constructor takes a sum in place of the rows, and no supplied sum can disagree with the rows it came from. _refuse_shadowing gives BatchView and RequestShape the __init_subclass__ guard StepCost already has: a subclass that redefines a reader, or __post_init__ (where a row's integers and a rung's width are checked), raises TypeError at class creation. Four tests assert the refusals, including that the message names the member. Not closed, and said in the module:

A test prices two batches that collapse identically (same token total, same context total) and asserts the answers differ: that shows they are distinguishable, not that the wrong one is unreachable, and the test file says so. Another checks the sums equal what Scheduler.compute_detailed_aggregates publishes, against ATOM's own method.

BatchView is defined in backends/, so the no-engine-imports scan reaches it; test_the_projection_type_is_where_the_scan_can_reach_it resolves inspect.getfile for BatchView and RequestShape and asserts each is inside the scanned package.

Decided, not covered by #70:

  1. Species. None of the five fits: nothing was measured or regressed, and ANALYTICAL is reserved for something computed without measuring the subject. FITTED with the origin in detail is the smallest misstatement and matches the landed package's example in test_backend_interface.py; a test asserts analytical never appears. detail is dropped by both aggregating views, so step.seconds_by_species() and ProvenanceMix().seconds_by_species() each report {'fitted': 0.00014931072}: 100% of this backend's seconds reported as regressed over measured steps. Nothing in this PR consumes ProvenanceMix. The enum is the owner's call: Species has no word for declared coefficients, so a stub that measured nothing reports as fitted #87.
  2. Tier COARSE, which says step granularity. ANALYTIC is tier 0's roofline and would read as the later backend arriving early.
  3. A rung on a prefill step is refused, not ignored, and a rung narrower than the decode row count is refused. The padding is the rung's rectangle; a test asserts it exceeds the batch's own.
  4. No ladder: with one source, Resolver would wrap a single rung, so the zero-priced-term collision is never reached. A second source will meet it.
  5. One quadratic coefficient for the whole batch and no layer kind, so a stack whose layers do not all grow with history is priced as if they all did; stated in the module. Parallelism.collectives() is necessary and not sufficient, so its members are charged as candidates and each term says so.
  6. The refusals are TypeError and ValueError, not CostRefused, because they are caller defects. A Resolver wrapping this backend catches CostRefused and falls through; it will not catch these. Recorded in estimate's docstring.

Design text this PR changes. D12's Option B bullet and #70's exit criterion now state the doubling term by term (below). The stale scheduler.py:790-792 citation for detailed_sqsq/detailed_sqsk/detailed_sk (now 801-803) is corrected in 01_execution_and_time_model.md, 02_model_runner_and_cost_backend.md, 03_memory_and_kv_model.md and 14_speculative_decoding.md, with D12's gate span (:2819-2820) and method span (:2788-2841). 01 and 09_fitting_and_law_selection.md also read :2788-2841 now; 09 is outside the original file set. 04_model_capture_and_cost_ir.md is not touched; another task holds it.

Surprised. ATOM's Scheduler constructs against conftest.MockConfig, and schedule()/postprocess() drive chunked prefill for real without a driver, so the integration property is a 30 ms test and not a GPU job.

Left undone.

  • The coefficients are declared round numbers, chosen so a test's arithmetic can be checked by hand.
  • No term reads bandwidth or FLOPs from the geometry. bytes_per_block and bytes_per_token_per_layer are deliberately unused: that is the roofline, and it belongs to the analytical milestone.
  • The collective term is coefficient x tokens x layers, one all-reduce per layer, using the geometry's layer count and nothing else. The stub charges the collective on paged layers, and the stack is four times deeper #123: that count is per paged layer (geometry.layers) where the all-reduce runs on the whole stack, 16 against 64 on the hybrid; it has no referent until compass(backends): the stand-in model, its widths declared once (M1-3, #24) #120 lands stage_layers.
  • ProvenanceMix's unenforced counters are not depended on; this backend produces StepCost only.
  • The only + on a number is the integer accumulator in _per_request; each term is one product and the total is what StepCost folds. The bit-for-bit re-fold is asserted.

Named result

A 512-token request arriving at virtual 748.00 µs joins batch 6 under the full form and batch 7 with the cross term d set to 0: one term, 0.35% of elapsed time, changes which batch ATOM's real Scheduler builds. Predicted equals observed in all four cases.

Break point, predicted against observed

The batch sequence is built by ATOM's real Scheduler (chunked prefill, real admission, the real block manager). The only channel from the cost model to the scheduler is a virtual clock advanced by nothing but this backend's predictions; a request is handed to scheduler.add when that clock reaches its arrival time. A 2048-token prompt at a 256-token per-request chunk cap and a 512-token batch budget, so a second request that has become admittable can share the step.

predicted observed clock at that step
full form 6 6 749.18 µs
cross term d set to 0 7 7 746.55 µs at step 6, 1.45 µs short
chunk cap 512 instead of 256 4 4 846.87 µs
constant pricing, 250 µs/step 4 4 750.00 µs

The prediction is a closed sum over the declared coefficients and consults no scheduler; the observation is the first batch whose req_ids has two entries. Over the five steps before the arrival the cross term is worth 2.62 µs against 749.18 µs elapsed, 0.35%, and that is the whole difference between making the sixth batch and missing it.

Row four is why constants are not the default: under a constant price the break is at step 4 whether each step computes 256 tokens or 512. The work in a step does not enter, and a run cannot show that anything downstream reacted to the model.

Doubling the chunk, term by term. The tests assert each part:

chunk prefill.step prefill.tokens prefill.query_square total
256 20.00 µs 128.00 µs 1.31 µs 149.31 µs
512 20.00 µs 256.00 µs 5.24 µs 281.24 µs
ratio 1.00 2.00 4.00 1.88

#70's original "doubling the chunk doubles the step" is exactly true of the token term and cannot be true of the total, because the mandated form carries a quadratic query term. It is asserted literally where it can hold (quadratic and cross coefficients zero, no intercept: long == 2 * short exactly) and term by term everywhere else. D12's bullet and #70's exit criterion now say so: the total doubles only for a = c = d = 0.

Gates

  • 58cb9dcce: 5329 passed, 155 skipped, 3 xfailed, rc 0; tip c92e4c1c1 5281; +48 node ids, all this PR's (base update). GPU tier not required.

Effort

Estimate 200 lines. Production physical non-blank 443 (2.22x); by AST 146 (0.73x). Over 2x is an escalation; the unit is #89's ruling.

Effort, three instruments

Measured at 0fc02f458 against d175b03c6. AST = parse, strip module/class/function docstrings, unparse, count non-blank.

AST physical, non-blank physical, all
production (shape.py + the __init__.py delta) 146 443 517
test (test_backend_shape_stub.py + the test_backend_interface.py delta) 300 572 696
both 446 1015 1213

Against the 200: production 0.73x by AST, 2.22x physical non-blank, 2.59x physical-all; the pair 2.23x by AST and 5.08x physical non-blank. Under the unit #70 now states (production, physical non-blank, excluding tests) the reading went 362 (1.81x) -> 425 (2.12x) -> 443 (2.22x); the last +18 are itemised in the review-reply comment. Reported, not trimmed; no comment or docstring was cut to offset it. A review that asks for explanation moves this number up.

The physical-to-AST ratio is 3.03x production and 1.91x test, against 2.3x for the parent and 2.03x for the landed backends/ package: the gap is the project's documentation style. The 200 is already 3.3x D12's sizing of the same option (~60 lines), so a 2x rule applied to it partly measures the estimator. Whichever unit the owner picks, #70 should state it ("200 production LOC, physical, excluding tests"); the ambiguity, not the code, produced the split.

Generated with Claude Code

Comment thread atom/compass/backends/shape.py Outdated
class. Whoever holds a scheduled batch reads the integers off it and
builds these rows.

There is no field and no constructor argument for a batch-level sum. The

@jgong5 jgong5 Sep 21, 2026 •

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The collapsed form is reachable three ways this sentence does not cover (module docstring, line 48: "cannot be handed in even by accident"). Non-blocking. Fix: __init_subclass__ guards on BatchView and RequestShape, as StepCost has (~10 lines), and soften this docstring and the PR body to what was measured.

The three routes, run on this branch

True, and worth keeping: no constructor path takes tokens/history in place of rows, and no supplied sum can disagree with the rows it came from.

  1. A one-row BatchView is tokens x history. BatchView((RequestShape(512, 2048, False),)) is legal and gives Sum N_Q.N_KV = 512 x 2048 = 1048576, against 524288 for the two-row batch it collapses (RequestShape(256,1024) twice). They price differently, 284.389 us vs 280.194 us, as test_two_batches_that_collapse_alike_are_priced_apart shows, but that proves they are distinguishable, not that the wrong one is unreachable. Nothing ties len(requests) to the number of requests the scheduler scheduled.
  2. A subclass passes estimate's isinstance check and can override the prefill property to synthesise a collapsed row from fields it added. I built one; its price is bit-equal to (1). StepCost in this package refuses exactly this with __init_subclass__ ("an answer that does not come from the terms is the thing this class exists to prevent"); the projection does not inherit it.
  3. RequestShape the same way: a subclass overriding cached_tokens fabricates the cross term directly.

What actually prevents collapsing is the projection step, one row per scheduled request, and that lives as _project in tests/compass/test_backend_shape_stub.py (lines 396-413), with no home in the package. #70 asked for the type to sit under backends/ so the scan reaches it; the builder that makes the type honest is still unowned.

@jgong5 jgong5 Sep 21, 2026 •

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Round 2 pushed 28592b983: both halves applied. _refuse_shadowing refuses routes 2 and 3 at class creation, and also a subclass dropping __post_init__, where row integers and rung width are checked. The protected set is read off the owner class, as StepCost does, so a later reader is covered. The docstrings and PR body are softened; test_one_row_is_a_legal_batch_and_is_the_collapsed_form asserts route 1. Owning the projection stays the successor's task.

Refusals and what the docstrings say
route 2: BatchView subclass overriding prefill
  REFUSED: TypeError: CollapsedView redefines prefill of BatchView; the shape sums are
  functions of the rows, and a reader that answers from anything else is what summing
  per request exists to prevent
route 3: RequestShape subclass overriding cached_tokens
  REFUSED: TypeError: Fabricated redefines cached_tokens of RequestShape; ...

__post_init__ is the only non-public member whose replacement changes what the object reports: a subclass that drops it admits shapes the sums are not defined over. A subclass that adds a field and redefines nothing is still accepted and its sums still come from the rows, so the guard is not a ban on subclassing.

The module docstring, the class docstring, the test-file docstring and the PR body now say: every shape sum is a function of the rows, no constructor takes a sum in place of them, and no supplied sum can disagree with the rows it came from. Then what that does not close: one row is a legal batch, and a one-row batch is tokens x history (1 048 576 against 524 288, 284.389 µs against 280.194 µs, reproduced). The module says the projection, one row per scheduled request, is what prevents collapsing, and that its builder does not live in this package yet.

"""

requests: tuple[RequestShape, ...]
capture_rung: int | None = None

@jgong5 jgong5 Sep 21, 2026 •

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

capture_rung is a batch-level scalar field, so "no field and no constructor argument for a batch-level sum" (line 176) and the PR body's "no field ... accepts any batch-level scalar" are false as written. It is the right scalar to carry, but it is unbounded above and enters the price linearly. Non-blocking. Fix: say "the shape sums are functions of the rows", here and in the PR body.

Measured

The rung is a property of the replayed graph, not of the rows, and __post_init__ guards the two directions that matter (a rung with no decode rows, a rung narrower than the decode count).

BatchView((RequestShape(1, 100, True),), capture_rung=10**9)
  -> graph_padding = 99999999900  ->  decode.graph_padding = 99.9999999 s

One decode row, one hundred seconds. No caller would pass that; the point is the wording.

@jgong5 jgong5 Sep 21, 2026 •

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Round 2 pushed 28592b983: fixed. The module docstring and PR body now say capture_rung is the right scalar (a property of the replayed graph), checked by __post_init__ in both directions against the decode rows, and unbounded above, entering the price linearly. Your capture_rung=10**9 figure reproduces on the new head and is in both. The structural claim everywhere is now "the shape sums are functions of the rows".

name,
count * coefficient,
Provenance(
Species.FITTED,

@jgong5 jgong5 Sep 21, 2026 •

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

FITTED is the right local call and still leaves a false aggregate: detail is dropped by both aggregating views, so 100% of this backend's seconds report as regressed over measured steps. Non-blocking, since nothing here consumes ProvenanceMix. Fix: register it in 12_open_items.md or a follow-up issue; adding DECLARED = "declared" to a closed vocabulary is the owner's call.

Measured

Accepted locally: ANALYTICAL is reserved by provenance.py, the landed package's own example is Provenance(Species.FITTED, "declared coefficients") in test_backend_interface.py, and test_the_species_is_not_analytical pins it. The disclosure in detail reaches rows() and describe(); checked against the real output.

step.seconds_by_species()           -> {'fitted': 0.00014931072}
ProvenanceMix().record(step); .seconds_by_species() -> {'fitted': 0.00014931072}

FITTED means "a form chosen, coefficients regressed over measured steps", and nothing was regressed, with no decomposition carrying the correction. That breaks two design principles: never report an aggregate without its decomposition, and every claim carries its measurement (here a number with a false source). The backend emits only StepCost, verified. The fix is one enum member plus a sentence in provenance.py, but the design calls the vocabulary "closed and small". Left in a PR body, the next backend copies FITTED from here for the same reason this one did.

@jgong5 jgong5 Sep 21, 2026 •

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Round 2 pushed 28592b983: registered as #87, "Species has no word for declared coefficients, so a stub that measured nothing reports as fitted", referenced from the PR body. No enum member here and no 12_open_items.md row: another PR is guarding that file's counts. FITTED stays, pinned by test_the_species_is_not_analytical. The aggregate is unchanged, and the PR body now states it in those terms.

Re-measured on 28592b9
step.seconds_by_species() -> {<Species.FITTED: 'fitted'>: 0.00014931072}

whatever `StepCost` folds from them, so there is no second summation
to disagree with the first.
"""
if not isinstance(batch_view, BatchView):

@jgong5 jgong5 Sep 21, 2026 •

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

isinstance rather than type(...) is lets route 2 of the line-176 comment through: a BatchView subclass overriding prefill is accepted here. The refusal message is good: it names the type and who builds the projection. Not a change request: an __init_subclass__ guard on BatchView closes it with no change on this line. Accepted with reservation: CostBackend.estimate names CostRefused for refusals, and this backend raises TypeError here and ValueError in BatchView.__post_init__; state it, because a Resolver catches only CostRefused.

@jgong5 jgong5 Sep 21, 2026 •

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Round 2 pushed 28592b983: both remarks taken. The isinstance needs no change: BatchView now refuses the subclass at class creation (refusal text in the line-176 reply). The asymmetry is recorded in estimate's docstring, where a caller looks, and in the PR body's "left undone" list.

The docstring sentence

The two refusals below and in BatchView.__post_init__ are TypeError and
ValueError rather than the CostRefused the seam names, because a batch that is not
a batch is a caller defect and not a step this backend declines to price; the asymmetry
is worth stating because a resolver that wraps this backend catches CostRefused and
falls through to the next rung, and will not catch these.

@jgong5

jgong5 commented Sep 21, 2026 •

Copy link
Copy Markdown
Owner Author

APPROVE at af97a2bf7 (posted as COMMENT): 5 findings and 3 reservations, none blocking the code. Finding 5 blocks landing: the tip moved and this PR now conflicts with feature/atomcompass_new.

Checked: the diff is the four files claimed; the named result reproduces, 0.35% included; the doubling ratios, sqsk - sqsq, the ungating claim and the disclaimer in the output all reproduce; branch af97a2bf7 4279 passed against 4239 at ff0cc3e4f; ruff clean.

Findings:

  1. The "inexpressible" claim is overstated: four routes reach the collapsed form (routes 1-3, the rung). Medium, non-blocking.
  2. FITTED for a declared coefficient survives into an aggregate the disclaimer does not reach (inline). Medium, non-blocking.
  3. The profile RPC is not equivalent to assigning profile_active. Low.
  4. M1-2 — the shape-analytic cost stub and its closed loop through the scheduler #70 and D12 overstate the chunk doubling the mandated form forbids; amend both.
  5. The integration tip moved; restack onto 14a197b07 and patch the base. Blocks landing.

Accepted with reservation: the 0.35% is a constructed margin; the rung path has no integration coverage; CostRefused is never raised (inline).
Watch next: own the projection, not just the type; set profile_active directly; Species has no word for "declared"; one global quadratic coefficient; the zero-priced-term collision.

Findings 3-5 in full

3. The RPC path. _handle_start_profile in atom/model_engine/engine_utility.py also calls runner_mgr.call_func("start_profiler"), and ModelRunner.start_profiler starts a real torch.profiler and installs a trace-export callback whenever profiler_dir is set. So "flag plus one assignment, no code" holds only for the direct assignment; the RPC drags a profiler and its trace files into a simulated run. Worth one clause in the handoff so a successor does not reach for the RPC.

4. The doubling criterion. #70's exit criterion and D12's Option B bullet say "double the chunk size and the step time doubles". The form the same documents mandate is a + b.tokens + c.Sum N_Q^2 + d.Sum N_Q.N_KV_cached: token term 2.00x, quadratic 4.00x, intercept 1.00x, total 1.8836x. The total doubles only for a = c = d = 0, the degenerate case test_the_step_doubles_exactly_where_the_form_is_linear asserts, so the criterion as written is satisfiable only by deleting the quadratic term #70 requires. Testing it literally where it can hold and term by term elsewhere is right, but a flag in a PR body is lost: amend D12's bullet and #70's exit criterion to say the token term doubles and the total rises by the form's terms. Otherwise the milestone-2 analytical backend inherits the contradiction.

5. The conflict. fork/feature/atomcompass_new is now 14a197b07: the parent landed as #74, squashed, so the merge-base fell back to a8ebf8c47.

git merge-tree --write-tree 14a197b07 af97a2bf7
CONFLICT (content): Merge conflict in atom/compass/backends/__init__.py    rc=1

A history artifact, not a content disagreement: the parent's hunks arrive twice by two routes. Taking this branch's __init__.py gives a tree that differs from the tip by exactly the four files, gated at 4439 below. Restack onto 14a197b07, patch the base with gh api -X PATCH repos/jgong5/ATOM/pulls/81 -f base=feature/atomcompass_new, and restate the gate as +40 over 4399; the +59 line is history.

Checked, with measurements

Diff: git diff --name-status ff0cc3e4f...HEAD gives M atom/compass/backends/__init__.py, A atom/compass/backends/shape.py, M tests/compass/test_backend_interface.py, A tests/compass/test_backend_shape_stub.py; merge-base ff0cc3e4f.

Closed loop, with my own driver against ATOM's Scheduler and MockConfig, reading the clock at each step rather than calling _drive/_predicted_break:

predicted observed clock entering that step
full form 6 6 749.175 µs
cross term d = 0 7 7 746.554 µs at step 6, 1.446 µs short of the 748.00 µs arrival
chunk cap 512 4 4 846.874 µs
constant 250 µs/step, chunk 256 4 4 750.000 µs
constant 250 µs/step, chunk 512 4 4 750.000 µs

Cross-term contribution over the five steps before the arrival: 749.1750 - 746.5536 = 2.6214 µs, 0.3499% of 749.1750 µs. The arrival sits inside [746.554, 749.175], so with the term the second request makes batch 6 and without it batch 7. Under constant pricing the break is step 4 at both chunk sizes.

Doubling: prefill.step 20.00 -> 20.00 µs (1.00x), prefill.tokens 128.00 -> 256.00 µs (2.00x), prefill.query_square 1.3107 -> 5.2429 µs (4.00x), total 149.31 -> 281.24 µs (1.8836x).

sqsk - sqsq: Scheduler.compute_detailed_aggregates computes sqsq += nq*nq and sqsk += nq*nkv from the same nq and nkv per request on both branches, so Sum N_Q.(N_KV - N_Q) = sqsk - sqsq identically, and #64's fourth reading is not needed. The identity holds whatever nkv means; what needs checking is that the projection's N_Q/N_KV match ATOM's per-branch definitions, which test_the_sums_are_the_ones_the_scheduler_publishes does against ATOM's own method. The design's scheduler.py citation for the three fields (790-792) is stale; they are at 801-803 today.

Ungating: EngineCore._process_engine_step_inner calls compute_detailed_aggregates unconditionally under has_seqs; the gate is inside the method; ATOM_ENABLE_DETAILED_ANNOTATION is a real env flag in atom/utils/envs.py, cached in Scheduler.__init__; profile_active is public, initialised False, flipped only by _handle_start_profile/_handle_stop_profile. No ATOM edit needed or made.

The disclaimer reaches the output:

prefill.tokens  1.280e-04  fitted (5e-07 s x 256; declared, not measured -- a plumbing
                           figure, not an accuracy claim) via shape-stub
describe(): step-level stand-in, shape-read coefficients, tp1 pp1 dp1, no collectives
            -- declared, not measured -- a plumbing figure, not an accuracy claim

Gates on node 18 (xiaobizh_n18_cpu):

tree commit passed skipped xfail rc
control (branch point) ff0cc3e4f 4239 149 3 0
this branch af97a2bf7 4279 149 3 0
integration tip (moved) 14a197b07 4399 149 3 0
merged (branch + new tip) e0013f61c 4439 149 3 0

+40 decomposes: 37 tests in test_backend_shape_stub.py, and test_backend_interface.py 10 -> 13 (2 parametrized cases + 1 import-scan case for shape.py). The two files: 50 passed in 0.19 s. Skips and xfails identical across all four trees, so the CPU flake did not fire. gpu: not required (.compass-changed stamp); none of the four paths is in gpu_gate_triggers.txt. ruff check and ruff format --check clean on all four files. The PR's merged 4439 is confirmed, but the parent landed as #74, so its +19 is now on the tip (4380 + 19 = 4399) and the merged delta is +40 over 4399, not +59 over 4380.

No design-document identifiers, task labels or doc file names in either new file.

Reservations and watch-next in full

Accepted with reservation:

  1. The 0.35% is a constructed margin: the arrival at 748.00 µs sits inside a 2.62 µs window, and any arrival outside it breaks at the same step with and without the term. M1-2 — the shape-analytic cost stub and its closed loop through the scheduler #70 asked for a batch sequence built to order and the PR says so, but the result shows sensitivity of the schedule to one term, not that 0.35% generally decides a batch. Do not cite it as a generic figure.
  2. capture_rung is never set in the closed loop, so decode.graph_padding is unit-tested only. D12 names ForwardMode.decide as the rule that selects the graph and nothing here consults it. Correct scope for a stub.
  3. CostBackend.estimate's docstring says a refusal is CostRefused; this backend raises TypeError and ValueError. Arguably right for caller errors, but worth one sentence: a Resolver wrapping this backend catches CostRefused and not these.

Watch next:

  1. Own the projection. The builder that guarantees one row per scheduled request makes finding 1 moot and exists only as _project in a test. CompassModelRunner is where it belongs, with the compute_detailed_aggregates cross-check as its test.
  2. Set scheduler.profile_active = True directly, not via the profile RPC.
  3. Species has no word for "declared": the first backend that feeds ProvenanceMix will report a false species mixture.
  4. The single global quadratic coefficient is already wrong for the 27B hybrid target (D12's own open issue); the term-per-kind repair belongs to 04_model_capture_and_cost_ir.md, not to a different number.
  5. The zero-priced-term collision in the ladder is unreached with one source, and prefill.query_cached is legitimately 0.0 on a first chunk; the second source will meet it.
Effort: my counts and my reading

Same AST method (parse, strip module/class/function docstrings, unparse, count non-blank), plus physical:

AST physical non-blank physical all
shape.py 133 344 407
__init__.py delta ~1 +18 +18
production ~134 362 425
test_backend_shape_stub.py 238 461 561
test_backend_interface.py delta (net) ~6 +16 +19
test ~244 477 580
both ~378 839 1005

Every physical figure in the PR body reproduces exactly; the AST figures to within a line. Production 0.67x by AST, 1.81x physical; the pair 1.89x by AST, 4.20x physical.

My reading for the owner: the 200 is production physical non-blank, 362 = 1.81x, inside the halt line. (a) D12 sizes the options physically and in production only ("about 40 lines", "about twenty lines more than A"). (b) An AST count is invariant to the docstrings that are this house style's main cost, so it will never fire. (c) Counting tests against a production estimate makes the 2x rule punish test coverage.

For the owner, since four tasks have now reported the same split: the physical-to-AST ratio here, 2.70x production and 1.95x test, is in line with the parent's 2.3x and the landed package's 2.03x, so the gap is the project's documentation style. #70's 200 is already 3.3x D12's sizing of the same option (~60 lines), so a 2x rule applied to it measures the estimator. Whichever unit is chosen, #70 should state it ("200 production LOC, physical, excluding tests").

Not checked
  • The GPU tier: not required for these paths; scope, not a gap.
  • The PR's 68ef4f329 = 4380 row: the tip had moved past it. Consistent with 4380 + 19 = 4399.
  • Accuracy: there is nothing here to be accurate about, and the output says so in every row.

jgong5 pushed a commit that referenced this pull request Sep 21, 2026
…m, and state what the type does not close

Review round 2 on #81. The module claimed the collapsed form "cannot be handed
in even by accident". Four routes to it were measured. Two are closed here and
two are stated rather than claimed away.

Closed at the type. `_refuse_shadowing` gives `BatchView` and `RequestShape`
the `__init_subclass__` guard `StepCost` in this package already has: a
subclass that redefines a reader can answer from a field it added rather than
from the rows, and it passes the backend's `isinstance` check. Overriding
`prefill` synthesises one row holding a whole batch's tokens against a whole
batch's history, which is the collapsed form exactly; overriding
`cached_tokens` fabricates the cross term directly. `__post_init__` is
protected by name as the only non-public member whose replacement also changes
what the object reports.

Stated, not closed. A one-row `BatchView` is `tokens x history` by
construction, and nothing in the type ties the row count to the number of
requests a scheduler scheduled -- the projection does that, and the projection
is not in this package yet. `capture_rung` is a batch-level scalar field,
correctly so because a rung is a property of the replayed graph rather than of
the rows, but it is unbounded above and enters the price linearly. The
docstrings now say the property that is true -- every shape sum is a function
of the rows -- rather than the stronger one that is not.

The refusals here are `TypeError` and `ValueError` where the seam names
`CostRefused`; `estimate` now records that a resolver wrapping this backend
will not catch them.

Design text, where a flag in a PR body would have been lost. D12's Option B
bullet said "double the chunk size, the step time doubles", which the form the
same section mandates forbids: measured, the intercept is 1.00x, the token
term 2.00x, the quadratic 4.00x and the total 1.8836x, and the total doubles
only for a = c = d = 0. The bullet now states the term-by-term property with
those figures. The stale `scheduler.py:790-792` citation for the three
detailed aggregate fields is corrected to `801-803` in all four design files
that carry it, and D12's gate and method spans to `:2819-2820` and
`:2788-2841`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@jgong5
jgong5 force-pushed the compass/m1-2-cost-stub branch from af97a2b to f8ad2a0 Compare September 21, 2026 18:18
@jgong5
jgong5 changed the base branch from compass/m1-1-kv-geometry to feature/atomcompass_new September 21, 2026 18:18
jgong5 pushed a commit that referenced this pull request Sep 21, 2026
…m, and state what the type does not close

Review round 2 on #81. The module claimed the collapsed form "cannot be handed
in even by accident". Four routes to it were measured. Two are closed here and
two are stated rather than claimed away.

Closed at the type. `_refuse_shadowing` gives `BatchView` and `RequestShape`
the `__init_subclass__` guard `StepCost` in this package already has: a
subclass that redefines a reader can answer from a field it added rather than
from the rows, and it passes the backend's `isinstance` check. Overriding
`prefill` synthesises one row holding a whole batch's tokens against a whole
batch's history, which is the collapsed form exactly; overriding
`cached_tokens` fabricates the cross term directly. `__post_init__` is
protected by name as the only non-public member whose replacement also changes
what the object reports.

Stated, not closed. A one-row `BatchView` is `tokens x history` by
construction, and nothing in the type ties the row count to the number of
requests a scheduler scheduled -- the projection does that, and the projection
is not in this package yet. `capture_rung` is a batch-level scalar field,
correctly so because a rung is a property of the replayed graph rather than of
the rows, but it is unbounded above and enters the price linearly. The
docstrings now say the property that is true -- every shape sum is a function
of the rows -- rather than the stronger one that is not.

The refusals here are `TypeError` and `ValueError` where the seam names
`CostRefused`; `estimate` now records that a resolver wrapping this backend
will not catch them.

Design text, where a flag in a PR body would have been lost. D12's Option B
bullet said "double the chunk size, the step time doubles", which the form the
same section mandates forbids: measured, the intercept is 1.00x, the token
term 2.00x, the quadratic 4.00x and the total 1.8836x, and the total doubles
only for a = c = d = 0. The bullet now states the term-by-term property with
those figures. The stale `scheduler.py:790-792` citation for the three
detailed aggregate fields is corrected to `801-803` in all four design files
that carry it, and D12's gate and method spans to `:2819-2820` and
`:2788-2841`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@jgong5
jgong5 force-pushed the compass/m1-2-cost-stub branch from f8ad2a0 to 28592b9 Compare September 21, 2026 18:18
@jgong5

jgong5 commented Sep 21, 2026 •

Copy link
Copy Markdown
Owner Author

Round 2 pushed 28592b983, restacked onto 669dc3f9d; base is now feature/atomcompass_new. Fixed: findings 1 (shadowing refused at the type, other routes stated), 2 (#87), 4 (design bullet and #70 amended), 5 (restack), reservation 3; replies inline. Gate: 4541 passed against 4496 at 669dc3f9d. Effort now 2.12x: over the halt line, flagged, not trimmed.

Restack

The head had moved past 14a197b07 by the time of the restack: SPEC-1 landed as #79, making the head 669dc3f9d.

git rebase --onto 14a197b07 ff0cc3e4f compass/m1-2-cost-stub   # #74
git rebase --onto 669dc3f9d 14a197b07 compass/m1-2-cost-stub   # #79
git merge-base --is-ancestor 669dc3f9d HEAD   ->  PASS
git merge-tree --write-tree  669dc3f9d HEAD   ->  rc=0, no conflict

Both replayed clean; a plain git rebase conflicts in backends/__init__.py and --onto does not. Base retargeted by REST (gh pr edit --base is broken on gh 2.45.0). Force-push was the one permitted case; no gh stack. The branch is a descendant of the head, so the merged tree is the branch tree. git diff --name-status 669dc3f9d..HEAD is eight files: the original four plus four design files.

Finding 1: the routes, re-run on 28592b9
route result
1. one-row BatchView = tokens x history still legal: 1 048 576 vs 524 288, 284.389 µs vs 280.194 µs. Stated, not closed.
2. BatchView subclass overriding prefill refused at class creation
3. RequestShape subclass overriding cached_tokens refused at class creation
3b. subclass dropping __post_init__ refused
4. capture_rung=10**9 on one decode row still 99.9999999 s. Stated, not closed.
control: subclass adding a field, redefining nothing accepted, sums still from the rows

The docstrings and the PR body claim only that the shape sums are functions of the rows, and say that a one-row batch is the collapsed form, that the row count is the projection's guarantee, and that _project has no home in the package: the successor's task, not closed here.

Findings 2 and 4, and the stale citations

Finding 2: Species gains nothing and 12_open_items.md gains no row (another PR is guarding that file's counts). The PR body references #87 and states the measurement, re-measured on the new head and unchanged.

Finding 4: D12's Option B bullet now states the term-by-term property with the figures: intercept 1.00x, token term 2.00x, quadratic 4.00x, total 1.8836x, and the total doubles only for a = c = d = 0. #70's exit criterion is amended the same way, with a note that the brief was wrong and this is the correction. 04_model_capture_and_cost_ir.md deliberately untouched.

Stale citations: scheduler.py 790-792 -> 801-803 in all four design files that carry it (01_execution_and_time_model.md, 02_model_runner_and_cost_backend.md, 03_memory_and_kv_model.md, 14_speculative_decoding.md), plus two more in D12: the gate is at 2819-2820 (cited 2820-2821) and compute_detailed_aggregates spans 2788-2841 (cited 2788-2842). No design references in code, comments or emitted strings, re-grepped.

Gate and effort on 28592b9

Node 18, xiaobizh_n18_cpu:

tree commit passed skipped xfail rc
control (integration head) 669dc3f9d 4496 149 3 0
this branch = merged 28592b983 4541 149 3 0

+45: 42 tests in test_backend_shape_stub.py (37 at round 1, six new replacing one), and test_backend_interface.py 10 -> 13. The two files: 55 passed in 0.20 s. Skips and xfails identical, so the ±1 flake did not fire. gpu: not required (.compass-changed stamp). ruff check and ruff format --check clean on all changed files. The round-1 4439 is superseded: two PRs landed under it, so the control is a measured 4496, not a read 4399.

AST physical non-blank physical all
production 145 425 497
test 281 533 650
both 426 958 1147

Under the round-1 reading (production physical non-blank) this is 2.12x, up from 1.81x. Round 2 added +63 physical non-blank to shape.py, mostly prose the review asked for: the softened claim in two docstrings, _refuse_shadowing's rationale, the CostRefused sentence. More than ~2x is a halt-and-discuss event, so it is flagged rather than trimmed. By AST it is 0.72x; physical-to-AST 2.93x production, 1.90x test.

Not done: no enum member or 12_open_items.md row (#87 instead); the projection has no home in the package; not merged, not undrafted, no reviewer spawned.

loop — step cost to queueing to a different batch — is what the whole design rests on,
and constants cannot exercise it.
- *Not* "the step time doubles". Measured against the form this section mandates, at
`b`=5e-7, `c`=2e-8 and a 256→512 single-request chunk: the intercept is 1.00x, the

@jgong5 jgong5 Sep 22, 2026 •

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Finding 1: this bullet states c=2e-8; the backend's c is 2e-11, and with 2e-8 the bullet's own 1.8836x becomes 3.7834x. a (2e-5) is omitted, and the total ratio depends on it. This amendment retires an unreproducible claim, and the milestone-2 analytical backend will read it. Fix: 2e-8 -> 2e-11, plus one clause naming a=2e-5.

Computed on 28592b9
shape.py Coefficients.prefill_query_square = 2.0e-11   <- the backend's c
c=2e-11 : short=149.3107 us  long=281.2429 us  ratio=1.8836x   <- the figure this bullet quotes
c=2e-8  : short=1458.7200 us long=5518.8800 us ratio=3.7834x   <- the figure this bullet's own c gives

The per-term ratios (1.00x / 2.00x / 4.00x) are independent of the coefficients and correct: prefill.step 1.000000, prefill.tokens 2.000000, prefill.query_square 4.000000, total 1.883608. Only the total is unreproducible from the stated inputs.

@jgong5 jgong5 Sep 22, 2026 •

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Round 3 pushed 0fc02f458: fixed. This bullet now reads a=2e-5, b=5e-7, c=2e-11, names the chunk as a first chunk so d's absence is visible, and carries both absolute step times so the ratio reproduces from what is written. Your 3.7834x for c=2e-8 reproduces too (1458.72 µs -> 5518.88 µs). The per-term ratios are unchanged. #70's exit criterion states no coefficients, so it needed no change; checked.

The arithmetic now in the bullet
a=2e-5, b=5e-7, c=2e-11, first chunk 256 -> 512, single request
  short = 2e-5 + 5e-7*256 + 2e-11*65536  = 149.31072 us
  long  = 2e-5 + 5e-7*512 + 2e-11*262144 = 281.24288 us
  ratio = 1.883608

Comment thread atom/compass/backends/shape.py Outdated
correctly so, because a rung is a property of the replayed graph rather than
of the rows; it is checked against the decode rows in both directions but is
unbounded above and enters the price linearly, so a rung nobody would pass
prices a single decode row at any duration one likes. Subclassing is refused

@jgong5 jgong5 Sep 22, 2026 •

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Finding 2: "Subclassing is refused outright" is not what the guard does. Shadowing is refused; a BatchView subclass adding a field and redefining nothing is accepted, as your round-1 reply and _refuse_shadowing's docstring both say. The type moved and the prose overshot it again. Fix: "Shadowing a reader is refused outright", no net lines.

Measured on 28592b9
BatchView subclass overriding prefill        -> REFUSED at class creation
RequestShape subclass overriding cached_tokens -> REFUSED
RequestShape subclass replacing __post_init__  -> REFUSED
BatchView subclass adding a field, redefining nothing -> ACCEPTED, sums still from the rows

@jgong5 jgong5 Sep 22, 2026 •

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Round 3 pushed 0fc02f458: fixed. The sentence now reads: "Shadowing a reader is refused outright (see _refuse_shadowing); a subclass that adds a field and redefines nothing is accepted, and its sums still come from its rows." Pinned: test_a_subclass_that_redefines_nothing_is_accepted adds a label field to a frozen BatchView subclass, asserts Sum N_Q.N_KV still comes from its rows, and asserts it prices like the plain batch. The protected sets are unchanged.

Comment thread atom/compass/backends/shape.py Outdated
a reading from a field a subclass added; the row count is not closed here and
cannot be. What closes it is the projection -- one row per scheduled request,
built by whoever holds the scheduled batch -- and that builder does not live
in this package yet.

@jgong5 jgong5 Sep 22, 2026 •

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Finding 3: both unclosed routes (the row count and the rung) are named here and registered nowhere, and the successor named as their owner has landed: #80, #90 and #97 reference neither type. Round 1's standard for the species defect was "a registered open item or a follow-up issue", and #87 met it. Fix, no diff lines: open one issue naming the projection builder and its cross-check against Scheduler.compute_detailed_aggregates, and reference it as #87 is.

Measured on d175b03
git grep -n "BatchView\|RequestShape\|_project\|CostBackend" d175b03c6 -- atom/compass/runner/   ->  no matches

No open issue names the projection builder. The next agent to build a projection will read the seam's signature, not this docstring. My answer to the review question is in the review record: for the row count, stating it is a close; what is missing is the schedule.

@jgong5 jgong5 Sep 22, 2026 •

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Round 3 pushed 0fc02f458: registered as #124, carrying both unclosed routes, since whoever builds the projection reads the rows off the scheduled batch and the rung off the capture ladder. The gap reproduces at d175b03c6. One deviation to rule on: #124 is cited in the PR body, not in shape.py, as #87 is. Say so if the finding requires it in the source.

#124's exit criteria, and why the number stays out of the source

Exit criteria: one row per scheduled request over a batch built by ATOM's real Scheduler, plus a test that a two-request batch does not project to one row; the sums cross-checked against Scheduler.compute_detailed_aggregates (scheduler.py:2788-2841), ATOM's own method rather than a re-derivation; the rung from the ladder, None where no graph replayed. It also carries round-1 finding 3 as a criterion, so "assign the attribute, do not issue the RPC" outlives this PR body.

Why not in shape.py: AI_DEV_RULES bans design-doc identifiers in code and says the durable task record is the issue and its PR; a number in a docstring is the same habit in another vocabulary, and it would add lines to a diff already over the halt line.

"""

requests: tuple[RequestShape, ...]
capture_rung: int | None = None

@jgong5 jgong5 Sep 22, 2026 •

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Finding 4: the third rung check, a rung far wider than the rows, is absent, the asymmetry is not argued, and nothing pins the statement, unlike route 1. The provenance's declared, not measured covers the coefficient, not the count; CANDIDATE exists here for a count less certain than its reading. Not blocking: the closed loop never sets capture_rung. Fix: close it, or argue it and pin it with a test; register it at minimum.

Measured on 28592b9
BatchView((RequestShape(1, 100, True),), capture_rung=10**9)
  graph_padding      = 99999999900
  decode.graph_padding = 99.9999999 s
  total                = 100.0000223 s      (one decode row)

__post_init__ refuses a rung with no decode rows and a rung narrower than the decode count; both check a rung against the rows. The open direction is the one that yields a fictional number rather than a negative one. grep capture_rung tests/compass/test_backend_shape_stub.py returns the two refusal tests, the field-list test and two pricing tests; nothing pins the unbounded-above behaviour, so bounding the rung later would leave this paragraph silently stale.

@jgong5 jgong5 Sep 22, 2026 •

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Round 3 pushed 0fc02f458: argued and pinned, not closed. The module docstring now argues the asymmetry; two tests pin both directions; the padding term's provenance carries UNCHECKED_RUNG whenever a rung was supplied. The close, a rung read off the real capture ladder, is the second half of #124. Cost: +18 physical non-blank lines.

The argument, the tests, the qualifier

The argument: the two refused directions are refusable from the rows. A rung with no decode rows names a graph that replayed nothing; a rung narrower than the decode row count makes the padding negative. A rung far wider than the rows is not that kind of error: rows cannot contradict a width, only a ladder of captured widths can, and this backend has no ladder (the "no ladder" decision in the dev record). Refusing on a guessed bound would invent the ladder, the kind of decision this project treats as a defect. So it is priced at face value.

The tests:

  • test_a_rung_no_ladder_would_hold_is_priced_and_labelled asserts graph_padding == 10**9 * 100 - 100, decode.graph_padding == 99.9999999 s, and the qualifier in the provenance.
  • test_with_no_rung_the_padding_term_carries_no_rung_qualifier asserts the qualifier is absent with no rung, so it cannot become decoration on every batch.

Bounding the rung later now breaks a test.

The qualifier:

UNCHECKED_RUNG = "padding from a supplied rung width; nothing here bounds it above"

It is attached only when a rung was supplied: on a zero term it would claim a width was supplied when none was. CANDIDATE was not reused: its text says a collective may not run, a different uncertainty from a count nothing bounds. The +18: +9 in the docstring paragraph, +9 for the constant and the conditional term; nothing was trimmed to offset it.

- `ScheduledBatch` fields — `scheduler.py:579-820`; notably `detailed_sqsq` /
`detailed_sqsk` / `detailed_sk` at `:790-792`, which are sum(N_Q^2), sum(N_Q * N_KV),
`detailed_sqsk` / `detailed_sk` at `:801-803`, which are sum(N_Q^2), sum(N_Q * N_KV),
sum(N_KV) per batch, computed by `compute_detailed_aggregates` (`:2788-2842`) and

@jgong5 jgong5 Sep 22, 2026 •

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Finding 5, low: after this PR three design documents cite compute_detailed_aggregates with two end lines. This line, edited here on its adjacent half, still reads :2788-2842; 02_model_runner_and_cost_backend.md now says :2788-2841, which is right. 09_fitting_and_law_selection.md is outside the file set. Fix: one character here.

Measured

def compute_detailed_aggregates is at scheduler.py:2788 and its last statement (scheduled_batch.detailed_sk = sk) at :2841; :801-803 and the gate at :2819-2820 also check out.

01_execution_and_time_model.md:1399   compute_detailed_aggregates (`:2788-2842`)
02_model_runner_and_cost_backend.md:468  compute_detailed_aggregates (`:2788-2841`)
09_fitting_and_law_selection.md:96       scheduler.py:2788-2842

@jgong5 jgong5 Sep 22, 2026 •

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Round 3 pushed 0fc02f458: fixed here and in 09_fitting_and_law_selection.md, both now :2788-2841. def compute_detailed_aggregates is at 2788, its last statement at 2841, blank at 2842, def _connector_flag at 2843; the gate at :2819-2820 and the fields at :801-803 check out. 09 is outside the original file set: one character, and one of three documents left disagreeing is worse than the conflict risk. My call to take.

@jgong5

jgong5 commented Sep 22, 2026 •

Copy link
Copy Markdown
Owner Author

REQUEST CHANGES at 28592b983: round 3 required for 5 new findings and round-1 finding 3, still open. Nothing wrong in the algorithm; the fixes are mostly text, about five net lines.

Checked: the eight-file diff; the four routes, the control and the guard's full protected sets; the species aggregate; #87; the named result by closed arithmetic; the four inline replies against the tree; ruff; the gate, +45 at 28592b983 over 669dc3f9d.

Findings, inline:

  1. D12's new bullet states c=2e-8; the backend's is 2e-11, and its 1.8836x needs 2e-11.
  2. "Subclassing is refused outright" over-claims.
  3. Both unclosed routes are registered nowhere, and the runner landed without them.
  4. capture_rung's third check is absent and unpinned.
  5. Three design documents cite one method with two end lines.

Round-1 finding 3 (the profile RPC) is not yet a two-cycle halt, because its handoff remedy was not due; if still open after round 3, the loop halts on it.
Accepted with reservation: none new.
Watch next: own the projection; bound or qualify the rung; #87; the global quadratic coefficient; the zero-priced-term collision.

Round-1 finding 3, still open

The RPC path _handle_start_profile in atom/model_engine/engine_utility.py calls runner_mgr.call_func("start_profiler") before it flips the flag, and ModelRunner.start_profiler starts a real torch.profiler whenever profiler_dir is set; re-read, confirmed. The round-2 comment has no section for it, and the PR body still reads "ungating means setting it from outside or issuing the existing RPC. Neither needs an edit to ATOM's serving path", with no clause about the profiler.

Not treated as the two-cycle halt trigger, stated so it can be overruled: the remedy round 1 asked for was a handoff clause, and the handoff is a closing comment on #70, which is open. So the finding has not survived a cycle in which its remedy was due. It is due now: one sentence in the PR body or in #70's handoff.

The review question: is stating an open hole in a docstring a close, or a guess wearing a disclaimer?

For route 1 it is a close. Refusing rather than guessing forbids answering without a basis; it does not forbid accepting a legal input. A one-row BatchView is a real batch (a single-request step is the common case, and five of the six steps in this PR's named result), and the backend prices exactly the rows it was handed. Refusing one row would decline a true shape. The collapse can only be committed by the projection, which #70's file set does not give this PR, and the statement carries its measurement and an asserting test.

What is missing is the schedule, not the honesty. A named gap becomes a close when it becomes work somebody can claim. Round 1 made that argument about the species word, it was accepted, and #87 is the result. It was not applied here: finding 3.

For route 4 it is weaker, closer to a disclaimer. The type already refuses two of the three ways a rung can disagree with its rows, which concedes that a rung is checked against the rows; the third turns one decode row into any duration and alone is handled by a paragraph. No test measures it, and the provenance carries no qualifier although this file has CANDIDATE for a count less certain than the reading behind it: finding 4.

Checked, with measurements

File list from git diff --name-only 669dc3f9d...28592b983, not a glob: backends/{__init__,shape}.py, design/{01,02,03,14}*.md, tests/compass/test_backend_{interface,shape_stub}.py.

route result
1. one-row BatchView legal: Sum N_Q.N_KV 1 048 576 vs 524 288, 284.389 µs vs 280.194 µs
2. BatchView subclass overriding prefill refused at class creation
3. RequestShape subclass overriding cached_tokens refused
3b. subclass replacing __post_init__ refused
4. capture_rung=10**9, one decode row legal: graph_padding 99 999 999 900, term 99.9999999 s, step 100.0000223 s
control: subclass adding a field, redefining nothing accepted, sums still from the rows

The guard is complete, not a listed subset: BatchView protects {requests, capture_rung, prefill, decode, graph_padding, __post_init__} and RequestShape protects {query_tokens, context_tokens, decode, cached_tokens, __post_init__}; all six probes refused. Better than round 1 asked for.

step.seconds_by_species()             -> {'fitted': 0.00014931072}
ProvenanceMix().seconds_by_species()  -> {'fitted': 0.00014931072}
fraction of this backend's seconds reported as fitted -> 1.0

#87 is open, unlabelled, and carries the measurement and three options. The bound on the harm: grep -rn ProvenanceMix --include=*.py outside backends/cost.py matches only the package re-export and two test files, so no production code reads a species mixture today.

Doubling recomputed: prefill.step 1.000000, prefill.tokens 2.000000, prefill.query_square 4.000000, total 1.883608; the linear regime gives long == 2 * short exactly.

Named result by closed arithmetic, independent of this PR's helpers: with a=2e-5, b=5e-7, c=2e-11, d=4e-12, a 2048-token prompt at a 256-token cap, step k costs 1.4931072e-4 + 2.62144e-7.(k-1), so the clock entering step 6 is 749.17504 µs with the cross term and 746.5536 µs without; the 748.00 µs arrival falls between. At a 512-token cap the clock entering step 4 is 846.874368 µs; at a flat 250 µs, 750.000 µs. Cross term 2.62144 µs against 749.17504 µs = 0.3499%.

The four inline replies land where they claim: the guards exist and refuse; the softening is in the module docstring, the class docstring and the body; #87 exists; the CostRefused asymmetry is in estimate's docstring. Reservation 3 is closed. No design-document identifiers in the four code files; ruff check and ruff format --check clean.

Accepted, and not to revisit in round 3
Drift and gate
  • git merge-base --is-ancestor 669dc3f9d 28592b983 -> PASS; the merge-base is 669dc3f9d.
  • git merge-base --is-ancestor d175b03c6 28592b983 -> FAIL. The integration head moved to d175b03c6, so "the merged tree is the branch tree" is no longer true. It was true of the ref it named; the PR body states it without one. Phrase it as of a named ref.
  • git merge-tree --write-tree d175b03c6 28592b983 -> rc=0, no conflict: staleness, not a disagreement.
  • The drift touches scripts/compass/{gate_cpu.sh,_lib.sh,snapshot.sh} and adds two test files, so the gate on the current head is a different instrument; the table below is against the stated base.

Each tree gated with its own scripts/compass (identical tree object ddb69e7aa on both sides), node 18, xiaobizh_n18_cpu:

tree commit passed skipped xfail rc
control (stated base) 669dc3f9d 4496 149 3 0
this branch 28592b983 4541 149 3 0

+45 passed, 0 skipped, 0 xfail; skips and xfails identical, so the ±1 flake did not fire. 42 tests in test_backend_shape_stub.py plus test_backend_interface.py 10 -> 13. gpu: not required (.compass-changed stamp); none of the eight paths is in gpu_gate_triggers.txt. The round-2 table reproduces exactly.

Effort

AST = parse, strip docstrings, unparse, count non-blank; __init__.py and test_backend_interface.py as deltas against 669dc3f9d.

AST statement lines physical non-blank
production (shape.py 144 + __init__.py +1) 145 425
test (test_backend_shape_stub.py 275 + test_backend_interface.py +6) 281 533

Both reproduce the PR body. Against #70's 300-450 envelope: production AST 145 below, production physical 425 inside, the pair by AST 426 inside, the pair physical 958 outside. Against #70's stated unit ("200 LOC, production, physical non-blank lines, excluding tests") production is 2.13x: a halt-and-discuss event. The developer flagged it rather than trimming prose. Not decided here: the instrument question is escalated to the owner on #89, which carries need human. I have not applied need human to this PR: the escalation exists, and labelling here would stop a review loop one small round from done. The main agent should decide whether to label.

The round-2 growth was +63 physical non-blank on shape.py, almost all prose round 1 asked for; that is why every finding here is scoped to a text substitution or an issue.

Watch next, and not checked
  1. Own the projection (finding 3): the runner landed without it. It belongs wherever the scheduled batch is held, cross-checked against Scheduler.compute_detailed_aggregates.
  2. capture_rung needs a bound or a qualifier, and a test either way.
  3. Species has no word for declared coefficients, so a stub that measured nothing reports as fitted #87 decides a closed vocabulary; until it lands a second backend will copy FITTED.
  4. The single global quadratic coefficient is still wrong for the hybrid target; the repair is a term per kind.
  5. The zero-priced-term collision in the ladder is unreached; the second source will meet it.

Not checked: the GPU tier (not required; scope, not a gap); the round-1 gate rows (the tip moved twice past them); accuracy (nothing here to be accurate about, and every row says so).

root and others added 2 commits September 22, 2026 00:59
… schedule that moves with it (M1-2)

A stand-in cost backend whose price is a declared linear form over the shapes
a scheduled batch already carries, replacing the alternative of two constants.
Constants price every step the same, so nothing downstream can depend on the
price and no run can show whether it did; the prior effort measured what that
hides, at +47.9% time-to-first-token once shapes varied.

The form is a launch intercept, a token term, a quadratic query term and a
query-by-history cross term for prefill; an intercept, a request count, a
context sum and the graph rung's padding rectangle for decode; plus one term
per collective the deployment widths admit, charged as a candidate and saying
so in its own provenance because two of the conditions that gate one are not
widths. Constant pricing is retained as a coefficient set with every shape
coefficient zero, which needs no branch in the estimator and leaves the zeros
visible in the breakdown.

Three properties are structural rather than asked for in prose. The attention
sums are computed from one row per request and the projection has no field
that could carry a collapsed scalar instead. The projection type lives in this
package, so the scan that forbids engine imports covers it, and a test fails
if it moves out. And every term's provenance, plus the backend's own
description, says the coefficient was declared rather than measured and is not
an accuracy claim -- in the record an artifact writes, not only in a docstring.

The tests drive ATOM's own scheduler, chunked prefill and admission included,
with a virtual clock advanced by nothing but these predictions, and check the
batch it builds against what the form predicts.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…m, and state what the type does not close

Review round 2 on #81. The module claimed the collapsed form "cannot be handed
in even by accident". Four routes to it were measured. Two are closed here and
two are stated rather than claimed away.

Closed at the type. `_refuse_shadowing` gives `BatchView` and `RequestShape`
the `__init_subclass__` guard `StepCost` in this package already has: a
subclass that redefines a reader can answer from a field it added rather than
from the rows, and it passes the backend's `isinstance` check. Overriding
`prefill` synthesises one row holding a whole batch's tokens against a whole
batch's history, which is the collapsed form exactly; overriding
`cached_tokens` fabricates the cross term directly. `__post_init__` is
protected by name as the only non-public member whose replacement also changes
what the object reports.

Stated, not closed. A one-row `BatchView` is `tokens x history` by
construction, and nothing in the type ties the row count to the number of
requests a scheduler scheduled -- the projection does that, and the projection
is not in this package yet. `capture_rung` is a batch-level scalar field,
correctly so because a rung is a property of the replayed graph rather than of
the rows, but it is unbounded above and enters the price linearly. The
docstrings now say the property that is true -- every shape sum is a function
of the rows -- rather than the stronger one that is not.

The refusals here are `TypeError` and `ValueError` where the seam names
`CostRefused`; `estimate` now records that a resolver wrapping this backend
will not catch them.

Design text, where a flag in a PR body would have been lost. D12's Option B
bullet said "double the chunk size, the step time doubles", which the form the
same section mandates forbids: measured, the intercept is 1.00x, the token
term 2.00x, the quadratic 4.00x and the total 1.8836x, and the total doubles
only for a = c = d = 0. The bullet now states the term-by-term property with
those figures. The stale `scheduler.py:790-792` citation for the three
detailed aggregate fields is corrected to `801-803` in all four design files
that carry it, and D12's gate and method spans to `:2819-2820` and
`:2788-2841`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… not, and fix the coefficient a document quotes

Round 3 of the review loop, against the five inline findings on the round-2
head plus the round-1 finding left open.

The module docstring claimed subclassing was refused outright. It is not:
shadowing a reader is, and a subclass that adds a field and redefines nothing
is accepted with its sums still running over its rows. The sentence now says
that, and a test asserts the accepted case so the corrected claim is pinned
rather than restated.

The rung asymmetry is argued rather than noted. Two ways a capture rung can
disagree with the decode rows are refused because the rows settle them: a rung
with no rows to pad, and a rung narrower than the rows it pads. A rung far
wider than the rows is not refusable from the rows -- only a ladder of
captured widths could refuse it, and this backend has none -- so it is priced
at face value and the padding term's own provenance now says the width was
supplied and nothing here bounds it. Two tests pin both directions, so
bounding the rung later breaks a test instead of leaving the paragraph stale.

The design text carried a coefficient three orders of magnitude from the
backend's. D12's doubling bullet said c=2e-8, which gives 3.7834x, not the
1.8836x printed beside it; the backend's c is 2e-11, a=2e-5 was omitted
although the total depends on it, and neither absolute step time was stated.
All three fixed. The span of compute_detailed_aggregates was re-measured --
def at scheduler.py:2788, last statement at :2841 -- and the two documents
still citing :2788-2842 now agree with the one that was corrected.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@jgong5

jgong5 commented Sep 22, 2026 •

Copy link
Copy Markdown
Owner Author

Round 3 pushed 0fc02f458 by force-push after a restack; #120, stacked on this branch, should restack from the old head 28592b983:

git fetch fork
git rebase --onto 0fc02f458 28592b983 <your-branch>

A plain git rebase conflicts in backends/__init__.py on this branch and --onto does not (three times of three). #120's branch is untouched; if this branch moves again, the SHAs go here the same way.

What moved
commit
old head (what #120 branched from) 28592b983631ba558586044317810d95287689cf
new head 0fc02f458344dadd1b0b8b7bea4885cfa92a02c7

First a restack onto the integration head: the old base 669dc3f9d had been overtaken by d175b03c6, whose drift touches scripts/compass/{gate_cpu.sh,_lib.sh,snapshot.sh}, so the gate on the old base is a different instrument.

git rebase --onto d175b03c6 669dc3f9d compass/m1-2-cost-stub     # rc=0, clean

Then one new commit, 0fc02f458, applying the round-3 review. The diff against d175b03c6 is nine files, none #120 is likely to hold: atom/compass/backends/{__init__,shape}.py, atom/compass/design/{01,02,03,09,14}*.md, tests/compass/test_backend_{interface,shape_stub}.py.

@jgong5

jgong5 commented Sep 22, 2026 •

Copy link
Copy Markdown
Owner Author

Round 3 pushed 0fc02f458, restacked onto d175b03c6. Fixed: findings 1, 2, 3 (#124), 5, and round-1 finding 3; finding 4 argued and pinned with two tests. Two deviations to rule on: #124 is not cited in shape.py, and 09_fitting_and_law_selection.md was edited outside the file set. #123 not taken. Gate: 4647 passed against 4599, rc 0. No blocking issues in the code; the effort overrun (2.22x) stays open.

Findings, and where each reply is

Round-1 finding 3 (the profile RPC): its remedy was a handoff clause, and #70's handoff is written on closure, which is why it survived two rounds without falling due. Closed now rather than deferred again, in three places: the PR body's dev record (the two ungating paths are not interchangeable), the comment above PUBLISHING in test_backend_shape_stub.py (it used to call profile_active "the public flag the profiler RPC flips"; it now says to assign the attribute and not the RPC, and why), and exit criterion 2 of #124. It still goes into #70's closing handoff. Nothing in ATOM's serving path was edited.

#123 not taken: the tensor-parallel collective is charged per paged layer (geometry.layers) where the all-reduce runs on the whole stack, 16 against 64 on the hybrid and invisible on a dense model. It is real and filed; this is the loop's last cycle, the correct count has no referent until #120 lands stage_layers, and a seventh finding is more likely to halt this PR than improve it.

Not done: not merged or undrafted; no need human; no reviewer spawned; no enum member or 12_open_items.md row (#87, #124); 04_model_capture_and_cost_ir.md untouched.

Restack
git rebase --onto d175b03c6 669dc3f9d compass/m1-2-cost-stub     # rc=0, replayed clean
git merge-base --is-ancestor d175b03c6 0fc02f458                 # PASS
git merge-tree --write-tree d175b03c6 0fc02f458                  # rc=0, no conflict

As of d175b03c6, this branch is a descendant of the integration head, so the merged tree is the branch tree; that expires the next time the head moves. The head was re-read at the moment of the restack; it had not moved. Force-push for the restack only; no gh stack; old and new heads posted for #120 in the previous comment. git diff --name-only d175b03c6...0fc02f458 is nine files: the four code files and design/{01,02,03,09,14}*.md.

Gate on 0fc02f4

Node 18, xiaobizh_n18_cpu, each tree gated with its own scripts/compass (identical tree object ba42b43d8 on both sides):

tree commit passed skipped xfail rc
control (integration head) d175b03c6 4599 149 3 0
this branch = merged 0fc02f458 4647 149 3 0

+48, by collection on each tree: test_backend_shape_stub.py 0 -> 45, test_backend_interface.py 10 -> 13. The three new tests are the accepted-subclass control and the two rung tests. The two files: 58 passed in 0.25 s. Skips and xfails identical, so the ±1 flake did not fire. gpu: not required (.compass-changed stamp). ruff check and ruff format --check clean on the four code files. The round-2 rows (4496 / 4541) are superseded: the base moved and its drift changed the gate scripts.

Effort on 0fc02f4
AST statement lines physical non-blank
production (shape.py 145 + __init__.py +1) 146 443
test (test_backend_shape_stub.py 294 + test_backend_interface.py +6) 300 572

Under #70's unit (production, physical non-blank, excluding tests): 362 (1.81x) -> 425 (2.12x) -> 443 (2.22x). More than ~2x is a halt-and-discuss event; reported, not trimmed. This round's +18 production physical non-blank, all from findings 2 and 4:

hunk physical non-blank
the module docstring paragraph: the corrected shadowing sentence and the rung argument, one rewrite +9
UNCHECKED_RUNG and the comment saying what it qualifies +4
the padding term carrying that qualifier only when a rung was supplied +5
the counted annotation widening to admit a qualifier 0

Finding 4's remedy was "argue it explicitly and pin it with a test"; an argument is prose. No comment or docstring was trimmed to offset it. By AST it is 0.73x, so the instrument decides the verdict: the open owner decision.

f"capture rung {self.capture_rung} is smaller than the "
f"{len(decode)} decode requests it pads, which would make the "
"padding negative"
)

@jgong5 jgong5 Sep 22, 2026 •

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note, not a finding: this refusal is right for a stronger reason than its message gives. A rung narrower than the decode rows need not make the padding negative; the real reason is that a graph captured at width 2 cannot replay three rows, the rows-settle-it argument the module docstring already uses. The check is unchanged across all three commits and accepted by two cycles, so it is recorded for #124, where the rung gets a real source.

Measured on 0fc02f4
rows contexts (1000, 1, 1), capture_rung=2
  REFUSED: capture rung 2 is smaller than the 3 decode requests it pads,
           which would make the padding negative
  but rung*max - sum = 2*1000 - 1002 = 998, which is NOT negative

@jgong5

jgong5 commented Sep 22, 2026 •

Copy link
Copy Markdown
Owner Author

APPROVE at 0fc02f458: the loop converges at cycle 3. All six findings are closed and nothing survived two cycles; round-1 finding 3 is closed in the tree, so the PR does not halt. Both deviations are upheld. No need human applied or warranted from this review; that call is the main agent's.

Checked: each finding re-measured; the restack replayed byte-identical; the gate, 4647 passed at 0fc02f458 against 4599 at d175b03c6; the effort figures.
Rulings: (A) do not write #124 into shape.py; (B) editing 09_fitting_and_law_selection.md outside the file set was correct.
Asks, not findings: correct #70's body, which still carries the three stale scheduler.py citations, before #70 closes; the round-3 comment's 3.783380 is 3.783371723. One inline note on a refusal message.
Accepted with reservation: the constructed 0.35%; no rung integration coverage; no CostRefused; FITTED (#87); #123 not taken.
Watch next: #124 closes both routes; set profile_active directly; #123 lands with stage_layers; #87; the rung message.

The six findings, re-measured

Range pinned to head 0fc02f458 against fork/feature/atomcompass_new = d175b03c6, re-read (not moved since the restack). File list from git diff --name-only d175b03c6...0fc02f458: nine files.

  1. D12's coefficient, closed. Coefficients() at 0fc02f458 is a=2e-05, b=5e-07, c=2e-11, d=4e-12. Priced through ShapeStubBackend.estimate on a single-request first chunk (RequestShape(n, n, False), so d cannot enter):
    prefill.step             20.000000 us ->    20.000000 us   ratio 1.000000
    prefill.tokens          128.000000 us ->   256.000000 us   ratio 2.000000
    prefill.query_square      1.310720 us ->     5.242880 us   ratio 4.000000
    prefill.query_cached      0.000000 us ->     0.000000 us   (d term, zero as claimed)
    total                   149.310720 us ->   281.242880 us   ratio 1.883608089
    
    D12 now reads a=2e-5, b=5e-7, c=2e-11, names a first chunk, and carries 149.31 µs -> 281.24 µs beside 1.8836x: every figure reproduces from the bullet itself. At c=2e-8: 1458.720000 µs -> 5518.880000 µs.
  2. The shadowing sentence, closed. Built the subclass myself:
    @dataclass(frozen=True)
    class Tagged(BatchView):
        label: str = ""
    Tagged((RequestShape(256,1024,False),)*2, label="x")   -> ACCEPTED
    sum_query_context(view.requests) = 524288   (== the plain BatchView, == 2*256*1024)
    price sub = 0.000280194304  plain = 0.000280194304   identical = True
    
    A RequestShape subclass adding a field is accepted the same way (cached_tokens still 768). All six shadowing probes still refuse at class creation (prefill, graph_padding, requests, __post_init__ on BatchView; cached_tokens, __post_init__ on RequestShape).
  3. Registered as compass(runner): the projection that builds a BatchView, and the rung it supplies #124, closed. compass(runner): the projection that builds a BatchView, and the rung it supplies #124 is open and unlabelled, carries both routes under one holder, and has the three exit criteria the round-3 record claims. git grep -n "BatchView\|RequestShape\|_project\|CostBackend" d175b03c6 -- atom/compass/runner/ -> no matches.
  4. capture_rung's third check, closed: argued and pinned, which is refusing rather than guessing correctly applied, not a disclaimer. (a) No check is available that would not be an invention: refusing a rung of 10^9 needs a bound, and the only honest source is the ladder of widths a graph was captured at, which this backend lacks; a guessed bound is a number with no source. The two refused directions are refused because the rows contradict them; both still refuse. (b) The uncertainty is in the record: UNCHECKED_RUNG reaches StepCost.rows() only when a rung was supplied:
    rung supplied:  fitted (1e-09 s x 99999999900; declared, not measured -- a plumbing figure,
                    not an accuracy claim; padding from a supplied rung width; nothing here
                    bounds it above) via shape-stub
    no rung:        fitted (1e-09 s x 0; declared, not measured -- a plumbing figure, not an
                    accuracy claim) via shape-stub
    
    (c) It cannot go stale silently: 99.9999999 s for the term and 100.0000223 s for the step at capture_rung=10**9, reproduced, and two tests assert each direction. The real close is the second half of compass(runner): the projection that builds a BatchView, and the rung it supplies #124.
  5. The anchors, closed, in atom/model_engine/scheduler.py at d175b03c6:
    2788  def compute_detailed_aggregates(
    2819      if not self.profile_active or not self._detailed_annotation_enabled:
    2820          return
    2841      scheduled_batch.detailed_sk = sk
    2842  (blank)
    2843  def _connector_flag(self, name: str) -> bool:
     801  self.detailed_sqsq  802  self.detailed_sqsk  803  self.detailed_sk
    
    At the head, 01, 02 and 09 all read :2788-2841; no 2788-2842 and no 790-792 remain anywhere in atom/compass/. The neighbouring anchors in 01 this PR did not claim also check out (ScheduledBatch, produces_output(), ScheduledBatchOutput).
  6. Round-1 finding 3, closed; the PR does not halt.
    engine_utility.py:252   result = self.runner_mgr.call_func("start_profiler", wait_out=True)
    engine_utility.py:256       self.scheduler.profile_active = True          <- after, not before
    model_runner.py :1067   def start_profiler(self, trace_name: str | None = None):
    model_runner.py :1075       if self.profiler_dir is not None and self.profiler is None:
    
    Two of the three places are durable without this PR body: the comment above PUBLISHING in tests/compass/test_backend_shape_stub.py and exit criterion 2 of compass(runner): the projection that builds a BatchView, and the rung it supplies #124, which names both call sites.
Rulings on the two deviations

A. #124 not written into shape.py: upheld; do not add it.

  1. The rule bans design-doc identifiers, and an issue is not one. AI_DEV_RULES enumerates them (decision ids, TODO ids, phase labels, backticked document numbers, "principle N", "Gate N"), and also says "nothing durable lives in the tree: the task record is the GitHub issue and its PR". Reading the ban to cover issue numbers puts the two halves in conflict.
  2. Precedent: git grep -E "#[0-9]{1,4}" over atom/compass/**/*.py and tests/compass/ at d175b03c6 returns nothing; Species has no word for declared coefficients, so a stub that measured nothing reports as fitted #87 is not cited in code either.
  3. A squashed PR body loses the pointer, but the prose names what closes each gap (the projection, built by whoever holds the scheduled batch; a ladder of captured widths), and the squashed commit carries its PR, which carries compass(runner): the projection that builds a BatchView, and the rung it supplies #124.

B. 09_fitting_and_law_selection.md outside the file set: upheld. AI_DEV_RULES makes a merge conflict the agent's call ("don't bother me on merge conflict, it's on you"). The edit is one character, and leaving one of three documents at :2788-2842 would re-create round-2 finding 5. No open PR touches 09: the file lists of all eighteen open compass PRs (#125, #120, #119, #118, #115, #109, #96, #95, #91, #86, #85, #75, #67, #66, #63, #62, #61, #59) were scanned.

Restack, gate, effort
git merge-base --is-ancestor d175b03c6 0fc02f458       -> PASS
git merge-tree --write-tree d175b03c6 0fc02f458        -> rc=0, no conflict
git log --oneline d175b03c6..0fc02f458                 -> 3 commits (6175bfb12, bfb246b5b, 0fc02f458)
git diff 669dc3f9d...28592b983  >  /tmp/o.diff     (the two pre-restack commits)
git diff d175b03c6...bfb246b5b  >  /tmp/n.diff     (the same two, replayed)
diff /tmp/o.diff /tmp/n.diff    ->  rc=0, 0 lines

The two original commits replayed without loss; 0fc02f458 is the only new one.

Node 18, xiaobizh_n18_cpu, each tree gated with its own scripts/compass (tree object ba42b43d8 on both sides), COMPASS_INTEGRATION_REF=fork/feature/atomcompass_new:

tree commit passed skipped xfail rc
control (integration head) d175b03c6 4599 149 3 0
this branch 0fc02f458 4647 149 3 0

+48, GATE_CPU_RC=0 both sides, by collection: test_backend_shape_stub.py absent -> 45, test_backend_interface.py 10 -> 13. The two files: 58 passed in 0.21 s. Skips and xfails identical, so the ±1 flake did not fire. ruff check and ruff format --check clean on the four code files; no design-document identifiers in them.

Effort: production AST 146, physical non-blank 443 (425 + 18); test AST 300, physical non-blank 572 (556 + 16). All four reproduce the PR body. 443 = 2.215x the 200; the series 362 -> 425 -> 443 reproduces; the +18 decomposes as the PR states. Production AST went 145 -> 146 across round 3: one statement line for eighteen physical ones (a four-tuple, one qualifier argument through _term, a widened annotation). The rest is the argument finding 4 asked for and the sentence finding 2 asked for. Not a finding; the unit is the owner's and escalated elsewhere. Two consecutive review cycles have pushed this number up by asking for prose.

Notes, reservations, watch next

Note A: the round-3 comment prints ratio 3.783380 for c=2e-8; through the backend it is 3.783371723 (5518.88 / 1458.72). The absolute times are exact, the PR body's 3.7834x is right, and D12 carries neither figure. Not a finding.

Note B: #70's body, in the "ATOM already computes the attention terms" paragraph, still reads scheduler.py:790-792, 2788-2842 and :2820-2821, the set this PR corrected in five design files. The developer checked #70's exit criterion, not this paragraph. An issue-body edit: no diff, no re-gate, no round. Worth doing before #70 closes, since #124 lists #70 as a predecessor.

Note C, inline: the rung refusal is right for a stronger reason than its message gives. Unchanged by round 3 and accepted by two prior cycles, so not a finding; recorded for #124.

Accepted with reservation:

  1. The 0.35% is a constructed margin; do not quote it as generic.
  2. Nothing sets capture_rung in the closed loop; compass(runner): the projection that builds a BatchView, and the rung it supplies #124's.
  3. CostRefused is never raised, and estimate's docstring says so; a Resolver will not catch the TypeError/ValueError.
  4. FITTED still aggregates to {'fitted': 0.00014931072} with no decomposition carrying the correction; Species has no word for declared coefficients, so a stub that measured nothing reports as fitted #87, the enum is the owner's.
  5. The stub charges the collective on paged layers, and the stack is four times deeper #123 not taken: the count has no referent until compass(backends): the stand-in model, its widths declared once (M1-3, #24) #120 lands stage_layers, and the issue says what a fix must show.

Watch next:

  1. compass(runner): the projection that builds a BatchView, and the rung it supplies #124 is the real close for both routes: the row count needs the projection, the rung needs the ladder. Neither is closable inside backends/.
  2. Set scheduler.profile_active = True directly; the RPC starts a profiler first.
  3. The stub charges the collective on paged layers, and the stack is four times deeper #123's paged-vs-stack layer count lands with stage_layers, not a coefficient.
  4. Species has no word for declared coefficients, so a stub that measured nothing reports as fitted #87 decides the vocabulary; until then a second backend copies FITTED.
  5. Note C, when compass(runner): the projection that builds a BatchView, and the rung it supplies #124 gives the rung a source.

Not checked: the GPU tier (gpu: not required (.compass-changed stamp); scope, not a gap); the round-1 and round-2 gate rows (the base moved twice and the instrument changed); accuracy (nothing here to be accurate about).

jgong5 pushed a commit that referenced this pull request Sep 22, 2026
Restacked onto #81's round-3 head and amended for the six inline findings of
round 1.

The three assertions that pinned the collective's layer count as a literal no
longer do. The count is recovered from the priced step -- the collective term
is tokens x layers x a coefficient -- and the step seconds and the ratio are
looked up in tables keyed on it, carrying both answers: 1.030496x charged per
paged layer as it is today, 1.121986x charged on the depth an all-reduce
actually runs on. The second row is priced rather than worked out in a
comment: a uniform 64-layer stand-in charges what the published hybrid would
if the charge moved, and the same read-back recovers 64 from it. A count that
is neither is refused by name. So when the charge moves the assertions move
with it instead of going red with nothing in them to say why.

The seam test's docstring claimed a route it does not close. Measured: a
config type defined beside the engine reaches FakeModel through its
duck-typed config argument and neither the import scan nor this test fires.
The docstring now claims only what the test does -- these types are defined
in files the scan reads, and the case only this test catches is a move to a
package that is not atom.* at all, which the scan skips by its own filter --
and names the duck-typed route as open, in the module docstring as well.

head_dim now defaults to unset, which reaches the reader's one fallback and
makes hidden_size decide something: unset, the block is hidden_size over the
query head count, and 4096 and 8192 no longer produce an equal geometry. A
dial that is not a whole number is refused by name rather than by a bare
comparison error, and an unset head_dim over too few hidden units is refused
by name too.

stages() checks that the spans partition the stack, not only that there are
as many as there are stages. Four identical spans, four spans that tile a
quarter of the model, and a gap between two spans were all accepted; each of
them sizes pools that hold a fraction of the KV with nothing saying so.
Sorted, the spans must start at 0, meet end to start, and end at the declared
depth.

The layer-kind spellings are one declaration again: the dial reads them from
the module that decides what they mean rather than repeating two string
literals across a module boundary. geometry.py is a landed PR's file and is
not touched.

Decode-per-rung is filed as #128 rather than left in a PR body a squash
erases, and #24 now carries a named result.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@jgong5
jgong5 added this pull request to stack #138 September 22, 2026 01:59
@jgong5

jgong5 commented Sep 22, 2026 •

Copy link
Copy Markdown
Owner Author

No blocking issues: APPROVE at 0fc02f458 (review) stands after a pin re-verification. The PR is still a draft, so it cannot land. Not a review cycle or a delta review.

Checked: the published counts reproduce (the two files 58 passed; full gate 4647 passed, 149 skipped, 3 xfailed, rc 0). Both fixes made under review have pins that fail when the real pre-fix code is put back. Of 37 mutants that introduce a defect, 31 bite. Five leave the whole gate green, and a pre-existing sibling test catches the sixth. 0fc02f458 is one commit behind 92f1fdafe; the merge is clean and the files are disjoint.
Accepted with reservation: RequestShape.__post_init__'s refusals and the row type check have no witness; many shape-stub properties have exactly one test; UNCHECKED_RUNG and CANDIDATE are pinned only through themselves.
Watch next: test_the_charged_collectives_are_the_ones_the_widths_produce's self-derived assertion pins naming and order, not membership.

State and staleness
item value
head 0fc02f458344dadd1b0b8b7bea4885cfa92a02c7 (compass/m1-2-cost-stub)
base feature/atomcompass_new at d175b03c6ed9f05222ca66bb128cb622363d0cc8: API .base.sha, git merge-base, and the round-3 restack target agree
draft true
state / mergeable / labels open / true / none
ahead / behind 3 commits ahead, 1 behind

Integration is now 92f1fdafe ("compass: withdraw the capture module and its tests" (#129)). git merge-tree --write-tree 92f1fdafe 0fc02f458 -> tree 4e75c33c4, rc 0, no conflict markers; git diff --name-only d175b03c6 92f1fdafe touches atom/compass/capture/*, design/04, design/12 and three tests/compass/test_capture_* files, disjoint from this PR's nine. The restack is still valid.

Method

Branch tree staged into xiaobizh_n18_cpu on node 18, gated with its own scripts/compass (tree object ba42b43d81cc053b920f5b554611e062788001ce on branch and base). Every run bounded by a timeout; GATE_CPU_RC read from the captured file. One mutation at a time on one tree, restored from a pristine copy between runs, __pycache__ cleared before each; every mutation called invisible was then run through the full CPU gate on its own copy.

Harness conditions, and what each did:

  • A mutant that would change a file's line count is refused: fired on X01 (1 -> 2 newlines).
  • Each target string must be unique: fired on X04 (occurs 4 times) and on X02/X03 (occurs 0 times). X02 was meant as the duplicate case and turned out absent; X04 is the deliberate duplicate. Four refusals, four runs that never happened.
  • Null controls: ten null runs gave 58 passed, rc 0, every time; two inert mutants through the same write path, N01 (a comment appended to c = self.coefficients) and M38 (collectives() * 1), gave 58 passed, rc 0.

Full gate on the unmutated tree: 4647 passed, 149 skipped, 3 xfailed, GATE_CPU_RC=0, 39.27 s, matching the approving review. Seven full-gate runs; the only excursion was the mutant meant to fail.

Denominators: 58 collected = 45 in tests/compass/test_backend_shape_stub.py + 13 in tests/compass/test_backend_interface.py (10 before; +2 projection parameters, +1 path parameter for shape.py). No empty parametrize: path collects 7 cases, projection 2, checked by node id.

A. The two fixes made under review, pre-fix state recovered from this PR's commits

Recovered with git show <sha>:atom/compass/backends/shape.py and reinstated by removing exactly the text the fixing commit added (byte-for-byte, asserted before the write).

pin defect reinstated reinstated node id / assertion verdict
test_a_projection_subclass_that_redefines_a_reader_is_refused RA: the shadowing guard as at 6175bfb12 (-43 lines) 54 passed / 4 failed / rc 1 :188 with pytest.raises(TypeError, match="functions of the rows") -> DID NOT RAISE bites
test_a_row_subclass_that_redefines_a_reader_is_refused same 54/4 :201 -> DID NOT RAISE bites
test_the_guard_names_what_the_subclass_redefined same 54/4 :211 -> DID NOT RAISE bites
test_dropping_the_row_checks_is_refused_too same 54/4 :225 match="__post_init__" -> DID NOT RAISE bites
test_a_subclass_that_redefines_nothing_is_accepted same passes none inert by construction: it pins the non-refusal. Correct as written
test_a_rung_no_ladder_would_hold_is_priced_and_labelled RB: the decode.graph_padding entry as at bfb246b5b, verbatim (-5 lines) 57 passed / 1 failed / rc 1 :340 assert UNCHECKED_RUNG in padding["decode.graph_padding"] bites
test_with_no_rung_the_padding_term_carries_no_rung_qualifier same passes none inert against this reinstatement (the pre-fix state also lacks the qualifier); bites the inverted qualifier, M27

Published count 58 passed in every row. Two whole-blob reverts, RA_full (shape.py at 6175bfb12) and RB_full (at bfb246b5b), give 0 passed, 1 collection error, rc 2: the head test file imports UNCHECKED_RUNG, which those blobs lack. A total failure proves nothing about any one pin, so neither is counted.

B. Every other pinned behaviour, one mutant at a time

42 mutants; published 58 passed / rc 0 for all.

# defect introduced counts first node id and assertion that fired verdict
M01 token term counts rows, not tokens 50/8 test_a_prefill_step_is_its_four_named_parts :91 assert named(step) == {…} (+7 more, incl. all four closed-loop tests) bites
M02 Σ N_Q² -> (Σ N_Q)², the collapsed form 50/8 test_the_sums_are_the_ones_the_scheduler_publishes :291 assert sum_query_square(rows) == batch.detailed_sqsq -> 64 == 26 bites
M32 _per_request evaluates one row and multiplies 49/9 :567 assert sum_query_context(view.requests) == batch.detailed_sqsk -> 786432 == 458752 bites
M03 cross term counts N_KV, not N_KV - N_Q 56/2 test_the_cached_cross_term_is_a_difference_of_two_published_sums :298 bites
M04 cached_tokens returns the whole context 56/2 :298 same bites
M05 padding is the batch's rectangle, not the rung's (the defect the module docstring names) 55/3 test_the_padding_is_the_rungs_rectangle_not_the_batchs :311 assert view.graph_padding == 16*1000-1400 -> 1600 == 14600 bites
M06 an empty batch is priced, not refused 57/1 test_a_step_with_no_requests_is_refused :127 -> DID NOT RAISE bites
M07 estimate accepts anything 57/1 test_something_that_is_not_a_projection_is_refused :444 -> AttributeError: 'dict' object has no attribute 'prefill' bites
M11 negative / infinite coefficients admitted 57/1 test_a_coefficient_that_is_not_a_duration_is_refused :140 -> DID NOT RAISE bites
M12 a rung on a prefill-only step admitted 57/1 test_a_step_that_replayed_nothing_has_no_rung :315 -> DID NOT RAISE bites
M13 a rung narrower than its rows admitted 57/1 test_a_rung_narrower_than_the_batch_is_refused :319 -> DID NOT RAISE bites
M14 a successor bounds the rung above (the check the round-3 argument says must not be added) 56/2 test_a_rung_no_ladder_would_hold_is_priced_and_labelled -> ValueError: capture rung 1000000000 is smaller than the 1 decode requests it pads; also test_a_rung_narrower_than_the_batch_is_refused :319 bites: the argument is enforced, not just written
M15 #123-shaped: the collective charged on a different layer count (layers*4) 57/1 test_the_charged_collectives_are_the_ones_the_widths_produce :437 -> 8.192e-06 == 2.048e-06 ± 2e-12 bites; the gate-level failure #120's audit saw, on the second assertion
M16 collective charged per token, not per token-layer 57/1 same test, :437 -> 1.28e-07 == 2.048e-06 bites
M17 a collective priced with no geometry 57/1 test_a_collective_with_no_geometry_to_price_it_is_refused :442 -> DID NOT RAISE bites
M18 a charged collective drops the candidate qualifier 57/1 test_a_charged_collective_says_it_is_a_candidate :459 bites
M19 DECLARED leaves the record 56/2 test_every_term_says_the_coefficient_was_declared :393 bites
M20 species claims ANALYTICAL 57/1 test_the_species_is_not_analytical :407 bites
M21 tier claims OP_LEVEL 57/1 test_the_backend_answers_at_step_granularity :133 bites
M22 prefill terms named on a decode-only step 57/1 test_a_decode_step_is_its_four_named_parts :103 bites
M23 decode terms named on a prefill-only step 51/7 test_a_step_names_only_the_groups_it_ran :117 bites
M24 constant pricing no longer reports itself blind 56/2 test_constant_pricing_is_a_coefficient_set_not_a_mode :361 assert flat.shape_blind bites
M25 constant() puts the seconds on the token coefficient 54/4 :359 assert short.seconds == long.seconds == 0.05 -> 0.8 == 806.4 bites
M26 describe() names the other pricing 57/1 test_the_description_says_it_and_which_pricing_produced_it :411 bites
M27 the unchecked-rung note goes on batches that supplied no rung 56/2 :340 and :346: both halves of the pair fire bites
M28 an all-to-all charged at one data-parallel rank 57/1 test_a_single_rank_deployment_is_charged_for_no_collective :428 bites
M30 tp-all-reduce never produced 56/2 same collectives test, :437 -> KeyError: 'collective.tp-all-reduce' bites (at 437, never at 436)
M35 the step total is a plain sum, not the fold 57/1 test_the_total_is_folded_from_the_parts_and_not_stored :124 -> 0.00019399999999999997 == 0.000194 bites; the mutated file, backends/cost.py, is pre-existing, not this PR's
M36 the charged term named collective_x instead of collective.x 56/2 :436 assert charged == [f"collective.{n}" …] bites
M37 the charged collectives in the opposite order 57/1 :436 bites
M39 the DECLARED constant's text reversed to claim measurement 57/1 :394 assert "not an accuracy claim" in provenance, the literal, not the constant bites, by one literal only
M08 BatchView rows need not be RequestShape 58/0 none blind spot
M09 a zero-token request admitted (query_tokens < 1 dropped) 58/0 none blind spot
M10 context may be shorter than the query 58/0 none blind spot
M29 #140's case: moe-all-to-all never produced 58/0 none inert here; a sibling catches it
M40 the UNCHECKED_RUNG text reversed to say the width was checked 58/0 none blind spot
M41 the CANDIDATE text reversed to say the collective was observed 58/0 none blind spot
N01, M38 inert edits (null controls) 58/0 none null, as intended

Of the 42 mutants, 37 introduce a defect (the other five are the three tripwires below and the two null controls); 31 bite and 6 leave both files green. Over 43 mutants and 4 recovered pre-fix states (2 discarded as total failures), every pin this PR presents as guarding a specific behaviour bites, and none of the six blind rows is a pin this PR claims.

C. What no test in the tree catches, and the one a sibling does

Each run through the full CPU gate on its own copy; control on the same tree and procedure: 4647 passed, GATE_CPU_RC=0.

mutation full gate sibling
M08 row type check dropped 4647 passed, 149 skipped, 3 xfailed, GATE_CPU_RC=0 none
M09 query_tokens < 1 refusal dropped 4647 passed, rc 0 none
M10 context_tokens < query_tokens refusal dropped 4647 passed, rc 0 none
M40 UNCHECKED_RUNG text reversed 4647 passed, rc 0 none
M41 CANDIDATE text reversed 4647 passed, rc 0 none
M29 moe-all-to-all silenced 1 failed, 4646 passed, GATE_CPU_RC=1 tests/compass/test_backend_kv_geometry.py::test_expert_parallelism_alone_builds_no_all_to_all, pre-existing; it bites because it asserts against literal tuples rather than collectives()

Notes on reach, not findings against the approved change:

  1. RequestShape.__post_init__ is pinned as a member, not as a behaviour. test_dropping_the_row_checks_is_refused_too proves a subclass may not drop it; nothing proves it does anything. Both of its ValueError refusals (M09, M10) and the isinstance(row, RequestShape) check (M08) can be deleted with the full gate unmoved: refusals with no witness. A pin over the guard is not a pin over the thing guarded.
  2. RequestShape, BatchView and ShapeStubBackend are named in exactly two files in the tree, tests/compass/test_backend_shape_stub.py and tests/compass/test_backend_interface.py, so every shape-stub property is single-sourced by construction. Eleven mutants (M06, M07, M11, M12, M13, M17, M18, M20, M21, M26, M28) each fail exactly one test.
  3. UNCHECKED_RUNG and CANDIDATE appear on both sides of every assertion that mentions them, so a qualifier saying the opposite passes the whole gate (M40, M41). DECLARED escapes only because :394 also asserts the literal "not an accuracy claim": one literal anchors the disclaimer this backend exists to carry. One more literal beside each of the other two would close all three; a note, not an ask.
D. Assertions that read the same expression on both sides; E. tripwires; F. prose

By AST over both files, with local single-assignments inlined three levels, checked against what 43 mutants did:

  • Literally identical on both sides: 0 of 74 assert statements (65 in the shape-stub file, 9 in the interface file).
  • Exactly 1 derives both sides from the same source: test_the_charged_collectives_are_the_ones_the_widths_produce, line 436, assert charged == [f"collective.{n}" for n in widths.collectives()]. The backend builds the left side by iterating self.parallelism.collectives(), the identical call on the identical object. This PR's own, pre-existing; compass(backends): charge the collective on the layers a worker runs #140 only moved a count beside it. It passes on every change to which collectives exist (M15, M16, M28, M29, M30 leave it green) and catches only M36 (name format) and M37 (order): a naming-and-ordering pin, not the membership pin its name suggests. Membership is held by line 437's literal 128 * 16 * Coefficients().collective_token_layer (M15, M16, M30) and by the sibling in test_backend_kv_geometry.py (M29).
  • 16 further == assertions share a sub-expression across the sides with different arguments (the chunk-size ratios, the two collapse-alike batches, the two constant-pricing runs, observed-vs-predicted break): relational properties, confirmed by M01 at :476, M02 at :291/:298, M23 at :501, M32 at :567.

Tripwires (raise replaced by a branch-local 1 / 0):

tripwire reached by
the shadowing refusal branch (if shadowed:) 4 tests, all four of section A's guard pins
the empty-batch refusal (if not rows:) exactly 1, test_a_step_with_no_requests_is_refused
the missing-geometry refusal (if named and geometry is None:) exactly 1, test_a_collective_with_no_geometry_to_price_it_is_refused

No refusal branch in this PR's file is unreached.

Prose: nothing in the tree reads a docstring (grep -rln "__doc__" tests/ atom/compass/ returns nothing, #243), so every explanatory paragraph in shape.py and the test module, including the module docstring's worked figures, is unpinned by construction. A property of the tree, not of this PR.

@jgong5

jgong5 commented Sep 23, 2026 •

Copy link
Copy Markdown
Owner Author

Blocked on #89: need human applied, because this PR overran its estimate by more than 2x.

Under AI_DEV_RULES.md as landed in #137, that overrun is an escalation, and "when the ruling lives on another issue, label each PR it holds and name that issue". #137's body decides #250 as option 1, which is this labelling, and the owner accepted that by landing #137.

The label also holds every PR stacked above this one. No agent will land, amend or review this chain until the owner rules on #89 and removes the label.

Base update under the need-human exception in AI_DEV_RULES.md: the branch
conflicted with the integration tip in one file.

tests/compass/test_backend_interface.py: both sides kept. The branch's
reformatted import assertion and its projection-location test stay; the
tip's test_every_module_the_walk_returns_is_a_case is appended after them.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@jgong5

jgong5 commented Sep 24, 2026 •

Copy link
Copy Markdown
Owner Author

Blocked on #89; base update pushed 58cb9dcce, merging fork/feature/atomcompass_new at c92e4c1c1 into 0fc02f458 under the base-update exception in AI_DEV_RULES.md. One conflict, in tests/compass/test_backend_interface.py, resolved keeping both sides. No rebase, amend or force-push; base unchanged; the label stays. Gate: 5329 passed, rc 0; +48 node ids over the tip, none removed.

Resolved file

tests/compass/test_backend_interface.py: one conflict hunk, at the end of test_the_package_imports_nothing_from_the_engine.

  • This branch reformatted the assert module.startswith("atom.compass.backends"), (...) line and added test_the_projection_type_is_where_the_scan_can_reach_it. Both kept as the branch had them.
  • The tip added test_every_module_the_walk_returns_is_a_case, appended after the branch's test, unchanged.
  • The tip's other edits to this file merged without conflict: the _backend_modules() helper, its three guard tests, and the parametrisation that calls it.

The four design docs merged cleanly. The remerge diff (git show --remerge-diff 58cb9dcce) contains only this hunk.

Dropped-change check

git diff c92e4c1c1 58cb9dcce against the PR's own diff at its old merge-base, git diff d175b03c6 0fc02f458, with index and hunk-header lines stripped: both touch the same 9 files and differ only in two context lines, both the tip's own:

  • In 02_model_runner_and_cost_backend.md, one context line now reads `forward(batch: ScheduledBatch) -> ScheduledBatchOutput` where it read `model_runner.py:3233-3320`.
  • The tip's new test appears as trailing context.
Gate

scripts/compass/gate_cpu.sh from each tree itself, node 18 (xiaobizh_n18_cpu), with .compass-commit and .compass-changed stamps; atom resolved under each staged root.

tree commit passed skipped xfailed failed verdict
tip c92e4c1c1 5281 155 3 0 GATE_CPU_RC=0 PASSED
old head 0fc02f458 4647 149 3 0 GATE_CPU_RC=0
new head 58cb9dcce 5329 155 3 0 GATE_CPU_RC=0 PASSED

Node-id delta, from junit XML:

  • New head vs tip: 48 added, all passing: this PR's own test_backend_shape_stub classes, plus test_backend_interface's [shape.py] case and the two projection-location cases. None removed, no outcome changed.
  • New head vs old head: 743 added, 55 removed, no outcome changed. The 55 are exactly the node ids the tip no longer collects (test_capture_*, test_clock_* and five others); none is this PR's.

GPU tier not required, per the .compass-changed stamp.

@jgong5

jgong5 commented Sep 24, 2026 •

Copy link
Copy Markdown
Owner Author

No blocking issues: the conflict resolutions of merge 58cb9dcce (parents 0fc02f458, c92e4c1c1) are sound. This is the one delta review the base-update exception allows. It covers only the resolutions, neither approves nor rejects the PR, and leaves the label untouched.

Checked: the merge keeps every change from both sides and changes nothing else; the remerge diff is one hunk; both dropped-change directions are clean; the merged file behaves as both sides intended; the gate logs agree.
Accepted with reservation: the gate's commit: line is the staging agent's stamp and the staged tree is gone, so it cannot be re-hashed; my own staged run agrees on this file.
Watch next: test_the_projection_type_is_where_the_scan_can_reach_it misses a scan narrowed to drop shape.py (M3 below). Pre-existing at 0fc02f458, not caused by the merge. Fix it in this PR's own review once the owner releases the label.

1. Remerge diff

git show --remerge-diff 58cb9dcce lists one file, tests/compass/test_backend_interface.py, one conflict hunk, byte-identical to the agent's remerge-81.diff. git merge-tree --write-tree 0fc02f458 c92e4c1c1 differs from the committed tree (ca5a0df82) only in that file, by 6 deleted lines: the three conflict markers and the tip's copy of the three-line assert.

  • The branch reformatted assert module.startswith("atom.compass.backends"), (...) and added test_the_projection_type_is_where_the_scan_can_reach_it.
  • The tip left that assert as it was at the merge-base d175b03c6 and appended test_every_module_the_walk_returns_is_a_case.

The tip's three assert lines are base context, so keeping the branch's formatting drops nothing of the tip's. Per file, diff(0fc02f458, 58cb9dcce) is exactly the tip's change diff(d175b03c6, c92e4c1c1), and diff(c92e4c1c1, 58cb9dcce) is exactly the branch's change diff(d175b03c6, 0fc02f458).

2. Dropped-change check, both directions

index and @@ lines stripped; merge-base d175b03c6.

Check Files --numstat Patch delta
The PR survives: git diff c92e4c1c1 58cb9dcce vs git diff d175b03c6 0fc02f458 9 = 9, same set identical in every file 8 lines, 0 of them +/-: all context
The tip survives: git diff 0fc02f458 58cb9dcce vs git diff d175b03c6 c92e4c1c1 105 = 105, same set identical in every file 12 lines, 0 of them +/-: all context

PR direction: the two context differences the agent reported. In 01_execution_and_time_model.md the tip replaced `model_runner.py:3233-3320` with the forward(...) signature, context next to this PR's re-pinned :801-803; and the tip's new test trails the PR's hunk. Tip direction: this PR's :801-803 and its projection-test lines. The tip leaves scheduler.py untouched, so :801-803 still points at the three detailed_* fields in the merged tree.

3. The merged file, run and mutated

58cb9dcce staged into my own path in xiaobizh_n18_cpu; atom resolved under the staged root.

Unmutated: 17 collected, 17 passed, rc 0, on both of two runs. test_the_package_imports_nothing_from_the_engine has 7 cases, [shape.py] included. mark.args[1] == _backend_modules() holds, so the tip's test reads the parametrisation this PR's reformatted assertion runs under.

Mutant What changes Result Caught by
M1: narrow the parametrisation _backend_modules() -> [p for p in _backend_modules() if p.name != "shape.py"] 1 failed, 15 passed, rc 1 test_every_module_the_walk_returns_is_a_case
M2: non-recursive walk rglob -> glob in _backend_modules() 1 failed, 16 passed, rc 1 test_the_walk_returns_every_module_under_the_root
M3: narrow the walk itself _backend_modules() excludes shape.py 16 passed, rc 0 nothing
M4: the reformatted assertion still fires shape.py gains def _mutant(): import atom.model_engine.scheduler 1 failed, 16 passed, rc 1 test_the_package_imports_nothing_from_the_engine[shape.py]

test_the_projection_type_is_where_the_scan_can_reach_it catches none of M1-M3: it checks defined_in in set(PACKAGE.rglob("*.py")), its own copy of the walk, not the walk the cases come from. Its docstring says "moving either type out of the package fails here rather than quietly thinning what the scan means", but a scan thinned to drop shape.py removes the [shape.py] case and stays green. Control on 0fc02f458, staged the same way: 13 passed; with its inline parametrisation narrowed to exclude shape.py, 12 passed, rc 0. The fix is to read the walk the cases use, for example set(_backend_modules()) or the parametrisation's mark.args[1]; that changes a line neither side wrote, which the exception forbids.

4. Gate, read from the logs

The node-18 staging dir is gone; read from agent_scratch/compass_dev/baseupdate-C/xml/:

Log atom: commit: Result
new81.log /tmp/buC-gates/new81/ATOM/atom/__init__.py 58cb9dcce (stamp) 5329 passed, 155 skipped, 3 xfailed, pytest: rc=0, GATE_CPU_RC=0 PASSED
tip.log under its own staged root c92e4c1c1 5281 passed, 155 skipped, 3 xfailed, GATE_CPU_RC=0 PASSED

Both show gpu: not required (.compass-changed stamp). Junit XML recomputed with grep, not the agent's script: new81.xml has 5487 testcase ids (5329 + 155 + 3), tip.xml 5439, neither with a <failure> or <error>. new81 minus tip: +48 ids, 0 removed; 45 are test_backend_shape_stub (8 classes) and 3 are test_backend_interface's [shape.py], [BatchView] and [RequestShape]. test_backend_interface has 17 ids in new81.xml, 14 in the tip and 13 in the old head, matching my own run.

Evidence: agent_scratch/compass_dev/bu81-review/ (remerge.diff, pr.delta, tip.delta, run_tests.out, run_control.out).

Agent-authored delta review under the base-update exception. No push, merge, label or draft change.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

need human Automation stopped; needs owner judgement. Agents apply when escalating, only the owner removes.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant