Skip to content

compass: the grant rule and the clock state machine (CA-2, #45) - #59

Closed
jgong5 wants to merge 4 commits into
feature/atomcompass_newfrom
compass/ca-2-grant-rule
Closed

jgong5 wants to merge 4 commits into
feature/atomcompass_newfrom
compass/ca-2-grant-rule

Conversation

@jgong5

@jgong5 jgong5 commented Sep 21, 2026 •

Copy link
Copy Markdown
Owner

Closes #45 when landed. Builds on #52 (compass/ca-1-clock-identity), which is the base of this PR.

Pure logic: the grant rule and the state machine that carries it. No transport, no endpoint resolution, no process management.

Review round 1 returned REQUEST CHANGES on one blocking finding. Fixed, and the fix found a second case the report did not reach. Details under The bookkeeping defect below. The branch was amended and force-pushed (bbd09f89b → 4c538fdb1), so CA-3 needs to rebase.

The interface

atom/compass/clock/state.py

LpStatus RUNNING, GRANTED, BLOCKED_ON_MESSAGE; .may_produce_events is false only for the last
LpState frozen (lp_id, now, next_event, status)
Grant frozen (lp_id, advance_from, advance_to, bound, bound_from); .seconds is the span. advance_to == advance_from is a real grant — it says a parked participant's wait is over, not that time moved

atom/compass/clock/authority.py — ClockAuthority(registry, lookahead, start_time=0.0)

Call Returns Notes
request_advance(lp_id, next_event=inf) tuple[Grant, ...] every grant the call released, for any participant, in registry order. Empty means the caller stays parked
take_up_grant(lp_id) Grant collect and start executing; GRANTED → RUNNING; consumes the events the grant reached
schedule_event(source, target, timestamp) tuple[Grant, ...] both safety checks, records the event against the target, resolves
grant_bound(lp_id) float +inf with no peers
earliest_emission_times() dict[LpId, float] the quantity the bound is a minimum over
peers(lp_id) tuple[LpId, ...] registry-derived; the seam the rule ranges over
require_sized_peers() — delegates to LookaheadMatrix.require_complete(); called once from __init__
state / states / now / held_grant / grants_issued / lp_table see T84 on grants_issued
registry / lookahead properties

Exceptions: ClockAbort (base, carries .reason and .table), BackdatedEvent, ClockDeadlock.

State machine: RUNNING --request_advance--> GRANTED or BLOCKED_ON_MESSAGE; BLOCKED_ON_MESSAGE --(another participant's call releases it)--> GRANTED; GRANTED --take_up_grant--> RUNNING.

The bookkeeping defect, and the second case behind it

Reported (blocking). request_advance assigned the declared horizon over the clock's record, so a participant asking for time after a peer had scheduled an event on it erased the only record of that event and was released straight past it. Reproduced here exactly as reported: zero floors, schedule_event(traffic→engine, 0.0), request_advance(engine, 10.0), request_advance(traffic, inf) grants engine 0 s → 10 s over an accepted event at 0 s, with nothing raised.

I did not take the suggested fix, because measuring it showed it incomplete. Three candidates, measured:

steps over the event a run that reaches its horizon an event at 5 s beyond a grant that reached 2 s
as committed yes progresses record lost
min only no freezes at 0.5 s kept
min + clear the horizon on take-up no progresses record lost (5.0 → inf)
min + drop only what the grant reached no progresses kept

The suggested fix closes the reported case and re-opens the same hole one grant further out: an event accepted beyond the grant is thrown away, and the participant cannot re-declare it because it has not been told about it.

Then the fuzzer found a third case, which neither the report nor my fix covered. A single scalar horizon holds only the earliest accepted event, so reaching the first of two events in flight forgets the second. 41 step-overs across 400 runs.

So the horizon is now two records. _declared is what the participant itself last said — only ever events it has seen. _accepted is a list of every event a peer has placed on it that it has not yet been released to reach. next_event is the minimum of the two. Taking up a grant drops the accepted entries the grant reached and the declared horizon if the grant reached it, and nothing else. Both halves are load-bearing and both are tested.

The fuzzer. 400 generated runs over two to five participants and three floors, driving legal operations and asserting that no grant carries a participant past an accepted event it has not been released to reach. Nothing in it relies on an abort being raised, because the failure raises nothing. Not vacuous — mutated three ways, it catches each:

mutant step-overs found in 400 runs
the committed defect (declared horizon replaces the record) 245
drop every accepted event on take-up 100
accepted side is one slot 25
as fixed 0

The named result

Two participants, engine and traffic-source, zero floor in both directions, both clocks at 0. The engine knows of an event of its own at 20 s. The traffic source is executing at 0 s and knows of no event of its own. The rejected rule is a test subclass overriding only earliest_emission_times, so the real state machine and both safety checks run underneath it.

next[j] rule now[j] rule
grant_bound(engine) +inf 0.0
engine granted to 20.0 s not granted; stays at 0.0 s
traffic source schedules an event at 0.0 s aborts — 20 s into the engine's past accepted; releases the engine

Repeated at a 1 ms floor: the clock rule releases the engine by exactly one floor, the horizon rule still to 20.0 s.

The sibling result: an unsized pair

Bound on decode, all clocks at 0
minimum over the pairs that were sized 10.0 s
minimum over registered peers, leg declared at 0.0 0.0 s

An absent floor is not a cautious version of a zero floor. A zero adds a term to the minimum and can only lower it; an absence removes a term, which raises it. Measured end to end, a peer set taken from the sized pairs grants decode to 10.0 s while the traffic source stands at 0.0 s, and the first event on that leg is 10 s into decode's past. Three doors are shut on it now: construction calls require_complete(), inbound() refuses a short row, and the walk is over peers(). Reproducing it needs two deliberate overrides.

The three invariants

Enforced by Covered by
Safety schedule_event raises BackdatedEvent on ts < now[source] + L and on ts < now[target], always, with the full table test_an_event_earlier_than_the_senders_own_floor_aborts, test_the_abort_carries_every_participants_clock_and_status, both discrimination classes
Deadlock _refuse_to_stall raises ClockDeadlock whenever a resolve issues nothing and everything is parked. No timeout exists in the package test_everyone_waiting_with_no_known_event_aborts_loudly, test_one_participant_with_nothing_left_to_do_is_the_same_stall, test_a_floor_above_zero_delays_the_stall_by_one_step_and_no_more
Determinism every walk over registry.ids(); grants issued and returned in that order; ties keep the first peer. No set constructed test_grants_come_out_in_participant_order_not_arrival_order (four permutations, identical ledgers), test_a_tie_on_the_bound_keeps_the_first_peer_in_the_total_order

Neither safety check is an assert — python -O deletes those. test_the_safety_check_is_not_an_assert_statement walks the AST and asserts the module has no Assert node.

Both degenerate cases

  • One participant — bound is +inf, grants land exactly on the declared horizons, bound_from is None. A local clock, no special case.
  • Every floor at zero — three participants with events at 5 s, 1 s and 7 s: nobody moves until the last declares, then all three advance to 1 s, the earliest event anywhere. Correct, serialized, and not an error.

Design amendments, in this PR

12_open_items.md §5 exists for corrections not yet applied; these are applied, so they go in the documents, with T-rows for what is not mine to settle. Register goes 81 rows → 83, 76 open → 78; README.md and 12_open_items.md §-intro counts updated together.

01 D3, grant rule — amendment. Two claims were false as written and both were found by implementing them. The literal rule cannot make a move at a zero floor: every clock starts equal, so min over j≠i of now[j] equals now[i]. It is not a start-up wrinkle — the condition recurs every step, because that minimum is over clocks, not over next events. The three-LP configuration D3 itself calls a global event loop (zero floors, events at 5/1/7 s) deadlocks under it. So it also does not degenerate to a global event loop; it degenerates to a stall. The amendment records the correction, and the reviewer's independent measurements of it: the fixpoint agrees with an unbounded relaxation on adversarial reverse-order chains at n=5…65 (exactly n passes, every time) and on 200 random topologies with no disagreement; and both halves are load-bearing, since "parked contributes its own next" without the fixpoint grants b to 15.0 s where the full rule gives 5.0 s and an event at 5.0 s then lands 10 s in b's past.

01 D3, sizing — amendment. The 240k grant estimate assumes a driver discipline nothing states. Measured here, two LPs, one event at 20 s, identical matrices:

How the LPs drive the clock 1 ms floor 1 µs floor
idle LP parks before the busy one asks 3 grants, reaches 20 s 3 grants, reaches 20 s
every LP takes up and immediately asks again 20,002 grants, reaches 20 s >2,000,000 grants, reaches 2 s

Two things the numbers say that the intuition does not: the efficient discipline is lookahead-independent, and a tighter floor makes the inefficient one worse, not better.

T83 — no way for a participant to say it has finished. A clean end of run is indistinguishable from a deadlock and aborts. Not theoretical: the harness I wrote to measure T84 hung on exactly this, because a drained participant could neither park (aborting the run) nor keep asking (pinning every peer at its stale clock). Owner is whoever takes #46 and #47; I deliberately did not invent a convention.

T84 — grants_issued() is arrival-order dependent while the event schedule is not: 12 distinct totals across 24 ask orders, 1457–1484, against one distinct set of 1024 events. A requirement on the run summary (#50). Measured by the reviewer and attributed, not reproduced here — my harness hit T83 and hung, and I judged reproducing someone else's measurement not worth further time against a registered row. Flagged rather than quietly dropped.

Carried, not fixed here

Cost at a large participant count — measured and mostly removed. It was not the relaxation. The row each walk takes a minimum over is now built once at construction, since membership is fixed and a floor is a declared constant:

participants, all parked per-walk hoisted
3 0.031 ms 0.019 ms
10 0.183 ms 0.074 ms
65 5.631 ms 1.682 ms

Identical grants in every case. (Measured in the local container, not on node 18, so the absolute numbers are not comparable with the reviewer's 8.49 → 1.59 ms; the ratio is.) Still above D3's 0.050 ms RPC estimate at 65, which is T70/PP territory and not this task's.

Effort

AST lines
atom/compass/clock/state.py 36
atom/compass/clock/authority.py 199
production total 235 (estimate 250–350)
tests/compass/test_clock_grant_rule.py 385

Plus 20 insertions / 6 deletions in __init__.py (additive exports), and the design amendments above. No CA-1 source file was otherwise touched. Method, matching CA-1's: ast.parse, strip docstrings from every module, class and function, ast.unparse, count non-blank lines.

Gates

Node 18, container xiaobizh_n18_cpu. Both trees staged by git archive + docker cp into paths of my own — the shared mount untouched — tarball md5 verified on both ends, .compass-commit and .compass-changed stamps written by snapshot.sh and confirmed present in the container before running, and import atom resolved under the tree before any count was read.

Control Branch
Commit 5b11e82ca (this PR's base) 4c538fdb1
Result 4084 passed, 149 skipped, 3 xfailed 4131 passed, 149 skipped, 3 xfailed
pytest: rc= 0 0
GATE_CPU_RC= 0 0
GPU tier not required not required

Delta: +47 passed, 0 failed, skips and xfails unchanged.

Arithmetic: 43 tests in the new file, plus 4 from CA-1's two package-wide parametrized checks — the standard-library allowlist and the no-set check are parametrized over clock/*.py, and this PR adds two modules, so each gains two cases. 43 + 4 = 47.

On the reviewer's GATE_CPU_RC=98 on both trees: that is gate_cpu.sh refusing a bare git archive tree with no .compass-changed stamp, not a difference in the trees. Building the snapshot with scripts/compass/snapshot.sh writes it; both runs above report gpu: not required (.compass-changed stamp).

ruff check and black --check clean on all touched files, against a repository baseline that is dirty.

🤖 Generated with Claude Code

jgong5 and others added 2 commits September 21, 2026 13:56
Adds atom/compass/clock with three things and nothing that uses them: the
identity of a logical process, the registry that holds a total order over
those identities, and the lookahead matrix addressed by identity in both
directions. The rule that reads them, the per-participant state and the
transport are separate.

The order is by name, which is a code-point comparison and therefore the
same in every process. Arrival order is never consulted. The registry
holds membership in a dict with None values rather than a set, and the
package builds no set at all, so no ordered answer it gives can depend on
a per-process hash seed.

A lookahead floor is declared per directed link and carries the class of
path it represents. An undeclared pair is refused rather than read as
zero: zero is a decision to serialize a pair and stays correct, silence is
a link nobody sized. Zero floors are accepted and reported by
LookaheadMatrix.serializing.

Tests are CPU-only. test_clock_order_across_processes.py spawns three
interpreters with different PYTHONHASHSEED values, registers the same
twelve names rotated differently in each, and compares the registry order
against a plain set of the same names built in the same children as a
control.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A caller minimises over one participant's whole row, quantified over every
registered peer. inbound() filtered the row to the links that happened to
exist, so a peer with no declared floor contributed no term at all: the
minimum came out higher, not lower, and more virtual time was handed out
than the peer allowed. Measured on a three-participant registry with one
leg undeclared, the walk gave 5.002 s where the safe answer was 0.0 -- the
grant lands past the peer's current time and the peer's next event lands
in the past. An omission reads as an unbounded lookahead, which is the one
direction that is unsafe; a zero would have been the conservative mistake.

inbound() now refuses a row missing any registered peer, naming each
missing pair, and require_complete() refuses an incomplete matrix at
set-up so the failure happens before anything advances rather than one
step later as a backdated event. tightest() says in its docstring that it
reports over declared links only.

InterLpLink is now a frozen dataclass. The accessors hand out the object
itself, so a writable floor_seconds let a holder rewrite a declared floor
and reach around every refusal declare() performs -- negative, NaN,
infinite, already declared. LpId was frozen for this reason; the object
carrying the number was not.

LpRegistry.require() type-checks its argument. It is what declare(),
lookahead() and inbound() funnel through, and without the check a bare
"decode" was reported as not registered beside the registered decode.

The source guards now rglob the package, so a subpackage added later
cannot escape the import allowlist or the no-set check silently, and the
allowlist drops typing, which nothing imports. A signature and field scan
covers the requirement that no public call names a host, port, endpoint,
address, url, level, depth, parent, rank, node or socket -- the import
allowlist proves the package reaches no device, clock or socket, which is
a different claim.

The cross-process control now builds its set from the unrotated name list,
so the hash seed is the only thing varying between children, and each
child prints the atom it resolved.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Comment thread atom/compass/clock/authority.py Outdated
f"{self._seconds(self._now[lp_id])}",
self.lp_table(),
)
self._next[lp_id] = horizon

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocking. This assignment discards an event the clock has already accepted, and the run then walks over it silently.

schedule_event lowers a participant's horizon with min (:356). request_advance replaces it. So any event a peer scheduled on i is erased the next time i asks for time — and i asking for time after a peer scheduled something on it is the ordinary case, not an edge case: the message is still in the transport, so the participant declares the event it knows about, which is its own.

Two participants, both floors zero, both clocks at 0:

schedule_event(traffic -> engine, 0.0)   -> engine's horizon = 0.0
request_advance(engine, 10.0)            -> engine's horizon = 10.0   the 0.0 is gone
request_advance(traffic, inf)            -> engine 0s -> 10s, traffic-source 0s -> 10s

The engine now stands at 10.0 s with an accepted, undelivered event at 0.0 s. Nothing aborts: schedule_event returned long ago and there is no later check. This is the failure the module exists to prevent, in the shape the module docstring describes — it does not crash, and the run produces a plausible answer.

It is not a zero-floor artefact. At a 1 ms floor, with traffic at 0.001 s and the engine at 0.002 s, schedule_event(traffic -> engine, 3.0) then request_advance(engine, 10.0) then request_advance(traffic, inf) grants engine 0.002s -> 10s, seven seconds past the event. With the engine declaring the horizon the clock already held, the same sequence grants engine 0.002s -> 3s.

Fuzzed, 400 randomised runs, 2–6 participants, floors drawn from {0, 1e-6, 1e-3, 0.25, 1.0}, 200 legal actions each, counting any grant whose advance_to steps past an accepted event on that participant:

driver runs that stepped over an accepted event
declares its own next event (in protocol) 378 / 400
also carries the horizon the clock already holds 0 / 400

No run in either arm raised. The control arm is what says the rule itself is sound: the defect is here, in the bookkeeping, not in earliest_emission_times.

The obvious one-line fix is not enough, and I measured that too. self._next[lp_id] = min(self._next[lp_id], horizon) alone closes the break and then freezes the run, because nothing ever raises the horizon again — six consecutive rounds with both clocks stuck at 1.0 s. min here plus clearing self._next[lp_id] to math.inf in take_up_grant closes the break and leaves an ordinary run advancing 1.0, 2.0, 3.0 … exactly as it does today:

break case, engine's clock ordinary run
as written 10.0 s (event at 0.0 s) 1, 2, 3, 4, 5, 6
min only 0.0 s 1, 1, 1, 1, 1, 1
min + clear on take-up 0.0 s 1, 2, 3, 4, 5, 6

Clearing on take-up is also the honest reading of the field: a participant that has collected a grant has seen everything up to its clock, and will re-declare on its next call.

Whatever shape you pick, the property to test is the one the fuzzer checks: no grant may move a participant past a timestamp schedule_event has already accepted for it. That is a different statement from either of the two checks in schedule_event, and nothing in the file tests it today.

same sequence however the requests happened to interleave in real time.
"""
self._participant(lp_id)
if self._status[lp_id] is LpStatus.GRANTED:

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Related to the horizon finding on :288, and part of the same fix: a second request_advance is refused while the participant is GRANTED but accepted while it is BLOCKED_ON_MESSAGE, where it runs straight into the assignment two lines below and silently raises a horizon a peer lowered.

schedule_event(traffic -> engine, 0.0)      horizon 0.0
request_advance(engine, 5.0)                blocked-on-message, horizon 5.0
request_advance(engine, inf)                blocked-on-message, horizon inf

The state machine in the PR body has no BLOCKED_ON_MESSAGE --request_advance--> ... edge, so this call is out of protocol; the authority is the component that should say so. A raise here costs one branch and removes a whole class of driver bug from CA-3 and CA-4 before either is written.

else:
earliest[lp_id] = max(self._now[lp_id], self._next[lp_id])
waiting.append(lp_id)
for _ in waiting:

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The cost the dev record left unmeasured, measured — and it is not the relaxation.

Rule arithmetic only (one earliest_emission_times() plus one _bound() per parked participant), a complete matrix, mean over 200–300 repetitions, in the container's Python 3.12:

registered parked as written peer row hoisted
2 2 0.013 ms 0.005 ms
3 3 0.029 ms 0.008 ms
10 10 0.226 ms 0.051 ms
33 33 2.260 ms 0.423 ms
65 8 1.065 ms 0.227 ms
65 33 4.628 ms 0.823 ms
65 65 8.493 ms 1.588 ms

For comparison, the design sizes one grant's IPC round trip at 0.050 ms and calls 240k grants ~12 s of overhead. At 10 participants this loop is ~5x that round trip; at 65 it is ~170x. So at the largest topology the rule, not the transport, is the cost.

The decomposition matters more than the total. Counting passes: the loop runs 2 passes on every realistic shape I tried at n=65 — equal floors, a cheap chain, a 1 µs ring with 10 ms elsewhere. The for _ in waiting bound is reached only by an adversarial topology where the cheap chain runs against the registry order, and there it uses exactly n passes and still agrees with an unbounded relaxation (checked at n = 5, 10, 20, 40, 65, and on 200 random topologies, zero disagreements, max 5 passes). The relaxation depth is not the cost. The cost is the O(parked x peers) walk underneath it, and specifically two things inside it: self.peers(target) rebuilds an (n-1)-tuple and re-runs _participant on every target of every pass, and lookahead(source, target) hashes a fresh (LpId, LpId) tuple of two unslotted frozen dataclasses on every element.

The "hoisted" column above is the same arithmetic with the inbound row and its floors read once per matrix — 4 to 5x, no change to the rule, no change to any result. LookaheadMatrix.inbound() already returns exactly that row, so the material is there.

Not blocking: at the 2–3 participants of the near milestones this is 0.03 ms and nothing. But the dev record names 65 as the number worth measuring before trusting, so here it is, and the answer is that it is affordable once the row stops being rebuilt.

Comment thread atom/compass/clock/authority.py Outdated
settled = True
for target in waiting:
best = earliest[target]
for source in self.peers(target):

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For CA-3 and CA-4, because this line is where it bites and nothing in the module says it.

Whether a run skips idle or crawls is decided by the driver, by four to six orders of magnitude. A peer that is RUNNING or GRANTED contributes its current clock, so it pins every other participant at that clock plus one floor. The fixpoint above only gets to jump to a distant event when every peer is parked at the moment of the resolve. Two participants, one event at 20 s, nothing else in the run:

floor driver grants clock reached
9 ms take up the grant, then let the peer ask 2,224 20 s
9 ms ask until the clock refuses, then park 3 20 s
1 ms take up, then alternate 20,000 20 s
1 ms ask until refused 3 20 s
1 µs take up, then alternate 800,000 (capped) 0.8 s
1 µs ask until refused 3 20 s

The first discipline is the natural one to write — ask, get time, take it, run, find nothing, ask again — and at the pipeline-stage floor it turns a 20-second idle stretch into 40 million grants. The second reaches the same schedule in three. The difference is only that in the second, a participant that has been refused stays refused until a peer's state changes, so the all-parked state is reachable.

Nothing is wrong with the rule here: every row above is safe, and both disciplines that finished reached the same event at the same virtual time (20.0 s). But the sizing this design rests on assumes the second discipline, and the module gives a driver author no signal about which they are in. Worth a sentence in the module docstring and a hard requirement on the harness.

self._participant(lp_id)
return self._held[lp_id]

def grants_issued(self, lp_id: LpId | None = None) -> int:

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

grants_issued() depends on real arrival order. Four participants, unequal floors and unequal horizons, one fixed workload, driven to 1024 events under all 24 permutations of the arrival order:

  • distinct (virtual time, participant) event sets: 1 — the D3.4 property holds, and holds exactly
  • distinct grants_issued() totals: 12, from 1457 to 1484

That split is the right one and it is worth stating in the interface: the simulated schedule is invariant under arrival order, the protocol's own cost is not. Since this counter is destined for an always-written run summary, CA-7 should keep it out of anything byte-diffed for reproducibility, or record it beside the diff rather than inside it.

The same effect shows up one level down: with equal floors and equal horizons the grant ledger is identical under permutation, and with unequal ones it is not — 24 distinct ledgers from 24 permutations, differing in advance_to and in bound_from. Not a defect (the events still land on the same times), but see the note on the determinism test.

def test_the_safety_check_is_not_an_assert_statement():
# `python -O` removes an `assert`. A check that is specified as always on,
# not behind a flag, cannot be one -- so the module contains none at all.
tree = ast.parse((CLOCK_PACKAGE / "authority.py").read_text())

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This test parses authority.py only, while the PR body states the property as "the module has no Assert node". Measured across the package at bbd09f89b: no Assert node in any of the six modules — __init__, authority, identity, lookahead, registry, state — so the property is true, wider than the test.

CA-1 made its two source guards package-wide and parametrised (rglob over clock/) for exactly this reason, and the round-2 reply says why. This one is the check with the sharpest failure mode of the three — python -O deleting a check specified as always on — and it is the one that is not package-wide. A parametrised version over the same glob costs the same line count and covers the case where a safety check moves into a helper module, which is a plausible CA-5 refactor.

class LpState:
"""One participant's row: its clock, its horizon, and what it is doing.

`next_event` is the earliest future event the participant knows of, and is

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This sentence is no longer what the field holds, and the gap is the root of the :288 finding in authority.py.

next_event starts as "the earliest future event the participant knows of", but schedule_event lowers it with min on the strength of a peer's event the participant has not been told about yet. So the field is the earliest event the clock knows of for that participant, which is a strictly larger set than what the participant knows. Once it is that, an assignment from the participant's own declaration is losing information, not refreshing it.

Worth saying here, since this docstring is what a CA-3 or CA-4 author will read before writing the call.

# --- determinism -------------------------------------------------------------


def test_grants_come_out_in_participant_order_not_arrival_order():

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sound for what it tests — mutation-checked: a _resolve that serves the caller first and the rest in order gives 3 distinct ledgers over these 4 permutations and fails both assertions — but the configuration hides a real property, and it is worth knowing which one.

All three participants declare the same horizon (5.0) at the same floor, so none of them can be granted until the last has parked, and every permutation therefore resolves in one pass. Order within a resolve is covered. Order across resolves is not: with unequal floors and unequal horizons, four participants give 24 distinct grant ledgers from the 24 request permutations, differing in advance_to and bound_from.

That is not a bug — driven to completion, all 24 arrival orders produce the same 1024 (virtual time, participant) events, which is the property the design actually asks for. But the invariant the design names is the event sequence, and this file asserts the ledger, which is the stronger claim and the one that is false in general. A second case that permutes arrival order on an asymmetric configuration and asserts the event sequence would pin the property that holds, and would catch a regression that this one cannot see.

@jgong5

jgong5 commented Sep 21, 2026

Copy link
Copy Markdown
Owner Author

Review record — round 1, agent-authored

Verdict: REQUEST CHANGES. One blocking finding, in ClockAuthority.request_advance: the assignment at authority.py:288 discards an event the clock has already accepted, and the run then walks over it silently. 378 of 400 randomised legal runs reach it. Everything else below is non-blocking, and the largest part of it is agreement.

GitHub refuses APPROVE / REQUEST_CHANGES on a self-authored PR, so the verdict is that sentence.

The departure from the design's stated rule is correct, necessary and minimal, and I could not break it. That is the thing this review was pointed at hardest and it is the strongest part of the PR. Details below, with the measurements. The blocking finding is one line away from it and is not about the rule at all.


The blocking finding, in one paragraph

schedule_event lowers a participant's horizon with min. request_advance replaces it. A participant that asks for time after a peer scheduled an event on it — the ordinary case, since the message is still in the transport and the participant declares what it knows — erases the clock's record of that event, and is then granted straight past it. Two participants, zero floors, both at 0: schedule_event(traffic -> engine, 0.0), request_advance(engine, 10.0), request_advance(traffic, inf) grants engine 0s -> 10s with an accepted, undelivered event at 0.0 s and no abort anywhere. At a 1 ms floor the same shape grants engine 0.002s -> 10s over an event at 3.0 s. Fuzzed over 400 runs, 2–6 participants, floors from {0, 1e-6, 1e-3, 0.25, 1.0}, 200 legal actions each: 378/400 step over an accepted event with a driver that declares its own horizon, 0/400 with a driver forced to carry the horizon the clock already holds. No run in either arm raised — it is the silent shape, exactly. The measured fix and the measured failure of the obvious one-line version are in the inline comment.


The central question: the rule departs from what the design states

1. Is the problem real?

Yes, and it is larger than the dev record claims. I implemented the stated rule verbatim as a subclass returning now[j] for every participant regardless of status, and ran it.

At a zero floor, three participants holding events at 5 s, 1 s and 7 s — the configuration the design calls a single global event loop:

literal rule : alpha -> [] ,  beta -> [] ,  gamma -> ClockDeadlock
               "every participant is waiting and none can be granted time,
                though alpha knows of an event at 5s, beta at 1s, gamma at 7s"
this PR      : alpha -> [] ,  beta -> [] ,  gamma -> alpha 0s->1s, beta 0s->1s, gamma 0s->1s

So the stated rule does not degenerate to a single global event loop at zero lookahead. It stops. And it is not a bootstrap problem: a global event loop advances to the earliest next event, and min over j != i of now[j] is the earliest clock, which at a zero floor equals the asker's own clock forever. The condition recurs at every step, not only at step one — the dev record understates its own finding.

The design document therefore states a rule that cannot run a configuration the same document requires to work. That is a defect in the document.

2. Is the fix safe?

I could not break it, and I tried the fixpoint first, as instructed. The property the design's justification rests on — nothing a participant generates can land in anyone's past — survives the substitution, and three independent attacks left it standing.

  • The fixpoint agrees with an unbounded relaxation. for _ in waiting is exactly the Bellman–Ford bound: |parked| − 1 propagation passes plus one confirmation. Checked against a relaxation run to true convergence on an adversarial chain whose cheap links run against the registry order, at n = 5, 10, 20, 40, 65 — it uses exactly n passes at every size and matches every time — and on 200 random topologies (2–9 participants, mixed floors and horizons, 80% parked): zero disagreements, maximum 5 passes.
  • Relaxing from above is the right direction, and it is what makes the substitution sound. Starting at max(now, next) and only lowering converges to the greatest fixpoint, which excludes circular support: two parked participants cannot talk each other into an earlier emission time, because a value only descends when it is grounded in a participant that can actually move or in a horizon somebody actually declared. Two parked participants with zero floors and no horizons keep +inf, produce no grant, and the stall check fires — which is correct, and is the case a least-fixpoint would have got wrong in the safe-but-useless direction.
  • The fuzzer, with the horizon defect neutralised, never produces a grant past an accepted event — 0/400 runs, and the second safety check never fires, which corroborates the "unreachable under this rule" claim in the dev record rather than merely restating it.

The inductive reason it holds: for a running peer the contributed value is its clock, which is monotone; for a parked peer the contributed value is exactly the time it would itself be granted to in that same resolve, so granting it changes nothing a second pass would read differently — and an event a peer schedules on a parked participant already clears both the sender's floor and the recipient's clock, so it cannot drag the contributed value below what a previous bound was computed from. The one thing that can drag it is the assignment at :288, which is the blocking finding.

3. Is it minimal?

Yes, and both halves are load-bearing. Measured.

  • A bootstrap grant or an initial condition cannot work — the stall recurs at every step, not only the first (above).
  • "A parked participant contributes its own next", without the fixpoint, is unsafe. Three participants a, b, c; L[a->b] = 10, L[a->c] = 0, L[c->b] = 0, everything else 10; a knows of an event at 5 s, b and c of none:
b granted to c -> b at 5.0 s
this PR (relaxed to a fixpoint) 5.0 s accepted
parked contributes only its own next 15.0 s abort — 10 s into b's past

Without the relaxation a parked participant with no horizon contributes +inf, which raises everyone's bound — the same shape as CA-1's absent-floor finding, one level up. So the fixpoint is not decoration.

The smallest correct change is the one that was made: keep now[j] wherever the peer can emit, substitute only for a peer that provably cannot, and relax because those substitutions refer to each other.

Should the design be amended, and what guarantees it?

Yes, here, and a PR-body note does not guarantee it. Two things are wrong in D3 as written, both measured above: the grant rule cannot start or continue a run at a zero floor, and the claim that it degenerates to a single global event loop at zero lookahead is false of the rule as stated (it is true of the rule in this PR). A third is weaker but real: the sizing rests on a driver discipline the document never states — see below.

The project has a mechanism for exactly this and a precedent from yesterday: 12_open_items.md section 4, and T82, opened 2026-09-21 by P0.5, which registers a design-document contradiction found while writing code, names the amendment, and explicitly does not make it. That is the shape to copy. I would hold this PR only for the blocking finding; the amendment can be a registered row landing with it. What I would not accept is the current state, where the only record that the document's rule does not run is a paragraph in a PR body — the development rules name "a design document that contradicts the code" as a stop-and-discuss, and this is one.


The discrimination test: is the rejected-rule subclass honest?

Yes. Checked mechanically rather than read: _HorizonRuleAuthority overrides exactly one method, earliest_emission_times. schedule_event, _resolve, _bound and _refuse_to_stall are all inherited, so both safety checks, the state machine and the stall check run underneath it, and the abort in the rejected column comes out of the real check at authority.py:346. _SizedPairsAsPeersAuthority overrides two, require_sized_peers and peers, and the dev record says why both are needed and that needing two is itself the measure.

I reproduced both results:

next[j] now[j]
grant_bound(engine) +inf 0.0
engine granted to 20.0 s not granted
event at 0.0 s abort, "20s into its past" accepted, releases the engine

and the 1 ms repetition (1.0e-3 vs 20.0), and the sibling result (10.0 s over sized pairs vs 0.0 s over registered peers; decode 0s -> 10s bounded by prefill while the traffic source stands at 0.0, then BackdatedEvent "10s into its past").

One note on what the discrimination turns on, which strengthens it: in that topology the engine is parked and the traffic source is running, so the whole difference between the two columns comes from a peer the new rule does not treat specially. The test measures the one case where the substitution is not in play, which is the right case to measure and is what makes "this is not the next[j] rule wearing a disguise" a demonstration rather than an assertion.

The three invariants

Safety — holds, and is not an assert. schedule_event raises BackdatedEvent unconditionally on both conditions, with lp_table() attached, and the table is in str(exc). AST walk across the package at bbd09f89b: no Assert node in any of the six modules, so the property is wider than the test, which parses authority.py only — inline note at test_clock_grant_rule.py:343. The path I found to an event that reaches neither check is not a path through schedule_event — it is the horizon being discarded after schedule_event has already returned, which is the blocking finding.

Deadlock — holds, and there is no timeout. Grepped the whole package for timeout, sleep, monotonic, perf_counter, time(), datetime, threading, asyncio, wait_for, select, socket: the only hit in code is the word "timeout" inside the ClockDeadlock docstring explaining why there is none. Imports across all six modules are enum, math, dataclasses and relative siblings, nothing else.

I also checked the "impossible" branch of _refuse_to_stall — every participant parked, at least one finite horizon, no grant issued. It is genuinely unreachable under this rule, not merely untriggered: if any parked participant has a finite horizon then every bound is finite, and a state in which every advance_to equals its own standing requires a cycle of zero floors among participants at the same clock, where the relaxation collapses the whole cycle to the minimum declared horizon and issues a grant. The zero-floor three-participant run above is that case, and it grants. The branch is correctly kept: if it ever fires, the table is the evidence the argument was wrong.

Determinism — holds where the design asks for it, and the counter does not. Mutation-checked that the test is not decorative: a _resolve that serves the caller first and then the registry order gives 3 distinct ledgers over the test's 4 permutations and fails both assertions. Permutations the test does not try:

  • 24 request permutations of 4 participants with unequal floors and unequal horizons: 24 distinct grant ledgers.
  • Registration order permuted with request order fixed: no effect — ids() sorts, as CA-1 intended.
  • Three PYTHONHASHSEED values in child processes on a fixed sequence: byte-identical output.
  • The one that matters: the same 4-participant workload driven to completion under all 24 arrival orders — 1 distinct (virtual time, participant) event set, 1024 events, and 12 distinct grants_issued() totals, 1457 to 1484.

So the simulated schedule is invariant under arrival order and the protocol's own cost is not. That split is right, and it is a requirement on CA-7: authority.py:176.

Both degenerate cases

One participant reduces to a local clock with no special case in the code — confirmed, bound is +inf, bound_from is None, three grants land exactly on the declared horizons. Every floor at zero reduces to a single global event loop — confirmed, and confirmed against the stated rule, which does not.

The cost at 65, which the dev record left unmeasured

Measured, and the decomposition is the interesting part: it is not the relaxation. Two passes on every realistic shape at n=65; the for _ in waiting bound is reached only by an adversarial reverse-order chain. The cost is the O(parked × peers) walk underneath, where peers() rebuilds a 64-tuple per target per pass and every element re-hashes a (LpId, LpId) key. 8.49 ms per grant's arithmetic at 65/65 parked, 0.23 ms at 10, 0.03 ms at 3; hoisting the inbound row gives 1.59 / 0.05 / 0.008 with identical results. For scale, D3 sizes one grant's IPC round trip at 0.050 ms. Full table at authority.py:209.

What I did not expect to find, and what CA-3 and CA-4 must do about it

Whether a run skips idle or crawls is decided by the driver, by four to six orders of magnitude. A peer that is RUNNING or GRANTED pins everyone else at its clock plus one floor, so the fixpoint only gets to jump to a distant event when every peer is parked at the moment of the resolve. Two participants, one event at 20 s:

floor take up the grant, then let the peer ask ask until the clock refuses, then park
9 ms 2,224 grants 3 grants
1 ms 20,000 grants 3 grants
1 µs 800,000 grants, reached 0.8 s of 20 3 grants

The first discipline is the natural one to write. Both are safe; only one is affordable, and the design's 240k-grant sizing assumes the second. Nothing in the module tells a driver author which they are in. authority.py:213.

Effort and gates

Independently re-measured with my own ast.parse / strip module-class-function docstrings / ast.unparse / count non-blank script:

file ast non-blank raw
clock/state.py 36 81 104
clock/authority.py 186 406 469
production total 222
clock/__init__.py (whole file) 6 40 44
tests/compass/test_clock_grant_rule.py 308 463 577

Every figure in the PR body reproduces exactly, including the choice to quote 222 for the two new modules and state the __init__.py change separately as 20 insertions / 6 deletions — git diff --numstat confirms 20 6. 222 against 250–350 is an underrun of about 11%; not a halt-and-discuss event, that rule is about overruns, and the same posture was taken on CA-1.

tests/compass/test_clock_grant_rule.py contains 38 test_ functions and no parametrize decorator, so the 38 half of the 38 + 4 = 42 arithmetic is exact.

Both trees re-measured on node 18, container xiaobizh_n18_cpu, staged by git archive + docker cp into /opt/ca2rev/{control,branch} of my own, tarball md5 matched on both ends, atom.__file__ asserted under each tree before any count was read. Nothing was written into the shared mount and no rsync was used; both staged trees were removed afterwards.

control 5b11e82ca branch bbd09f89b
pytest 4084 passed, 149 skipped, 3 xfailed 4126 passed, 149 skipped, 3 xfailed
pytest rc 0 0
exclusions 29 files + tests/plugin same
GATE_CPU_RC 98 98

+42 passed, nothing else moved — reproduces exactly. So does the decomposition, checked by node id rather than by count: test_clock_grant_rule.py collects 38; test_clock_lp_identity.py goes 49 → 53, and the four new ids are test_the_package_imports_only_the_standard_library_it_names[authority.py], [state.py] and test_the_package_builds_no_set_at_all[authority.py], [state.py] — the two CA-1 package-wide checks picking up exactly the two modules this PR adds. test_clock_order_across_processes.py is 5 on both and contributes nothing, which rules out any other source for the delta. 38 + 4 = 42, with no unaccounted test. ruff check and black --check clean on all four touched files, rc=0.

One thing I did not reproduce: GATE_CPU_RC=0. I got 98 on both trees, so it carries no differential signal, and it is my staging rather than the branch: gate_cpu.sh:147 reports gpu: UNKNOWN -- no git, no COMPASS_CHANGED_FILES, no .compass-changed and exits 98 on a bare git archive tree, which has none of the three. Its own message names the fix — stage with scripts/compass/snapshot.sh, which stamps .compass-changed. The PR's GATE_CPU_RC=0 is consistent with that route and with CA-1's review, which quotes the gate printing gpu: not required (.compass-changed stamp). Stated because a GATE_CPU_RC reported without saying how the tree was staged is not reproducible, and mine was not.

Smaller things

  • schedule_event(source, target, math.inf) is accepted and is a silent no-op — _simulated_seconds refuses -inf and NaN but not +inf, and min(next, inf) changes nothing while the call reports success. One line, same guard.
  • A negative start_time is accepted (math.isfinite(-5.0) is true). Probably harmless, possibly deliberate; worth a word either way.
  • LpState.next_event's docstring no longer describes what the field holds once schedule_event has lowered it, and that gap is the conceptual root of the blocking finding: state.py:55.

Accepted with reservation

  • No retire call — I agree with the decision, and disagree with where it is recorded. A run that ends with every participant parked and no horizon is the deadlock condition; inventing a retire call here would be guessing at an interface with no caller, and CA-3 and CA-4 genuinely have the better view. But this is not a note-in-the-PR-body item: the first integration hits it at the end of every clean run, it changes the exception surface CA-3 and CA-4 code against, and a PR body is not somewhere a later agent looks. It belongs in 12_open_items.md section 4 as a registered row, alongside the D3 amendment, and linked from CA-3 — both transports from one implementation, in-process and socket #46 and CA-4 — synthetic-LP harness over ATOM's verified topology, 2 to 65 LPs #47. The same argument applies to the D3 amendment, and both are cheap to do together.
  • lp_table() as the minimum an abort needs rather than a dump format. Right call, right boundary, and it already carries the bound and the pinning peer, which is more than the minimum. CA-7 owns the rest.
  • The zero-span grant. The guard (advance_to == standing only fires when horizon <= standing) is correct and it is the right shape — a parked participant whose wait is over does need to be told, and Grant.seconds == 0.0 says exactly that. I checked it cannot spin.
  • earliest_emission_times as a public method. It is what makes the discrimination test honest, which is worth the surface. But it is now the seam any subclass can move the safety argument through, and CA-5's adversarial scenarios will be tempted to. Worth a sentence saying that overriding it is a test device, not an extension point.
  • peers() derived from the registry, not from the sized pairs. This is CA-1's blocking finding answered properly and then measured a second time from the other end. Good.

What CA-3, CA-4 and CA-7 should watch

  1. The horizon contract. Whatever shape the :288 fix takes, CA-3 and CA-4 must know whether a participant is expected to re-declare events a peer scheduled on it. Today the answer is "yes, and it cannot", which is why the defect exists.
  2. Poll until refused, then park. Not an optimisation — the difference between 3 grants and 800,000. It belongs in the transport's contract, not in a driver's discretion.
  3. grants_issued() is not reproducible and the event schedule is. Keep the counter out of anything byte-diffed.
  4. ClockDeadlock at the end of every clean run until a retire convention exists. Decide it in CA-3/CA-4, do not catch it.
  5. The rule is the cost at scale, not the transport — 170x the design's per-grant RPC estimate at 65 participants, before any socket exists. Fix the row walk before measuring the transport, or the transport measurement will be swamped.

What I could not check

  • The GPU tier. Correctly not required; not run; no claim made.
  • The behaviour of any real participant. There is none yet — everything above is the rule driven by harnesses I wrote, and my harnesses are a model of a driver, not a driver.
  • Whether the relaxation's worst case is reachable from a real topology. I reached |parked| passes only by constructing the link order adversarially against the registry order; whether ATOM's actual link structure can do that is a CA-4 question.
  • The provenance of the design's 50 µs per-grant RPC estimate, which I quoted for scale. The design itself records it as an estimate that has not been measured on this box.

Round 1. No escalation and no need human: the blocking finding is an actionable review finding, and the design-document amendment is a registered-item call, not a scope call.

The rule that hands out simulated time. A participant may move its clock to
min(bound, the earliest event the clock knows of for it), where the bound is the
minimum over every other participant of that participant's earliest emission
time plus the declared floor into this one.

The bound is taken over the peer's current clock, not over the peer's next
event. A peer executing at 0 while knowing of nothing until 10 can still emit at
0; bounding on the 10 releases the other side to 10 and the first event it
receives is ten seconds into its past. Nothing crashes -- the run finishes and
reports a plausible answer -- so both rules are measured against each other in
the tests rather than argued about in a comment.

A participant that has asked to advance and been refused is the one case where
the current clock is not the answer, since it emits nothing until granted. Its
earliest emission is relaxed to a fixpoint over the declared floors. Without
that, every floor at zero cannot make a move at all -- not once, but at every
step, because the minimum over clocks is not the minimum over next events.

The minimum ranges over registered participants, never over declared links: an
unsized pair is a term missing from a minimum, which lifts the bound instead of
tightening it. Refused when the clock is built.

The horizon is two records and not one. A participant declares only the events
it has seen; a peer may already have placed one on it that is still in the hands
of whatever carries messages. Folding those into one slot loses whichever was
written second, and losing the accepted one releases a participant straight past
a timestamp it has not reached. The accepted side is a list, because reaching
the first of two events in flight must not forget the second. Taking up a grant
consumes what the grant reached, and only that.

Three checks, none optional and none an assert statement, since -O deletes
those: an event behind the sender's floor aborts; an event behind the
recipient's clock aborts; nothing able to move with nothing known anywhere
aborts. There is no timeout. Ties break by participant name, so grants come out
in the same order however the requests interleaved.

One participant reduces to a local clock. Every floor at zero reduces to one
event loop spread over several processes -- correct, serialized, and not an
error.

400 generated runs assert that no grant carries a participant past an event the
clock has already accepted.

The row each walk takes a minimum over is built once, since membership is fixed
and a floor is a declared constant.

The execution and time model is amended where implementing it showed two of its
claims to be false, and the open-items register gains the missing retire call
and the arrival-order dependence of the grant count.

No transport, no endpoint resolution, no process management.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
)
self._held[lp_id] = None
self._status[lp_id] = LpStatus.RUNNING
self._accepted[lp_id] = [

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocking. A grant still erases an accepted event, one window later than round 1's — this filter drops events the grant was never computed against.

_resolve computes advance_to and moves _now[lp_id] immediately. The drop happens here, at take-up. Everything a peer schedules onto the participant in between — while it is GRANTED — is appended to _accepted after the grant was decided, and is then thrown away by this filter because it is <= grant.advance_to. The participant was never held at it, never told about it, and the clock keeps no record.

It is not a zero-floor tie. Asymmetric non-zero floors, L[p->q] = 4.0, L[q->p] = 1.0:

q asks, is granted 0s -> 4s, takes it up            q RUNNING at 4.0
p asks declaring its own event at 10.0
  -> p 0s -> 5s   (bound 5s by q)                   truncated by the BOUND, not by p's horizon
q schedules an event on p at exactly 5.0            legal: 4.0 + 1.0 floor == 5.0, and p's clock is 5.0
  -> p._accepted == [5.0]
p takes up its grant                                -> p._accepted == [],  next_event back to 10.0
q asks;  p asks 10.0                                -> p 5s -> 10s

p finishes at 10.0 s with an accepted event at 5.0 s, five seconds in its past, and nothing raised. That is the silent shape this module exists to prevent, and the grant that got truncated at 5.0 s was a bound truncation — p had no reason to expect anything at 5.0 s and the clock discarded the only record of it.

Fuzzed, with an oracle that (a) observes the grants schedule_event returns and (b) releases only what each grant was computed against:

step-overs in 400 legal runs
as committed, 4c538fdb1 18
grant remembers the accepted entries it was computed against; take-up drops exactly those 0

Same generator, same floors {0, 1e-3, 1.0}, same 2–5 participants, same 24 actions as _legal_runs.

The direction that measured clean is the one the docstring above already describes but the code does not implement — "taking up a grant consumes the events the grant reached" has become "consumes every event at or before where the clock now stands", and those are different sets once a peer can write into _accepted after _resolve has run. Snapshotting [t for t in self._accepted[lp_id] if t <= advance_to] at grant issue and dropping that list (not a threshold) at take-up is a four-line change and keeps both properties the current code was built for: test_a_run_that_reaches_its_horizon_keeps_moving still passes, because the entry that caused the grant is in the snapshot.

Principle 6: the failure produces a plausible answer instead of a refusal.

"so this run no longer describes the system being modelled",
self.lp_table(),
)
self._accepted[target].append(when)

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A round-1 finding that survived, and that this PR's change turned from a no-op into unbounded growth.

Round 1: "schedule_event(source, target, math.inf) is accepted and is a silent no-op — _simulated_seconds refuses -inf and NaN but not +inf ... One line, same guard." The guard was not added.

With a scalar horizon that was harmless: min(next, inf) changed nothing. With _accepted as a list it is not. take_up_grant keeps every when > grant.advance_to, and inf is greater than every finite advance_to, so an inf entry can never be dropped by any grant. Measured:

20,001 x schedule_event(y, x, math.inf)
  ->  len(_accepted[x]) == 20,002
  ->  one further schedule_event costs 1.716 ms

because _restate_horizon re-mins the whole list on every schedule_event, request_advance and take_up_grant. That is quadratic in a module whose PR body measures a 3.3x hoist at 65 participants to keep a single walk under 2 ms.

self._simulated_seconds(timestamp, "timestamp") at :380 is where it belongs — refusing +inf there costs the same line as refusing -inf does now, and it is the honest answer: "an event that never happens" is not an event, and a caller that means "I know of nothing" says that by not calling.

Principle 6.

bookkeeping instead of through the rule.
"""
self._participant(lp_id)
if self._status[lp_id] is LpStatus.GRANTED:

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Round-1 finding, partly answered. The harm it named is gone — _declared is a separate record now, so a second request_advance can no longer raise a horizon a peer lowered; I checked, and an event accepted at 1.0 s survives request_advance(engine, 5.0) followed by request_advance(engine, inf).

The call is still accepted, and it is still out of protocol: the state machine in the PR body has no BLOCKED_ON_MESSAGE --request_advance--> ... edge. Only GRANTED is refused here. Non-blocking, unchanged in substance from round 1: one branch, and CA-3 and CA-4 get told rather than diverging quietly.

class LpState:
"""One participant's row: its clock, its horizon, and what it is doing.

`next_event` is the earliest future event the participant knows of, and is

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Unchanged from round 1, and now further from the truth than it was then.

Round 1 asked for this sentence to be corrected because next_event is the earliest event the clock knows of, not what the participant knows. This PR makes that split explicit in authority.py — _declared is what the participant said, _accepted is what peers placed on it, next_event is their minimum — and the field this docstring describes is the minimum, i.e. exactly the larger set round 1 named.

This docstring is what a CA-3 or CA-4 author reads before writing the call, and the whole of the blocking finding in authority.py is about which of those two sets a given line is operating on. One sentence.

Principle 7 is not the right number here; it is just that the two-record design is now load-bearing and the public type does not mention it.

when = clock.now(actor) + floor_out + ((state >> 4) % 5) * 0.25
clock.schedule_event(actor, target, when)
pending[target].append(when)
pending[target] = [

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This line is why the fuzzer reports 0. It deletes exactly the class of event the remaining defect is made of.

Two things compound:

  1. clock.schedule_event(...) at :447 returns grants and the return value is dropped, so the shadow ledger never drains on a grant an event released.
  2. This line compensates for that by clearing every pending entry at or before the target's clock — which is also every event accepted onto a participant that is currently GRANTED, because _resolve has already moved its clock to advance_to.

So the ledger forgets precisely the events that the implementation also forgets, and the assertion compares two records of the same omission. Measured, on the same generator, floors and action stream:

oracle step-overs in 400 runs
as written 0
this line removed, schedule_event's grants still dropped 32 — but 14 of them are false, an artefact of (1)
schedule_event's grants observed and each grant releasing only what it was computed against 18, all true

The mutation table in the PR body is real — I did not re-run it, and the three mutants it names would still be caught. A fuzzer can be non-vacuous against the mutants its author imagined and blind to the case none of them models; that is what happened here.

The invariant to assert is the one the docstring already states — "must not be carried past an event that has already been accepted on it and that it has not yet been released to reach" — with "released to reach" resolved at the moment the grant was computed, not at the moment it is collected.

def test_the_safety_check_is_not_an_assert_statement():
# `python -O` removes an `assert`. A check that is specified as always on,
# not behind a flag, cannot be one -- so the module contains none at all.
tree = ast.parse((CLOCK_PACKAGE / "authority.py").read_text())

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Round-1 finding, not taken and not answered in the PR body. Still authority.py only, while the property the PR body claims is "the module has no Assert node" and the property that actually matters is the package one.

state.py is now a second production module in this package and it is not covered by this test. CA-1 made its two source guards package-wide and parametrised over clock/*.py for exactly this reason — and the +4 in this PR's own gate delta is those two guards picking up authority.py and state.py automatically. This is the third guard, it has the sharpest failure mode of the three (python -O silently deleting a check specified as always on), and it is the one that has to be edited by hand every time a module is added.

Same glob, same line count.

# --- determinism -------------------------------------------------------------


def test_grants_come_out_in_participant_order_not_arrival_order():

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Round-1 finding, not taken. All three participants still declare 5.0 at the same floor, so every permutation resolves in one pass and only order within a resolve is covered.

Restating the measurement rather than the argument: four participants, unequal floors and unequal horizons, driven to completion under all 24 arrival permutations give 24 distinct grant ledgers and 1 distinct (virtual time, participant) event set of 1024 events. This file asserts the ledger, which is the stronger claim and the one that is false in general; the design asks for the event sequence. Non-blocking, and the same sentence as round 1 — a second case on an asymmetric configuration asserting the event sequence would pin the property that holds.

It is also the property T84 is registered about, so the test and the register row currently disagree about which invariant is the real one.

Comment thread atom/compass/design/12_open_items.md Outdated
| **T49** | The prefix-index *lookup* cost is charged to nobody — ~1,387 blocks hashed and probed per request at the cc-traces p50, magnitude unmeasured | `03` |
| **T50** | Whether runtime memory constants transfer across dies (the working assumption says yes within a software generation) | `03`, `05` |
| **T53** | Whether tokenizer throughput transfers across CPU classes (the working assumption says yes, adjusted by derate) | `05` |
| T83 | **The virtual-time protocol has no way for a logical process to say it has finished, so a clean end of run is indistinguishable from a deadlock.** Opened 2026-09-21 by CA-2 (#45, PR #59). `01` D3's deadlock invariant is "all LPs blocked and none holding a finite `next` -> abort loudly", and the Clock Authority implements it exactly. But that state is also what a *correct, completed* run looks like — every LP has drained its work and parked with an empty horizon — so an integration reaches it at the end of every clean run and aborts. Asserted as the specified behaviour in `tests/compass/test_clock_grant_rule.py::test_everyone_waiting_with_no_known_event_aborts_loudly` and `::test_one_participant_with_nothing_left_to_do_is_the_same_stall`. The gap is not theoretical: a harness written while measuring the row below hung rather than finishing, because an LP that had drained its own events had no third option — it could neither park (which would abort the run) nor keep asking (which pins every peer at its stale clock), and the run spun with no grant issued and no abort raised. The item is the missing third case: either a retire call on the protocol, so a departed LP leaves the minimum and the rest carry on, or a stated convention that the run is torn down before the last LP parks. It is **not** a choice the clock can make alone, because "the workload is over" is a statement about the traffic source and the engines, not about time. Owner: whoever takes the transports (#46) and the traffic source (#47); CA-2 deliberately did not invent one, since the wrong convention here is more expensive than the abort. The abort must not be softened into a timeout while this is open — that invariant exists because a prior 120 s arrival barrier released on an invalid run and the client reported "0 failed". | `01` |

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocking, and visible only on the merged tree: T83 and T84 are already allocated.

P0.6 (#51) landed on feature/atomcompass_new before this PR's base did, and it registered T83–T87. The integration head f67618eb9 already carries:

T83  The EP group's size and moe_parallel_config.ep_size are two different numbers ...
T84  `15` D94's M1 EP2 leg is degenerate as written ...
T85  Multi-node EP rank-to-node mapping is assumed, not verified ...
T86  aiter is a second executed source root, and no artifact key names it ...
T87  `07`'s price-list table states the MoE all-to-all's block-count cap as one number ...

This branch's base is 5b11e82ca, which does not contain c19710bcd (#51), so nothing on the branch can see the collision. Cherry-picking 4c538fdb1 onto f67618eb9 conflicts in both 12_open_items.md and README.md, in exactly the register-count paragraphs:

<<<<<<< HEAD
   3. TODO register - 86 rows, T1-T80 and T82-T87, of which 80 are open
=======
   3. TODO register - 83 rows, T1-T80 and T82-T84, of which 78 are open
>>>>>>> 4c538fdb1

Resolved as a union the register holds two T83 rows and two T84 rows with unrelated content, and 83/78, 86/80 and the counting rule the paragraph itself states (grep -oE '^\| *~*\**T[0-9]+' over section 3) all disagree. The true merged figures are 88 rows and 82 open. The rows themselves are good — T83 in particular is the strongest thing in this PR's documentation, and the harness that hung on it is the right kind of evidence — they just need T88 and T89, and the two count lines need to be written against the integration head rather than against the base.

It also bites T84's attribution, which is otherwise accurate: I checked the numbers against round 1's review and "12 distinct totals spanning 1457 to 1484" and "one distinct set of 1024 events" are exactly what was measured and are correctly credited — but on the merged tree "T84" names an EP topology item, so the citation lands on the wrong row.

Process, not a finding: this PR's base branch is still compass/ca-1-clock-identity, which merged at 14:46 today. It needs retargeting to feature/atomcompass_new before any of this can land.

Principle 8: the counts are the measurement, and they are stated against a tree that is no longer the one this lands on.

truncates more grants, so the crawl is finer. The rule a transport has to implement is
therefore: *an LP asks again when it has work or has been refused, and an LP with nothing
to do parks rather than stepping its clock forward one floor at a time.* That decision, not
PP degree, is what decides whether the grant traffic is affordable.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No blank line between the end of the sizing amendment and **PP is therefore an efficiency concern ...** on the next line, so Markdown renders the amendment's closing sentence and D3's original PP conclusion as one paragraph. (There is also a doubled blank line above the amendment heading.) One newline.

The amendment itself reproduces to the digit — two LPs, one event at 20 s, identical matrices, my own driver:

discipline 1 ms floor 1 us floor
idle LP asks until refused, then parks 3 grants, reached 20 s 3 grants, reached 20 s
take up and immediately re-ask 20,002 grants, reached 20 s 2,000,000 grants (cap), reached 2 s

The inversion holds and it is the right thing to have written down: the efficient discipline is lookahead-independent, and a tighter floor makes the inefficient one strictly worse. Worth noting for whoever implements the transport that the efficient discipline is not "park instead of asking" — it is ask, be refused, and stay refused, which needs the idle LP to ask twice before the busy one asks at all. That ordering is the whole of the three-grant result and nothing in D3 or in the module says it.

@jgong5

jgong5 commented Sep 21, 2026

Copy link
Copy Markdown
Owner Author

Review record — round 2, agent-authored

Range: 4c538fdb1 against its base 5b11e82ca, and the merged tree (4c538fdb1 cherry-picked onto the integration head f67618eb9) as 92121524c. The head at the time of writing is 4c538fdb1. The branch was force-pushed, so round 1's head no longer exists and this is not a diff against it — every round-1 finding below was re-checked by re-running round 1's own reproduction, not by reading a diff.

Verdict: CHANGES REQUESTED. Two blocking findings, both new, neither of them the one round 1 raised:

  1. authority.py:350 — take-up drops every accepted event at or before the clock, including events a peer wrote into _accepted after the grant was computed. A grant truncated by the bound still erases an event landing on its end and walks past it, silently. 18 step-overs in 400 legal runs, 0 with the snapshot fix.
  2. 12_open_items.md:226 — T83 and T84 are already allocated. P0.6 (compass: state EP group membership per configuration (P0.6, T65) #51) registered T83–T87 and landed on feature/atomcompass_new before this branch's base did. This PR does not apply to the integration head: the cherry-pick conflicts in both README.md and 12_open_items.md, on the count paragraphs themselves.

need human applied, under AI_DEV_RULES.md's review-loop stop. Reasons in The loop stop below. GitHub refuses APPROVE / REQUEST_CHANGES on a self-authored PR, so the verdict is that line.

Round 1's blocking finding is fixed, and the fix is right. That is the first thing to say and it is not grudging — the two-record horizon is a better answer than the one round 1 suggested, the author's reasoning for rejecting the suggestion is correct, and I confirmed the rejection by measurement rather than by reading it. The remaining hole is one window further out again, in the same three lines. That recurrence, not any single defect, is what the loop stop is for.


Round 1's blocking finding: fixed, verified by re-running its own reproductions

round 1's reproduction round 1's result at 4c538fdb1
zero floors; schedule_event(traffic->engine, 0.0), request_advance(engine, 10.0), request_advance(traffic, inf) engine 0 s -> 10 s over the event, nothing raised engine held at 0.0 s; the event releases it as a zero-span grant engine 0s -> 0s (bound 0s by traffic-source)
1 ms floor; traffic at 0.001, engine at 0.002, event at 3.0, engine asks 10.0 engine 0.002 s -> 10 s, seven seconds past the event engine 0.002s -> 3s (bound 3.002s by traffic-source) — exactly the answer round 1 named as correct
the freeze that min-only produced clocks 1, 1, 1, 1, 1, 1 1, 2, 3, 4, 5, 6

The author's claim that the suggested fix was incomplete is correct, and the third case they found — a scalar horizon holding only the earliest of two events in flight — is real. The two-record split is the right shape.

The blocking finding: what survives it

_resolve moves _now at grant issue; _accepted is drained at take-up. Anything a peer schedules in that window is appended after the grant was decided and is then dropped by a threshold that was computed without it. Detail, the non-zero-floor reproduction and the fuzz numbers are inline at authority.py:350.

Why the fuzzer reports 0 while this is present — test_clock_grant_rule.py:449. The shadow ledger drops the grants schedule_event returns, and the ts > clock.now(target) line that compensates for that erases exactly the events accepted onto a GRANTED participant. The ledger forgets what the implementation forgets, and the assertion compares an omission with itself. The mutation table is not thereby worthless — the three mutants it names would still be caught — but a fuzzer can be non-vacuous against the mutants its author imagined and blind to the case none of them models.

What I could not break

Everything else I threw at the two-record horizon held, and I went looking specifically where the task pointed.

attack result
three events in flight on one participant from three different peers _accepted == [3.0, 1.0, 2.0], reached in order 1.0, 2.0, 3.0, list drained to empty
take-up out of registry order (o, m, n on a three-participant resolve) no effect; each participant's horizon restated correctly
accept an event, park, be granted, take up, accept another, park, be granted again clean; _declared and _accepted stay separate across the cycle
out-of-protocol second request_advance while BLOCKED_ON_MESSAGE the event a peer lowered survives it — round 1's named harm is gone; see :306 for what is left
does anything iterate _accepted order-dependently no — _restate_horizon is a min and take-up is a filter; both are order-invariant
can _accepted grow without bound yes, on +inf — see below

+inf is the one that broke, and it is a round-1 finding: authority.py:409. 20,001 calls to schedule_event(y, x, math.inf) leave 20,002 entries that no grant can ever drop, and one further call then costs 1.716 ms because _restate_horizon re-mins the list. Round 1 called this a silent no-op needing one line of guard; with _accepted as a list it is quadratic growth in a module that hoisted a row to keep one walk under 2 ms.

The D3 amendment, judged as a document change

The grant-rule amendment is right, and it is in the right place. Round 1 said a PR-body paragraph does not guarantee a design correction and named 12_open_items.md §4 and T82 as the shape to copy. The author went further and amended D3 itself, which is better: the document no longer states a rule that cannot run a configuration the same document requires. Both false claims are named as false, the recurrence argument is stated correctly (the minimum is over clocks, not next events, so it is not a start-up wrinkle), and the reviewer's measurements are cited as the reviewer's. The replaced paragraph is left in place with a pointer rather than deleted, which is the honest edit.

The sizing amendment inverts D3's intuition, and the inversion holds. Re-measured with my own driver, two LPs, one event at 20 s, identical matrices — every figure reproduces to the digit:

discipline 1 ms floor 1 us floor
idle LP asks until refused, then parks 3 grants, reached 20 s 3 grants, reached 20 s
take up and immediately re-ask 20,002 grants, reached 20 s 2,000,000 (cap), reached 2 s

Lookahead-independent in the first row, and a tighter floor makes the second strictly worse. One thing the amendment does not say and a transport author needs: the efficient discipline is not "park instead of asking", it is ask, be refused, stay refused — the idle LP has to ask twice, and get a one-floor crawl grant the first time, before the all-parked state that produces the jump is reachable at all. That ordering is the whole of the three-grant result. 01:520, where there is also a missing blank line that merges the amendment into D3's PP conclusion.

T83 is the strongest piece of documentation in the PR. A registered row for a gap the author hit in their own harness, with the harness hang recorded as the evidence, and a deliberate refusal to invent the convention. That is the right call and the right record. T84's attribution is accurate — I checked "12 distinct totals spanning 1457 to 1484" and "one distinct set of 1024 events" against round 1's review and both are exactly what was measured and are correctly credited to the reviewer rather than claimed. Both rows are undermined only by the number they were given; see the blocking finding.

Gates — three trees, including the merged one

Node 18, xiaobizh_n18_cpu. All three staged by scripts/compass/snapshot.sh (so .compass-commit and .compass-changed come from one rev-parse), docker cp into /opt/ca2rev2/ of my own, tarball md5 matched on both ends, import atom asserted under each tree before any count was read. The shared mount was not touched and no rsync was used.

Control Branch Merged
Commit 5b11e82ca 4c538fdb1 92121524c = f67618eb9 + this PR
pytest 4084 passed, 149 skipped, 3 xfailed 4131 passed, 149 skipped, 3 xfailed 4131 passed, 149 skipped, 3 xfailed
pytest: rc= 0 0 0
GATE_CPU_RC= 0 0 0

+47, nothing else moved — reproduces exactly, and so does the decomposition by node id: test_clock_lp_identity.py 49 -> 53, the four new ids being test_the_package_imports_only_the_standard_library_it_names[authority.py], [state.py] and test_the_package_builds_no_set_at_all[authority.py], [state.py]; test_clock_grant_rule.py collects 43 with no parametrize. 43 + 4 = 47.

The merged tree is green. CA-1's two package-wide globs over atom/compass/clock/ pick up authority.py and state.py and pass on the merged tree as well as on the branch, so the failure mode that has CA-0 red does not occur here. The merged tree is genuinely a different tree — the branch's base does not contain P0.6 (#51) — which is also how the T83/T84 collision was found, and is the argument for gating the merged tree rather than only the branch.

The author's diagnosis of round 1's GATE_CPU_RC=98 is confirmed: stamping with snapshot.sh gives 0. Round 1's 98 was its staging, as round 1 said.

ruff check and black --check clean on all four touched files, rc=0, in the same container.

Effort — reproduces exactly

My own ast.parse / strip module-class-function docstrings / ast.unparse / count non-blank:

file ast non-blank raw
clock/state.py 36 81 104
clock/authority.py 199 457 523
production total 235
clock/__init__.py (whole file) 6 40 44
tests/compass/test_clock_grant_rule.py 385 570 700

235 against 250–350 is a ~6% underrun; the halt rule is about overruns and CA-1 took the same posture. git diff --numstat on __init__.py confirms 20 6.

Cost at scale — grants really are identical

Confirmed the thing the hoist has to be judged on. _row replaced with a mapping that rebuilds the inbound row on every access (the pre-hoist behaviour: peers() rebuilt, (LpId, LpId) re-hashed per element), driven over the same workload:

participants grants compared identical
3 21 yes
10 70 yes
65 455 yes

Rule arithmetic with every participant parked, my container: 0.032 -> 0.009 ms at 3, 0.252 -> 0.070 at 10, 9.575 -> 2.511 at 65 (3.8x). The PR's 5.631 -> 1.682 and round 1's node-18 8.49 -> 1.59 are the same shape at a different relaxation depth; the ratio band 3.3–5.3 is consistent and the absolute numbers are not comparable across containers, as the PR body says. Still ~50x D3's 0.050 ms per-grant RPC estimate at 65, which remains T70/PP territory.

Round-1 findings that survived

round-1 finding state at 4c538fdb1
schedule_event(..., +inf) accepted, "one line, same guard" not fixed, and now worse — unbounded _accepted growth, quadratic restate. :409
LpState.next_event docstring no longer describes the field not fixed, and the two-record split makes it more wrong. state.py:55
the Assert-node test covers authority.py only, not the package not fixed. :343
the determinism test's configuration is symmetric and hides the property not fixed. :536
a second request_advance while BLOCKED_ON_MESSAGE is out of protocol harm fixed, call still accepted. :306
a negative start_time is accepted; worth a word either way no word, in code or in the PR body

None of these is individually blocking and round 1 said so. They are listed because the loop stop counts findings, not severities.

The loop stop

AI_DEV_RULES.md: "If the same finding survives two cycles, or the loop passes three cycles, it halts and goes to the owner and applies need human to the PR: a task that cannot converge is mis-cut, not under-worked."

Applied. Three things, in the order they weigh:

  1. The defect class has not converged. Round 1 found one silent erasure of an accepted event in this bookkeeping. Fixing it surfaced a second (the suggested fix losing an event beyond the grant) and a third (one slot holding only the earliest of two). This round finds a fourth, in the same three lines. Four cases of one shape across two cycles is the signal the rule names, and it is not a criticism of the work — each fix was correct and each was found by measurement. It says the seam between "what the clock knows" and "when the clock forgets it" wants settling once, as a stated contract with a snapshot, rather than case by case.
  2. A round-1 finding survived and regressed. The +inf guard was named with its one-line fix; not taking it was reasonable when _accepted was a scalar and is not now.
  3. The T83/T84 collision is not resolvable inside the review loop. T-numbers are allocated across parallel branches, two branches have now allocated the same two, and which pair renumbers is an owner's call, not a reviewer's finding.

I am not escalating the blocking bookkeeping finding as such — it is an actionable review finding and the fix direction measured clean at 0/400.

What CA-3, CA-4 and CA-7 should watch

  1. The horizon contract, restated. Round 1 asked whether a participant must re-declare events a peer scheduled on it; the answer is now "no, the clock keeps them" — except in the grant window. CA-3 and CA-4 must not assume next_event is complete between _resolve and take_up_grant.
  2. Ask, be refused, stay refused — and ask twice. Three grants against 2,000,000 on one topology. It belongs in the transport's contract, and the "twice" is load-bearing.
  3. ClockDeadlock at the end of every clean run until T83 has a convention. Decide it; do not catch it.
  4. A grant count is a property of a run, not of a configuration (T84). Keep it out of anything byte-diffed.
  5. Gate the merged tree, not only the branch. Two of this round's findings — the T-number collision and the merged-tree green — exist only there.
  6. The rule is still the cost at scale, not the transport, at ~50x the design's per-grant RPC estimate at 65 participants with the hoist in.

What I could not check

  • The GPU tier. Correctly not required, not run, no claim made.
  • The author's mutation table. I did not re-run the three mutants; I measured the oracle's blind spot instead, which is a different question and does not contradict it.
  • Any real participant. There is none. Everything here is the rule driven by harnesses I wrote.
  • Whether the grant-window race is reachable from ATOM's real topology. It needs a peer whose clock plus floor lands exactly on a truncated grant. I produced it at a tie and at asymmetric non-zero floors; whether ATOM's link structure produces ties is a CA-4 question.
  • Whether P0.6's T83–T87 or CA-2's T83–T84 should renumber. Owner's call.

Round 2. need human applied per the loop stop; the label is on the PR and stops further agent action until the owner removes it.

…rule

Base update authorized by the owner on 2026-09-24 ("go ahead and update
the clock chain"). The PR keeps its need human label. This commit only
resolves the merge.

Resolved files:
- atom/compass/clock/identity.py, lookahead.py, registry.py,
  tests/compass/test_clock_lp_identity.py,
  tests/compass/test_clock_order_across_processes.py: took the
  integration side whole. This branch carried each one byte-identical to
  the pre-#410 tip (the #52 landing), so every differing line came from
  #410 or #417.
- atom/compass/clock/__init__.py: this branch's docstring, imports and
  __all__ (ClockAuthority, the state types), plus the paragraph #53 added
  on the tip. That paragraph holds for authority.py and state.py: both
  import only math, enum and dataclasses, and build no set.
- atom/compass/design/12_open_items.md: kept both sides' rows. The tip
  had meanwhile landed P0.6's T83 and T84, so this branch's two rows are
  renumbered T89 (no retire call on the protocol) and T90 (grant count
  depends on arrival order), and T90's pointer to "T83 above" now reads
  T89. The register header now reads 90 rows, T1-T90 with no gaps, 84
  open, counted with the header's own grep rule; the allocation sentence
  names T88 with #196 and T89-T90 with CA-2.
- atom/compass/design/README.md: the TODO count line reads 90 registered,
  T1-T90 with no gaps, 84 open, with the tip's struck list.

No member removed by #410 appears anywhere in the merged tree.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
jgong5 added a commit that referenced this pull request Sep 24, 2026
Base update authorized by the owner on 2026-09-24 ("go ahead and update
the clock chain"). This commit only resolves the merge, which brings in
#59's merge of fork/feature/atomcompass_new (3a3267a).

Resolved files:
- atom/compass/design/12_open_items.md: conflicted on the two TODO rows
  this branch edits. #59's merge renumbered them T83 -> T89 and T84 -> T90,
  because the tip had landed P0.6's T83 and T84. Kept this branch's edited
  text of both rows under the new numbers, after the tip's T88. T90's
  pointer to "T83 above" now reads T89. This branch's T70 edit merged
  cleanly. The register still counts 90 rows, T1-T90, 5 struck.
- atom/compass/design/01_execution_and_time_model.md (merged without a
  conflict, edited here so it stays true): this branch's two citations of
  those rows now use the new numbers, "(T90 records CA-2's reviewer ...)"
  and "T89: a run whose work is finished ...". The other T83/T84 mentions
  in the tree are P0.6's and are unchanged.

No member removed by #410 appears anywhere in the merged tree.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
jgong5 added a commit that referenced this pull request Sep 24, 2026
…arness

Base update authorized by the owner on 2026-09-24 ("go ahead and update
the clock chain"). This commit only records the merge, which brings in
#59's merge of fork/feature/atomcompass_new (3a3267a).

Resolved files: none. The merge had no conflicts, and no file was edited
beyond what git merged. This branch's own files (tests/compass/clock/)
read TRAFFIC_TO_ENGINE_FLOOR_SECONDS and the LinkClass members, all of
which #410 kept.

No member removed by #410 appears anywhere in the merged tree.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@jgong5

jgong5 commented Sep 24, 2026

Copy link
Copy Markdown
Owner Author

Base update (owner-authorized, 2026-09-24)

The owner authorized this update on 2026-09-24 ("go ahead and update the clock chain"). It only merges the integration branch in. Labels, draft state and review state are unchanged, and nothing was rebased or force-pushed.

What was merged: freshly fetched fork/feature/atomcompass_new at 3a3267a84.

  • Old head: 4c538fdb1
  • New head: a386039dd, the merge commit itself (parents 4c538fdb1, 3a3267a84)

Conflicted files and how each was resolved (8)

File Resolution
atom/compass/clock/identity.py, lookahead.py, registry.py Took the integration side whole. This branch carried each one byte-identical to the pre-#410 tip (the #52 landing; checked with git diff against f67618eb9 and 9a6317927~1). So every differing line came from #410.
tests/compass/test_clock_lp_identity.py, tests/compass/test_clock_order_across_processes.py Took the integration side whole, for the same reason. The differing lines are #410's and #417's.
atom/compass/clock/__init__.py This branch's docstring, imports and __all__ (ClockAuthority, BackdatedEvent, ClockAbort, ClockDeadlock, Grant, LpState, LpStatus), plus the paragraph #53 added on the tip: every .py under the package is held to {dataclasses, enum, math} and builds no set. That paragraph holds for this branch: authority.py imports only math, state.py only enum, math and dataclasses, and the tip's rglob-parametrised tests pass on both new modules.
atom/compass/design/12_open_items.md Kept both sides' rows, and renumbered this branch's two. Meanwhile the tip landed P0.6's own T83 and T84 (the EP-group items). So this branch's rows are now T89 (the protocol has no retire call) and T90 (the grant count depends on arrival order). T90's pointer to "T83 above" now reads T89. The header reads 90 rows, T1–T90 with no gaps, 84 open, counted with the header's own grep -oE '^| *~*\**T[0-9]+' rule over section 3 (90 rows, 5 struck, T77 closed). The allocation sentence now names T88 with #196, and T89–T90 with CA-2. No open PR on the fork claims T89 or above.
atom/compass/design/README.md The TODO count line now reads 90 registered, T1–T90 with no gaps, 84 open, with the tip's struck list.

01_execution_and_time_model.md merged without a conflict. The tip's hunks there are citation-line fixes that do not overlap this branch's D3 amendment.

The resolutions are all in git show --remerge-diff a386039dd.

Removed-member grep

I grepped atom/, tests/, scripts/ and the design docs of the merged tree for every member #410 removed: .tightest, serializing, .links(, _missing_into, scale_seconds, ._ordered, len(matrix), LinkClass(…), LinkClass.X.label/.value.

  • 0 hits in the merged tree.
  • 21 hits in an archive of the old head 4c538fdb1, which shows the grep fires.

Gate (node 18, xiaobizh_n18_cpu)

The tree was a git archive of a386039dd, with .compass-commit and .compass-changed written from the same ref. It ran with its own scripts/compass/gate_cpu.sh, bounded by timeout -k 10 3600 and not piped, one gate at a time. atom.__file__ resolved under the staged root.

Tree commit: Result GATE_CPU_RC
new head a386039dd a386039dd (stamp) 5325 passed, 155 skipped, 3 xfailed (193 s) 0
  • gpu: not required.
  • No failures, so no timing flake fired.

Collected node ids (gate_cpu.sh --collect-only)

  • Old head → new head: 4262 → 5462.
  • New head against the tip (5415): +47, −0.
    • The +47 are this PR's test_clock_grant_rule.py (43), plus 4 cases of the tip's two rglob-parametrised package tests over authority.py and state.py.

No blocking issues from this update.

@jgong5

jgong5 commented Sep 24, 2026

Copy link
Copy Markdown
Owner Author

Base-update resolution review

Verdict: the resolutions in merge a386039dd are sound. There are no blocking issues. One non-blocking finding is posted inline: a sentence the resolution wrote into 12_open_items.md gets the order of events wrong.

This review was written by an agent. It covers only the conflict resolutions of the owner-authorized base update (git show --remerge-diff a386039dd), not the PR's own content. The need human label stays.

I read AI_DEV_RULES.md at the tip, the base-update exception in #437, the Design principles in design/README.md, and the #410 review before the diff.

1. Resolutions (parents 4c538fdb1 ← PR, 3a3267a84 ← tip; merge base 7fc7a5ddd)

I compared the patch for every file, diff(tip, merge) against the PR's own diff(base, 4c538fdb1). 4 of 13 files are identical patches. The other 9 are the 8 conflicted files, plus atom/compass/__init__.py, which both sides add byte-identically.

File Check Result
clock/identity.py, lookahead.py, registry.py, tests/compass/test_clock_lp_identity.py, test_clock_order_across_processes.py blob at 4c538fdb1 vs f67618eb9 (#52) and 9a6317927~1 (pre-#410); blob at merge vs tip PR side equals the pre-#410 tip in all 5 files, so the PR's diff changed none of these lines. The merge blob equals the tip blob in all 5. Toward the integration side, correctly.
clock/__init__.py 3-way against #52's version 46fe09db The merge is exactly the union: the PR's docstring, authority/state imports and __all__, plus #53's enforcement paragraph. No tip line dropped.
design/12_open_items.md word diff of tip → merge Only additions: rows T89/T90, the counts 88→90 and 82→84, "; T89 and T90 make it 90", and the allocation sentence. All of the tip's text is kept. The one tip phrase reworded, "and T88 arrives here", becomes "T88 with #196". That keeps the fact, because the phrase is no longer true on this branch.
same, rows PR rows vs merged rows After renaming T83→T89 and T84→T90 and "T83 above"→"T89 above", the merged rows equal the PR's rows byte for byte.
design/README.md tip → merge The count line changes and nothing else. The tip's struck list (including T65) is kept.
01_execution_and_time_model.md patch-id Identical to the PR's own diff. It merged cleanly, with no hand edit.

Did any tip change get dropped? No, in any hunk.

2. Removed-member grep

  • What I grepped: git grep over the merge tree (all paths, *.md included) for .tightest, serializing, .links(), _missing_into, scale_seconds, ._ordered, len(matrix…), LinkClass(, LinkClass.X.label/.value/.scale_seconds, <link>.label, repr/str of a registry, and truth tests on a matrix.
  • Result: 0 hits beyond the tip's own 13. The 13 are ATOM comments containing "serializing", len(m) in unrelated code, and class LinkClass(. The hit list is identical to the tip's line for line.
  • The grep fires: 35 hits at the old head 4c538fdb1.

3. clock/__init__.py docstring against this PR's modules

  • Imports: I parsed every .py under atom/compass/clock/ at the merge with ast. authority.py imports math plus .identity, .lookahead, .registry and .state. state.py imports dataclasses, enum and math plus .identity. Neither has a stray module against {dataclasses, enum, math}.
  • Sets: no set/frozenset call, Set or SetComp anywhere.
  • The rest of the docstring: "the rule that reads all three" is true, because authority.py imports LpId, LpRegistry and LookaheadMatrix. "The transport … sits elsewhere" is true: nothing in the package is a transport.
  • Tests: the tip's rglob-parametrised tests collect [authority.py] and [state.py] cases, and they passed in the gate below.

4. T83/T84 → T89/T90 renumbering

5. Gate

I read n59.gate.log in full and did not re-gate.

  • Run details: commit: a386039dd (stamp), atom.__file__ under /tmp/xiaobizh_chainupd/n59/ATOM, gpu: not required, timeout -k 10 3600, not piped, and a pgrep wait with a re-check before launch.
  • Result: 5325 passed, 155 skipped, 3 xfailed, GATE_CPU_RC=0.
  • Node ids from old head to merge: exactly compass(clock): apply the ponytail-audit of atom/compass/clock/ (#406 phase 2) #410's 7 removed. All 1207 added ids are collected at the tip. Against the tip it is +47 and −0.

Blocking: none.
Non-blocking: 1, inline on 12_open_items.md line 26.

T88 arrives here. It happens to be complete at this head; that is an observation about
what has landed, not a property to rely on
branches was in flight: T73–T76 arrived with P0.3, T83–T87 with P0.6, T81 with P0.4,
T88 with #196, and T89–T90 with CA-2, which first numbered them T83–T84 before P0.6's

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Non-blocking (base-update resolution): the order of events here is wrong.

"CA-2, which first numbered them T83–T84 before P0.6's landed under those numbers" does not match the history:

So P0.6's T83/T84 landed first. CA-2 reused the numbers because its branch was cut from 7fc7a5ddd, before P0.6 landed.

Suggested wording: "…and T89–T90 with CA-2, which had numbered them T83–T84 on a branch cut before P0.6's rows landed under those numbers."

This wording carries into #67 and #75 through this merge. A base update may change nothing beyond its conflicts, so under need human the fix waits until the owner lifts the label.

jgong5 added a commit that referenced this pull request Sep 24, 2026
…ers the REST base patch (#434)

Review cycle 1 on #437.

- A merge that cannot keep both sides' changes in a conflict hunk now
  commits nothing and names the hunk in a PR comment. The old text
  ("toward the integration side unless the PR's own diff changes those
  lines") gave no side for the usual case, so a literal reader could
  drop a change the tip made.
- The exception now covers the REST base patch that an unlinked child
  whose parent landed also gets (live on #61; #63 once #59 lands).
- The merge sources are no longer restated; the branch-update rule
  names them.
- Deleted "It lands nothing, reviews nothing and leaves the label." and
  the back-reference in the branch-update rule. Both restated other
  text, and the second read as a duty.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@jgong5

jgong5 commented Sep 28, 2026

Copy link
Copy Markdown
Owner Author

Design-alignment review against the PDES time model (#443) — agent-authored, requested by the owner.
Verdict: CHANGES NEEDED — the grant is partial and non-strict, a stall aborts instead of being recovered or ending the run, and the clock tracks events by (target, timestamp) with no send log and no (channel, seq); these are the rules v0.20 replaced, and three open findings on this PR (round 2's step-over, T89, #94's crawl) all come from them.

This is a design-alignment review, not a gate-4 round. It approves nothing and changes no label. The design text is carried by the doc PRs under #443 (doc 01 §D1-§D3.5 is #446; §D4, §D5, §D9 and the log are #450; §D6-§D8 are #448); section numbers below are the v0.20 design's. Head a386039dd.

Measured at the head (git archive of a386039dd, local container, a 40-line probe driving ClockAuthority directly):

P1  idle asks next_event=inf while busy is RUNNING at 0, 1 ms floors   -> idle 0s -> 0.001s (bound 0.001s by busy)
P4  engine asks for a 10 ms step, frontend RUNNING at 0, 1 ms floors   -> engine 0s -> 0.001s (bound 0.001s by frontend)
P2  round 2's case, L[p->q]=4, L[q->p]=1: p granted 0 -> 5 (bound 5), q schedules on p at 5.0, p takes up
                                                                        -> p next_event 10.0; the 5.0 event is gone

Corrections

  1. A grant is partial and reaches the bound inclusively. atom/compass/clock/authority.py:428-443, specifically :430 advance_to = min(horizon, bound). A waiting participant is granted whatever part of its request the bound allows (P1, P4), and a grant may land exactly on the bound (P2). The design (§4.1 try_grant/grant, §4.2, §4.2.2) grants a waiting LP i only when N_i < LBTS_i, strictly, and then grants exactly N_i: the TAR target, or for an NER the smaller of its t and the earliest registered arrival. Otherwise the request is held and nothing is sent. Three things follow from this one change:

    • Round 2's blocking step-over has no window left. After a strict grant to G, every later message has arrival >= N_j + D(j->i) > G, so the take-up threshold never meets an event that arrived after the grant. The zero-lookahead case goes through the recovery branch in correction 3.
    • The grant discipline that terminates is a global service order, and only the clock can enforce it #94's crawl goes away. An idle LP (NER(inf) with nothing registered) is never granted until a message is registered for it, so no one-floor increments are handed out.
    • A step is no longer split. A TAR is one compound event (§3.1.1, §4.7), and the grant to a TAR always equals T (§4.1 advance_to). P4 shows a 10 ms forward currently gets 1 ms.

    Smallest change: in _resolve, if n < bound: grant(n), otherwise leave the participant held.

  2. One request kind, and a fixpoint where the design uses a static distance. request_advance (authority.py:282-322) has no TAR/NER distinction. earliest_emission_times (:212-246) relaxes every parked participant to max(now, min(declared, accepted)) and then lowers it through its peers.

    • The design (§4.2.2 table) sets N_j by state: now[j] if running; the target T_j if waiting in TAR, by I1 and I4, whether or not messages are pending for it (they are released inside the grant and stepped through by _step_through, §4.4); min(t_j, earliest pending arrival) if waiting in NER.
    • The design's LBTS_i = min_j (N_j + D(j->i)) uses D, the all-pairs shortest path over channel lookaheads. D is computed once, because lookahead is static (§4.1 all_pairs_shortest).
    • The fixpoint is safe: round 1 verified it, and every term it adds is a path term that D already contains. But it stops a TAR at a pending message, which is what the horizon-min at :324-329 does. It also contributes that message's arrival to peers instead of T, so it holds back peers that §4.2.2 releases. The walk it needs is the 1.6-2.5 ms per resolve measured at 65 LPs in both review rounds. With a precomputed D, each bound is one pass over the peers.

    Change: add a kind (TAR | NER) to request_advance, compute N_j from the state table, precompute D, and delete the relaxation.

  3. A stall aborts. authority.py:450-451 and :471-494 (_refuse_to_stall raises ClockDeadlock). The behaviour is pinned by tests/compass/test_clock_grant_rule.py:370, :379 and :393. The design (§4.1 try_grant, §4.2.3) never aborts on a stall, and handles it in three branches:

    • Every LP is waiting, the strict rule grants none, and some N is finite: this is Chandy-Misra recovery. Grant the (N, LP id) minimum. It is reachable only through a zero-lookahead cycle.
    • Every N is infinite: this is the end of the run. Grant INF to all (finish).
    • The traffic LP calls end_workload (an END request): the run ends the same way. END is needed because periodic timers keep N finite forever (§4.14).

    A stall the CA cannot see is a fault. The design waits for it indefinitely and prints one diagnostic after DIAG_S = 30 wall seconds without a request. That wall-clock check belongs in the serving loop (compass: both clock transports from one implementation (CA-3, #46) #63), not in this package, which reads no clock. Change: replace _refuse_to_stall with the recovery branch and finish, add END, and drop ClockDeadlock. BackdatedEvent stays as the one abort. The 120 s-barrier concern behind the original invariant is still met, because nothing is ever released on a timer.

  4. The clock records events per target, with no send log and no (channel, seq), and there is nothing in the grant to deliver. schedule_event (authority.py:358-411) is a per-event call that appends a bare timestamp to _accepted[target] (:128, :409). take_up_grant (:331-356) drops entries by threshold, and Grant (state.py:71-91) carries no message set. The design works in four steps:

    • The send log [(channel, seq, arrival)] rides on the TAR/NER request and is registered before the state changes (§4.1 on_request, §4.2.1 second invariant).
    • The CA keeps arrivals[channel][seq].
    • A grant to G carries, for each inbound channel, the (seq, arrival) pairs with arrival <= G not yet promised (§4.1 grant, §4.2).
    • The receiver counts by (channel, seq), and FIFO is not assumed (§3.3).

    That promised set is the snapshot round 2 asked for, so it closes the take-up window structurally. Keep both BackdatedEvent checks per log entry, with the channel's lookahead as the floor. Keep the source-may-not-be-parked refusal (:381-387), which is I1/I4 seen from the CA. GRANTED and take_up_grant become unnecessary once the grant is the reply to the held request (see the compass: both clock transports from one implementation (CA-3, #46) #63 review). Depends on the channel table, compass(clock): a channel table with per-channel lookahead and path distances between LPs #453.

  5. The design amendments encode the rule v0.20 replaces. atom/compass/design/01_execution_and_time_model.md:409-464 holds the D3 fixpoint amendment and :499-520 the sizing "driver discipline" amendment. 12_open_items.md:238-239 holds T89 and T90.

Missing in scope

  • The END request and finish (correction 3). Add them here, since they are CA state-machine logic.
  • The zero-lookahead same-instant rule (§4.2.3). A message at a == now[dst] is accepted as the next round at that instant, and released in the next grant because it is not yet in promised. It becomes reachable once correction 3 lands. One test with two zero-lookahead LPs, both waiting in TAR at 10.3, recovered one at a time (§4.2.3 example), belongs here.

Filed separately

Matches

  • The safety check is a raised BackdatedEvent, not an assert, and it checks both ts >= now[src] + L and ts >= now[dst] (authority.py:388-408). This matches the design's CA check a >= now[dst] (§4.1) and principle 6.
  • A running peer contributes its current clock (:229-230). TestTheRuleReadsAClockAndNotAHorizon stays valid under v0.20, since its peer is running.
  • The peer set is every registered LP, not only direct senders (:160-170). So the removed "direct senders only" rule is not encoded here.
  • Grants are issued in LP-id order (LpRegistry.ids() sorts by name), which matches the LP-id term of the §4.8 tie-break. The channel and seq terms arrive with correction 4.
  • Nothing in the package reads a wall clock or times out (round 1's grep). That is consistent with "wait, never release on a timer" (§4.2.3).

@jgong5

jgong5 commented Sep 29, 2026

Copy link
Copy Markdown
Owner Author

Closed on the owner's ruling of 2026-09-29. Asked "for the stale PRs on the clock, shall we close them and start from scratch according to the updated PDES design, or revise them?", the owner answered: "close and re-cut". This PR is not revised; its scope is re-briefed against the v0.20 design (#443). Agent-authored.

Replacement: #480, the CA grant rule and state machine (strict LBTS over D, TAR/NER/END with the send log, recovery, end of run).

Carries over (pinned to a386039dd): ClockAbort / BackdatedEvent with the LP table, the no-assert AST test, both degenerate cases and the generated-run pattern. The partial grant, the fixpoint relaxation, the two-record horizon, take-up and ClockDeadlock are not carried. The design-alignment review above is the record of why.

The branch stays on the fork, so the head the salvage lists cite stays reachable. The need human label is left for the owner.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

need human Automation stopped; needs owner judgement. Agents apply when escalating, only the owner removes.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant