Conversation
Adds atom/compass/clock with three things and nothing that uses them: the identity of a logical process, the registry that holds a total order over those identities, and the lookahead matrix addressed by identity in both directions. The rule that reads them, the per-participant state and the transport are separate. The order is by name, which is a code-point comparison and therefore the same in every process. Arrival order is never consulted. The registry holds membership in a dict with None values rather than a set, and the package builds no set at all, so no ordered answer it gives can depend on a per-process hash seed. A lookahead floor is declared per directed link and carries the class of path it represents. An undeclared pair is refused rather than read as zero: zero is a decision to serialize a pair and stays correct, silence is a link nobody sized. Zero floors are accepted and reported by LookaheadMatrix.serializing. Tests are CPU-only. test_clock_order_across_processes.py spawns three interpreters with different PYTHONHASHSEED values, registers the same twelve names rotated differently in each, and compares the registry order against a plain set of the same names built in the same children as a control. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A caller minimises over one participant's whole row, quantified over every registered peer. inbound() filtered the row to the links that happened to exist, so a peer with no declared floor contributed no term at all: the minimum came out higher, not lower, and more virtual time was handed out than the peer allowed. Measured on a three-participant registry with one leg undeclared, the walk gave 5.002 s where the safe answer was 0.0 -- the grant lands past the peer's current time and the peer's next event lands in the past. An omission reads as an unbounded lookahead, which is the one direction that is unsafe; a zero would have been the conservative mistake. inbound() now refuses a row missing any registered peer, naming each missing pair, and require_complete() refuses an incomplete matrix at set-up so the failure happens before anything advances rather than one step later as a backdated event. tightest() says in its docstring that it reports over declared links only. InterLpLink is now a frozen dataclass. The accessors hand out the object itself, so a writable floor_seconds let a holder rewrite a declared floor and reach around every refusal declare() performs -- negative, NaN, infinite, already declared. LpId was frozen for this reason; the object carrying the number was not. LpRegistry.require() type-checks its argument. It is what declare(), lookahead() and inbound() funnel through, and without the check a bare "decode" was reported as not registered beside the registered decode. The source guards now rglob the package, so a subpackage added later cannot escape the import allowlist or the no-set check silently, and the allowlist drops typing, which nothing imports. A signature and field scan covers the requirement that no public call names a host, port, endpoint, address, url, level, depth, parent, rank, node or socket -- the import allowlist proves the package reaches no device, clock or socket, which is a different claim. The cross-process control now builds its set from the unrotated name list, so the hash seed is the only thing varying between children, and each child prints the atom it resolved. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
| f"{self._seconds(self._now[lp_id])}", | ||
| self.lp_table(), | ||
| ) | ||
| self._next[lp_id] = horizon |
There was a problem hiding this comment.
Blocking. This assignment discards an event the clock has already accepted, and the run then walks over it silently.
schedule_event lowers a participant's horizon with min (:356). request_advance replaces it. So any event a peer scheduled on i is erased the next time i asks for time — and i asking for time after a peer scheduled something on it is the ordinary case, not an edge case: the message is still in the transport, so the participant declares the event it knows about, which is its own.
Two participants, both floors zero, both clocks at 0:
schedule_event(traffic -> engine, 0.0) -> engine's horizon = 0.0
request_advance(engine, 10.0) -> engine's horizon = 10.0 the 0.0 is gone
request_advance(traffic, inf) -> engine 0s -> 10s, traffic-source 0s -> 10s
The engine now stands at 10.0 s with an accepted, undelivered event at 0.0 s. Nothing aborts: schedule_event returned long ago and there is no later check. This is the failure the module exists to prevent, in the shape the module docstring describes — it does not crash, and the run produces a plausible answer.
It is not a zero-floor artefact. At a 1 ms floor, with traffic at 0.001 s and the engine at 0.002 s, schedule_event(traffic -> engine, 3.0) then request_advance(engine, 10.0) then request_advance(traffic, inf) grants engine 0.002s -> 10s, seven seconds past the event. With the engine declaring the horizon the clock already held, the same sequence grants engine 0.002s -> 3s.
Fuzzed, 400 randomised runs, 2–6 participants, floors drawn from {0, 1e-6, 1e-3, 0.25, 1.0}, 200 legal actions each, counting any grant whose advance_to steps past an accepted event on that participant:
| driver | runs that stepped over an accepted event |
|---|---|
| declares its own next event (in protocol) | 378 / 400 |
| also carries the horizon the clock already holds | 0 / 400 |
No run in either arm raised. The control arm is what says the rule itself is sound: the defect is here, in the bookkeeping, not in earliest_emission_times.
The obvious one-line fix is not enough, and I measured that too. self._next[lp_id] = min(self._next[lp_id], horizon) alone closes the break and then freezes the run, because nothing ever raises the horizon again — six consecutive rounds with both clocks stuck at 1.0 s. min here plus clearing self._next[lp_id] to math.inf in take_up_grant closes the break and leaves an ordinary run advancing 1.0, 2.0, 3.0 … exactly as it does today:
| break case, engine's clock | ordinary run | |
|---|---|---|
| as written | 10.0 s (event at 0.0 s) | 1, 2, 3, 4, 5, 6 |
min only |
0.0 s | 1, 1, 1, 1, 1, 1 |
min + clear on take-up |
0.0 s | 1, 2, 3, 4, 5, 6 |
Clearing on take-up is also the honest reading of the field: a participant that has collected a grant has seen everything up to its clock, and will re-declare on its next call.
Whatever shape you pick, the property to test is the one the fuzzer checks: no grant may move a participant past a timestamp schedule_event has already accepted for it. That is a different statement from either of the two checks in schedule_event, and nothing in the file tests it today.
| same sequence however the requests happened to interleave in real time. | ||
| """ | ||
| self._participant(lp_id) | ||
| if self._status[lp_id] is LpStatus.GRANTED: |
There was a problem hiding this comment.
Related to the horizon finding on :288, and part of the same fix: a second request_advance is refused while the participant is GRANTED but accepted while it is BLOCKED_ON_MESSAGE, where it runs straight into the assignment two lines below and silently raises a horizon a peer lowered.
schedule_event(traffic -> engine, 0.0) horizon 0.0
request_advance(engine, 5.0) blocked-on-message, horizon 5.0
request_advance(engine, inf) blocked-on-message, horizon inf
The state machine in the PR body has no BLOCKED_ON_MESSAGE --request_advance--> ... edge, so this call is out of protocol; the authority is the component that should say so. A raise here costs one branch and removes a whole class of driver bug from CA-3 and CA-4 before either is written.
| else: | ||
| earliest[lp_id] = max(self._now[lp_id], self._next[lp_id]) | ||
| waiting.append(lp_id) | ||
| for _ in waiting: |
There was a problem hiding this comment.
The cost the dev record left unmeasured, measured — and it is not the relaxation.
Rule arithmetic only (one earliest_emission_times() plus one _bound() per parked participant), a complete matrix, mean over 200–300 repetitions, in the container's Python 3.12:
| registered | parked | as written | peer row hoisted |
|---|---|---|---|
| 2 | 2 | 0.013 ms | 0.005 ms |
| 3 | 3 | 0.029 ms | 0.008 ms |
| 10 | 10 | 0.226 ms | 0.051 ms |
| 33 | 33 | 2.260 ms | 0.423 ms |
| 65 | 8 | 1.065 ms | 0.227 ms |
| 65 | 33 | 4.628 ms | 0.823 ms |
| 65 | 65 | 8.493 ms | 1.588 ms |
For comparison, the design sizes one grant's IPC round trip at 0.050 ms and calls 240k grants ~12 s of overhead. At 10 participants this loop is ~5x that round trip; at 65 it is ~170x. So at the largest topology the rule, not the transport, is the cost.
The decomposition matters more than the total. Counting passes: the loop runs 2 passes on every realistic shape I tried at n=65 — equal floors, a cheap chain, a 1 µs ring with 10 ms elsewhere. The for _ in waiting bound is reached only by an adversarial topology where the cheap chain runs against the registry order, and there it uses exactly n passes and still agrees with an unbounded relaxation (checked at n = 5, 10, 20, 40, 65, and on 200 random topologies, zero disagreements, max 5 passes). The relaxation depth is not the cost. The cost is the O(parked x peers) walk underneath it, and specifically two things inside it: self.peers(target) rebuilds an (n-1)-tuple and re-runs _participant on every target of every pass, and lookahead(source, target) hashes a fresh (LpId, LpId) tuple of two unslotted frozen dataclasses on every element.
The "hoisted" column above is the same arithmetic with the inbound row and its floors read once per matrix — 4 to 5x, no change to the rule, no change to any result. LookaheadMatrix.inbound() already returns exactly that row, so the material is there.
Not blocking: at the 2–3 participants of the near milestones this is 0.03 ms and nothing. But the dev record names 65 as the number worth measuring before trusting, so here it is, and the answer is that it is affordable once the row stops being rebuilt.
| settled = True | ||
| for target in waiting: | ||
| best = earliest[target] | ||
| for source in self.peers(target): |
There was a problem hiding this comment.
For CA-3 and CA-4, because this line is where it bites and nothing in the module says it.
Whether a run skips idle or crawls is decided by the driver, by four to six orders of magnitude. A peer that is RUNNING or GRANTED contributes its current clock, so it pins every other participant at that clock plus one floor. The fixpoint above only gets to jump to a distant event when every peer is parked at the moment of the resolve. Two participants, one event at 20 s, nothing else in the run:
| floor | driver | grants | clock reached |
|---|---|---|---|
| 9 ms | take up the grant, then let the peer ask | 2,224 | 20 s |
| 9 ms | ask until the clock refuses, then park | 3 | 20 s |
| 1 ms | take up, then alternate | 20,000 | 20 s |
| 1 ms | ask until refused | 3 | 20 s |
| 1 µs | take up, then alternate | 800,000 (capped) | 0.8 s |
| 1 µs | ask until refused | 3 | 20 s |
The first discipline is the natural one to write — ask, get time, take it, run, find nothing, ask again — and at the pipeline-stage floor it turns a 20-second idle stretch into 40 million grants. The second reaches the same schedule in three. The difference is only that in the second, a participant that has been refused stays refused until a peer's state changes, so the all-parked state is reachable.
Nothing is wrong with the rule here: every row above is safe, and both disciplines that finished reached the same event at the same virtual time (20.0 s). But the sizing this design rests on assumes the second discipline, and the module gives a driver author no signal about which they are in. Worth a sentence in the module docstring and a hard requirement on the harness.
| self._participant(lp_id) | ||
| return self._held[lp_id] | ||
|
|
||
| def grants_issued(self, lp_id: LpId | None = None) -> int: |
There was a problem hiding this comment.
grants_issued() depends on real arrival order. Four participants, unequal floors and unequal horizons, one fixed workload, driven to 1024 events under all 24 permutations of the arrival order:
- distinct
(virtual time, participant)event sets: 1 — the D3.4 property holds, and holds exactly - distinct
grants_issued()totals: 12, from 1457 to 1484
That split is the right one and it is worth stating in the interface: the simulated schedule is invariant under arrival order, the protocol's own cost is not. Since this counter is destined for an always-written run summary, CA-7 should keep it out of anything byte-diffed for reproducibility, or record it beside the diff rather than inside it.
The same effect shows up one level down: with equal floors and equal horizons the grant ledger is identical under permutation, and with unequal ones it is not — 24 distinct ledgers from 24 permutations, differing in advance_to and in bound_from. Not a defect (the events still land on the same times), but see the note on the determinism test.
| def test_the_safety_check_is_not_an_assert_statement(): | ||
| # `python -O` removes an `assert`. A check that is specified as always on, | ||
| # not behind a flag, cannot be one -- so the module contains none at all. | ||
| tree = ast.parse((CLOCK_PACKAGE / "authority.py").read_text()) |
There was a problem hiding this comment.
This test parses authority.py only, while the PR body states the property as "the module has no Assert node". Measured across the package at bbd09f89b: no Assert node in any of the six modules — __init__, authority, identity, lookahead, registry, state — so the property is true, wider than the test.
CA-1 made its two source guards package-wide and parametrised (rglob over clock/) for exactly this reason, and the round-2 reply says why. This one is the check with the sharpest failure mode of the three — python -O deleting a check specified as always on — and it is the one that is not package-wide. A parametrised version over the same glob costs the same line count and covers the case where a safety check moves into a helper module, which is a plausible CA-5 refactor.
| class LpState: | ||
| """One participant's row: its clock, its horizon, and what it is doing. | ||
|
|
||
| `next_event` is the earliest future event the participant knows of, and is |
There was a problem hiding this comment.
This sentence is no longer what the field holds, and the gap is the root of the :288 finding in authority.py.
next_event starts as "the earliest future event the participant knows of", but schedule_event lowers it with min on the strength of a peer's event the participant has not been told about yet. So the field is the earliest event the clock knows of for that participant, which is a strictly larger set than what the participant knows. Once it is that, an assignment from the participant's own declaration is losing information, not refreshing it.
Worth saying here, since this docstring is what a CA-3 or CA-4 author will read before writing the call.
| # --- determinism ------------------------------------------------------------- | ||
|
|
||
|
|
||
| def test_grants_come_out_in_participant_order_not_arrival_order(): |
There was a problem hiding this comment.
Sound for what it tests — mutation-checked: a _resolve that serves the caller first and the rest in order gives 3 distinct ledgers over these 4 permutations and fails both assertions — but the configuration hides a real property, and it is worth knowing which one.
All three participants declare the same horizon (5.0) at the same floor, so none of them can be granted until the last has parked, and every permutation therefore resolves in one pass. Order within a resolve is covered. Order across resolves is not: with unequal floors and unequal horizons, four participants give 24 distinct grant ledgers from the 24 request permutations, differing in advance_to and bound_from.
That is not a bug — driven to completion, all 24 arrival orders produce the same 1024 (virtual time, participant) events, which is the property the design actually asks for. But the invariant the design names is the event sequence, and this file asserts the ledger, which is the stronger claim and the one that is false in general. A second case that permutes arrival order on an asymmetric configuration and asserts the event sequence would pin the property that holds, and would catch a regression that this one cannot see.
Review record — round 1, agent-authoredVerdict: REQUEST CHANGES. One blocking finding, in GitHub refuses APPROVE / REQUEST_CHANGES on a self-authored PR, so the verdict is that sentence. The departure from the design's stated rule is correct, necessary and minimal, and I could not break it. That is the thing this review was pointed at hardest and it is the strongest part of the PR. Details below, with the measurements. The blocking finding is one line away from it and is not about the rule at all. The blocking finding, in one paragraph
The central question: the rule departs from what the design states1. Is the problem real?Yes, and it is larger than the dev record claims. I implemented the stated rule verbatim as a subclass returning At a zero floor, three participants holding events at 5 s, 1 s and 7 s — the configuration the design calls a single global event loop: So the stated rule does not degenerate to a single global event loop at zero lookahead. It stops. And it is not a bootstrap problem: a global event loop advances to the earliest next event, and The design document therefore states a rule that cannot run a configuration the same document requires to work. That is a defect in the document. 2. Is the fix safe?I could not break it, and I tried the fixpoint first, as instructed. The property the design's justification rests on — nothing a participant generates can land in anyone's past — survives the substitution, and three independent attacks left it standing.
The inductive reason it holds: for a running peer the contributed value is its clock, which is monotone; for a parked peer the contributed value is exactly the time it would itself be granted to in that same resolve, so granting it changes nothing a second pass would read differently — and an event a peer schedules on a parked participant already clears both the sender's floor and the recipient's clock, so it cannot drag the contributed value below what a previous bound was computed from. The one thing that can drag it is the assignment at 3. Is it minimal?Yes, and both halves are load-bearing. Measured.
Without the relaxation a parked participant with no horizon contributes The smallest correct change is the one that was made: keep Should the design be amended, and what guarantees it?Yes, here, and a PR-body note does not guarantee it. Two things are wrong in D3 as written, both measured above: the grant rule cannot start or continue a run at a zero floor, and the claim that it degenerates to a single global event loop at zero lookahead is false of the rule as stated (it is true of the rule in this PR). A third is weaker but real: the sizing rests on a driver discipline the document never states — see below. The project has a mechanism for exactly this and a precedent from yesterday: The discrimination test: is the rejected-rule subclass honest?Yes. Checked mechanically rather than read: I reproduced both results:
and the 1 ms repetition ( One note on what the discrimination turns on, which strengthens it: in that topology the engine is parked and the traffic source is running, so the whole difference between the two columns comes from a peer the new rule does not treat specially. The test measures the one case where the substitution is not in play, which is the right case to measure and is what makes "this is not the The three invariantsSafety — holds, and is not an Deadlock — holds, and there is no timeout. Grepped the whole package for I also checked the "impossible" branch of Determinism — holds where the design asks for it, and the counter does not. Mutation-checked that the test is not decorative: a
So the simulated schedule is invariant under arrival order and the protocol's own cost is not. That split is right, and it is a requirement on CA-7: Both degenerate casesOne participant reduces to a local clock with no special case in the code — confirmed, The cost at 65, which the dev record left unmeasuredMeasured, and the decomposition is the interesting part: it is not the relaxation. Two passes on every realistic shape at n=65; the What I did not expect to find, and what CA-3 and CA-4 must do about itWhether a run skips idle or crawls is decided by the driver, by four to six orders of magnitude. A peer that is
The first discipline is the natural one to write. Both are safe; only one is affordable, and the design's 240k-grant sizing assumes the second. Nothing in the module tells a driver author which they are in. Effort and gatesIndependently re-measured with my own
Every figure in the PR body reproduces exactly, including the choice to quote 222 for the two new modules and state the
Both trees re-measured on node 18, container
+42 passed, nothing else moved — reproduces exactly. So does the decomposition, checked by node id rather than by count: One thing I did not reproduce: Smaller things
Accepted with reservation
What CA-3, CA-4 and CA-7 should watch
What I could not check
Round 1. No escalation and no |
The rule that hands out simulated time. A participant may move its clock to min(bound, the earliest event the clock knows of for it), where the bound is the minimum over every other participant of that participant's earliest emission time plus the declared floor into this one. The bound is taken over the peer's current clock, not over the peer's next event. A peer executing at 0 while knowing of nothing until 10 can still emit at 0; bounding on the 10 releases the other side to 10 and the first event it receives is ten seconds into its past. Nothing crashes -- the run finishes and reports a plausible answer -- so both rules are measured against each other in the tests rather than argued about in a comment. A participant that has asked to advance and been refused is the one case where the current clock is not the answer, since it emits nothing until granted. Its earliest emission is relaxed to a fixpoint over the declared floors. Without that, every floor at zero cannot make a move at all -- not once, but at every step, because the minimum over clocks is not the minimum over next events. The minimum ranges over registered participants, never over declared links: an unsized pair is a term missing from a minimum, which lifts the bound instead of tightening it. Refused when the clock is built. The horizon is two records and not one. A participant declares only the events it has seen; a peer may already have placed one on it that is still in the hands of whatever carries messages. Folding those into one slot loses whichever was written second, and losing the accepted one releases a participant straight past a timestamp it has not reached. The accepted side is a list, because reaching the first of two events in flight must not forget the second. Taking up a grant consumes what the grant reached, and only that. Three checks, none optional and none an assert statement, since -O deletes those: an event behind the sender's floor aborts; an event behind the recipient's clock aborts; nothing able to move with nothing known anywhere aborts. There is no timeout. Ties break by participant name, so grants come out in the same order however the requests interleaved. One participant reduces to a local clock. Every floor at zero reduces to one event loop spread over several processes -- correct, serialized, and not an error. 400 generated runs assert that no grant carries a participant past an event the clock has already accepted. The row each walk takes a minimum over is built once, since membership is fixed and a floor is a declared constant. The execution and time model is amended where implementing it showed two of its claims to be false, and the open-items register gains the missing retire call and the arrival-order dependence of the grant count. No transport, no endpoint resolution, no process management. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
bbd09f8 to
4c538fd
Compare
| ) | ||
| self._held[lp_id] = None | ||
| self._status[lp_id] = LpStatus.RUNNING | ||
| self._accepted[lp_id] = [ |
There was a problem hiding this comment.
Blocking. A grant still erases an accepted event, one window later than round 1's — this filter drops events the grant was never computed against.
_resolve computes advance_to and moves _now[lp_id] immediately. The drop happens here, at take-up. Everything a peer schedules onto the participant in between — while it is GRANTED — is appended to _accepted after the grant was decided, and is then thrown away by this filter because it is <= grant.advance_to. The participant was never held at it, never told about it, and the clock keeps no record.
It is not a zero-floor tie. Asymmetric non-zero floors, L[p->q] = 4.0, L[q->p] = 1.0:
q asks, is granted 0s -> 4s, takes it up q RUNNING at 4.0
p asks declaring its own event at 10.0
-> p 0s -> 5s (bound 5s by q) truncated by the BOUND, not by p's horizon
q schedules an event on p at exactly 5.0 legal: 4.0 + 1.0 floor == 5.0, and p's clock is 5.0
-> p._accepted == [5.0]
p takes up its grant -> p._accepted == [], next_event back to 10.0
q asks; p asks 10.0 -> p 5s -> 10s
p finishes at 10.0 s with an accepted event at 5.0 s, five seconds in its past, and nothing raised. That is the silent shape this module exists to prevent, and the grant that got truncated at 5.0 s was a bound truncation — p had no reason to expect anything at 5.0 s and the clock discarded the only record of it.
Fuzzed, with an oracle that (a) observes the grants schedule_event returns and (b) releases only what each grant was computed against:
| step-overs in 400 legal runs | |
|---|---|
as committed, 4c538fdb1 |
18 |
| grant remembers the accepted entries it was computed against; take-up drops exactly those | 0 |
Same generator, same floors {0, 1e-3, 1.0}, same 2–5 participants, same 24 actions as _legal_runs.
The direction that measured clean is the one the docstring above already describes but the code does not implement — "taking up a grant consumes the events the grant reached" has become "consumes every event at or before where the clock now stands", and those are different sets once a peer can write into _accepted after _resolve has run. Snapshotting [t for t in self._accepted[lp_id] if t <= advance_to] at grant issue and dropping that list (not a threshold) at take-up is a four-line change and keeps both properties the current code was built for: test_a_run_that_reaches_its_horizon_keeps_moving still passes, because the entry that caused the grant is in the snapshot.
Principle 6: the failure produces a plausible answer instead of a refusal.
| "so this run no longer describes the system being modelled", | ||
| self.lp_table(), | ||
| ) | ||
| self._accepted[target].append(when) |
There was a problem hiding this comment.
A round-1 finding that survived, and that this PR's change turned from a no-op into unbounded growth.
Round 1: "schedule_event(source, target, math.inf) is accepted and is a silent no-op — _simulated_seconds refuses -inf and NaN but not +inf ... One line, same guard." The guard was not added.
With a scalar horizon that was harmless: min(next, inf) changed nothing. With _accepted as a list it is not. take_up_grant keeps every when > grant.advance_to, and inf is greater than every finite advance_to, so an inf entry can never be dropped by any grant. Measured:
20,001 x schedule_event(y, x, math.inf)
-> len(_accepted[x]) == 20,002
-> one further schedule_event costs 1.716 ms
because _restate_horizon re-mins the whole list on every schedule_event, request_advance and take_up_grant. That is quadratic in a module whose PR body measures a 3.3x hoist at 65 participants to keep a single walk under 2 ms.
self._simulated_seconds(timestamp, "timestamp") at :380 is where it belongs — refusing +inf there costs the same line as refusing -inf does now, and it is the honest answer: "an event that never happens" is not an event, and a caller that means "I know of nothing" says that by not calling.
Principle 6.
| bookkeeping instead of through the rule. | ||
| """ | ||
| self._participant(lp_id) | ||
| if self._status[lp_id] is LpStatus.GRANTED: |
There was a problem hiding this comment.
Round-1 finding, partly answered. The harm it named is gone — _declared is a separate record now, so a second request_advance can no longer raise a horizon a peer lowered; I checked, and an event accepted at 1.0 s survives request_advance(engine, 5.0) followed by request_advance(engine, inf).
The call is still accepted, and it is still out of protocol: the state machine in the PR body has no BLOCKED_ON_MESSAGE --request_advance--> ... edge. Only GRANTED is refused here. Non-blocking, unchanged in substance from round 1: one branch, and CA-3 and CA-4 get told rather than diverging quietly.
| class LpState: | ||
| """One participant's row: its clock, its horizon, and what it is doing. | ||
|
|
||
| `next_event` is the earliest future event the participant knows of, and is |
There was a problem hiding this comment.
Unchanged from round 1, and now further from the truth than it was then.
Round 1 asked for this sentence to be corrected because next_event is the earliest event the clock knows of, not what the participant knows. This PR makes that split explicit in authority.py — _declared is what the participant said, _accepted is what peers placed on it, next_event is their minimum — and the field this docstring describes is the minimum, i.e. exactly the larger set round 1 named.
This docstring is what a CA-3 or CA-4 author reads before writing the call, and the whole of the blocking finding in authority.py is about which of those two sets a given line is operating on. One sentence.
Principle 7 is not the right number here; it is just that the two-record design is now load-bearing and the public type does not mention it.
| when = clock.now(actor) + floor_out + ((state >> 4) % 5) * 0.25 | ||
| clock.schedule_event(actor, target, when) | ||
| pending[target].append(when) | ||
| pending[target] = [ |
There was a problem hiding this comment.
This line is why the fuzzer reports 0. It deletes exactly the class of event the remaining defect is made of.
Two things compound:
clock.schedule_event(...)at:447returns grants and the return value is dropped, so the shadow ledger never drains on a grant an event released.- This line compensates for that by clearing every pending entry at or before the target's clock — which is also every event accepted onto a participant that is currently
GRANTED, because_resolvehas already moved its clock toadvance_to.
So the ledger forgets precisely the events that the implementation also forgets, and the assertion compares two records of the same omission. Measured, on the same generator, floors and action stream:
| oracle | step-overs in 400 runs |
|---|---|
| as written | 0 |
this line removed, schedule_event's grants still dropped |
32 — but 14 of them are false, an artefact of (1) |
schedule_event's grants observed and each grant releasing only what it was computed against |
18, all true |
The mutation table in the PR body is real — I did not re-run it, and the three mutants it names would still be caught. A fuzzer can be non-vacuous against the mutants its author imagined and blind to the case none of them models; that is what happened here.
The invariant to assert is the one the docstring already states — "must not be carried past an event that has already been accepted on it and that it has not yet been released to reach" — with "released to reach" resolved at the moment the grant was computed, not at the moment it is collected.
| def test_the_safety_check_is_not_an_assert_statement(): | ||
| # `python -O` removes an `assert`. A check that is specified as always on, | ||
| # not behind a flag, cannot be one -- so the module contains none at all. | ||
| tree = ast.parse((CLOCK_PACKAGE / "authority.py").read_text()) |
There was a problem hiding this comment.
Round-1 finding, not taken and not answered in the PR body. Still authority.py only, while the property the PR body claims is "the module has no Assert node" and the property that actually matters is the package one.
state.py is now a second production module in this package and it is not covered by this test. CA-1 made its two source guards package-wide and parametrised over clock/*.py for exactly this reason — and the +4 in this PR's own gate delta is those two guards picking up authority.py and state.py automatically. This is the third guard, it has the sharpest failure mode of the three (python -O silently deleting a check specified as always on), and it is the one that has to be edited by hand every time a module is added.
Same glob, same line count.
| # --- determinism ------------------------------------------------------------- | ||
|
|
||
|
|
||
| def test_grants_come_out_in_participant_order_not_arrival_order(): |
There was a problem hiding this comment.
Round-1 finding, not taken. All three participants still declare 5.0 at the same floor, so every permutation resolves in one pass and only order within a resolve is covered.
Restating the measurement rather than the argument: four participants, unequal floors and unequal horizons, driven to completion under all 24 arrival permutations give 24 distinct grant ledgers and 1 distinct (virtual time, participant) event set of 1024 events. This file asserts the ledger, which is the stronger claim and the one that is false in general; the design asks for the event sequence. Non-blocking, and the same sentence as round 1 — a second case on an asymmetric configuration asserting the event sequence would pin the property that holds.
It is also the property T84 is registered about, so the test and the register row currently disagree about which invariant is the real one.
| | **T49** | The prefix-index *lookup* cost is charged to nobody — ~1,387 blocks hashed and probed per request at the cc-traces p50, magnitude unmeasured | `03` | | ||
| | **T50** | Whether runtime memory constants transfer across dies (the working assumption says yes within a software generation) | `03`, `05` | | ||
| | **T53** | Whether tokenizer throughput transfers across CPU classes (the working assumption says yes, adjusted by derate) | `05` | | ||
| | T83 | **The virtual-time protocol has no way for a logical process to say it has finished, so a clean end of run is indistinguishable from a deadlock.** Opened 2026-09-21 by CA-2 (#45, PR #59). `01` D3's deadlock invariant is "all LPs blocked and none holding a finite `next` -> abort loudly", and the Clock Authority implements it exactly. But that state is also what a *correct, completed* run looks like — every LP has drained its work and parked with an empty horizon — so an integration reaches it at the end of every clean run and aborts. Asserted as the specified behaviour in `tests/compass/test_clock_grant_rule.py::test_everyone_waiting_with_no_known_event_aborts_loudly` and `::test_one_participant_with_nothing_left_to_do_is_the_same_stall`. The gap is not theoretical: a harness written while measuring the row below hung rather than finishing, because an LP that had drained its own events had no third option — it could neither park (which would abort the run) nor keep asking (which pins every peer at its stale clock), and the run spun with no grant issued and no abort raised. The item is the missing third case: either a retire call on the protocol, so a departed LP leaves the minimum and the rest carry on, or a stated convention that the run is torn down before the last LP parks. It is **not** a choice the clock can make alone, because "the workload is over" is a statement about the traffic source and the engines, not about time. Owner: whoever takes the transports (#46) and the traffic source (#47); CA-2 deliberately did not invent one, since the wrong convention here is more expensive than the abort. The abort must not be softened into a timeout while this is open — that invariant exists because a prior 120 s arrival barrier released on an invalid run and the client reported "0 failed". | `01` | |
There was a problem hiding this comment.
Blocking, and visible only on the merged tree: T83 and T84 are already allocated.
P0.6 (#51) landed on feature/atomcompass_new before this PR's base did, and it registered T83–T87. The integration head f67618eb9 already carries:
T83 The EP group's size and moe_parallel_config.ep_size are two different numbers ...
T84 `15` D94's M1 EP2 leg is degenerate as written ...
T85 Multi-node EP rank-to-node mapping is assumed, not verified ...
T86 aiter is a second executed source root, and no artifact key names it ...
T87 `07`'s price-list table states the MoE all-to-all's block-count cap as one number ...
This branch's base is 5b11e82ca, which does not contain c19710bcd (#51), so nothing on the branch can see the collision. Cherry-picking 4c538fdb1 onto f67618eb9 conflicts in both 12_open_items.md and README.md, in exactly the register-count paragraphs:
<<<<<<< HEAD
3. TODO register - 86 rows, T1-T80 and T82-T87, of which 80 are open
=======
3. TODO register - 83 rows, T1-T80 and T82-T84, of which 78 are open
>>>>>>> 4c538fdb1
Resolved as a union the register holds two T83 rows and two T84 rows with unrelated content, and 83/78, 86/80 and the counting rule the paragraph itself states (grep -oE '^\| *~*\**T[0-9]+' over section 3) all disagree. The true merged figures are 88 rows and 82 open. The rows themselves are good — T83 in particular is the strongest thing in this PR's documentation, and the harness that hung on it is the right kind of evidence — they just need T88 and T89, and the two count lines need to be written against the integration head rather than against the base.
It also bites T84's attribution, which is otherwise accurate: I checked the numbers against round 1's review and "12 distinct totals spanning 1457 to 1484" and "one distinct set of 1024 events" are exactly what was measured and are correctly credited — but on the merged tree "T84" names an EP topology item, so the citation lands on the wrong row.
Process, not a finding: this PR's base branch is still compass/ca-1-clock-identity, which merged at 14:46 today. It needs retargeting to feature/atomcompass_new before any of this can land.
Principle 8: the counts are the measurement, and they are stated against a tree that is no longer the one this lands on.
| truncates more grants, so the crawl is finer. The rule a transport has to implement is | ||
| therefore: *an LP asks again when it has work or has been refused, and an LP with nothing | ||
| to do parks rather than stepping its clock forward one floor at a time.* That decision, not | ||
| PP degree, is what decides whether the grant traffic is affordable. |
There was a problem hiding this comment.
No blank line between the end of the sizing amendment and **PP is therefore an efficiency concern ...** on the next line, so Markdown renders the amendment's closing sentence and D3's original PP conclusion as one paragraph. (There is also a doubled blank line above the amendment heading.) One newline.
The amendment itself reproduces to the digit — two LPs, one event at 20 s, identical matrices, my own driver:
| discipline | 1 ms floor | 1 us floor |
|---|---|---|
| idle LP asks until refused, then parks | 3 grants, reached 20 s | 3 grants, reached 20 s |
| take up and immediately re-ask | 20,002 grants, reached 20 s | 2,000,000 grants (cap), reached 2 s |
The inversion holds and it is the right thing to have written down: the efficient discipline is lookahead-independent, and a tighter floor makes the inefficient one strictly worse. Worth noting for whoever implements the transport that the efficient discipline is not "park instead of asking" — it is ask, be refused, and stay refused, which needs the idle LP to ask twice before the busy one asks at all. That ordering is the whole of the three-grant result and nothing in D3 or in the module says it.
Review record — round 2, agent-authoredRange: Verdict: CHANGES REQUESTED. Two blocking findings, both new, neither of them the one round 1 raised:
Round 1's blocking finding is fixed, and the fix is right. That is the first thing to say and it is not grudging — the two-record horizon is a better answer than the one round 1 suggested, the author's reasoning for rejecting the suggestion is correct, and I confirmed the rejection by measurement rather than by reading it. The remaining hole is one window further out again, in the same three lines. That recurrence, not any single defect, is what the loop stop is for. Round 1's blocking finding: fixed, verified by re-running its own reproductions
The author's claim that the suggested fix was incomplete is correct, and the third case they found — a scalar horizon holding only the earliest of two events in flight — is real. The two-record split is the right shape. The blocking finding: what survives it
Why the fuzzer reports 0 while this is present — What I could not breakEverything else I threw at the two-record horizon held, and I went looking specifically where the task pointed.
The D3 amendment, judged as a document changeThe grant-rule amendment is right, and it is in the right place. Round 1 said a PR-body paragraph does not guarantee a design correction and named The sizing amendment inverts D3's intuition, and the inversion holds. Re-measured with my own driver, two LPs, one event at 20 s, identical matrices — every figure reproduces to the digit:
Lookahead-independent in the first row, and a tighter floor makes the second strictly worse. One thing the amendment does not say and a transport author needs: the efficient discipline is not "park instead of asking", it is ask, be refused, stay refused — the idle LP has to ask twice, and get a one-floor crawl grant the first time, before the all-parked state that produces the jump is reachable at all. That ordering is the whole of the three-grant result. T83 is the strongest piece of documentation in the PR. A registered row for a gap the author hit in their own harness, with the harness hang recorded as the evidence, and a deliberate refusal to invent the convention. That is the right call and the right record. T84's attribution is accurate — I checked "12 distinct totals spanning 1457 to 1484" and "one distinct set of 1024 events" against round 1's review and both are exactly what was measured and are correctly credited to the reviewer rather than claimed. Both rows are undermined only by the number they were given; see the blocking finding. Gates — three trees, including the merged oneNode 18,
+47, nothing else moved — reproduces exactly, and so does the decomposition by node id: The merged tree is green. CA-1's two package-wide globs over The author's diagnosis of round 1's
Effort — reproduces exactlyMy own
235 against 250–350 is a ~6% underrun; the halt rule is about overruns and CA-1 took the same posture. Cost at scale — grants really are identicalConfirmed the thing the hoist has to be judged on.
Rule arithmetic with every participant parked, my container: 0.032 -> 0.009 ms at 3, 0.252 -> 0.070 at 10, 9.575 -> 2.511 at 65 (3.8x). The PR's 5.631 -> 1.682 and round 1's node-18 8.49 -> 1.59 are the same shape at a different relaxation depth; the ratio band 3.3–5.3 is consistent and the absolute numbers are not comparable across containers, as the PR body says. Still ~50x D3's 0.050 ms per-grant RPC estimate at 65, which remains T70/PP territory. Round-1 findings that survived
None of these is individually blocking and round 1 said so. They are listed because the loop stop counts findings, not severities. The loop stop
Applied. Three things, in the order they weigh:
I am not escalating the blocking bookkeeping finding as such — it is an actionable review finding and the fix direction measured clean at 0/400. What CA-3, CA-4 and CA-7 should watch
What I could not check
Round 2. |
…rule
Base update authorized by the owner on 2026-09-24 ("go ahead and update
the clock chain"). The PR keeps its need human label. This commit only
resolves the merge.
Resolved files:
- atom/compass/clock/identity.py, lookahead.py, registry.py,
tests/compass/test_clock_lp_identity.py,
tests/compass/test_clock_order_across_processes.py: took the
integration side whole. This branch carried each one byte-identical to
the pre-#410 tip (the #52 landing), so every differing line came from
#410 or #417.
- atom/compass/clock/__init__.py: this branch's docstring, imports and
__all__ (ClockAuthority, the state types), plus the paragraph #53 added
on the tip. That paragraph holds for authority.py and state.py: both
import only math, enum and dataclasses, and build no set.
- atom/compass/design/12_open_items.md: kept both sides' rows. The tip
had meanwhile landed P0.6's T83 and T84, so this branch's two rows are
renumbered T89 (no retire call on the protocol) and T90 (grant count
depends on arrival order), and T90's pointer to "T83 above" now reads
T89. The register header now reads 90 rows, T1-T90 with no gaps, 84
open, counted with the header's own grep rule; the allocation sentence
names T88 with #196 and T89-T90 with CA-2.
- atom/compass/design/README.md: the TODO count line reads 90 registered,
T1-T90 with no gaps, 84 open, with the tip's struck list.
No member removed by #410 appears anywhere in the merged tree.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Base update authorized by the owner on 2026-09-24 ("go ahead and update
the clock chain"). This commit only resolves the merge, which brings in
#59's merge of fork/feature/atomcompass_new (3a3267a).
Resolved files:
- atom/compass/design/12_open_items.md: conflicted on the two TODO rows
this branch edits. #59's merge renumbered them T83 -> T89 and T84 -> T90,
because the tip had landed P0.6's T83 and T84. Kept this branch's edited
text of both rows under the new numbers, after the tip's T88. T90's
pointer to "T83 above" now reads T89. This branch's T70 edit merged
cleanly. The register still counts 90 rows, T1-T90, 5 struck.
- atom/compass/design/01_execution_and_time_model.md (merged without a
conflict, edited here so it stays true): this branch's two citations of
those rows now use the new numbers, "(T90 records CA-2's reviewer ...)"
and "T89: a run whose work is finished ...". The other T83/T84 mentions
in the tree are P0.6's and are unchanged.
No member removed by #410 appears anywhere in the merged tree.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…arness
Base update authorized by the owner on 2026-09-24 ("go ahead and update
the clock chain"). This commit only records the merge, which brings in
#59's merge of fork/feature/atomcompass_new (3a3267a).
Resolved files: none. The merge had no conflicts, and no file was edited
beyond what git merged. This branch's own files (tests/compass/clock/)
read TRAFFIC_TO_ENGINE_FLOOR_SECONDS and the LinkClass members, all of
which #410 kept.
No member removed by #410 appears anywhere in the merged tree.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Base update (owner-authorized, 2026-09-24)The owner authorized this update on 2026-09-24 ("go ahead and update the clock chain"). It only merges the integration branch in. Labels, draft state and review state are unchanged, and nothing was rebased or force-pushed. What was merged: freshly fetched
Conflicted files and how each was resolved (8)
The resolutions are all in Removed-member grepI grepped
Gate (node 18,
|
| Tree | commit: |
Result | GATE_CPU_RC |
|---|---|---|---|
new head a386039dd |
a386039dd (stamp) |
5325 passed, 155 skipped, 3 xfailed (193 s) | 0 |
gpu: not required.- No failures, so no timing flake fired.
Collected node ids (gate_cpu.sh --collect-only)
- Old head → new head: 4262 → 5462.
- −7: exactly the 7 tests compass(clock): apply the ponytail-audit of atom/compass/clock/ (#406 phase 2) #410 removed (six in
test_clock_lp_identity.py, plustest_the_children_really_did_register_in_different_orders). - +1207: every one of them is also collected at the tip
3a3267a84.
- −7: exactly the 7 tests compass(clock): apply the ponytail-audit of atom/compass/clock/ (#406 phase 2) #410 removed (six in
- New head against the tip (5415): +47, −0.
- The +47 are this PR's
test_clock_grant_rule.py(43), plus 4 cases of the tip's two rglob-parametrised package tests overauthority.pyandstate.py.
- The +47 are this PR's
No blocking issues from this update.
Base-update resolution reviewVerdict: the resolutions in merge This review was written by an agent. It covers only the conflict resolutions of the owner-authorized base update ( I read 1. Resolutions (parents
|
| File | Check | Result |
|---|---|---|
clock/identity.py, lookahead.py, registry.py, tests/compass/test_clock_lp_identity.py, test_clock_order_across_processes.py |
blob at 4c538fdb1 vs f67618eb9 (#52) and 9a6317927~1 (pre-#410); blob at merge vs tip |
PR side equals the pre-#410 tip in all 5 files, so the PR's diff changed none of these lines. The merge blob equals the tip blob in all 5. Toward the integration side, correctly. |
clock/__init__.py |
3-way against #52's version 46fe09db |
The merge is exactly the union: the PR's docstring, authority/state imports and __all__, plus #53's enforcement paragraph. No tip line dropped. |
design/12_open_items.md |
word diff of tip → merge | Only additions: rows T89/T90, the counts 88→90 and 82→84, "; T89 and T90 make it 90", and the allocation sentence. All of the tip's text is kept. The one tip phrase reworded, "and T88 arrives here", becomes "T88 with #196". That keeps the fact, because the phrase is no longer true on this branch. |
| same, rows | PR rows vs merged rows | After renaming T83→T89 and T84→T90 and "T83 above"→"T89 above", the merged rows equal the PR's rows byte for byte. |
design/README.md |
tip → merge | The count line changes and nothing else. The tip's struck list (including T65) is kept. |
01_execution_and_time_model.md |
patch-id | Identical to the PR's own diff. It merged cleanly, with no hand edit. |
Did any tip change get dropped? No, in any hunk.
2. Removed-member grep
- What I grepped:
git grepover the merge tree (all paths,*.mdincluded) for.tightest,serializing,.links(),_missing_into,scale_seconds,._ordered,len(matrix…),LinkClass(,LinkClass.X.label/.value/.scale_seconds,<link>.label,repr/strof a registry, and truth tests on a matrix. - Result: 0 hits beyond the tip's own 13. The 13 are ATOM comments containing "serializing",
len(m)in unrelated code, andclass LinkClass(. The hit list is identical to the tip's line for line. - The grep fires: 35 hits at the old head
4c538fdb1.
3. clock/__init__.py docstring against this PR's modules
- Imports: I parsed every
.pyunderatom/compass/clock/at the merge withast.authority.pyimportsmathplus.identity,.lookahead,.registryand.state.state.pyimportsdataclasses,enumandmathplus.identity. Neither has a stray module against{dataclasses, enum, math}. - Sets: no
set/frozensetcall,SetorSetCompanywhere. - The rest of the docstring: "the rule that reads all three" is true, because
authority.pyimportsLpId,LpRegistryandLookaheadMatrix. "The transport … sits elsewhere" is true: nothing in the package is a transport. - Tests: the tip's rglob-parametrised tests collect
[authority.py]and[state.py]cases, and they passed in the gate below.
4. T83/T84 → T89/T90 renumbering
- Section 3 by the header's rule: 90 rows, T1–T90, no gaps, no duplicates. 5 are struck and T77 is closed, so 84 are open. That matches the header and
README.md. - Other references to the four numbers: at this merge, compass: what a run of the clock says about itself (CA-7, #50) #67
8222b3a93, compass: ATOM's deployments driven against the clock, synthetically (CA-4, #47) #7519b674275, and the local compass: the causality detectors, and three injected violations that fire them (CA-5, #48) #91/compass: a determinism harness two processes wide, and the set-iteration lint #96 merges00d96edeb/8e2c99a27, every T83/T84 is P0.6's:12L219–221 and15_parallelism_support.md. Every T89/T90 is CA-2's. Nothing in code or tests cites any of the four. - Open PRs: I scanned every open PR's patch to
12,READMEand01for T83–T99. Only compass(design): guard the open-items register's own counts #85 adds any, and its "T83–T87 with P0.6" is P0.6's. No other PR claims T89 or T90. - compass: both clock transports from one implementation (CA-3, #46) #63 (
ea5f78f36, not updated): it still carries the rows as T83/T84. Its own base update will have to renumber them the same way.
5. Gate
I read n59.gate.log in full and did not re-gate.
- Run details:
commit: a386039dd (stamp),atom.__file__under/tmp/xiaobizh_chainupd/n59/ATOM,gpu: not required,timeout -k 10 3600, not piped, and apgrepwait with a re-check before launch. - Result: 5325 passed, 155 skipped, 3 xfailed,
GATE_CPU_RC=0. - Node ids from old head to merge: exactly compass(clock): apply the ponytail-audit of atom/compass/clock/ (#406 phase 2) #410's 7 removed. All 1207 added ids are collected at the tip. Against the tip it is +47 and −0.
Blocking: none.
Non-blocking: 1, inline on 12_open_items.md line 26.
| T88 arrives here. It happens to be complete at this head; that is an observation about | ||
| what has landed, not a property to rely on | ||
| branches was in flight: T73–T76 arrived with P0.3, T83–T87 with P0.6, T81 with P0.4, | ||
| T88 with #196, and T89–T90 with CA-2, which first numbered them T83–T84 before P0.6's |
There was a problem hiding this comment.
Non-blocking (base-update resolution): the order of events here is wrong.
"CA-2, which first numbered them T83–T84 before P0.6's landed under those numbers" does not match the history:
- P0.6 (compass: state EP group membership per configuration (P0.6, T65) #51,
c19710bcd) landed on 2026-09-21 at 14:14:22 UTC. - The CA-2 commit that first carries these rows as T83/T84,
4c538fdb1, was authored at 14:27:03. Neither earlier commit on the branch (b94c50a55,5b11e82ca) has them.
So P0.6's T83/T84 landed first. CA-2 reused the numbers because its branch was cut from 7fc7a5ddd, before P0.6 landed.
Suggested wording: "…and T89–T90 with CA-2, which had numbered them T83–T84 on a branch cut before P0.6's rows landed under those numbers."
This wording carries into #67 and #75 through this merge. A base update may change nothing beyond its conflicts, so under need human the fix waits until the owner lifts the label.
…ers the REST base patch (#434) Review cycle 1 on #437. - A merge that cannot keep both sides' changes in a conflict hunk now commits nothing and names the hunk in a PR comment. The old text ("toward the integration side unless the PR's own diff changes those lines") gave no side for the usual case, so a literal reader could drop a change the tip made. - The exception now covers the REST base patch that an unlinked child whose parent landed also gets (live on #61; #63 once #59 lands). - The merge sources are no longer restated; the branch-update rule names them. - Deleted "It lands nothing, reviews nothing and leaves the label." and the back-reference in the branch-update rule. Both restated other text, and the second read as a duty. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
Design-alignment review against the PDES time model (#443) — agent-authored, requested by the owner. This is a design-alignment review, not a gate-4 round. It approves nothing and changes no label. The design text is carried by the doc PRs under #443 (doc 01 §D1-§D3.5 is #446; §D4, §D5, §D9 and the log are #450; §D6-§D8 are #448); section numbers below are the v0.20 design's. Head Measured at the head ( Corrections
Missing in scope
Filed separately
Matches
|
|
Closed on the owner's ruling of 2026-09-29. Asked "for the stale PRs on the clock, shall we close them and start from scratch according to the updated PDES design, or revise them?", the owner answered: "close and re-cut". This PR is not revised; its scope is re-briefed against the v0.20 design (#443). Agent-authored. Replacement: #480, the CA grant rule and state machine (strict LBTS over Carries over (pinned to The branch stays on the fork, so the head the salvage lists cite stays reachable. The |
Closes #45 when landed. Builds on #52 (
compass/ca-1-clock-identity), which is the base of this PR.Pure logic: the grant rule and the state machine that carries it. No transport, no endpoint resolution, no process management.
Review round 1 returned REQUEST CHANGES on one blocking finding. Fixed, and the fix found a second case the report did not reach. Details under The bookkeeping defect below. The branch was amended and force-pushed (
bbd09f89b→4c538fdb1), so CA-3 needs to rebase.The interface
atom/compass/clock/state.pyLpStatusRUNNING,GRANTED,BLOCKED_ON_MESSAGE;.may_produce_eventsis false only for the lastLpState(lp_id, now, next_event, status)Grant(lp_id, advance_from, advance_to, bound, bound_from);.secondsis the span.advance_to == advance_fromis a real grant — it says a parked participant's wait is over, not that time movedatom/compass/clock/authority.py—ClockAuthority(registry, lookahead, start_time=0.0)request_advance(lp_id, next_event=inf)tuple[Grant, ...]take_up_grant(lp_id)GrantGRANTED→RUNNING; consumes the events the grant reachedschedule_event(source, target, timestamp)tuple[Grant, ...]grant_bound(lp_id)float+infwith no peersearliest_emission_times()dict[LpId, float]peers(lp_id)tuple[LpId, ...]require_sized_peers()LookaheadMatrix.require_complete(); called once from__init__state/states/now/held_grant/grants_issued/lp_tablegrants_issuedregistry/lookaheadExceptions:
ClockAbort(base, carries.reasonand.table),BackdatedEvent,ClockDeadlock.State machine:
RUNNING--request_advance-->GRANTEDorBLOCKED_ON_MESSAGE;BLOCKED_ON_MESSAGE--(another participant's call releases it)-->GRANTED;GRANTED--take_up_grant-->RUNNING.The bookkeeping defect, and the second case behind it
Reported (blocking).
request_advanceassigned the declared horizon over the clock's record, so a participant asking for time after a peer had scheduled an event on it erased the only record of that event and was released straight past it. Reproduced here exactly as reported: zero floors,schedule_event(traffic→engine, 0.0),request_advance(engine, 10.0),request_advance(traffic, inf)grants engine 0 s → 10 s over an accepted event at 0 s, with nothing raised.I did not take the suggested fix, because measuring it showed it incomplete. Three candidates, measured:
minonlymin+ clear the horizon on take-upmin+ drop only what the grant reachedThe suggested fix closes the reported case and re-opens the same hole one grant further out: an event accepted beyond the grant is thrown away, and the participant cannot re-declare it because it has not been told about it.
Then the fuzzer found a third case, which neither the report nor my fix covered. A single scalar horizon holds only the earliest accepted event, so reaching the first of two events in flight forgets the second. 41 step-overs across 400 runs.
So the horizon is now two records.
_declaredis what the participant itself last said — only ever events it has seen._acceptedis a list of every event a peer has placed on it that it has not yet been released to reach.next_eventis the minimum of the two. Taking up a grant drops the accepted entries the grant reached and the declared horizon if the grant reached it, and nothing else. Both halves are load-bearing and both are tested.The fuzzer. 400 generated runs over two to five participants and three floors, driving legal operations and asserting that no grant carries a participant past an accepted event it has not been released to reach. Nothing in it relies on an abort being raised, because the failure raises nothing. Not vacuous — mutated three ways, it catches each:
The named result
Two participants,
engineandtraffic-source, zero floor in both directions, both clocks at 0. The engine knows of an event of its own at 20 s. The traffic source is executing at 0 s and knows of no event of its own. The rejected rule is a test subclass overriding onlyearliest_emission_times, so the real state machine and both safety checks run underneath it.next[j]rulenow[j]rulegrant_bound(engine)+inf0.020.0 s0.0 s0.0 sRepeated at a 1 ms floor: the clock rule releases the engine by exactly one floor, the horizon rule still to
20.0 s.The sibling result: an unsized pair
decode, all clocks at 010.0 s0.00.0 sAn absent floor is not a cautious version of a zero floor. A zero adds a term to the minimum and can only lower it; an absence removes a term, which raises it. Measured end to end, a peer set taken from the sized pairs grants decode to
10.0 swhile the traffic source stands at0.0 s, and the first event on that leg is 10 s into decode's past. Three doors are shut on it now: construction callsrequire_complete(),inbound()refuses a short row, and the walk is overpeers(). Reproducing it needs two deliberate overrides.The three invariants
schedule_eventraisesBackdatedEventonts < now[source] + Land onts < now[target], always, with the full tabletest_an_event_earlier_than_the_senders_own_floor_aborts,test_the_abort_carries_every_participants_clock_and_status, both discrimination classes_refuse_to_stallraisesClockDeadlockwhenever a resolve issues nothing and everything is parked. No timeout exists in the packagetest_everyone_waiting_with_no_known_event_aborts_loudly,test_one_participant_with_nothing_left_to_do_is_the_same_stall,test_a_floor_above_zero_delays_the_stall_by_one_step_and_no_moreregistry.ids(); grants issued and returned in that order; ties keep the first peer. Nosetconstructedtest_grants_come_out_in_participant_order_not_arrival_order(four permutations, identical ledgers),test_a_tie_on_the_bound_keeps_the_first_peer_in_the_total_orderNeither safety check is an
assert—python -Odeletes those.test_the_safety_check_is_not_an_assert_statementwalks the AST and asserts the module has noAssertnode.Both degenerate cases
+inf, grants land exactly on the declared horizons,bound_fromisNone. A local clock, no special case.Design amendments, in this PR
12_open_items.md§5 exists for corrections not yet applied; these are applied, so they go in the documents, with T-rows for what is not mine to settle. Register goes 81 rows → 83, 76 open → 78;README.mdand12_open_items.md§-intro counts updated together.01D3, grant rule — amendment. Two claims were false as written and both were found by implementing them. The literal rule cannot make a move at a zero floor: every clock starts equal, somin over j≠i of now[j]equalsnow[i]. It is not a start-up wrinkle — the condition recurs every step, because that minimum is over clocks, not over next events. The three-LP configuration D3 itself calls a global event loop (zero floors, events at 5/1/7 s) deadlocks under it. So it also does not degenerate to a global event loop; it degenerates to a stall. The amendment records the correction, and the reviewer's independent measurements of it: the fixpoint agrees with an unbounded relaxation on adversarial reverse-order chains at n=5…65 (exactly n passes, every time) and on 200 random topologies with no disagreement; and both halves are load-bearing, since "parked contributes its ownnext" without the fixpoint grantsbto 15.0 s where the full rule gives 5.0 s and an event at 5.0 s then lands 10 s inb's past.01D3, sizing — amendment. The 240k grant estimate assumes a driver discipline nothing states. Measured here, two LPs, one event at 20 s, identical matrices:Two things the numbers say that the intuition does not: the efficient discipline is lookahead-independent, and a tighter floor makes the inefficient one worse, not better.
T83 — no way for a participant to say it has finished. A clean end of run is indistinguishable from a deadlock and aborts. Not theoretical: the harness I wrote to measure T84 hung on exactly this, because a drained participant could neither park (aborting the run) nor keep asking (pinning every peer at its stale clock). Owner is whoever takes #46 and #47; I deliberately did not invent a convention.
T84 —
grants_issued()is arrival-order dependent while the event schedule is not: 12 distinct totals across 24 ask orders, 1457–1484, against one distinct set of 1024 events. A requirement on the run summary (#50). Measured by the reviewer and attributed, not reproduced here — my harness hit T83 and hung, and I judged reproducing someone else's measurement not worth further time against a registered row. Flagged rather than quietly dropped.Carried, not fixed here
Cost at a large participant count — measured and mostly removed. It was not the relaxation. The row each walk takes a minimum over is now built once at construction, since membership is fixed and a floor is a declared constant:
Identical grants in every case. (Measured in the local container, not on node 18, so the absolute numbers are not comparable with the reviewer's 8.49 → 1.59 ms; the ratio is.) Still above D3's 0.050 ms RPC estimate at 65, which is T70/PP territory and not this task's.
Effort
atom/compass/clock/state.pyatom/compass/clock/authority.pytests/compass/test_clock_grant_rule.pyPlus 20 insertions / 6 deletions in
__init__.py(additive exports), and the design amendments above. No CA-1 source file was otherwise touched. Method, matching CA-1's:ast.parse, strip docstrings from every module, class and function,ast.unparse, count non-blank lines.Gates
Node 18, container
xiaobizh_n18_cpu. Both trees staged bygit archive+docker cpinto paths of my own — the shared mount untouched — tarball md5 verified on both ends,.compass-commitand.compass-changedstamps written bysnapshot.shand confirmed present in the container before running, andimport atomresolved under the tree before any count was read.5b11e82ca(this PR's base)4c538fdb1pytest: rc=GATE_CPU_RC=Delta: +47 passed, 0 failed, skips and xfails unchanged.
Arithmetic: 43 tests in the new file, plus 4 from CA-1's two package-wide parametrized checks — the standard-library allowlist and the no-
setcheck are parametrized overclock/*.py, and this PR adds two modules, so each gains two cases. 43 + 4 = 47.On the reviewer's
GATE_CPU_RC=98on both trees: that isgate_cpu.shrefusing a baregit archivetree with no.compass-changedstamp, not a difference in the trees. Building the snapshot withscripts/compass/snapshot.shwrites it; both runs above reportgpu: not required (.compass-changed stamp).ruff checkandblack --checkclean on all touched files, against a repository baseline that is dirty.🤖 Generated with Claude Code