compass(design): doc 15 models PP send completion with a rendezvous ack channel - #451
Conversation
…th and makes the idle PP flush an event cost D90: a DP group's step is a compound event priced by the per-layer critical path over ranks; max over whole-rank step times is the special case with only step-level sync and underestimates with per-layer barriers. The lockstep collective is part of that cost, not ignored. D91 Q3: the async PP send surfaces in two waits, the send wait inside the forward and the idle-loop flush_pp_send; the simulated runner answers the remaining transfer time and the call site advances by it. Part of #443. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…of category A/B Aligns with the doc 01 D4 K1-K9 mechanisms: the DP collectives are internal to the LP and neither is ignored; the lockstep all_reduce carries the batch descriptions the per-layer critical path is computed from. Part of #443. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
D93's formula added one API-server LP per run, so 1P1D came to 4 LPs. Doc 01 D3 and D3.1 (#446) give every deployment its own frontend LP (the API server's event loop is a clock owner) beside its engine LP, plus one traffic LP per run: 3 aggregated, 5 for 1P1D (traffic, frontend-P, engine-P, frontend-D, engine-D), 6 for PP4. The formula, its worked examples, D91 Q1's "multiplies the LP count by P" and the D93 log row now say that. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
Follow-up commit, Left as is, not covered by the design: D91 Q1's hierarchical-CA rule (the design keeps one CA), and the router LP for more than one replica per role (deferred in the design). Edits
|
…ck channel The owner ruled on 2026-09-30 to model PP send completion with a rendezvous ack channel. D91 Q3 now prices a stage send with a LogGP-style model per carrier (gloo, RCCL): an eager send completes on the sender's clock, a rendezvous send at max(t_send, t_recv_posted) + T, which the receiving stage returns on stage(k+1)->stage(k):pp_ack#dp0. The send itself becomes stage(k)->stage(k+1):pp_data#dp0. The section derives the ack's lookahead, T_min - L_data, and refuses a declaration where it is not positive. It sets the stage loop's TAR/NER order that keeps the forward's compute overlap, makes the idle flush a K5 receive, names the gloo carrier on the PP CPU group, splits the send's parts by carrier, and lists the per-carrier calibration parameters without values. The D91 decision-log row is revised to match. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
D91 Q3 now names where the PP send model's parameters live: five optional fields per carrier under interconnect.intra_node.pp.<carrier>, with the rendezvous duration written as rendezvous_fixed_s + bytes x G, so T_min is the duration at the eager threshold. A PP run requires them and is refused by name until the two-GPU measurement supplies values. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
Round 2 pushed
What the reviewer should check first
|
The carrier paragraph of D91 Q3 now states the two rules that keep the worker's frame moves visible to the Clock Authority: the send posts the isend and never waits for it, so the stage loop never waits on the receiving stage outside a grant; the receive runs only after a grant released the frame, so the frame is registered and in flight and the wait is bounded. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
| as a priced collective (Q3 below). | ||
| 3. **Both collectives are internal to the LP, and neither is ignored** (revised | ||
| 2026-09-28, #443). Neither is a cross-LP wait the CA sees. The lockstep `all_reduce` | ||
| (`engine_core.py:770`) is an event cost (`01` D4, K1): the ranks exchange their batch |
There was a problem hiding this comment.
Required 3 of 3. A line number in a doc goes stale; the rules ask for path and symbol. Cite DPEngineCoreProc._sync_dp_state in atom/model_engine/engine_core.py. Whether this call can carry the batch descriptions at all is escalated in #518; no change asked here for that.
| adjacent stages, each on a path the real system has: the stage-to-stage send itself, and | ||
| the completion that the receiving transport (RCCL or gloo) returns to the sender. PP runs | ||
| only with one DP rank (`CoreManager.__init__` rejects PP with DP), so `#dp0` is the only | ||
| instance. The three `pp_transport.py` channels (`meta`, `tokens`, `kv_status`) are listed |
There was a problem hiding this comment.
Required 3 of 3. "The three pp_transport.py channels" states a set's size, which the rules forbid in docs. Drop "three": the names follow it.
| all-gather before the send when enabled, which is an ordinary priced collective. | ||
| materialises.** So its size is computed from geometry, as for KV transfer (`01` D6). The | ||
| single intra-node latency and bandwidth of the machine spec do not tell carrier from | ||
| protocol, so each carrier (`rccl`, `gloo`) gets five fields under |
There was a problem hiding this comment.
Required 3 of 3. "gets five fields" states a set's size. Drop "five": the field names follow it.
| tighter than the rendezvous itself. `L_data` is declared as the smallest carrier latency | ||
| `L`. It moves no modelled time, because the receiver's clock goes on to `done` (or, for an | ||
| eager send, to `t_send + L + o + bytes × G`), and both are at least `t_send + L`; it only | ||
| decides how the cycle's lookahead is split. If `T_min <= L_data`, the calibration says a |
There was a problem hiding this comment.
Required 1 of 3. The refusal covers L_ack <= 0 but not L_data <= 0. A carrier with latency_s 0 makes pp_data a zero-lookahead channel, and the reason given for refusing a zero L_ack (a zero-lookahead coupling belongs inside one LP) applies to it word for word. Refuse a non-positive latency_s on any carrier as well, naming it; #468's exit then lists both refusals.
| site. | ||
| - a CPU tensor in the list goes on the PP CPU group, the gloo carrier. | ||
|
|
||
| The send completes with its last part. It takes an ack when any part is rendezvous, and |
There was a problem hiding this comment.
Required 2 of 3. "done covers every part" leaves the composition undefined, so #468's recv_done has nothing to implement. State it from ATOM's receiver: parts complete in GroupCoordinator.recv_tensor_dict order (metadata first, then each tensor), each part's receive posted when the previous part completes. Fix the sparse_kv_indices bytes too.
Evidence
- aiter
GroupCoordinator.recv_tensor_dictrunsrecv_object(two blocking gloo receives), then receives each tensor in list order, all-gathering each whenpp_send_allgather_groupis on. The sender posts every part at once (async_send_intermediate_tensors). So a tensor part's receive is posted at the previous part's completion, not att_recv_posted. L_ackstill holds: the last rendezvous part ends at leastT_minaftermax(t_send, t_recv_posted).sparse_kv_indicesistokens x _pp_index_topkindices (ModelRunner.run_model), nottokens x hidden x dtype; the shard applies only whennumeldivides by the all-gather width.
| optional in the schema and required by a PP run. None has a value until a two-GPU | ||
| measurement on RCCL and gloo supplies it (README principle 8); until then a PP run is | ||
| refused by name. The `pp_send_allgather_group` path adds a TP-wide all-gather | ||
| before the send when enabled, which is an ordinary priced collective. |
There was a problem hiding this comment.
Part of required 2. The all-gather does not run before the send. The sender sends its shard (async_send_intermediate_tensors) and the receiver all-gathers after each tensor's receive (GroupCoordinator.recv_tensor_dict), so the collective belongs to the receiving stage's time, between parts.
| | D90 | DP's two collectives **run for real** — both reduce over scheduling metadata, never over model outputs, so the real reduction is more faithful than a model and free. The DP group stays one LP. Step duration is `max` over ranks, computed not rank-0-sourced, and idle ranks cost a dummy batch. | 2026-09-19 | | ||
| | D91 | PP is one LP per stage at microsecond lookahead, and PP boundaries are never a hierarchical-CA cut point. The inter-stage transfer is a **size from the machine spec**, like KV transfer. Layer split comes from `get_pp_indices`, never re-derived; weights shard by that range but **KV shards by the paged-layer count inside it**, which on a hybrid is not proportional to it. Memory readings gain a PP-degree key. | 2026-09-19 | | ||
| | D90 | DP's two collectives **run for real** — both reduce over scheduling metadata, never over model outputs, so the real reduction is more faithful than a model and free. The DP group stays one LP. Step duration is `max` over ranks, computed not rank-0-sourced, and idle ranks cost a dummy batch. Revised: the step is a compound event priced by the per-layer critical path over ranks (`max` is its step-sync-only special case). | 2026-09-19, revised 2026-09-28 | | ||
| | D91 | PP is one LP per stage at microsecond lookahead, and PP boundaries are never a hierarchical-CA cut point. The inter-stage transfer is a **size from the machine spec**, like KV transfer. Layer split comes from `get_pp_indices`, never re-derived; weights shard by that range but **KV shards by the paged-layer count inside it**, which on a hybrid is not proportional to it. Memory readings gain a PP-degree key. Revised: the send wait inside forward and the idle `flush_pp_send` are event costs. Revised again: a send completes on the sender's clock when eager and at `max(t_send, t_recv_posted) + T` when rendezvous, a per-carrier size threshold deciding which; the receiver returns a rendezvous completion on `stage(k+1)->stage(k):pp_ack#dp0`, lookahead `T_min - L_data`, and both waits receive it (K5). | 2026-09-19, revised 2026-09-28 and 2026-09-30 | |
There was a problem hiding this comment.
ponytail: delete: "Revised: the send wait inside forward and the idle flush_pp_send are event costs." It is superseded and now false (both are pp_ack receives). Keep the date in the last column; nothing replaces the clause.
|
Review round 1. Verdict: CHANGES NEEDED. REQUEST CHANGES: 3 blocking, at
Beyond the ruling, for the owner: Checked: every Q3 code claim against the tip Beyond-ruling decisions
Derivation and code checks
Generated with Claude Code |
…nd refuses a zero carrier latency Review round 1 on #451: - refuse a non-positive carrier latency_s (a zero-lookahead pp_data) beside L_ack <= 0; - a send's parts complete in GroupCoordinator.recv_tensor_dict order, each receive posted when the previous part is done, the receiving stage's all-gather between parts; done is the latest sender-side completion, and the L_ack bound still holds; - sparse_kv_indices bytes are tokens x _pp_index_topk int32; - cite _sync_dp_state by symbol, drop two set counts, and drop the superseded clause from the D91 log row. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
Round 2 pushed
Also changed
|
| overhead `o`, per-byte time `G`, bytes from geometry (below). | ||
| - **Eager** (bytes up to the carrier's `eager_threshold_bytes`): the send completes on the | ||
| sender's clock at `t_send + o + bytes × G`. The receiver plays no part. | ||
| - **Rendezvous** (larger): the send completes at `done = max(t_send, t_recv_posted) + T`, |
There was a problem hiding this comment.
No blocking issue here: this gives done in its one-part form. Under the fold below, a multi-part send's done is the latest part's completion, and a rendezvous part ends at max(t_send, r_i) + T with r_i >= t_recv_posted. Say a rendezvous part completes at that time (see How the parts combine); the D91 log row carries the same formula.
|
Review round 2. Verdict: APPROVE. APPROVE at Landing this PR records two decisions beyond the owner's 2026-09-30 ruling that the owner has not ruled on: the Checked: delta Round-1 findings
Derivation
ponytail-reviewLean already. Ship. Generated with Claude Code |
|
Needs owner ruling: confirm the The owner ruled the rendezvous ack channel on 2026-09-30. This PR also records two decisions beyond that ruling, both judged sound in review:
Recommendation: confirm both. APPROVE covers |
|
Owner ruling, 2026-10-01, affecting D90 in this PR. A DP group's step is now priced as the This does not hold this PR, which carries the PP ruling. The D90 text, its decision-log row and the matching PDF sections are rewritten in #532 after this PR lands. |
Confirmed both. Also, please check the latest related updates to see if any more changes are needed. |
…ass/doc15-q3-pp-flush
… replaced by the DP ruling - D91 log row: the owner confirmed on 2026-10-02 the stage(k)->stage(k+1):pp_data#dp0 channel and the TP-rank-0 workers carrying both channels' frames over the PP CPU group. - Q3 Rendezvous bullet and the D91 log row give a rendezvous part's completion at max(t_send, r_i) + T, with done the latest over the send's parts. - Q3 shutdown calls: they run after the +inf grant that ends the run, where the loop makes no clock call, so they settle nothing. - D90: a note and its log row say the owner's DP ruling of 2026-10-01 replaces the step pricing and the exchange; the sentences putting the exchange on the lockstep all_reduce are dropped. The full rewrite is #532. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
Developer round 3. Round 3 pushed
Changes and what was left
|
| |---|---|---| | ||
| | **TP** | LP collapse; width as an artifact key | priced collectives per width; T21 | | ||
| | **DP** | both collectives run for real; `max`-over-ranks step duration; dummy-batch pricing for idle ranks | the collectives' own cost | | ||
| | **DP** | both collectives run for real; step priced by the per-layer critical path over ranks; dummy-batch pricing for idle ranks | the collectives' own cost | |
There was a problem hiding this comment.
Required 1 of 2. The tip's cell already says what the 2026-10-01 DP ruling says, and this hunk replaces it with the critical-path pricing the ruling replaced; the note on the DP section does not reach this table. Restore the tip's cell: both collectives run for real; + "max-over-ranks step duration; dummy-batch pricing for idle ranks". The PR body's first line then drops the milestone section.
| measured whether the Class-C constants move with it. One startup per PP degree settles | ||
| it; recorded as **T66**. | ||
| - **The DP `max`-over-ranks rule is unmeasured.** `01`'s 0.06% figure is a TP result. What | ||
| - **The DP critical-path rule is unmeasured.** `01`'s 0.06% figure is a TP result. What |
There was a problem hiding this comment.
Required 1 of 2. Same as the M1 table's DP cell: restore the tip's heading, "The DP max-over-ranks rule is unmeasured."
| advance to its local completion if it was eager, nothing if no send is pending. ATOM's own | ||
| `flush_pp_send` call then runs, and the simulated runner answers it at once (`02` D10). The | ||
| shutdown call in `_downstream_busy_loop` is the same call as the idle one and shares its | ||
| answer; it runs after the `+inf` grant that ends the run, where the loop makes no clock |
There was a problem hiding this comment.
Required 2 of 2. "Where the loop makes no clock call" is not #533's contract: after +inf, advance_to and next_event raise, but stamp_send keeps answering shutdown sends, and the finally of _downstream_busy_loop sends a last KV status report (_poll_and_send_kv_status) just before this call. Say what holds, e.g. "it runs after the +inf grant that closes the simulation window, where advance_to and next_event raise (#533), so it settles nothing". #468's shutdown bullet ("where a clock call raises") takes the same fix.
| - **The LP's step duration is `max` over DP ranks, and it must be computed, not | ||
| approximated by rank 0.** Rank-0 single-sourcing is a TP result and does not transfer to | ||
| DP. | ||
| - **The LP's step is a compound event priced by the per-layer critical path over DP |
There was a problem hiding this comment.
ponytail: shrink: this bullet is this PR's own 2026-09-28 text and is the pricing the DP ruling replaced. The tip's bullet ("The LP's step duration is max over DP ranks, and it must be computed, not approximated by rank 0") states the ruling in 3 lines and leaves #532 less to rewrite. Non-blocking.
| | D89 | TP is the settled instance and supplies the per-width discipline: width is a key, not a parameter. | 2026-09-19 | | ||
| | D90 | DP's two collectives **run for real** — both reduce over scheduling metadata, never over model outputs, so the real reduction is more faithful than a model and free. The DP group stays one LP. Step duration is `max` over ranks, computed not rank-0-sourced, and idle ranks cost a dummy batch. | 2026-09-19 | | ||
| | D91 | PP is one LP per stage at microsecond lookahead, and PP boundaries are never a hierarchical-CA cut point. The inter-stage transfer is a **size from the machine spec**, like KV transfer. Layer split comes from `get_pp_indices`, never re-derived; weights shard by that range but **KV shards by the paged-layer count inside it**, which on a hybrid is not proportional to it. Memory readings gain a PP-degree key. | 2026-09-19 | | ||
| | D90 | DP's two collectives **run for real** — both reduce over scheduling metadata, never over model outputs, so the real reduction is more faithful than a model and free. The DP group stays one LP. Step duration is `max` over ranks, computed not rank-0-sourced, and idle ranks cost a dummy batch. Revised: the step is a compound event priced by the per-layer critical path over ranks (`max` is its step-sync-only special case). Replaced in part by the owner's DP ruling ([#470](https://github.com/jgong5/ATOM/issues/470#issuecomment-5933154215)); #532 rewrites this row. | 2026-09-19, revised 2026-09-28; ruling 2026-10-01 | |
There was a problem hiding this comment.
ponytail: delete: the Revised: ... critical path ... clause, if the DP step bullet reverts to the tip's. Non-blocking.
|
Review round 3. Verdict: CHANGES NEEDED. REQUEST CHANGES: 2 blocking, at
Checked: delta ponytail-review
net: -10 lines possible. Generated with Claude Code |
|
Needs owner ruling: allow a fourth review cycle? No finding repeated. Round 3's two findings are line edits from today's ruling follow-ups. Recommendation: allow it; the 2026-10-01 and 2026-10-02 rulings grew the scope, the task is not mis-cut. Cycles and findingsCycles: round 1 at
|
|
Allow 3 more rounds of review. |
…mes what raises after +inf Review round 3 of #451: - The M1 table's DP cell and the open-issue bullet on the DP rule go back to the integration tip's "max over ranks" text, which the 2026-10-01 DP ruling states. - The DP section's step bullet goes back to the tip's text, and the D90 log row drops its critical-path clause; the note that the ruling replaces the section stays. - Q3: the shutdown flush runs after the +inf grant, where advance_to and next_event raise (#533), so it settles nothing; stamp_send still answers shutdown sends. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
Developer round 4. Round 4 pushed
Also changed
|
|
Review round 4. Verdict: APPROVE. APPROVE at Checked: delta Round-3 findings
ponytail-reviewThe delta reverts 17 lines to the tip's shorter text and adds 9. Lean already. Ship. Generated with Claude Code |
Part of #443. Ready for review. No blocking issues.
Moves D91 Q3 and D93 of
15_parallelism_support.mdand their decision-log rows to the PDES time model: a PP send part completes on the sender's clock when eager or by a rendezvous ack that the next stage returns onpp_ack, per the owner's 2026-09-30 ruling. D90 carries a note that the owner's DP ruling of 2026-10-01 (#470) replaces its step pricing; #532 rewrites it. Design-doc change only; no code.Dev record
ignorelabel is inatom/compass/audit/sync_sites.json, not in the doc; the LP formula gave 4 LPs for 1P1D against the design's 5 (fixed).pp_datachannel and TP-rank-0 workers carrying the frames over the PP CPU group, both confirmed by the owner on 2026-10-02; a send's parts fold inGroupCoordinator.recv_tensor_dictorder; a non-positive carrierlatency_sis refused besideL_ack <= 0.Not in this PR:
flush_pp_send's reply (#445); the K1-K9 table (#450); the code (#468); the carrier values (#516); the PP channels in the channel table (#446).Register impact
T67 stays open; #532 restates its target under the DP ruling. Possibly new: Q1's hierarchical CA.
Named result
none: design-doc only.
Gates
none apply: at
5c5d5fbe9no test or script reads15_parallelism_support.md, and the branch merges cleanly into the tipaa034ed67.What changes, by section
t_send + o + bytes x Gon the sender's clock, a rendezvous part atmax(t_send, r_i) + Twithr_ithe time its receive is posted. Parts fold in the receiver's order (r_1 = t_recv_posted, each next receive posted at the previous part's end plus the receiver's all-gather of a sharded tensor);doneis the latest sender-side completion, returned onstage(k+1)->stage(k):pp_ack#dp0afterstage(k)->stage(k+1):pp_data#dp0.L_ack = T_min - L_data; the channel table refusesL_ack <= 0and a carrierlatency_s <= 0. Both waits are K5 receives; the shutdown calls run after the+infgrant, whereadvance_toandnext_eventraise, and settle nothing (compass(clock): drop END; finish when no essential work is left, with daemon deadlines #533). The carrier is gloo on the PP CPU group; the calibration fields are named without values.all_reducean event cost (K1) only.1 + deployments x (1 + pp_size)LPs: 3 aggregated, 5 for 1P1D, 6 for PP4 aggregated, matching compass(design): doc 01 D1-D3.5 protocol to HLA-style PDES grants #446.Evidence
GroupCoordinator.recv_tensor_dict(recv_object, then each tensor, all-gather after each sharded receive);async_send_intermediate_tensorsinatom/distributed/pp_comm.py(every part posted at once);_sparse_kv_indices_gpuis int32 (atom/model_ops/attentions/aiter_mla.py).Generated with Claude Code