Skip to content

compass(design): doc 15 models PP send completion with a rendezvous ack channel - #451

Merged
jgong5 merged 10 commits into
feature/atomcompass_newfrom
compass/doc15-q3-pp-flush
Oct 2, 2026
Merged

jgong5 merged 10 commits into
feature/atomcompass_newfrom
compass/doc15-q3-pp-flush

Conversation

@jgong5

@jgong5 jgong5 commented Sep 28, 2026 •

Copy link
Copy Markdown
Owner

Part of #443. Ready for review. No blocking issues.

Moves D91 Q3 and D93 of 15_parallelism_support.md and their decision-log rows to the PDES time model: a PP send part completes on the sender's clock when eager or by a rendezvous ack that the next stage returns on pp_ack, per the owner's 2026-09-30 ruling. D90 carries a note that the owner's DP ruling of 2026-10-01 (#470) replaces its step pricing; #532 rewrites it. Design-doc change only; no code.

Dev record

Not in this PR: flush_pp_send's reply (#445); the K1-K9 table (#450); the code (#468); the carrier values (#516); the PP channels in the channel table (#446).

Register impact

T67 stays open; #532 restates its target under the DP ruling. Possibly new: Q1's hierarchical CA.

Named result

none: design-doc only.

Gates

none apply: at 5c5d5fbe9 no test or script reads 15_parallelism_support.md, and the branch merges cleanly into the tip aa034ed67.

What changes, by section
  • D91 Q3, PP send completion: a LogGP-style model per carrier (gloo, RCCL); an eager part completes at t_send + o + bytes x G on the sender's clock, a rendezvous part at max(t_send, r_i) + T with r_i the time its receive is posted. Parts fold in the receiver's order (r_1 = t_recv_posted, each next receive posted at the previous part's end plus the receiver's all-gather of a sharded tensor); done is the latest sender-side completion, returned on stage(k+1)->stage(k):pp_ack#dp0 after stage(k)->stage(k+1):pp_data#dp0. L_ack = T_min - L_data; the channel table refuses L_ack <= 0 and a carrier latency_s <= 0. Both waits are K5 receives; the shutdown calls run after the +inf grant, where advance_to and next_event raise, and settle nothing (compass(clock): drop END; finish when no essential work is left, with daemon deadlines #533). The carrier is gloo on the PP CPU group; the calibration fields are named without values.
  • D90: a note and the log row say the 2026-10-01 ruling replaces the step pricing and the exchange; consequence 3 calls the lockstep all_reduce an event cost (K1) only.
  • D93: 1 + deployments x (1 + pp_size) LPs: 3 aggregated, 5 for 1P1D, 6 for PP4 aggregated, matching compass(design): doc 01 D1-D3.5 protocol to HLA-style PDES grants #446.
  • Decision log: D90, D91 and D93 rows restated with their dates; D91 records the 2026-10-02 confirmation.
Evidence
  • Prose fixes carry no pin: no test reads the doc.
  • Code checked for the part order: aiter GroupCoordinator.recv_tensor_dict (recv_object, then each tensor, all-gather after each sharded receive); async_send_intermediate_tensors in atom/distributed/pp_comm.py (every part posted at once); _sparse_kv_indices_gpu is int32 (atom/model_ops/attentions/aiter_mla.py).
  • The owner's design PDF v0.21 section 4.15 carries the same composition, refusal, bytes, shutdown rule and confirmation.

Generated with Claude Code

jgong5 and others added 2 commits September 28, 2026 14:40
…th and makes the idle PP flush an event cost

D90: a DP group's step is a compound event priced by the per-layer critical
path over ranks; max over whole-rank step times is the special case with only
step-level sync and underestimates with per-layer barriers. The lockstep
collective is part of that cost, not ignored.

D91 Q3: the async PP send surfaces in two waits, the send wait inside the
forward and the idle-loop flush_pp_send; the simulated runner answers the
remaining transfer time and the call site advances by it.

Part of #443.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…of category A/B

Aligns with the doc 01 D4 K1-K9 mechanisms: the DP collectives are internal
to the LP and neither is ignored; the lockstep all_reduce carries the batch
descriptions the per-layer critical path is computed from.

Part of #443.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
D93's formula added one API-server LP per run, so 1P1D came to 4 LPs. Doc 01
D3 and D3.1 (#446) give every deployment its own frontend LP (the API
server's event loop is a clock owner) beside its engine LP, plus one traffic
LP per run: 3 aggregated, 5 for 1P1D (traffic, frontend-P, engine-P,
frontend-D, engine-D), 6 for PP4. The formula, its worked examples, D91 Q1's
"multiplies the LP count by P" and the D93 log row now say that.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@jgong5

jgong5 commented Sep 28, 2026 •

Copy link
Copy Markdown
Owner Author

Follow-up commit, 15_parallelism_support.md only: D93's LP formula is now 1 + deployments x (1 + pp_size), matching #446's milestone table: 3 aggregated, 5 for 1P1D, 6 for PP4 aggregated.

Left as is, not covered by the design: D91 Q1's hierarchical-CA rule (the design keeps one CA), and the router LP for more than one replica per role (deferred in the design).

Edits
  • D93 formula: it counted one API-server LP per run, which gives 4 LPs for 1P1D. 01 D3 and D3.1 (compass(design): doc 01 D1-D3.5 protocol to HLA-style PDES grants #446) give each deployment its own frontend LP (the API server's event loop is a clock owner, overlapping the engine's step loop) next to its engine LP, plus one traffic LP per run. 1P1D is traffic, frontend-P, engine-P, frontend-D, engine-D.
  • D93 worked example: "a PP4 deployment is four times that" (12) replaced by 6.
  • D91 Q1: "PP degree P multiplies the LP count by P" -> "turns each engine LP into P LPs", since the traffic and frontend LPs do not multiply.
  • Decision log, D93 row: restated with the new count, marked revised 2026-09-28.
  • D91 Q1's rule left as is: PP stages stay under one node-local CA, and the standalone CA sits at the PD boundary.

…ck channel

The owner ruled on 2026-09-30 to model PP send completion with a rendezvous
ack channel. D91 Q3 now prices a stage send with a LogGP-style model per
carrier (gloo, RCCL): an eager send completes on the sender's clock, a
rendezvous send at max(t_send, t_recv_posted) + T, which the receiving stage
returns on stage(k+1)->stage(k):pp_ack#dp0. The send itself becomes
stage(k)->stage(k+1):pp_data#dp0. The section derives the ack's lookahead,
T_min - L_data, and refuses a declaration where it is not positive. It sets
the stage loop's TAR/NER order that keeps the forward's compute overlap, makes
the idle flush a K5 receive, names the gloo carrier on the PP CPU group,
splits the send's parts by carrier, and lists the per-carrier calibration
parameters without values. The D91 decision-log row is revised to match.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
D91 Q3 now names where the PP send model's parameters live: five optional
fields per carrier under interconnect.intra_node.pp.<carrier>, with the
rendezvous duration written as rendezvous_fixed_s + bytes x G, so T_min is
the duration at the eager threshold. A PP run requires them and is refused
by name until the two-GPU measurement supplies values.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@jgong5

jgong5 commented Sep 30, 2026 •

Copy link
Copy Markdown
Owner Author

Round 2 pushed 9c1f533ad: PP send completion by a rendezvous ack channel (owner ruling 2026-09-30).

  • fb441c6e5: D91 Q3: eager or rendezvous completion, pp_data and pp_ack, L_ack = T_min - L_data (refused if not positive), both waits K5.
  • 2902af898: names the per-carrier calibration fields.
  • 9c1f533ad: the worker never waits on a frame send, and receives only released frames.

test_pp_kv_shard_key.py @ 9c1f533ad: 4 passed, rc 0.

What the reviewer should check first
  • Decided beyond the ruling: the stage(k)->stage(k+1):pp_data#dp0 channel. The receiver needs t_send to compute done, and no channel carried it: meta leaves the head before its own forward, so a downstream forward was never tied to the upstream send. It is the path of the real hidden-state send.
  • Lookahead: the receiver acts at now_r = max(t_recv_posted, t_send + L_data); done - now_r is least, T - L_data, when the receive is posted no later than the send. So L_ack = T_min - L_data, and the pair sums to T_min. L_data is the smallest carrier latency and moves no modelled time.
  • Eager sends take no ack: their completion can precede the receiver's clock.
  • Carrier: gloo point-to-point on the PP CPU group between TP-rank-0 workers, moved by two runner methods the stage loop calls; the loop stamps, registers and releases. A socket between stage loops would add a channel ATOM does not have.
  • Regime per part: _async_send_object's size and metadata on gloo; device tensors on RCCL, decided per send by size. 01 D3's channel table needs both channels; that is compass(design): doc 01 D1-D3.5 protocol to HLA-style PDES grants #446's, held while it carries need human.

@jgong5 jgong5 changed the title compass(design): doc 15 prices a DP step by its per-layer critical path and makes the idle PP flush an event cost compass(design): doc 15 prices a DP step by its per-layer critical path and models PP send completion with a rendezvous ack channel Sep 30, 2026
The carrier paragraph of D91 Q3 now states the two rules that keep the
worker's frame moves visible to the Clock Authority: the send posts the isend
and never waits for it, so the stage loop never waits on the receiving stage
outside a grant; the receive runs only after a grant released the frame, so
the frame is registered and in flight and the wait is bounded.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
as a priced collective (Q3 below).
3. **Both collectives are internal to the LP, and neither is ignored** (revised
2026-09-28, #443). Neither is a cross-LP wait the CA sees. The lockstep `all_reduce`
(`engine_core.py:770`) is an event cost (`01` D4, K1): the ranks exchange their batch

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Required 3 of 3. A line number in a doc goes stale; the rules ask for path and symbol. Cite DPEngineCoreProc._sync_dp_state in atom/model_engine/engine_core.py. Whether this call can carry the batch descriptions at all is escalated in #518; no change asked here for that.

adjacent stages, each on a path the real system has: the stage-to-stage send itself, and
the completion that the receiving transport (RCCL or gloo) returns to the sender. PP runs
only with one DP rank (`CoreManager.__init__` rejects PP with DP), so `#dp0` is the only
instance. The three `pp_transport.py` channels (`meta`, `tokens`, `kv_status`) are listed

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Required 3 of 3. "The three pp_transport.py channels" states a set's size, which the rules forbid in docs. Drop "three": the names follow it.

all-gather before the send when enabled, which is an ordinary priced collective.
materialises.** So its size is computed from geometry, as for KV transfer (`01` D6). The
single intra-node latency and bandwidth of the machine spec do not tell carrier from
protocol, so each carrier (`rccl`, `gloo`) gets five fields under

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Required 3 of 3. "gets five fields" states a set's size. Drop "five": the field names follow it.

tighter than the rendezvous itself. `L_data` is declared as the smallest carrier latency
`L`. It moves no modelled time, because the receiver's clock goes on to `done` (or, for an
eager send, to `t_send + L + o + bytes × G`), and both are at least `t_send + L`; it only
decides how the cycle's lookahead is split. If `T_min <= L_data`, the calibration says a

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Required 1 of 3. The refusal covers L_ack <= 0 but not L_data <= 0. A carrier with latency_s 0 makes pp_data a zero-lookahead channel, and the reason given for refusing a zero L_ack (a zero-lookahead coupling belongs inside one LP) applies to it word for word. Refuse a non-positive latency_s on any carrier as well, naming it; #468's exit then lists both refusals.

site.
- a CPU tensor in the list goes on the PP CPU group, the gloo carrier.

The send completes with its last part. It takes an ack when any part is rendezvous, and

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Required 2 of 3. "done covers every part" leaves the composition undefined, so #468's recv_done has nothing to implement. State it from ATOM's receiver: parts complete in GroupCoordinator.recv_tensor_dict order (metadata first, then each tensor), each part's receive posted when the previous part completes. Fix the sparse_kv_indices bytes too.

Evidence
  • aiter GroupCoordinator.recv_tensor_dict runs recv_object (two blocking gloo receives), then receives each tensor in list order, all-gathering each when pp_send_allgather_group is on. The sender posts every part at once (async_send_intermediate_tensors). So a tensor part's receive is posted at the previous part's completion, not at t_recv_posted.
  • L_ack still holds: the last rendezvous part ends at least T_min after max(t_send, t_recv_posted).
  • sparse_kv_indices is tokens x _pp_index_topk indices (ModelRunner.run_model), not tokens x hidden x dtype; the shard applies only when numel divides by the all-gather width.

optional in the schema and required by a PP run. None has a value until a two-GPU
measurement on RCCL and gloo supplies it (README principle 8); until then a PP run is
refused by name. The `pp_send_allgather_group` path adds a TP-wide all-gather
before the send when enabled, which is an ordinary priced collective.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Part of required 2. The all-gather does not run before the send. The sender sends its shard (async_send_intermediate_tensors) and the receiver all-gathers after each tensor's receive (GroupCoordinator.recv_tensor_dict), so the collective belongs to the receiving stage's time, between parts.

| D90 | DP's two collectives **run for real** — both reduce over scheduling metadata, never over model outputs, so the real reduction is more faithful than a model and free. The DP group stays one LP. Step duration is `max` over ranks, computed not rank-0-sourced, and idle ranks cost a dummy batch. | 2026-09-19 |
| D91 | PP is one LP per stage at microsecond lookahead, and PP boundaries are never a hierarchical-CA cut point. The inter-stage transfer is a **size from the machine spec**, like KV transfer. Layer split comes from `get_pp_indices`, never re-derived; weights shard by that range but **KV shards by the paged-layer count inside it**, which on a hybrid is not proportional to it. Memory readings gain a PP-degree key. | 2026-09-19 |
| D90 | DP's two collectives **run for real** — both reduce over scheduling metadata, never over model outputs, so the real reduction is more faithful than a model and free. The DP group stays one LP. Step duration is `max` over ranks, computed not rank-0-sourced, and idle ranks cost a dummy batch. Revised: the step is a compound event priced by the per-layer critical path over ranks (`max` is its step-sync-only special case). | 2026-09-19, revised 2026-09-28 |
| D91 | PP is one LP per stage at microsecond lookahead, and PP boundaries are never a hierarchical-CA cut point. The inter-stage transfer is a **size from the machine spec**, like KV transfer. Layer split comes from `get_pp_indices`, never re-derived; weights shard by that range but **KV shards by the paged-layer count inside it**, which on a hybrid is not proportional to it. Memory readings gain a PP-degree key. Revised: the send wait inside forward and the idle `flush_pp_send` are event costs. Revised again: a send completes on the sender's clock when eager and at `max(t_send, t_recv_posted) + T` when rendezvous, a per-carrier size threshold deciding which; the receiver returns a rendezvous completion on `stage(k+1)->stage(k):pp_ack#dp0`, lookahead `T_min - L_data`, and both waits receive it (K5). | 2026-09-19, revised 2026-09-28 and 2026-09-30 |

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ponytail: delete: "Revised: the send wait inside forward and the idle flush_pp_send are event costs." It is superseded and now false (both are pp_ack receives). Keep the date in the last column; nothing replaces the clause.

@jgong5

jgong5 commented Sep 30, 2026

Copy link
Copy Markdown
Owner Author

Review round 1. Verdict: CHANGES NEEDED. REQUEST CHANGES: 3 blocking, at 9c1f533ad. Each is a sentence-level edit; the ruled design stands.

  1. Refuse a non-positive carrier latency_s (a zero-lookahead pp_data) as well as L_ack <= 0.
  2. Define how a send's parts compose into done, in the receiver's order; fix the sparse_kv_indices bytes and the all-gather's side.
  3. A line number and two set counts added to the doc.

Beyond the ruling, for the owner: pp_data is sound and necessary. Worker-moved frames are sound (see details).

Checked: every Q3 code claim against the tip 3ca6ab9f9; the L_ack derivation re-derived; the compute order in ModelRunner.run_model; #445, #450, #468, #516 and PDF v0.21 section 4.15 agree. Gates 1-2: not applicable (no test on the tip reads this doc).
Accepted with reservation: the DP exchange on _sync_dp_state restates the ruling; it runs before the step it prices, escalated as #518 (need human), not held here.
Watch next: #468's frame Works must never enter _pp_pending_send.

Beyond-ruling decisions
  • pp_data: PPStageTransport has head-to-stage meta, last-to-head tokens, stage-to-head kv_status, and nothing between stages k and k+1. PPEngineCoreProc._pp_head_step calls send_metadata before the head's forward, so meta carries neither stage k's send time nor, for k > 0, anything from stage k. The real downstream forward blocks in recv_intermediate_tensors until the upstream data arrives, so without pp_data a downstream forward has no dependency on the upstream compute at all, eager or rendezvous. Necessary independent of the ack.
  • Worker frames: the engine process is not in the torch.distributed world, so a frame on the real path must be moved by a worker; a stage-loop socket would add a channel, and relaying through the head would put the head LP on the path. The text keeps stamping, registration and release in the loop, sends never wait, and receives run only after release. Three things compass(runner): PP send completion by a rendezvous ack channel; the idle flush_pp_send receives the pp_ack #468 must hold for that to be true: frame Works retained without waiting and kept out of _pp_pending_send (else ATOM's flush_pp_send waits on the receiver outside a grant); register a frame only after its isend is posted; read every released frame once, in seq order, including an ack released inside the compute TAR.
Derivation and code checks
  • now_r = max(t_recv_posted, t_send + L_data), done = max(t_send, t_recv_posted) + T. Receive posted late: done - now_r = T. Otherwise max(0, t_recv_posted - t_send) + T - L_data, least at T - L_data. Minimum over sends T_min - L_data; L_data + L_ack = T_min. Correct, and it survives any monotone part composition.
  • Regime split: _async_send_object issues two gloo isends on pp.cpu_group; each tensor goes to pp.device_group unless on the CPU; the regime is per send by size. Correct.
  • Overlap: ModelRunner.run_model receives, runs the model, commit_pp_send_work on the previous send, then async_send_intermediate_tensors. The TAR, ack receive, send sequence keeps that order.
  • CoreManager.__init__ asserts PP out with DP; send_intermediate_tensors has callers only in tests/test_pp.py; PP addresses come from get_open_zmq_ipc_path.
  • K5 for the idle and downstream-shutdown flushes, K9 for the head shutdown and the runner's wait(): same in compass(design): doc 01 D4, D5, D9 and the decision log move to the PDES mechanisms #450's table, compass(design): doc 02 D10 keeps ATOM's flush_pp_send answer behind the stage loop's pp_ack receive and refuses nine control-command methods #445's D10 and the PDF.
  • ponytail-review: one superseded clause in the decision log. net: -1 lines possible.

Generated with Claude Code

…nd refuses a zero carrier latency

Review round 1 on #451:
- refuse a non-positive carrier latency_s (a zero-lookahead pp_data)
  beside L_ack <= 0;
- a send's parts complete in GroupCoordinator.recv_tensor_dict order,
  each receive posted when the previous part is done, the receiving
  stage's all-gather between parts; done is the latest sender-side
  completion, and the L_ack bound still holds;
- sparse_kv_indices bytes are tokens x _pp_index_topk int32;
- cite _sync_dp_state by symbol, drop two set counts, and drop the
  superseded clause from the D91 log row.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@jgong5

jgong5 commented Sep 30, 2026

Copy link
Copy Markdown
Owner Author

Round 2 pushed b38ea5967: answers review round 1.

  1. A carrier latency_s <= 0 is refused by name beside L_ack <= 0, both in compass(runner): PP send completion by a rendezvous ack channel; the idle flush_pp_send receives the pp_ack #468's exit.
  2. Parts fold in GroupCoordinator.recv_tensor_dict order, each receive posted at the previous part's end plus the receiver's all-gather; done is the latest sender-side completion, and L_ack still holds. sparse_kv_indices is tokens x _pp_index_topk int32.
  3. _sync_dp_state cited by symbol; set counts and the superseded log clause dropped.

test_pp_kv_shard_key.py @ b38ea5967: 4 passed, rc 0.

Also changed

overhead `o`, per-byte time `G`, bytes from geometry (below).
- **Eager** (bytes up to the carrier's `eager_threshold_bytes`): the send completes on the
sender's clock at `t_send + o + bytes × G`. The receiver plays no part.
- **Rendezvous** (larger): the send completes at `done = max(t_send, t_recv_posted) + T`,

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No blocking issue here: this gives done in its one-part form. Under the fold below, a multi-part send's done is the latest part's completion, and a rendezvous part ends at max(t_send, r_i) + T with r_i >= t_recv_posted. Say a rendezvous part completes at that time (see How the parts combine); the D91 log row carries the same formula.

@jgong5

jgong5 commented Sep 30, 2026

Copy link
Copy Markdown
Owner Author

Review round 2. Verdict: APPROVE. APPROVE at b38ea5967d6935dcd01b502113d36d0f130b852b: the three round-1 findings are closed.

Landing this PR records two decisions beyond the owner's 2026-09-30 ruling that the owner has not ruled on: the stage(k)->stage(k+1):pp_data#dp0 channel, and frames moved by the TP-rank-0 workers over the PP CPU group. The owner rules on both before or by landing.

Checked: delta 9c1f533ad..b38ea5967 (one commit, 15_parallelism_support.md only); the part fold against aiter GroupCoordinator.recv_tensor_dict and ATOM async_send_intermediate_tensors at the tip d629ddd40; the L_ack bound re-derived; #468's body; PDF v0.21 section 4.15; the DP section changes only the _sync_dp_state citation; the branch merges cleanly into the tip.
Accepted with reservation: the Q3 Rendezvous bullet and the D91 log row still give done in its one-part form (inline, non-blocking).
Watch next: the gate line credits test_pp_kv_shard_key.py, which on this branch's base still reads the doc; #492 deleted those asserts on the tip, so no test reads the doc after landing.

Round-1 findings
  1. Closed. Q3 refuses by name L_ack <= 0 (naming T_min and L_data) and any carrier latency_s <= 0 (naming the carrier and value); compass(runner): PP send completion by a rendezvous ack channel; the idle flush_pp_send receives the pp_ack #468's exit lists both.
  2. Closed. The fold matches the code: recv_object makes two blocking gloo receives, then each non-empty tensor is received in list order, a shard followed by the receiving stage's TP all-gather, before the next receive is posted. The sender posts every part without waiting and commit_pp_send_work waits on all of them, so done is the latest. Empty-tensor skip and the shard test (numel divisible by the all-gather width) are the same on both sides. sparse_kv_indices is num_tokens * _pp_index_topk int32 (ModelRunner.run_model, _sparse_kv_indices_gpu).
  3. Closed. The delta adds no line number or set count; the superseded log clause is gone.
Derivation
  • a_i >= r_i in both regimes, so r_i never decreases and every r_i >= t_recv_posted.
  • A send that takes an ack has a rendezvous part j: done >= a_j = max(t_send, r_j) + T_j >= max(t_send, t_recv_posted) + T_min, since a rendezvous part's T is at least T at its carrier's threshold.
  • With now_r = max(t_recv_posted, t_send + L_data), both branches give done - now_r >= T_min - L_data. L_ack = T_min - L_data stands.
  • The receive's end, the last a_i + g_i, is at least a_1 >= t_send + L_data (eager: L >= L_data; rendezvous: T_min > L_data once refused otherwise) and at least done (eager sender-side completion is at most its a_i). So L_data moves no time and the receiver never advances to a point before done.
  • The metadata is one part in the doc and two receives in recv_object. The two agree while both are eager, since an eager arrival depends on r_i only through the max; they differ only for a pickled list above gloo's threshold. No change asked.
ponytail-review

Lean already. Ship.

Generated with Claude Code

@jgong5 jgong5 added the need human Automation stopped; needs owner judgement. Agents apply when escalating, only the owner removes. label Sep 30, 2026
@jgong5

jgong5 commented Sep 30, 2026

Copy link
Copy Markdown
Owner Author

Needs owner ruling: confirm the pp_data channel and worker-carried frames. need human applied.

The owner ruled the rendezvous ack channel on 2026-09-30. This PR also records two decisions beyond that ruling, both judged sound in review:

  • (a) stage(k)->stage(k+1):pp_data#dp0 carries t_send; without it a downstream forward never waits on upstream compute.
  • (b) The TP-rank-0 workers carry these frames over the PP CPU group; the stage loop still stamps, registers and releases them.

Recommendation: confirm both. APPROVE covers b38ea5967; this PR lands once the label is lifted.

@jgong5

jgong5 commented Oct 1, 2026

Copy link
Copy Markdown
Owner Author

Owner ruling, 2026-10-01, affecting D90 in this PR. A DP group's step is now priced as the max over ranks of each rank's own step cost, exchanged by one all_reduce(MAX) inside the forward (#518 option (b)), with the MoE segment priced from max_tokens_across_dp. Under uniform routing this equals the per-layer critical path that D90 states here. Every DP rank also holds its own LP runtime and the CA joins their requests. The full ruling is in #470's first comment.

This does not hold this PR, which carries the PP ruling. The D90 text, its decision-log row and the matching PDF sections are rewritten in #532 after this PR lands.

@jgong5

jgong5 commented Oct 2, 2026

Copy link
Copy Markdown
Owner Author

Needs owner ruling: confirm the pp_data channel and worker-carried frames. need human applied.

The owner ruled the rendezvous ack channel on 2026-09-30. This PR also records two decisions beyond that ruling, both judged sound in review:

  • (a) stage(k)->stage(k+1):pp_data#dp0 carries t_send; without it a downstream forward never waits on upstream compute.
  • (b) The TP-rank-0 workers carry these frames over the PP CPU group; the stage loop still stamps, registers and releases them.

Recommendation: confirm both. APPROVE covers b38ea5967; this PR lands once the label is lifted.

Confirmed both. Also, please check the latest related updates to see if any more changes are needed.

@jgong5 jgong5 removed the need human Automation stopped; needs owner judgement. Agents apply when escalating, only the owner removes. label Oct 2, 2026
jgong5 and others added 2 commits October 2, 2026 14:20
… replaced by the DP ruling

- D91 log row: the owner confirmed on 2026-10-02 the stage(k)->stage(k+1):pp_data#dp0
  channel and the TP-rank-0 workers carrying both channels' frames over the PP CPU group.
- Q3 Rendezvous bullet and the D91 log row give a rendezvous part's completion at
  max(t_send, r_i) + T, with done the latest over the send's parts.
- Q3 shutdown calls: they run after the +inf grant that ends the run, where the loop
  makes no clock call, so they settle nothing.
- D90: a note and its log row say the owner's DP ruling of 2026-10-01 replaces the step
  pricing and the exchange; the sentences putting the exchange on the lockstep
  all_reduce are dropped. The full rewrite is #532.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@jgong5 jgong5 changed the title compass(design): doc 15 prices a DP step by its per-layer critical path and models PP send completion with a rendezvous ack channel compass(design): doc 15 models PP send completion with a rendezvous ack channel Oct 2, 2026
@jgong5

jgong5 commented Oct 2, 2026

Copy link
Copy Markdown
Owner Author

Developer round 3. Round 3 pushed e440a6b98, with the tip 78b0b7ce9 merged. No blocking issues.

Changes and what was left

|---|---|---|
| **TP** | LP collapse; width as an artifact key | priced collectives per width; T21 |
| **DP** | both collectives run for real; `max`-over-ranks step duration; dummy-batch pricing for idle ranks | the collectives' own cost |
| **DP** | both collectives run for real; step priced by the per-layer critical path over ranks; dummy-batch pricing for idle ranks | the collectives' own cost |

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Required 1 of 2. The tip's cell already says what the 2026-10-01 DP ruling says, and this hunk replaces it with the critical-path pricing the ruling replaced; the note on the DP section does not reach this table. Restore the tip's cell: both collectives run for real; + "max-over-ranks step duration; dummy-batch pricing for idle ranks". The PR body's first line then drops the milestone section.

measured whether the Class-C constants move with it. One startup per PP degree settles
it; recorded as **T66**.
- **The DP `max`-over-ranks rule is unmeasured.** `01`'s 0.06% figure is a TP result. What
- **The DP critical-path rule is unmeasured.** `01`'s 0.06% figure is a TP result. What

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Required 1 of 2. Same as the M1 table's DP cell: restore the tip's heading, "The DP max-over-ranks rule is unmeasured."

advance to its local completion if it was eager, nothing if no send is pending. ATOM's own
`flush_pp_send` call then runs, and the simulated runner answers it at once (`02` D10). The
shutdown call in `_downstream_busy_loop` is the same call as the idle one and shares its
answer; it runs after the `+inf` grant that ends the run, where the loop makes no clock

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Required 2 of 2. "Where the loop makes no clock call" is not #533's contract: after +inf, advance_to and next_event raise, but stamp_send keeps answering shutdown sends, and the finally of _downstream_busy_loop sends a last KV status report (_poll_and_send_kv_status) just before this call. Say what holds, e.g. "it runs after the +inf grant that closes the simulation window, where advance_to and next_event raise (#533), so it settles nothing". #468's shutdown bullet ("where a clock call raises") takes the same fix.

- **The LP's step duration is `max` over DP ranks, and it must be computed, not
approximated by rank 0.** Rank-0 single-sourcing is a TP result and does not transfer to
DP.
- **The LP's step is a compound event priced by the per-layer critical path over DP

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ponytail: shrink: this bullet is this PR's own 2026-09-28 text and is the pricing the DP ruling replaced. The tip's bullet ("The LP's step duration is max over DP ranks, and it must be computed, not approximated by rank 0") states the ruling in 3 lines and leaves #532 less to rewrite. Non-blocking.

| D89 | TP is the settled instance and supplies the per-width discipline: width is a key, not a parameter. | 2026-09-19 |
| D90 | DP's two collectives **run for real** — both reduce over scheduling metadata, never over model outputs, so the real reduction is more faithful than a model and free. The DP group stays one LP. Step duration is `max` over ranks, computed not rank-0-sourced, and idle ranks cost a dummy batch. | 2026-09-19 |
| D91 | PP is one LP per stage at microsecond lookahead, and PP boundaries are never a hierarchical-CA cut point. The inter-stage transfer is a **size from the machine spec**, like KV transfer. Layer split comes from `get_pp_indices`, never re-derived; weights shard by that range but **KV shards by the paged-layer count inside it**, which on a hybrid is not proportional to it. Memory readings gain a PP-degree key. | 2026-09-19 |
| D90 | DP's two collectives **run for real** — both reduce over scheduling metadata, never over model outputs, so the real reduction is more faithful than a model and free. The DP group stays one LP. Step duration is `max` over ranks, computed not rank-0-sourced, and idle ranks cost a dummy batch. Revised: the step is a compound event priced by the per-layer critical path over ranks (`max` is its step-sync-only special case). Replaced in part by the owner's DP ruling ([#470](https://github.com/jgong5/ATOM/issues/470#issuecomment-5933154215)); #532 rewrites this row. | 2026-09-19, revised 2026-09-28; ruling 2026-10-01 |

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ponytail: delete: the Revised: ... critical path ... clause, if the DP step bullet reverts to the tip's. Non-blocking.

@jgong5

jgong5 commented Oct 2, 2026

Copy link
Copy Markdown
Owner Author

Review round 3. Verdict: CHANGES NEEDED. REQUEST CHANGES: 2 blocking, at e440a6b98dfa2ab18846ee3e6a44a43f4ac2ffdc. The PP design stands.

  1. This PR moves the M1 table's DP cell and the open-issue bullet on the DP rule from the tip's max over ranks, which the 2026-10-01 DP ruling states, to the per-layer critical path it replaced. The DP section's note covers only that section. Restore the tip's text in both; the PR body's first line then drops the milestone section.
  2. The shutdown sentence says the loop makes no clock call after +inf. compass(clock): drop END; finish when no essential work is left, with daemon deadlines #533 keeps stamp_send for shutdown sends, and the finally of _downstream_busy_loop sends a KV status report just before the flush. Name what raises instead.

Checked: delta b38ea5967..e440a6b98 with remerge; the round-2 reservation is closed; the PP log row records both confirmations; the DP note matches the ruling; no text relies on END or decides whether PP bounded polls are daemon deadlines; no test or script reads the doc.
Accepted with reservation: #468's first line still gives the one-part done, and its shutdown bullet repeats item 2; fix alongside.
Watch next: after this PR only the DP section and its row name the critical path, for #532.

ponytail-review
  • shrink: the DP section's critical-path bullet, this PR's own 2026-09-28 text, is the pricing the ruling replaced; the tip's max-over-ranks bullet states the ruling in 3 lines. Non-blocking.
  • delete: the matching Revised: clause in the DP log row, if the bullet reverts. Non-blocking.

net: -10 lines possible.

Generated with Claude Code

@jgong5 jgong5 added the need human Automation stopped; needs owner judgement. Agents apply when escalating, only the owner removes. label Oct 2, 2026
@jgong5

jgong5 commented Oct 2, 2026

Copy link
Copy Markdown
Owner Author

Needs owner ruling: allow a fourth review cycle? need human applied: the review loop passed three cycles (AI_DEV_RULES.md).

No finding repeated. Round 3's two findings are line edits from today's ruling follow-ups. Recommendation: allow it; the 2026-10-01 and 2026-10-02 rulings grew the scope, the task is not mis-cut.

Cycles and findings

Cycles: round 1 at 9c1f533ad (changes), round 2 at b38ea5967 (approve), round 3 at e440a6b98 (changes).

@jgong5

jgong5 commented Oct 2, 2026

Copy link
Copy Markdown
Owner Author

Allow 3 more rounds of review.

@jgong5 jgong5 removed the need human Automation stopped; needs owner judgement. Agents apply when escalating, only the owner removes. label Oct 2, 2026
…mes what raises after +inf

Review round 3 of #451:
- The M1 table's DP cell and the open-issue bullet on the DP rule go back to the
  integration tip's "max over ranks" text, which the 2026-10-01 DP ruling states.
- The DP section's step bullet goes back to the tip's text, and the D90 log row
  drops its critical-path clause; the note that the ruling replaces the section stays.
- Q3: the shutdown flush runs after the +inf grant, where advance_to and next_event
  raise (#533), so it settles nothing; stamp_send still answers shutdown sends.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@jgong5

jgong5 commented Oct 2, 2026

Copy link
Copy Markdown
Owner Author

Developer round 4. Round 4 pushed 5c5d5fbe9 for review round 3. No blocking issues.

  1. The M1 table's DP cell and the open-issue bullet are the tip's again; the PR body drops the milestone section.
  2. Q3: after the +inf grant advance_to and next_event raise (compass(clock): drop END; finish when no essential work is left, with daemon deadlines #533), so the shutdown flush settles nothing. compass(runner): PP send completion by a rendezvous ack channel; the idle flush_pp_send receives the pp_ack #468 likewise, with the per-part done.
  3. Ponytail: the DP section's step bullet is the tip's, and its log row drops the critical-path clause.
Also changed
  • 15_parallelism_support.md only: 9 lines added, 17 removed. Against the tip only D90's note and consequence 3, D91 Q1 and Q3, D93 and the log rows differ; the doc no longer names the critical path. D90's ruling note stays.
  • The branch merges cleanly into the tip aa034ed67; no test or script reads the doc. Prose fixes carry no pin.
  • The owner's PDF v0.21 section 4.15 shutdown sentence amended in place and rebuilt; page count unchanged.
  • compass(runner): PP send completion by a rendezvous ack channel; the idle flush_pp_send receives the pp_ack #468: body text only; its open question stays with the owner. Its body was over the issue word budget before this edit (250 of 150) and is 258 now.

@jgong5

jgong5 commented Oct 2, 2026

Copy link
Copy Markdown
Owner Author

Review round 4. Verdict: APPROVE. APPROVE at 5c5d5fbe9d629ac2dd7ba24a95fc28a14b6378c7: both round-3 findings are closed.

Checked: delta e440a6b98..5c5d5fbe9 (one commit, 15_parallelism_support.md only) and the whole file against the tip aa034ed67; #468's body edit history and its open question; #533's contract; the branch merges cleanly into the tip and no test or script reads the doc.
Accepted with reservation: the DP log row's date column still says "revised 2026-09-28", while the row no longer states a 2026-09-28 revision (that revision now lives only in consequence 3). #532 rewrites the row.
Watch next: #532 should state consequence 3's event-cost rule in the DP log row, or drop the date; the PR is still a draft, which merge-async refuses.

Round-3 findings
  1. Closed. The M1 table's DP cell, the open-issue bullet on the DP rule and the DP section's step bullet are byte-identical to the tip. The DP log row is the tip's text plus the note that the 2026-10-01 ruling replaces it in part (compass(design): write the DP ruling of 2026-10-01 into docs 15, 01, 09, 08 and the PDES PDF #532). The doc no longer names the critical path. The PR body's first line no longer names the milestone section.
  2. Closed. Q3 now says the shutdown call runs after the +inf grant that closes the simulation window, where advance_to and next_event raise, so it settles nothing. This is compass(clock): drop END; finish when no essential work is left, with daemon deadlines #533's rule word for word, and it no longer claims that no clock call is made, so stamp_send answering shutdown sends and the last _poll_and_send_kv_status in the finally of _downstream_busy_loop are consistent with it.
  3. compass(runner): PP send completion by a rendezvous ack channel; the idle flush_pp_send receives the pp_ack #468: only two body lines changed text, the first line (per-part max(t_send, r_i) + T) and the shutdown bullet (same fix as item 2). The owner question comment is unedited and the need human label stays on compass(runner): PP send completion by a rendezvous ack channel; the idle flush_pp_send receives the pp_ack #468.
ponytail-review

The delta reverts 17 lines to the tip's shorter text and adds 9. Lean already. Ship.

Generated with Claude Code

@jgong5
jgong5 marked this pull request as ready for review October 2, 2026 14:52
@jgong5
jgong5 merged commit 248c4c6 into feature/atomcompass_new Oct 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant