Skip to content

compass(design): doc 01 D6-D8 on the PDES time model: Mooncake scheduler-side hooks, router as a channel, counted arrivals - #448

Draft
jgong5 wants to merge 1 commit into
feature/atomcompass_newfrom
compass/doc01-d6-d8-pdes
Draft

jgong5 wants to merge 1 commit into
feature/atomcompass_newfrom
compass/doc01-d6-d8-pdes

Conversation

@jgong5

@jgong5 jgong5 commented Sep 28, 2026 •

Copy link
Copy Markdown
Owner

Part of #443. REQUEST CHANGES at 0ed78c98a: 1 blocking (R1, the PP head).

Moves D6-D8 of 01_execution_and_time_model.md onto the PDES time model (design sections 4.2, 4.9, 4.9.1, 4.10, 4.11, 7 and 10.2 of the owner's v0.20 PDES design that #443 summarizes). One file, no code. Every code citation it adds was checked at 52c37317b.

Dev record

Register impact

T13 (the simulated KV connector's completion semantic) is answered by D6 and can be closed. Found 3(d) and 4 are new items for the lead to file.

Named result

none: design-doc only.

Gates

Branch: tests/compass/test_kv_blob_doc_table.py 10 passed, tests/compass/test_sync_inventory.py 26 passed. Both parse this doc.

What changes, by section (old -> new, why)
  • D6 Design. Old: the worker's get_finished() releases a transfer on the virtual clock; Mooncake vs MoRI-IO undecided. New: Mooncake, the backend ATOM's PD CI deploys (pd_server_atom.sh:560/612). The connector implements scheduler-side hooks only:
    • decode posts the write request in update_state_after_alloc (scheduler.py:2196) and computes its ready time as a + T + notify latency;
    • prefill records r in request_finished (:2708/:2768), takes the write request inline in process_completions (:3002-3004) and frees the parked blocks at max(a, r) + T;
    • write-done is not a message; each end computes the time. Prefill asserts r <= a.
    • Why: it keeps the real write-request channel (the owner's rule to use only real channels) and shares no LP state with a worker. Ceiling: modelling bandwidth contention would need the write-done message back, plus a handler thread allowed to send.
  • D6 connector landscape: adds that process_completions is not on the ABC (Found 1).
  • D6 park bullet: cites Mooncake (:320-328, flag cleared at :383) instead of MoRI-IO.
  • D6 blob prose: the simulated connector emits the 17-field Mooncake blob, not the 13-field one. The table is unchanged.
  • D6 Cons: the MoRI-IO last-status bullet is removed. "Pick one and declare it" is resolved: one event per request, unequal rank shares not modelled, PP's MSG_RELEASE out of scope.
  • D7. Old: the router has no virtual clock. New: in 1P1D the router is one segment of three channels, each with a real LP at both ends. Its per-request delay is those channels' lookahead, and (channel, seq) counting covers its wall time. Timestamps travel in carriers the router already forwards, each checked in the router source and in openai-protocol 1.0.0: requests in the compass entry of the tracestate header; the relay in a field of kv_transfer_params; streamed output in SSE comment lines. None needs a router change (the owner's rules to follow atomesh's routing as-is and keep the deployment form). xPyD is deferred: there the router would become an LP, because its policies read load, randomness or wall-clock time.
  • D8. Old: delivery lag is handled by an enqueue ack, only for open-loop arrivals; arrival time in the compass_arrival body field. New: delivery lag is the general transient-message case, handled by D3's send/receive counting for closed and open workloads. Arrival time moves to tracestate, because the router drops unknown top-level fields from a chat request. "Engine LP bounded by T + L[traffic->engine]" now reads as a lookahead distance, because M1-M3 now have a frontend LP.
Found while drafting

Citations at 52c37317b.

  1. process_completions is not a Mooncake hook. It is not on KVConnectorSchedulerBase (base.py:77-104), and neither Mooncake nor MoRI-IO defines it. The scheduler calls it by getattr (scheduler.py:3002-3004); only the offload (offload/connector.py:141) and multi (multi_connector.py:414) connectors define it. Design section 4.9 implies it is a Mooncake hook. It still works as the inline receive point, but the simulated connector has to define it. The doc says this.
  2. The write-request endpoints move by process. Real Mooncake sends MSG_WRITE_REQUEST from the decode worker's start_load_kv (mooncake_connector.py:1048) to the prefill worker's listener thread (:1139). The simulated version sends from the decode engine process to the prefill engine process: same LP pair, message and trigger, different process endpoints. Design section 3.3 says this. Flagged for the reviewer against the owner's rule to use only real channels.
  3. The landed code differs from the new D6 (checked in compass: a simulated KV connector on an injected clock (#156) #162 b7430a841, compass: park a remote fill and hand back what the router relays (#160) #170 5a07d4489, compass(runner): build the worker-side KV connector from the Compass allocate_kv_cache #428 ad5d505d0):
    • (a) timing sits in the worker half: SimulatedKVConnector.start_load_kv/get_finished read an injected compass_clock; the scheduler half holds no clock. compass(runner): build the worker-side KV connector from the Compass allocate_kv_cache #428 builds that worker half. The new D6 has no worker participation;
    • (b) the blob is MoRI-IO's 13 fields (handoff.transfer_params). Mooncake's 17 are needed; hash_block_size enables the incremental path;
    • (c) the claim mark is MoRI-IO's kv_async_tagged, and do_remote_prefill is not cleared, which is why the landed code offers the transfer only at alloc and announces it in build_connector_meta. Posting in update_state_after_alloc is safe only with Mooncake's flag clearing (:383);
    • (d) latent producer block leak. The scheduler half never announces reqs_to_save. On the producer, update_state_after_alloc returns early without do_remote_prefill, and build_connector_meta builds only recv entries, so an engine-P worker never starts a send and finished_sending never arrives. Every finished prefill stays in deferred_free_blocks (scheduler.py:2772-2774) and is never freed (:3038-3046). Tests drive the producer worker only with hand-built metadata (test_kv_simulated_connector.py:188). Nothing reaches this until M4 runs end to end. It needs an issue; the new D6 removes it by design;
    • (e) transfers are priced from each side's own worker announcement, not from a + T + notify at decode and max(a, r) + T at prefill.
  4. --dp-aware breaks the 1P1D premise. CI passes --dp-aware with DP attention (pd_server_atom.sh:649-655). The router then builds one worker per DP rank (create_worker.rs:227), and the policy chooses among them (dp_sticky reads Instant::now()). So design section 4.9.1's condition that the router has nothing to choose fails, even at 1P1D. The design does not cover this; the doc puts it in the deferred xPyD case.
  5. request_finished runs twice per finished request in one postprocess (scheduler.py:2708, :2768). Both calls are at the same virtual time, so the first should win.
  6. Sibling-doc staleness, not edited here; it belongs to the 06 PR and the README owner:
    • 06_workload_harness_contract.md (:181) still describes take2's compass_arrival request half;
    • its D30 says the timeline rides only kv_transfer_params, and so does the README Scope row "Atomesh router";
    • D8's open issue on measurement egress is partly answered by the SSE comment stamps.
  7. Design section 10.2's description of the current text matches the tip in substance; its line numbers drifted by about 40. One existing cite was corrected: http_pd_router.rs:1073-1078 is now :1067-1078.

Generated with Claude Code

… router as a channel, counted arrivals

D6: the simulated connector reproduces Mooncake, the backend ATOM's PD CI
deploys, through scheduler-side hooks only. Decode posts the write request
in update_state_after_alloc and computes its ready time a + T + notify
latency; prefill records r in request_finished, takes the write request in
process_completions and frees its blocks at max(a, r) + T. Write-done is
computed at both ends rather than sent; prefill asserts r <= a. States the
ceiling: bandwidth contention would need the write-done message back and a
handler thread allowed to send. Notes that process_completions is not on
the connector ABC.

D7: in 1P1D the router is one segment of three channels, its overhead is
their lookahead, and the timestamps ride carriers it already forwards
(tracestate entry, a kv_transfer_params field, SSE comment lines). A
--dp-aware launch and xPyD make the router an LP; deferred.

D8: delivery lag is the general transient-message case, covered by
(channel, seq) counting for closed and open workloads alike; the enqueue
acknowledgement is dropped and the arrival time moves from the body field
compass_arrival, which the router drops on chat requests, to tracestate.

Part of #443.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
which the step loop receives inline. At `max(a, r) + T` it reports
`finished_sending`, and the scheduler frees the blocks it parked in
`deferred_free_blocks` (`:2772-2774`, freed at `:3038-3046`). A local event of
engine-P. Scheduler code is unchanged.

@jgong5 jgong5 Sep 29, 2026 •

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Required (R1). The release this text describes happens on EngineCore and DPEngineCoreProc, never on the PP head: PPEngineCoreProc._poll_kv_transfer_progress calls Scheduler._update_from_kv_xfer_finished only for non-empty worker output, so prefill never frees and decode never gets finished_recving. Say so in the Cons bullet at L1188-1192, with what M7 does: (a) the PP head passes its output on every poll, empty included; or (b) the simulated connector refuses PP with a named reason. #466 carries the choice.

Evidence

pp_engine_core.py:261-262 for recv; it returns at :272-273 otherwise. Measured at this head with the real PPEngineCoreProc._poll_kv_transfer_progress and Scheduler._update_from_kv_xfer_finished, an empty KVConnectorOutput from the worker, one request parked in deferred_free_blocks: 50 polls, 0 process_completions calls, block still held. The same probe through the real EngineCore.busy_loop gets one call per idle iteration and frees on the priced poll (review body).

PP is in scope at M7, and the NER list of the owner's v0.20 PDES design names the PP head loop (pp_engine_core.py:67) among the loops whose next event includes KV completion. (a) is a small refactor of a reused module, which README design principle 1 allows; (b) follows principle 6.

which prefill asserts and refuses otherwise; and `T` depends only on this request's
bytes and the configuration, not on other writes in flight. No channel is added; one
message whose time both ends already know is not simulated.
- **Why this is the real channel.** A simulated component keeps the real one's

@jgong5 jgong5 Sep 29, 2026 •

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ponytail shrink: the first two sentences restate D3's rule that a replaced component reproduces the real one's channel (#446, "Lookahead sources"). Keep the Mooncake-specific endpoint sentence and cite D3. Saves about 2 lines.

`hash_block_size`, `local_slot_index`. They are a second backend's blob, not optional
fields of one, and a simulated connector standing in for `moriio` emits the thirteen.
fields of one, and the simulated connector, standing in for `mooncake`, emits the
seventeen.

@jgong5 jgong5 Sep 29, 2026 •

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reservation: D7 adds a relay timestamp inside kv_transfer_params, making the relayed blob Mooncake's 17 fields plus one. Say here that frontend-P writes the stamp outside the connector's blob, or the handoff field-set test in #466's file set will pin the wrong set (17 or 18).

read reaches its result, and the router is **one segment of a channel**, not an LP. Each
of the three channels through it has a real LP at both ends:

| Channel | Registered by |

@jgong5 jgong5 Sep 29, 2026 •

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ponytail shrink: this table repeats the sender column of D3's channel table (#446 already has frontend-P->frontend-D:relay with its sender and lookahead). Replace it with one sentence naming the three channels and pointing to D3. Saves about 6 lines.

|---|---|---|
| `traffic->frontend-P:http` | Parses the body into a typed request and re-serializes it. `CompletionRequest` keeps unknown fields (`#[serde(flatten)]`, `completion.rs:146`); `ChatCompletionRequest` has no such field (`chat.rs:151`), so **an unknown top-level chat field is dropped**. Request headers are forwarded by allow-list, `tracestate` included (`header_utils.rs:51-63`), to decode as well. | The `compass` entry of the W3C `tracestate` header, e.g. `tracestate: compass=a:12.345;s:17`, appended after any existing entry so real tracing is unaffected. The same for chat and completion. |
| `frontend-P->frontend-D:relay` | Takes `kv_transfer_params` from prefill's JSON, lets the ATOM adapter insert fields, and writes it into the decode body (`http_pd_router.rs:1067-1114`). | A field inside `kv_transfer_params`. frontend-D also receives the original `tracestate`, so a request carrying `kv_transfer_params` takes its stamp from there (relay channel), any other from `tracestate` (traffic channel). |
| `frontend-D->traffic:stream` | Passes the ATOM stream through byte for byte: `create_streaming_response` (`http_pd_router.rs:1587`) rewrites only with `return_logprob` and prefill logprobs, and the ATOM path has neither. | An SSE comment line before each event, e.g. `: compass a=12.345 s=18`. SSE clients ignore lines starting with a colon. |

@jgong5 jgong5 Sep 29, 2026 •

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reservation: this row covers streamed responses only. A non-streaming carrier exists: the ATOM relay returns decode's body byte for byte and keeps its response headers (http_pd_router.rs:1182-1192, header_utils.rs:19-31). The D8 / 06 traffic adapter will need it named. Add one sentence: the traffic LP always requests stream: true (a stated rule), or name the non-streaming carrier, e.g. a response header.

A single deployment (M1-M3, no router) uses the same carriers.

**xPyD is deferred.** With several instances per role, or `--dp-aware`, the policies
this router ships read state that changes over time — `power_of_two` and `cache_aware`

@jgong5 jgong5 Sep 29, 2026 •

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reservation: round_robin and prefix_hash also ship in atom/mesh/src/policies/. round_robin's counter (round_robin.rs:43) decides by arrival order at the router, which is wall-clock order under concurrency; prefix_hash falls back on load() (prefix_hash.rs:12-13). Both support the conclusion, but "the policies this router ships" reads as a complete list. Name all six or write "e.g.".

@jgong5

jgong5 commented Sep 29, 2026 •

Copy link
Copy Markdown
Owner Author

REQUEST CHANGES at 0ed78c98a: 1 blocking. Under PP, process_completions is never reached, so the D6 release does not happen (R1, inline).

Agent-authored review, requested by the owner. Base 52c37317b; code claims checked at the integration tip f87413a7a, which differs from the base only in AI_DEV_RULES.md.

R1: PPEngineCoreProc._poll_kv_transfer_progress passes worker output to the scheduler only when it is non-empty; measured 50 polls, 0 calls, block still held. PP is in scope at M7. Fix: state the limit beside the MSG_RELEASE bullet and choose (a) the PP head passes empty output through, or (b) the connector refuses PP.

Checked: every cite listed below holds; process_completions runs on an idle EngineCore and DPEngineCoreProc (measured), so #466 closes #465 on the non-PP loops, and fully only with R1's PP decision. Siblings #446, #450, #466, #469 are consistent, with two notes on #450. Gates: branch and control both 36 doc-test passes and 5281 passed in gate_cpu.sh, no delta.
Accepted with reservation: six items below, three inline.
Watch next: #466 takes R1's PP decision into its exit criteria; the evidence supports #465's candidate 3 (close as superseded by #466) for non-PP deployments.

The idle question behind #465: does `process_completions` run on an idle engine?

Yes on EngineCore and DPEngineCoreProc; no on the PP head (R1). The chain at f87413a7a:

  1. A finished prefill leaves its sequence in deferred_free_blocks (scheduler.py:2772-2774). Scheduler.is_finished() counts deferred_free_blocks (scheduler.py:1219-1224), so an idle producer is not "finished".
  2. busy_loop therefore keeps calling _process_engine_step (engine_core.py:326-327). schedule() returns None because waiting and running are empty (scheduler.py:1437-1438).
  3. _process_engine_step_inner then calls _advance_idle_kv_transfer (engine_core.py:363-365), paced by KV_IDLE_DRAIN_INTERVAL_S (:55, gate :454-456), which calls _poll_kv_transfer_progress (:459).
  4. _poll_kv_transfer_progress passes the aggregated output on unconditionally (:485-489). _update_from_kv_xfer_finished then calls process_completions (scheduler.py:3002-3004).
  5. The same holds when is_finished() is true and only has_pending_kv_work() is: engine_core.py:328-329, :441-442. For DP: engine_core.py:724-731.

Measured at this head in the gpu_docker container, no GPU. The probe ran the real EngineCore.busy_loop, Scheduler.is_finished, Scheduler.schedule (up to its None return) and Scheduler._update_from_kv_xfer_finished. Fakes stood in for the worker aggregation (empty KVConnectorOutput), the block manager and two schedule() preamble helpers (_promote_ready_remote_kv_requests, _park_ready_offload_partial_prefills). One sequence was parked in deferred_free_blocks; the connector's process_completions returns finished_sending on its k-th call.

run idle loop iterations process_completions calls in the loop result
k = 3 50 3, then idle (the loop stops polling once nothing is pending) freed on call 3
k = 1000 (control: priced time not reached in the loop) 50 50, one per iteration; the exit drain supplied the rest still held at loop exit
k = 100000 50 50, then 1894 in the 2.0 s shutdown drain never freed; drain warns and gives up, as expected
PP head, _poll_kv_transfer_progress x50 none 0 never freed

The idle gap the #465 developer found belongs to the worker-announced design, where metadata reaches the worker only with a scheduled batch and the idle path dispatches it only for offload connectors (engine_core.py:491-499). Releasing from process_completions does not depend on that dispatch.

One precondition the doc could name: _update_from_kv_xfer_finished returns early on None (scheduler.py:2995-2996), and call_func_with_aggregation returns None when a worker has not answered within 10 s (async_proc.py:456-464). A worker with no connector still answers with an empty KVConnectorOutput (model_runner.py:3380-3382), so "no worker takes part" is safe as long as the worker keeps answering the poll.

Claims checked (all hold at `f87413a7a` unless noted)
  • process_completions is not on the ABC. It is defined only by offload/connector.py:141, offload/_offload_common.py:268 and multi_connector.py:414, not by Mooncake or MoRI-IO, and is called by getattr at scheduler.py:3002-3004.
  • Decode posts in update_state_after_alloc, called via scheduler.py:1614 -> :2194-2196. Mooncake clears do_remote_prefill at mooncake_connector.py:383; the park follows at scheduler.py:1616-1622.
  • request_finished runs at scheduler.py:2708 (streaming path, leave_reason set) and at :2768. The deferral is at :2772-2774, the free at :3038-3046.
  • Mooncake endpoints: MSG_WRITE_REQUEST sent at :1048, the listener at :1139, write-done at :1313, MSG_RELEASE at :1759, the all-pairs rule at :1761-1832.
  • Router: the missing blob is a hard error at http_pd_router.rs:1067-1078; kv_transfer_params is relayed into the decode body at :1067-1114. enrich_decode_kv only inserts fields (atom.rs:37-59). tracestate is forwarded via build_worker_post_with_headers (:1829-1836) -> header_utils.rs:51-63 on every worker post, and nothing in atom/mesh/src rewrites it. The ATOM stream is passed through (:1587-1615).
  • openai-protocol 1.0.0: CompletionRequest.other is #[serde(flatten)] (completion.rs:145-147); ChatCompletionRequest (chat.rs:151) has no flatten field.
  • CI: Mooncake on both sides (pd_server_atom.sh:560, :612). Single-node PD is 1P1D only (:270-271). --dp-aware is set with agentic DP attention or a rank-mapping policy (:649-655). create_dp_aware_workers is at create_worker.rs:227.
  • Policies: dp_sticky reads Instant::now() (dp_sticky.rs:119); random / power_of_two / cache_aware use unseeded rand::rng(); power_of_two / cache_aware use load().
  • Found 2, the endpoint move from worker to engine process: design section 3.3 of the owner's v0.20 PDES design carries it in the channel table, and compass(design): doc 01 D1-D3.5 protocol to HLA-style PDES grants #446's table matches. No owner ruling needed.
Sibling consistency
Gates

The PR changes one .md that two tests read, so gate 1 applies. Node 18 xiaobizh_n18_cpu, each tree's own scripts, stamps written, atom resolved under each root. Control is the base 52c37317b.

  • tests/compass/test_kv_blob_doc_table.py + tests/compass/test_sync_inventory.py: branch 36 passed, rc=0; control 36 passed, rc=0. Matches the PR body's 10 + 26.
  • scripts/compass/gate_cpu.sh: branch commit: 0ed78c98a (stamp), 5281 passed, 155 skipped, 3 xfailed, GATE_CPU_RC=0; control commit: 52c37317b (stamp), 5281 passed, 155 skipped, 3 xfailed, GATE_CPU_RC=0. No delta.

No new tests and no test credited with pinning a fix, so no reinstatement applies. The PR adds no code, so the design-reference grep has nothing to scan.

R1 options and reservations (not blocking)

R1: the NER list of the owner's v0.20 PDES design names the PP head loop. (a) is a small refactor of a reused module, which README design principle 1 allows; (b) refuses with a named reason (principle 6).

  1. Non-streaming responses have no named carrier (inline).
  2. The relay stamp field versus "the seventeen" (inline).
  3. The xPyD policy list omits round_robin and prefix_hash (inline).
  4. The 1P1D channel premise also assumes the router's limiter is off. --max-concurrent-requests defaults to -1 and the token bucket to unset (cliargs.rs:336-352, :804-807), and CI sets neither. With them on, a wall-time queue or refill decides admission and the router is no longer a pure channel segment. One sentence would cover it.
  5. Router-originated responses have no registered sender, e.g. 502 on a missing blob (http_pd_router.rs:1067-1078) or a decode error (:1154-1167). They are faults under the "stalls the CA cannot see" rule; saying so stops an implementer treating them as messages on frontend-D->traffic:stream.
  6. "Recording r must be idempotent (first call wins)" is in the PR body and in compass(kv): time the simulated KV transfer on Mooncake's scheduler-side hooks instead of the worker #466 but not in the doc. Both calls happen in one postprocess at the same virtual time, so it matters only to avoid double registration.
ponytail-review and next-task notes
  • shrink: the three-row "Registered by" table repeats D3's channel table (compass(design): doc 01 D1-D3.5 protocol to HLA-style PDES grants #446); one sentence plus a pointer (inline).
  • shrink: the replaced-component rule restates D3; keep only the Mooncake-specific sentence (inline).
  • The design itself (no worker half, no write-done message, no router change) is a net simplification.

net: -8 lines possible.

#466: add the free/ready times to the LP's next-event time, so an idle LP's NER target is e. The idle drain pace (KV_IDLE_DRAIN_INTERVAL_S) is a K7 timer in #450, so a report can land up to one drain interval after e, as it does in real ATOM.

Generated with Claude Code

jgong5 added a commit that referenced this pull request Sep 30, 2026
- D4 table loses its Sites column; the crosswalk and the count
  paragraphs go. Each site's mechanism is the `mechanism` field of its
  row in sync_sites.json (#476). No prose states a site count.
- D4, D5 and D9 cite code by path and symbol, not by line; D5's
  cite-audit bullet is deleted.
- D5's K8 table and D9 item 9 state one atomesh bound, the
  --worker-request-timeout-secs option (#478); the 30 s client default
  is never reached (#477). The D7 log row takes #448's decision.
- D4's zero-lookahead and DP rules keep the decisions and link D3 for
  the sizing and the critical-path example.
- "Mistakes fall on the loud side" is limited to sends and receives;
  the clock-source lint and validation catch the rest.
- D4's TSO handler case drops the rejected-design history.
- README: D4 headline row and the doc 01 index row rewritten from the
  log; the Atomesh and wall-clock Scope rows follow D5; the straggler is
  checked on receipt; the header states no decision count.
- "simulated window" becomes "simulation window", "stall report"
  becomes "stall diagnostic".

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@jgong5 jgong5 added the M1 M1: GPU-free as deployed on the PDES clock, scheduling parity with the real engine (D95) label Oct 3, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

M1 M1: GPU-free as deployed on the PDES clock, scheduling parity with the real engine (D95)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant