compass(design): doc 01 D6-D8 on the PDES time model: Mooncake scheduler-side hooks, router as a channel, counted arrivals - #448
Conversation
… router as a channel, counted arrivals D6: the simulated connector reproduces Mooncake, the backend ATOM's PD CI deploys, through scheduler-side hooks only. Decode posts the write request in update_state_after_alloc and computes its ready time a + T + notify latency; prefill records r in request_finished, takes the write request in process_completions and frees its blocks at max(a, r) + T. Write-done is computed at both ends rather than sent; prefill asserts r <= a. States the ceiling: bandwidth contention would need the write-done message back and a handler thread allowed to send. Notes that process_completions is not on the connector ABC. D7: in 1P1D the router is one segment of three channels, its overhead is their lookahead, and the timestamps ride carriers it already forwards (tracestate entry, a kv_transfer_params field, SSE comment lines). A --dp-aware launch and xPyD make the router an LP; deferred. D8: delivery lag is the general transient-message case, covered by (channel, seq) counting for closed and open workloads alike; the enqueue acknowledgement is dropped and the arrival time moves from the body field compass_arrival, which the router drops on chat requests, to tracestate. Part of #443. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
| which the step loop receives inline. At `max(a, r) + T` it reports | ||
| `finished_sending`, and the scheduler frees the blocks it parked in | ||
| `deferred_free_blocks` (`:2772-2774`, freed at `:3038-3046`). A local event of | ||
| engine-P. Scheduler code is unchanged. |
There was a problem hiding this comment.
Required (R1). The release this text describes happens on EngineCore and DPEngineCoreProc, never on the PP head: PPEngineCoreProc._poll_kv_transfer_progress calls Scheduler._update_from_kv_xfer_finished only for non-empty worker output, so prefill never frees and decode never gets finished_recving. Say so in the Cons bullet at L1188-1192, with what M7 does: (a) the PP head passes its output on every poll, empty included; or (b) the simulated connector refuses PP with a named reason. #466 carries the choice.
Evidence
pp_engine_core.py:261-262 for recv; it returns at :272-273 otherwise. Measured at this head with the real PPEngineCoreProc._poll_kv_transfer_progress and Scheduler._update_from_kv_xfer_finished, an empty KVConnectorOutput from the worker, one request parked in deferred_free_blocks: 50 polls, 0 process_completions calls, block still held. The same probe through the real EngineCore.busy_loop gets one call per idle iteration and frees on the priced poll (review body).
PP is in scope at M7, and the NER list of the owner's v0.20 PDES design names the PP head loop (pp_engine_core.py:67) among the loops whose next event includes KV completion. (a) is a small refactor of a reused module, which README design principle 1 allows; (b) follows principle 6.
| which prefill asserts and refuses otherwise; and `T` depends only on this request's | ||
| bytes and the configuration, not on other writes in flight. No channel is added; one | ||
| message whose time both ends already know is not simulated. | ||
| - **Why this is the real channel.** A simulated component keeps the real one's |
| `hash_block_size`, `local_slot_index`. They are a second backend's blob, not optional | ||
| fields of one, and a simulated connector standing in for `moriio` emits the thirteen. | ||
| fields of one, and the simulated connector, standing in for `mooncake`, emits the | ||
| seventeen. |
There was a problem hiding this comment.
| read reaches its result, and the router is **one segment of a channel**, not an LP. Each | ||
| of the three channels through it has a real LP at both ends: | ||
|
|
||
| | Channel | Registered by | |
| |---|---|---| | ||
| | `traffic->frontend-P:http` | Parses the body into a typed request and re-serializes it. `CompletionRequest` keeps unknown fields (`#[serde(flatten)]`, `completion.rs:146`); `ChatCompletionRequest` has no such field (`chat.rs:151`), so **an unknown top-level chat field is dropped**. Request headers are forwarded by allow-list, `tracestate` included (`header_utils.rs:51-63`), to decode as well. | The `compass` entry of the W3C `tracestate` header, e.g. `tracestate: compass=a:12.345;s:17`, appended after any existing entry so real tracing is unaffected. The same for chat and completion. | | ||
| | `frontend-P->frontend-D:relay` | Takes `kv_transfer_params` from prefill's JSON, lets the ATOM adapter insert fields, and writes it into the decode body (`http_pd_router.rs:1067-1114`). | A field inside `kv_transfer_params`. frontend-D also receives the original `tracestate`, so a request carrying `kv_transfer_params` takes its stamp from there (relay channel), any other from `tracestate` (traffic channel). | | ||
| | `frontend-D->traffic:stream` | Passes the ATOM stream through byte for byte: `create_streaming_response` (`http_pd_router.rs:1587`) rewrites only with `return_logprob` and prefill logprobs, and the ATOM path has neither. | An SSE comment line before each event, e.g. `: compass a=12.345 s=18`. SSE clients ignore lines starting with a colon. | |
There was a problem hiding this comment.
Reservation: this row covers streamed responses only. A non-streaming carrier exists: the ATOM relay returns decode's body byte for byte and keeps its response headers (http_pd_router.rs:1182-1192, header_utils.rs:19-31). The D8 / 06 traffic adapter will need it named. Add one sentence: the traffic LP always requests stream: true (a stated rule), or name the non-streaming carrier, e.g. a response header.
| A single deployment (M1-M3, no router) uses the same carriers. | ||
|
|
||
| **xPyD is deferred.** With several instances per role, or `--dp-aware`, the policies | ||
| this router ships read state that changes over time — `power_of_two` and `cache_aware` |
There was a problem hiding this comment.
Reservation: round_robin and prefix_hash also ship in atom/mesh/src/policies/. round_robin's counter (round_robin.rs:43) decides by arrival order at the router, which is wall-clock order under concurrency; prefix_hash falls back on load() (prefix_hash.rs:12-13). Both support the conclusion, but "the policies this router ships" reads as a complete list. Name all six or write "e.g.".
|
REQUEST CHANGES at Agent-authored review, requested by the owner. Base R1: Checked: every cite listed below holds; The idle question behind #465: does `process_completions` run on an idle engine?Yes on
Measured at this head in the
The idle gap the #465 developer found belongs to the worker-announced design, where metadata reaches the worker only with a scheduled batch and the idle path dispatches it only for offload connectors ( One precondition the doc could name: Claims checked (all hold at `f87413a7a` unless noted)
Sibling consistency
GatesThe PR changes one
No new tests and no test credited with pinning a fix, so no reinstatement applies. The PR adds no code, so the design-reference grep has nothing to scan. R1 options and reservations (not blocking)R1: the NER list of the owner's v0.20 PDES design names the PP head loop. (a) is a small refactor of a reused module, which README design principle 1 allows; (b) refuses with a named reason (principle 6).
ponytail-review and next-task notes
net: -8 lines possible. #466: add the free/ready times to the LP's next-event time, so an idle LP's NER target is Generated with Claude Code |
- D4 table loses its Sites column; the crosswalk and the count paragraphs go. Each site's mechanism is the `mechanism` field of its row in sync_sites.json (#476). No prose states a site count. - D4, D5 and D9 cite code by path and symbol, not by line; D5's cite-audit bullet is deleted. - D5's K8 table and D9 item 9 state one atomesh bound, the --worker-request-timeout-secs option (#478); the 30 s client default is never reached (#477). The D7 log row takes #448's decision. - D4's zero-lookahead and DP rules keep the decisions and link D3 for the sizing and the critical-path example. - "Mistakes fall on the loud side" is limited to sends and receives; the clock-source lint and validation catch the rest. - D4's TSO handler case drops the rejected-design history. - README: D4 headline row and the doc 01 index row rewritten from the log; the Atomesh and wall-clock Scope rows follow D5; the straggler is checked on receipt; the header states no decision count. - "simulated window" becomes "simulation window", "stall report" becomes "stall diagnostic". Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Part of #443. REQUEST CHANGES at
0ed78c98a: 1 blocking (R1, the PP head).Moves D6-D8 of
01_execution_and_time_model.mdonto the PDES time model (design sections 4.2, 4.9, 4.9.1, 4.10, 4.11, 7 and 10.2 of the owner's v0.20 PDES design that #443 summarizes). One file, no code. Every code citation it adds was checked at52c37317b.Dev record
process_completionsis not a Mooncake hook, so the simulated connector must define it (Found 1); the landed connector differs from the new D6 and leaks producer blocks (Found 3);--dp-awarebreaks the 1P1D premise (Found 4).request_finishedruns twice per finished request, so recordingrmust be idempotent, first call wins (Found 5).11_metrics_support.md(compass(design): doc 11 D72/D74 - metrics clocks on virtual time, sim-time observer #444);02,06and15(their own PRs); sibling-doc staleness (Found 6); landed-code findings go to issues.Register impact
T13 (the simulated KV connector's completion semantic) is answered by D6 and can be closed. Found 3(d) and 4 are new items for the lead to file.
Named result
none: design-doc only.
Gates
Branch:
tests/compass/test_kv_blob_doc_table.py10 passed,tests/compass/test_sync_inventory.py26 passed. Both parse this doc.What changes, by section (old -> new, why)
get_finished()releases a transfer on the virtual clock; Mooncake vs MoRI-IO undecided. New: Mooncake, the backend ATOM's PD CI deploys (pd_server_atom.sh:560/612). The connector implements scheduler-side hooks only:update_state_after_alloc(scheduler.py:2196) and computes its ready time asa + T + notify latency;rinrequest_finished(:2708/:2768), takes the write request inline inprocess_completions(:3002-3004) and frees the parked blocks atmax(a, r) + T;r <= a.process_completionsis not on the ABC (Found 1).:320-328, flag cleared at:383) instead of MoRI-IO.MSG_RELEASEout of scope.(channel, seq)counting covers its wall time. Timestamps travel in carriers the router already forwards, each checked in the router source and inopenai-protocol 1.0.0: requests in thecompassentry of thetracestateheader; the relay in a field ofkv_transfer_params; streamed output in SSE comment lines. None needs a router change (the owner's rules to follow atomesh's routing as-is and keep the deployment form). xPyD is deferred: there the router would become an LP, because its policies read load, randomness or wall-clock time.compass_arrivalbody field. New: delivery lag is the general transient-message case, handled by D3's send/receive counting for closed and open workloads. Arrival time moves totracestate, because the router drops unknown top-level fields from a chat request. "Engine LP bounded byT + L[traffic->engine]" now reads as a lookahead distance, because M1-M3 now have a frontend LP.Found while drafting
Citations at
52c37317b.process_completionsis not a Mooncake hook. It is not onKVConnectorSchedulerBase(base.py:77-104), and neither Mooncake nor MoRI-IO defines it. The scheduler calls it bygetattr(scheduler.py:3002-3004); only the offload (offload/connector.py:141) andmulti(multi_connector.py:414) connectors define it. Design section 4.9 implies it is a Mooncake hook. It still works as the inline receive point, but the simulated connector has to define it. The doc says this.MSG_WRITE_REQUESTfrom the decode worker'sstart_load_kv(mooncake_connector.py:1048) to the prefill worker's listener thread (:1139). The simulated version sends from the decode engine process to the prefill engine process: same LP pair, message and trigger, different process endpoints. Design section 3.3 says this. Flagged for the reviewer against the owner's rule to use only real channels.b7430a841, compass: park a remote fill and hand back what the router relays (#160) #1705a07d4489, compass(runner): build the worker-side KV connector from the Compass allocate_kv_cache #428ad5d505d0):SimulatedKVConnector.start_load_kv/get_finishedread an injectedcompass_clock; the scheduler half holds no clock. compass(runner): build the worker-side KV connector from the Compass allocate_kv_cache #428 builds that worker half. The new D6 has no worker participation;handoff.transfer_params). Mooncake's 17 are needed;hash_block_sizeenables the incremental path;kv_async_tagged, anddo_remote_prefillis not cleared, which is why the landed code offers the transfer only at alloc and announces it inbuild_connector_meta. Posting inupdate_state_after_allocis safe only with Mooncake's flag clearing (:383);reqs_to_save. On the producer,update_state_after_allocreturns early withoutdo_remote_prefill, andbuild_connector_metabuilds only recv entries, so an engine-P worker never starts a send andfinished_sendingnever arrives. Every finished prefill stays indeferred_free_blocks(scheduler.py:2772-2774) and is never freed (:3038-3046). Tests drive the producer worker only with hand-built metadata (test_kv_simulated_connector.py:188). Nothing reaches this until M4 runs end to end. It needs an issue; the new D6 removes it by design;a + T + notifyat decode andmax(a, r) + Tat prefill.--dp-awarebreaks the 1P1D premise. CI passes--dp-awarewith DP attention (pd_server_atom.sh:649-655). The router then builds one worker per DP rank (create_worker.rs:227), and the policy chooses among them (dp_stickyreadsInstant::now()). So design section 4.9.1's condition that the router has nothing to choose fails, even at 1P1D. The design does not cover this; the doc puts it in the deferred xPyD case.request_finishedruns twice per finished request in one postprocess (scheduler.py:2708,:2768). Both calls are at the same virtual time, so the first should win.06PR and the README owner:06_workload_harness_contract.md(:181) still describes take2'scompass_arrivalrequest half;kv_transfer_params, and so does the README Scope row "Atomesh router";http_pd_router.rs:1073-1078is now:1067-1078.Generated with Claude Code