diff --git a/atom/compass/design/01_execution_and_time_model.md b/atom/compass/design/01_execution_and_time_model.md index e007bc1eef..babc763da6 100644 --- a/atom/compass/design/01_execution_and_time_model.md +++ b/atom/compass/design/01_execution_and_time_model.md @@ -46,9 +46,9 @@ one of those branches, on this hardware. These are measurements, not opinions, and they are load-bearing below: - **The seam needs no ATOM change.** `Config.runner_qualname` (`atom/config.py:1595`), - consumed at `engine_core.py:129` and `async_proc.py:166`. Two in-tree precedents: + consumed at `engine_core.py:128` and `async_proc.py:166`. Two in-tree precedents: `RLHFModelRunner` (`atom/rollout/async_engine.py:26-32`) and `RapidServeModelRunner` - (`config.py:1729-1736`). + (`Config.__post_init__`, `config.py:1730-1736`). - **Rank-0 single-sourcing of the clock is correct for symmetric TP.** TP=2 over 1727 steps: per-step rank difference median 0.03%, worst 0.82%, rank 1 slower on 51% of steps. TP=4 over 2295 steps: rank totals within ±0.02%; charging every step to its @@ -941,11 +941,10 @@ that assumes the enqueue is the only thing in that function will be surprised. VRAM for KV sizing only after that ACK. It must stay on the real clock. Nothing in ATOM couples it to whether weights are real, but `Config` keeps a simulated runner from it: `--enable-rapidserve` selects `RapidServeModelRunner` only when `runner_qualname` is - still the default (`config.py:1727-1736`), and otherwise `Config` raises `ValueError` - unless `runner_qualname` is in `RAPIDSERVE_RUNNERS` (`config.py:1737-1745`), before - `LLMEngine.__init__` constructs any engine core. The cost, for a runner that list names, - is two real seconds of startup and no modelled time, because - it runs before READY and therefore before any arrival. + still the default (`config.py:1730-1736`), and otherwise `Config` raises `ValueError` + unless `runner_qualname` is in `RAPIDSERVE_RUNNERS`. The cost, for a runner that list + names, is two real seconds of startup and no modelled time, because it runs before + READY and therefore before any arrival. - The scanner's boundary is a list of directories, not a graph. It reads every `.py` file under `SCANNED_ROOTS`, so a module added beside a scanned one is caught; but a blocking call under one of the three directories `UNSCANNED_ROOTS` names is invisible @@ -1415,7 +1414,7 @@ Facts this design leans on, with their source, so a later reader can re-check ra re-derive. **The seam** -- `Config.runner_qualname` — `atom/config.py:1595`; consumed `engine_core.py:129`, +- `Config.runner_qualname` — `atom/config.py:1595`; consumed `engine_core.py:128`, `async_proc.py:166-169` - `model_runner.py::ModelRunner.forward`, whose signature is `forward(batch: ScheduledBatch) -> ScheduledBatchOutput` diff --git a/atom/compass/design/02_model_runner_and_cost_backend.md b/atom/compass/design/02_model_runner_and_cost_backend.md index 54aa7cafbd..423df19f9c 100644 --- a/atom/compass/design/02_model_runner_and_cost_backend.md +++ b/atom/compass/design/02_model_runner_and_cost_backend.md @@ -36,9 +36,9 @@ enough that it does not become a maintenance burden against upstream ATOM. `Config.runner_qualname` (`atom/config.py:1595`) is consumed at `engine_core.py:125-130` and `async_proc.py:166-169`. It already has two in-tree users — -`atom/rollout/async_engine.py:26-32` injects `RLHFModelRunner`, and `config.py:1729-1736` -swaps in `RapidServeModelRunner` automatically. **The injection itself requires no ATOM -change.** +`atom/rollout/async_engine.py:26-32` injects `RLHFModelRunner`, and `Config.__post_init__` +(`config.py:1730-1736`) swaps in `RapidServeModelRunner` automatically. **The injection +itself requires no ATOM change.** `model_runner.py::RapidServeModelRunner` is a working template for a non-allocating runner already in the tree. It overrides exactly the memory-owning diff --git a/atom/compass/runner/overrides.py b/atom/compass/runner/overrides.py index 7c4e49844d..f0a57e95d0 100644 --- a/atom/compass/runner/overrides.py +++ b/atom/compass/runner/overrides.py @@ -117,8 +117,7 @@ # RapidServe runner (`config.py:1730-1736`) fires only while `runner_qualname` # is still ATOM's default -- which Compass overwrites. `Config` therefore raises # `ValueError` for `enable_rapidserve=True` with any runner not in -# `RAPIDSERVE_RUNNERS` (`config.py:1737-1745`), this one included, and -# `LLMEngine` builds that `Config` (`llm_engine.py:43`) before either class. +# `RAPIDSERVE_RUNNERS`, this one included. # # Also outside the table, and outside anything a broadcast-derived enumeration # can see: three of these twelve are called in-process on the runner itself,