Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
25 changes: 13 additions & 12 deletions atom/compass/design/01_execution_and_time_model.md
Original file line number Diff line number Diff line change
Expand Up @@ -244,8 +244,8 @@ Recorded as **T47**.
- Simulation wall-clock cost is higher than Option A's. The >=5x target is stated as
negotiable with a bottom line of "faster than real runs". Under saturation the prior
design was 0.30x. This must be measured early, not assumed.
- `torch.cuda.set_device` (`model_runner.py:972`, in `_setup_device_and_distributed`) and
`torch.cuda.mem_get_info` (`model_runner.py:1676`, in `_read_device_memory`) are the two
- `torch.cuda.set_device` (in `model_runner.py::ModelRunner._setup_device_and_distributed`) and
`torch.cuda.mem_get_info` (in `model_runner.py::ModelRunner._read_device_memory`) are the two
hard GPU dependencies a simulated runner must not inherit.

---
Expand Down Expand Up @@ -1077,7 +1077,7 @@ The ABC is small and has **no `send_kv` / `recv_kv` verb** to fake
Every connector's completion reaches the scheduler through **one** method:

```
ModelRunner.async_proc_aggregation model_runner.py:3378-3399
model_runner.py::ModelRunner.async_proc_aggregation
-> EngineCore._poll_kv_transfer_progress engine_core.py:485-489
-> Scheduler._update_from_kv_xfer_finished scheduler.py:2989-3053
```
Expand Down Expand Up @@ -1415,8 +1415,8 @@ re-derive.
**The seam**
- `Config.runner_qualname` — `atom/config.py:1595`; consumed `engine_core.py:129`,
`async_proc.py:166-169`
- `ModelRunner.forward(batch: ScheduledBatch) -> ScheduledBatchOutput` —
`model_runner.py:3262-3350`
- `model_runner.py::ModelRunner.forward`, whose signature is
`forward(batch: ScheduledBatch) -> ScheduledBatchOutput`
- the RPC boundary — `engine_core.py:386-388`
- `ScheduledBatch` fields — `scheduler.py:579-820`; notably `detailed_sqsq` /
`detailed_sqsk` / `detailed_sk` at `:790-792`, which are sum(N_Q^2), sum(N_Q * N_KV),
Expand All @@ -1427,11 +1427,12 @@ re-derive.
**Existing simulation-shaped hooks in ATOM**
- `--load_dummy {empty,zero,xavier}` — `config.py:1556`, `arg_utils.py:260`,
`loader.py:179-227,309-310`, `loading_core.py:266-291`
- meta-device model construction — `RapidServeModelRunner._init_weight_params_on_meta`,
`model_runner.py:4216-4239`
- a working non-allocating runner template — `RapidServeModelRunner` overrides at
`model_runner.py:4245,4259,4266,4272,4288,4296`
- `ModelRunner.dummy_execution()` — `model_runner.py:1191-1231`, shows how to hand-build
- meta-device model construction —
`model_runner.py::RapidServeModelRunner._init_weight_params_on_meta`
- a working non-allocating runner template — `model_runner.py::RapidServeModelRunner`,
which overrides `_build_and_load_model`, `_maybe_warmup`, `_kv_budget_extra_reserve`,
`get_num_blocks`, `allocate_kv_cache` and `forward`
- `model_runner.py::ModelRunner.dummy_execution` shows how to hand-build
a `ScheduledBatch`
- `ScheduledBatch.is_dummy_run` — `scheduler.py:589,781`
- simulated TP (`--fake-eplb`) — `atom/distributed/simulated_tp.py`; explicit precedent
Expand All @@ -1442,8 +1443,8 @@ re-derive.
`tools/parse_trace.py`

**Memory sizing (needed because it decides which configurations exist)**
- `ModelRunner.get_num_blocks()` — `model_runner.py:1686-1899`, with its four `torch.cuda`
reads in `_read_device_memory` (`model_runner.py:1666-1684`). Five device readings plus
- `model_runner.py::ModelRunner.get_num_blocks`, with its four `torch.cuda`
reads in `model_runner.py::ModelRunner._read_device_memory`. Five device readings plus
arithmetic: `mem_get_info`, `allocated_bytes.all.peak`,
`(total - free) - memory_reserved()`, `_estimate_cudagraph_overhead()`, a 2% safety
margin, then `min(budget - ..., free)` and `plan_pools`. Consumed
Expand Down
23 changes: 11 additions & 12 deletions atom/compass/design/02_model_runner_and_cost_backend.md
Original file line number Diff line number Diff line change
Expand Up @@ -26,8 +26,8 @@ enough that it does not become a maintenance burden against upstream ATOM.

| Cut | Where | Replaces | Verdict |
|---|---|---|---|
| `ModelRunner.run_model()` | `model_runner.py:2867-3087` | `model(input_ids, positions)` + `compute_logits` | **Rejected.** Returns `(logits, hidden_states)` as real tensors that `postprocess` indexes, samples and gathers logprobs from. Faking it still requires a `[bs, vocab]` device allocation. |
| `ModelRunner.forward()` | `model_runner.py:3262-3350` | the whole forward, sampling included | **Chosen.** Its contract is `ScheduledBatch` in, `ScheduledBatchOutput` out — both pure numpy/list/int, already pickled across the worker boundary. |
| `ModelRunner.run_model()` | `model_runner.py::ModelRunner.run_model` | `model(input_ids, positions)` + `compute_logits` | **Rejected.** Returns `(logits, hidden_states)` as real tensors that `postprocess` indexes, samples and gathers logprobs from. Faking it still requires a `[bs, vocab]` device allocation. |
| `ModelRunner.forward()` | `model_runner.py::ModelRunner.forward` | the whole forward, sampling included | **Chosen.** Its contract is `ScheduledBatch` in, `ScheduledBatchOutput` out — both pure numpy/list/int, already pickled across the worker boundary. |
| whole-runner replacement via `Config.runner_qualname` | `config.py:1595` | weights, KV tensors, CUDA graphs, sampling | **Chosen as the delivery mechanism** for the above. |

### Decision
Expand All @@ -40,12 +40,11 @@ and `async_proc.py:166-169`. It already has two in-tree users —
swaps in `RapidServeModelRunner` automatically. **The injection itself requires no ATOM
change.**

`RapidServeModelRunner` (`model_runner.py:4195-4710`) is a working template for a
`model_runner.py::RapidServeModelRunner` is a working template for a
non-allocating runner already in the tree. It overrides exactly the memory-owning
methods: `_build_and_load_model` (`:4245`), `_maybe_warmup` (`:4259`),
`_kv_budget_extra_reserve` (`:4266`), `get_num_blocks` (`:4272`),
`allocate_kv_cache` (`:4288`), `forward` (`:4296`) — and constructs parameters on meta
through `_init_weight_params_on_meta` (`:4216-4239`).
methods: `_build_and_load_model`, `_maybe_warmup`, `_kv_budget_extra_reserve`,
`get_num_blocks`, `allocate_kv_cache` and `forward` — and constructs parameters on meta
through `model_runner.py::RapidServeModelRunner._init_weight_params_on_meta`.

### The RPC surface that must be honoured

Expand Down Expand Up @@ -82,8 +81,8 @@ third case — a method that raises — and the table of which names wait.
step for step (mean 8.99 s vs a real 9.00 s).
3. **`produces_output()`** (`scheduler.py:823-840`): a pure-middle-chunk prefill batch
must return an **empty** `token_ids` list with the same `req_ids`, mirroring
the early `return ScheduledBatchOutput(...)` in `ModelRunner.forward`
(`model_runner.py:3329-3335`).
the early `return ScheduledBatchOutput(...)` in
`model_runner.py::ModelRunner.forward`.

Speculative decoding adds `num_rejected` / `num_bonus` sized `batch.total_seqs_num` and
`draft_token_ids` shaped `[bs, mtp_k]`. ATOM's existing `synthetic_acceptance_rates` path
Expand Down Expand Up @@ -118,9 +117,9 @@ missing is a statement of which combination Compass uses, and when.
| Piece | Where | What it does |
|---|---|---|
| `--load_dummy {empty,zero,xavier}` | `config.py:1556`, `arg_utils.py:69,260`, `loader.py:126,234,289-307` | skips the checkpoint read; `empty` leaves params uninitialised, the others fill them with finite values in place |
| `_init_weight_params_on_meta` | `model_runner.py:4216-4239` | wraps `Module.register_parameter` so every `nn.Parameter` is replaced by a meta tensor as it is registered |
| `_init_weight_params_on_meta` | `model_runner.py::RapidServeModelRunner._init_weight_params_on_meta` | wraps `Module.register_parameter` so every `nn.Parameter` is replaced by a meta tensor as it is registered |
| `no_init_weights` | `models/utils.py:457-496` | uses `torch.device("meta")` as a **context manager**, so construction itself lands on meta - no transient, no GPU, and it covers buffers. **Currently unused in ATOM.** |
| `RapidServeModelRunner._build_and_load_model` | `model_runner.py:4245` | the override point where a runner declines to load |
| `RapidServeModelRunner._build_and_load_model` | `model_runner.py::RapidServeModelRunner._build_and_load_model` | the override point where a runner declines to load |

#### Why `_init_weight_params_on_meta` allocates on the real device, despite its name

Expand All @@ -146,7 +145,7 @@ CUDA IPC and recomputes RoPE caches locally.
**So it is not a bug.** It is a deliberate trade — a one-parameter transient bought in
exchange for real buffers and unchanged init branches — and for its use case the trade is
clearly right. It avoids what the call site calls *"the transient 2x-weights peak that
OOMs at TP=4"* (`model_runner.py:4250-4254`), which is the thing that mattered there.
OOMs at TP=4"* (in `model_runner.py::RapidServeModelRunner._build_and_load_model`), which is the thing that mattered there.

Worth noting for anyone reading that code: the call site says construction *"allocates
zero GPU bytes"* while the helper says one parameter is transiently real. Both are true of
Expand Down
51 changes: 26 additions & 25 deletions atom/compass/design/03_memory_and_kv_model.md
Original file line number Diff line number Diff line change
Expand Up @@ -38,7 +38,7 @@ with exactly one point of GPU contact.
| Prefix-cache hash (`compute_hash`, xxhash xxh64 chained with the parent) | `block_manager.py:233-245` | **No** |
| Prefix-cache scan and claim (`can_allocate`, `allocate`) | `block_manager.py:469-561`, `:563-597` | **No** |
| Prefix publish, deferred until the forward computed the KV (`hash_blocks`) | `block_manager.py:696-771` | **No** |
| **`allocate_kv_cache`** — creating the actual tensors | **`model_runner.py:1901-2105`** | **Yes** |
| **`allocate_kv_cache`** — creating the actual tensors | **`model_runner.py::ModelRunner.allocate_kv_cache`** | **Yes** |

`BlockManager.__init__` asserts `num_blocks > 0` (`block_manager.py:77`) and nothing else
about the device.
Expand All @@ -50,7 +50,8 @@ about the device.
Two changes only:

1. `allocate_kv_cache` becomes a no-op in the simulated runner (`RapidServeModelRunner`
already does exactly this at `model_runner.py:4288-4293`), so no tensor is created.
already does exactly this in `model_runner.py::RapidServeModelRunner.allocate_kv_cache`),
so no tensor is created.
2. `get_num_blocks` returns a block count produced from a device model rather than from
device readings. See D14.

Expand Down Expand Up @@ -126,10 +127,10 @@ probed per request. Whether that is 0.1 ms or 10 ms is unmeasured. Recorded as *

### Open issues

- `allocate_kv_cache` also registers tensors globally via `set_kv_cache_data`
(`model_runner.py:2060-2065`) and cross-validates expected against actual bytes
(`:2067-2094`). The no-op must keep whatever downstream code reads from that registry
satisfied, or supply a descriptor-shaped stand-in.
- `model_runner.py::ModelRunner.allocate_kv_cache` also registers tensors globally via
`set_kv_cache_data` and cross-validates expected against actual bytes. The no-op must
keep whatever downstream code reads from that registry satisfied, or supply a
descriptor-shaped stand-in.
- `BlockManager.hash_block_size = block_size * dcp_world_size`
(`block_manager.py:93`) — decode context parallelism changes the hash granularity.
Out of scope now; noted so it is not discovered later.
Expand All @@ -140,28 +141,28 @@ probed per request. Whether that is 0.1 ms or 10 ms is unmeasured. Recorded as *

### Problem

`ModelRunner.get_num_blocks()` (`model_runner.py:1686-1899`) is **five device readings
`model_runner.py::ModelRunner.get_num_blocks` is **five device readings
plus arithmetic**:

```
# in _read_device_memory (:1666-1684), which get_num_blocks calls at :1693
free, total = torch.cuda.mem_get_info() # :1676
peak_torch = max(allocated_bytes.all.peak, .all.current) # :1677-1680
non_torch = max((total - free) - torch.cuda.memory_reserved(), 0) # :1683
# in get_num_blocks (:1686-1899)
cudagraph_overhead = self._estimate_cudagraph_overhead() # :1695
safety_margin = int(total * 0.02) # :1696
budget = int(total * config.gpu_memory_utilization) # :1698
# in _read_device_memory, which get_num_blocks calls before any budget arithmetic
free, total = torch.cuda.mem_get_info()
peak_torch = max(allocated_bytes.all.peak, .all.current)
non_torch = max((total - free) - torch.cuda.memory_reserved(), 0)
# in get_num_blocks
cudagraph_overhead = self._estimate_cudagraph_overhead()
safety_margin = int(total * 0.02)
budget = int(total * config.gpu_memory_utilization)
available_for_kv = min(budget - (peak_torch + non_torch + cudagraph_overhead
+ safety_margin)
- self._kv_budget_extra_reserve(total), free) # :1699-1706
- self._kv_budget_extra_reserve(total), free)
plan = plan_pools(self._sub_pool_specs(), available_for_kv,
config.max_num_seqs) # :1721
num_kvcache_blocks = plan.paged_entries # :1764
# under PP: all_reduce MIN across stages # :1765-1771
config.max_num_seqs)
num_kvcache_blocks = plan.paged_entries
# under PP: all_reduce MIN across stages
```

Line numbers are `model_runner.py` at `feature/atomcompass_new` `75a265a3f`.
`_read_device_memory` is `model_runner.py::ModelRunner._read_device_memory`.

Consumed at `engine_core.py:132-145`, which sets `config.num_kvcache_blocks` before the
`Scheduler` and `BlockManager` are constructed at `:170`.
Expand Down Expand Up @@ -225,15 +226,15 @@ is the right thing to model and it is stated here so it is not discovered as a g
### The pipeline minimum is not inert

An earlier note here said that under PP the `all_reduce(MIN)` across stages
(`model_runner.py:1764-1771`) is inert, *because every stage computes the same number*,
and that it *still needs a live process group or a stub*. Both halves are wrong — and the
engine's own comment on that reduce has said so all along (`model_runner.py:1759-1760`):
(in `model_runner.py::ModelRunner.get_num_blocks`) is inert, *because every stage computes
the same number*, and that it *still needs a live process group or a stub*. Both halves are
wrong — and the engine's own comment on that reduce has said so all along:
*"PP stages compute different block counts; block ids must be valid on every stage's KV
tensor, so reduce to the global minimum."* The code said what this document denied.

**The stages do not compute the same number.** The five readings *are* identical across
stages — nothing in the memory model varies with pipeline rank. The layer count is not:
`_get_total_num_layers` (`model_runner.py:1515-1538`) takes a `get_pp_indices` slice, so
`model_runner.py::ModelRunner._get_total_num_layers` takes a `get_pp_indices` slice, so
each stage sizes its pool from the layers it actually holds. Re-derived from ATOM's own
partitioner at `feature/atomcompass_new` `92f1fdafe`, over the 64-layer hybrid vendored at
`tests/compass/qwen3_5_27b_config.json` — one full-attention layer in four, so 16 of the 64
Expand Down Expand Up @@ -268,7 +269,7 @@ holds the most paged layers — which is not in general the last one: at pp = 5
3 of 5, and at pp = 6 it is four stages of the six.

**Neither a live process group nor a stub is needed.** The reduce is already guarded by
`torch.distributed.is_initialized()` (`model_runner.py:1765`): with no group it does not
`torch.distributed.is_initialized()` (in `model_runner.py::ModelRunner.get_num_blocks`): with no group it does not
run, and with one it runs ATOM's own code unchanged. Building a stub for it is building
something nothing asks for.

Expand Down
6 changes: 3 additions & 3 deletions atom/compass/design/04_model_capture_and_cost_ir.md
Original file line number Diff line number Diff line change
Expand Up @@ -198,9 +198,9 @@ fills and slices, nothing else:
| where | lines | conversions |
|---|---|---|
| `aiter_attention.py in prepare_decode` | 1106, 1115, 1121, 1122, 1123, 1131, 1132 | 10 |
| `model_runner.py in prepare_inputs` | 2468, 2479, 2481 | 4 |
| `model_runner.py in prepare_input_ids` | 510, 513 | 2 |
| `model_runner.py in prepare_sample` | 2564 | 1 |
| `model_runner.py::ModelRunner.prepare_inputs` | 2468, 2479, 2481 | 4 |
| `model_runner.py::tokenIDProcessor.prepare_input_ids` | 510, 513 | 2 |
| `model_runner.py::ModelRunner.prepare_sample` | 2564 | 1 |
| `backends.py in _mrope_cpu_view` | 398, 400 | 2 |
| `gdn_attn.py in _attach_gdn_decode_metadata` | 1237 | 1 |

Expand Down
4 changes: 2 additions & 2 deletions atom/compass/design/07_calibration_toolchain.md
Original file line number Diff line number Diff line change
Expand Up @@ -311,8 +311,8 @@ forward context only the runner establishes, and building it by hand means reimp

The constraints point at the *runner*, not the server — a single-process `ModelRunner` **is**
the worker, so Phase 1b's placement requirement is satisfied. ATOM already shows how to
drive one without a scheduler: `dummy_execution` (`model_runner.py:1191-1231`) and
`warmup_model` (`:1233-1298`) both fabricate `ScheduledBatch`es by hand.
drive one without a scheduler: `model_runner.py::ModelRunner.dummy_execution` and
`model_runner.py::ModelRunner.warmup_model` both fabricate `ScheduledBatch`es by hand.

```
Phase 0 (device-free) standalone bench (GPU, one process)
Expand Down
Loading