Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
23 changes: 7 additions & 16 deletions atom/compass/artifacts/matrix.py
Original file line number Diff line number Diff line change
Expand Up @@ -5,17 +5,8 @@
artifact records the fingerprint of **its own dependency row**, so a bump
invalidates the rows that name it and leaves the rest alone.

The table lives in `atom/compass/design/07_calibration_toolchain.md`, in the
section headed *Invalidation*. `test_the_matrix_is_the_documents_table` opens
that file by path, finds the section by its heading and parses the table
back out, so the document is a functional dependency of that test and of
nothing this module emits.

`MATRIX` below is that table transcribed, and it is deliberately a table: the
deliverable is that a reader can hold the document beside the code and check
them cell by cell. `test_the_matrix_is_the_documents_table` does the same check
mechanically, parsing the document's own rows, marks and parentheses, so the
two cannot drift.
`MATRIX` below holds the dependency rows, and it is deliberately a table: the
deliverable is that a reader can check it cell by cell.

**A uniform matrix has lost its point**, and two cells are the ones a tidying
hand would take away:
Expand All @@ -29,11 +20,11 @@
`machine_spec` is three rows here rather than one, and a device change
invalidates two of them while a tokenizer change invalidates the third.

**Two of the seven artifacts this store holds have no row at all.**
`shape_population` and `coverage_hull` appear in the artifact key table and
nowhere in this one, and the two "sevens" are not the same seven (#174). Guessing a row for them is the failure
this package exists to prevent -- an artifact answering under conditions
nobody checked -- so `rows_for` refuses and names what is missing.
**Two artifacts this store holds have no row at all.** `shape_population` and
`coverage_hull` appear in the artifact key table and nowhere in this one.
Guessing a row for them is the failure this package exists to prevent -- an
artifact answering under conditions nobody checked -- so `rows_for` refuses and
names what is missing.
"""

import enum
Expand Down
9 changes: 0 additions & 9 deletions atom/compass/audit/sync_scan.py
Original file line number Diff line number Diff line change
Expand Up @@ -505,15 +505,6 @@ def load_inventory(path: str | os.PathLike | None = None) -> dict:
return json.loads(Path(path or INVENTORY_PATH).read_text(encoding="utf-8"))


def category_counts(inventory: dict | None = None) -> dict[str, int]:
"""How many classified sites sit in each category."""
inv = inventory if inventory is not None else load_inventory()
counts: dict[str, int] = {}
for row in inv["sites"] + inv["anchors"]:
counts[row["category"]] = counts.get(row["category"], 0) + 1
return counts


def repo_root_from_here() -> Path:
"""The checkout this module was imported from."""
return Path(__file__).resolve().parents[3]
32 changes: 17 additions & 15 deletions atom/compass/memory/graph_pool.py
Original file line number Diff line number Diff line change
Expand Up @@ -6,10 +6,11 @@
This module keeps two graph-pool functions apart on purpose:

- **`reserves()` mirrors ATOM's own `_estimate_cudagraph_overhead`**
(`model_runner.py:1560-1655`). It is not the better number and it is not meant
to be. It is the number that *actually reserves the memory*, because it is
what `get_num_blocks` subtracts from the budget, so substituting anything else
there would predict a block count ATOM would never produce.
(`atom/model_engine/model_runner.py`). It is not the better number and it is
not meant to be. It is the number that *actually reserves the memory*,
because it is what `get_num_blocks` subtracts from the budget, so
substituting anything else there would predict a block count ATOM would
never produce.
- **`predicts()` is what the pool really costs**, from the measured form in the
spec: `91.1 MiB + 0.3033 MiB per captured token` at width 1, and flat above it.

Expand Down Expand Up @@ -52,20 +53,21 @@
#: copy of that flag and must not be allowed to disagree with it in silence.
EAGER_SOURCE = "config.enforce_eager"

#: ATOM's declared live-tensors-per-layer coefficient (`model_runner.py:3628`).
#: ATOM's declared live-tensors-per-layer coefficient, in
#: `ModelRunner._piecewise_per_token_bytes`.
#: Mirrored rather than re-derived: it is what the engine spends.
LIVE_TENSORS_PER_LAYER = 2.8
#: ATOM's whole-graph estimate, as a fraction of peak activations (`:1636`).
#: ATOM's whole-graph estimate, as a fraction of peak activations.
ACTIVATION_FRACTION = 0.2
#: The fraction of the utilisation budget the piecewise branch will reserve
#: before it stops taking buckets (`:1605`).
#: before it stops taking buckets.
TARGET_RESERVE_FRACTION = 0.15


def piecewise_per_token_bytes(
*, hidden_size: int, layers: int, dtype_bytes: int, dp_size: int = 1
) -> float:
"""ATOM's per-token retained estimate, from model geometry (`:3616-3639`).
"""ATOM's per-token retained estimate, from model geometry.

Mirrored including the sub-linear `dp ** 0.6` amplification, which is there
because the MoE all-gathers hidden to about `dp_size` times the local tokens
Expand Down Expand Up @@ -102,7 +104,7 @@ class PiecewiseCapture:
budget_bytes: int

def taken(self) -> tuple[tuple[int, ...], int]:
"""The buckets the capture loop keeps, and their token sum (`:1619-1625`).
"""The buckets the capture loop keeps, and their token sum.

Greedy in ascending order and the first bucket is always taken, exactly
as ATOM does it -- the cap is on how much of the budget the reservation
Expand Down Expand Up @@ -139,10 +141,10 @@ def reserves(

- A DSpark confidence-schedule drafter rescales the whole-graph branch by
the captured bucket count, `activation_bytes * 0.2 * n_buckets`
(`model_runner.py:1640-1646`). Without it this **under-reserves** by that
factor and so predicts more KV blocks than ATOM would.
(`_estimate_cudagraph_overhead`). Without it this **under-reserves** by
that factor and so predicts more KV blocks than ATOM would.
- The piecewise branch drops buckets over `ATOM_PIECEWISE_DP_MAX_TOKENS`
when `dp_size > 1` and a drafter is attached (`:1616-1618`).
when `dp_size > 1` and a drafter is attached.
`capture_token_shapes` does not, so that configuration **over-reserves**
and predicts fewer.
"""
Expand All @@ -156,7 +158,7 @@ def reserves(
Basis.DEPLOYMENT,
EAGER_SOURCE,
"ATOM captures no graph under enforce_eager and reserves "
"nothing for one (model_runner.py:1570)",
"nothing for one (ModelRunner._estimate_cudagraph_overhead)",
),
),
)
Expand All @@ -177,7 +179,7 @@ def reserves(
int(activation_bytes * ACTIVATION_FRACTION),
Basis.DECLARED,
f"{ACTIVATION_FRACTION} x {activation_bytes} activation bytes "
"of the warmup shape (model_runner.py:1636)",
"of the warmup shape (ModelRunner._estimate_cudagraph_overhead)",
"ATOM's coefficient, mirrored because it is what reserves; "
"`predicts()` is the measured pool and disagrees by 4-19x",
),
Expand All @@ -194,7 +196,7 @@ def reserves(
f"{piecewise.per_token_bytes / (1 << 20):.3f} MiB/token x "
f"{tokens} tokens over {len(taken)}/{len(piecewise.token_shapes)} "
f"buckets, capped at {TARGET_RESERVE_FRACTION} x budget "
"(model_runner.py:1579-1626)",
"(ModelRunner._estimate_cudagraph_overhead)",
"ATOM's geometry-derived coefficient, mirrored because it is "
"what reserves; `predicts()` is the measured pool",
),
Expand Down
21 changes: 12 additions & 9 deletions atom/compass/memory/readings.py
Original file line number Diff line number Diff line change
Expand Up @@ -4,11 +4,11 @@
"""The five readings `get_num_blocks` takes off a card, taken off a spec instead.

The rule is **substitute the readings, never the arithmetic**. ATOM's
`get_num_blocks` (`model_runner.py:1686-1900`) is five device readings and then
some arithmetic over them; this module owns the five, and the budget formula,
the 2% margin, the `min(budget, free)` clamp and `plan_pools` stay ATOM's. A
copy of that formula here is exactly the drift the substitution exists to avoid,
and upstream will change it.
`ModelRunner.get_num_blocks` (`atom/model_engine/model_runner.py`) is five
device readings and then some arithmetic over them; this module owns the five,
and the budget formula, the 2% margin, the `min(budget, free)` clamp and
`plan_pools` stay ATOM's. A copy of that formula here is exactly the drift the
substitution exists to avoid, and upstream will change it.

| Reading | Where it comes from now |
|---|---|
Expand Down Expand Up @@ -169,9 +169,11 @@ def from_declared_config(
"not a recording -- it is derived from ATOM's own rotary source "
"and validated against no card. cos and sin together are "
"positions x rotary_dim elements, because inv_freq holds "
"rotary_dim/2 of them (model_ops/rotary_embedding.py:58-80), and "
"rotary_dim/2 of them (RotaryEmbedding._compute_inv_freq in "
"atom/model_ops/rotary_embedding.py), and "
"they are resident at the model dtype they are cast to, not the "
"fp32 they are computed in (:39-49, set at model_runner.py:714)",
"fp32 they are computed in (RotaryEmbedding.__init__, with the "
"default dtype set in ModelRunner.__init__)",
)
activations = Term(
"activations",
Expand All @@ -189,7 +191,8 @@ def from_declared_config(
f"{_LIVE_INTERMEDIATE} x {intermediate} intermediate)",
"a liveness walk over a traced op graph, plus the per-leaf "
"invisible-scratch constants, replace this; the per-layer "
"coefficient is ATOM's own (model_runner.py:3628), over one live "
"coefficient is ATOM's own "
"(ModelRunner._piecewise_per_token_bytes), over one live "
"layer rather than all of them",
)
return cls(weights, buffers, activations)
Expand Down Expand Up @@ -334,7 +337,7 @@ def device_readings(
"this configuration does not start on this card. Even "
"--gpu-memory-utilization 1.0 is insufficient -- the non-KV terms "
f"alone are {needed:.2f} of total -- so the lever ATOM names on "
"its own version of this failure (model_runner.py:1725-1733) will "
"its own version of this failure (ModelRunner.get_num_blocks) will "
"not reach it; reduce the width, the model or the warmup shape, "
"or name a larger card. A clamped zero here would be a free "
"reading nobody could read as a refusal",
Expand Down
9 changes: 5 additions & 4 deletions atom/compass/runner/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -29,16 +29,17 @@ class is checked against `overrides.RPC_SURFACE`, the table of names a worker
the table marks unwaited, a hole parks nobody, and what is lost is the work the
name stood for:

- `exit` (`engine_core.py:260`) never reaches `ModelRunner.exit`; the comment
- `exit` (`EngineCore.exit`) never reaches `ModelRunner.exit`; the comment
on the `RPC_SURFACE` check in `model_runner` says what that loses here.
The worker still leaves its loop -- `busy_loop` breaks
on the dispatched name, in a statement beside the per-runner loop rather than
inside it -- so the symptom is what shutdown failed to release, not a hang.
- `process_kvconnector_output` never starts the asynchronous KV transfer its
metadata was built for (a consumer's load, a producer's send, an offload
save). It is broadcast five times and waited for at none of them:
`engine_core.py:378` and `engine_core.py:500`, `pp_engine_core.py:113`,
`pp_engine_core.py:232` and `pp_engine_core.py:369`.
save). It is broadcast from `EngineCore._process_engine_step_inner`,
`EngineCore._dispatch_idle_offload_work`, `PPEngineCoreProc._pp_head_step`,
`PPEngineCoreProc._dispatch_connector_only_batch` and
`PPEngineCoreProc._downstream_busy_loop`, and waited for at none of them.

`overrides` states the reply contract both of these follow from, including why
a method that is present and answers None is the same event to a caller as one
Expand Down
4 changes: 2 additions & 2 deletions atom/compass/runner/model_runner.py
Original file line number Diff line number Diff line change
Expand Up @@ -80,8 +80,8 @@ class CompassModelRunner(NonAllocatingRunner, ModelRunner):
# architecture probe and needs a driver, so a green CPU gate is not
# evidence that the composed class answers the surface -- only an import on
# a machine with a GPU is. And when it does fire inside a worker,
# `AsyncIOProc.__init__` resolves the runner class (`async_proc.py:166`)
# before assigning `self.runners = []` (`:167`), so the atexit finalizer
# `AsyncIOProc.__init__` (`atom/model_engine/async_proc.py`) resolves the
# runner class before assigning `self.runners = []`, so the atexit finalizer
# then fails on the half-built object and the worker log *ends* with
# `AttributeError: 'AsyncIOProc' object has no attribute 'runners'`. The
# refusal is the traceback above that one.
Expand Down
Loading