Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 3 additions & 3 deletions atom/compass/audit/sync_sites.json
Original file line number Diff line number Diff line change
Expand Up @@ -584,7 +584,7 @@
"peer": "none",
"why": "the metrics scrape cadence is deliberately left on the real clock",
"mechanism": "K7",
"mechanism_why": "the metrics refresh cadence runs on virtual time once the frontend's event loop does, with no change at the call"
"mechanism_why": "the metrics refresh cadence is a daemon timer on simulated time once the frontend's event loop runs on the LP clock, with no change at the call"
},
{
"id": "atom/entrypoints/openai/api_server.py::chat_completions::loop.run_in_executor( None, _prepare_multimodal_inputs, messages, merged_kwargs, request.tools, )::executor_handoff#0",
Expand Down Expand Up @@ -2955,7 +2955,7 @@
"peer": "none",
"why": "the metrics push cadence is deliberately left on the real clock; the reported timeline comes from sampling, not from this",
"mechanism": "K7",
"mechanism_why": "the metrics push cadence becomes a virtual timer, folded into the loop's idle point"
"mechanism_why": "the metrics push cadence becomes a daemon timer on simulated time, folded into the loop's idle point"
},
{
"file": "atom/model_engine/engine_core.py",
Expand All @@ -2965,7 +2965,7 @@
"peer": "none",
"why": "the same cadence in the data-parallel loop, deliberately left on the real clock",
"mechanism": "K7",
"mechanism_why": "the metrics push cadence becomes a virtual timer, folded into the loop's idle point"
"mechanism_why": "the metrics push cadence becomes a daemon timer on simulated time, folded into the loop's idle point"
},
{
"file": "atom/model_engine/engine_core.py",
Expand Down
25 changes: 16 additions & 9 deletions atom/compass/design/01_execution_and_time_model.md
Original file line number Diff line number Diff line change
Expand Up @@ -1316,9 +1316,13 @@ The pacing timers are local next events the clock owner hands to
engine's metrics push (`EngineCore.busy_loop` and `DPEngineCoreProc.busy_loop`), and
the API server's refresh loop
(`_metrics_refresh_loop` in `api_server.py`). **Metric cadence is virtual time again**,
which revises `11` D72's decision to keep it on the real clock: an observer in the
traffic LP scrapes `/metrics` every `scrape_interval` of simulated time. `11` D72
carries the observer and the reasons.
as revised `11` D72 also decides (its first version kept it on the real clock): an
observer in the traffic LP scrapes `/metrics` every `scrape_interval` of simulated
time. `11` D72 carries the observer and the reasons.
- uvicorn's server tick, a 0.1 s `asyncio.sleep` in `Server.main_loop`, and its
keep-alive timeout are daemon deadlines too. They are uvicorn's code, not ATOM's, so the
sync inventory (`atom/compass/audit/sync_sites.json`) has no row for them; the frontend
declares them with the refresh loop in `DAEMON_TIMERS` (`atom/utils/compass_loop.py`).

### Must read the LP clock — these three change scheduling (K2)

Expand Down Expand Up @@ -1478,9 +1482,9 @@ the producer's `hash_block_size` against its own and falls back to a full transf
(`mooncake_connector.py::MooncakeConnectorScheduler.update_state_after_alloc`), so a blob carrying only the thirteen can never
take the incremental path.

`tests/compass/test_kv_blob_doc_table.py` re-derives both sets from the connectors and
fails naming the field that differs, so this table cannot drift from the source the
way its twelve-field predecessor did.
`tests/compass/test_kv_blob_site.py` holds the one-site claim above. Nothing re-checks
the field sets: they describe the connectors' code as read at `92f1fdafe`, and where the
two differ the code is right.

### Pros

Expand Down Expand Up @@ -1548,9 +1552,12 @@ router-side latency injection is ever wanted, this is where it goes.

### Open issues

- The two hardcoded Rust timeouts (`worker.rs::DEFAULT_WORKER_HTTP_TIMEOUT_SECS`, `worker_manager.rs::DEFAULT_WORKER_REQUEST_TIMEOUT_SECS`) require
touching the Rust crate, which means `ATOM_MESH_BUILD=1` and a Rust toolchain in the
loop. Confirm the container has one.
- `worker.rs::DEFAULT_WORKER_HTTP_TIMEOUT_SECS` is the one Atomesh timeout still compiled
in; changing it means touching the Rust crate, with `ATOM_MESH_BUILD=1` and a Rust
toolchain in the loop. D5 finds a simulated launch never reaches it, so this stays open
only if that launch ever keeps the health check on.
`worker_manager.rs::DEFAULT_WORKER_REQUEST_TIMEOUT_SECS` is only the default of
`--worker-request-timeout-secs` (#478), which D5 sets large.
- Mesh-only mode has not been exercised by this project. Confirm it serves the ATOM
relay path correctly against two ATOM servers before depending on it.

Expand Down
77 changes: 40 additions & 37 deletions atom/compass/design/11_metrics_support.md
Original file line number Diff line number Diff line change
Expand Up @@ -45,7 +45,7 @@ influence how they land rather than retrofit them. See D73.

---

## D72. Counters and gauges are valid by construction; their clocks stay real
## D72. Counters and gauges are valid by construction; their cadence is simulated time

### Why they are valid

Expand All @@ -54,47 +54,50 @@ influence how they land rather than retrofit them. See D73.
run unmodified under simulation** (`03` D13). Nothing there is a timing read, and there are
no rate denominators to correct.

### The two clock reads, and why both stay on the real clock
### The two timers and the stamp, all on the LP clock

```python
# engine_core.py:317-320 <- inside EngineCore.busy_loop, the process that OWNS the clock
now = time.monotonic()
# EngineCore.busy_loop and DPEngineCoreProc.busy_loop (atom/model_engine/engine_core.py)
now = clock.now(time.monotonic)
if now >= next_metrics_push:
next_metrics_push = now + METRICS_PUSH_INTERVAL_S # 5.0, engine_core.py:50
next_metrics_push = now + METRICS_PUSH_INTERVAL_S # 5.0
self.utility_handler.push_metrics()

# metrics.py:408 <- in the API-server process
self._last_refresh = time.time() # "when was this snapshot taken", returned by read()
```

**Decision: both stay real, and both go on the allowlist of `01` D9's AST test.**
# _metrics_refresh_loop (atom/entrypoints/openai/api_server.py), in the API-server process
await asyncio.sleep(_METRICS_REFRESH_INTERVAL_SECONDS) # 5.0
await _refresh_metrics_once()

*"Leave them alone"* is not sufficient, because both sit inside processes whose clocks are
virtualized:
# AtomMetricsExporter.update (atom/entrypoints/openai/metrics.py)
self._last_refresh = clock.now(time.time) # "when was this snapshot taken", returned by read()
```

- `engine_core.py:317` is in the **clock owner's hot loop**. Redirect `time.monotonic()`
wholesale there and the metrics gate goes virtual whether or not that was intended.
- `metrics.py:408` is in the API-server process, which installs a virtual clock for arrival
stamping. Redirect `time.time()` broadly and `_last_refresh` silently reports a snapshot
*"taken"* at virtual t=180 s while wall time is t=4 s — a scraper's staleness check then
reads nonsense.
**Decision (owner ruling, 2026-10-02): the push and the refresh are timers on simulated
time, declared daemon.** They follow `01` D5's rule like every other timer the CA can
reach. The push reads the engine LP's clock through `atom.utils.clock.now`. The refresh
sleeps on the frontend's event loop, whose `time()` is the frontend LP's clock (`01` D5.1),
and `_last_refresh` stamps the simulated instant of the snapshot. As daemon deadlines they
fire as usual but never keep a run alive: the run finishes when no essential work is left
(`01` D3, #533). The frontend names its daemon timers in `DAEMON_TIMERS`
(`atom/utils/compass_loop.py`).

The allowlist is the mechanism; these are two of its entries; the AST test is what stops a
later refactor from quietly virtualizing them.
The observer is in the simulation too: the traffic LP scrapes `/metrics` every
`scrape_interval` of simulated time, a daemon timer of its own (`01` D9, item 13), so the
scrape, the refresh, the push and the stamp share one clock.

### This corrects `01` D5
### This revises the first version of this decision

That document lists `METRICS_PUSH_INTERVAL_S` among the things that *"become a virtual
timer, not disabled — metrics should be on the virtual timeline"*. **That is wrong.** A
virtual cadence is harmful in both directions:
The first version kept both clock reads on the real clock. Its reasons do not survive the
ruling:

| virtual clock | 5 virtual s becomes | consequence |
|---|---|---|
| 100× fast (idle-skipping) | ~50 ms wall | **100× more pushes** over the ZMQ output socket and through `output_queue.get()` — real load, perturbing the measurement (D77 of the previous draft, now D75) |
| 0.3× (saturated) | ~16.7 s wall | metrics stale past the API server's own 5 s refresh, which re-reads the same values |
| reason it gave | why it no longer holds |
|---|---|
| at 100× (idle-skipping) a 5 s virtual push is ~50 ms wall: **100× more pushes**, real load perturbing the measurement | emission never advances the virtual clock (D75), so the extra pushes cost wall time and change no simulated result |
| at 0.3× (saturated) a 5 s virtual push is ~16.7 s wall, staler than the API server's real 5 s refresh | the refresh is on simulated time as well, so push and refresh keep their ratio at every speed |
| a virtual `_last_refresh` read against wall time makes a scraper's staleness check nonsense | the scraper reads simulated time too (above) |
| a redirected clock read in the engine's hot loop virtualizes the push whether or not that was intended | it is intended; the clock-source lint (`01` D1, detector (3)) treats both reads as substituted |

A wall cadence avoids both. The simulated *timeline* comes from D74's sampling, not from the
push cadence.
The simulated *timeline* still comes from D74's per-step sampling, not from the push
cadence.

---

Expand Down Expand Up @@ -227,8 +230,8 @@ not an approximation of it**.

The prior 27B cc-traces run was **106 prefill + 4,346 decode = ~4,450 steps over 267 s of
virtual time**. At ~40 series that is ~178k data points for a whole run — trivial to buffer
and to write. The transport still drains on the wall-clock cadence of D72, so the ZMQ
message **count** is unchanged; only the payload gets wider.
and to write. The transport still drains on D72's push cadence, so the ZMQ message
**count** is unchanged; only the payload gets wider.

```
EngineCore (owns the virtual clock)
Expand All @@ -237,14 +240,14 @@ message **count** is unchanged; only the payload gets wider.
| snapshot = collect_metrics() <- cheap |
| buffer.append((virtual_ts, snapshot)) |
| |
| wall timer, every 5 REAL seconds (D72) |
| daemon timer, every 5 SIMULATED s (D72) |
| push_metrics(buffer.drain()) <- costly |
+----------------------+-------------------------+
| same ZMQ, same message COUNT, bigger payload
v
+---------------------------------------------+
| API server |
| /metrics -> current values, wall scrape | <- liveness, unchanged
| /metrics -> values, simulated-time scrape | <- the traffic LP's daemon scrape (D72)
| run end -> OpenMetrics text w/ virtual ts| <- the simulated timeline
+---------------------+-----------------------+
v
Expand Down Expand Up @@ -359,7 +362,7 @@ attach to where the number came from, not to how it is exposed.
| **E — event tally** | a count incremented on an occurrence: requests finished, preemptions, tokens, histogram `_bucket` and `_count` | **valid by construction.** Monotone; sample per step. |
| **D — duration** | a clock delta: TTFT, TPOT, step time, queue wait, histogram `_sum` over durations | **must be a simulated duration.** Audit the observation argument (D73). |
| **R — rate** | a count divided by elapsed time | **do not export.** Export the underlying counter and let PromQL `rate()` compute it over virtual timestamps. If a rate must be exported, its denominator is virtual elapsed. |
| **T — timestamp** | a wall-clock instant exported as a value: `process_start_time_seconds`, `_created`, `_last_refresh`, exemplar timestamps | **must declare its clock.** Usually real, because it describes the *process*; virtual if it describes the *run*. Never left implicit. |
| **T — timestamp** | an instant exported as a value: `process_start_time_seconds`, `_created`, exemplar timestamps; `_last_refresh`, which stamps simulated time (D72) | **must declare its clock.** Usually real, because it describes the *process*; virtual if it describes the *run*. Never left implicit. |
| **X — external** | measured outside the engine: GPU telemetry, host stats, the `server_metrics/` scraper | **invalid under simulation.** Refuse. |

ATOM's current twenty metrics are all **S** or **E**, which is why D72's "valid by
Expand Down Expand Up @@ -440,9 +443,9 @@ infrastructure and add nothing" should mean in practice.
| # | Decision | Date |
|---|---|---|
| D71 | Scope is ATOM's Prometheus engine metrics only. Harness metrics belong to `06`; simulator observability is run-artifact fields, not a metrics subsystem. | 2026-09-19 |
| D72 | Counters and gauges are valid by construction. **Both metrics clock reads stay on the real clock** and go on `01` D9's allowlist. This corrects `01` D5, which made the push cadence virtual. | 2026-09-19 |
| D72 | Counters and gauges are valid by construction. **The metrics push and refresh are timers on simulated time, declared daemon**, so they never keep a run alive (`01` D3, D5); `_last_refresh` stamps simulated time. | 2026-09-19; revised 2026-10-02 |
| D73 | Histograms observe durations, so the clock audit extends from clock *reads* to observation *arguments*. Ask for classic histograms and for instrumentation that is handed a duration rather than computing one inline. | 2026-09-19 |
| D74 | **Sample once per engine step** — virtual time is discrete-event, so state changes only at step boundaries and the per-step series is the ground truth. Transport on the wall clock; write OpenMetrics with virtual timestamps; backfill into TSDB. Treat the real run identically, for overlay. | 2026-09-19 |
| D74 | **Sample once per engine step** — virtual time is discrete-event, so state changes only at step boundaries and the per-step series is the ground truth. Transport on D72's push cadence; write OpenMetrics with virtual timestamps; backfill into TSDB. Treat the real run identically, for overlay. | 2026-09-19 |
| D75 | Four classes of invalid metric, refused rather than reported. Emission never advances the virtual clock, and that is asserted. | 2026-09-19 |
| D76 | The gauges supply `08`'s counting invariants for free; a histogram's `_count` is one too, and `_sum/_count` is a mean comparable at the noise floor. | 2026-09-19 |
| D77 | Classify every metric by the **provenance of its value** (state, event tally, duration, rate, timestamp, external), not by its Prometheus type. Declare the class at construction; an unclassified metric refuses to export under simulation. The backfill writer uses the OpenMetrics serializer and the `timestamp=` argument that `add_metric` already provides. | 2026-09-19 |
Expand Down
2 changes: 1 addition & 1 deletion atom/compass/design/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -528,7 +528,7 @@ The documents use these precisely; a reader will bounce off without them.
|---|---|---|
| [`06`](06_workload_harness_contract.md) | Workload Harness Contract | A three-part contract, not a bespoke client. agentx-harness reused with **zero edits** via an out-of-tree plugin. One namespaced additive field each direction, audited for minimality. Timeline piggybacked on `kv_transfer_params` so Atomesh needs no change. Tokenizer cost is a queue, not a constant. |
| [`08`](08_validation_protocol.md) | Validation Protocol | ATOM's own test suite as the first validation layer, in two tiers: a CPU tier over every test file outside `tests/plugin/` and `cpu_gate_exclude.txt`, driver-free **as a batch** and held to green, and a GPU superset judged as a **delta** against **4779 / 5** (`fe9ea043c`, torch 2.10.0+rocm7.2.4, ROCm 7.2.4, AITER v0.1.21.dev0-49-gf4e7c7509, all five failing node-ids on file). Three separable results, never one number. **The real-vs-real spread is the tolerance.** A metric is admissible only if stable *and* sensitive. |
| [`11`](11_metrics_support.md) | Engine Metrics under Virtual Time | ATOM's Prometheus exporter under a virtual clock. Metrics are classified by the **provenance of their value**, not their type. Sample once per engine step — virtual time is discrete-event. Both metrics clock reads stay real. |
| [`11`](11_metrics_support.md) | Engine Metrics under Virtual Time | ATOM's Prometheus exporter under a virtual clock. Metrics are classified by the **provenance of their value**, not their type. Sample once per engine step — virtual time is discrete-event. The metrics push and refresh are daemon timers on simulated time. |

### Part V — Cross-cutting

Expand Down