diff --git a/atom/compass/audit/sync_sites.json b/atom/compass/audit/sync_sites.json index ad79687198..aa2af26f87 100644 --- a/atom/compass/audit/sync_sites.json +++ b/atom/compass/audit/sync_sites.json @@ -584,7 +584,7 @@ "peer": "none", "why": "the metrics scrape cadence is deliberately left on the real clock", "mechanism": "K7", - "mechanism_why": "the metrics refresh cadence runs on virtual time once the frontend's event loop does, with no change at the call" + "mechanism_why": "the metrics refresh cadence is a daemon timer on simulated time once the frontend's event loop runs on the LP clock, with no change at the call" }, { "id": "atom/entrypoints/openai/api_server.py::chat_completions::loop.run_in_executor( None, _prepare_multimodal_inputs, messages, merged_kwargs, request.tools, )::executor_handoff#0", @@ -2955,7 +2955,7 @@ "peer": "none", "why": "the metrics push cadence is deliberately left on the real clock; the reported timeline comes from sampling, not from this", "mechanism": "K7", - "mechanism_why": "the metrics push cadence becomes a virtual timer, folded into the loop's idle point" + "mechanism_why": "the metrics push cadence becomes a daemon timer on simulated time, folded into the loop's idle point" }, { "file": "atom/model_engine/engine_core.py", @@ -2965,7 +2965,7 @@ "peer": "none", "why": "the same cadence in the data-parallel loop, deliberately left on the real clock", "mechanism": "K7", - "mechanism_why": "the metrics push cadence becomes a virtual timer, folded into the loop's idle point" + "mechanism_why": "the metrics push cadence becomes a daemon timer on simulated time, folded into the loop's idle point" }, { "file": "atom/model_engine/engine_core.py", diff --git a/atom/compass/design/01_execution_and_time_model.md b/atom/compass/design/01_execution_and_time_model.md index 51fb564d75..1c94bdb4df 100644 --- a/atom/compass/design/01_execution_and_time_model.md +++ b/atom/compass/design/01_execution_and_time_model.md @@ -1316,9 +1316,13 @@ The pacing timers are local next events the clock owner hands to engine's metrics push (`EngineCore.busy_loop` and `DPEngineCoreProc.busy_loop`), and the API server's refresh loop (`_metrics_refresh_loop` in `api_server.py`). **Metric cadence is virtual time again**, - which revises `11` D72's decision to keep it on the real clock: an observer in the - traffic LP scrapes `/metrics` every `scrape_interval` of simulated time. `11` D72 - carries the observer and the reasons. + as revised `11` D72 also decides (its first version kept it on the real clock): an + observer in the traffic LP scrapes `/metrics` every `scrape_interval` of simulated + time. `11` D72 carries the observer and the reasons. +- uvicorn's server tick, a 0.1 s `asyncio.sleep` in `Server.main_loop`, and its + keep-alive timeout are daemon deadlines too. They are uvicorn's code, not ATOM's, so the + sync inventory (`atom/compass/audit/sync_sites.json`) has no row for them; the frontend + declares them with the refresh loop in `DAEMON_TIMERS` (`atom/utils/compass_loop.py`). ### Must read the LP clock — these three change scheduling (K2) @@ -1478,9 +1482,9 @@ the producer's `hash_block_size` against its own and falls back to a full transf (`mooncake_connector.py::MooncakeConnectorScheduler.update_state_after_alloc`), so a blob carrying only the thirteen can never take the incremental path. -`tests/compass/test_kv_blob_doc_table.py` re-derives both sets from the connectors and -fails naming the field that differs, so this table cannot drift from the source the -way its twelve-field predecessor did. +`tests/compass/test_kv_blob_site.py` holds the one-site claim above. Nothing re-checks +the field sets: they describe the connectors' code as read at `92f1fdafe`, and where the +two differ the code is right. ### Pros @@ -1548,9 +1552,12 @@ router-side latency injection is ever wanted, this is where it goes. ### Open issues -- The two hardcoded Rust timeouts (`worker.rs::DEFAULT_WORKER_HTTP_TIMEOUT_SECS`, `worker_manager.rs::DEFAULT_WORKER_REQUEST_TIMEOUT_SECS`) require - touching the Rust crate, which means `ATOM_MESH_BUILD=1` and a Rust toolchain in the - loop. Confirm the container has one. +- `worker.rs::DEFAULT_WORKER_HTTP_TIMEOUT_SECS` is the one Atomesh timeout still compiled + in; changing it means touching the Rust crate, with `ATOM_MESH_BUILD=1` and a Rust + toolchain in the loop. D5 finds a simulated launch never reaches it, so this stays open + only if that launch ever keeps the health check on. + `worker_manager.rs::DEFAULT_WORKER_REQUEST_TIMEOUT_SECS` is only the default of + `--worker-request-timeout-secs` (#478), which D5 sets large. - Mesh-only mode has not been exercised by this project. Confirm it serves the ATOM relay path correctly against two ATOM servers before depending on it. diff --git a/atom/compass/design/11_metrics_support.md b/atom/compass/design/11_metrics_support.md index 451b6d44ce..220e22817b 100644 --- a/atom/compass/design/11_metrics_support.md +++ b/atom/compass/design/11_metrics_support.md @@ -45,7 +45,7 @@ influence how they land rather than retrofit them. See D73. --- -## D72. Counters and gauges are valid by construction; their clocks stay real +## D72. Counters and gauges are valid by construction; their cadence is simulated time ### Why they are valid @@ -54,47 +54,50 @@ influence how they land rather than retrofit them. See D73. run unmodified under simulation** (`03` D13). Nothing there is a timing read, and there are no rate denominators to correct. -### The two clock reads, and why both stay on the real clock +### The two timers and the stamp, all on the LP clock ```python -# engine_core.py:317-320 <- inside EngineCore.busy_loop, the process that OWNS the clock -now = time.monotonic() +# EngineCore.busy_loop and DPEngineCoreProc.busy_loop (atom/model_engine/engine_core.py) +now = clock.now(time.monotonic) if now >= next_metrics_push: - next_metrics_push = now + METRICS_PUSH_INTERVAL_S # 5.0, engine_core.py:50 + next_metrics_push = now + METRICS_PUSH_INTERVAL_S # 5.0 self.utility_handler.push_metrics() -# metrics.py:408 <- in the API-server process -self._last_refresh = time.time() # "when was this snapshot taken", returned by read() -``` - -**Decision: both stay real, and both go on the allowlist of `01` D9's AST test.** +# _metrics_refresh_loop (atom/entrypoints/openai/api_server.py), in the API-server process +await asyncio.sleep(_METRICS_REFRESH_INTERVAL_SECONDS) # 5.0 +await _refresh_metrics_once() -*"Leave them alone"* is not sufficient, because both sit inside processes whose clocks are -virtualized: +# AtomMetricsExporter.update (atom/entrypoints/openai/metrics.py) +self._last_refresh = clock.now(time.time) # "when was this snapshot taken", returned by read() +``` -- `engine_core.py:317` is in the **clock owner's hot loop**. Redirect `time.monotonic()` - wholesale there and the metrics gate goes virtual whether or not that was intended. -- `metrics.py:408` is in the API-server process, which installs a virtual clock for arrival - stamping. Redirect `time.time()` broadly and `_last_refresh` silently reports a snapshot - *"taken"* at virtual t=180 s while wall time is t=4 s — a scraper's staleness check then - reads nonsense. +**Decision (owner ruling, 2026-10-02): the push and the refresh are timers on simulated +time, declared daemon.** They follow `01` D5's rule like every other timer the CA can +reach. The push reads the engine LP's clock through `atom.utils.clock.now`. The refresh +sleeps on the frontend's event loop, whose `time()` is the frontend LP's clock (`01` D5.1), +and `_last_refresh` stamps the simulated instant of the snapshot. As daemon deadlines they +fire as usual but never keep a run alive: the run finishes when no essential work is left +(`01` D3, #533). The frontend names its daemon timers in `DAEMON_TIMERS` +(`atom/utils/compass_loop.py`). -The allowlist is the mechanism; these are two of its entries; the AST test is what stops a -later refactor from quietly virtualizing them. +The observer is in the simulation too: the traffic LP scrapes `/metrics` every +`scrape_interval` of simulated time, a daemon timer of its own (`01` D9, item 13), so the +scrape, the refresh, the push and the stamp share one clock. -### This corrects `01` D5 +### This revises the first version of this decision -That document lists `METRICS_PUSH_INTERVAL_S` among the things that *"become a virtual -timer, not disabled — metrics should be on the virtual timeline"*. **That is wrong.** A -virtual cadence is harmful in both directions: +The first version kept both clock reads on the real clock. Its reasons do not survive the +ruling: -| virtual clock | 5 virtual s becomes | consequence | -|---|---|---| -| 100× fast (idle-skipping) | ~50 ms wall | **100× more pushes** over the ZMQ output socket and through `output_queue.get()` — real load, perturbing the measurement (D77 of the previous draft, now D75) | -| 0.3× (saturated) | ~16.7 s wall | metrics stale past the API server's own 5 s refresh, which re-reads the same values | +| reason it gave | why it no longer holds | +|---|---| +| at 100× (idle-skipping) a 5 s virtual push is ~50 ms wall: **100× more pushes**, real load perturbing the measurement | emission never advances the virtual clock (D75), so the extra pushes cost wall time and change no simulated result | +| at 0.3× (saturated) a 5 s virtual push is ~16.7 s wall, staler than the API server's real 5 s refresh | the refresh is on simulated time as well, so push and refresh keep their ratio at every speed | +| a virtual `_last_refresh` read against wall time makes a scraper's staleness check nonsense | the scraper reads simulated time too (above) | +| a redirected clock read in the engine's hot loop virtualizes the push whether or not that was intended | it is intended; the clock-source lint (`01` D1, detector (3)) treats both reads as substituted | -A wall cadence avoids both. The simulated *timeline* comes from D74's sampling, not from the -push cadence. +The simulated *timeline* still comes from D74's per-step sampling, not from the push +cadence. --- @@ -227,8 +230,8 @@ not an approximation of it**. The prior 27B cc-traces run was **106 prefill + 4,346 decode = ~4,450 steps over 267 s of virtual time**. At ~40 series that is ~178k data points for a whole run — trivial to buffer -and to write. The transport still drains on the wall-clock cadence of D72, so the ZMQ -message **count** is unchanged; only the payload gets wider. +and to write. The transport still drains on D72's push cadence, so the ZMQ message +**count** is unchanged; only the payload gets wider. ``` EngineCore (owns the virtual clock) @@ -237,14 +240,14 @@ message **count** is unchanged; only the payload gets wider. | snapshot = collect_metrics() <- cheap | | buffer.append((virtual_ts, snapshot)) | | | - | wall timer, every 5 REAL seconds (D72) | + | daemon timer, every 5 SIMULATED s (D72) | | push_metrics(buffer.drain()) <- costly | +----------------------+-------------------------+ | same ZMQ, same message COUNT, bigger payload v +---------------------------------------------+ | API server | - | /metrics -> current values, wall scrape | <- liveness, unchanged + | /metrics -> values, simulated-time scrape | <- the traffic LP's daemon scrape (D72) | run end -> OpenMetrics text w/ virtual ts| <- the simulated timeline +---------------------+-----------------------+ v @@ -359,7 +362,7 @@ attach to where the number came from, not to how it is exposed. | **E — event tally** | a count incremented on an occurrence: requests finished, preemptions, tokens, histogram `_bucket` and `_count` | **valid by construction.** Monotone; sample per step. | | **D — duration** | a clock delta: TTFT, TPOT, step time, queue wait, histogram `_sum` over durations | **must be a simulated duration.** Audit the observation argument (D73). | | **R — rate** | a count divided by elapsed time | **do not export.** Export the underlying counter and let PromQL `rate()` compute it over virtual timestamps. If a rate must be exported, its denominator is virtual elapsed. | -| **T — timestamp** | a wall-clock instant exported as a value: `process_start_time_seconds`, `_created`, `_last_refresh`, exemplar timestamps | **must declare its clock.** Usually real, because it describes the *process*; virtual if it describes the *run*. Never left implicit. | +| **T — timestamp** | an instant exported as a value: `process_start_time_seconds`, `_created`, exemplar timestamps; `_last_refresh`, which stamps simulated time (D72) | **must declare its clock.** Usually real, because it describes the *process*; virtual if it describes the *run*. Never left implicit. | | **X — external** | measured outside the engine: GPU telemetry, host stats, the `server_metrics/` scraper | **invalid under simulation.** Refuse. | ATOM's current twenty metrics are all **S** or **E**, which is why D72's "valid by @@ -440,9 +443,9 @@ infrastructure and add nothing" should mean in practice. | # | Decision | Date | |---|---|---| | D71 | Scope is ATOM's Prometheus engine metrics only. Harness metrics belong to `06`; simulator observability is run-artifact fields, not a metrics subsystem. | 2026-09-19 | -| D72 | Counters and gauges are valid by construction. **Both metrics clock reads stay on the real clock** and go on `01` D9's allowlist. This corrects `01` D5, which made the push cadence virtual. | 2026-09-19 | +| D72 | Counters and gauges are valid by construction. **The metrics push and refresh are timers on simulated time, declared daemon**, so they never keep a run alive (`01` D3, D5); `_last_refresh` stamps simulated time. | 2026-09-19; revised 2026-10-02 | | D73 | Histograms observe durations, so the clock audit extends from clock *reads* to observation *arguments*. Ask for classic histograms and for instrumentation that is handed a duration rather than computing one inline. | 2026-09-19 | -| D74 | **Sample once per engine step** — virtual time is discrete-event, so state changes only at step boundaries and the per-step series is the ground truth. Transport on the wall clock; write OpenMetrics with virtual timestamps; backfill into TSDB. Treat the real run identically, for overlay. | 2026-09-19 | +| D74 | **Sample once per engine step** — virtual time is discrete-event, so state changes only at step boundaries and the per-step series is the ground truth. Transport on D72's push cadence; write OpenMetrics with virtual timestamps; backfill into TSDB. Treat the real run identically, for overlay. | 2026-09-19 | | D75 | Four classes of invalid metric, refused rather than reported. Emission never advances the virtual clock, and that is asserted. | 2026-09-19 | | D76 | The gauges supply `08`'s counting invariants for free; a histogram's `_count` is one too, and `_sum/_count` is a mean comparable at the noise floor. | 2026-09-19 | | D77 | Classify every metric by the **provenance of its value** (state, event tally, duration, rate, timestamp, external), not by its Prometheus type. Declare the class at construction; an unclassified metric refuses to export under simulation. The backfill writer uses the OpenMetrics serializer and the `timestamp=` argument that `add_metric` already provides. | 2026-09-19 | diff --git a/atom/compass/design/README.md b/atom/compass/design/README.md index d42b108456..83271e2fa4 100644 --- a/atom/compass/design/README.md +++ b/atom/compass/design/README.md @@ -528,7 +528,7 @@ The documents use these precisely; a reader will bounce off without them. |---|---|---| | [`06`](06_workload_harness_contract.md) | Workload Harness Contract | A three-part contract, not a bespoke client. agentx-harness reused with **zero edits** via an out-of-tree plugin. One namespaced additive field each direction, audited for minimality. Timeline piggybacked on `kv_transfer_params` so Atomesh needs no change. Tokenizer cost is a queue, not a constant. | | [`08`](08_validation_protocol.md) | Validation Protocol | ATOM's own test suite as the first validation layer, in two tiers: a CPU tier over every test file outside `tests/plugin/` and `cpu_gate_exclude.txt`, driver-free **as a batch** and held to green, and a GPU superset judged as a **delta** against **4779 / 5** (`fe9ea043c`, torch 2.10.0+rocm7.2.4, ROCm 7.2.4, AITER v0.1.21.dev0-49-gf4e7c7509, all five failing node-ids on file). Three separable results, never one number. **The real-vs-real spread is the tolerance.** A metric is admissible only if stable *and* sensitive. | -| [`11`](11_metrics_support.md) | Engine Metrics under Virtual Time | ATOM's Prometheus exporter under a virtual clock. Metrics are classified by the **provenance of their value**, not their type. Sample once per engine step — virtual time is discrete-event. Both metrics clock reads stay real. | +| [`11`](11_metrics_support.md) | Engine Metrics under Virtual Time | ATOM's Prometheus exporter under a virtual clock. Metrics are classified by the **provenance of their value**, not their type. Sample once per engine step — virtual time is discrete-event. The metrics push and refresh are daemon timers on simulated time. | ### Part V — Cross-cutting