Skip to content

feat(OMN-19981): the six local dashboard pages read served exposures, live, with typed empty states - #349

Merged
Patel230 merged 7 commits into
lakshman/OMN-19981-local-pages-overview-runsfrom
lakshman/OMN-19981-overview-savings
Oct 3, 2026
Merged

Patel230 merged 7 commits into
lakshman/OMN-19981-local-pages-overview-runsfrom
lakshman/OMN-19981-overview-savings

Conversation

@Patel230

@Patel230 Patel230 commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Stacked on omnidash#346 (base is #346's branch; this retargets to dev when #346 merges). On the lab, OMN-19981's Overview rendered nothing: it was bound to onex.snapshot.projection.baselines.roi.v1, which the catalogue marks degraded (not_yet_bus_backed) and whose table holds 0 rows, so the page said "Projection exposure is not served". This binds Overview to the served cost.savings-overview.v1 (the source named in the ticket's savings handoff, comment aac9032d). It also takes the ticket's strangler step for the two topics the local pages now read from served exposures.

What changed

  • Overview (overview.contracts.yaml, overview.page.yaml): three metric cards on cost.savings-overview.v1:

    • Spend (total_cost_usd);
    • Savings (total_savings_usd, with total_baseline_cost_usd required);
    • Tokens (tokens_total).

    local_token_pct is not read: the view reports it as a literal 0 it does not measure (OMN-20320).

  • Renderer (LocalDashboardPage.tsx):

    • MetricCard renders the field its widget names. It previously always rendered roi_percent.
    • A missing baseline renders "Baseline unresolved"; any other missing field renders "Not measured"; neither ever renders 0 (ticket AC4).
    • The page document resets on a page switch. Overview opened from Runs used to show the header "Runs".
  • Ticket AC2 strangler ("Reader folds are removed page by page"):

    • Both server readers drop their hand-written SQL folds for cost.savings-overview.v1 and delegation.savings.v1. fix(OMN-19981): show delegation savings and focus workbench evidence #346's Runs already reads the second one, so AC2 was false for Runs too; the existing test checked Overview against one reader only.
    • Each reader keeps its other 15 folds.
    • The helpers only those folds used go with them: buildCostSavingsOverviewResult, plus the readers' mergeDelegationSessions wrappers and sessionKey. The shared mergeDelegationSessions stays, since projection-utils.ts uses it.
  • Effect: in sqlite and postgres modes the two topics now answer [], the readers' unknown-topic path. The default http mode, which reads served exposures (ruling D1), is unchanged.

  • Tests (pages.test.ts):

    • The served list comes from a captured lab catalogue (served-catalogue.lab.json, GET /projections, 69 exposures, 21 ok) instead of a hand-written list. The hand-written list called baselines.roi.v1 served.
    • AC2 is checked against both readers for every local page.
    • New LocalDashboardPage.overview.test.tsx covers the cards and the header. Each reader's 3 or 5 tests of the removed folds are replaced by one test that the topics answer [], with a staying fold as the control.

Second commit, 82abf05: the dashboard requirements this page must meet. These come from beta/requirements/2026-09-28-dashboard-requirements.md (sections 2.1 and 5, linked from the ticket) and Jonah's handoffs on OMN-19981:

  • SV-3 ("token in and token out are separate numbers everywhere, never a single tokens total"): the single Tokens card is removed. cost.savings-overview.v1 serves no in/out split and the browser may not compute one, so a card bound to no exposure reads "Not served yet: waits on metering-summary.v1".
  • OV-3 / SV-2 ("the savings tile shows the baseline model string and pricing manifest version under the number"): the Savings card's second binding reads baseline_model and pricing_manifest_version from delegation.savings.v1's served row. A null baseline reads "Baseline unresolved".
  • OMN-20008 AC4 (a measured count beside the total): a Measured runs card (measured_run_count), with zero_token_run_count beside it.
  • Jonah's handoff: the placeholder model_name delegate-skill and an empty model render "unknown" in Runs.
  • local-page-loader.ts: the page empty-state rules look only at each component's first binding. The caption binding had switched Overview into the Runs session-row rules and silenced BASELINE_UNRESOLVED until this was fixed.

Third commit, a821c9a: last run, recent runs, and run status, cause and filters (requirements OV-4, OV-5, RU-1, RU-4):

  • Overview, Last run card (OV-4): built from the newest delegation.decisions.v1 row by written_at. It shows:
    • status: passed, or failed with the typed cause from quality_gate_detail;
    • run id, model (a placeholder reads "unknown"), duration, tokens in and tokens out (separate, SV-3), task type, route tier (cost_tier_name) and age;
    • cost, from the delegation.savings.v1 session with the same id (session_id = correlation_id, handoff aac9032d);
    • backend reads "Not served (OMN-20162)". Nothing is blank or 0 when unmeasured.
  • Overview, Recent runs (OV-5): the ten newest decisions, newest first, with the run columns. A session with no baseline and 0 savings reads "Baseline unresolved", never 0. Its cells wrap, so every panel fits at 1440x900 (ticket AC4).
  • Runs (RU-1):
    • each session looks up its decision for Status, Cause, Route tier and Quality score ("Not recorded" when it has none);
    • status, model and window (today, UTC) filters apply to the served rows only.
  • Empty Runs page (RU-4): NO_RUNS_YET names the first-run command, onex delegate "ping" (local MVP plan, verify step).
  • Label: a measured saving whose baseline_model is null reads "Not recorded", not "Not measured".
  • Ticket AC2 strangler: both readers stop folding onex.snapshot.projection.delegation.decisions.v1, which the local pages now read from the served exposure. The legacy delegation name keeps its SQL.
  • Effect: in sqlite and postgres modes the full topic answers [], so the legacy Lab, Event Bus, Experiments, Delegation Evidence and routing-table views get no decisions rows in those modes. The default http mode is unchanged.

Fourth commit, 3fe6d89: the six local pages are the dashboard, read live, and Runs lists every run (requirements FR-3, the widget-state table, RU-1, FR-2 row badge):

  • Navigation (FR-3):
    • the sidebar opens with a Local group: Overview, Runs, Workflow, Usage, Credentials, API Keys, in that order;
    • the app opens on Overview instead of the editable canvas; the existing pages stay reachable below.
  • Partial pages: Workflow, Usage and API Keys carry a "partial" chip and name what they wait on:
    • Workflow: run-trace.v1, OMN-19987;
    • Usage: usage-by-model-day.v1, OMN-20006;
    • API Keys: local-identity.v1, OMN-19986, with CLOUD_NOT_LINKED for cloud keys (AK-3).
  • Credentials (CR-1 to CR-3): reads tenant-credentials.v1 and shows provider, key ref, set and revoked, never a value. With no key it shows onex secret set. There is no input field.
  • Live data: each local page re-reads its exposures every declared refresh_interval_seconds (30), keeping the last good rows on screen. Each panel shows "As of ", highlighted once stale.
  • Widget states: an unserved exposure, an HTTP error, or no answer in 5 s fails only the widgets bound to it, naming the exposure, the status and when it was last good. A failed read is never NO_RUNS_YET.
  • Runs (RU-1):
    • rows are delegation.decisions.v1, so a failed run with no savings row is still listed; the savings session with the same id supplies cost, savings and baseline;
    • 25 per page, the declared page_size;
    • a cause filter;
    • a fixture badge on rows whose served data_source is not real (FR-2, row part).

Fifth commit, e187e72: every local page reads a served exposure, AC4 has its Playwright proof, and the sqlite reader is gone (the rest of OMN-19981: AC1, AC3, AC4, and comments 3376eac0 and aac9032d):

  • AC1: all six pages now name a served exposure.
    • Usage reads usage-by-model-day.v1, showing this tenant's rows only (the exposure declares no tenant column), tokens in and out apart. With no rows it names the event it waits on (llm-call-completed, OMN-20006).
    • Workflow reads delegation.decisions.v1 and shows the newest run's recorded steps: request, routing, quality gate, terminal. The full path waits on run-trace.v1 (OMN-19987).
    • API Keys reads the tenant id from delegation.savings.v1 (AK-1). Minted-at waits on local-identity.v1 (OMN-19986). Cloud keys show CLOUD_NOT_LINKED.
  • AC3: server/sqlite-projection-reader.ts and its tests are deleted. OMNIDASH_DATA_SOURCE=sqlite now answers 503 sqlite_reader_retired, naming onex dashboard (OMN-19976).
  • AC4: tests/e2e/local-dashboard-contracts.spec.ts is rewritten for the current pages, and its screenshots are refreshed. The old spec targeted baselines.roi.v1 and failed.
  • First load: a page that mounts twice (React development mode) shares each read in flight. Before, the second copy queued past the 5 s limit: 0.9 s, then 5.8 s.
  • Savings label (aac9032d, item 5): "Modelled: the runs' tokens priced at the baseline model's list price. The baseline never ran."
  • Postgres reader (3376eac0): it reads model_cloud_baseline AS baseline_model, and sums savings once per run (newest row per session_id) in both savings totals.

How it was verified

Red first, each test seen failing before the change:

Green:

  • vitest 1823 passed, 15 skipped (198 files).
  • tsc --noEmit for both tsconfig.json and tsconfig.node.json. The root npm run check does not cover server/, so the node config was run explicitly; it caught the four orphaned helpers on the first pass.
  • eslint --max-warnings=0 clean on every changed file.

Fold removal is exact: each reader's case list before and after differs by exactly these two topics; nothing else was removed or added.

Lab first, on the lakshman lane on .201:

  • 20 real delegations (08:29Z to 08:35Z, onex delegate --locus deployed-lane in the lane's runtime, Qwen3.8-27B) each landed in delegation_events as data_source = real.
  • Runs (fix(OMN-19981): show delegation savings and focus workbench evidence #346's head): shows them, and the paged Runs session ids equal the store's 41 ids for the tenant, with 0 differences either way.
  • Overview (this head, 154bf0c): at 08:53:21Z it renders Spend $0.0069, Savings $1.2235, Tokens 56,800, equal to the API's cost.savings-overview.v1 row read at the same moment (total_cost_usd 0.00686, total_savings_usd 1.22354, tokens_total 56800).
  • Not a real-data claim: on this lane the savings figure comes from 20 seeded fixture rows. The 21 real runs carry no premium_counterfactual, so their savings are unmeasured. Why this lane prices no baseline is not investigated here.

For 82abf05:

  • Red first: 7 new cases failed before the change (2 contract, 4 renderer, 1 Runs). Then the whole suite: 1829 passed, 15 skipped (198 files). tsc --noEmit for both configs and eslint --max-warnings=0 are clean.

  • Lab, lakshman lane, 10:59:04Z: the cards equal the API's served rows read at the same moment.

    Card Rendered Served
    Spend $0.0069 total_cost_usd 0.00686
    Savings $1.2252, "Baseline unresolved · pricing manifest v1" total_savings_usd 1.225228; baseline_model null, pricing_manifest_version "1"
    Measured runs 46, "Zero-token runs: 0" measured_run_count 46, zero_token_run_count 0
    Tokens in and out "Not served yet: waits on metering-summary.v1" no in/out split served

    The served baseline is null on that lane because it runs no savings writer, so savings_estimates has no rows for real runs.

For a821c9a:

  • Red first:

    • 20 new page and contract tests failed before the change (no components, no bindings, no filters, the old label);
    • the AC2 test failed for both pages once they bound decisions;
    • 5 reader tests asserted the old fold;
    • a lakshman-lane read showed 0 savings for unpriced runs in Recent runs, and a test for that failed before the fix.
  • Green: vitest 1850 passed, 15 skipped. tsc --noEmit for both configs and eslint on every changed file are clean.

  • Mutation: limiting Recent runs to 11 rows, or forcing a failed run's cause to "none", fails 4 tests.

  • Related suites (server, src/pages, src/layout): origin/dev 48e0028 379 passed; this head 414 passed; 0 failed on both.

  • Lab, lakshman lane, API read at 12:23:23Z (46 decisions, not truncated):

    Check Rendered Served
    Last run 81971f1c…, passed, Qwen3.8-27B, 813 ms, 162 in / 52 out, cost 0, local newest decision by written_at, same values
    Runs, status = failed 3 rows: 429 quota, citation gate, 30s timeout 3 rows with quality_gate_passed false, same quality_gate_detail
    Runs, window = today (UTC) 25 rows 25 created 2026-10-02
    Runs, model = deepseek/deepseek-chat-v3 6 rows 6
    Layout at 1440x900 every panel's content fits its width (Recent runs 1098 of 1098 px, was 1883) n/a

    No console errors.

For 3fe6d89:

  • Red first: 57 new or changed tests failed before the change. Green: vitest 1900 passed, 15 skipped; tsc --noEmit for both configs and eslint clean.

  • Mutation: removing the 30 s re-read and disabling the fixture badge fails 3 tests.

  • Related suites (server, src/pages, src/layout, src/components/frame, App): origin/dev 48e0028 403 passed; this head 488 passed; 0 failed on both.

  • Lab, lakshman lane, API read at 13:17:04Z:

    Check Rendered Served
    Opening page Overview; Local group first; partial chip on the last four n/a
    Runs pages "Runs 1–25 of 46", then "26–46 of 46"; 46 unique ids; newest first 81971f1c… 46 decisions; newest by written_at 81971f1c…
    Fixture badge 20 rows 20 rows data_source = fixture, 26 real
    Cause = "provider timeout after 30s" 1 row, 48f88588… the same row
    Live re-read one new decisions read within 33 s; the open Runs page is kept n/a
    Credentials "No provider key set", onex secret set 0 rows for the tenant

    No console errors; every Overview panel fits at 1440x900.

For e187e72:

  • Red first: 14 AC1 tests, the shared-read and StrictMode tests, the savings label, 3 postgres tests, 4 AC3 tests, and both old Playwright tests failed before the change.

  • Green: vitest 1896 passed, 15 skipped; Playwright 3 passed (Chrome); tsc --noEmit for both configs and eslint clean.

  • Mutation: showing 0 for the unmeasured saving fails the Overview Playwright test.

  • Related suites: origin/dev 48e0028 403 passed; this head 484 passed; 0 failed on both.

  • Real data, lakshman lane, 14:58 to 14:59Z: the six pages were read on localhost and compared with the API, 685 values in all:

    • the Overview cards;
    • the Last run card;
    • the 10 Recent runs;
    • all 46 Runs rows over both pages, in order, including 20 fixture badges equal to the API's 20 data_source = fixture rows;
    • the Workflow run, the API Keys tenant, and the Usage and Credentials empty states (0 rows served).

    Every value matches. First load reads each exposure once, the slowest in 847 ms, with no alerts.

  • Lane Postgres, read-only, 14:42:44Z: the old savings.v1 select fails with column "baseline_model" does not exist; the new one runs. That lane has 0 savings_estimates rows, so the per-run sum is shown to run, not shown to dedupe real rows.

Lab, 20 real delegations (the ticket's How-to), lakshman lane, 15:19 to 15:27Z:

  • Command: each run inside the lane's runtime, as onex delegate "<prompt> (OMN-19981 lab check NN)" --task-type <type> --bus kafka --kafka-bootstrap redpanda:9092 --locus deployed-lane, with mixed task types.
  • Results: 16 succeeded on Qwen3.8-27B (local-heavy-reasoning, per run 01's receipt.json). 4 failed, with every local rung refusing, erroring or climbing.
  • Recorded: all 20 reached delegation.decisions.v1 and delegation.savings.v1 with data_source = real; the lane went from 46 to 66 runs.
  • Dashboard, 15:28:30Z, picked up by the 30 s re-read with no reload: 826 of 826 rendered values match the API across all 66 runs and the six pages.
    • Savings moved $1.2252 → $1.2536, exactly the $0.028398 the 16 successful runs saved.
    • Measured runs moved 46 → 66.
    • Last run is lab check 20 (11e8fb56, failed). Runs reads "Runs 51–66 of 66" on its last page.
    • Status = failed with window = today shows exactly the 4 lab failures, each with savings "Baseline unresolved", never 0.

Failure paths

n/a: this change writes no verdict, receipt or report. It changes what two pages read and render, and removes two reader folds.

Open defects

  • Failed runs carry no usable cause. The 4 failed lab runs reach delegation.decisions.v1 with quality_gate_detail = [REDACTED - potentially sensitive data]; the writer redacts the detail, so Runs can only show that. A typed cause (RU-2) needs terminal_failure_cause served, OMN-19448.

  • Fixture rows on the lakshman lane: 20 of its 46 decisions are fixture rows, including all 3 failed runs. The failed-run rows in the a821c9a table above are fixture data. Runs now badges them; the Overview totals still include them until metering-summary.v1 (OMN-19977) can exclude them (FR-2 totals).

  • On the lakshman lane the real delegation runs carry no priced baseline (premium_counterfactual empty on all 21), so Overview's savings figure there is fixture-only. On the shared dev lane about 64% of runs are priced (handoff aac9032d).

  • The pages render BASELINE_UNRESOLVED for a null baseline; per-run unresolved rendering in the Runs table is fix(OMN-19981): show delegation savings and focus workbench evidence #346's.

  • The Runs table (18 columns) still scrolls sideways inside its panel. AC4 names the default page, Overview, which fits.

Local gates

  • No gate was bypassed — no --no-verify, no SKIP=, no --no-gpg-sign, no core.hooksPath override. Every pre-commit and pre-push hook ran.
  • Any hook that failed was fixed at the input, not worked around.

Not in this change

  • From the same requirements, not in this PR: fixture exclusion from totals (FR-2; needs metering-summary.v1), the environment strip, services grid, alerts and errors (OV-1, OV-2, OV-6, OV-7), the backend column and filter (not served, OMN-20162), RU-2, run detail (RU-3), OMN-20225, OMN-20009, the rest of OMN-20008, and metering-summary.v1 (OMN-19977).

  • Real data behind Workflow's full path, Usage rows, Credentials events and the local identity (OMN-19987, OMN-19978, OMN-19985, OMN-19986).

  • Why the lakshman lane prices no counterfactual.

  • local_token_pct (OMN-20320).

Overlap-Reviewed: #346 — this PR is stacked on #346 (its base is #346's branch). Both touch src/pages/LocalDashboardPage.tsx and src/pages/local/pages.test.ts: #346 adds the Runs table and Runs tests; this adds the Overview metric cards, the page reset and the catalogue and two-reader checks, on top of #346's versions of both files. Each is needed; this one cannot merge before #346.

Sixth commit: canonical navigation and refresh

Commit 5062a111fe382e5d0d00721b51d37c0f79f6d7b6 keeps this work in PR #349 as requested:

  • Gives all six local pages canonical browser URLs.
  • Makes sidebar links, breadcrumbs, reload, Back, Forward, redirects, and unknown-page recovery follow the URL.
  • Persists Runs status, cause, model, window, and pagination state in the URL.
  • Makes Refresh reload the runtime exposures used by the visible local page.
  • Improves light and dark navigation selection, focus, controls, row hover, sidebar scrolling, and compact-header actions.
  • Documents each page and its Projection API exposures.

The commit is a content-preserving rebuild of the previously verified navigation tree. Fifteen Chrome cases, 216 related tests, 1,899 full-suite tests with 15 existing skips, both TypeScript configurations, lint, and six pre-change mutations were recorded for the identical product paths. git diff --exit-code 5062a11 ade66de -- . ':(exclude).github' ':(exclude)tests/ci/test_sibling_reads_pinned.py' exits 0, and git verify-commit 5062a11 reports a valid ED25519 signature.

The full reviewed change is 7,084 lines, below the operator-set 10,000-line gate; no size exemptions are used.

Overlap-Reviewed: #349 — this is an update to the existing #349 head, prepared from the clean reconstruction worktree; the apparent overlap is the PR with itself, not a competing change.

Overlap-Reviewed: #337 — #337 adds a parked execution-graph route to src/App.tsx; this PR adds canonical local-page routing in the same file. The route families are separate and both changes remain necessary.

…ved savings overview

Overview was bound to onex.snapshot.projection.baselines.roi.v1, which the
lab catalogue marks degraded (not_yet_bus_backed) and whose table holds no
rows, so on the lab the page said "Projection exposure is not served".

- overview.contracts.yaml / overview.page.yaml: three metric cards on
  cost.savings-overview.v1 (spend total_cost_usd, savings
  total_savings_usd with total_baseline_cost_usd required, tokens
  tokens_total). local_token_pct is not read; the view does not measure it.
- LocalDashboardPage.tsx: MetricCard renders the field its widget names
  (it always rendered roi_percent). A missing baseline renders "Baseline
  unresolved" and any other missing field "Not measured", never 0. The
  page document resets on a page switch, so Overview opened from Runs no
  longer shows the header "Runs".
- Ticket AC2 strangler: both server readers drop their hand-written SQL
  folds for cost.savings-overview.v1 and delegation.savings.v1, which the
  two local pages now read from served exposures. Each reader keeps its
  other 15 folds; the now-unused buildCostSavingsOverviewResult,
  mergeDelegationSessions wrappers and sessionKey helpers go with them.
  In sqlite and postgres modes the two topics answer [] (the readers'
  unknown-topic path); the default http mode is unchanged.
- pages.test.ts reads the served list from a captured lab catalogue
  (served-catalogue.lab.json) instead of a hand-written one, and checks
  both readers for every local page (it checked one reader for Overview).

Red before: Overview bound a degraded exposure; Runs and Overview named
SQL-answered topics; the header test read "Runs"; the reader tests for
the removed topics pass only against the readers with folds. Green after:
vitest 1823 passed, tsc for both configs and eslint clean.
@onexbot-occ-writer

Copy link
Copy Markdown
Contributor

OCC autobind did not mint a companion for this PR: no changed-file candidate could be proven RED against the merge base, and emitting a PR-existence probe instead would be non-falsifiable evidence (OMN-15247). Hand-authored evidence is required.

Considered 12 changed file(s):

  • server/__tests__/postgres-projection-reader.test.ts: no candidate grammar reads this file type (only Python declarations, uv.lock lines, release-artefact lines and contract-pin lines are proposed)
  • server/postgres-projection-reader.ts: no candidate grammar reads this file type (only Python declarations, uv.lock lines, release-artefact lines and contract-pin lines are proposed)
  • server/projection-reader-shared.test.ts: no candidate grammar reads this file type (only Python declarations, uv.lock lines, release-artefact lines and contract-pin lines are proposed)
  • server/projection-reader-shared.ts: no candidate grammar reads this file type (only Python declarations, uv.lock lines, release-artefact lines and contract-pin lines are proposed)
  • server/sqlite-projection-reader.test.ts: no candidate grammar reads this file type (only Python declarations, uv.lock lines, release-artefact lines and contract-pin lines are proposed)
  • server/sqlite-projection-reader.ts: no candidate grammar reads this file type (only Python declarations, uv.lock lines, release-artefact lines and contract-pin lines are proposed)
  • src/pages/LocalDashboardPage.overview.test.tsx: no candidate grammar reads this file type (only Python declarations, uv.lock lines, release-artefact lines and contract-pin lines are proposed)
  • src/pages/LocalDashboardPage.tsx: no candidate grammar reads this file type (only Python declarations, uv.lock lines, release-artefact lines and contract-pin lines are proposed)
  • src/pages/local/overview.contracts.yaml: no candidate grammar reads this file type (only Python declarations, uv.lock lines, release-artefact lines and contract-pin lines are proposed)
  • src/pages/local/overview.page.yaml: no candidate grammar reads this file type (only Python declarations, uv.lock lines, release-artefact lines and contract-pin lines are proposed)
  • src/pages/local/pages.test.ts: no candidate grammar reads this file type (only Python declarations, uv.lock lines, release-artefact lines and contract-pin lines are proposed)
  • src/pages/local/served-catalogue.lab.json: no candidate grammar reads this file type (only Python declarations, uv.lock lines, release-artefact lines and contract-pin lines are proposed)

This verdict is occ-autobind's, derived from this head and its merge base alone. A second producer, the occ-companion-effect backstop, consumes its own command on the lab bus and may still mint a companion under the generic binding. Whether it does depends on that consumer running, not on this diff (OMN-18876).

…verview-runs' into lakshman/OMN-19981-overview-savings
…older model as unknown

Against the dashboard requirements the ticket links and Jonah's handoffs:
- SV-3 (token in and out are separate, never one total): the single
  Tokens card is gone. cost.savings-overview.v1 serves no in/out split
  and the browser may not compute one, so a card bound to nothing says
  "Not served yet: waits on metering-summary.v1".
- OV-3 / SV-2 (the savings figure names its baseline model and pricing
  manifest version): a second binding on the Savings card reads both
  from delegation.savings.v1's served row; a null baseline reads
  "Baseline unresolved".
- OMN-20008 AC4: a Measured runs card (measured_run_count) with the
  zero-token run count beside it, both served.
- Jonah's handoff: model_name 'delegate-skill' (a placeholder) and an
  empty model render "unknown".

The page empty-state rules now look only at each component's first
binding, so the caption binding does not switch Overview into the Runs
session-row rules (it silenced BASELINE_UNRESOLVED until fixed).
Tests first: 7 new cases red, then 1829 passed; tsc both configs and
eslint clean.
…ows status, cause and filters

Requirements OV-4, OV-5, RU-1 and RU-4 (dashboard requirements, section 2),
built from exposures the runtime serves, with ticket AC2 kept by the
strangler the ticket names.

- Overview: a Last run card from the newest delegation.decisions.v1 row by
  written_at. It shows status (passed, or failed with the typed cause from
  quality_gate_detail), run id, model (a placeholder reads unknown),
  duration, tokens in and out separately, cost from the savings session
  with the same id, task type, route tier and age. Backend reads
  "Not served (OMN-20162)"; nothing is blank or 0 when unmeasured.
- Overview: Recent runs, the ten newest decisions, newest first, with the
  run columns. A session with no baseline and 0 savings reads Baseline
  unresolved, never 0. Its cells wrap so every panel fits at 1440x900
  (ticket AC4).
- Runs: each savings session looks up its decision by
  session_id = correlation_id for Status, Cause, Route tier and Quality
  score ("Not recorded" when there is none). Status, model and window
  (today, UTC) filters apply to the served rows only.
- NO_RUNS_YET names the first-run command, onex delegate "ping" (local MVP
  plan, verify step).
- A measured saving with a null baseline_model reads "Not recorded", not
  "Not measured".
- Ticket AC2 strangler: both server readers stop folding
  onex.snapshot.projection.delegation.decisions.v1, which the local pages
  now read from the served exposure. The legacy 'delegation' name keeps its
  SQL. In sqlite and postgres modes the full topic answers [], so the
  legacy Lab, Event Bus, Experiments, Delegation Evidence and routing-table
  views get no decisions rows in those modes; the default http mode is
  unchanged.

Red before: 20 new page and contract tests failed (no components, no
bindings, no filters, the old label), the AC2 test failed for both pages
once they bound decisions, the reader tests asserted the old fold, and a
lakshman-lane read showed 0 savings for unpriced runs in Recent runs.
Green after: vitest 1850 passed, 15 skipped; tsc for both configs and
eslint clean.
…d Runs lists every run

Requirements FR-3, the widget-state table, RU-1 and FR-2 (row badge), from
the dashboard requirements (knowledge-base-internal main 5d4f2d8).

- Navigation (FR-3): the sidebar opens with a Local group listing
  Overview, Runs, Workflow, Usage, Credentials and API Keys, in that order.
  The app opens on Overview instead of the editable canvas. The existing
  pages stay reachable below.
- Partial pages: Workflow, Usage and API Keys carry a "partial" chip and
  name what they wait on (run-trace.v1 OMN-19987, usage-by-model-day.v1
  OMN-20006, local-identity.v1 OMN-19986). API Keys shows CLOUD_NOT_LINKED
  for cloud keys (AK-3).
- Credentials: reads tenant-credentials.v1. It shows provider, key ref,
  set and revoked (never a value), or "No provider key set" with
  `onex secret set` (CR-1 to CR-3). It has no input field.
- Live data: each local page re-reads its exposures every declared
  refresh_interval_seconds (30) and keeps the last good rows on screen
  while it reads. Each panel shows "As of <age>", highlighted once older
  than twice the interval.
- Widget states: an exposure the census does not serve, an HTTP error, or
  no answer in 5 s fails only the widgets bound to it. Each names the
  exposure, the status or "no answer in 5 s", and when it was last good.
  The rest of the page renders, and a failed read is never NO_RUNS_YET.
- Runs (RU-1) lists delegation.decisions.v1 rows, so a failed run with
  no savings row is still listed. The savings session with the same id
  supplies cost, savings, baseline and usage source. It pages at the
  declared page_size (25), adds a cause filter, and badges rows whose
  served data_source is not real (FR-2, row part).

Red before: 57 new or changed tests failed (no local nav, no default
Overview, no partial pages, no polling, a page-wide failure, Runs bound
to savings, no pagination, no cause filter, no badge). Green after:
vitest 1900 passed, 15 skipped; tsc for both configs and eslint clean.
…s Playwright proof, and the sqlite reader is gone

Implements the rest of OMN-19981 (description, AC1 to AC4, and the asks
in Jonah's comments 3376eac0 and aac9032d).

- AC1: every one of the six pages names at least one served exposure.
  - Usage reads usage-by-model-day.v1, showing this tenant's rows only
    (the exposure declares no tenant column), tokens in and out apart. With
    no rows it names the event it waits on (llm-call-completed, OMN-20006).
  - Workflow reads delegation.decisions.v1 and shows the newest run's
    recorded steps (request, routing, quality gate, terminal). The full
    path waits on run-trace.v1 (OMN-19987).
  - API Keys reads the tenant id from delegation.savings.v1 (AK-1).
    Minted-at waits on local-identity.v1 (OMN-19986). Cloud keys show
    CLOUD_NOT_LINKED, with no form.
- AC3: server/sqlite-projection-reader.ts and its tests are deleted. sqlite
  mode now answers 503 sqlite_reader_retired, naming onex dashboard
  (OMN-19976), instead of folding topics by hand-written SQL.
- AC4: the Playwright spec is rewritten for the current pages. Overview
  opens by default at 1440x900 with no clipped panel and no 0 for the
  unmeasured run; Runs and the other four pages fit too. Screenshots are
  refreshed.
- First load: a page that mounts twice (React development mode) shares
  each read in flight instead of sending a second copy that queued past
  the 5 s limit (lakshman lane: 0.9 s, then 5.8 s for the copy).
- Savings is labelled a modelled counterfactual, priced at the baseline's
  list price with no baseline run (aac9032d, item 5).
- Postgres reader (3376eac0): it reads model_cloud_baseline AS
  baseline_model, the column savings_estimates has, and sums savings once
  per run (newest row per session_id) in both savings totals.

Red before: 14 AC1 tests, the shared-read and StrictMode tests, the
modelled label, 3 postgres tests, 4 AC3 tests, and both old Playwright
tests (they targeted baselines.roi.v1). Green after: vitest 1896 passed,
15 skipped; Playwright 3 passed; tsc for both configs and eslint clean.
A mutation that shows 0 for the unmeasured saving fails the Overview
Playwright test.
@Patel230 Patel230 changed the title fix(OMN-19981): Overview shows spend, savings and tokens from the served savings overview feat(OMN-19981): the six local dashboard pages read served exposures, live, with typed empty states Oct 2, 2026
@Patel230
Patel230 marked this pull request as draft October 2, 2026 16:21
Onex-Lane: codex-01a0fd03
Onex-Session: d862cd8315844a4a8e16acf6ff2ba4bb
@Patel230
Patel230 marked this pull request as ready for review October 3, 2026 05:18
@Patel230
Patel230 merged commit 152bf9f into lakshman/OMN-19981-local-pages-overview-runs Oct 3, 2026
18 checks passed
@Patel230
Patel230 deleted the lakshman/OMN-19981-overview-savings branch October 3, 2026 05:26
Patel230 added a commit that referenced this pull request Oct 3, 2026
…346)

* fix(OMN-19981): read Runs from delegation savings

The previous Runs binding targeted a projection that delegation does not populate. Read the tenant-scoped savings envelope, render its sessions, and preserve unresolved baselines as typed state so the Friday dashboard reports real delegation data without inventing zero savings.

Onex-Lane: codex-01a0f6a9
Onex-Session: 5e3c5377abcf4109b5fe3f99fad91c16

* fix(OMN-19981): focus Agent Workbench on evidence panels

Onex-Lane: codex-omn-19981
Onex-Session: 311b6aca89d245a8a13cad9792ad94f9

* test(OMN-19981): add Runs E2E evidence

Onex-Lane: codex-omn-19981
Onex-Session: 311b6aca89d245a8a13cad9792ad94f9

* feat(OMN-19981): the six local dashboard pages read served exposures, live, with typed empty states (#349)

* fix(OMN-19981): Overview shows spend, savings and tokens from the served savings overview

Overview was bound to onex.snapshot.projection.baselines.roi.v1, which the
lab catalogue marks degraded (not_yet_bus_backed) and whose table holds no
rows, so on the lab the page said "Projection exposure is not served".

- overview.contracts.yaml / overview.page.yaml: three metric cards on
  cost.savings-overview.v1 (spend total_cost_usd, savings
  total_savings_usd with total_baseline_cost_usd required, tokens
  tokens_total). local_token_pct is not read; the view does not measure it.
- LocalDashboardPage.tsx: MetricCard renders the field its widget names
  (it always rendered roi_percent). A missing baseline renders "Baseline
  unresolved" and any other missing field "Not measured", never 0. The
  page document resets on a page switch, so Overview opened from Runs no
  longer shows the header "Runs".
- Ticket AC2 strangler: both server readers drop their hand-written SQL
  folds for cost.savings-overview.v1 and delegation.savings.v1, which the
  two local pages now read from served exposures. Each reader keeps its
  other 15 folds; the now-unused buildCostSavingsOverviewResult,
  mergeDelegationSessions wrappers and sessionKey helpers go with them.
  In sqlite and postgres modes the two topics answer [] (the readers'
  unknown-topic path); the default http mode is unchanged.
- pages.test.ts reads the served list from a captured lab catalogue
  (served-catalogue.lab.json) instead of a hand-written one, and checks
  both readers for every local page (it checked one reader for Overview).

Red before: Overview bound a degraded exposure; Runs and Overview named
SQL-answered topics; the header test read "Runs"; the reader tests for
the removed topics pass only against the readers with folds. Green after:
vitest 1823 passed, tsc for both configs and eslint clean.

* fix(OMN-19981): Overview meets SV-3 and OV-3, and Runs shows a placeholder model as unknown

Against the dashboard requirements the ticket links and Jonah's handoffs:
- SV-3 (token in and out are separate, never one total): the single
  Tokens card is gone. cost.savings-overview.v1 serves no in/out split
  and the browser may not compute one, so a card bound to nothing says
  "Not served yet: waits on metering-summary.v1".
- OV-3 / SV-2 (the savings figure names its baseline model and pricing
  manifest version): a second binding on the Savings card reads both
  from delegation.savings.v1's served row; a null baseline reads
  "Baseline unresolved".
- OMN-20008 AC4: a Measured runs card (measured_run_count) with the
  zero-token run count beside it, both served.
- Jonah's handoff: model_name 'delegate-skill' (a placeholder) and an
  empty model render "unknown".

The page empty-state rules now look only at each component's first
binding, so the caption binding does not switch Overview into the Runs
session-row rules (it silenced BASELINE_UNRESOLVED until fixed).
Tests first: 7 new cases red, then 1829 passed; tsc both configs and
eslint clean.

* feat(OMN-19981): Overview shows the last run and recent runs, Runs shows status, cause and filters

Requirements OV-4, OV-5, RU-1 and RU-4 (dashboard requirements, section 2),
built from exposures the runtime serves, with ticket AC2 kept by the
strangler the ticket names.

- Overview: a Last run card from the newest delegation.decisions.v1 row by
  written_at. It shows status (passed, or failed with the typed cause from
  quality_gate_detail), run id, model (a placeholder reads unknown),
  duration, tokens in and out separately, cost from the savings session
  with the same id, task type, route tier and age. Backend reads
  "Not served (OMN-20162)"; nothing is blank or 0 when unmeasured.
- Overview: Recent runs, the ten newest decisions, newest first, with the
  run columns. A session with no baseline and 0 savings reads Baseline
  unresolved, never 0. Its cells wrap so every panel fits at 1440x900
  (ticket AC4).
- Runs: each savings session looks up its decision by
  session_id = correlation_id for Status, Cause, Route tier and Quality
  score ("Not recorded" when there is none). Status, model and window
  (today, UTC) filters apply to the served rows only.
- NO_RUNS_YET names the first-run command, onex delegate "ping" (local MVP
  plan, verify step).
- A measured saving with a null baseline_model reads "Not recorded", not
  "Not measured".
- Ticket AC2 strangler: both server readers stop folding
  onex.snapshot.projection.delegation.decisions.v1, which the local pages
  now read from the served exposure. The legacy 'delegation' name keeps its
  SQL. In sqlite and postgres modes the full topic answers [], so the
  legacy Lab, Event Bus, Experiments, Delegation Evidence and routing-table
  views get no decisions rows in those modes; the default http mode is
  unchanged.

Red before: 20 new page and contract tests failed (no components, no
bindings, no filters, the old label), the AC2 test failed for both pages
once they bound decisions, the reader tests asserted the old fold, and a
lakshman-lane read showed 0 savings for unpriced runs in Recent runs.
Green after: vitest 1850 passed, 15 skipped; tsc for both configs and
eslint clean.

* feat(OMN-19981): the six local pages are the dashboard, read live, and Runs lists every run

Requirements FR-3, the widget-state table, RU-1 and FR-2 (row badge), from
the dashboard requirements (knowledge-base-internal main 5d4f2d8).

- Navigation (FR-3): the sidebar opens with a Local group listing
  Overview, Runs, Workflow, Usage, Credentials and API Keys, in that order.
  The app opens on Overview instead of the editable canvas. The existing
  pages stay reachable below.
- Partial pages: Workflow, Usage and API Keys carry a "partial" chip and
  name what they wait on (run-trace.v1 OMN-19987, usage-by-model-day.v1
  OMN-20006, local-identity.v1 OMN-19986). API Keys shows CLOUD_NOT_LINKED
  for cloud keys (AK-3).
- Credentials: reads tenant-credentials.v1. It shows provider, key ref,
  set and revoked (never a value), or "No provider key set" with
  `onex secret set` (CR-1 to CR-3). It has no input field.
- Live data: each local page re-reads its exposures every declared
  refresh_interval_seconds (30) and keeps the last good rows on screen
  while it reads. Each panel shows "As of <age>", highlighted once older
  than twice the interval.
- Widget states: an exposure the census does not serve, an HTTP error, or
  no answer in 5 s fails only the widgets bound to it. Each names the
  exposure, the status or "no answer in 5 s", and when it was last good.
  The rest of the page renders, and a failed read is never NO_RUNS_YET.
- Runs (RU-1) lists delegation.decisions.v1 rows, so a failed run with
  no savings row is still listed. The savings session with the same id
  supplies cost, savings, baseline and usage source. It pages at the
  declared page_size (25), adds a cause filter, and badges rows whose
  served data_source is not real (FR-2, row part).

Red before: 57 new or changed tests failed (no local nav, no default
Overview, no partial pages, no polling, a page-wide failure, Runs bound
to savings, no pagination, no cause filter, no badge). Green after:
vitest 1900 passed, 15 skipped; tsc for both configs and eslint clean.

* feat(OMN-19981): every local page reads a served exposure, AC4 has its Playwright proof, and the sqlite reader is gone

Implements the rest of OMN-19981 (description, AC1 to AC4, and the asks
in Jonah's comments 3376eac0 and aac9032d).

- AC1: every one of the six pages names at least one served exposure.
  - Usage reads usage-by-model-day.v1, showing this tenant's rows only
    (the exposure declares no tenant column), tokens in and out apart. With
    no rows it names the event it waits on (llm-call-completed, OMN-20006).
  - Workflow reads delegation.decisions.v1 and shows the newest run's
    recorded steps (request, routing, quality gate, terminal). The full
    path waits on run-trace.v1 (OMN-19987).
  - API Keys reads the tenant id from delegation.savings.v1 (AK-1).
    Minted-at waits on local-identity.v1 (OMN-19986). Cloud keys show
    CLOUD_NOT_LINKED, with no form.
- AC3: server/sqlite-projection-reader.ts and its tests are deleted. sqlite
  mode now answers 503 sqlite_reader_retired, naming onex dashboard
  (OMN-19976), instead of folding topics by hand-written SQL.
- AC4: the Playwright spec is rewritten for the current pages. Overview
  opens by default at 1440x900 with no clipped panel and no 0 for the
  unmeasured run; Runs and the other four pages fit too. Screenshots are
  refreshed.
- First load: a page that mounts twice (React development mode) shares
  each read in flight instead of sending a second copy that queued past
  the 5 s limit (lakshman lane: 0.9 s, then 5.8 s for the copy).
- Savings is labelled a modelled counterfactual, priced at the baseline's
  list price with no baseline run (aac9032d, item 5).
- Postgres reader (3376eac0): it reads model_cloud_baseline AS
  baseline_model, the column savings_estimates has, and sums savings once
  per run (newest row per session_id) in both savings totals.

Red before: 14 AC1 tests, the shared-read and StrictMode tests, the
modelled label, 3 postgres tests, 4 AC3 tests, and both old Playwright
tests (they targeted baselines.roi.v1). Green after: vitest 1896 passed,
15 skipped; Playwright 3 passed; tsc for both configs and eslint clean.
A mutation that shows 0 for the unmeasured saving fails the Overview
Playwright test.

* feat(OMN-19981): add page URLs and polish navigation states

Onex-Lane: codex-01a0fd03
Onex-Session: d862cd8315844a4a8e16acf6ff2ba4bb
@Patel230

Patel230 commented Oct 4, 2026

Copy link
Copy Markdown
Contributor Author

Close-out for omnidash#349 (OMN-19981). Merged 2026-10-03T05:26:04Z, merge commit 152bf9f. It landed Overview bound to the served cost.savings-overview.v1 exposure; a missing baseline renders "Baseline unresolved", any other missing field "Not measured", never 0.

Verified: 12 check-runs succeeded, 5 skipped, 1 neutral, none failing. Read live 2026-10-04T15:48Z; no review comments, and the only comments are bot notices.

Unblocks: local_token_pct is deliberately not read (it is a literal 0 in the view; OMN-20320).

Left on OMN-19981: In Review; Jonah was asked at 2026-10-04T09:33Z to review and move it to Done. The Workflow, Usage, Credentials and API Keys pages and AC3 (deleting the sqlite reader) are scoped to the 10/5 sprint by the ticket description.

Thank you!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant