Repository navigation
fix(OMN-19981): show delegation savings and focus workbench evidence - #346
Conversation
The previous Runs binding targeted a projection that delegation does not populate. Read the tenant-scoped savings envelope, render its sessions, and preserve unresolved baselines as typed state so the Friday dashboard reports real delegation data without inventing zero savings. Onex-Lane: codex-01a0f6a9 Onex-Session: 5e3c5377abcf4109b5fe3f99fad91c16
Onex-Lane: codex-omn-19981 Onex-Session: 311b6aca89d245a8a13cad9792ad94f9
|
OCC autobind did not mint a companion for this PR: no changed-file candidate could be proven RED against the merge base, and emitting a PR-existence probe instead would be non-falsifiable evidence (OMN-15247). Hand-authored evidence is required. Considered 8 changed file(s):
This verdict is occ-autobind's, derived from this head and its merge base alone. A second producer, the occ-companion-effect backstop, consumes its own command on the lab bus and may still mint a companion under the generic binding. Whether it does depends on that consumer running, not on this diff (OMN-18876). |
|
The hand-authored companion requested here is onex_change_control#12229. It pins this PR's exact head |
Onex-Lane: codex-omn-19981 Onex-Session: 311b6aca89d245a8a13cad9792ad94f9
…ved savings overview Overview was bound to onex.snapshot.projection.baselines.roi.v1, which the lab catalogue marks degraded (not_yet_bus_backed) and whose table holds no rows, so on the lab the page said "Projection exposure is not served". - overview.contracts.yaml / overview.page.yaml: three metric cards on cost.savings-overview.v1 (spend total_cost_usd, savings total_savings_usd with total_baseline_cost_usd required, tokens tokens_total). local_token_pct is not read; the view does not measure it. - LocalDashboardPage.tsx: MetricCard renders the field its widget names (it always rendered roi_percent). A missing baseline renders "Baseline unresolved" and any other missing field "Not measured", never 0. The page document resets on a page switch, so Overview opened from Runs no longer shows the header "Runs". - Ticket AC2 strangler: both server readers drop their hand-written SQL folds for cost.savings-overview.v1 and delegation.savings.v1, which the two local pages now read from served exposures. Each reader keeps its other 15 folds; the now-unused buildCostSavingsOverviewResult, mergeDelegationSessions wrappers and sessionKey helpers go with them. In sqlite and postgres modes the two topics answer [] (the readers' unknown-topic path); the default http mode is unchanged. - pages.test.ts reads the served list from a captured lab catalogue (served-catalogue.lab.json) instead of a hand-written one, and checks both readers for every local page (it checked one reader for Overview). Red before: Overview bound a degraded exposure; Runs and Overview named SQL-answered topics; the header test read "Runs"; the reader tests for the removed topics pass only against the readers with folds. Green after: vitest 1823 passed, tsc for both configs and eslint clean.
…al-pages-overview-runs
…verview-runs' into lakshman/OMN-19981-overview-savings
…older model as unknown Against the dashboard requirements the ticket links and Jonah's handoffs: - SV-3 (token in and out are separate, never one total): the single Tokens card is gone. cost.savings-overview.v1 serves no in/out split and the browser may not compute one, so a card bound to nothing says "Not served yet: waits on metering-summary.v1". - OV-3 / SV-2 (the savings figure names its baseline model and pricing manifest version): a second binding on the Savings card reads both from delegation.savings.v1's served row; a null baseline reads "Baseline unresolved". - OMN-20008 AC4: a Measured runs card (measured_run_count) with the zero-token run count beside it, both served. - Jonah's handoff: model_name 'delegate-skill' (a placeholder) and an empty model render "unknown". The page empty-state rules now look only at each component's first binding, so the caption binding does not switch Overview into the Runs session-row rules (it silenced BASELINE_UNRESOLVED until fixed). Tests first: 7 new cases red, then 1829 passed; tsc both configs and eslint clean.
…ows status, cause and filters
Requirements OV-4, OV-5, RU-1 and RU-4 (dashboard requirements, section 2),
built from exposures the runtime serves, with ticket AC2 kept by the
strangler the ticket names.
- Overview: a Last run card from the newest delegation.decisions.v1 row by
written_at. It shows status (passed, or failed with the typed cause from
quality_gate_detail), run id, model (a placeholder reads unknown),
duration, tokens in and out separately, cost from the savings session
with the same id, task type, route tier and age. Backend reads
"Not served (OMN-20162)"; nothing is blank or 0 when unmeasured.
- Overview: Recent runs, the ten newest decisions, newest first, with the
run columns. A session with no baseline and 0 savings reads Baseline
unresolved, never 0. Its cells wrap so every panel fits at 1440x900
(ticket AC4).
- Runs: each savings session looks up its decision by
session_id = correlation_id for Status, Cause, Route tier and Quality
score ("Not recorded" when there is none). Status, model and window
(today, UTC) filters apply to the served rows only.
- NO_RUNS_YET names the first-run command, onex delegate "ping" (local MVP
plan, verify step).
- A measured saving with a null baseline_model reads "Not recorded", not
"Not measured".
- Ticket AC2 strangler: both server readers stop folding
onex.snapshot.projection.delegation.decisions.v1, which the local pages
now read from the served exposure. The legacy 'delegation' name keeps its
SQL. In sqlite and postgres modes the full topic answers [], so the
legacy Lab, Event Bus, Experiments, Delegation Evidence and routing-table
views get no decisions rows in those modes; the default http mode is
unchanged.
Red before: 20 new page and contract tests failed (no components, no
bindings, no filters, the old label), the AC2 test failed for both pages
once they bound decisions, the reader tests asserted the old fold, and a
lakshman-lane read showed 0 savings for unpriced runs in Recent runs.
Green after: vitest 1850 passed, 15 skipped; tsc for both configs and
eslint clean.
…d Runs lists every run Requirements FR-3, the widget-state table, RU-1 and FR-2 (row badge), from the dashboard requirements (knowledge-base-internal main 5d4f2d8). - Navigation (FR-3): the sidebar opens with a Local group listing Overview, Runs, Workflow, Usage, Credentials and API Keys, in that order. The app opens on Overview instead of the editable canvas. The existing pages stay reachable below. - Partial pages: Workflow, Usage and API Keys carry a "partial" chip and name what they wait on (run-trace.v1 OMN-19987, usage-by-model-day.v1 OMN-20006, local-identity.v1 OMN-19986). API Keys shows CLOUD_NOT_LINKED for cloud keys (AK-3). - Credentials: reads tenant-credentials.v1. It shows provider, key ref, set and revoked (never a value), or "No provider key set" with `onex secret set` (CR-1 to CR-3). It has no input field. - Live data: each local page re-reads its exposures every declared refresh_interval_seconds (30) and keeps the last good rows on screen while it reads. Each panel shows "As of <age>", highlighted once older than twice the interval. - Widget states: an exposure the census does not serve, an HTTP error, or no answer in 5 s fails only the widgets bound to it. Each names the exposure, the status or "no answer in 5 s", and when it was last good. The rest of the page renders, and a failed read is never NO_RUNS_YET. - Runs (RU-1) lists delegation.decisions.v1 rows, so a failed run with no savings row is still listed. The savings session with the same id supplies cost, savings, baseline and usage source. It pages at the declared page_size (25), adds a cause filter, and badges rows whose served data_source is not real (FR-2, row part). Red before: 57 new or changed tests failed (no local nav, no default Overview, no partial pages, no polling, a page-wide failure, Runs bound to savings, no pagination, no cause filter, no badge). Green after: vitest 1900 passed, 15 skipped; tsc for both configs and eslint clean.
…s Playwright proof, and the sqlite reader is gone
Implements the rest of OMN-19981 (description, AC1 to AC4, and the asks
in Jonah's comments 3376eac0 and aac9032d).
- AC1: every one of the six pages names at least one served exposure.
- Usage reads usage-by-model-day.v1, showing this tenant's rows only
(the exposure declares no tenant column), tokens in and out apart. With
no rows it names the event it waits on (llm-call-completed, OMN-20006).
- Workflow reads delegation.decisions.v1 and shows the newest run's
recorded steps (request, routing, quality gate, terminal). The full
path waits on run-trace.v1 (OMN-19987).
- API Keys reads the tenant id from delegation.savings.v1 (AK-1).
Minted-at waits on local-identity.v1 (OMN-19986). Cloud keys show
CLOUD_NOT_LINKED, with no form.
- AC3: server/sqlite-projection-reader.ts and its tests are deleted. sqlite
mode now answers 503 sqlite_reader_retired, naming onex dashboard
(OMN-19976), instead of folding topics by hand-written SQL.
- AC4: the Playwright spec is rewritten for the current pages. Overview
opens by default at 1440x900 with no clipped panel and no 0 for the
unmeasured run; Runs and the other four pages fit too. Screenshots are
refreshed.
- First load: a page that mounts twice (React development mode) shares
each read in flight instead of sending a second copy that queued past
the 5 s limit (lakshman lane: 0.9 s, then 5.8 s for the copy).
- Savings is labelled a modelled counterfactual, priced at the baseline's
list price with no baseline run (aac9032d, item 5).
- Postgres reader (3376eac0): it reads model_cloud_baseline AS
baseline_model, the column savings_estimates has, and sums savings once
per run (newest row per session_id) in both savings totals.
Red before: 14 AC1 tests, the shared-read and StrictMode tests, the
modelled label, 3 postgres tests, 4 AC3 tests, and both old Playwright
tests (they targeted baselines.roi.v1). Green after: vitest 1896 passed,
15 skipped; Playwright 3 passed; tsc for both configs and eslint clean.
A mutation that shows 0 for the unmeasured saving fails the Overview
Playwright test.
Onex-Lane: codex-01a0fd03 Onex-Session: d862cd8315844a4a8e16acf6ff2ba4bb
… live, with typed empty states (#349) * fix(OMN-19981): Overview shows spend, savings and tokens from the served savings overview Overview was bound to onex.snapshot.projection.baselines.roi.v1, which the lab catalogue marks degraded (not_yet_bus_backed) and whose table holds no rows, so on the lab the page said "Projection exposure is not served". - overview.contracts.yaml / overview.page.yaml: three metric cards on cost.savings-overview.v1 (spend total_cost_usd, savings total_savings_usd with total_baseline_cost_usd required, tokens tokens_total). local_token_pct is not read; the view does not measure it. - LocalDashboardPage.tsx: MetricCard renders the field its widget names (it always rendered roi_percent). A missing baseline renders "Baseline unresolved" and any other missing field "Not measured", never 0. The page document resets on a page switch, so Overview opened from Runs no longer shows the header "Runs". - Ticket AC2 strangler: both server readers drop their hand-written SQL folds for cost.savings-overview.v1 and delegation.savings.v1, which the two local pages now read from served exposures. Each reader keeps its other 15 folds; the now-unused buildCostSavingsOverviewResult, mergeDelegationSessions wrappers and sessionKey helpers go with them. In sqlite and postgres modes the two topics answer [] (the readers' unknown-topic path); the default http mode is unchanged. - pages.test.ts reads the served list from a captured lab catalogue (served-catalogue.lab.json) instead of a hand-written one, and checks both readers for every local page (it checked one reader for Overview). Red before: Overview bound a degraded exposure; Runs and Overview named SQL-answered topics; the header test read "Runs"; the reader tests for the removed topics pass only against the readers with folds. Green after: vitest 1823 passed, tsc for both configs and eslint clean. * fix(OMN-19981): Overview meets SV-3 and OV-3, and Runs shows a placeholder model as unknown Against the dashboard requirements the ticket links and Jonah's handoffs: - SV-3 (token in and out are separate, never one total): the single Tokens card is gone. cost.savings-overview.v1 serves no in/out split and the browser may not compute one, so a card bound to nothing says "Not served yet: waits on metering-summary.v1". - OV-3 / SV-2 (the savings figure names its baseline model and pricing manifest version): a second binding on the Savings card reads both from delegation.savings.v1's served row; a null baseline reads "Baseline unresolved". - OMN-20008 AC4: a Measured runs card (measured_run_count) with the zero-token run count beside it, both served. - Jonah's handoff: model_name 'delegate-skill' (a placeholder) and an empty model render "unknown". The page empty-state rules now look only at each component's first binding, so the caption binding does not switch Overview into the Runs session-row rules (it silenced BASELINE_UNRESOLVED until fixed). Tests first: 7 new cases red, then 1829 passed; tsc both configs and eslint clean. * feat(OMN-19981): Overview shows the last run and recent runs, Runs shows status, cause and filters Requirements OV-4, OV-5, RU-1 and RU-4 (dashboard requirements, section 2), built from exposures the runtime serves, with ticket AC2 kept by the strangler the ticket names. - Overview: a Last run card from the newest delegation.decisions.v1 row by written_at. It shows status (passed, or failed with the typed cause from quality_gate_detail), run id, model (a placeholder reads unknown), duration, tokens in and out separately, cost from the savings session with the same id, task type, route tier and age. Backend reads "Not served (OMN-20162)"; nothing is blank or 0 when unmeasured. - Overview: Recent runs, the ten newest decisions, newest first, with the run columns. A session with no baseline and 0 savings reads Baseline unresolved, never 0. Its cells wrap so every panel fits at 1440x900 (ticket AC4). - Runs: each savings session looks up its decision by session_id = correlation_id for Status, Cause, Route tier and Quality score ("Not recorded" when there is none). Status, model and window (today, UTC) filters apply to the served rows only. - NO_RUNS_YET names the first-run command, onex delegate "ping" (local MVP plan, verify step). - A measured saving with a null baseline_model reads "Not recorded", not "Not measured". - Ticket AC2 strangler: both server readers stop folding onex.snapshot.projection.delegation.decisions.v1, which the local pages now read from the served exposure. The legacy 'delegation' name keeps its SQL. In sqlite and postgres modes the full topic answers [], so the legacy Lab, Event Bus, Experiments, Delegation Evidence and routing-table views get no decisions rows in those modes; the default http mode is unchanged. Red before: 20 new page and contract tests failed (no components, no bindings, no filters, the old label), the AC2 test failed for both pages once they bound decisions, the reader tests asserted the old fold, and a lakshman-lane read showed 0 savings for unpriced runs in Recent runs. Green after: vitest 1850 passed, 15 skipped; tsc for both configs and eslint clean. * feat(OMN-19981): the six local pages are the dashboard, read live, and Runs lists every run Requirements FR-3, the widget-state table, RU-1 and FR-2 (row badge), from the dashboard requirements (knowledge-base-internal main 5d4f2d8). - Navigation (FR-3): the sidebar opens with a Local group listing Overview, Runs, Workflow, Usage, Credentials and API Keys, in that order. The app opens on Overview instead of the editable canvas. The existing pages stay reachable below. - Partial pages: Workflow, Usage and API Keys carry a "partial" chip and name what they wait on (run-trace.v1 OMN-19987, usage-by-model-day.v1 OMN-20006, local-identity.v1 OMN-19986). API Keys shows CLOUD_NOT_LINKED for cloud keys (AK-3). - Credentials: reads tenant-credentials.v1. It shows provider, key ref, set and revoked (never a value), or "No provider key set" with `onex secret set` (CR-1 to CR-3). It has no input field. - Live data: each local page re-reads its exposures every declared refresh_interval_seconds (30) and keeps the last good rows on screen while it reads. Each panel shows "As of <age>", highlighted once older than twice the interval. - Widget states: an exposure the census does not serve, an HTTP error, or no answer in 5 s fails only the widgets bound to it. Each names the exposure, the status or "no answer in 5 s", and when it was last good. The rest of the page renders, and a failed read is never NO_RUNS_YET. - Runs (RU-1) lists delegation.decisions.v1 rows, so a failed run with no savings row is still listed. The savings session with the same id supplies cost, savings, baseline and usage source. It pages at the declared page_size (25), adds a cause filter, and badges rows whose served data_source is not real (FR-2, row part). Red before: 57 new or changed tests failed (no local nav, no default Overview, no partial pages, no polling, a page-wide failure, Runs bound to savings, no pagination, no cause filter, no badge). Green after: vitest 1900 passed, 15 skipped; tsc for both configs and eslint clean. * feat(OMN-19981): every local page reads a served exposure, AC4 has its Playwright proof, and the sqlite reader is gone Implements the rest of OMN-19981 (description, AC1 to AC4, and the asks in Jonah's comments 3376eac0 and aac9032d). - AC1: every one of the six pages names at least one served exposure. - Usage reads usage-by-model-day.v1, showing this tenant's rows only (the exposure declares no tenant column), tokens in and out apart. With no rows it names the event it waits on (llm-call-completed, OMN-20006). - Workflow reads delegation.decisions.v1 and shows the newest run's recorded steps (request, routing, quality gate, terminal). The full path waits on run-trace.v1 (OMN-19987). - API Keys reads the tenant id from delegation.savings.v1 (AK-1). Minted-at waits on local-identity.v1 (OMN-19986). Cloud keys show CLOUD_NOT_LINKED, with no form. - AC3: server/sqlite-projection-reader.ts and its tests are deleted. sqlite mode now answers 503 sqlite_reader_retired, naming onex dashboard (OMN-19976), instead of folding topics by hand-written SQL. - AC4: the Playwright spec is rewritten for the current pages. Overview opens by default at 1440x900 with no clipped panel and no 0 for the unmeasured run; Runs and the other four pages fit too. Screenshots are refreshed. - First load: a page that mounts twice (React development mode) shares each read in flight instead of sending a second copy that queued past the 5 s limit (lakshman lane: 0.9 s, then 5.8 s for the copy). - Savings is labelled a modelled counterfactual, priced at the baseline's list price with no baseline run (aac9032d, item 5). - Postgres reader (3376eac0): it reads model_cloud_baseline AS baseline_model, the column savings_estimates has, and sums savings once per run (newest row per session_id) in both savings totals. Red before: 14 AC1 tests, the shared-read and StrictMode tests, the modelled label, 3 postgres tests, 4 AC3 tests, and both old Playwright tests (they targeted baselines.roi.v1). Green after: vitest 1896 passed, 15 skipped; Playwright 3 passed; tsc for both configs and eslint clean. A mutation that shows 0 for the unmeasured saving fails the Overview Playwright test. * feat(OMN-19981): add page URLs and polish navigation states Onex-Lane: codex-01a0fd03 Onex-Session: d862cd8315844a4a8e16acf6ff2ba4bb
Onex-Lane: codex-omn-19981 Onex-Session: 311b6aca89d245a8a13cad9792ad94f9
Onex-Lane: codex-omn-19981 Onex-Session: 311b6aca89d245a8a13cad9792ad94f9
…erved (#363) On a fresh install (.201, 2026-10-05) the store served one decision, but delegation.savings.v1 answered 503 projection_table_missing: the local SQLite store has no savings table and nothing local writes one. The Runs page hid the decision and printed "HTTP 503 Service Unavailable", so AC1 ("Runs rows after one delegation") could not pass on any fresh store. Two changes. projection_table_missing is a refusal the read node's contract declares, so the source now carries the backend's named code and the loader files that one code as a typed not-served state instead of an HTTP error; every other non-2xx stays an error. And a table's rows are now gated only by its first binding (RU-1's exposure is decisions.v1); a second binding is a lookup for cost and basis, so when it is not served the rows stay and the lookup is named under them as a typed status. A failed lookup still shows its error beside the rows. Metric cards keep needing all of their bindings. This reverses one case of OMN-19981's widget-state rule (#346): an unserved secondary binding used to blank a table. That test now asserts the new behaviour.
Summary
Follow-up for OMN-19981 after omnidash#344. Runs now reads the tenant-scoped delegation-savings exposure that
onex delegatepopulates, preserves unresolved baselines as typed state, and the default Agent Workbench focuses on delegation traces, token savings, and event-bus traces.What changed
onex.snapshot.projection.swarm.runs.v1toonex.snapshot.projection.delegation.savings.v1.sessions[]into one table row per session and declared the consumed session fields in the page contract.BASELINE_UNRESOLVEDwhenbaseline_modelis null and savings is unmeasured, and retainedNO_RUNS_YETfor absent or empty sessions.control-planeanddelegate-taskfrom the default Agent Workbench.live-event-stream,delegation-token-usage, andconsumer-flow.How it was verified
Current base:
origin/devcc628179717f4830bdd6c5a5e1b8d0c7bc4042ff.src/pages/local/pages.test.tsandsrc/pages/LocalDashboardPage.test.tsx.src/templates/templates.test.ts.control-planemade 3/10 focused tests fail; restoring the implementation returned 10/10.npm test -- --run src/templates src/components/dashboardpassed 78 files / 533 tests on clean dev and 78 files / 534 tests on the branch.npm run check,npm run lint,npm run build, andgit diff --checkpassed.OMN-19981-AC-lab-20261001T172544Z.jsonpassed against product head752307696c73741fd7d23b71ebb50b077fb61e3f: the runtime/skilledge and direct projection returned identical rows (SHA-256471cc5046999d913c7ef13384a4cf2cea4a994c816a4d77bf93525c59b575580), and the product transform preserved all 21 unique sessions, including 4 unresolved baselines.fe24ddfef652363dccfb10fae66278cbf361ef22and752307696c73741fd7d23b71ebb50b077fb61e3fare signed.Failure paths
NO_RUNS_YET.Baseline unresolvedand does not expose$0as measured savings.Open defects
2da199e618221e02222182a60467eaf85207512e. A follow-up OCC repair must bind this PR's reachable reviewed head and regenerate the native test receipts before AC-E3 is complete.Local gates
origin/dev.20261001T170304Z-codex-01a0f6a9.md.Not in this change
server/sqlite-projection-reader.ts.Evidence-Ticket: OMN-19981
Evidence-Source: OCC#12229