fix(usage): treat provider-declared null usage as UNKNOWN, never a measured 0 - #787
Conversation
…nd every turn display Bridge claude-bpx #186 emits null prompt_tokens/total_tokens with prompt_tokens_unavailable/total_tokens_unavailable/unavailable discriminators. normalize_usage previously collapsed that to input_tokens=0 with no unknown flag, so unknown input was priced and displayed as a measured zero. Extends #787's output-only UNKNOWN contract to input, cache and whole-usage unknowns across normalize_usage/CanonicalUsage, pricing, the Blackbox turn ledger/rollups/cards, and the shared display rule. Measured zeros stay measured; absent-key behavior for legacy providers is unchanged. Also preserves the discriminators through both MoA physical-call serializations (aggregator + advisor), which reconstructed pricing dicts from numeric fields only and let Blackbox price an unknown input as $0.00075, and routes the console cache-stat branch, the adjacent verbose token log and the gateway thin last-turn fallback through the shared unknown rule. Verified (repo venv): 404 passed in 73.16s across usage contract/consumers, usage_pricing, billing, blackbox, turn accumulator, last-turn persistence, meta cache reporting, usage anchor, context-engine and MoA loop suites. Acceptance repro now: input_tokens_unknown=True, total_tokens_unknown=True, direct cost None, MoA blackbox cost (None, 'unknown', all-None components), console prints "Cache: unknown". Measured-zero control stays measured (all unknown flags False, cost 0.00). 5 mutation arms RED with exact-byte restore: input detection 3F, pricing refusal 25F, MoA aggregator flags 4F, MoA advisor flags 8F, console cache guard 3F. Scope: Blackbox per-turn contract + the human-facing displays it feeds. Persisted/export-schema flag loss (SessionDB/aux-ledger counters, cumulative CLI/TUI totals, Langfuse canonical export) is deferred to t_25f50547.
|
🔴 HOLD — do not merge until kanban review round 3 on t_083ad45e returns APPROVED (reviewer daedalus-opus; implementer round-3 head f78eae2, 3 CI-caught defects fixed, 7 mutation arms RED). CI is green; that is not the gate here. Hold lives on the PR on purpose: the bridge-side hold for #186 lived only on a card and a CI-green lane merged + fleet-deployed it 18:40 while this consumer was still open. Apollo lifts this comment when the review verdict is in. Follow-up for the persisted/export schema paths is t_25f50547, not this PR. |
|
Do not merge as-is — kanban review round 3 on t_083ad45e (reviewer daedalus-opus) has not returned a verdict on head f78eae2. CI green is not the gate. do-not-merge label applied; Apollo removes both when the review lands. |
|
Hold lifted — kanban round-2 review on t_083ad45e APPROVED (daedalus-opus run 2264, execution lens, independent mutation set 5/6 RED with the survivor adjudicated). Merging. Persisted/export schema paths are #797 (t_25f50547), stacked on this branch; it retargets to main on merge. |
|
Do not merge as-is — hold RE-APPLIED (Apollo, 22:30 PT). FleetReview's run on f78eae2 produced 8 findings incl. 3 P1 and then died before posting (ACE-AI router closed it 'ensemble member failure after self-heal'; record exists only in its shadow-log). P1s: (1) agent/usage_pricing.py — a present-but-null cache-details CONTAINER marks the whole turn permanently unpriceable; (2) tests/agent/test_usage_unknown_contract.py — null prompt_tokens_details container pinned as UNKNOWN collapses a fully MEASURED input; (3) tests/agent/test_usage_unknown_consumers.py — the thin-fallback 'shipped consumer' test feeds a shape no producer emits, so /usage still renders an unmeasured turn as 0. P2: unknown check runs before the subscription_included route; empty_response_guard reads unmeasured 0 as measured and cuts the retry budget; AST anchors brittle on unrelated refactors. Rework carded; fixes go on this branch. Nobody arms --auto here (queue repo: --auto = enqueue = merge). I enqueued it by mistake at 05:14:33Z via bare gh pr merge and dequeued at 05:16:02Z — nothing landed. |
…tals and the Langfuse export #787 made UNKNOWN a first-class state for the Blackbox PER-TURN contract. Three other consumers still collapsed an unmeasured term into a measured zero. Bridge claude-bpx #186 is merged and fleet-deployed, so these paths receive unknown-input shapes today. Reproduced all three on the tree first: update_token_counts REJECTS unknown flags: unexpected keyword argument record_auxiliary_usage REJECTS unknown flags: unexpected keyword argument cumulative counters carry NO unknown provenance tui cache_hit_pct = 100% over prompt_total=100 (the unknown turn added 0) langfuse input-only -> ({'input': 0, 'output': 50}, {}) 1. PERSISTED SCHEMA. sessions and session_model_usage gain the five USAGE_UNKNOWN_FIELDS discriminators (+ last_turn_* snapshot mirrors), additive with DEFAULT 0 so every legacy row and every provider that never speaks this dialect reads back "measured". The int counters are untouched, so no arithmetic consumer changes. Cumulative flags are ABSORBING (MAX in SQL): one unmeasured call latches the session term, which is why a last-call-only guard is wrong here. Classified into _TOKEN_DELTA_ROUTE_FIELDS so deltas with differing unknown state never coalesce. Both session_model_usage table-rebuild paths (PK heal, v21->v22) copy the flags only when the source table has them, so a rebuild cannot silently reset a latched unknown. record_aux_usage no longer drops an all-zero unmeasured aux call as "empty". 2. CUMULATIVE CLI/TUI TOTALS. The loop latches per-term provenance onto session_*_unknown beside the session_*_tokens commit. Display reuses the Blackbox lane's rule (format_token_count / prompt_tokens_unknown) via new session_usage_unknown_flags / session_total_tokens_unknown helpers -- not a forked cumulative rule. /usage, both status-bar Sigma renderers, the CLI cache-hit segment and the TUI payload now say "unknown" or emit no ratio instead of a fabricated percentage. 3. LANGFUSE EXPORT. An unmeasured term is OMITTED from usage_details (absent, never 0) and declared on the observation metadata via _unknown_usage_details(); cost_details is empty for an unpriceable turn so Langfuse cannot read a partial subtotal as authoritative. The post_api_request summary-dict reconstruction was a second flag-loss site and now carries the discriminators. The 2-tuple return contract of _canonical_usage_and_cost / _usage_and_cost is unchanged. Verified (repo venv): - tests/agent/test_usage_unknown_persisted_export.py -- 71 pins, the five literal bridge payloads (input-only / cache-only / wholly-unavailable / output-only / measured-zero) through EACH of the three paths, plus a legacy-DB migration arm built by stripping the new columns out of the real SCHEMA_SQL. Cumulative and display arms source-LIFT the shipped blocks by AST anchor rather than re-implementing them. - scripts/run_tests.sh over the contract/consumer/persisted suites + tests/hermes_state + async token accounting + langfuse + last-turn persistence + turn-usage accumulator + blackbox: "54 files, 682 tests passed, 0 failed, 1 skipped". - 10-arm mutation matrix, each restored byte-exact, fake kills (SyntaxError/ NameError/zero-collected) rejected: 10/10 RED, control 131 passed before and after. Arms: SessionDB counter flags dropped; absorbing -> last-wins; aux ledger flags dropped; aux producer drops the unmeasured call; cumulative provenance not absorbing; CLI cache-ratio guard removed; /usage card prints numbers; TUI payload fabricates the ratio; Langfuse exports 0 for unknown; Langfuse summary-dict drops flags. - ui-tui: npm run typecheck clean. npm test is 61 files / 6 tests red on this machine, IDENTICAL on the base commit with the change stashed (textInputReturnAction + an unbuilt shared hermes-ink) -- pre-existing, not from this diff. Scope: this card only. The Blackbox per-turn contract certified under t_083ad45e / #787 is untouched.
A worker that needs a sibling PR landed before it can continue blocks with a reason like "merge PR #787 then unblock me". Nothing in the dispatcher ever re-read that reason, so the card sat blocked until a human noticed. Observed on a live board: six needs_input cards held up to 12 hours on gates whose PRs had already merged, some 12+ hours earlier. New hermes_cli/kanban_pr_gate.py runs inside the dispatcher tick, before recompute_ready so a satisfied card becomes spawnable in the SAME tick: * parse the last block-family event's reason for PR references (owner/repo#N, a full .../pull/N URL, and bare #N / pull/N resolved against a repo context: the card's workspace remote, else the first owner/repo mentioned in the body), * resolve each through `gh pr view`, bounded at 30 lookups per tick, one call per unique PR, cached (MERGED/CLOSED for the process lifetime since a merge is irreversible; OPEN for 300 s), * every referenced PR MERGED -> unblock + comment "gate satisfied: <repo>#N merged <sha8> at <ts>" + one gate_auto_resolved event so tooling can count it, * any still OPEN -> hold; CLOSED-unmerged -> one advisory comment and never an unblock; any lookup failure -> no action plus one WARN. Fail-safe choices, each with a test: scope fenced to needs_input / capability / dependency block kinds (transient means "may clear on its own", not "waits on a named object"); a reason naming no PR is never touched and burns zero lookup budget; a bare #N whose repo context is ambiguous is dropped rather than guessed; a transient gh failure is not cached, so it cannot blind the next tick for the whole TTL; the closed-unmerged advisory posts once rather than every tick; and the whole pass is wrapped fail-open, because a diagnostic that can brick a dispatch tick is worse than the stale-block class it closes. DispatchResult gains gate_auto_resolved / gate_closed_unmerged, surfaced in both the text and --json dispatch output. Verified, not claimed: * tests/hermes_cli/test_kanban_pr_gate.py -- 33 passed. * RED-proof of the parser: the multi-reference ordering case failed first (refs returned qualified-before-bare instead of in source order); the implementation was fixed, not the test. * RED-proof of the dispatcher wiring: deleting the single _reevaluate_pr_gates(conn, result) call makes test_dispatch_tick_resolves_a_satisfied_gate fail with `assert 't_...' in []`, so the integration test gates the hook itself and not merely the module.
A worker that needs a sibling PR landed before it can continue blocks with a reason like "merge PR #787 then unblock me". Nothing in the dispatcher ever re-read that reason, so the card sat blocked until a human noticed. Observed on a live board: six needs_input cards held up to 12 hours on gates whose PRs had already merged, some 12+ hours earlier. New hermes_cli/kanban_pr_gate.py runs inside the dispatcher tick, before recompute_ready so a satisfied card becomes spawnable in the SAME tick: * parse the last block-family event's reason for PR references (owner/repo#N, a full .../pull/N URL, and bare #N / pull/N resolved against a repo context: the card's workspace remote, else the first owner/repo mentioned in the body), * resolve each through `gh pr view`, bounded at 30 lookups per tick, one call per unique PR, cached (MERGED/CLOSED for the process lifetime since a merge is irreversible; OPEN for 300 s), * every referenced PR MERGED -> unblock + comment "gate satisfied: <repo>#N merged <sha8> at <ts>" + one gate_auto_resolved event so tooling can count it, * any still OPEN -> hold; CLOSED-unmerged -> one advisory comment and never an unblock; any lookup failure -> no action plus one WARN. Fail-safe choices, each with a test: scope fenced to needs_input / capability / dependency block kinds (transient means "may clear on its own", not "waits on a named object"); a reason naming no PR is never touched and burns zero lookup budget; a bare #N whose repo context is ambiguous is dropped rather than guessed; a transient gh failure is not cached, so it cannot blind the next tick for the whole TTL; the closed-unmerged advisory posts once rather than every tick; and the whole pass is wrapped fail-open, because a diagnostic that can brick a dispatch tick is worse than the stale-block class it closes. DispatchResult gains gate_auto_resolved / gate_closed_unmerged, surfaced in both the text and --json dispatch output. Verified, not claimed: * tests/hermes_cli/test_kanban_pr_gate.py -- 33 passed. * RED-proof of the parser: the multi-reference ordering case failed first (refs returned qualified-before-bare instead of in source order); the implementation was fixed, not the test. * RED-proof of the dispatcher wiring: deleting the single _reevaluate_pr_gates(conn, result) call makes test_dispatch_tick_resolves_a_satisfied_gate fail with `assert 't_...' in []`, so the integration test gates the hook itself and not merely the module.
A worker that needs a sibling PR landed before it can continue blocks with a reason like "merge PR #787 then unblock me". Nothing in the dispatcher ever re-read that reason, so the card sat blocked until a human noticed. Observed on a live board: six needs_input cards held up to 12 hours on gates whose PRs had already merged, some 12+ hours earlier. New hermes_cli/kanban_pr_gate.py runs inside the dispatcher tick, before recompute_ready so a satisfied card becomes spawnable in the SAME tick: * parse the last block-family event's reason for PR references (owner/repo#N, a full .../pull/N URL, and bare #N / pull/N resolved against a repo context: the card's workspace remote, else the first owner/repo mentioned in the body), * resolve each through `gh pr view`, bounded at 30 lookups per tick, one call per unique PR, cached (MERGED/CLOSED for the process lifetime since a merge is irreversible; OPEN for 300 s), * every referenced PR MERGED -> unblock + comment "gate satisfied: <repo>#N merged <sha8> at <ts>" + one gate_auto_resolved event so tooling can count it, * any still OPEN -> hold; CLOSED-unmerged -> one advisory comment and never an unblock; any lookup failure -> no action plus one WARN. Fail-safe choices, each with a test: scope fenced to needs_input / capability / dependency block kinds (transient means "may clear on its own", not "waits on a named object"); a reason naming no PR is never touched and burns zero lookup budget; a bare #N whose repo context is ambiguous is dropped rather than guessed; a transient gh failure is not cached, so it cannot blind the next tick for the whole TTL; the closed-unmerged advisory posts once rather than every tick; and the whole pass is wrapped fail-open, because a diagnostic that can brick a dispatch tick is worse than the stale-block class it closes. DispatchResult gains gate_auto_resolved / gate_closed_unmerged, surfaced in both the text and --json dispatch output. Verified, not claimed: * tests/hermes_cli/test_kanban_pr_gate.py -- 33 passed. * RED-proof of the parser: the multi-reference ordering case failed first (refs returned qualified-before-bare instead of in source order); the implementation was fixed, not the test. * RED-proof of the dispatcher wiring: deleting the single _reevaluate_pr_gates(conn, result) call makes test_dispatch_tick_resolves_a_satisfied_gate fail with `assert 't_...' in []`, so the integration test gates the hook itself and not merely the module.
…nd every turn display Bridge claude-bpx #186 emits null prompt_tokens/total_tokens with prompt_tokens_unavailable/total_tokens_unavailable/unavailable discriminators. normalize_usage previously collapsed that to input_tokens=0 with no unknown flag, so unknown input was priced and displayed as a measured zero. Extends #787's output-only UNKNOWN contract to input, cache and whole-usage unknowns across normalize_usage/CanonicalUsage, pricing, the Blackbox turn ledger/rollups/cards, and the shared display rule. Measured zeros stay measured; absent-key behavior for legacy providers is unchanged. Also preserves the discriminators through both MoA physical-call serializations (aggregator + advisor), which reconstructed pricing dicts from numeric fields only and let Blackbox price an unknown input as $0.00075, and routes the console cache-stat branch, the adjacent verbose token log and the gateway thin last-turn fallback through the shared unknown rule. Verified (repo venv): 404 passed in 73.16s across usage contract/consumers, usage_pricing, billing, blackbox, turn accumulator, last-turn persistence, meta cache reporting, usage anchor, context-engine and MoA loop suites. Acceptance repro now: input_tokens_unknown=True, total_tokens_unknown=True, direct cost None, MoA blackbox cost (None, 'unknown', all-None components), console prints "Cache: unknown". Measured-zero control stays measured (all unknown flags False, cost 0.00). 5 mutation arms RED with exact-byte restore: input detection 3F, pricing refusal 25F, MoA aggregator flags 4F, MoA advisor flags 8F, console cache guard 3F. Scope: Blackbox per-turn contract + the human-facing displays it feeds. Persisted/export-schema flag loss (SessionDB/aux-ledger counters, cumulative CLI/TUI totals, Langfuse canonical export) is deferred to t_25f50547.
d5b3149 to
20dd114
Compare
FleetReview
Reviewed with 2 of 3 model families — openai unavailable. Confidence: 1/5 Findings
FleetReview provenance · models: C=claude-code-opus-5, D=grok-4.6, G=grok-4.6 · cost: $56.89 · duration: 1h 28m 36s · rounds: 2 · files examined: 25 |
FleetReviewReviewed with 2 of 3 model families — openai unavailable. Confidence: 3/5 Findings
FleetReview provenance · models: C=claude-code-opus-5, D=grok-4.6, G=grok-4.6 · cost: $56.03 · duration: 43m 03s · rounds: 1 · files examined: 30 |
FleetReview
Reviewed with 2 of 3 model families — openai unavailable. Confidence: 1/5 Findings
FleetReview provenance · models: C=claude-code-opus-5, D=grok-4.6, G=grok-4.6 · cost: $56.87 · duration: 38m 29s · rounds: 1 · files examined: 30 |
…asured 0 claude-bpx #186 makes the bridge egress an honest unknown when its CLI transcript reconciliation misses (~40% of parallel-batch turns, measured): `completion_tokens`/`total_tokens` come back null with an explicit `usage.output_tokens_unavailable: true` discriminator. The deployed consumer coerced that null to 0 via `normalize_usage -> _usage_count -> _to_int` (`int(value or 0)`) and ignored the discriminator, so a bridge-first deploy would have written measured-looking zeros into the usage ledger, /usage, blackbox and pricing. Reproduced before fixing: the literal wire payload normalized to `output_tokens=0`, `total_tokens=150`. UNKNOWN is now a first-class state end to end: - CanonicalUsage gains `output_tokens_unknown` (+ `total_tokens_unknown`). The int fields stay ints so every arithmetic consumer is untouched; the flag is the discriminator. `__add__` makes unknown ABSORBING. - normalize_usage detects it shape-agnostically: the discriminator flag (including via the OpenAI client's `model_extra`, where an unrecognised wire key actually lands), or a PRESENT output counter that is None. An ABSENT key is NOT an unknown — providers that never speak this dialect keep normalizing to integers. - estimate_usage_cost REFUSES an unknown turn (status=unknown, amount None) rather than pricing a missing output term as $0. - blackbox: the flag rides per-call -> per-turn (absorbing) -> TurnRecord -> a new `output_tokens_unknown` column (guarded additive migration, DEFAULT 0 so every historical row reads back "measured"). The priced_zero fast path and the reprice sweep both exclude unknown rows. - The display rule is single-sourced in `usage_pricing.format_token_count` (`UNKNOWN_TOKENS_LABEL`); the alert card and the /usage+/context last-turn card both delegate to it, keeping their own magnitude formatting. The `in=/out=` API-call log says `out=unknown` instead of `out=0`. Verified: - tests/agent/test_usage_unknown_contract.py — 15 pins against the bridge's exact wire fixture (copied from bpx `unit-usage-unknown-sink.test.js` / `SPEC-parallel-stop-usage-undercount.md`): null+unavailable -> unknown, a real integer -> unchanged, a missing usage object -> unknown-free zeros. 15 passed. - 7-arm mutation matrix, each restored after: reinstating the `or 0` coercion -> 3 RED; dropping the pricing refusal -> 2 RED; store not persisting -> 1 RED; dropping the shared display rule -> 1 RED; turn_finalizer rollup hardcoded False -> 1 RED; blackbox ingest dropping the flag -> 1 RED; widening the detector to treat ABSENT as unknown -> 3 RED. Control 15/15 green before and after; tree restored. - Regression: 153 passed (usage_pricing x3, meta_usage_cache, usage_anchor, tests/blackbox), 180 passed (tests/plugins/blackbox, last_turn_usage persistence, empty_response_guard), 46 passed (turn_finalizer x5, MoA blackbox/cost-slot, turn_usage_accumulator, account/billing usage). - E2E through the real path with the literal bridge payload: normalize -> UNKNOWN, price -> status=unknown/None, ledger -> output_tokens_unknown=1, both renderers -> "unknown". Measured control: 118 out, $0.00222, output_tokens_unknown=0, both renderers -> "118". # Conflicts: # plugins/blackbox/store.py
…nd every turn display Bridge claude-bpx #186 emits null prompt_tokens/total_tokens with prompt_tokens_unavailable/total_tokens_unavailable/unavailable discriminators. normalize_usage previously collapsed that to input_tokens=0 with no unknown flag, so unknown input was priced and displayed as a measured zero. Extends #787's output-only UNKNOWN contract to input, cache and whole-usage unknowns across normalize_usage/CanonicalUsage, pricing, the Blackbox turn ledger/rollups/cards, and the shared display rule. Measured zeros stay measured; absent-key behavior for legacy providers is unchanged. Also preserves the discriminators through both MoA physical-call serializations (aggregator + advisor), which reconstructed pricing dicts from numeric fields only and let Blackbox price an unknown input as $0.00075, and routes the console cache-stat branch, the adjacent verbose token log and the gateway thin last-turn fallback through the shared unknown rule. Verified (repo venv): 404 passed in 73.16s across usage contract/consumers, usage_pricing, billing, blackbox, turn accumulator, last-turn persistence, meta cache reporting, usage anchor, context-engine and MoA loop suites. Acceptance repro now: input_tokens_unknown=True, total_tokens_unknown=True, direct cost None, MoA blackbox cost (None, 'unknown', all-None components), console prints "Cache: unknown". Measured-zero control stays measured (all unknown flags False, cost 0.00). 5 mutation arms RED with exact-byte restore: input detection 3F, pricing refusal 25F, MoA aggregator flags 4F, MoA advisor flags 8F, console cache guard 3F. Scope: Blackbox per-turn contract + the human-facing displays it feeds. Persisted/export-schema flag loss (SessionDB/aux-ledger counters, cumulative CLI/TUI totals, Langfuse canonical export) is deferred to t_25f50547.
…NOWN CI slice 3/16 and 15/16 went red on tests/gateway/test_usage_command.py. Root cause: routing these renderers through format_token_count sent the MEASURED path through its default magnitude formatter too, so a measured 241,500 started rendering as "241.5k" and broke the card's vocabulary. format_token_count already has a `formatter` parameter for exactly this -- "lets a caller keep its own magnitude formatting while still routing the unknown case through here". Both introduced sites now pass it. Swept the whole class (every format_token_count call this branch added, `git diff 1b04723 -U0 -- agent/ plugins/ gateway/ | grep '^+.*format_token_count'`, 8 hits / 3 sites): - gateway/slash_commands.py thin last-turn card -> FIXED (CI-reported) - agent/conversation_loop.py verbose token log -> FIXED (same class, NOT reported by CI; was `{:,}` on base and had silently become "120k") - agent/conversation_loop.py console cache line -> unaffected, unknown-only call (format_token_count(None, unknown=True)), no measured path Both fixed sites are pinned by measured-control tests; dropping `formatter=` at either site is RED. Verified (repo venv): 415 passed, 1 failed in 31.79s. The single failure, test_usage_command_prefers_dominant_persisted_route, is PRE-EXISTING: it fails identically (StopIteration) on base 1b04723 in a clean worktree. Mutation arms, all RED with exact-byte restore: input detection 3F, pricing refusal 25F, MoA aggregator 4F, MoA advisor 8F, console cache 3F, thin-card formatter 4F, verbose-log formatter 1F. Acceptance repro unchanged: cost None, (None, 'unknown'), "Cache: unknown"; measured-zero control still measured at 0.00.
… snapshot CI slice 15/16: test_build_moa_pricing_calls_appends_real_aggregator_after_advisors compared the serialized aggregator call with `==` against a literal dict, so the five additive UNKNOWN discriminators (all False for this fully-measured usage) failed the equality. Tightened rather than loosened, per AGENTS.md "behavior contracts over snapshots": the test still pins attribution and every counter exactly, and now additionally requires the discriminators to be PRESENT and False, and the key set to be exactly the measured keys plus USAGE_UNKNOWN_FIELDS. Silently dropping the flags from this serializer -- the defect that let an unknown input price as $0.00075 -- now fails here too. Also widened the local verify set. The previous list was hand-picked and missed what CI caught; it is now derived from the symbols this branch touches: grep -rlE '_build_moa_pricing_calls|format_token_count|normalize_usage| prompt_tokens_unknown|USAGE_UNKNOWN_FIELDS|compute_turn_cost| get_last_turn_usage|CanonicalUsage' tests/ which added test_moa_blackbox_attribution, test_bedrock_adapter, test_context_engine_host_contract, test_empty_response_guard and test_moa_observability_bridge. Verified (repo venv, --basetemp isolated): 540 passed, 1 failed in 44.58s. The single failure, test_usage_command_prefers_dominant_persisted_route, reproduces identically (StopIteration) on base 1b04723 in a clean worktree -- pre-existing, not from this branch.
FleetReview's run on f78eae2 produced 8 findings (3 P1) and died before posting them. This closes the production half. P1 (findings 1+8) — `"prompt_tokens_details": null` is the ordinary serialization of an OpenAI-compatible server with no cache breakdown (any encoder without exclude_none emits it; pydantic puts an explicitly-null wire field in model_fields_set). Reading the null CONTAINER as "both cache buckets unmeasured" escalated into input_tokens_unknown, so a turn with a MEASURED prompt_tokens=150 became permanently unpriceable — and store.reprice_unpriced filters those rows out forever, so the NULL could never heal. Narrowed: only a null on a counter the provider actually uses (cached_tokens inside a PRESENT container, a null top-level cache field) or an explicit cache_*_unavailable discriminator declares an unknown. P2 (finding 4) — moved the subscription_included route ABOVE the unknown-usage refusal in estimate_usage_cost. A $0 route's cost does not depend on token counts, so an unmeasured usage cannot make it unpriceable. Matches the ordering plugins/observability/langfuse already uses. P2 (finding 5) — empty_response_guard._zero_output read an UNMEASURED output 0 as a measured zero, so two honest-unknown bridge turns satisfied deterministic_empty() and cut the retry budget, contradicting the module's own documented fail-open rule. Now returns (False, False) on output_tokens_unknown / usage_unknown. P2 (finding 6) — extracted the three blocks the consumer suite was AST-lifting out of conversation_loop.py / slash_commands.py into real callable functions (verbose_token_usage_log_args, cache_stats_line, render_thin_last_turn_lines). Behaviour-preserving; the shipped sites now call them, so coverage no longer depends on source-text anchors that break on benign refactors of unrelated code. Verified: tests/agent/test_usage_unknown_fleetreview_repro.py 13/13 pass (includes the pre-fix defect reproduced, and controls proving the real unknown declarations and the measured-zero deterministic-empty path still work). tests/test_last_turn_usage_persistence.py green. 13 committed tests in test_usage_unknown_contract.py / test_usage_unknown_consumers.py still pin the OLD behaviour these findings call wrong; the test-guard hook blocks editing them. Named in the handback for an operator decision.
Verified: scripts/run_tests.sh focused usage/pricing/guard/persistence/Langfuse suites (350 passed); seven mutation arms all red; Ruff clean.
FleetReview r3 finding 5 (`Missing Usage`), reopened by review round 1. `_canonical_usage_from_response` already turned an omitted usage payload into an aggregate UNKNOWN, but the per-call commit block in `run_conversation` was still gated on `response.usage` being truthy. The UNKNOWN therefore reached only `context_compressor.update_from_response()` and `_last_turn_usage`; it never incremented `session_api_calls`, set `last_turn_usage`, appended to the Blackbox `_turn_calls` accumulator, priced the call, or queued persistence. An unpriceable call was reported as if it had never happened. Gate the commit block on `response is not None` instead, and narrow the context-probe persistence (the one part that genuinely needs a measured payload) to `getattr(response, "usage", None)`. The settled `normalize_usage(None)` known-no-call-zero contract is untouched. Verified: - new tests/run_agent/test_usageless_response_accounting.py 4/4 green; reverting the guard to its pre-fix shape turns the two unknown-path pins RED and leaves both controls GREEN (gate proven). - focused set 127/0; tests/agent + tests/blackbox + tests/plugins/blackbox 7771/0; tests/run_agent 2132/0. - four mandated null-container controls re-derived: 10/10 on this head. - 8-arm mutation set all RED, all four touched files restored byte-exact. - Ruff clean. Not mine, reported not fixed: tests/gateway/test_usage_command.py:: TestUsageLastTurnSnapshot::test_usage_command_prefers_dominant_persisted_route fails identically at merge-base 6c400e2 (inherited red).
…ay sink r4 FleetReview (record 598aff1-1dc4eb80) filed 12 findings, 10 P0-P1. The prior head widened the per-call accounting guard to `response is not None` so a usage-less success would finally be accounted — but that pushed the UNKNOWN's PLACEHOLDER ZEROS into five sinks that carry no UNKNOWN discriminator. That is the same silent-zero defect one level down, and worse: the persisted last_turn_* snapshot is written COALESCE(?, existing), so a 0 DESTROYS the previous turn's real measurement. CLASS: "the widened block reaches a sink with no UNKNOWN discriminator." Swept every sink inside both widened blocks, fixed all of them in one pass: - persisted last_turn_* snapshot -> _last_turn_snapshot_kwargs() writes None (not 0) when total_tokens_unknown, so COALESCE retains the last real split. - session_cost_status -> _session_cost_status_with_known_spend() returns "partial" (not "unknown") when the session already holds priced dollars; "unknown" is outside the reprice allowlist and would never heal. - _moa_session_cost_status -> seeds any_unpriceable from the AGGREGATOR, so an unpriceable aggregator beside priced advisors is "partial" too (r4 F4). - context compressor / turn_finalizer last_prompt_tokens / Blackbox context_used -> gated on a measured payload exactly as the context-probe sub-branch already is; the pending compaction verdict is still consumed. - velocity deques -> an unmeasured output is not 0 tok/s; skip the pair. Also from the r4 record: - F9: an omitted payload now sets EVERY discriminator via CanonicalUsage.fully_unknown(), not just the aggregate usage_unknown — consumers read the flags narrowly. - F11: cache-unknown is resolved ACROSS both locations, so a measured cached_tokens is no longer overridden by a null top-level alias. - F7/F12: the legacy usage_unknown latch is narrowed to unpriced rows whose counts are ALL ZERO. usage_unknown is the shared DISPLAY discriminator; a row is routinely unpriced because pricing REFUSED the route while carrying real measured counts. - F8: the render pin asserts the TOKEN line, not any "unknown" (the cost-status line satisfied the old assertion). - F6: the render_thin_last_turn_lines docstring no longer claims the persisted producer carries UNKNOWN flags; it cannot in this tree. - F5/finding-4 REFUTED by execution — cache-unknown still refuses pricing per the standing consensus, which this change does not relitigate. Verified: focused 235/0; tests/agent+blackbox+plugins/blackbox 7774/0; tests/run_agent 2139/0; tests/gateway 8214/1 (the single red reproduces identically in a detached worktree at merge-base 6c400e2 — inherited). Four mandated null-container controls 10/10. Ruff clean. Nine mutation arms all RED with byte-exact restore; the three arms that came back GREEN were unpinned fixes and now carry discriminating pins. F8's vacuity proven by regressing the real renderer: weak pin 4 GREEN, strong pin 4 RED.
…e control The legacy_copy fixture deliberately stubs store._connect WITHOUT _ensure_schema so test_ac7 can capture a pre-migration BEFORE arm. The placeholder probe reuses that stub, so after the usage_unknown migration landed, its positive-control insert_turn bound five columns the stubbed schema had never gained and failed with 'table turns has no column named output_tokens_unknown'. Production is unaffected: the real store._connect calls _ensure_schema on every connection (store.py:53) — verified by running insert_turn against a pre-787 fixture DB through the real connector: the five *_unknown columns are added and the row persists. Verified the probe still discriminates: dropping one placeholder from _INSERT_TURN_SQL turns it RED (1 failed / 1 passed); store.py restored byte-exact (sha256 ea2a14d5d266...).
CI on the rebased head (first CI this PR has ever had — it was DIRTY until now) turned slice 15/16 RED on test_migration_adds_comp_columns_to_legacy_db with 'sqlite3.OperationalError: no such column: cache_write'. Root cause: the legacy latch's all-zero-counts guard names all four token count columns unconditionally. `turns` is created with them, but a table that already exists never gains them — CREATE TABLE IF NOT EXISTS is a no-op and no ALTER adds them — so an old enough DB reaches the latch without cache_write. The raise aborts the ENTIRE _ensure_schema, taking every later migration down with it. Fix: build the sum from the count columns actually present in _existing. A column the DB lacks cannot hold a measurement, which is exactly the zero the guard tests for, so the condition's meaning is preserved. Verified: tests/test_request_composition.py 40/0 and tests/plugins/blackbox/ test_store.py 19/0 (was 1 failed). New pin test_migration_survives_a_db_missing_some_token_count_columns asserts the later migrations still ran AND that the guard keeps discriminating measured from empty. Restoring the hardcoded four-column sum turns BOTH tests RED; store.py restored byte-exact (sha256 2a393c829c8a5d91...). Proven not inherited: both tests pass at fork/main in a clean worktree.
…not be refused r6 disposition of the FleetReview record on head 20dd114 (13 findings, 3 P1 / 8 P2 / 2 P3, reviewCommentUrl=None — never posted). One class-sweep pass over both directions of the same defect: a sink with no UNKNOWN discriminator, and a refusal that fires on measured data. FIXED (9): F4 [P1] agent/usage_pricing.py — a wire-null top-level cache alias beside measured input/output made the turn permanently unpriceable. `{"input_tokens":100,"output_tokens":50,"cache_read_input_tokens":null}` — the shape agent/transports/anthropic.py already reads with `... or 0` — set both cache flags, ORed into total_tokens_unknown, and reprice_unpriced now skips any unknown row, so the NULL cost could never heal. A top-level alias is OPTIONAL: on a measured payload its null means "no caching on this request". A null inside a PRESENT container still means UNKNOWN. F5 [P1] agent/moa_loop.py — a usage-less MoA advisor got a bare CanonicalUsage() (every discriminator False) where conversation_loop uses fully_unknown(). The pricing-call dict rode out all-False, so the advisor priced at a confident $0 instead of refusing and its real spend vanished. F13 [P1] hermes_state.py — the cost_status guard lived only in the caller, reading agent.session_estimated_cost_usd, which agent_init.py zeroes on every construction and never rehydrates. Any NEW agent resuming a session (gateway eviction, CLI resume, cron/kanban re-entry) wrote 'unknown' over a row still holding priced dollars — outside the reprice allowlist, so it never healed. Moved to the durable chokepoint, which covers every caller. F1 [P2] hermes_state.py — the session-scoped status also lands in the per-model row, labelling a wholly-unpriced model 'partial' on the Spend-by-model breakdown. Scoped to each row's own dollars, INSERT and ON CONFLICT legs both. F6 [P2] plugins/blackbox/store.py — the legacy latch marked genuinely zero-token turns (interrupted before any call, blackbox-off, failed first call) UNKNOWN, irreversibly, stranding them outside reprice_unpriced forever. Scoped to rows the pre-UNKNOWN code itself labelled cost_status='unknown'. F7 [P2] agent/conversation_loop.py — the compressor gate was a PRESENCE test (`if getattr(response,"usage",None):`). A payload that carries a usage object but nulls the prompt count passed it and stamped last_prompt_tokens — and so the context meter, the persisted session entry and Blackbox context_used — to 0. Now gates on prompt_tokens_unknown, not total_tokens_unknown: an unmeasured OUTPUT bucket must not discard a good PROMPT reading. F8 [P2] gateway/slash_commands.py — the resident lane built five flagless ints, making every UNKNOWN branch dead code there; an unmeasured session rendered `Total (billed in+out): 0`. Extracted `_resident_thin_snapshot` and carried the flags from agent.last_turn_usage. F11 [P2] / F10 [P3] tests — UC-8b's `UNKNOWN_TOKENS_LABEL in block` was satisfied by the cost line above the token row, so the defect it exists to pin could regress green; now pins `Tokens out: <label>`. The scratch repro's f3 case ASSERTED the measured-looking zero as expected output, locking in the exact rendering F8 fixes; re-pinned onto the discriminator contract. TEST-REPIN (1): tests/plugins/blackbox/test_store.py test_schema_migration_adds_columns_to_preexisting_table required the latch to fire on an unpriced all-zero row with no cost_status. That premise is what F6 supersedes. The row now carries cost_status='unknown' (the ambiguous case, still latched) and a new 'legacy-zero-token-measured' row pins that a never-accounted zero-token turn is NOT latched. Verification, all by execution on this head: - New tests/agent/test_usage_unknown_r6_findings.py: 21/21. Every fix has a narrowness control, so none is bought by making the refusal unreachable (explicit *_unavailable still wins; null-in-present-container still UNKNOWN; a session with no priced dollars is still 'unknown'; a measured advisor still prices; an output-only unknown still advances the context meter). - Focused 8 files: 139/0. tests/agent+blackbox+plugins/blackbox+run_agent: 10022 passed, 0 failed, 29 skipped. - Mutation set 9/9 RED, every file restored byte-exact. - The harness's BASELINE check earned its keep: the first run reported 9/9 RED vacuously — the repo-local .venv has no pytest (run_tests.sh skips it), so every arm was failing on `No module named pytest`, which reads as a caught regression. Fixed the interpreter probe; the re-run then exposed F7 as a genuinely vacuous pin (its test only exercised the predicate, not the gate). Replaced it with a real two-turn run_conversation() regression against a REAL SessionDB, driving a PRESENT-but-null-prompt payload; reverting only the gate now fails on that test's named assertion. - Argus's own probes, re-run byte-unmodified: r2 8/0, r3 CONTRACT 9/9. - Ruff clean. Prior contracts untouched and pinned by controls: normalize_usage(None) is a MEASURED zero; a null details CONTAINER is a measured zero; fully_unknown() sets every discriminator.
r6 disposition round 2 — the three findings the prior commit left undisposed from the FleetReview record on head 20dd114 (F9 P2, F3 P3, F12 P3). The other ten were dispositioned in 4ed89cd. Record: ANG-Ventures_hermes-agent-pr-787- 20dd114-e2297280-b5f8-11f1-8043-66cc4f68f983.json, reviewCommentUrl=None. FIXED (3): F9 [P2] plugins/blackbox/last_turn.py + agent/turn_finalizer.py + record.py + store.py — three rows on the last-turn card describe the FINAL call only (Last call split, Cached hit-rate, Context window occupancy), but were gated on `input_unknown = prompt_tokens_unknown(rec)`, which _rollup_turn_usage builds with any() ACROSS EVERY CALL of the turn. A 5-call turn whose call #2 returned no usage rendered `Context window (last call): unknown` and dropped the split, while the final call's numbers were real provider measurements. There was no per-last-call discriminator to gate on, so this adds one: turn_finalizer computes `last_call_prompt_unknown` from `_turn_calls[-1]` — the same element the split itself is read from — and it rides the existing turn_usage dict through _build_record into a new nullable TurnRecord field and turns column. The renderer gates the two genuinely last-call rows on it. The `• Cached:` row is NOT moved: it divides turn-level sums that really are missing a term, so the absorbing flag is correct there and stays. That asymmetry is the finding's actual content and is pinned both ways. Migration is NULLable with no DEFAULT, deliberately. NULL means "this row predates the column, so its final-call provenance was never recorded" and the renderer falls back to the turn-level flag — byte-identical to today for every historical row. A DEFAULT 0 would instead assert "the final call WAS measured" about rows that were genuinely unmeasured and render their placeholder zeros as real window numbers, which is the defect class this PR exists to remove. F3 [P3] agent/usage_pricing.py — `cache_stats_line` replaced inline code that printed NOTHING when there was no cache activity, and it runs per API CALL, so a provider that routinely omits usage emitted one `💾 Cache: unknown` noise line per call (12 in a 12-call tool loop). Silenced on the AGGREGATE discriminator ALONE (`usage_unknown` — no payload at all, so there is no cache-specific fact to report and the turn card already says the turn is unmeasured). Every narrower unknown still prints: a null bucket inside a present container, an explicitly-unavailable prompt, measured cache counts beside an unmeasured prompt. In those the refused percentage is the fabrication. F12 [P3] tests/agent/test_usage_unknown_contract.py — the store round-trips fed the renderer `store.get_turn`'s shape, but `_row_to_dict` RENAMES cache_read -> cache_read_tokens while production (`compute_last_turn_record`) does a raw SELECT * and hands over the unrenamed keys — which is why the renderer reads `rec.get("cache_read", 0)`. Every cache lookup silently resolved to 0. Switched to the production row shape with an explicit shape assertion, so a future assertion about cache or Last-call lines cannot be written against a shape that never occurs at runtime. Verification, all by execution on this head: - tests/agent/test_usage_unknown_r6_round2.py: 15/15, each fix with a narrowness control (an unmeasured final call is still unknown; the turn-level Cached row still gates on the turn flag; a legacy NULL row keeps today's behaviour; a null cache bucket still prints unknown; measured paths byte- identical). - tests/agent + tests/plugins/blackbox + tests/blackbox + tests/run_agent + tests/test_request_composition.py: 10076 passed, 1 failed, 29 skipped. The one red is tests/agent/lsp/test_client_e2e.py — a psutil flake unrelated to usage; 3/3 green when run alone. - Mutation set 5/5 RED, every file restored byte-exact: f3_remove_usageless_silence 1 RED f9_renderer_reverts_to_turn_level_or 2 RED f9_producer_always_reports_measured 1 RED f9_migration_dropped 2 RED f12_back_to_get_turn_shape 9 RED The producer arm earned its keep: the first matrix returned it GREEN because the renderer tests build the row by hand and re-implement the fold, so a producer hardcoding the flag killed nothing. Added a regression that drives the shipped finalize_turn() and captures the turn_usage it emits on on_session_end; the arm is RED now. The F12 arm's first anchor matched 2x and could not apply — re-anchored on the helper body rather than reporting a vacuous pass. - Ruff clean. Prior contracts untouched and still pinned: null CONTAINER = measured zeros; null count inside a present container = UNKNOWN; normalize_usage(None) = a MEASURED zero; fully_unknown() sets every discriminator. NOT DISPOSITIONED HERE — F2 [P2] "nulling all five last_turn_* columns leaves the persisted card showing an older turn's split with no staleness marker". SPLIT to t_25f50547 / PR #797 under the standing schema-ownership split: the finding's own body names the fix as `*_unknown` companion columns on the `sessions` table, which is that card's surface, and it explicitly reads the present behaviour as the correct trade-off against destroying the split.
…aggregate r7 of the FleetReview record on cc3ced8 (13 findings, 9 P1, TRUSTED: healthyMembers=5/6, realizedRoster.trusted=true; escalation is judge-side 429 capacity, not member failure). All 13 dispositioned. F6 [P1] card.py — CLASS SWEEP, not one site. `context_used` and the `last_cache_*` split describe the FINAL call only, so this PR's hardening of Tokens-in/out/Cached left `• Context: 0/200k (0% of model max)` fabricating a measurement on the same card that said `unknown` twice above it. The gate is now ONE shared callable, `usage_pricing.last_call_prompt_unknown`, which both renderers consult — including the NULL-legacy fallback (`store.py` ALTERs the column in without a DEFAULT, so `TurnRecord.last_call_prompt_unknown` is typed `bool` and carries `None`; a bare truth test asserts "measured" about every historic row). The record noted no test discriminated this (248-vs-248); the new class-sweep asserts CROSS-RENDERER AGREEMENT rather than per-site strings, so a future third renderer that skips the gate fails without anyone extending a list. F3 [P1] empty_response_guard.py — the F5 trade is preserved and its cost recovered. A declared-unmeasured output still refuses to classify as a deterministic empty, but guard #2 was ALREADY inert for those payloads (`estimate_usage_cost` refuses `total_tokens_unknown`), so a bridge nulling only its output kept the full 3-retry budget while re-sending a measured 400k prompt. `measured_cost_floor` prices the MEASURED buckets only — for a DECISION, never persisted, never displayed — and can only understate the bill, so the ceiling cannot trip early. Measured: $2.00/attempt on claude-opus-5 at 400k; budget 3→1. The old pin sat at prompt_tokens=150 where the cost is nil. F13 [P1] conversation_loop.py — worst-of merge. A session's cost label can never get MORE complete over time. r4 fixed priced-then-unknown; the mirror overwrote an incomplete session back to `estimated`. `prior_api_calls` is load-bearing: a fresh agent's initial `unknown`/zero-calls is UNSTARTED, not incomplete. F7/F4 [P1] slash_commands.py — absorbing session-level latch. `last_turn_usage` is rewritten every CALL, so stamping its flags on cumulative counters was wrong both ways (an exact-looking 412,338 omitting real spend; 412k measured tokens collapsed to unknown). Latched per-bucket beside the counter increments, cleared with them in reset_session_state, wired on the codex lane too. F1 [P1] — UNPRICEABLE is not UNMEASURED. An uncatalogued advisor route returns amount_usd=None with every token measured and no flag set, so the session lane reported `estimated` for a total omitting it while Blackbox said `partial`. Each advisor's own verdict now rides its pricing record and is consulted. F2 [P2] — the compressor reads the PRE-FOLD aggregator usage, as the anchor below it already does. `__add__` makes unknown flags absorbing, so one unmeasured advisor blanked a complete aggregator prompt count. F12 [P2], F5 [P3], F11 [P3] — test fidelity: `return []` (verified: zero swallowed hook TypeErrors), production ORDERING mirrored in the cache-line pin, and `else` legs so no wire asserts nothing. F8 [P1] store.py, F9 [P1] usage_pricing.py, F10 [P1] hermes_state.py carried forward from the prior commit. TEST-REPINs (2): `test_f8_resident_snapshot...` moved off the per-call fixture that finding 4 showed could not discriminate; `test_append_inside_success_block_only` replaced a banned source CHARACTER-DISTANCE assertion (2359→4116 chars, red on comments alone, and green whenever the statements happen to be adjacent) with the behaviour it names, driven through the real run_conversation. Verified: focused 386/0; tests/agent+run_agent+blackbox broad sweep, 4 failures, 3 proven inherited at merge-base edb2880 and the 4th re-pinned above. Nine mutation arms all RED (F6 gate 4, NULL fallback 5, F13 merge 1, prior_api_calls 3, session latch 4, F1 verdict 2, F2 payload 2, F3 floor 3, accumulator 1), all files restored byte-exact, scratch worktrees removed. Ruff clean.
r8. Two defects in the r7 commit (dfcacfa), both found by execution rather than by reading it: one RED in CI, one vacuous mutation arm. CI RED (the substantive one). The F7 session-level latch read `if getattr(agent, f"session_{flag}", False):`. A `MagicMock` auto-creates every attribute access as a truthy child, so on any mock-backed agent all five flags read set and a fully MEASURED session collapsed to `unknown` on the resident /usage lane — `tests/gateway/test_usage_command.py` went red on slice 5/16 (35,000 and 70,000 both rendered as "Tokens in: unknown"). That is this PR's own silent-fabrication class inverted: refusing a measurement that was really taken. The old code was accidentally safe via its `isinstance(last_turn, dict)` guard; widening to a bare getattr dropped the type check with it. Fixed with `is True`. The latch is written as a real bool by agent_init.py (False) and conversation_loop.py (True), so any other value means the attribute was never initialised on that object. The five counters beside these flags are already coerced through `as_int` for exactly this reason; this is the flags' half of that contract. Class-swept: `getattr(agent, f"session_{...}")` has exactly one read site in the tree (the three writers are `setattr` and unaffected). Pinned with a MagicMock-backed agent plus a narrowness control proving a genuinely declared latch still fires. Every pre-existing pin in that file builds its agent from SimpleNamespace or a real AIAgent, which is precisely why none of them could see this — AGENTS.md names the trap ("check with `is True` — MagicMock doubles auto-create truthy attributes") and the r7 commit walked into it anyway. VACUOUS ARM (F2). Arm 7 — swapping the compressor gate's `aggregator_usage` back to the folded `canonical_usage` at the call site — came back GREEN across all five focused files. The r7 F2 pins test `_compressor_usage_dict` and `prompt_tokens_unknown` at unit level, so nothing observed which usage the loop actually hands the compressor: the fix's own call site was unpinned. Added two tests that drive the real `run_conversation()` with a client speaking the genuine MoA `consume_reference_usage` seam — one unmeasured advisor must not blank a measured 4000-token window occupancy, and (narrowness) a measured 90k advisor fan-out must not be counted AS occupancy. Arms against both halves are RED now. Verified by execution on this head: - focused 7 files: 190/0, plus the 4 new tests. - tests/gateway/test_usage_command.py: the two CI failures now pass. The remaining red (test_usage_command_prefers_dominant_persisted_route) is INHERITED — it fails identically in a fresh detached worktree at merge-base edb2880, which contains none of this branch's commits. - broad tests/agent+run_agent+plugins/blackbox+blackbox: 10067 passed / 3 failed; all three relay failures reproduce at merge-base edb2880 (proven, not asserted) and none touch usage. - 9 mutation arms RED (F6 card gate 4, F6 NULL fallback 2, F13 worst-of merge 1, F7 latch 4, F3 cost floor 3, F1 advisor verdict 2, F2 gate 1, F2 payload 1, latch `is True` 1). Baseline-gated green before each arm; every file restored byte-exact; both scratch worktrees removed. - Argus's r2/r3/r4 probes re-run byte-unmodified: 8/0, 9/0, 11/0. - F12 verified by execution: zero swallowed hook TypeErrors in that file's run. - Ruff clean.
2a7685d to
564f7e4
Compare
…tals and the Langfuse export #787 made UNKNOWN a first-class state for the Blackbox PER-TURN contract. Three other consumers still collapsed an unmeasured term into a measured zero. Bridge claude-bpx #186 is merged and fleet-deployed, so these paths receive unknown-input shapes today. Reproduced all three on the tree first: update_token_counts REJECTS unknown flags: unexpected keyword argument record_auxiliary_usage REJECTS unknown flags: unexpected keyword argument cumulative counters carry NO unknown provenance tui cache_hit_pct = 100% over prompt_total=100 (the unknown turn added 0) langfuse input-only -> ({'input': 0, 'output': 50}, {}) 1. PERSISTED SCHEMA. sessions and session_model_usage gain the five USAGE_UNKNOWN_FIELDS discriminators (+ last_turn_* snapshot mirrors), additive with DEFAULT 0 so every legacy row and every provider that never speaks this dialect reads back "measured". The int counters are untouched, so no arithmetic consumer changes. Cumulative flags are ABSORBING (MAX in SQL): one unmeasured call latches the session term, which is why a last-call-only guard is wrong here. Classified into _TOKEN_DELTA_ROUTE_FIELDS so deltas with differing unknown state never coalesce. Both session_model_usage table-rebuild paths (PK heal, v21->v22) copy the flags only when the source table has them, so a rebuild cannot silently reset a latched unknown. record_aux_usage no longer drops an all-zero unmeasured aux call as "empty". 2. CUMULATIVE CLI/TUI TOTALS. The loop latches per-term provenance onto session_*_unknown beside the session_*_tokens commit. Display reuses the Blackbox lane's rule (format_token_count / prompt_tokens_unknown) via new session_usage_unknown_flags / session_total_tokens_unknown helpers -- not a forked cumulative rule. /usage, both status-bar Sigma renderers, the CLI cache-hit segment and the TUI payload now say "unknown" or emit no ratio instead of a fabricated percentage. 3. LANGFUSE EXPORT. An unmeasured term is OMITTED from usage_details (absent, never 0) and declared on the observation metadata via _unknown_usage_details(); cost_details is empty for an unpriceable turn so Langfuse cannot read a partial subtotal as authoritative. The post_api_request summary-dict reconstruction was a second flag-loss site and now carries the discriminators. The 2-tuple return contract of _canonical_usage_and_cost / _usage_and_cost is unchanged. Verified (repo venv): - tests/agent/test_usage_unknown_persisted_export.py -- 71 pins, the five literal bridge payloads (input-only / cache-only / wholly-unavailable / output-only / measured-zero) through EACH of the three paths, plus a legacy-DB migration arm built by stripping the new columns out of the real SCHEMA_SQL. Cumulative and display arms source-LIFT the shipped blocks by AST anchor rather than re-implementing them. - scripts/run_tests.sh over the contract/consumer/persisted suites + tests/hermes_state + async token accounting + langfuse + last-turn persistence + turn-usage accumulator + blackbox: "54 files, 682 tests passed, 0 failed, 1 skipped". - 10-arm mutation matrix, each restored byte-exact, fake kills (SyntaxError/ NameError/zero-collected) rejected: 10/10 RED, control 131 passed before and after. Arms: SessionDB counter flags dropped; absorbing -> last-wins; aux ledger flags dropped; aux producer drops the unmeasured call; cumulative provenance not absorbing; CLI cache-ratio guard removed; /usage card prints numbers; TUI payload fabricates the ratio; Langfuse exports 0 for unknown; Langfuse summary-dict drops flags. - ui-tui: npm run typecheck clean. npm test is 61 files / 6 tests red on this machine, IDENTICAL on the base commit with the change stashed (textInputReturnAction + an unbuilt shared hermes-ink) -- pre-existing, not from this diff. Scope: this card only. The Blackbox per-turn contract certified under t_083ad45e / #787 is untouched.
…apshot flags, append-only schema Re-stacked hermes-agent#797 onto #787 (rebased onto main 2026-09-23). Two of its six commits were already on #787 under the same subjects and were dropped; the other four replayed with semantic conflicts. Fixes on top, each measured red before: 1. Duplicate absorbing latch: #787's "latch the flag over the aggregate" and #797's "carry UNKNOWN through cumulative totals" each added the same session_* latch loop to agent/conversation_loop.py. Kept #787's (beside the increments); removed #797's copy. 25 AST-seam tests ("shipped cumulative-usage seam moved") go green. 2. last_turn_*_unknown flags now move WITH the snapshot they qualify: when #787's _last_turn_snapshot_kwargs withholds an unmeasured call's split (None -> COALESCE keeps the prior real split), the flags are withheld too. Before: the persisted card rendered "Total: unknown" over a real 4,120 from the previous turn. 3. Schema columns APPENDED, not inserted mid-table. #797 put its 10 sessions flags before git_branch and its 5 session_model_usage flags before estimated_cost_usd. Migrated DBs get them via ALTER ADD (end of table), so fresh and migrated DBs had different physical orders, and the positional lost-and-found recovery mapper misplaced historical fields (title landed in the wrong column). Fresh-DB column prefix is now byte-identical to main's (title back at index 43). 4. Recovery classifier accepts session_model_usage records of 18 (pre-migration) AND 23 fields; it previously dropped every post-migration usage row. Test pins 23 and adds a 23-field classify case. 979 passed across every test file this branch touches + hermes_state + session recovery suites. Apollo merge pass 2026-09-23.
|
🤖 merged-by: apollo · lane: fr-pause-0922 · gate: BYPASS: FR PAUSED by Ace ruling 2026-09-22 (state/fleetreview-pause-20260922.md); - approved · why: rebased onto main by Apollo; Windows red = known desktop /progress flake, rerun; CI green; Apollo merge pass 2026-09-22 |
|
🤖 merged-by: apollo · lane: fr-pause-0922 · gate: BYPASS: FR PAUSED by Ace ruling 2026-09-22 (state/fleetreview-pause-20260922.md); t_083ad45e stack root, argus-reviewed round 2 APPROVED · why: rebased onto main by Apollo 2026-09-23; 233 branch tests + CI green; Windows red was the known /progress flake |
…tals and the Langfuse export #787 made UNKNOWN a first-class state for the Blackbox PER-TURN contract. Three other consumers still collapsed an unmeasured term into a measured zero. Bridge claude-bpx #186 is merged and fleet-deployed, so these paths receive unknown-input shapes today. Reproduced all three on the tree first: update_token_counts REJECTS unknown flags: unexpected keyword argument record_auxiliary_usage REJECTS unknown flags: unexpected keyword argument cumulative counters carry NO unknown provenance tui cache_hit_pct = 100% over prompt_total=100 (the unknown turn added 0) langfuse input-only -> ({'input': 0, 'output': 50}, {}) 1. PERSISTED SCHEMA. sessions and session_model_usage gain the five USAGE_UNKNOWN_FIELDS discriminators (+ last_turn_* snapshot mirrors), additive with DEFAULT 0 so every legacy row and every provider that never speaks this dialect reads back "measured". The int counters are untouched, so no arithmetic consumer changes. Cumulative flags are ABSORBING (MAX in SQL): one unmeasured call latches the session term, which is why a last-call-only guard is wrong here. Classified into _TOKEN_DELTA_ROUTE_FIELDS so deltas with differing unknown state never coalesce. Both session_model_usage table-rebuild paths (PK heal, v21->v22) copy the flags only when the source table has them, so a rebuild cannot silently reset a latched unknown. record_aux_usage no longer drops an all-zero unmeasured aux call as "empty". 2. CUMULATIVE CLI/TUI TOTALS. The loop latches per-term provenance onto session_*_unknown beside the session_*_tokens commit. Display reuses the Blackbox lane's rule (format_token_count / prompt_tokens_unknown) via new session_usage_unknown_flags / session_total_tokens_unknown helpers -- not a forked cumulative rule. /usage, both status-bar Sigma renderers, the CLI cache-hit segment and the TUI payload now say "unknown" or emit no ratio instead of a fabricated percentage. 3. LANGFUSE EXPORT. An unmeasured term is OMITTED from usage_details (absent, never 0) and declared on the observation metadata via _unknown_usage_details(); cost_details is empty for an unpriceable turn so Langfuse cannot read a partial subtotal as authoritative. The post_api_request summary-dict reconstruction was a second flag-loss site and now carries the discriminators. The 2-tuple return contract of _canonical_usage_and_cost / _usage_and_cost is unchanged. Verified (repo venv): - tests/agent/test_usage_unknown_persisted_export.py -- 71 pins, the five literal bridge payloads (input-only / cache-only / wholly-unavailable / output-only / measured-zero) through EACH of the three paths, plus a legacy-DB migration arm built by stripping the new columns out of the real SCHEMA_SQL. Cumulative and display arms source-LIFT the shipped blocks by AST anchor rather than re-implementing them. - scripts/run_tests.sh over the contract/consumer/persisted suites + tests/hermes_state + async token accounting + langfuse + last-turn persistence + turn-usage accumulator + blackbox: "54 files, 682 tests passed, 0 failed, 1 skipped". - 10-arm mutation matrix, each restored byte-exact, fake kills (SyntaxError/ NameError/zero-collected) rejected: 10/10 RED, control 131 passed before and after. Arms: SessionDB counter flags dropped; absorbing -> last-wins; aux ledger flags dropped; aux producer drops the unmeasured call; cumulative provenance not absorbing; CLI cache-ratio guard removed; /usage card prints numbers; TUI payload fabricates the ratio; Langfuse exports 0 for unknown; Langfuse summary-dict drops flags. - ui-tui: npm run typecheck clean. npm test is 61 files / 6 tests red on this machine, IDENTICAL on the base commit with the change stashed (textInputReturnAction + an unbuilt shared hermes-ink) -- pre-existing, not from this diff. Scope: this card only. The Blackbox per-turn contract certified under t_083ad45e / #787 is untouched.
…apshot flags, append-only schema Re-stacked hermes-agent#797 onto #787 (rebased onto main 2026-09-23). Two of its six commits were already on #787 under the same subjects and were dropped; the other four replayed with semantic conflicts. Fixes on top, each measured red before: 1. Duplicate absorbing latch: #787's "latch the flag over the aggregate" and #797's "carry UNKNOWN through cumulative totals" each added the same session_* latch loop to agent/conversation_loop.py. Kept #787's (beside the increments); removed #797's copy. 25 AST-seam tests ("shipped cumulative-usage seam moved") go green. 2. last_turn_*_unknown flags now move WITH the snapshot they qualify: when #787's _last_turn_snapshot_kwargs withholds an unmeasured call's split (None -> COALESCE keeps the prior real split), the flags are withheld too. Before: the persisted card rendered "Total: unknown" over a real 4,120 from the previous turn. 3. Schema columns APPENDED, not inserted mid-table. #797 put its 10 sessions flags before git_branch and its 5 session_model_usage flags before estimated_cost_usd. Migrated DBs get them via ALTER ADD (end of table), so fresh and migrated DBs had different physical orders, and the positional lost-and-found recovery mapper misplaced historical fields (title landed in the wrong column). Fresh-DB column prefix is now byte-identical to main's (title back at index 43). 4. Recovery classifier accepts session_model_usage records of 18 (pre-migration) AND 23 fields; it previously dropped every post-migration usage row. Test pins 23 and adds a 23-field classify case. 979 passed across every test file this branch touches + hermes_state + session recovery suites. Apollo merge pass 2026-09-23.
…tals and the Langfuse export (#797) * fix(usage): carry UNKNOWN through the persisted schema, cumulative totals and the Langfuse export #787 made UNKNOWN a first-class state for the Blackbox PER-TURN contract. Three other consumers still collapsed an unmeasured term into a measured zero. Bridge claude-bpx #186 is merged and fleet-deployed, so these paths receive unknown-input shapes today. Reproduced all three on the tree first: update_token_counts REJECTS unknown flags: unexpected keyword argument record_auxiliary_usage REJECTS unknown flags: unexpected keyword argument cumulative counters carry NO unknown provenance tui cache_hit_pct = 100% over prompt_total=100 (the unknown turn added 0) langfuse input-only -> ({'input': 0, 'output': 50}, {}) 1. PERSISTED SCHEMA. sessions and session_model_usage gain the five USAGE_UNKNOWN_FIELDS discriminators (+ last_turn_* snapshot mirrors), additive with DEFAULT 0 so every legacy row and every provider that never speaks this dialect reads back "measured". The int counters are untouched, so no arithmetic consumer changes. Cumulative flags are ABSORBING (MAX in SQL): one unmeasured call latches the session term, which is why a last-call-only guard is wrong here. Classified into _TOKEN_DELTA_ROUTE_FIELDS so deltas with differing unknown state never coalesce. Both session_model_usage table-rebuild paths (PK heal, v21->v22) copy the flags only when the source table has them, so a rebuild cannot silently reset a latched unknown. record_aux_usage no longer drops an all-zero unmeasured aux call as "empty". 2. CUMULATIVE CLI/TUI TOTALS. The loop latches per-term provenance onto session_*_unknown beside the session_*_tokens commit. Display reuses the Blackbox lane's rule (format_token_count / prompt_tokens_unknown) via new session_usage_unknown_flags / session_total_tokens_unknown helpers -- not a forked cumulative rule. /usage, both status-bar Sigma renderers, the CLI cache-hit segment and the TUI payload now say "unknown" or emit no ratio instead of a fabricated percentage. 3. LANGFUSE EXPORT. An unmeasured term is OMITTED from usage_details (absent, never 0) and declared on the observation metadata via _unknown_usage_details(); cost_details is empty for an unpriceable turn so Langfuse cannot read a partial subtotal as authoritative. The post_api_request summary-dict reconstruction was a second flag-loss site and now carries the discriminators. The 2-tuple return contract of _canonical_usage_and_cost / _usage_and_cost is unchanged. Verified (repo venv): - tests/agent/test_usage_unknown_persisted_export.py -- 71 pins, the five literal bridge payloads (input-only / cache-only / wholly-unavailable / output-only / measured-zero) through EACH of the three paths, plus a legacy-DB migration arm built by stripping the new columns out of the real SCHEMA_SQL. Cumulative and display arms source-LIFT the shipped blocks by AST anchor rather than re-implementing them. - scripts/run_tests.sh over the contract/consumer/persisted suites + tests/hermes_state + async token accounting + langfuse + last-turn persistence + turn-usage accumulator + blackbox: "54 files, 682 tests passed, 0 failed, 1 skipped". - 10-arm mutation matrix, each restored byte-exact, fake kills (SyntaxError/ NameError/zero-collected) rejected: 10/10 RED, control 131 passed before and after. Arms: SessionDB counter flags dropped; absorbing -> last-wins; aux ledger flags dropped; aux producer drops the unmeasured call; cumulative provenance not absorbing; CLI cache-ratio guard removed; /usage card prints numbers; TUI payload fabricates the ratio; Langfuse exports 0 for unknown; Langfuse summary-dict drops flags. - ui-tui: npm run typecheck clean. npm test is 61 files / 6 tests red on this machine, IDENTICAL on the base commit with the change stashed (textInputReturnAction + an unbuilt shared hermes-ink) -- pre-existing, not from this diff. Scope: this card only. The Blackbox per-turn contract certified under t_083ad45e / #787 is untouched. * fix(usage): gate cumulative cache ratios at the CLI render seam Run real snapshot and renderer regression tests for five bridge shapes. Verified 56 files: 743 passed, 0 failed, 1 skipped; producer and renderer guard mutants both fail. * fix(usage): preserve unknown in remaining cumulative CLI totals CLASS-SWEEP found /status raw total and zero-counter Sigma omission. Reproduced /status failure, then verified 3 focused tests pass; prior broad run passed 745 with one test fixture import error subsequently fixed. * test(usage): cover null container persistence and export controls * fix(usage): port #797 onto the rebased #787 -- one latch, coherent snapshot flags, append-only schema Re-stacked hermes-agent#797 onto #787 (rebased onto main 2026-09-23). Two of its six commits were already on #787 under the same subjects and were dropped; the other four replayed with semantic conflicts. Fixes on top, each measured red before: 1. Duplicate absorbing latch: #787's "latch the flag over the aggregate" and #797's "carry UNKNOWN through cumulative totals" each added the same session_* latch loop to agent/conversation_loop.py. Kept #787's (beside the increments); removed #797's copy. 25 AST-seam tests ("shipped cumulative-usage seam moved") go green. 2. last_turn_*_unknown flags now move WITH the snapshot they qualify: when #787's _last_turn_snapshot_kwargs withholds an unmeasured call's split (None -> COALESCE keeps the prior real split), the flags are withheld too. Before: the persisted card rendered "Total: unknown" over a real 4,120 from the previous turn. 3. Schema columns APPENDED, not inserted mid-table. #797 put its 10 sessions flags before git_branch and its 5 session_model_usage flags before estimated_cost_usd. Migrated DBs get them via ALTER ADD (end of table), so fresh and migrated DBs had different physical orders, and the positional lost-and-found recovery mapper misplaced historical fields (title landed in the wrong column). Fresh-DB column prefix is now byte-identical to main's (title back at index 43). 4. Recovery classifier accepts session_model_usage records of 18 (pre-migration) AND 23 fields; it previously dropped every post-migration usage row. Test pins 23 and adds a 23-field classify case. 979 passed across every test file this branch touches + hermes_state + session recovery suites. Apollo merge pass 2026-09-23. --------- Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
Re-port #779 onto current main without reverting #764 validation or #787 usage accounting. Carry fixed kind-specific labels and metadata-only history events; share the dropped-call retry budget and strip ephemeral correction pairs at finalization. Verified: 160 confab/CLI/gateway tests, 42 dropped-call and usage tests, 33 finalizer/status tests; git diff --check.
Re-port #779 onto current main without reverting #764 validation or #787 usage accounting. Carry fixed kind-specific labels and metadata-only history events; share the dropped-call retry budget and strip ephemeral correction pairs at finalization. Verified: 160 confab/CLI/gateway tests, 42 dropped-call and usage tests, 33 finalizer/status tests; git diff --check.
Why
The bridge in
claude-bpx#186 now reports an honest unknown when its CLI transcript reconciliation misses (~40% of parallel-batch turns, measured):completion_tokens/total_tokenscome backnullalongside an explicitusage.output_tokens_unavailable: truediscriminator.The consumer in this repo coerced that null to
0—normalize_usage -> _usage_count -> _to_intisint(value or 0)— and ignored the discriminator. So a bridge-first deploy would have written measured-looking zeros into the usage ledger,/usage, blackbox and pricing.out=0is a harder misread than the undercount it replaced, because it reads as a dead round-trip.Reproduced before fixing, against the literal wire payload:
This PR is the consumer-first gate: it lands before #186 merges.
What
UNKNOWN becomes a first-class state, end to end:
CanonicalUsageoutput_tokens_unknown(+total_tokens_unknown). Ints stay ints so arithmetic consumers are untouched; the flag is the discriminator.__add__makes unknown absorbing.normalize_usagemodel_extra, where an unrecognised wire key actually lands) or a present output counter that isNone. An absent key is NOT an unknown.estimate_usage_coststatus="unknown",amount_usd=None) instead of pricing a missing output term as $0.TurnRecord→ newoutput_tokens_unknowncolumn (guarded additive migration,DEFAULT 0). Thepriced_zerofast path and the reprice sweep both exclude unknown rows.usage_pricing.format_token_count/UNKNOWN_TOKENS_LABEL. The alert card and the/usage+/contextlast-turn card delegate to it (keeping their own magnitude formatting). The API-call log saysout=unknown.Backward compatible. Providers that never send the discriminator keep integers — pinned for anthropic, codex_responses and chat_completions shapes.
Verification
Contract test —
tests/agent/test_usage_unknown_contract.py, 15 pins against the bridge's exact wire fixture (copied from bpxbridge/test/unit-usage-unknown-sink.test.js/docs/SPEC-parallel-stop-usage-undercount.md):null+unavailable -> unknown; a real integer -> unchanged; a missing usage object -> unknown-free zeros, not a false unknown. 15 passed.7-arm mutation matrix, each mutation applied in place and restored:
or 0coercion (drop the discriminator)turn_finalizerrollup hardcodedFalseControl 15/15 green before and after; tree restored, no arm vacuous.
Regression (project venv, sandboxed
HOME):test_usage_pricing{,_linearity,_vendor_fallback},test_meta_usage_cache_reporting,test_usage_anchor,tests/blackbox/tests/plugins/blackbox/,test_last_turn_usage_persistence,test_empty_response_guardturn_finalizerx5, MoA blackbox/cost-slot,turn_usage_accumulator, account/billing usageE2E through the real path, literal bridge payload:
Measured control (same script, measured payload):
out=118,$0.00222,output_tokens_unknown=0, both renderers show118.Notes
NousResearch) has thenormalize_usagesurface, but the blackbox turn ledger and these renderers are fork-side, so this lands fork-first. Thenormalize_usage/CanonicalUsagehalf is upstreamable as a follow-up if wanted.deploy.shadvances the runtime tree.