Skip to content

fix(usage): treat provider-declared null usage as UNKNOWN, never a measured 0 - #787

Merged
Kyzcreig merged 16 commits into
mainfrom
fix/usage-unknown-end-to-end
Sep 23, 2026
Merged

Kyzcreig merged 16 commits into
mainfrom
fix/usage-unknown-end-to-end

Conversation

@Kyzcreig

Copy link
Copy Markdown
Collaborator

Why

The bridge in claude-bpx #186 now reports an honest unknown when its CLI transcript reconciliation misses (~40% of parallel-batch turns, measured): completion_tokens / total_tokens come back null alongside an explicit usage.output_tokens_unavailable: true discriminator.

The consumer in this repo coerced that null to 0 — normalize_usage -> _usage_count -> _to_int is int(value or 0) — and ignored the discriminator. So a bridge-first deploy would have written measured-looking zeros into the usage ledger, /usage, blackbox and pricing. out=0 is a harder misread than the undercount it replaced, because it reads as a dead round-trip.

Reproduced before fixing, against the literal wire payload:

UNKNOWN wire -> CanonicalUsage(input_tokens=150, output_tokens=0, ...)
  output_tokens = 0   total_tokens = 150

This PR is the consumer-first gate: it lands before #186 merges.

What

UNKNOWN becomes a first-class state, end to end:

layer change
CanonicalUsage new output_tokens_unknown (+ total_tokens_unknown). Ints stay ints so arithmetic consumers are untouched; the flag is the discriminator. __add__ makes unknown absorbing.
normalize_usage shape-agnostic detection: the discriminator flag (incl. via the OpenAI client's model_extra, where an unrecognised wire key actually lands) or a present output counter that is None. An absent key is NOT an unknown.
estimate_usage_cost refuses an unknown turn (status="unknown", amount_usd=None) instead of pricing a missing output term as $0.
blackbox flag rides per-call → per-turn (absorbing) → TurnRecord → new output_tokens_unknown column (guarded additive migration, DEFAULT 0). The priced_zero fast path and the reprice sweep both exclude unknown rows.
renderers display rule single-sourced in usage_pricing.format_token_count / UNKNOWN_TOKENS_LABEL. The alert card and the /usage+/context last-turn card delegate to it (keeping their own magnitude formatting). The API-call log says out=unknown.

Backward compatible. Providers that never send the discriminator keep integers — pinned for anthropic, codex_responses and chat_completions shapes.

Verification

Contract test — tests/agent/test_usage_unknown_contract.py, 15 pins against the bridge's exact wire fixture (copied from bpx bridge/test/unit-usage-unknown-sink.test.js / docs/SPEC-parallel-stop-usage-undercount.md): null+unavailable -> unknown; a real integer -> unchanged; a missing usage object -> unknown-free zeros, not a false unknown. 15 passed.

7-arm mutation matrix, each mutation applied in place and restored:

arm result
reinstate the or 0 coercion (drop the discriminator) 3 RED
drop the pricing refusal 2 RED
store stops persisting the flag 1 RED
drop the shared display rule 1 RED
turn_finalizer rollup hardcoded False 1 RED
blackbox ingest drops the flag 1 RED
widen the detector so ABSENT counts as unknown 3 RED

Control 15/15 green before and after; tree restored, no arm vacuous.

Regression (project venv, sandboxed HOME):

  • 153 passed — test_usage_pricing{,_linearity,_vendor_fallback}, test_meta_usage_cache_reporting, test_usage_anchor, tests/blackbox/
  • 180 passed — tests/plugins/blackbox/, test_last_turn_usage_persistence, test_empty_response_guard
  • 46 passed — turn_finalizer x5, MoA blackbox/cost-slot, turn_usage_accumulator, account/billing usage

E2E through the real path, literal bridge payload:

1. normalize   -> UNKNOWN
2. price       -> status=unknown amount=None
3. ledger      -> output_tokens=0 output_tokens_unknown=1
4. renderer A  -> '• Tokens out: unknown'
5. renderer B  -> '• Tokens: 150 in + unknown out'

Measured control (same script, measured payload): out=118, $0.00222, output_tokens_unknown=0, both renderers show 118.

Notes

  • Upstream (NousResearch) has the normalize_usage surface, but the blackbox turn ledger and these renderers are fork-side, so this lands fork-first. The normalize_usage/CanonicalUsage half is upstreamable as a follow-up if wanted.
  • No live gateway is affected until deploy.sh advances the runtime tree.

@Kyzcreig
Kyzcreig added this pull request to the merge queue Sep 21, 2026
@github-merge-queue
github-merge-queue Bot removed this pull request from the merge queue due to failed status checks Sep 21, 2026
Kyzcreig added a commit that referenced this pull request Sep 21, 2026
…nd every turn display

Bridge claude-bpx #186 emits null prompt_tokens/total_tokens with
prompt_tokens_unavailable/total_tokens_unavailable/unavailable discriminators.
normalize_usage previously collapsed that to input_tokens=0 with no unknown
flag, so unknown input was priced and displayed as a measured zero.

Extends #787's output-only UNKNOWN contract to input, cache and whole-usage
unknowns across normalize_usage/CanonicalUsage, pricing, the Blackbox turn
ledger/rollups/cards, and the shared display rule. Measured zeros stay
measured; absent-key behavior for legacy providers is unchanged.

Also preserves the discriminators through both MoA physical-call
serializations (aggregator + advisor), which reconstructed pricing dicts from
numeric fields only and let Blackbox price an unknown input as $0.00075, and
routes the console cache-stat branch, the adjacent verbose token log and the
gateway thin last-turn fallback through the shared unknown rule.

Verified (repo venv):
  404 passed in 73.16s across usage contract/consumers, usage_pricing,
  billing, blackbox, turn accumulator, last-turn persistence, meta cache
  reporting, usage anchor, context-engine and MoA loop suites.
  Acceptance repro now: input_tokens_unknown=True, total_tokens_unknown=True,
  direct cost None, MoA blackbox cost (None, 'unknown', all-None components),
  console prints "Cache: unknown". Measured-zero control stays measured
  (all unknown flags False, cost 0.00).
  5 mutation arms RED with exact-byte restore: input detection 3F,
  pricing refusal 25F, MoA aggregator flags 4F, MoA advisor flags 8F,
  console cache guard 3F.

Scope: Blackbox per-turn contract + the human-facing displays it feeds.
Persisted/export-schema flag loss (SessionDB/aux-ledger counters, cumulative
CLI/TUI totals, Langfuse canonical export) is deferred to t_25f50547.
@Kyzcreig

Copy link
Copy Markdown
Collaborator Author

🔴 HOLD — do not merge until kanban review round 3 on t_083ad45e returns APPROVED (reviewer daedalus-opus; implementer round-3 head f78eae2, 3 CI-caught defects fixed, 7 mutation arms RED). CI is green; that is not the gate here. Hold lives on the PR on purpose: the bridge-side hold for #186 lived only on a card and a CI-green lane merged + fleet-deployed it 18:40 while this consumer was still open. Apollo lifts this comment when the review verdict is in. Follow-up for the persisted/export schema paths is t_25f50547, not this PR.

@Kyzcreig Kyzcreig added the do-not-merge Hold: a lane must not merge this PR (Apollo lifts it) label Sep 21, 2026
@Kyzcreig

Copy link
Copy Markdown
Collaborator Author

Do not merge as-is — kanban review round 3 on t_083ad45e (reviewer daedalus-opus) has not returned a verdict on head f78eae2. CI green is not the gate. do-not-merge label applied; Apollo removes both when the review lands.

@Kyzcreig Kyzcreig removed the do-not-merge Hold: a lane must not merge this PR (Apollo lifts it) label Sep 21, 2026
@Kyzcreig

Copy link
Copy Markdown
Collaborator Author

Hold lifted — kanban round-2 review on t_083ad45e APPROVED (daedalus-opus run 2264, execution lens, independent mutation set 5/6 RED with the survivor adjudicated). Merging. Persisted/export schema paths are #797 (t_25f50547), stacked on this branch; it retargets to main on merge.

@Kyzcreig
Kyzcreig added this pull request to the merge queue Sep 21, 2026
@Kyzcreig
Kyzcreig removed this pull request from the merge queue due to a manual request Sep 21, 2026
@Kyzcreig Kyzcreig added the do-not-merge Hold: a lane must not merge this PR (Apollo lifts it) label Sep 21, 2026
@Kyzcreig

Copy link
Copy Markdown
Collaborator Author

Do not merge as-is — hold RE-APPLIED (Apollo, 22:30 PT). FleetReview's run on f78eae2 produced 8 findings incl. 3 P1 and then died before posting (ACE-AI router closed it 'ensemble member failure after self-heal'; record exists only in its shadow-log). P1s: (1) agent/usage_pricing.py — a present-but-null cache-details CONTAINER marks the whole turn permanently unpriceable; (2) tests/agent/test_usage_unknown_contract.py — null prompt_tokens_details container pinned as UNKNOWN collapses a fully MEASURED input; (3) tests/agent/test_usage_unknown_consumers.py — the thin-fallback 'shipped consumer' test feeds a shape no producer emits, so /usage still renders an unmeasured turn as 0. P2: unknown check runs before the subscription_included route; empty_response_guard reads unmeasured 0 as measured and cuts the retry budget; AST anchors brittle on unrelated refactors. Rework carded; fixes go on this branch. Nobody arms --auto here (queue repo: --auto = enqueue = merge). I enqueued it by mistake at 05:14:33Z via bare gh pr merge and dequeued at 05:16:02Z — nothing landed.

Kyzcreig added a commit that referenced this pull request Sep 21, 2026
…tals and the Langfuse export

#787 made UNKNOWN a first-class state for the Blackbox PER-TURN contract.
Three other consumers still collapsed an unmeasured term into a measured
zero. Bridge claude-bpx #186 is merged and fleet-deployed, so these paths
receive unknown-input shapes today. Reproduced all three on the tree first:

  update_token_counts REJECTS unknown flags: unexpected keyword argument
  record_auxiliary_usage REJECTS unknown flags: unexpected keyword argument
  cumulative counters carry NO unknown provenance
  tui cache_hit_pct = 100% over prompt_total=100 (the unknown turn added 0)
  langfuse input-only -> ({'input': 0, 'output': 50}, {})

1. PERSISTED SCHEMA. sessions and session_model_usage gain the five
   USAGE_UNKNOWN_FIELDS discriminators (+ last_turn_* snapshot mirrors),
   additive with DEFAULT 0 so every legacy row and every provider that
   never speaks this dialect reads back "measured". The int counters are
   untouched, so no arithmetic consumer changes. Cumulative flags are
   ABSORBING (MAX in SQL): one unmeasured call latches the session term,
   which is why a last-call-only guard is wrong here. Classified into
   _TOKEN_DELTA_ROUTE_FIELDS so deltas with differing unknown state never
   coalesce. Both session_model_usage table-rebuild paths (PK heal, v21->v22)
   copy the flags only when the source table has them, so a rebuild cannot
   silently reset a latched unknown. record_aux_usage no longer drops an
   all-zero unmeasured aux call as "empty".

2. CUMULATIVE CLI/TUI TOTALS. The loop latches per-term provenance onto
   session_*_unknown beside the session_*_tokens commit. Display reuses the
   Blackbox lane's rule (format_token_count / prompt_tokens_unknown) via new
   session_usage_unknown_flags / session_total_tokens_unknown helpers -- not
   a forked cumulative rule. /usage, both status-bar Sigma renderers, the CLI
   cache-hit segment and the TUI payload now say "unknown" or emit no ratio
   instead of a fabricated percentage.

3. LANGFUSE EXPORT. An unmeasured term is OMITTED from usage_details
   (absent, never 0) and declared on the observation metadata via
   _unknown_usage_details(); cost_details is empty for an unpriceable turn
   so Langfuse cannot read a partial subtotal as authoritative. The
   post_api_request summary-dict reconstruction was a second flag-loss site
   and now carries the discriminators. The 2-tuple return contract of
   _canonical_usage_and_cost / _usage_and_cost is unchanged.

Verified (repo venv):
- tests/agent/test_usage_unknown_persisted_export.py -- 71 pins, the five
  literal bridge payloads (input-only / cache-only / wholly-unavailable /
  output-only / measured-zero) through EACH of the three paths, plus a
  legacy-DB migration arm built by stripping the new columns out of the real
  SCHEMA_SQL. Cumulative and display arms source-LIFT the shipped blocks by
  AST anchor rather than re-implementing them.
- scripts/run_tests.sh over the contract/consumer/persisted suites +
  tests/hermes_state + async token accounting + langfuse + last-turn
  persistence + turn-usage accumulator + blackbox:
  "54 files, 682 tests passed, 0 failed, 1 skipped".
- 10-arm mutation matrix, each restored byte-exact, fake kills (SyntaxError/
  NameError/zero-collected) rejected: 10/10 RED, control 131 passed before
  and after. Arms: SessionDB counter flags dropped; absorbing -> last-wins;
  aux ledger flags dropped; aux producer drops the unmeasured call;
  cumulative provenance not absorbing; CLI cache-ratio guard removed; /usage
  card prints numbers; TUI payload fabricates the ratio; Langfuse exports 0
  for unknown; Langfuse summary-dict drops flags.
- ui-tui: npm run typecheck clean. npm test is 61 files / 6 tests red on
  this machine, IDENTICAL on the base commit with the change stashed
  (textInputReturnAction + an unbuilt shared hermes-ink) -- pre-existing,
  not from this diff.

Scope: this card only. The Blackbox per-turn contract certified under
t_083ad45e / #787 is untouched.
Kyzcreig added a commit that referenced this pull request Sep 21, 2026
A worker that needs a sibling PR landed before it can continue blocks with a
reason like "merge PR #787 then unblock me". Nothing in the dispatcher ever
re-read that reason, so the card sat blocked until a human noticed. Observed on
a live board: six needs_input cards held up to 12 hours on gates whose PRs had
already merged, some 12+ hours earlier.

New hermes_cli/kanban_pr_gate.py runs inside the dispatcher tick, before
recompute_ready so a satisfied card becomes spawnable in the SAME tick:

  * parse the last block-family event's reason for PR references
    (owner/repo#N, a full .../pull/N URL, and bare #N / pull/N resolved
    against a repo context: the card's workspace remote, else the first
    owner/repo mentioned in the body),
  * resolve each through `gh pr view`, bounded at 30 lookups per tick, one
    call per unique PR, cached (MERGED/CLOSED for the process lifetime since
    a merge is irreversible; OPEN for 300 s),
  * every referenced PR MERGED -> unblock + comment
    "gate satisfied: <repo>#N merged <sha8> at <ts>" + one gate_auto_resolved
    event so tooling can count it,
  * any still OPEN -> hold; CLOSED-unmerged -> one advisory comment and never
    an unblock; any lookup failure -> no action plus one WARN.

Fail-safe choices, each with a test: scope fenced to needs_input / capability /
dependency block kinds (transient means "may clear on its own", not "waits on a
named object"); a reason naming no PR is never touched and burns zero lookup
budget; a bare #N whose repo context is ambiguous is dropped rather than
guessed; a transient gh failure is not cached, so it cannot blind the next tick
for the whole TTL; the closed-unmerged advisory posts once rather than every
tick; and the whole pass is wrapped fail-open, because a diagnostic that can
brick a dispatch tick is worse than the stale-block class it closes.

DispatchResult gains gate_auto_resolved / gate_closed_unmerged, surfaced in
both the text and --json dispatch output.

Verified, not claimed:
  * tests/hermes_cli/test_kanban_pr_gate.py -- 33 passed.
  * RED-proof of the parser: the multi-reference ordering case failed first
    (refs returned qualified-before-bare instead of in source order); the
    implementation was fixed, not the test.
  * RED-proof of the dispatcher wiring: deleting the single
    _reevaluate_pr_gates(conn, result) call makes
    test_dispatch_tick_resolves_a_satisfied_gate fail with
    `assert 't_...' in []`, so the integration test gates the hook itself and
    not merely the module.
Kyzcreig added a commit that referenced this pull request Sep 21, 2026
A worker that needs a sibling PR landed before it can continue blocks with a
reason like "merge PR #787 then unblock me". Nothing in the dispatcher ever
re-read that reason, so the card sat blocked until a human noticed. Observed on
a live board: six needs_input cards held up to 12 hours on gates whose PRs had
already merged, some 12+ hours earlier.

New hermes_cli/kanban_pr_gate.py runs inside the dispatcher tick, before
recompute_ready so a satisfied card becomes spawnable in the SAME tick:

  * parse the last block-family event's reason for PR references
    (owner/repo#N, a full .../pull/N URL, and bare #N / pull/N resolved
    against a repo context: the card's workspace remote, else the first
    owner/repo mentioned in the body),
  * resolve each through `gh pr view`, bounded at 30 lookups per tick, one
    call per unique PR, cached (MERGED/CLOSED for the process lifetime since
    a merge is irreversible; OPEN for 300 s),
  * every referenced PR MERGED -> unblock + comment
    "gate satisfied: <repo>#N merged <sha8> at <ts>" + one gate_auto_resolved
    event so tooling can count it,
  * any still OPEN -> hold; CLOSED-unmerged -> one advisory comment and never
    an unblock; any lookup failure -> no action plus one WARN.

Fail-safe choices, each with a test: scope fenced to needs_input / capability /
dependency block kinds (transient means "may clear on its own", not "waits on a
named object"); a reason naming no PR is never touched and burns zero lookup
budget; a bare #N whose repo context is ambiguous is dropped rather than
guessed; a transient gh failure is not cached, so it cannot blind the next tick
for the whole TTL; the closed-unmerged advisory posts once rather than every
tick; and the whole pass is wrapped fail-open, because a diagnostic that can
brick a dispatch tick is worse than the stale-block class it closes.

DispatchResult gains gate_auto_resolved / gate_closed_unmerged, surfaced in
both the text and --json dispatch output.

Verified, not claimed:
  * tests/hermes_cli/test_kanban_pr_gate.py -- 33 passed.
  * RED-proof of the parser: the multi-reference ordering case failed first
    (refs returned qualified-before-bare instead of in source order); the
    implementation was fixed, not the test.
  * RED-proof of the dispatcher wiring: deleting the single
    _reevaluate_pr_gates(conn, result) call makes
    test_dispatch_tick_resolves_a_satisfied_gate fail with
    `assert 't_...' in []`, so the integration test gates the hook itself and
    not merely the module.
Kyzcreig added a commit that referenced this pull request Sep 21, 2026
A worker that needs a sibling PR landed before it can continue blocks with a
reason like "merge PR #787 then unblock me". Nothing in the dispatcher ever
re-read that reason, so the card sat blocked until a human noticed. Observed on
a live board: six needs_input cards held up to 12 hours on gates whose PRs had
already merged, some 12+ hours earlier.

New hermes_cli/kanban_pr_gate.py runs inside the dispatcher tick, before
recompute_ready so a satisfied card becomes spawnable in the SAME tick:

  * parse the last block-family event's reason for PR references
    (owner/repo#N, a full .../pull/N URL, and bare #N / pull/N resolved
    against a repo context: the card's workspace remote, else the first
    owner/repo mentioned in the body),
  * resolve each through `gh pr view`, bounded at 30 lookups per tick, one
    call per unique PR, cached (MERGED/CLOSED for the process lifetime since
    a merge is irreversible; OPEN for 300 s),
  * every referenced PR MERGED -> unblock + comment
    "gate satisfied: <repo>#N merged <sha8> at <ts>" + one gate_auto_resolved
    event so tooling can count it,
  * any still OPEN -> hold; CLOSED-unmerged -> one advisory comment and never
    an unblock; any lookup failure -> no action plus one WARN.

Fail-safe choices, each with a test: scope fenced to needs_input / capability /
dependency block kinds (transient means "may clear on its own", not "waits on a
named object"); a reason naming no PR is never touched and burns zero lookup
budget; a bare #N whose repo context is ambiguous is dropped rather than
guessed; a transient gh failure is not cached, so it cannot blind the next tick
for the whole TTL; the closed-unmerged advisory posts once rather than every
tick; and the whole pass is wrapped fail-open, because a diagnostic that can
brick a dispatch tick is worse than the stale-block class it closes.

DispatchResult gains gate_auto_resolved / gate_closed_unmerged, surfaced in
both the text and --json dispatch output.

Verified, not claimed:
  * tests/hermes_cli/test_kanban_pr_gate.py -- 33 passed.
  * RED-proof of the parser: the multi-reference ordering case failed first
    (refs returned qualified-before-bare instead of in source order); the
    implementation was fixed, not the test.
  * RED-proof of the dispatcher wiring: deleting the single
    _reevaluate_pr_gates(conn, result) call makes
    test_dispatch_tick_resolves_a_satisfied_gate fail with
    `assert 't_...' in []`, so the integration test gates the hook itself and
    not merely the module.
Kyzcreig added a commit that referenced this pull request Sep 21, 2026
…nd every turn display

Bridge claude-bpx #186 emits null prompt_tokens/total_tokens with
prompt_tokens_unavailable/total_tokens_unavailable/unavailable discriminators.
normalize_usage previously collapsed that to input_tokens=0 with no unknown
flag, so unknown input was priced and displayed as a measured zero.

Extends #787's output-only UNKNOWN contract to input, cache and whole-usage
unknowns across normalize_usage/CanonicalUsage, pricing, the Blackbox turn
ledger/rollups/cards, and the shared display rule. Measured zeros stay
measured; absent-key behavior for legacy providers is unchanged.

Also preserves the discriminators through both MoA physical-call
serializations (aggregator + advisor), which reconstructed pricing dicts from
numeric fields only and let Blackbox price an unknown input as $0.00075, and
routes the console cache-stat branch, the adjacent verbose token log and the
gateway thin last-turn fallback through the shared unknown rule.

Verified (repo venv):
  404 passed in 73.16s across usage contract/consumers, usage_pricing,
  billing, blackbox, turn accumulator, last-turn persistence, meta cache
  reporting, usage anchor, context-engine and MoA loop suites.
  Acceptance repro now: input_tokens_unknown=True, total_tokens_unknown=True,
  direct cost None, MoA blackbox cost (None, 'unknown', all-None components),
  console prints "Cache: unknown". Measured-zero control stays measured
  (all unknown flags False, cost 0.00).
  5 mutation arms RED with exact-byte restore: input detection 3F,
  pricing refusal 25F, MoA aggregator flags 4F, MoA advisor flags 8F,
  console cache guard 3F.

Scope: Blackbox per-turn contract + the human-facing displays it feeds.
Persisted/export-schema flag loss (SessionDB/aux-ledger counters, cumulative
CLI/TUI totals, Langfuse canonical export) is deferred to t_25f50547.
@Kyzcreig
Kyzcreig force-pushed the fix/usage-unknown-end-to-end branch from d5b3149 to 20dd114 Compare September 21, 2026 20:13
@Kyzcreig

Copy link
Copy Markdown
Collaborator Author

FleetReview

PARTIAL — ensemble escalated: judge transient failure

This review did not reach a trusted verdict, so it is not a gate pass and the findings below may be incomplete. They are posted so they can be read rather than lost in a terminal record.

Reviewed with 2 of 3 model families — openai unavailable.

Confidence: 1/5

Findings

  • P1 agent/conversation_loop.py:599 — _moa_session_cost_status misses route-unpriceable advisors, so the session lane still reports estimated for an incomplete MoA total
  • P2 agent/conversation_loop.py:4872 — Compressor gate uses the MoA-folded usage, so one unmeasured advisor discards a fully measured aggregator prompt reading
  • P1 agent/empty_response_guard.py:192 — Unknown-usage fail-open leaves BOTH empty-response guards inert, including the cost-aware one
  • P1 tests/agent/test_usage_unknown_r6_findings.py:408 — Resident /usage snapshot derives a CUMULATIVE total's unknown flags from one call; the new narrowness test locks that in
  • P3 tests/agent/test_usage_unknown_consumers.py:50 — test_console_cache_block_honors_unknown asserts a rule cache_stats_line deliberately breaks, and errors instead of failing
  • P1 plugins/blackbox/card.py:99 — Spending-alert card's • Context: line still renders an unmeasured final call's placeholder 0 as a measurement
  • P1 gateway/slash_commands.py:216 — _resident_thin_snapshot stamps the LAST CALL's unknown flags onto SESSION-CUMULATIVE counters
  • P1 plugins/blackbox/store.py:347 — Incomplete unknown latch
  • P1 agent/usage_pricing.py:2450 — Cache unknown false-positive
  • P1 hermes_state.py:9672 — Incomplete cost_status CASE
  • P3 tests/agent/test_usage_unknown_consumers.py:94 — test_verbose_log_honors_unknown has no unconditional assertion — vacuous for the all-zero wire
  • P2 tests/run_agent/test_usageless_response_accounting.py:91 — captured_turn_usage stub violates invoke_hook's List[Any] contract, forcing a swallowed TypeError on every turn
  • P1 agent/conversation_loop.py:576 — Session status clobber

FleetReview provenance · models: C=claude-code-opus-5, D=grok-4.6, G=grok-4.6 · cost: $56.89 · duration: 1h 28m 36s · rounds: 2 · files examined: 25

@Kyzcreig

Copy link
Copy Markdown
Collaborator Author

FleetReview

Reviewed with 2 of 3 model families — openai unavailable.

Confidence: 3/5

Findings

  • P1 agent/empty_response_guard.py:307 — Floor shown exact
  • P1 agent/conversation_loop.py:5293 — Stale cost_status persisted
  • P1 agent/conversation_loop.py:673 — Merge drops partial
  • P1 hermes_state.py:9672 — Status Ignores Retained Estimate
  • P2 agent/moa_loop.py:2200 — Failed/skipped MoA advisors now carry cost_usd=None and downgrade every such turn's session cost label to "partial"
  • P1 agent/codex_runtime.py:126 — Codex unknown latch and cost-status merge are wired into the one branch that can never need them
  • P1 agent/codex_runtime.py:207 — Codex unknown-latch is wired on the branch where it can never fire, and absent from the usage-less path where the counters really do lose a call
  • P1 plugins/blackbox/store.py:344 — Legacy unknown-latch drops its cost_status guard entirely when the column is absent, latching measured zeros

FleetReview provenance · models: C=claude-code-opus-5, D=grok-4.6, G=grok-4.6 · cost: $56.03 · duration: 43m 03s · rounds: 1 · files examined: 30

@Kyzcreig

Copy link
Copy Markdown
Collaborator Author

FleetReview

PARTIAL — ensemble escalated: judge transient failure

This review did not reach a trusted verdict, so it is not a gate pass and the findings below may be incomplete. They are posted so they can be read rather than lost in a terminal record.

Reviewed with 2 of 3 model families — openai unavailable.

Confidence: 1/5

Findings

  • P1 agent/codex_runtime.py:126 — Codex usage-less turn bypasses the new session unknown latch entirely
  • P2 agent/moa_loop.py:2211 — Failed MoA advisor now carries cost_usd=None, flipping the whole session's cost label to "partial"
  • P1 agent/conversation_loop.py:5293 — Worst-of session cost-status merge never reaches the persisted sessions row — DB heals back to a complete label
  • P2 agent/turn_finalizer.py:912 — last_call_prompt_unknown is derived from the MoA advisor-folded usage, not the final call's own prompt
  • P1 plugins/blackbox/store.py:327 — Migration now tolerates a DB missing token-count columns, but insert_turn still binds them unconditionally and swallows the failure
  • P2 plugins/blackbox/store.py:285 — CREATE TABLE omits four of the five unknown columns, so every brand-new DB takes the legacy-migration path
  • P2 tests/plugins/blackbox/test_last_call_renderer_class_sweep.py:209 — Unnecessary, unrestored importlib.reload of hermes_constants and blackbox.store leaks global state into the rest of the session
  • P2 tests/plugins/blackbox/test_api_calls_compat.py:124 — sqlite3 connection opened for the one-shot migration is committed but never closed
  • P1 agent/empty_response_guard.py:178 — Floor latch dropped
  • P1 hermes_state.py:9708 — SQL completeness overwrite
  • P1 hermes_state.py:9961 — Model status clobber
  • P1 agent/conversation_loop.py:673 — Unknown+partial merge
  • P1 hermes_state.py:9672 — Estimated null-keep wipe
  • P3 tests/plugins/blackbox/test_last_call_renderer_class_sweep.py:42 — _row() uses the get_turn/TurnRecord key shape, not the raw SELECT * shape render_last_turn_record sees in production
  • P3 tests/plugins/blackbox/test_last_call_renderer_class_sweep.py:30 — Unused sqlite3 and pytest imports
  • P1 agent/codex_runtime.py:207 — Codex session UNKNOWN latch is wired into the branch where it can never fire, and missing from the usage-less branch where it is always true

FleetReview provenance · models: C=claude-code-opus-5, D=grok-4.6, G=grok-4.6 · cost: $56.87 · duration: 38m 29s · rounds: 1 · files examined: 30

Kyzcreig and others added 7 commits September 23, 2026 08:18
…asured 0

claude-bpx #186 makes the bridge egress an honest unknown when its CLI
transcript reconciliation misses (~40% of parallel-batch turns, measured):
`completion_tokens`/`total_tokens` come back null with an explicit
`usage.output_tokens_unavailable: true` discriminator. The deployed consumer
coerced that null to 0 via `normalize_usage -> _usage_count -> _to_int`
(`int(value or 0)`) and ignored the discriminator, so a bridge-first deploy
would have written measured-looking zeros into the usage ledger, /usage,
blackbox and pricing. Reproduced before fixing: the literal wire payload
normalized to `output_tokens=0`, `total_tokens=150`.

UNKNOWN is now a first-class state end to end:

- CanonicalUsage gains `output_tokens_unknown` (+ `total_tokens_unknown`).
  The int fields stay ints so every arithmetic consumer is untouched; the
  flag is the discriminator. `__add__` makes unknown ABSORBING.
- normalize_usage detects it shape-agnostically: the discriminator flag
  (including via the OpenAI client's `model_extra`, where an unrecognised
  wire key actually lands), or a PRESENT output counter that is None.
  An ABSENT key is NOT an unknown — providers that never speak this dialect
  keep normalizing to integers.
- estimate_usage_cost REFUSES an unknown turn (status=unknown, amount None)
  rather than pricing a missing output term as $0.
- blackbox: the flag rides per-call -> per-turn (absorbing) -> TurnRecord ->
  a new `output_tokens_unknown` column (guarded additive migration, DEFAULT 0
  so every historical row reads back "measured"). The priced_zero fast path
  and the reprice sweep both exclude unknown rows.
- The display rule is single-sourced in `usage_pricing.format_token_count`
  (`UNKNOWN_TOKENS_LABEL`); the alert card and the /usage+/context last-turn
  card both delegate to it, keeping their own magnitude formatting. The
  `in=/out=` API-call log says `out=unknown` instead of `out=0`.

Verified:
- tests/agent/test_usage_unknown_contract.py — 15 pins against the bridge's
  exact wire fixture (copied from bpx `unit-usage-unknown-sink.test.js` /
  `SPEC-parallel-stop-usage-undercount.md`): null+unavailable -> unknown,
  a real integer -> unchanged, a missing usage object -> unknown-free zeros.
  15 passed.
- 7-arm mutation matrix, each restored after: reinstating the `or 0`
  coercion -> 3 RED; dropping the pricing refusal -> 2 RED; store not
  persisting -> 1 RED; dropping the shared display rule -> 1 RED;
  turn_finalizer rollup hardcoded False -> 1 RED; blackbox ingest dropping
  the flag -> 1 RED; widening the detector to treat ABSENT as unknown -> 3
  RED. Control 15/15 green before and after; tree restored.
- Regression: 153 passed (usage_pricing x3, meta_usage_cache, usage_anchor,
  tests/blackbox), 180 passed (tests/plugins/blackbox, last_turn_usage
  persistence, empty_response_guard), 46 passed (turn_finalizer x5, MoA
  blackbox/cost-slot, turn_usage_accumulator, account/billing usage).
- E2E through the real path with the literal bridge payload: normalize ->
  UNKNOWN, price -> status=unknown/None, ledger -> output_tokens_unknown=1,
  both renderers -> "unknown". Measured control: 118 out, $0.00222,
  output_tokens_unknown=0, both renderers -> "118".

# Conflicts:
#	plugins/blackbox/store.py
…nd every turn display

Bridge claude-bpx #186 emits null prompt_tokens/total_tokens with
prompt_tokens_unavailable/total_tokens_unavailable/unavailable discriminators.
normalize_usage previously collapsed that to input_tokens=0 with no unknown
flag, so unknown input was priced and displayed as a measured zero.

Extends #787's output-only UNKNOWN contract to input, cache and whole-usage
unknowns across normalize_usage/CanonicalUsage, pricing, the Blackbox turn
ledger/rollups/cards, and the shared display rule. Measured zeros stay
measured; absent-key behavior for legacy providers is unchanged.

Also preserves the discriminators through both MoA physical-call
serializations (aggregator + advisor), which reconstructed pricing dicts from
numeric fields only and let Blackbox price an unknown input as $0.00075, and
routes the console cache-stat branch, the adjacent verbose token log and the
gateway thin last-turn fallback through the shared unknown rule.

Verified (repo venv):
  404 passed in 73.16s across usage contract/consumers, usage_pricing,
  billing, blackbox, turn accumulator, last-turn persistence, meta cache
  reporting, usage anchor, context-engine and MoA loop suites.
  Acceptance repro now: input_tokens_unknown=True, total_tokens_unknown=True,
  direct cost None, MoA blackbox cost (None, 'unknown', all-None components),
  console prints "Cache: unknown". Measured-zero control stays measured
  (all unknown flags False, cost 0.00).
  5 mutation arms RED with exact-byte restore: input detection 3F,
  pricing refusal 25F, MoA aggregator flags 4F, MoA advisor flags 8F,
  console cache guard 3F.

Scope: Blackbox per-turn contract + the human-facing displays it feeds.
Persisted/export-schema flag loss (SessionDB/aux-ledger counters, cumulative
CLI/TUI totals, Langfuse canonical export) is deferred to t_25f50547.
…NOWN

CI slice 3/16 and 15/16 went red on tests/gateway/test_usage_command.py.
Root cause: routing these renderers through format_token_count sent the
MEASURED path through its default magnitude formatter too, so a measured
241,500 started rendering as "241.5k" and broke the card's vocabulary.

format_token_count already has a `formatter` parameter for exactly this --
"lets a caller keep its own magnitude formatting while still routing the
unknown case through here". Both introduced sites now pass it.

Swept the whole class (every format_token_count call this branch added,
`git diff 1b04723 -U0 -- agent/ plugins/ gateway/ | grep '^+.*format_token_count'`,
8 hits / 3 sites):
  - gateway/slash_commands.py thin last-turn card  -> FIXED (CI-reported)
  - agent/conversation_loop.py verbose token log   -> FIXED (same class, NOT
    reported by CI; was `{:,}` on base and had silently become "120k")
  - agent/conversation_loop.py console cache line  -> unaffected, unknown-only
    call (format_token_count(None, unknown=True)), no measured path

Both fixed sites are pinned by measured-control tests; dropping `formatter=`
at either site is RED.

Verified (repo venv):
  415 passed, 1 failed in 31.79s. The single failure,
  test_usage_command_prefers_dominant_persisted_route, is PRE-EXISTING: it
  fails identically (StopIteration) on base 1b04723 in a clean worktree.
  Mutation arms, all RED with exact-byte restore: input detection 3F,
  pricing refusal 25F, MoA aggregator 4F, MoA advisor 8F, console cache 3F,
  thin-card formatter 4F, verbose-log formatter 1F.
  Acceptance repro unchanged: cost None, (None, 'unknown'), "Cache: unknown";
  measured-zero control still measured at 0.00.
… snapshot

CI slice 15/16: test_build_moa_pricing_calls_appends_real_aggregator_after_advisors
compared the serialized aggregator call with `==` against a literal dict, so
the five additive UNKNOWN discriminators (all False for this fully-measured
usage) failed the equality.

Tightened rather than loosened, per AGENTS.md "behavior contracts over
snapshots": the test still pins attribution and every counter exactly, and
now additionally requires the discriminators to be PRESENT and False, and the
key set to be exactly the measured keys plus USAGE_UNKNOWN_FIELDS. Silently
dropping the flags from this serializer -- the defect that let an unknown
input price as $0.00075 -- now fails here too.

Also widened the local verify set. The previous list was hand-picked and
missed what CI caught; it is now derived from the symbols this branch
touches:
  grep -rlE '_build_moa_pricing_calls|format_token_count|normalize_usage|
  prompt_tokens_unknown|USAGE_UNKNOWN_FIELDS|compute_turn_cost|
  get_last_turn_usage|CanonicalUsage' tests/
which added test_moa_blackbox_attribution, test_bedrock_adapter,
test_context_engine_host_contract, test_empty_response_guard and
test_moa_observability_bridge.

Verified (repo venv, --basetemp isolated): 540 passed, 1 failed in 44.58s.
The single failure, test_usage_command_prefers_dominant_persisted_route,
reproduces identically (StopIteration) on base 1b04723 in a clean
worktree -- pre-existing, not from this branch.
FleetReview's run on f78eae2 produced 8 findings (3 P1) and died before
posting them. This closes the production half.

P1 (findings 1+8) — `"prompt_tokens_details": null` is the ordinary
serialization of an OpenAI-compatible server with no cache breakdown (any
encoder without exclude_none emits it; pydantic puts an explicitly-null
wire field in model_fields_set). Reading the null CONTAINER as "both cache
buckets unmeasured" escalated into input_tokens_unknown, so a turn with a
MEASURED prompt_tokens=150 became permanently unpriceable — and
store.reprice_unpriced filters those rows out forever, so the NULL could
never heal. Narrowed: only a null on a counter the provider actually uses
(cached_tokens inside a PRESENT container, a null top-level cache field) or
an explicit cache_*_unavailable discriminator declares an unknown.

P2 (finding 4) — moved the subscription_included route ABOVE the
unknown-usage refusal in estimate_usage_cost. A $0 route's cost does not
depend on token counts, so an unmeasured usage cannot make it unpriceable.
Matches the ordering plugins/observability/langfuse already uses.

P2 (finding 5) — empty_response_guard._zero_output read an UNMEASURED
output 0 as a measured zero, so two honest-unknown bridge turns satisfied
deterministic_empty() and cut the retry budget, contradicting the module's
own documented fail-open rule. Now returns (False, False) on
output_tokens_unknown / usage_unknown.

P2 (finding 6) — extracted the three blocks the consumer suite was
AST-lifting out of conversation_loop.py / slash_commands.py into real
callable functions (verbose_token_usage_log_args, cache_stats_line,
render_thin_last_turn_lines). Behaviour-preserving; the shipped sites now
call them, so coverage no longer depends on source-text anchors that break
on benign refactors of unrelated code.

Verified: tests/agent/test_usage_unknown_fleetreview_repro.py 13/13 pass
(includes the pre-fix defect reproduced, and controls proving the real
unknown declarations and the measured-zero deterministic-empty path still
work). tests/test_last_turn_usage_persistence.py green.

13 committed tests in test_usage_unknown_contract.py /
test_usage_unknown_consumers.py still pin the OLD behaviour these findings
call wrong; the test-guard hook blocks editing them. Named in the handback
for an operator decision.
Verified: scripts/run_tests.sh focused usage/pricing/guard/persistence/Langfuse suites (350 passed); seven mutation arms all red; Ruff clean.
Kyzcreig and others added 8 commits September 23, 2026 08:18
FleetReview r3 finding 5 (`Missing Usage`), reopened by review round 1.

`_canonical_usage_from_response` already turned an omitted usage payload
into an aggregate UNKNOWN, but the per-call commit block in
`run_conversation` was still gated on `response.usage` being truthy. The
UNKNOWN therefore reached only `context_compressor.update_from_response()`
and `_last_turn_usage`; it never incremented `session_api_calls`, set
`last_turn_usage`, appended to the Blackbox `_turn_calls` accumulator,
priced the call, or queued persistence. An unpriceable call was reported
as if it had never happened.

Gate the commit block on `response is not None` instead, and narrow the
context-probe persistence (the one part that genuinely needs a measured
payload) to `getattr(response, "usage", None)`. The settled
`normalize_usage(None)` known-no-call-zero contract is untouched.

Verified:
- new tests/run_agent/test_usageless_response_accounting.py 4/4 green;
  reverting the guard to its pre-fix shape turns the two unknown-path
  pins RED and leaves both controls GREEN (gate proven).
- focused set 127/0; tests/agent + tests/blackbox + tests/plugins/blackbox
  7771/0; tests/run_agent 2132/0.
- four mandated null-container controls re-derived: 10/10 on this head.
- 8-arm mutation set all RED, all four touched files restored byte-exact.
- Ruff clean.

Not mine, reported not fixed: tests/gateway/test_usage_command.py::
TestUsageLastTurnSnapshot::test_usage_command_prefers_dominant_persisted_route
fails identically at merge-base 6c400e2 (inherited red).
…ay sink

r4 FleetReview (record 598aff1-1dc4eb80) filed 12 findings, 10 P0-P1. The
prior head widened the per-call accounting guard to `response is not None` so
a usage-less success would finally be accounted — but that pushed the
UNKNOWN's PLACEHOLDER ZEROS into five sinks that carry no UNKNOWN
discriminator. That is the same silent-zero defect one level down, and worse:
the persisted last_turn_* snapshot is written COALESCE(?, existing), so a 0
DESTROYS the previous turn's real measurement.

CLASS: "the widened block reaches a sink with no UNKNOWN discriminator."
Swept every sink inside both widened blocks, fixed all of them in one pass:

- persisted last_turn_* snapshot -> _last_turn_snapshot_kwargs() writes None
  (not 0) when total_tokens_unknown, so COALESCE retains the last real split.
- session_cost_status -> _session_cost_status_with_known_spend() returns
  "partial" (not "unknown") when the session already holds priced dollars;
  "unknown" is outside the reprice allowlist and would never heal.
- _moa_session_cost_status -> seeds any_unpriceable from the AGGREGATOR, so an
  unpriceable aggregator beside priced advisors is "partial" too (r4 F4).
- context compressor / turn_finalizer last_prompt_tokens / Blackbox
  context_used -> gated on a measured payload exactly as the context-probe
  sub-branch already is; the pending compaction verdict is still consumed.
- velocity deques -> an unmeasured output is not 0 tok/s; skip the pair.

Also from the r4 record:
- F9: an omitted payload now sets EVERY discriminator via
  CanonicalUsage.fully_unknown(), not just the aggregate usage_unknown —
  consumers read the flags narrowly.
- F11: cache-unknown is resolved ACROSS both locations, so a measured
  cached_tokens is no longer overridden by a null top-level alias.
- F7/F12: the legacy usage_unknown latch is narrowed to unpriced rows whose
  counts are ALL ZERO. usage_unknown is the shared DISPLAY discriminator; a
  row is routinely unpriced because pricing REFUSED the route while carrying
  real measured counts.
- F8: the render pin asserts the TOKEN line, not any "unknown" (the
  cost-status line satisfied the old assertion).
- F6: the render_thin_last_turn_lines docstring no longer claims the
  persisted producer carries UNKNOWN flags; it cannot in this tree.
- F5/finding-4 REFUTED by execution — cache-unknown still refuses pricing per
  the standing consensus, which this change does not relitigate.

Verified: focused 235/0; tests/agent+blackbox+plugins/blackbox 7774/0;
tests/run_agent 2139/0; tests/gateway 8214/1 (the single red reproduces
identically in a detached worktree at merge-base 6c400e2 — inherited).
Four mandated null-container controls 10/10. Ruff clean. Nine mutation arms
all RED with byte-exact restore; the three arms that came back GREEN were
unpinned fixes and now carry discriminating pins. F8's vacuity proven by
regressing the real renderer: weak pin 4 GREEN, strong pin 4 RED.
…e control

The legacy_copy fixture deliberately stubs store._connect WITHOUT _ensure_schema
so test_ac7 can capture a pre-migration BEFORE arm. The placeholder probe reuses
that stub, so after the usage_unknown migration landed, its positive-control
insert_turn bound five columns the stubbed schema had never gained and failed
with 'table turns has no column named output_tokens_unknown'.

Production is unaffected: the real store._connect calls _ensure_schema on every
connection (store.py:53) — verified by running insert_turn against a pre-787
fixture DB through the real connector: the five *_unknown columns are added and
the row persists.

Verified the probe still discriminates: dropping one placeholder from
_INSERT_TURN_SQL turns it RED (1 failed / 1 passed); store.py restored
byte-exact (sha256 ea2a14d5d266...).
CI on the rebased head (first CI this PR has ever had — it was DIRTY until now)
turned slice 15/16 RED on test_migration_adds_comp_columns_to_legacy_db with
'sqlite3.OperationalError: no such column: cache_write'.

Root cause: the legacy latch's all-zero-counts guard names all four token count
columns unconditionally. `turns` is created with them, but a table that already
exists never gains them — CREATE TABLE IF NOT EXISTS is a no-op and no ALTER adds
them — so an old enough DB reaches the latch without cache_write. The raise
aborts the ENTIRE _ensure_schema, taking every later migration down with it.

Fix: build the sum from the count columns actually present in _existing. A column
the DB lacks cannot hold a measurement, which is exactly the zero the guard tests
for, so the condition's meaning is preserved.

Verified: tests/test_request_composition.py 40/0 and tests/plugins/blackbox/
test_store.py 19/0 (was 1 failed). New pin
test_migration_survives_a_db_missing_some_token_count_columns asserts the later
migrations still ran AND that the guard keeps discriminating measured from empty.
Restoring the hardcoded four-column sum turns BOTH tests RED; store.py restored
byte-exact (sha256 2a393c829c8a5d91...). Proven not inherited: both tests pass at
fork/main in a clean worktree.
…not be refused

r6 disposition of the FleetReview record on head 20dd114 (13 findings,
3 P1 / 8 P2 / 2 P3, reviewCommentUrl=None — never posted). One class-sweep
pass over both directions of the same defect: a sink with no UNKNOWN
discriminator, and a refusal that fires on measured data.

FIXED (9):

F4 [P1] agent/usage_pricing.py — a wire-null top-level cache alias beside
measured input/output made the turn permanently unpriceable.
`{"input_tokens":100,"output_tokens":50,"cache_read_input_tokens":null}` —
the shape agent/transports/anthropic.py already reads with `... or 0` — set
both cache flags, ORed into total_tokens_unknown, and reprice_unpriced now
skips any unknown row, so the NULL cost could never heal. A top-level alias
is OPTIONAL: on a measured payload its null means "no caching on this
request". A null inside a PRESENT container still means UNKNOWN.

F5 [P1] agent/moa_loop.py — a usage-less MoA advisor got a bare
CanonicalUsage() (every discriminator False) where conversation_loop uses
fully_unknown(). The pricing-call dict rode out all-False, so the advisor
priced at a confident $0 instead of refusing and its real spend vanished.

F13 [P1] hermes_state.py — the cost_status guard lived only in the caller,
reading agent.session_estimated_cost_usd, which agent_init.py zeroes on every
construction and never rehydrates. Any NEW agent resuming a session (gateway
eviction, CLI resume, cron/kanban re-entry) wrote 'unknown' over a row still
holding priced dollars — outside the reprice allowlist, so it never healed.
Moved to the durable chokepoint, which covers every caller.

F1 [P2] hermes_state.py — the session-scoped status also lands in the
per-model row, labelling a wholly-unpriced model 'partial' on the
Spend-by-model breakdown. Scoped to each row's own dollars, INSERT and
ON CONFLICT legs both.

F6 [P2] plugins/blackbox/store.py — the legacy latch marked genuinely
zero-token turns (interrupted before any call, blackbox-off, failed first
call) UNKNOWN, irreversibly, stranding them outside reprice_unpriced forever.
Scoped to rows the pre-UNKNOWN code itself labelled cost_status='unknown'.

F7 [P2] agent/conversation_loop.py — the compressor gate was a PRESENCE test
(`if getattr(response,"usage",None):`). A payload that carries a usage object
but nulls the prompt count passed it and stamped last_prompt_tokens — and so
the context meter, the persisted session entry and Blackbox context_used — to
0. Now gates on prompt_tokens_unknown, not total_tokens_unknown: an unmeasured
OUTPUT bucket must not discard a good PROMPT reading.

F8 [P2] gateway/slash_commands.py — the resident lane built five flagless
ints, making every UNKNOWN branch dead code there; an unmeasured session
rendered `Total (billed in+out): 0`. Extracted `_resident_thin_snapshot` and
carried the flags from agent.last_turn_usage.

F11 [P2] / F10 [P3] tests — UC-8b's `UNKNOWN_TOKENS_LABEL in block` was
satisfied by the cost line above the token row, so the defect it exists to
pin could regress green; now pins `Tokens out: <label>`. The scratch repro's
f3 case ASSERTED the measured-looking zero as expected output, locking in the
exact rendering F8 fixes; re-pinned onto the discriminator contract.

TEST-REPIN (1): tests/plugins/blackbox/test_store.py
test_schema_migration_adds_columns_to_preexisting_table required the latch to
fire on an unpriced all-zero row with no cost_status. That premise is what F6
supersedes. The row now carries cost_status='unknown' (the ambiguous case,
still latched) and a new 'legacy-zero-token-measured' row pins that a
never-accounted zero-token turn is NOT latched.

Verification, all by execution on this head:
- New tests/agent/test_usage_unknown_r6_findings.py: 21/21. Every fix has a
  narrowness control, so none is bought by making the refusal unreachable
  (explicit *_unavailable still wins; null-in-present-container still UNKNOWN;
  a session with no priced dollars is still 'unknown'; a measured advisor
  still prices; an output-only unknown still advances the context meter).
- Focused 8 files: 139/0. tests/agent+blackbox+plugins/blackbox+run_agent:
  10022 passed, 0 failed, 29 skipped.
- Mutation set 9/9 RED, every file restored byte-exact.
- The harness's BASELINE check earned its keep: the first run reported 9/9 RED
  vacuously — the repo-local .venv has no pytest (run_tests.sh skips it), so
  every arm was failing on `No module named pytest`, which reads as a caught
  regression. Fixed the interpreter probe; the re-run then exposed F7 as a
  genuinely vacuous pin (its test only exercised the predicate, not the gate).
  Replaced it with a real two-turn run_conversation() regression against a
  REAL SessionDB, driving a PRESENT-but-null-prompt payload; reverting only
  the gate now fails on that test's named assertion.
- Argus's own probes, re-run byte-unmodified: r2 8/0, r3 CONTRACT 9/9.
- Ruff clean.

Prior contracts untouched and pinned by controls: normalize_usage(None) is a
MEASURED zero; a null details CONTAINER is a measured zero; fully_unknown()
sets every discriminator.
r6 disposition round 2 — the three findings the prior commit left undisposed
from the FleetReview record on head 20dd114 (F9 P2, F3 P3, F12 P3). The other
ten were dispositioned in 4ed89cd. Record: ANG-Ventures_hermes-agent-pr-787-
20dd114-e2297280-b5f8-11f1-8043-66cc4f68f983.json, reviewCommentUrl=None.

FIXED (3):

F9 [P2] plugins/blackbox/last_turn.py + agent/turn_finalizer.py + record.py +
store.py — three rows on the last-turn card describe the FINAL call only (Last
call split, Cached hit-rate, Context window occupancy), but were gated on
`input_unknown = prompt_tokens_unknown(rec)`, which _rollup_turn_usage builds
with any() ACROSS EVERY CALL of the turn. A 5-call turn whose call #2 returned
no usage rendered `Context window (last call): unknown` and dropped the split,
while the final call's numbers were real provider measurements.

There was no per-last-call discriminator to gate on, so this adds one:
turn_finalizer computes `last_call_prompt_unknown` from `_turn_calls[-1]` — the
same element the split itself is read from — and it rides the existing
turn_usage dict through _build_record into a new nullable TurnRecord field and
turns column. The renderer gates the two genuinely last-call rows on it.

The `• Cached:` row is NOT moved: it divides turn-level sums that really are
missing a term, so the absorbing flag is correct there and stays. That
asymmetry is the finding's actual content and is pinned both ways.

Migration is NULLable with no DEFAULT, deliberately. NULL means "this row
predates the column, so its final-call provenance was never recorded" and the
renderer falls back to the turn-level flag — byte-identical to today for every
historical row. A DEFAULT 0 would instead assert "the final call WAS measured"
about rows that were genuinely unmeasured and render their placeholder zeros as
real window numbers, which is the defect class this PR exists to remove.

F3 [P3] agent/usage_pricing.py — `cache_stats_line` replaced inline code that
printed NOTHING when there was no cache activity, and it runs per API CALL, so
a provider that routinely omits usage emitted one `💾 Cache: unknown` noise line
per call (12 in a 12-call tool loop). Silenced on the AGGREGATE discriminator
ALONE (`usage_unknown` — no payload at all, so there is no cache-specific fact
to report and the turn card already says the turn is unmeasured). Every
narrower unknown still prints: a null bucket inside a present container, an
explicitly-unavailable prompt, measured cache counts beside an unmeasured
prompt. In those the refused percentage is the fabrication.

F12 [P3] tests/agent/test_usage_unknown_contract.py — the store round-trips fed
the renderer `store.get_turn`'s shape, but `_row_to_dict` RENAMES cache_read ->
cache_read_tokens while production (`compute_last_turn_record`) does a raw
SELECT * and hands over the unrenamed keys — which is why the renderer reads
`rec.get("cache_read", 0)`. Every cache lookup silently resolved to 0. Switched
to the production row shape with an explicit shape assertion, so a future
assertion about cache or Last-call lines cannot be written against a shape that
never occurs at runtime.

Verification, all by execution on this head:
- tests/agent/test_usage_unknown_r6_round2.py: 15/15, each fix with a
  narrowness control (an unmeasured final call is still unknown; the turn-level
  Cached row still gates on the turn flag; a legacy NULL row keeps today's
  behaviour; a null cache bucket still prints unknown; measured paths byte-
  identical).
- tests/agent + tests/plugins/blackbox + tests/blackbox + tests/run_agent +
  tests/test_request_composition.py: 10076 passed, 1 failed, 29 skipped. The
  one red is tests/agent/lsp/test_client_e2e.py — a psutil flake unrelated to
  usage; 3/3 green when run alone.
- Mutation set 5/5 RED, every file restored byte-exact:
    f3_remove_usageless_silence          1 RED
    f9_renderer_reverts_to_turn_level_or 2 RED
    f9_producer_always_reports_measured  1 RED
    f9_migration_dropped                 2 RED
    f12_back_to_get_turn_shape           9 RED
  The producer arm earned its keep: the first matrix returned it GREEN because
  the renderer tests build the row by hand and re-implement the fold, so a
  producer hardcoding the flag killed nothing. Added a regression that drives
  the shipped finalize_turn() and captures the turn_usage it emits on
  on_session_end; the arm is RED now. The F12 arm's first anchor matched 2x and
  could not apply — re-anchored on the helper body rather than reporting a
  vacuous pass.
- Ruff clean.

Prior contracts untouched and still pinned: null CONTAINER = measured zeros;
null count inside a present container = UNKNOWN; normalize_usage(None) = a
MEASURED zero; fully_unknown() sets every discriminator.

NOT DISPOSITIONED HERE — F2 [P2] "nulling all five last_turn_* columns leaves
the persisted card showing an older turn's split with no staleness marker".
SPLIT to t_25f50547 / PR #797 under the standing schema-ownership split: the
finding's own body names the fix as `*_unknown` companion columns on the
`sessions` table, which is that card's surface, and it explicitly reads the
present behaviour as the correct trade-off against destroying the split.
…aggregate

r7 of the FleetReview record on cc3ced8 (13 findings, 9 P1, TRUSTED:
healthyMembers=5/6, realizedRoster.trusted=true; escalation is judge-side 429
capacity, not member failure). All 13 dispositioned.

F6 [P1] card.py — CLASS SWEEP, not one site. `context_used` and the
`last_cache_*` split describe the FINAL call only, so this PR's hardening of
Tokens-in/out/Cached left `• Context: 0/200k (0% of model max)` fabricating a
measurement on the same card that said `unknown` twice above it. The gate is now
ONE shared callable, `usage_pricing.last_call_prompt_unknown`, which both
renderers consult — including the NULL-legacy fallback (`store.py` ALTERs the
column in without a DEFAULT, so `TurnRecord.last_call_prompt_unknown` is typed
`bool` and carries `None`; a bare truth test asserts "measured" about every
historic row). The record noted no test discriminated this (248-vs-248); the new
class-sweep asserts CROSS-RENDERER AGREEMENT rather than per-site strings, so a
future third renderer that skips the gate fails without anyone extending a list.

F3 [P1] empty_response_guard.py — the F5 trade is preserved and its cost
recovered. A declared-unmeasured output still refuses to classify as a
deterministic empty, but guard #2 was ALREADY inert for those payloads
(`estimate_usage_cost` refuses `total_tokens_unknown`), so a bridge nulling only
its output kept the full 3-retry budget while re-sending a measured 400k prompt.
`measured_cost_floor` prices the MEASURED buckets only — for a DECISION, never
persisted, never displayed — and can only understate the bill, so the ceiling
cannot trip early. Measured: $2.00/attempt on claude-opus-5 at 400k; budget 3→1.
The old pin sat at prompt_tokens=150 where the cost is nil.

F13 [P1] conversation_loop.py — worst-of merge. A session's cost label can never
get MORE complete over time. r4 fixed priced-then-unknown; the mirror overwrote
an incomplete session back to `estimated`. `prior_api_calls` is load-bearing: a
fresh agent's initial `unknown`/zero-calls is UNSTARTED, not incomplete.

F7/F4 [P1] slash_commands.py — absorbing session-level latch. `last_turn_usage`
is rewritten every CALL, so stamping its flags on cumulative counters was wrong
both ways (an exact-looking 412,338 omitting real spend; 412k measured tokens
collapsed to unknown). Latched per-bucket beside the counter increments, cleared
with them in reset_session_state, wired on the codex lane too.

F1 [P1] — UNPRICEABLE is not UNMEASURED. An uncatalogued advisor route returns
amount_usd=None with every token measured and no flag set, so the session lane
reported `estimated` for a total omitting it while Blackbox said `partial`. Each
advisor's own verdict now rides its pricing record and is consulted.

F2 [P2] — the compressor reads the PRE-FOLD aggregator usage, as the anchor
below it already does. `__add__` makes unknown flags absorbing, so one unmeasured
advisor blanked a complete aggregator prompt count.

F12 [P2], F5 [P3], F11 [P3] — test fidelity: `return []` (verified: zero
swallowed hook TypeErrors), production ORDERING mirrored in the cache-line pin,
and `else` legs so no wire asserts nothing.

F8 [P1] store.py, F9 [P1] usage_pricing.py, F10 [P1] hermes_state.py carried
forward from the prior commit.

TEST-REPINs (2): `test_f8_resident_snapshot...` moved off the per-call fixture
that finding 4 showed could not discriminate; `test_append_inside_success_block_only`
replaced a banned source CHARACTER-DISTANCE assertion (2359→4116 chars, red on
comments alone, and green whenever the statements happen to be adjacent) with
the behaviour it names, driven through the real run_conversation.

Verified: focused 386/0; tests/agent+run_agent+blackbox broad sweep, 4 failures,
3 proven inherited at merge-base edb2880 and the 4th re-pinned above. Nine
mutation arms all RED (F6 gate 4, NULL fallback 5, F13 merge 1, prior_api_calls
3, session latch 4, F1 verdict 2, F2 payload 2, F3 floor 3, accumulator 1), all
files restored byte-exact, scratch worktrees removed. Ruff clean.
r8. Two defects in the r7 commit (dfcacfa), both found by execution rather
than by reading it: one RED in CI, one vacuous mutation arm.

CI RED (the substantive one). The F7 session-level latch read
`if getattr(agent, f"session_{flag}", False):`. A `MagicMock` auto-creates every
attribute access as a truthy child, so on any mock-backed agent all five flags
read set and a fully MEASURED session collapsed to `unknown` on the resident
/usage lane — `tests/gateway/test_usage_command.py` went red on slice 5/16
(35,000 and 70,000 both rendered as "Tokens in: unknown"). That is this PR's own
silent-fabrication class inverted: refusing a measurement that was really taken.
The old code was accidentally safe via its `isinstance(last_turn, dict)` guard;
widening to a bare getattr dropped the type check with it.

Fixed with `is True`. The latch is written as a real bool by agent_init.py
(False) and conversation_loop.py (True), so any other value means the attribute
was never initialised on that object. The five counters beside these flags are
already coerced through `as_int` for exactly this reason; this is the flags' half
of that contract. Class-swept: `getattr(agent, f"session_{...}")` has exactly one
read site in the tree (the three writers are `setattr` and unaffected).

Pinned with a MagicMock-backed agent plus a narrowness control proving a
genuinely declared latch still fires. Every pre-existing pin in that file builds
its agent from SimpleNamespace or a real AIAgent, which is precisely why none of
them could see this — AGENTS.md names the trap ("check with `is True` — MagicMock
doubles auto-create truthy attributes") and the r7 commit walked into it anyway.

VACUOUS ARM (F2). Arm 7 — swapping the compressor gate's `aggregator_usage` back
to the folded `canonical_usage` at the call site — came back GREEN across all
five focused files. The r7 F2 pins test `_compressor_usage_dict` and
`prompt_tokens_unknown` at unit level, so nothing observed which usage the loop
actually hands the compressor: the fix's own call site was unpinned. Added two
tests that drive the real `run_conversation()` with a client speaking the genuine
MoA `consume_reference_usage` seam — one unmeasured advisor must not blank a
measured 4000-token window occupancy, and (narrowness) a measured 90k advisor
fan-out must not be counted AS occupancy. Arms against both halves are RED now.

Verified by execution on this head:
- focused 7 files: 190/0, plus the 4 new tests.
- tests/gateway/test_usage_command.py: the two CI failures now pass. The
  remaining red (test_usage_command_prefers_dominant_persisted_route) is
  INHERITED — it fails identically in a fresh detached worktree at merge-base
  edb2880, which contains none of this branch's commits.
- broad tests/agent+run_agent+plugins/blackbox+blackbox: 10067 passed / 3 failed;
  all three relay failures reproduce at merge-base edb2880 (proven, not
  asserted) and none touch usage.
- 9 mutation arms RED (F6 card gate 4, F6 NULL fallback 2, F13 worst-of merge 1,
  F7 latch 4, F3 cost floor 3, F1 advisor verdict 2, F2 gate 1, F2 payload 1,
  latch `is True` 1). Baseline-gated green before each arm; every file restored
  byte-exact; both scratch worktrees removed.
- Argus's r2/r3/r4 probes re-run byte-unmodified: 8/0, 9/0, 11/0.
- F12 verified by execution: zero swallowed hook TypeErrors in that file's run.
- Ruff clean.
@Kyzcreig
Kyzcreig force-pushed the fix/usage-unknown-end-to-end branch from 2a7685d to 564f7e4 Compare September 23, 2026 15:23
Kyzcreig added a commit that referenced this pull request Sep 23, 2026
…tals and the Langfuse export

#787 made UNKNOWN a first-class state for the Blackbox PER-TURN contract.
Three other consumers still collapsed an unmeasured term into a measured
zero. Bridge claude-bpx #186 is merged and fleet-deployed, so these paths
receive unknown-input shapes today. Reproduced all three on the tree first:

  update_token_counts REJECTS unknown flags: unexpected keyword argument
  record_auxiliary_usage REJECTS unknown flags: unexpected keyword argument
  cumulative counters carry NO unknown provenance
  tui cache_hit_pct = 100% over prompt_total=100 (the unknown turn added 0)
  langfuse input-only -> ({'input': 0, 'output': 50}, {})

1. PERSISTED SCHEMA. sessions and session_model_usage gain the five
   USAGE_UNKNOWN_FIELDS discriminators (+ last_turn_* snapshot mirrors),
   additive with DEFAULT 0 so every legacy row and every provider that
   never speaks this dialect reads back "measured". The int counters are
   untouched, so no arithmetic consumer changes. Cumulative flags are
   ABSORBING (MAX in SQL): one unmeasured call latches the session term,
   which is why a last-call-only guard is wrong here. Classified into
   _TOKEN_DELTA_ROUTE_FIELDS so deltas with differing unknown state never
   coalesce. Both session_model_usage table-rebuild paths (PK heal, v21->v22)
   copy the flags only when the source table has them, so a rebuild cannot
   silently reset a latched unknown. record_aux_usage no longer drops an
   all-zero unmeasured aux call as "empty".

2. CUMULATIVE CLI/TUI TOTALS. The loop latches per-term provenance onto
   session_*_unknown beside the session_*_tokens commit. Display reuses the
   Blackbox lane's rule (format_token_count / prompt_tokens_unknown) via new
   session_usage_unknown_flags / session_total_tokens_unknown helpers -- not
   a forked cumulative rule. /usage, both status-bar Sigma renderers, the CLI
   cache-hit segment and the TUI payload now say "unknown" or emit no ratio
   instead of a fabricated percentage.

3. LANGFUSE EXPORT. An unmeasured term is OMITTED from usage_details
   (absent, never 0) and declared on the observation metadata via
   _unknown_usage_details(); cost_details is empty for an unpriceable turn
   so Langfuse cannot read a partial subtotal as authoritative. The
   post_api_request summary-dict reconstruction was a second flag-loss site
   and now carries the discriminators. The 2-tuple return contract of
   _canonical_usage_and_cost / _usage_and_cost is unchanged.

Verified (repo venv):
- tests/agent/test_usage_unknown_persisted_export.py -- 71 pins, the five
  literal bridge payloads (input-only / cache-only / wholly-unavailable /
  output-only / measured-zero) through EACH of the three paths, plus a
  legacy-DB migration arm built by stripping the new columns out of the real
  SCHEMA_SQL. Cumulative and display arms source-LIFT the shipped blocks by
  AST anchor rather than re-implementing them.
- scripts/run_tests.sh over the contract/consumer/persisted suites +
  tests/hermes_state + async token accounting + langfuse + last-turn
  persistence + turn-usage accumulator + blackbox:
  "54 files, 682 tests passed, 0 failed, 1 skipped".
- 10-arm mutation matrix, each restored byte-exact, fake kills (SyntaxError/
  NameError/zero-collected) rejected: 10/10 RED, control 131 passed before
  and after. Arms: SessionDB counter flags dropped; absorbing -> last-wins;
  aux ledger flags dropped; aux producer drops the unmeasured call;
  cumulative provenance not absorbing; CLI cache-ratio guard removed; /usage
  card prints numbers; TUI payload fabricates the ratio; Langfuse exports 0
  for unknown; Langfuse summary-dict drops flags.
- ui-tui: npm run typecheck clean. npm test is 61 files / 6 tests red on
  this machine, IDENTICAL on the base commit with the change stashed
  (textInputReturnAction + an unbuilt shared hermes-ink) -- pre-existing,
  not from this diff.

Scope: this card only. The Blackbox per-turn contract certified under
t_083ad45e / #787 is untouched.
Kyzcreig added a commit that referenced this pull request Sep 23, 2026
…apshot flags, append-only schema

Re-stacked hermes-agent#797 onto #787 (rebased onto main 2026-09-23). Two of its six
commits were already on #787 under the same subjects and were dropped; the other four
replayed with semantic conflicts. Fixes on top, each measured red before:

1. Duplicate absorbing latch: #787's "latch the flag over the aggregate" and #797's
   "carry UNKNOWN through cumulative totals" each added the same session_* latch loop
   to agent/conversation_loop.py. Kept #787's (beside the increments); removed #797's
   copy. 25 AST-seam tests ("shipped cumulative-usage seam moved") go green.
2. last_turn_*_unknown flags now move WITH the snapshot they qualify: when #787's
   _last_turn_snapshot_kwargs withholds an unmeasured call's split (None -> COALESCE
   keeps the prior real split), the flags are withheld too. Before: the persisted card
   rendered "Total: unknown" over a real 4,120 from the previous turn.
3. Schema columns APPENDED, not inserted mid-table. #797 put its 10 sessions flags
   before git_branch and its 5 session_model_usage flags before estimated_cost_usd.
   Migrated DBs get them via ALTER ADD (end of table), so fresh and migrated DBs had
   different physical orders, and the positional lost-and-found recovery mapper
   misplaced historical fields (title landed in the wrong column). Fresh-DB column
   prefix is now byte-identical to main's (title back at index 43).
4. Recovery classifier accepts session_model_usage records of 18 (pre-migration) AND
   23 fields; it previously dropped every post-migration usage row. Test pins 23 and
   adds a 23-field classify case.

979 passed across every test file this branch touches + hermes_state + session
recovery suites. Apollo merge pass 2026-09-23.
@Kyzcreig

Copy link
Copy Markdown
Collaborator Author

🤖 merged-by: apollo · lane: fr-pause-0922 · gate: BYPASS: FR PAUSED by Ace ruling 2026-09-22 (state/fleetreview-pause-20260922.md); - approved · why: rebased onto main by Apollo; Windows red = known desktop /progress flake, rerun; CI green; Apollo merge pass 2026-09-22

@Kyzcreig
Kyzcreig added this pull request to the merge queue Sep 23, 2026
@github-merge-queue
github-merge-queue Bot removed this pull request from the merge queue due to failed status checks Sep 23, 2026
@Kyzcreig

Copy link
Copy Markdown
Collaborator Author

🤖 merged-by: apollo · lane: fr-pause-0922 · gate: BYPASS: FR PAUSED by Ace ruling 2026-09-22 (state/fleetreview-pause-20260922.md); t_083ad45e stack root, argus-reviewed round 2 APPROVED · why: rebased onto main by Apollo 2026-09-23; 233 branch tests + CI green; Windows red was the known /progress flake

@Kyzcreig
Kyzcreig added this pull request to the merge queue Sep 23, 2026
Merged via the queue into main with commit 5139318 Sep 23, 2026
54 checks passed
@Kyzcreig
Kyzcreig deleted the fix/usage-unknown-end-to-end branch September 23, 2026 16:38
Kyzcreig added a commit that referenced this pull request Sep 23, 2026
…tals and the Langfuse export

#787 made UNKNOWN a first-class state for the Blackbox PER-TURN contract.
Three other consumers still collapsed an unmeasured term into a measured
zero. Bridge claude-bpx #186 is merged and fleet-deployed, so these paths
receive unknown-input shapes today. Reproduced all three on the tree first:

  update_token_counts REJECTS unknown flags: unexpected keyword argument
  record_auxiliary_usage REJECTS unknown flags: unexpected keyword argument
  cumulative counters carry NO unknown provenance
  tui cache_hit_pct = 100% over prompt_total=100 (the unknown turn added 0)
  langfuse input-only -> ({'input': 0, 'output': 50}, {})

1. PERSISTED SCHEMA. sessions and session_model_usage gain the five
   USAGE_UNKNOWN_FIELDS discriminators (+ last_turn_* snapshot mirrors),
   additive with DEFAULT 0 so every legacy row and every provider that
   never speaks this dialect reads back "measured". The int counters are
   untouched, so no arithmetic consumer changes. Cumulative flags are
   ABSORBING (MAX in SQL): one unmeasured call latches the session term,
   which is why a last-call-only guard is wrong here. Classified into
   _TOKEN_DELTA_ROUTE_FIELDS so deltas with differing unknown state never
   coalesce. Both session_model_usage table-rebuild paths (PK heal, v21->v22)
   copy the flags only when the source table has them, so a rebuild cannot
   silently reset a latched unknown. record_aux_usage no longer drops an
   all-zero unmeasured aux call as "empty".

2. CUMULATIVE CLI/TUI TOTALS. The loop latches per-term provenance onto
   session_*_unknown beside the session_*_tokens commit. Display reuses the
   Blackbox lane's rule (format_token_count / prompt_tokens_unknown) via new
   session_usage_unknown_flags / session_total_tokens_unknown helpers -- not
   a forked cumulative rule. /usage, both status-bar Sigma renderers, the CLI
   cache-hit segment and the TUI payload now say "unknown" or emit no ratio
   instead of a fabricated percentage.

3. LANGFUSE EXPORT. An unmeasured term is OMITTED from usage_details
   (absent, never 0) and declared on the observation metadata via
   _unknown_usage_details(); cost_details is empty for an unpriceable turn
   so Langfuse cannot read a partial subtotal as authoritative. The
   post_api_request summary-dict reconstruction was a second flag-loss site
   and now carries the discriminators. The 2-tuple return contract of
   _canonical_usage_and_cost / _usage_and_cost is unchanged.

Verified (repo venv):
- tests/agent/test_usage_unknown_persisted_export.py -- 71 pins, the five
  literal bridge payloads (input-only / cache-only / wholly-unavailable /
  output-only / measured-zero) through EACH of the three paths, plus a
  legacy-DB migration arm built by stripping the new columns out of the real
  SCHEMA_SQL. Cumulative and display arms source-LIFT the shipped blocks by
  AST anchor rather than re-implementing them.
- scripts/run_tests.sh over the contract/consumer/persisted suites +
  tests/hermes_state + async token accounting + langfuse + last-turn
  persistence + turn-usage accumulator + blackbox:
  "54 files, 682 tests passed, 0 failed, 1 skipped".
- 10-arm mutation matrix, each restored byte-exact, fake kills (SyntaxError/
  NameError/zero-collected) rejected: 10/10 RED, control 131 passed before
  and after. Arms: SessionDB counter flags dropped; absorbing -> last-wins;
  aux ledger flags dropped; aux producer drops the unmeasured call;
  cumulative provenance not absorbing; CLI cache-ratio guard removed; /usage
  card prints numbers; TUI payload fabricates the ratio; Langfuse exports 0
  for unknown; Langfuse summary-dict drops flags.
- ui-tui: npm run typecheck clean. npm test is 61 files / 6 tests red on
  this machine, IDENTICAL on the base commit with the change stashed
  (textInputReturnAction + an unbuilt shared hermes-ink) -- pre-existing,
  not from this diff.

Scope: this card only. The Blackbox per-turn contract certified under
t_083ad45e / #787 is untouched.
Kyzcreig added a commit that referenced this pull request Sep 23, 2026
…apshot flags, append-only schema

Re-stacked hermes-agent#797 onto #787 (rebased onto main 2026-09-23). Two of its six
commits were already on #787 under the same subjects and were dropped; the other four
replayed with semantic conflicts. Fixes on top, each measured red before:

1. Duplicate absorbing latch: #787's "latch the flag over the aggregate" and #797's
   "carry UNKNOWN through cumulative totals" each added the same session_* latch loop
   to agent/conversation_loop.py. Kept #787's (beside the increments); removed #797's
   copy. 25 AST-seam tests ("shipped cumulative-usage seam moved") go green.
2. last_turn_*_unknown flags now move WITH the snapshot they qualify: when #787's
   _last_turn_snapshot_kwargs withholds an unmeasured call's split (None -> COALESCE
   keeps the prior real split), the flags are withheld too. Before: the persisted card
   rendered "Total: unknown" over a real 4,120 from the previous turn.
3. Schema columns APPENDED, not inserted mid-table. #797 put its 10 sessions flags
   before git_branch and its 5 session_model_usage flags before estimated_cost_usd.
   Migrated DBs get them via ALTER ADD (end of table), so fresh and migrated DBs had
   different physical orders, and the positional lost-and-found recovery mapper
   misplaced historical fields (title landed in the wrong column). Fresh-DB column
   prefix is now byte-identical to main's (title back at index 43).
4. Recovery classifier accepts session_model_usage records of 18 (pre-migration) AND
   23 fields; it previously dropped every post-migration usage row. Test pins 23 and
   adds a 23-field classify case.

979 passed across every test file this branch touches + hermes_state + session
recovery suites. Apollo merge pass 2026-09-23.
github-merge-queue Bot pushed a commit that referenced this pull request Sep 23, 2026
…tals and the Langfuse export (#797)

* fix(usage): carry UNKNOWN through the persisted schema, cumulative totals and the Langfuse export

#787 made UNKNOWN a first-class state for the Blackbox PER-TURN contract.
Three other consumers still collapsed an unmeasured term into a measured
zero. Bridge claude-bpx #186 is merged and fleet-deployed, so these paths
receive unknown-input shapes today. Reproduced all three on the tree first:

  update_token_counts REJECTS unknown flags: unexpected keyword argument
  record_auxiliary_usage REJECTS unknown flags: unexpected keyword argument
  cumulative counters carry NO unknown provenance
  tui cache_hit_pct = 100% over prompt_total=100 (the unknown turn added 0)
  langfuse input-only -> ({'input': 0, 'output': 50}, {})

1. PERSISTED SCHEMA. sessions and session_model_usage gain the five
   USAGE_UNKNOWN_FIELDS discriminators (+ last_turn_* snapshot mirrors),
   additive with DEFAULT 0 so every legacy row and every provider that
   never speaks this dialect reads back "measured". The int counters are
   untouched, so no arithmetic consumer changes. Cumulative flags are
   ABSORBING (MAX in SQL): one unmeasured call latches the session term,
   which is why a last-call-only guard is wrong here. Classified into
   _TOKEN_DELTA_ROUTE_FIELDS so deltas with differing unknown state never
   coalesce. Both session_model_usage table-rebuild paths (PK heal, v21->v22)
   copy the flags only when the source table has them, so a rebuild cannot
   silently reset a latched unknown. record_aux_usage no longer drops an
   all-zero unmeasured aux call as "empty".

2. CUMULATIVE CLI/TUI TOTALS. The loop latches per-term provenance onto
   session_*_unknown beside the session_*_tokens commit. Display reuses the
   Blackbox lane's rule (format_token_count / prompt_tokens_unknown) via new
   session_usage_unknown_flags / session_total_tokens_unknown helpers -- not
   a forked cumulative rule. /usage, both status-bar Sigma renderers, the CLI
   cache-hit segment and the TUI payload now say "unknown" or emit no ratio
   instead of a fabricated percentage.

3. LANGFUSE EXPORT. An unmeasured term is OMITTED from usage_details
   (absent, never 0) and declared on the observation metadata via
   _unknown_usage_details(); cost_details is empty for an unpriceable turn
   so Langfuse cannot read a partial subtotal as authoritative. The
   post_api_request summary-dict reconstruction was a second flag-loss site
   and now carries the discriminators. The 2-tuple return contract of
   _canonical_usage_and_cost / _usage_and_cost is unchanged.

Verified (repo venv):
- tests/agent/test_usage_unknown_persisted_export.py -- 71 pins, the five
  literal bridge payloads (input-only / cache-only / wholly-unavailable /
  output-only / measured-zero) through EACH of the three paths, plus a
  legacy-DB migration arm built by stripping the new columns out of the real
  SCHEMA_SQL. Cumulative and display arms source-LIFT the shipped blocks by
  AST anchor rather than re-implementing them.
- scripts/run_tests.sh over the contract/consumer/persisted suites +
  tests/hermes_state + async token accounting + langfuse + last-turn
  persistence + turn-usage accumulator + blackbox:
  "54 files, 682 tests passed, 0 failed, 1 skipped".
- 10-arm mutation matrix, each restored byte-exact, fake kills (SyntaxError/
  NameError/zero-collected) rejected: 10/10 RED, control 131 passed before
  and after. Arms: SessionDB counter flags dropped; absorbing -> last-wins;
  aux ledger flags dropped; aux producer drops the unmeasured call;
  cumulative provenance not absorbing; CLI cache-ratio guard removed; /usage
  card prints numbers; TUI payload fabricates the ratio; Langfuse exports 0
  for unknown; Langfuse summary-dict drops flags.
- ui-tui: npm run typecheck clean. npm test is 61 files / 6 tests red on
  this machine, IDENTICAL on the base commit with the change stashed
  (textInputReturnAction + an unbuilt shared hermes-ink) -- pre-existing,
  not from this diff.

Scope: this card only. The Blackbox per-turn contract certified under
t_083ad45e / #787 is untouched.

* fix(usage): gate cumulative cache ratios at the CLI render seam

Run real snapshot and renderer regression tests for five bridge shapes. Verified 56 files: 743 passed, 0 failed, 1 skipped; producer and renderer guard mutants both fail.

* fix(usage): preserve unknown in remaining cumulative CLI totals

CLASS-SWEEP found /status raw total and zero-counter Sigma omission. Reproduced /status failure, then verified 3 focused tests pass; prior broad run passed 745 with one test fixture import error subsequently fixed.

* test(usage): cover null container persistence and export controls

* fix(usage): port #797 onto the rebased #787 -- one latch, coherent snapshot flags, append-only schema

Re-stacked hermes-agent#797 onto #787 (rebased onto main 2026-09-23). Two of its six
commits were already on #787 under the same subjects and were dropped; the other four
replayed with semantic conflicts. Fixes on top, each measured red before:

1. Duplicate absorbing latch: #787's "latch the flag over the aggregate" and #797's
   "carry UNKNOWN through cumulative totals" each added the same session_* latch loop
   to agent/conversation_loop.py. Kept #787's (beside the increments); removed #797's
   copy. 25 AST-seam tests ("shipped cumulative-usage seam moved") go green.
2. last_turn_*_unknown flags now move WITH the snapshot they qualify: when #787's
   _last_turn_snapshot_kwargs withholds an unmeasured call's split (None -> COALESCE
   keeps the prior real split), the flags are withheld too. Before: the persisted card
   rendered "Total: unknown" over a real 4,120 from the previous turn.
3. Schema columns APPENDED, not inserted mid-table. #797 put its 10 sessions flags
   before git_branch and its 5 session_model_usage flags before estimated_cost_usd.
   Migrated DBs get them via ALTER ADD (end of table), so fresh and migrated DBs had
   different physical orders, and the positional lost-and-found recovery mapper
   misplaced historical fields (title landed in the wrong column). Fresh-DB column
   prefix is now byte-identical to main's (title back at index 43).
4. Recovery classifier accepts session_model_usage records of 18 (pre-migration) AND
   23 fields; it previously dropped every post-migration usage row. Test pins 23 and
   adds a 23-field classify case.

979 passed across every test file this branch touches + hermes_state + session
recovery suites. Apollo merge pass 2026-09-23.

---------

Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
@Kyzcreig Kyzcreig added the fleetreview:post-merge Ask FleetReview to review this MERGED pull (merge commit vs first parent) label Sep 24, 2026
Kyzcreig added a commit that referenced this pull request Sep 24, 2026
Re-port #779 onto current main without reverting #764 validation or #787 usage accounting. Carry fixed kind-specific labels and metadata-only history events; share the dropped-call retry budget and strip ephemeral correction pairs at finalization.

Verified: 160 confab/CLI/gateway tests, 42 dropped-call and usage tests, 33 finalizer/status tests; git diff --check.
Kyzcreig added a commit that referenced this pull request Sep 24, 2026
Re-port #779 onto current main without reverting #764 validation or #787 usage accounting. Carry fixed kind-specific labels and metadata-only history events; share the dropped-call retry budget and strip ephemeral correction pairs at finalization.

Verified: 160 confab/CLI/gateway tests, 42 dropped-call and usage tests, 33 finalizer/status tests; git diff --check.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

do-not-merge Hold: a lane must not merge this PR (Apollo lifts it) fleetreview:post-merge Ask FleetReview to review this MERGED pull (merge commit vs first parent)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants