fix(desktop): server-side pinned sessions (multi-device pin desync) - #186
Conversation
Add a server-owned pinned flag to SessionDB, expose it through REST/session lists and the session.pin gateway RPC, and derive desktop pinned rows from server session data instead of localStorage. Verified: scripts/run_tests.sh tests/test_hermes_state.py tests/test_tui_gateway_server.py tests/test_web_server_sessiondb_eventloop.py; npm --workspace apps/desktop run typecheck; npm --workspace apps/desktop run test:ui -- src/store/session.test.ts src/store/layout.test.ts src/store/updates.test.ts. Note: npm --workspace apps/desktop run test:ui was also attempted after a temporary worktree node_modules symlink; it failed in desktop suites outside the touched pinned-session files (46 files / 21 tests).
| -- Intentionally asymmetric: rewind_count bumps per rewind_to_message call; | ||
| -- redo_count bumps once per /redo command, regardless of M. | ||
| redo_count INTEGER, | ||
| pinned INTEGER NOT NULL DEFAULT 0, |
There was a problem hiding this comment.
Existing Databases Miss Column
This adds pinned only to the table creation SQL, so existing state.db files can still lack the column after upgrade. The first pin or unpin then reaches set_session_pinned and runs UPDATE sessions SET pinned = ?, which fails with no such column: pinned for existing users and makes the new pin feature unusable on those profiles.
| assert session is not None | ||
| db = _get_db() | ||
| if db is None: | ||
| return _db_unavailable_error(rid, code=5007) |
There was a problem hiding this comment.
| export const $pinnedSessionIds: ReadableAtom<string[]> = computed( | ||
| [$sessions, $legacyPinnedSessionIds], | ||
| (sessions, legacyPinnedSessionIds) => { | ||
| const serverCapable = sessions.some(session => 'pinned' in session) | ||
|
|
||
| if (!serverCapable) { | ||
| return legacyPinnedSessionIds | ||
| } | ||
|
|
||
| const ids: string[] = [] | ||
|
|
||
| for (const session of sessions) { | ||
| if (!session.pinned) { | ||
| continue | ||
| } | ||
|
|
||
| const id = pinIdForSession(session) | ||
|
|
||
| if (!ids.includes(id)) { | ||
| ids.push(id) | ||
| } | ||
| } | ||
|
|
||
| return ids | ||
| } | ||
| ) |
There was a problem hiding this comment.
The computed pin list only scans the currently loaded $sessions rows. After app start, reconnect, or a profile switch, a pinned session that is outside the first loaded page is not in $sessions, so sessionsToKeep() cannot preserve it and the sidebar drops that pin until a later page happens to load the row.
| export async function migrateLegacyPinnedSessions( | ||
| sessions: Array<{ id: string; _lineage_root_id?: null | string; pinned?: boolean }>, | ||
| pushPinned: (sessionId: string) => Promise<void>, | ||
| log: Pick<Console, 'info'> = console | ||
| ): Promise<number> { | ||
| const legacyIds = readLegacyPinnedSessionIds() | ||
|
|
||
| if (!arraysEqual(prev, next)) { | ||
| $pinnedSessionIds.set(next) | ||
| if (legacyIds.length === 0) { | ||
| return 0 | ||
| } | ||
|
|
||
| const serverPinnedIds = new Set(sessions.filter(session => session.pinned).map(pinIdForSession)) | ||
| const pushedIds = legacyIds.filter((id, index) => legacyIds.indexOf(id) === index && !serverPinnedIds.has(id)) |
There was a problem hiding this comment.
The migration pushes each raw local pin id through setSessionPinned(id, true) without carrying the profile scope from the session-list fetch. If the legacy pin belongs to another profile and the row is not already in $sessions, the REST fallback sends no profile, so migration can fail against the default profile or pin a same-id session in the wrong profile.
| @@ -912,7 +894,6 @@ export function useSessionActions({ | |||
| // that race after the mutation succeeds so right-click → Archive does | |||
There was a problem hiding this comment.
Archived Pins Stay Server-Side
Archiving now removes the row locally but no longer clears the server-side pinned flag. An archived session that was pinned can come back as pinned: true when a later refresh, search, or wider page includes that archived row, making the pinned section show a conversation the user just archived.
The server-side pin PR (#186) made $pinnedSessionIds a read-only computed atom, which dropped setPinnedSessionOrder + the sidebar's reorderPinned handler and hardcoded sortable={false} on the pinned list — losing drag-to-reorder. Restore it with a local $sidebarPinnedOrderIds order atom (mirrors $sidebarSessionOrderIds) layered over the server-synced pin SET via orderByIds: pin membership syncs across devices, visual order stays per-device. Keyed by durable lineage-root pin ids so order survives compression.
…nd every turn display Bridge claude-bpx #186 emits null prompt_tokens/total_tokens with prompt_tokens_unavailable/total_tokens_unavailable/unavailable discriminators. normalize_usage previously collapsed that to input_tokens=0 with no unknown flag, so unknown input was priced and displayed as a measured zero. Extends #787's output-only UNKNOWN contract to input, cache and whole-usage unknowns across normalize_usage/CanonicalUsage, pricing, the Blackbox turn ledger/rollups/cards, and the shared display rule. Measured zeros stay measured; absent-key behavior for legacy providers is unchanged. Also preserves the discriminators through both MoA physical-call serializations (aggregator + advisor), which reconstructed pricing dicts from numeric fields only and let Blackbox price an unknown input as $0.00075, and routes the console cache-stat branch, the adjacent verbose token log and the gateway thin last-turn fallback through the shared unknown rule. Verified (repo venv): 404 passed in 73.16s across usage contract/consumers, usage_pricing, billing, blackbox, turn accumulator, last-turn persistence, meta cache reporting, usage anchor, context-engine and MoA loop suites. Acceptance repro now: input_tokens_unknown=True, total_tokens_unknown=True, direct cost None, MoA blackbox cost (None, 'unknown', all-None components), console prints "Cache: unknown". Measured-zero control stays measured (all unknown flags False, cost 0.00). 5 mutation arms RED with exact-byte restore: input detection 3F, pricing refusal 25F, MoA aggregator flags 4F, MoA advisor flags 8F, console cache guard 3F. Scope: Blackbox per-turn contract + the human-facing displays it feeds. Persisted/export-schema flag loss (SessionDB/aux-ledger counters, cumulative CLI/TUI totals, Langfuse canonical export) is deferred to t_25f50547.
…tals and the Langfuse export #787 made UNKNOWN a first-class state for the Blackbox PER-TURN contract. Three other consumers still collapsed an unmeasured term into a measured zero. Bridge claude-bpx #186 is merged and fleet-deployed, so these paths receive unknown-input shapes today. Reproduced all three on the tree first: update_token_counts REJECTS unknown flags: unexpected keyword argument record_auxiliary_usage REJECTS unknown flags: unexpected keyword argument cumulative counters carry NO unknown provenance tui cache_hit_pct = 100% over prompt_total=100 (the unknown turn added 0) langfuse input-only -> ({'input': 0, 'output': 50}, {}) 1. PERSISTED SCHEMA. sessions and session_model_usage gain the five USAGE_UNKNOWN_FIELDS discriminators (+ last_turn_* snapshot mirrors), additive with DEFAULT 0 so every legacy row and every provider that never speaks this dialect reads back "measured". The int counters are untouched, so no arithmetic consumer changes. Cumulative flags are ABSORBING (MAX in SQL): one unmeasured call latches the session term, which is why a last-call-only guard is wrong here. Classified into _TOKEN_DELTA_ROUTE_FIELDS so deltas with differing unknown state never coalesce. Both session_model_usage table-rebuild paths (PK heal, v21->v22) copy the flags only when the source table has them, so a rebuild cannot silently reset a latched unknown. record_aux_usage no longer drops an all-zero unmeasured aux call as "empty". 2. CUMULATIVE CLI/TUI TOTALS. The loop latches per-term provenance onto session_*_unknown beside the session_*_tokens commit. Display reuses the Blackbox lane's rule (format_token_count / prompt_tokens_unknown) via new session_usage_unknown_flags / session_total_tokens_unknown helpers -- not a forked cumulative rule. /usage, both status-bar Sigma renderers, the CLI cache-hit segment and the TUI payload now say "unknown" or emit no ratio instead of a fabricated percentage. 3. LANGFUSE EXPORT. An unmeasured term is OMITTED from usage_details (absent, never 0) and declared on the observation metadata via _unknown_usage_details(); cost_details is empty for an unpriceable turn so Langfuse cannot read a partial subtotal as authoritative. The post_api_request summary-dict reconstruction was a second flag-loss site and now carries the discriminators. The 2-tuple return contract of _canonical_usage_and_cost / _usage_and_cost is unchanged. Verified (repo venv): - tests/agent/test_usage_unknown_persisted_export.py -- 71 pins, the five literal bridge payloads (input-only / cache-only / wholly-unavailable / output-only / measured-zero) through EACH of the three paths, plus a legacy-DB migration arm built by stripping the new columns out of the real SCHEMA_SQL. Cumulative and display arms source-LIFT the shipped blocks by AST anchor rather than re-implementing them. - scripts/run_tests.sh over the contract/consumer/persisted suites + tests/hermes_state + async token accounting + langfuse + last-turn persistence + turn-usage accumulator + blackbox: "54 files, 682 tests passed, 0 failed, 1 skipped". - 10-arm mutation matrix, each restored byte-exact, fake kills (SyntaxError/ NameError/zero-collected) rejected: 10/10 RED, control 131 passed before and after. Arms: SessionDB counter flags dropped; absorbing -> last-wins; aux ledger flags dropped; aux producer drops the unmeasured call; cumulative provenance not absorbing; CLI cache-ratio guard removed; /usage card prints numbers; TUI payload fabricates the ratio; Langfuse exports 0 for unknown; Langfuse summary-dict drops flags. - ui-tui: npm run typecheck clean. npm test is 61 files / 6 tests red on this machine, IDENTICAL on the base commit with the change stashed (textInputReturnAction + an unbuilt shared hermes-ink) -- pre-existing, not from this diff. Scope: this card only. The Blackbox per-turn contract certified under t_083ad45e / #787 is untouched.
…asured 0 claude-bpx #186 makes the bridge egress an honest unknown when its CLI transcript reconciliation misses (~40% of parallel-batch turns, measured): `completion_tokens`/`total_tokens` come back null with an explicit `usage.output_tokens_unavailable: true` discriminator. The deployed consumer coerced that null to 0 via `normalize_usage -> _usage_count -> _to_int` (`int(value or 0)`) and ignored the discriminator, so a bridge-first deploy would have written measured-looking zeros into the usage ledger, /usage, blackbox and pricing. Reproduced before fixing: the literal wire payload normalized to `output_tokens=0`, `total_tokens=150`. UNKNOWN is now a first-class state end to end: - CanonicalUsage gains `output_tokens_unknown` (+ `total_tokens_unknown`). The int fields stay ints so every arithmetic consumer is untouched; the flag is the discriminator. `__add__` makes unknown ABSORBING. - normalize_usage detects it shape-agnostically: the discriminator flag (including via the OpenAI client's `model_extra`, where an unrecognised wire key actually lands), or a PRESENT output counter that is None. An ABSENT key is NOT an unknown — providers that never speak this dialect keep normalizing to integers. - estimate_usage_cost REFUSES an unknown turn (status=unknown, amount None) rather than pricing a missing output term as $0. - blackbox: the flag rides per-call -> per-turn (absorbing) -> TurnRecord -> a new `output_tokens_unknown` column (guarded additive migration, DEFAULT 0 so every historical row reads back "measured"). The priced_zero fast path and the reprice sweep both exclude unknown rows. - The display rule is single-sourced in `usage_pricing.format_token_count` (`UNKNOWN_TOKENS_LABEL`); the alert card and the /usage+/context last-turn card both delegate to it, keeping their own magnitude formatting. The `in=/out=` API-call log says `out=unknown` instead of `out=0`. Verified: - tests/agent/test_usage_unknown_contract.py — 15 pins against the bridge's exact wire fixture (copied from bpx `unit-usage-unknown-sink.test.js` / `SPEC-parallel-stop-usage-undercount.md`): null+unavailable -> unknown, a real integer -> unchanged, a missing usage object -> unknown-free zeros. 15 passed. - 7-arm mutation matrix, each restored after: reinstating the `or 0` coercion -> 3 RED; dropping the pricing refusal -> 2 RED; store not persisting -> 1 RED; dropping the shared display rule -> 1 RED; turn_finalizer rollup hardcoded False -> 1 RED; blackbox ingest dropping the flag -> 1 RED; widening the detector to treat ABSENT as unknown -> 3 RED. Control 15/15 green before and after; tree restored. - Regression: 153 passed (usage_pricing x3, meta_usage_cache, usage_anchor, tests/blackbox), 180 passed (tests/plugins/blackbox, last_turn_usage persistence, empty_response_guard), 46 passed (turn_finalizer x5, MoA blackbox/cost-slot, turn_usage_accumulator, account/billing usage). - E2E through the real path with the literal bridge payload: normalize -> UNKNOWN, price -> status=unknown/None, ledger -> output_tokens_unknown=1, both renderers -> "unknown". Measured control: 118 out, $0.00222, output_tokens_unknown=0, both renderers -> "118".
…nd every turn display Bridge claude-bpx #186 emits null prompt_tokens/total_tokens with prompt_tokens_unavailable/total_tokens_unavailable/unavailable discriminators. normalize_usage previously collapsed that to input_tokens=0 with no unknown flag, so unknown input was priced and displayed as a measured zero. Extends #787's output-only UNKNOWN contract to input, cache and whole-usage unknowns across normalize_usage/CanonicalUsage, pricing, the Blackbox turn ledger/rollups/cards, and the shared display rule. Measured zeros stay measured; absent-key behavior for legacy providers is unchanged. Also preserves the discriminators through both MoA physical-call serializations (aggregator + advisor), which reconstructed pricing dicts from numeric fields only and let Blackbox price an unknown input as $0.00075, and routes the console cache-stat branch, the adjacent verbose token log and the gateway thin last-turn fallback through the shared unknown rule. Verified (repo venv): 404 passed in 73.16s across usage contract/consumers, usage_pricing, billing, blackbox, turn accumulator, last-turn persistence, meta cache reporting, usage anchor, context-engine and MoA loop suites. Acceptance repro now: input_tokens_unknown=True, total_tokens_unknown=True, direct cost None, MoA blackbox cost (None, 'unknown', all-None components), console prints "Cache: unknown". Measured-zero control stays measured (all unknown flags False, cost 0.00). 5 mutation arms RED with exact-byte restore: input detection 3F, pricing refusal 25F, MoA aggregator flags 4F, MoA advisor flags 8F, console cache guard 3F. Scope: Blackbox per-turn contract + the human-facing displays it feeds. Persisted/export-schema flag loss (SessionDB/aux-ledger counters, cumulative CLI/TUI totals, Langfuse canonical export) is deferred to t_25f50547.
…asured 0 claude-bpx #186 makes the bridge egress an honest unknown when its CLI transcript reconciliation misses (~40% of parallel-batch turns, measured): `completion_tokens`/`total_tokens` come back null with an explicit `usage.output_tokens_unavailable: true` discriminator. The deployed consumer coerced that null to 0 via `normalize_usage -> _usage_count -> _to_int` (`int(value or 0)`) and ignored the discriminator, so a bridge-first deploy would have written measured-looking zeros into the usage ledger, /usage, blackbox and pricing. Reproduced before fixing: the literal wire payload normalized to `output_tokens=0`, `total_tokens=150`. UNKNOWN is now a first-class state end to end: - CanonicalUsage gains `output_tokens_unknown` (+ `total_tokens_unknown`). The int fields stay ints so every arithmetic consumer is untouched; the flag is the discriminator. `__add__` makes unknown ABSORBING. - normalize_usage detects it shape-agnostically: the discriminator flag (including via the OpenAI client's `model_extra`, where an unrecognised wire key actually lands), or a PRESENT output counter that is None. An ABSENT key is NOT an unknown — providers that never speak this dialect keep normalizing to integers. - estimate_usage_cost REFUSES an unknown turn (status=unknown, amount None) rather than pricing a missing output term as $0. - blackbox: the flag rides per-call -> per-turn (absorbing) -> TurnRecord -> a new `output_tokens_unknown` column (guarded additive migration, DEFAULT 0 so every historical row reads back "measured"). The priced_zero fast path and the reprice sweep both exclude unknown rows. - The display rule is single-sourced in `usage_pricing.format_token_count` (`UNKNOWN_TOKENS_LABEL`); the alert card and the /usage+/context last-turn card both delegate to it, keeping their own magnitude formatting. The `in=/out=` API-call log says `out=unknown` instead of `out=0`. Verified: - tests/agent/test_usage_unknown_contract.py — 15 pins against the bridge's exact wire fixture (copied from bpx `unit-usage-unknown-sink.test.js` / `SPEC-parallel-stop-usage-undercount.md`): null+unavailable -> unknown, a real integer -> unchanged, a missing usage object -> unknown-free zeros. 15 passed. - 7-arm mutation matrix, each restored after: reinstating the `or 0` coercion -> 3 RED; dropping the pricing refusal -> 2 RED; store not persisting -> 1 RED; dropping the shared display rule -> 1 RED; turn_finalizer rollup hardcoded False -> 1 RED; blackbox ingest dropping the flag -> 1 RED; widening the detector to treat ABSENT as unknown -> 3 RED. Control 15/15 green before and after; tree restored. - Regression: 153 passed (usage_pricing x3, meta_usage_cache, usage_anchor, tests/blackbox), 180 passed (tests/plugins/blackbox, last_turn_usage persistence, empty_response_guard), 46 passed (turn_finalizer x5, MoA blackbox/cost-slot, turn_usage_accumulator, account/billing usage). - E2E through the real path with the literal bridge payload: normalize -> UNKNOWN, price -> status=unknown/None, ledger -> output_tokens_unknown=1, both renderers -> "unknown". Measured control: 118 out, $0.00222, output_tokens_unknown=0, both renderers -> "118". # Conflicts: # plugins/blackbox/store.py
…nd every turn display Bridge claude-bpx #186 emits null prompt_tokens/total_tokens with prompt_tokens_unavailable/total_tokens_unavailable/unavailable discriminators. normalize_usage previously collapsed that to input_tokens=0 with no unknown flag, so unknown input was priced and displayed as a measured zero. Extends #787's output-only UNKNOWN contract to input, cache and whole-usage unknowns across normalize_usage/CanonicalUsage, pricing, the Blackbox turn ledger/rollups/cards, and the shared display rule. Measured zeros stay measured; absent-key behavior for legacy providers is unchanged. Also preserves the discriminators through both MoA physical-call serializations (aggregator + advisor), which reconstructed pricing dicts from numeric fields only and let Blackbox price an unknown input as $0.00075, and routes the console cache-stat branch, the adjacent verbose token log and the gateway thin last-turn fallback through the shared unknown rule. Verified (repo venv): 404 passed in 73.16s across usage contract/consumers, usage_pricing, billing, blackbox, turn accumulator, last-turn persistence, meta cache reporting, usage anchor, context-engine and MoA loop suites. Acceptance repro now: input_tokens_unknown=True, total_tokens_unknown=True, direct cost None, MoA blackbox cost (None, 'unknown', all-None components), console prints "Cache: unknown". Measured-zero control stays measured (all unknown flags False, cost 0.00). 5 mutation arms RED with exact-byte restore: input detection 3F, pricing refusal 25F, MoA aggregator flags 4F, MoA advisor flags 8F, console cache guard 3F. Scope: Blackbox per-turn contract + the human-facing displays it feeds. Persisted/export-schema flag loss (SessionDB/aux-ledger counters, cumulative CLI/TUI totals, Langfuse canonical export) is deferred to t_25f50547.
…tals and the Langfuse export #787 made UNKNOWN a first-class state for the Blackbox PER-TURN contract. Three other consumers still collapsed an unmeasured term into a measured zero. Bridge claude-bpx #186 is merged and fleet-deployed, so these paths receive unknown-input shapes today. Reproduced all three on the tree first: update_token_counts REJECTS unknown flags: unexpected keyword argument record_auxiliary_usage REJECTS unknown flags: unexpected keyword argument cumulative counters carry NO unknown provenance tui cache_hit_pct = 100% over prompt_total=100 (the unknown turn added 0) langfuse input-only -> ({'input': 0, 'output': 50}, {}) 1. PERSISTED SCHEMA. sessions and session_model_usage gain the five USAGE_UNKNOWN_FIELDS discriminators (+ last_turn_* snapshot mirrors), additive with DEFAULT 0 so every legacy row and every provider that never speaks this dialect reads back "measured". The int counters are untouched, so no arithmetic consumer changes. Cumulative flags are ABSORBING (MAX in SQL): one unmeasured call latches the session term, which is why a last-call-only guard is wrong here. Classified into _TOKEN_DELTA_ROUTE_FIELDS so deltas with differing unknown state never coalesce. Both session_model_usage table-rebuild paths (PK heal, v21->v22) copy the flags only when the source table has them, so a rebuild cannot silently reset a latched unknown. record_aux_usage no longer drops an all-zero unmeasured aux call as "empty". 2. CUMULATIVE CLI/TUI TOTALS. The loop latches per-term provenance onto session_*_unknown beside the session_*_tokens commit. Display reuses the Blackbox lane's rule (format_token_count / prompt_tokens_unknown) via new session_usage_unknown_flags / session_total_tokens_unknown helpers -- not a forked cumulative rule. /usage, both status-bar Sigma renderers, the CLI cache-hit segment and the TUI payload now say "unknown" or emit no ratio instead of a fabricated percentage. 3. LANGFUSE EXPORT. An unmeasured term is OMITTED from usage_details (absent, never 0) and declared on the observation metadata via _unknown_usage_details(); cost_details is empty for an unpriceable turn so Langfuse cannot read a partial subtotal as authoritative. The post_api_request summary-dict reconstruction was a second flag-loss site and now carries the discriminators. The 2-tuple return contract of _canonical_usage_and_cost / _usage_and_cost is unchanged. Verified (repo venv): - tests/agent/test_usage_unknown_persisted_export.py -- 71 pins, the five literal bridge payloads (input-only / cache-only / wholly-unavailable / output-only / measured-zero) through EACH of the three paths, plus a legacy-DB migration arm built by stripping the new columns out of the real SCHEMA_SQL. Cumulative and display arms source-LIFT the shipped blocks by AST anchor rather than re-implementing them. - scripts/run_tests.sh over the contract/consumer/persisted suites + tests/hermes_state + async token accounting + langfuse + last-turn persistence + turn-usage accumulator + blackbox: "54 files, 682 tests passed, 0 failed, 1 skipped". - 10-arm mutation matrix, each restored byte-exact, fake kills (SyntaxError/ NameError/zero-collected) rejected: 10/10 RED, control 131 passed before and after. Arms: SessionDB counter flags dropped; absorbing -> last-wins; aux ledger flags dropped; aux producer drops the unmeasured call; cumulative provenance not absorbing; CLI cache-ratio guard removed; /usage card prints numbers; TUI payload fabricates the ratio; Langfuse exports 0 for unknown; Langfuse summary-dict drops flags. - ui-tui: npm run typecheck clean. npm test is 61 files / 6 tests red on this machine, IDENTICAL on the base commit with the change stashed (textInputReturnAction + an unbuilt shared hermes-ink) -- pre-existing, not from this diff. Scope: this card only. The Blackbox per-turn contract certified under t_083ad45e / #787 is untouched.
…tals and the Langfuse export #787 made UNKNOWN a first-class state for the Blackbox PER-TURN contract. Three other consumers still collapsed an unmeasured term into a measured zero. Bridge claude-bpx #186 is merged and fleet-deployed, so these paths receive unknown-input shapes today. Reproduced all three on the tree first: update_token_counts REJECTS unknown flags: unexpected keyword argument record_auxiliary_usage REJECTS unknown flags: unexpected keyword argument cumulative counters carry NO unknown provenance tui cache_hit_pct = 100% over prompt_total=100 (the unknown turn added 0) langfuse input-only -> ({'input': 0, 'output': 50}, {}) 1. PERSISTED SCHEMA. sessions and session_model_usage gain the five USAGE_UNKNOWN_FIELDS discriminators (+ last_turn_* snapshot mirrors), additive with DEFAULT 0 so every legacy row and every provider that never speaks this dialect reads back "measured". The int counters are untouched, so no arithmetic consumer changes. Cumulative flags are ABSORBING (MAX in SQL): one unmeasured call latches the session term, which is why a last-call-only guard is wrong here. Classified into _TOKEN_DELTA_ROUTE_FIELDS so deltas with differing unknown state never coalesce. Both session_model_usage table-rebuild paths (PK heal, v21->v22) copy the flags only when the source table has them, so a rebuild cannot silently reset a latched unknown. record_aux_usage no longer drops an all-zero unmeasured aux call as "empty". 2. CUMULATIVE CLI/TUI TOTALS. The loop latches per-term provenance onto session_*_unknown beside the session_*_tokens commit. Display reuses the Blackbox lane's rule (format_token_count / prompt_tokens_unknown) via new session_usage_unknown_flags / session_total_tokens_unknown helpers -- not a forked cumulative rule. /usage, both status-bar Sigma renderers, the CLI cache-hit segment and the TUI payload now say "unknown" or emit no ratio instead of a fabricated percentage. 3. LANGFUSE EXPORT. An unmeasured term is OMITTED from usage_details (absent, never 0) and declared on the observation metadata via _unknown_usage_details(); cost_details is empty for an unpriceable turn so Langfuse cannot read a partial subtotal as authoritative. The post_api_request summary-dict reconstruction was a second flag-loss site and now carries the discriminators. The 2-tuple return contract of _canonical_usage_and_cost / _usage_and_cost is unchanged. Verified (repo venv): - tests/agent/test_usage_unknown_persisted_export.py -- 71 pins, the five literal bridge payloads (input-only / cache-only / wholly-unavailable / output-only / measured-zero) through EACH of the three paths, plus a legacy-DB migration arm built by stripping the new columns out of the real SCHEMA_SQL. Cumulative and display arms source-LIFT the shipped blocks by AST anchor rather than re-implementing them. - scripts/run_tests.sh over the contract/consumer/persisted suites + tests/hermes_state + async token accounting + langfuse + last-turn persistence + turn-usage accumulator + blackbox: "54 files, 682 tests passed, 0 failed, 1 skipped". - 10-arm mutation matrix, each restored byte-exact, fake kills (SyntaxError/ NameError/zero-collected) rejected: 10/10 RED, control 131 passed before and after. Arms: SessionDB counter flags dropped; absorbing -> last-wins; aux ledger flags dropped; aux producer drops the unmeasured call; cumulative provenance not absorbing; CLI cache-ratio guard removed; /usage card prints numbers; TUI payload fabricates the ratio; Langfuse exports 0 for unknown; Langfuse summary-dict drops flags. - ui-tui: npm run typecheck clean. npm test is 61 files / 6 tests red on this machine, IDENTICAL on the base commit with the change stashed (textInputReturnAction + an unbuilt shared hermes-ink) -- pre-existing, not from this diff. Scope: this card only. The Blackbox per-turn contract certified under t_083ad45e / #787 is untouched.
…tals and the Langfuse export (#797) * fix(usage): carry UNKNOWN through the persisted schema, cumulative totals and the Langfuse export #787 made UNKNOWN a first-class state for the Blackbox PER-TURN contract. Three other consumers still collapsed an unmeasured term into a measured zero. Bridge claude-bpx #186 is merged and fleet-deployed, so these paths receive unknown-input shapes today. Reproduced all three on the tree first: update_token_counts REJECTS unknown flags: unexpected keyword argument record_auxiliary_usage REJECTS unknown flags: unexpected keyword argument cumulative counters carry NO unknown provenance tui cache_hit_pct = 100% over prompt_total=100 (the unknown turn added 0) langfuse input-only -> ({'input': 0, 'output': 50}, {}) 1. PERSISTED SCHEMA. sessions and session_model_usage gain the five USAGE_UNKNOWN_FIELDS discriminators (+ last_turn_* snapshot mirrors), additive with DEFAULT 0 so every legacy row and every provider that never speaks this dialect reads back "measured". The int counters are untouched, so no arithmetic consumer changes. Cumulative flags are ABSORBING (MAX in SQL): one unmeasured call latches the session term, which is why a last-call-only guard is wrong here. Classified into _TOKEN_DELTA_ROUTE_FIELDS so deltas with differing unknown state never coalesce. Both session_model_usage table-rebuild paths (PK heal, v21->v22) copy the flags only when the source table has them, so a rebuild cannot silently reset a latched unknown. record_aux_usage no longer drops an all-zero unmeasured aux call as "empty". 2. CUMULATIVE CLI/TUI TOTALS. The loop latches per-term provenance onto session_*_unknown beside the session_*_tokens commit. Display reuses the Blackbox lane's rule (format_token_count / prompt_tokens_unknown) via new session_usage_unknown_flags / session_total_tokens_unknown helpers -- not a forked cumulative rule. /usage, both status-bar Sigma renderers, the CLI cache-hit segment and the TUI payload now say "unknown" or emit no ratio instead of a fabricated percentage. 3. LANGFUSE EXPORT. An unmeasured term is OMITTED from usage_details (absent, never 0) and declared on the observation metadata via _unknown_usage_details(); cost_details is empty for an unpriceable turn so Langfuse cannot read a partial subtotal as authoritative. The post_api_request summary-dict reconstruction was a second flag-loss site and now carries the discriminators. The 2-tuple return contract of _canonical_usage_and_cost / _usage_and_cost is unchanged. Verified (repo venv): - tests/agent/test_usage_unknown_persisted_export.py -- 71 pins, the five literal bridge payloads (input-only / cache-only / wholly-unavailable / output-only / measured-zero) through EACH of the three paths, plus a legacy-DB migration arm built by stripping the new columns out of the real SCHEMA_SQL. Cumulative and display arms source-LIFT the shipped blocks by AST anchor rather than re-implementing them. - scripts/run_tests.sh over the contract/consumer/persisted suites + tests/hermes_state + async token accounting + langfuse + last-turn persistence + turn-usage accumulator + blackbox: "54 files, 682 tests passed, 0 failed, 1 skipped". - 10-arm mutation matrix, each restored byte-exact, fake kills (SyntaxError/ NameError/zero-collected) rejected: 10/10 RED, control 131 passed before and after. Arms: SessionDB counter flags dropped; absorbing -> last-wins; aux ledger flags dropped; aux producer drops the unmeasured call; cumulative provenance not absorbing; CLI cache-ratio guard removed; /usage card prints numbers; TUI payload fabricates the ratio; Langfuse exports 0 for unknown; Langfuse summary-dict drops flags. - ui-tui: npm run typecheck clean. npm test is 61 files / 6 tests red on this machine, IDENTICAL on the base commit with the change stashed (textInputReturnAction + an unbuilt shared hermes-ink) -- pre-existing, not from this diff. Scope: this card only. The Blackbox per-turn contract certified under t_083ad45e / #787 is untouched. * fix(usage): gate cumulative cache ratios at the CLI render seam Run real snapshot and renderer regression tests for five bridge shapes. Verified 56 files: 743 passed, 0 failed, 1 skipped; producer and renderer guard mutants both fail. * fix(usage): preserve unknown in remaining cumulative CLI totals CLASS-SWEEP found /status raw total and zero-counter Sigma omission. Reproduced /status failure, then verified 3 focused tests pass; prior broad run passed 745 with one test fixture import error subsequently fixed. * test(usage): cover null container persistence and export controls * fix(usage): port #797 onto the rebased #787 -- one latch, coherent snapshot flags, append-only schema Re-stacked hermes-agent#797 onto #787 (rebased onto main 2026-09-23). Two of its six commits were already on #787 under the same subjects and were dropped; the other four replayed with semantic conflicts. Fixes on top, each measured red before: 1. Duplicate absorbing latch: #787's "latch the flag over the aggregate" and #797's "carry UNKNOWN through cumulative totals" each added the same session_* latch loop to agent/conversation_loop.py. Kept #787's (beside the increments); removed #797's copy. 25 AST-seam tests ("shipped cumulative-usage seam moved") go green. 2. last_turn_*_unknown flags now move WITH the snapshot they qualify: when #787's _last_turn_snapshot_kwargs withholds an unmeasured call's split (None -> COALESCE keeps the prior real split), the flags are withheld too. Before: the persisted card rendered "Total: unknown" over a real 4,120 from the previous turn. 3. Schema columns APPENDED, not inserted mid-table. #797 put its 10 sessions flags before git_branch and its 5 session_model_usage flags before estimated_cost_usd. Migrated DBs get them via ALTER ADD (end of table), so fresh and migrated DBs had different physical orders, and the positional lost-and-found recovery mapper misplaced historical fields (title landed in the wrong column). Fresh-DB column prefix is now byte-identical to main's (title back at index 43). 4. Recovery classifier accepts session_model_usage records of 18 (pre-migration) AND 23 fields; it previously dropped every post-migration usage row. Test pins 23 and adds a 23-field classify case. 979 passed across every test file this branch touches + hermes_state + session recovery suites. Apollo merge pass 2026-09-23. --------- Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
Partial revert of 0df54d7 (audit t_712c13f2, DROP, D9 mixed desktop row): only the server half is in scope; apps/desktop is taken from upstream at sync. Of #186's server half, hermes_state.set_session_pinned / include_pinned and the web_server pinned PATCH are now upstream's own (hermes_state_sessions.py:952, web_routers/sessions.py:755 on upstream/main; fork callers sessions_cmd.py, web_routers/profiles.py, api_server.py, prune tests depend on them) — not revertable, superseded-by-upstream. The only fork-only remnant is the tui_gateway session.pin JSON-RPC method + its 3 tests. _session_live_pinned stays: session.info/session.list 'pinned' is in upstream's contract (tui_gateway/contracts/common.py:122). Verified: test-gate pytest tests/test_tui_gateway_server.py -> 669 passed.
Partial revert of 0df54d7 (audit t_712c13f2, DROP, D9 mixed desktop row): only the server half is in scope; apps/desktop is taken from upstream at sync. Of #186's server half, hermes_state.set_session_pinned / include_pinned and the web_server pinned PATCH are now upstream's own (hermes_state_sessions.py:952, web_routers/sessions.py:755 on upstream/main; fork callers sessions_cmd.py, web_routers/profiles.py, api_server.py, prune tests depend on them) — not revertable, superseded-by-upstream. The only fork-only remnant is the tui_gateway session.pin JSON-RPC method + its 3 tests. _session_live_pinned stays: session.info/session.list 'pinned' is in upstream's contract (tui_gateway/contracts/common.py:122). Verified: test-gate pytest tests/test_tui_gateway_server.py -> 669 passed.
Partial revert of 0df54d7 (audit t_712c13f2, DROP, D9 mixed desktop row): only the server half is in scope; apps/desktop is taken from upstream at sync. Of #186's server half, hermes_state.set_session_pinned / include_pinned and the web_server pinned PATCH are now upstream's own (hermes_state_sessions.py:952, web_routers/sessions.py:755 on upstream/main; fork callers sessions_cmd.py, web_routers/profiles.py, api_server.py, prune tests depend on them) — not revertable, superseded-by-upstream. The only fork-only remnant is the tui_gateway session.pin JSON-RPC method + its 3 tests. _session_live_pinned stays: session.info/session.list 'pinned' is in upstream's contract (tui_gateway/contracts/common.py:122). Verified: test-gate pytest tests/test_tui_gateway_server.py -> 669 passed.
Partial revert of 0df54d7 (audit t_712c13f2, DROP, D9 mixed desktop row): only the server half is in scope; apps/desktop is taken from upstream at sync. Of #186's server half, hermes_state.set_session_pinned / include_pinned and the web_server pinned PATCH are now upstream's own (hermes_state_sessions.py:952, web_routers/sessions.py:755 on upstream/main; fork callers sessions_cmd.py, web_routers/profiles.py, api_server.py, prune tests depend on them) — not revertable, superseded-by-upstream. The only fork-only remnant is the tui_gateway session.pin JSON-RPC method + its 3 tests. _session_live_pinned stays: session.info/session.list 'pinned' is in upstream's contract (tui_gateway/contracts/common.py:122). Verified: test-gate pytest tests/test_tui_gateway_server.py -> 669 passed.
Fixes the multi-device pin desync Ace flagged. docs/desktop/2026-07-02-pinned-session-sync-PRD.md (v0.2). Additive
pinnedcolumn (archived precedent, app migration), singlesession.pin{pinned}RPC mirroring session.title (+_emit_session_info live cross-device), client $pinnedSessionIds now a computed derived-view (server truth wins on reconnect), migration = per-session idempotent ADD (never set-replace) + push-count log + retry-on-fail. 621 gateway pytest + 25 desktop tests green, typecheck+eslint clean. Reviewed by Apollo (lint auto-fix amended).