Skip to content

feat(buzz-acp): give each channel thread its own agent session - #6732

Merged
salman1993 merged 11 commits into
mainfrom
codex/thread-scoped-acp-sessions
Aug 31, 2026
Merged

feat(buzz-acp): give each channel thread its own agent session#6732
salman1993 merged 11 commits into
mainfrom
codex/thread-scoped-acp-sessions

Conversation

@salman1993

@salman1993 salman1993 commented Aug 24, 2026

Copy link
Copy Markdown
Contributor

What this does

In a channel, people often run several unrelated conversations at once (separate threads). Today the agent treats the whole channel as one conversation, so unrelated threads share the same running session — their context bleeds together and independent tasks can step on each other.

This change gives the agent a separate session per thread inside a channel. Direct messages stay as one conversation (unchanged). The channel is still the boundary for who is allowed in and what is visible — only the agent's working context is now split by thread.

How it is turned on

Off by default. Operators opt in with one setting:

  • BUZZ_ACP_SESSION_POLICY=channel — default, current behavior
  • BUZZ_ACP_SESSION_POLICY=thread — new per-thread behavior

Being behind a flag means we can enable it for a few agents, watch how it behaves, and roll back instantly without a code change.

Key design decisions

  • Decide the thread once, up front. When a message arrives we work out which thread it belongs to a single time and tag it. Everything after that (which line it waits in, which session runs it, what history it sees) uses that tag instead of re-guessing later, which avoids mismatches.
  • Default stays identical to today. Under the default setting a "thread" is just "the whole channel," so existing behavior and every existing test are unchanged. The new, riskier behavior is strictly opt-in.
  • Give the agent only its thread's history. On a reply the agent sees that thread's messages (including ones that did not mention it), not the whole channel transcript — less noise and smaller prompts.
  • Don't let one channel use more memory than before. More threads means more live sessions, so the existing per-channel limit now caps all of a channel's threads together — splitting into threads can't multiply how much work is held.

Bugs found and fixed while iterating (from review)

  • Same thread, two sessions. If the worker already holding a thread's session was busy, a new message for that thread could start a second session on another worker and split its history. Now it waits for the right worker instead of forking.
  • Interrupting the wrong thread. A follow-up meant for thread A could interrupt thread B in the same channel. Interrupts now target the exact thread.
  • Stuck thread after a crash. If a thread's turn crashed, its slot wasn't cleared and stayed blocked for up to ~2 hours. It now clears right away and retries.
  • Lost the original request. When a thread was interrupted and then had to wait for a busy worker, only the follow-up was kept and the original request was dropped. The full request is now preserved on retry.
  • Same thread seen as two. Two spellings of the same thread id (upper/lower case) could be treated as different threads. Normalized so they count as one.

Not in this PR

Testing

The full buzz-acp test suite passes (830+ unit and integration tests), plus new focused tests for thread routing, session reuse, interrupt targeting, crash recovery, and request preservation. Behavior with the flag off is unchanged.

@salman1993 salman1993 changed the title feat(buzz-acp): thread-scoped ACP sessions — foundation (SessionScope + operator policy) feat(buzz-acp): thread-scoped ACP sessions — backend (Steps 1–4) Aug 25, 2026
@salman1993 salman1993 changed the title feat(buzz-acp): thread-scoped ACP sessions — backend (Steps 1–4) feat(buzz-acp): thread-scoped ACP sessions — backend Aug 25, 2026
@salman1993

Copy link
Copy Markdown
Contributor Author

🤖 Request changes at exact head 73e87eaa55c267d1cb6793c8db1eba0a001f911b.

Major — equivalent NIP-10 root spellings still split one relay thread into separate ACP sessions

crates/buzz-acp/src/scope.rs:115-119 stores the parsed root_event_id string verbatim in the SessionScope hash key. The shared parser accepts all ASCII hex, including uppercase, and preserves the original spelling (crates/buzz-core/src/nip10.rs:49-50,72-75). Relay ingest, however, decodes the same marker to bytes before resolving ancestry (crates/buzz-relay/src/handlers/ingest.rs:826-853), and accepted events are fanned out with their original signed tags intact.

Concrete scenario: a valid reply tags root AB… and a later reply tags the same root ab…. Relay ingest accepts both as the same root bytes, but ACP derives two unequal SessionScope::Thread keys. Under session_policy=thread, that splits queue/in-flight state, provider sessions, worker affinity, and delivery ledgers for one canonical relay thread. It violates this PR's core acceptance criterion that repeated activity under one canonical root reuses exactly one session and can also duplicate context delivery.

Normalize validated marker IDs before constructing the key (or key on nostr::EventId/bytes), and add a mixed-case regression covering scope equality/session reuse.

Verification

  • Full cargo test -p buzz-acp: pass — 827 lib + 9 integration tests.
  • cargo fmt --all -- --check: pass.
  • cargo clippy -p buzz-acp --all-targets -- -D warnings: pass.
  • Independently verified live-local at this exact SHA with a real relay and ACP harness under session_policy=thread: two lowercase canonical roots created distinct sessions and direct/nested replies reused the correct original session with scope-correct reply destinations and no observed cross-thread prompt mixing. Evidence: .scratch/fastvalidator-live-73e87eaa5/{live-observation.log,live-thread-scenario.log,fake-acp-wire.log,acp-launchd.log}. Artifacts self-attest the SHA and postdate the commit.
  • The mixed-case defect is proven by static runtime-path inspection; that exact edge case was not exercised live. DMs were not exercised live in this pass.

salman1993 added a commit that referenced this pull request Aug 25, 2026
…re one scope

The shared NIP-10 parser accepts and preserves uppercase ASCII hex in `e`-tag
marker ids (`is_ascii_hexdigit`), but the relay decodes event ids to bytes on
ingest — so a reply tagging its root as `AB…` and one tagging `ab…` are the
SAME accepted relay thread. `SessionScope::Thread` keyed on the raw string,
so under thread policy those equivalent spellings hashed to different keys and
split one thread across two ACP sessions (queue partitions, provider sessions,
worker affinity, and delivery ledgers), violating same-root reuse.

Normalize the resolved root id to lowercase in `SessionScope::derive` before it
becomes the scope key. `nostr::EventId::to_hex()` is already lowercase, so the
top-level-mention path is unaffected. Adds a mixed-case regression test.

Reported in review of PR #6732.

Signed-off-by: Salman Mohammed <smohammed@squareup.com>
@salman1993

Copy link
Copy Markdown
Contributor Author

🤖 Re-review: ready at exact head f8eaa71763baa9fa7d07be95a066aab383c7e0e2. (GitHub does not permit this account to formally approve its own PR.)

The prior mixed-case NIP-10 blocker is resolved. SessionScope::derive now canonicalizes the validated root ID before constructing the scope key, so case-equivalent relay roots share queue, in-flight, provider-session, affinity, and delivery state. The new regression test is causal: removing only the lowercase normalization makes scope::tests::mixed_case_root_spellings_share_one_thread_scope fail.

Fresh verification at this SHA:

  • cargo test -p buzz-acp: 828 lib + 9 integration passed
  • cargo fmt --all -- --check: passed
  • cargo clippy -p buzz-acp --all-targets -- -D warnings: passed
  • CI: all reported checks passed/skipped
  • Live local relay + ACP (BUZZ_ACP_SESSION_POLICY=thread): lowercase and uppercase spellings were both accepted, resolved to the same lowercase scope, and sequential turns reused the same provider session. Each trigger appeared once in its prompt.

No actionable findings remain. DMs were not re-exercised live in this pass.

salman1993 added a commit that referenced this pull request Aug 26, 2026
…ting

Addresses three correctness gaps found in review of PR #6732.

1. Duplicate provider session for a busy thread (pool.rs, lib.rs).
   Worker affinity only scanned idle slots, so while the worker that owns a
   thread's session was checked out, a new message for that thread could be
   handed to another idle worker, forking a second session and splitting the
   thread's history/tool context. Added an authoritative
   `SessionScope -> worker` directory (`session_owners`) that survives while a
   worker is checked out. `dispatch_pending` now holds a batch (leaves it
   queued) when its session owner is busy, instead of forking a duplicate; the
   held batch dispatches to that exact worker when it returns. The directory is
   pruned on channel-wide session invalidation, and stale entries (rotation /
   crash) self-heal on the next dispatch.

2. Mid-turn steer/interrupt could target the wrong thread (lib.rs).
   The native-steer fallback and the steer-ack fallback routed by channel via
   `signal_in_flight_task`, which picks the first task for the channel — so a
   message in thread A could interrupt thread B in the same channel. Added
   `signal_in_flight_task_for_scope` (exact `SessionScope` match) and used it
   for both mid-turn fallbacks. The deferred channel-level control paths
   (`!cancel`, `!rotate`, observer `cancel_turn` / `switch_model`) keep
   channel-targeting intentionally.

3. Panicked thread stayed "in flight" (lib.rs).
   `recover_panicked_agent` called `mark_complete(channel_id)`, which resolves
   to `Conversation(channel_id)` via IntoScope and, under thread policy, left
   the real `Thread(...)` entry wedged in-flight until the ~2h backstop —
   blocking the batch it had just requeued. It now uses `meta.scope`.

Tests: scope-exact signalling targets only the matching thread; a busy session
owner holds the batch instead of forking a session (and the directory prunes on
channel invalidation); panic recovery frees the exact Thread scope and requeues
its batch. Full buzz-acp suite (831 lib + integration) green; clippy and fmt
clean.

Signed-off-by: Salman Mohammed <smohammed@squareup.com>
salman1993 added a commit that referenced this pull request Aug 26, 2026
…hausted batch

`requeue_preserve_timestamps` restored only `batch.events`, silently dropping
`batch.cancelled_events` and `cancel_reason`. Both callers pass a full
`FlushBatch` that can carry an interrupted turn's original request:

- the new busy-owner affinity hold (thread scoping), and
- the pre-existing "no agent available" pool-exhausted path.

Reachable loss: thread A is interrupted (Buzz retains A's original request to
re-prompt "original + follow-up"); older work in thread B grabs A's
session-owning worker; A's merged batch is held; only the follow-up was put
back and the original request vanished.

The helper now restores the entire batch — cancelled carryover is returned to
the pending cancelled-batches (ahead of any concurrently staged carryover) with
its reason, so the next flush reconstructs the same merged prompt. Adds a
queue-level round-trip regression asserting events + cancelled_events +
cancel_reason all survive requeue -> mark_complete -> flush.

Reported in review of PR #6732. Full buzz-acp suite (832 lib + integration)
green; clippy and fmt clean.

Signed-off-by: Salman Mohammed <smohammed@squareup.com>
@salman1993

Copy link
Copy Markdown
Contributor Author

Good catch on the requeue retry path. Confirmed it's pre-existing (same behavior on main, and independent of the thread-scoping flag), so I'm keeping it out of this PR to avoid mixing concerns and will track it as a separate follow-up — factoring one shared "restore the whole batch" helper used by both retry paths, with the same round-trip test.

@salman1993 salman1993 changed the title feat(buzz-acp): thread-scoped ACP sessions — backend feat(buzz-acp): give each channel thread its own agent session Aug 26, 2026
@salman1993

Copy link
Copy Markdown
Contributor Author

Tested PR head 326bd455 against the live relay using the PR-built buzz-acp and an instrumented ACP provider. I ran once with BUZZ_ACP_SESSION_POLICY=channel and once with thread, sending two top-level messages plus a reply to the first thread via the CLI.

  • channel: 1 ACP session for both threads and the reply.
  • thread: 2 ACP sessions for the 2 threads; the reply reused the first thread’s session.

This validates the harness/session routing. I did not test the full Desktop UI path.

@salman1993
salman1993 marked this pull request as ready for review August 27, 2026 17:01
@salman1993
salman1993 requested a review from a team as a code owner August 27, 2026 17:01

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: f8f8808890

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

scope.telemetry_label()
);
agent.state.sessions.insert(*cid, sid.clone());
agent.state.sessions.insert(scope.clone(), sid.clone());

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Bound retained sessions under thread policy

When BUZZ_ACP_SESSION_POLICY=thread, every sequential top-level mention gets a unique scope and this insertion retains the resulting provider session indefinitely; no queue cap evicts entries from SessionState.sessions. A busy channel can therefore accumulate an unbounded number of ACP sessions even with no queued work, and providers such as buzz-agent spawn a separate MCP registry for every session/new, eventually exhausting processes and memory. Add a bounded per-channel/session eviction policy that also closes or invalidates the provider session.

Useful? React with 👍 / 👎.

Comment thread crates/buzz-acp/src/lib.rs Outdated
Comment thread crates/buzz-acp/src/lib.rs Outdated
@salman1993 salman1993 closed this Aug 27, 2026
@salman1993 salman1993 reopened this Aug 27, 2026
@github-actions

github-actions Bot commented Aug 27, 2026

Copy link
Copy Markdown

🔐 Codex Security Review

Status: review required for the current range.

The current range is 2f3dd850db3afe27e56f18cbcd3548eabdd9b9c2...a089e89c4ca340a07e795f879560477e4d5d5096.
A new review must complete for this exact range. When manual authorization
is required, a Block organization member must comment exactly
@buzz-security-review a089e89c4ca340a07e795f879560477e4d5d5096 to authorize a new review.
Any previous review applies only to its recorded range.

@salman1993

Copy link
Copy Markdown
Contributor Author

Thanks for the review. Addressed in 8daa4405f:

P2 — !cancel / !rotate wrong thread (fixed). Both routed through signal_in_flight_task, which matches on channel_id and takes the first map entry, so with two threads running in one channel they could tear down an arbitrary thread. They now derive the SessionScope from the command event's NIP-10 tags (same resolver admission uses) and target it via signal_in_flight_task_for_scope — the scope-exact primitive steering already uses. The idle !rotate path invalidates only that scope via the new invalidate_scope_session rather than the whole channel. signal_in_flight_task is now used only by the desktop observer frames (cancel_turn / switch_model), which carry a bare channelId and no thread context.

P2 — channel-keyed typing state (fixed). typing_channels is now keyed by SessionScope: dispatch_pending returns the scope, completion / panic / membership-removal clear the exact scope (new PromptSource::scope() accessor), and the refresh loop publishes one indicator per active thread with that thread's tags. Concurrent thread turns no longer overwrite or prematurely clear each other's indicator.

Both are only reachable under thread policy; the default channel policy is byte-for-byte unchanged (scope = the channel's sole conversation). New tests: invalidate_scope_session_targets_one_thread_and_drops_its_owner and prompt_source_scope_exposes_thread_scope_and_none_for_heartbeat.

P1 — unbounded session / provider retention (split out). Agreed this isn't a clean PR-local fix — partial eviction of Buzz-side IDs would leak provider resources. It's pre-existing (not thread-scoping-specific) and needs a designed lifecycle: idle-TTL / LRU on the scope maps, a real per-session close contract that tears down McpRegistry / provider resources, and a finite BUZZ_AGENT_MAX_SESSIONS default. Tracked separately in #6958.

@salman1993
salman1993 force-pushed the codex/thread-scoped-acp-sessions branch 2 times, most recently from c409bc9 to 1e2a507 Compare August 31, 2026 18:54
Lay the foundation for thread-scoped ACP sessions (rollout step 1: land
behind the operator policy with channel scope as the fallback).

- Add `scope` module with a hashable `SessionScope`
  (Conversation/Thread) and `SessionPolicy` (channel/thread). Scope is
  derived once at admission from policy + DM status + NIP-10 thread tags
  using the shared `buzz_core::nip10` canonical-root rules. DMs are always
  conversation-scoped; under the default `channel` policy every channel
  event collapses to a conversation scope, preserving today's behavior.
- Add `--session-policy` / `BUZZ_ACP_SESSION_POLICY` (default `channel`)
  wired through `CliArgs` -> `Config` and the config summary; document it
  in `.env.example`.
- Derive and log the resolved scope at event admission (telemetry only for
  now).

Unit tests cover scope derivation for top-level mentions, direct/nested
replies, repeated mentions, DMs, malformed thread tags, hashing/map-key
use, and policy parsing/defaults.

Follow-up (same ticket, later commits): partition queue/in-flight state by
scope (step 2), partition provider-session state by scope (step 3),
scope-correct context gathering (step 4), and the Settings->Experiments
toggle + managed-agent deployment wiring (step 5).

Signed-off-by: Salman Mohammed <smohammed@squareup.com>
…step 2)

Carry the admission-time SessionScope through the queue instead of keying
everything on channel_id. All EventQueue partitions — pending queues,
in-flight tracking, deadlines, retry counters/backoff, cancelled-batch
carryover, and the goose-native steer side table — are now keyed by
SessionScope. Under the default `channel` policy every scope is
Conversation{channel_id}, so behavior is byte-identical; under `thread`
policy, distinct canonical roots in one channel become independent
partitions.

- QueuedEvent and FlushBatch carry the resolved SessionScope; dispatch,
  completion, and requeue route by it. PromptSource::Channel now carries
  the scope (with a channel_id() accessor) so a completed turn marks the
  exact scope complete. No code rederives scope from the last event in a
  batch.
- The mid-turn steer gate is now scope-level (is_scope_in_flight), and the
  native-steer withhold/release/dedup + deadline extension target the
  scope, so an unrelated thread's in-flight turn never steers this one.
- drain_channel performs channel-wide cleanup across every child thread
  scope. Backlog protection is preserved with a per-scope cap plus an
  aggregate per-channel cap, so per-thread partitioning cannot multiply the
  admitted queue size.
- An IntoScope helper lets the queue API accept a bare channel Uuid
  (conversation scope) or an explicit SessionScope, keeping the existing
  channel-keyed unit tests intact.

Pool provider-session STATE is still channel-keyed in this commit (indexed
via scope.channel_id()); step 3 rekeys it by scope so repeated activity in
a thread reuses exactly that thread's provider session.

New queue tests: two threads in one channel are independent partitions,
events from different roots never share a batch, an in-flight scope blocks
only that scope, channel drain clears every child thread scope, and the
aggregate channel cap is not multiplied by threads. Full buzz-acp suite
(819 lib + integration) green; clippy and fmt clean.

Signed-off-by: Salman Mohammed <smohammed@squareup.com>
…p 3)

Key the pool's provider-session state by SessionScope instead of only
channel_id, so repeated activity in a thread reuses exactly that thread's
provider session and unrelated threads in one channel never share session
state, turn counters, context-delivery markers, or delivery-dedup state.

- SessionState maps (sessions, turn_counts, core_sections, canvas_sections,
  deliveries) are now keyed by SessionScope. run_prompt_task resolves the
  session, core/canvas sections, standing-context-sent marker, turn count,
  and delivery ledger by the batch's scope; channel-level fetches (canvas,
  huddle, title, resolve) still use scope.channel_id().
- Scope-to-worker affinity: has_session_for/try_claim match the exact scope,
  so a temporarily busy worker cannot cause another worker to open a
  duplicate session for the same thread. TaskMeta carries the scope;
  send_steer and record_successful_steer route by it.
- Channel-wide cleanup preserved: invalidate_channel clears every child
  thread scope for a channel (returns the count); invalidate_channel_sessions
  and the removed-channel path use it. Added invalidate_scope for
  single-session invalidation and mark_scope_delivery_success.
- Model-switch targeting stays channel-level with an explicit TODO for
  channel-vs-thread control targeting (same open question as top-level
  !cancel / !rotate, deferred per the ticket).

New pool tests: two threads in one channel get distinct sessions and reuse
per root, invalidate_scope leaves a sibling thread untouched, and
invalidate_channel clears every thread scope while sparing other channels.
Full buzz-acp suite (822 lib + integration) green; clippy and fmt clean.

Signed-off-by: Salman Mohammed <smohammed@squareup.com>
Drive conversation-context gathering from the batch's resolved SessionScope
instead of re-inferring it from whichever event is last in the batch.

- A Thread scope fetches only that canonical thread's history (all messages
  under the root, including intervening non-mention human messages), so no
  unrelated channel transcript is injected. A brand-new thread's first turn
  has no prior history — the trigger is delivered as the [Event] block.
- Conversation scope preserves current behavior: DMs (and legacy
  channel-policy channels) fetch the reply chain for a threaded reply or
  recent DM history for a non-reply; a plain top-level channel message gets
  no supplementary context.
- The existing scope-keyed delivery-delta filter (step 3) then strips events
  this session already received, so subsequent turns deliver only the
  intervening same-thread messages plus the trigger, without duplication.

The routing decision is extracted into a pure `resolve_context_target` so it
is unit-tested directly: thread scope wins over a divergent last-event tag,
a new top-level thread resolves to its own root, a plain conversation-scope
channel message gets no context, DM non-reply fetches DM history, and a
conversation-scope reply uses its reply chain.

Together with steps 1–3 this makes BUZZ_ACP_SESSION_POLICY=thread deliver
real end-to-end isolation: distinct roots in one channel get distinct queue
partitions (step 2), distinct provider sessions with scope-keyed worker
affinity (step 3), and distinct canonical-thread context (this step); while
the default `channel` policy is byte-identical to prior behavior. Full
buzz-acp suite (827 lib + integration) green; clippy and fmt clean.

Signed-off-by: Salman Mohammed <smohammed@squareup.com>
…re one scope

The shared NIP-10 parser accepts and preserves uppercase ASCII hex in `e`-tag
marker ids (`is_ascii_hexdigit`), but the relay decodes event ids to bytes on
ingest — so a reply tagging its root as `AB…` and one tagging `ab…` are the
SAME accepted relay thread. `SessionScope::Thread` keyed on the raw string,
so under thread policy those equivalent spellings hashed to different keys and
split one thread across two ACP sessions (queue partitions, provider sessions,
worker affinity, and delivery ledgers), violating same-root reuse.

Normalize the resolved root id to lowercase in `SessionScope::derive` before it
becomes the scope key. `nostr::EventId::to_hex()` is already lowercase, so the
top-level-mention path is unaffected. Adds a mixed-case regression test.

Reported in review of PR #6732.

Signed-off-by: Salman Mohammed <smohammed@squareup.com>
…ting

Addresses three correctness gaps found in review of PR #6732.

1. Duplicate provider session for a busy thread (pool.rs, lib.rs).
   Worker affinity only scanned idle slots, so while the worker that owns a
   thread's session was checked out, a new message for that thread could be
   handed to another idle worker, forking a second session and splitting the
   thread's history/tool context. Added an authoritative
   `SessionScope -> worker` directory (`session_owners`) that survives while a
   worker is checked out. `dispatch_pending` now holds a batch (leaves it
   queued) when its session owner is busy, instead of forking a duplicate; the
   held batch dispatches to that exact worker when it returns. The directory is
   pruned on channel-wide session invalidation, and stale entries (rotation /
   crash) self-heal on the next dispatch.

2. Mid-turn steer/interrupt could target the wrong thread (lib.rs).
   The native-steer fallback and the steer-ack fallback routed by channel via
   `signal_in_flight_task`, which picks the first task for the channel — so a
   message in thread A could interrupt thread B in the same channel. Added
   `signal_in_flight_task_for_scope` (exact `SessionScope` match) and used it
   for both mid-turn fallbacks. The deferred channel-level control paths
   (`!cancel`, `!rotate`, observer `cancel_turn` / `switch_model`) keep
   channel-targeting intentionally.

3. Panicked thread stayed "in flight" (lib.rs).
   `recover_panicked_agent` called `mark_complete(channel_id)`, which resolves
   to `Conversation(channel_id)` via IntoScope and, under thread policy, left
   the real `Thread(...)` entry wedged in-flight until the ~2h backstop —
   blocking the batch it had just requeued. It now uses `meta.scope`.

Tests: scope-exact signalling targets only the matching thread; a busy session
owner holds the batch instead of forking a session (and the directory prunes on
channel invalidation); panic recovery frees the exact Thread scope and requeues
its batch. Full buzz-acp suite (831 lib + integration) green; clippy and fmt
clean.

Signed-off-by: Salman Mohammed <smohammed@squareup.com>
…hausted batch

`requeue_preserve_timestamps` restored only `batch.events`, silently dropping
`batch.cancelled_events` and `cancel_reason`. Both callers pass a full
`FlushBatch` that can carry an interrupted turn's original request:

- the new busy-owner affinity hold (thread scoping), and
- the pre-existing "no agent available" pool-exhausted path.

Reachable loss: thread A is interrupted (Buzz retains A's original request to
re-prompt "original + follow-up"); older work in thread B grabs A's
session-owning worker; A's merged batch is held; only the follow-up was put
back and the original request vanished.

The helper now restores the entire batch — cancelled carryover is returned to
the pending cancelled-batches (ahead of any concurrently staged carryover) with
its reason, so the next flush reconstructs the same merged prompt. Adds a
queue-level round-trip regression asserting events + cancelled_events +
cancel_reason all survive requeue -> mark_complete -> flush.

Reported in review of PR #6732. Full buzz-acp suite (832 lib + integration)
green; clippy and fmt clean.

Signed-off-by: Salman Mohammed <smohammed@squareup.com>
Signed-off-by: Salman Mohammed <smohammed@squareup.com>
salman1993 and others added 3 commits August 31, 2026 16:40
Two thread-isolation correctness bugs surfaced in review, both only
reachable under BUZZ_ACP_SESSION_POLICY=thread; behavior under the default
channel policy is unchanged (the scope is the channel's sole conversation).

!cancel / !rotate could hit the wrong thread. Both routed through
signal_in_flight_task, which matches on channel_id and takes the first
HashMap entry — so with two threads running concurrently in one channel an
owner's !cancel/!rotate could tear down an arbitrary thread's turn. They
now derive the SessionScope from the command event's NIP-10 tags (the same
resolver admission uses) and target it via signal_in_flight_task_for_scope,
the scope-exact primitive steering already uses. The idle !rotate path now
invalidates only that scope via the new AgentPool::invalidate_scope_session
instead of the whole channel. signal_in_flight_task is now used only by the
desktop observer control frames (cancel_turn / switch_model), which carry a
bare channelId and no thread context.

Typing state was channel-keyed. typing_channels keyed by channel_id, so two
concurrent thread turns overwrote one entry and either turn's completion
removed it — the indicator could reflect the wrong thread or stop while a
sibling turn continued. It is now keyed by SessionScope: dispatch_pending
returns the scope, completion/panic/ownership-removal clear the exact scope
(via the new PromptSource::scope accessor), and the refresh loop publishes
one indicator per active thread carrying that thread's NIP-10 tags.

Tests: invalidate_scope_session targets one thread and drops its owner;
PromptSource::scope exposes the thread scope and None for heartbeats.

The reviewer's P1 (unbounded session/provider-resource retention) is a
pre-existing lifecycle-hardening concern, not thread-scoping-specific, and
is tracked separately rather than adding partial eviction here.

Signed-off-by: Salman Mohammed <smohammed@squareup.com>
Match session framing and titles to the configured scope. Reject ambiguous channel controls and correlate cancel acknowledgements with their requests.

Co-authored-by: Salman Mohammed <smohammed@squareup.com>
Signed-off-by: Leo <5faf251baee50ee6bcde338aef6acdd70bb3e60115664c2cd490d94a55dfc488@buzz.block.builderlab.xyz>
Signed-off-by: Salman Mohammed <smohammed@squareup.com>
@salman1993
salman1993 force-pushed the codex/thread-scoped-acp-sessions branch from 1e2a507 to a089e89 Compare August 31, 2026 20:44

@wpfleger96 wpfleger96 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🤖 I think the thread-scoping behavior and rollout gate are working as intended, but there is one lifecycle issue to address before enabling this.

BUZZ_ACP_SESSION_POLICY defaults to channel, so the existing one-session-per-channel behavior remains the default; thread is opt-in and DMs remain conversation-scoped. Runtime verification against a fresh relay confirmed both sides of that boundary: the default reused one provider session across two roots, while thread mode created distinct sessions, reused the correct session for a follow-up, and kept sibling-thread prompts isolated.

The blocker is retained-session cardinality. Thread mode changes the long-lived state from roughly one provider session per channel to one per encountered thread. SessionState retains sessions, turn_counts, context sections, and delivery state by SessionScope, and cleanup currently depends on explicit rotation/invalidation, channel removal, or process exit. There is no LRU, TTL, or maximum retained-scope policy. Because buzz-agent also permits unlimited sessions by default and ACP sessions own provider/MCP resources, a channel with many one-off roots can grow memory and child-process usage for the lifetime of the agent.

Please add bounded idle-session eviction with provider-side cleanup, or establish and enforce a safe end-to-end session cap before rollout. I’d also fix the small coverage gap while touching this: test_session_policy_env_var_parses currently passes --session-policy=thread rather than exercising BUZZ_ACP_SESSION_POLICY.

The remaining source paths and current CI checks look good at a089e89c4ca340a07e795f879560477e4d5d5096.

@wpfleger96 wpfleger96 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🤖 Approved. The missing LRU/TTL is pre-existing; this PR increases the reachable session cardinality under the opt-in thread policy, but the feature is default-off and its intended behavior passed source review, CI, and live relay/ACP verification. I’m treating lifecycle bounding as separate hardening rather than a merge blocker.

@salman1993
salman1993 merged commit 674c173 into main Aug 31, 2026
42 checks passed
@salman1993
salman1993 deleted the codex/thread-scoped-acp-sessions branch August 31, 2026 22:17
wpfleger96 pushed a commit that referenced this pull request Sep 1, 2026
…enericize

* origin/main:
  feat(desktop): add isolated named demo builds (#6407)
  fix(model-capabilities): humanize databricks goose model names (#7135)
  feat(db): add NIP-FI identity and final-admission schema foundation (#6994)
  feat(buzz-acp): give each channel thread its own agent session (#6732)
  docs: add review-proven failure-path & async-state rules to AGENTS.md (#7061)

Signed-off-by: Duncan <dcfd242e557282d7a1e2cf2e6877522682f1e5c6156dc92ca7d90eaedd3b0f95@buzz.block.builderlab.xyz>
wpfleger96 added a commit that referenced this pull request Sep 1, 2026
…h-coordinator

* origin/main:
  feat(desktop): add isolated named demo builds (#6407)
  fix(model-capabilities): humanize databricks goose model names (#7135)
  feat(db): add NIP-FI identity and final-admission schema foundation (#6994)
  feat(buzz-acp): give each channel thread its own agent session (#6732)
  docs: add review-proven failure-path & async-state rules to AGENTS.md (#7061)
  fix(desktop): back split thread headers (#7137)

Signed-off-by: Duncan <dcfd242e557282d7a1e2cf2e6877522682f1e5c6156dc92ca7d90eaedd3b0f95@buzz.block.builderlab.xyz>
Co-authored-by: Will Pfleger <pfleger.will@gmail.com>
Signed-off-by: Will Pfleger <pfleger.will@gmail.com>
johnmatthewtennant added a commit that referenced this pull request Sep 1, 2026
…-channel-permissions

* origin/main:
  fix(model-capabilities): humanize databricks goose model names (#7135)
  feat(db): add NIP-FI identity and final-admission schema foundation (#6994)
  feat(buzz-acp): give each channel thread its own agent session (#6732)
  docs: add review-proven failure-path & async-state rules to AGENTS.md (#7061)
  fix(desktop): back split thread headers (#7137)
  add public descriptions to agent personas (#7126)
  feat(desktop): add protected-build Bestie experiment (#6902)
  fix(relay): reject a frame on its own acknowledgement channel (#6961)
  fix(acp): wake agents from workflow messages (#6953)

Signed-off-by: John Tennant <jtennant@squareup.com>

# Conflicts:
#	crates/buzz-db/src/runtime/migration.rs
#	schema/schema.sql
shivchander added a commit to shivchander/os1 that referenced this pull request Sep 1, 2026
* fix: retrieving cold memories; add regression task (#6950)

## Why

Evaluating buzz agent memory retrieval by seeding a memory then asking
the buzz agent a question it needs that memory.

**Bug Found**: System prompt had no inclusion of retrieving cold
memories and suggested looking in a mem/*.md directory that does not
exist. Updated `system-prompt.md` to include memory CLI tools and usage.

Eval Before System Prompt Change: 0/3 
Eval After System Prompt Change: 3/3 

## What

- Add a `memory-retrieval` benchmark that seeds agent memory with `buzz
mem set` before asking a direct question.
- Grade the observable threaded answer without inspecting tool calls or
exposing the answer in channel history.
- Teach agents to use `buzz mem set`, `buzz mem ls`, and `buzz mem get`
for cold memory.
- Add a wire-debug endpoint configuration for diagnosing ACP tool calls
in local runs.
- Add fixture, seeding, verifier, and prompt coverage.

## Risk Assessment

Low. The runtime changes are limited to the benchmark harness. The
production-facing change clarifies existing memory commands in the base
prompt; it does not change memory storage, relay behavior, or
authorization.

## References

- Before the system-prompt changes, 0/3 attempts passed because agents
never invoked the `buzz mem` CLI and instead searched a non existent
filesystem
- After the changes, 3/3 attempts passed. ACP wire logs confirmed that
every agent ran `buzz mem ls` followed by `buzz mem get` and returned
`net_gpv`.

---------

Signed-off-by: Philip Azar <pazar@squareup.com>

* fix(ci): salvage Codex review output on PTY-shutdown hang (#7042)

Codex CLI can leave a PTY descendant holding the action's inherited
stdio after the turn completes. The `runCodexExec.ts` wrapper waits on a
`close` event that never fires, so the `Review pull request` step hangs
until the job timeout kills it — discarding the finished review the CLI
already wrote to disk.

The CLI writes the completed review to the `--output-last-message` file
(exposed as `output-file`) **before** the hang. This PR adds a salvage
step that recovers it, and sets the step and job timeouts to preserve
the full 30-minute Codex execution budget.

**Changes (`codex-security-review.yml`):**

- Add `output-file: ${{ runner.temp }}/codex-review.json` to the `Review
pull request` step so the CLI writes the result before the hang.
(`runner` context is valid in `steps.with`; not in `jobs.env`.)
- Add `timeout-minutes: 30` and `continue-on-error: true` to the Codex
step — a hang now costs ≤30 minutes instead of 40, and the salvage step
still runs.
- Set job `timeout-minutes: 40` to give setup, step cancellation, and
salvage sufficient headroom without colliding with the Codex execution
budget. The original 30-minute job timeout was too narrow: evidence from
run
[33114428326](https://github.com/block/buzz/actions/runs/33114428326/job/98665369165)
shows completed output appearing 28m46s after step start, meaning a
20-minute step timeout could kill a legitimate review before the salvage
file exists.
- Add a `Salvage review output` step with `if: always()`: prefers
`steps.run_codex.outputs.final-message` on a clean exit; falls back to
the output file when the step timed out. The output file path is set in
the step's own `env` block (`CODEX_OUTPUT_FILE: ${{ runner.temp
}}/codex-review.json`), where `runner` is valid. Validates shape
(non-empty JSON object, has `overall_risk`); fails the job hard if
neither source is present.
- Wire the job `outputs.review_json` to
`steps.salvage.outputs.review_json`.

**Changes (`Justfile`, `ci.yml`):**

- Add `actionlint .github/workflows/codex-security-review.yml` to
`security-review-check` so expression-validity errors are caught
locally.
- Provision `actionlint` via Hermit (pinned v1.7.12) rather than a
one-off `Install actionlint` curl step, so the same binary is used
locally and in CI.

**Security posture is unchanged:** the salvage step reads the action's
own output and a file written to `runner.temp` — neither is
PR-controlled. Credential-stripping env block on the Codex step is
untouched.


Note this is a temporary workaround until
https://github.com/openai/codex-action/issues/169 is addressed

---------

Signed-off-by: Will Pfleger <pfleger.will@gmail.com>
Co-authored-by: Duncan <dcfd242e557282d7a1e2cf2e6877522682f1e5c6156dc92ca7d90eaedd3b0f95@buzz.block.builderlab.xyz>

* feat: render agent avatars as squircles (#7106)

## Summary

- render every agent/AI identity as a 30% squircle across desktop and
mobile while keeping human avatars circular
- propagate agent identity through message, thread, profile, reaction,
member, DM, search, workflow, project, huddle, forum, pulse, and
agent-management surfaces
- preserve squircle geometry for fallbacks, focus/status treatments,
add-agent controls, and overlapping avatar outlines (`calc(30% + 2px)`
for the outer background)

### Related issue

None found. This change was requested and visually reviewed in the
originating Buzz thread.

### Testing

- `just desktop-test` — 5,799 passed
- `just mobile-test` — 2,008 passed
- pre-push gates passed at `0d59d77b120dcb90aac2f918e422c11c9fa5353b`:
desktop check, TypeScript typecheck, desktop full test suite, mobile
format/analyze and full test suite, Rust tests, Tauri checks, and
differential file-size gate
- deterministic desktop visual sweep covered channel messages/thread
summaries; thread, subthread, and sub-subthread depths; reactions and
reactor popovers; hover/full profiles; added-to-channel activity;
channel members/settings; agent library/team overlaps; agent creation;
mention autocomplete; and DM header/sidebar/settings

### UI evidence

The complete labeled visual matrix is available in the originating Buzz
review thread. GitHub-hosted copies will be added in a follow-up PR
comment using the repository screenshot script.

---------

Signed-off-by: Fizz <dae5f6af70b8695a8b83c8deae555f63be41630ec2b8cd493e41a439c9527dd8@buzz.block.builderlab.xyz>
Co-authored-by: Fizz <dae5f6af70b8695a8b83c8deae555f63be41630ec2b8cd493e41a439c9527dd8@buzz.block.builderlab.xyz>

* fix(acp): wake agents from workflow messages (#6953)

> Pinky, an AI agent, is opening this PR on Wes's behalf.

## Summary

Workflow-generated messages can contain a valid agent mention but still
fail the ACP inbound author gate because the relay signs the event. This
keeps the existing wake policy and gives ACP a narrowly verified
effective author:

- preserve the workflow owner's existing `p` tag and all
rendered-mention `p` tags
- add explicit `["buzz:workflow-owner", <owner hex>]` provenance to
relay-generated workflow messages
- add `["buzz:workflow-mention", <agent hex>]` authority only for
mentions resolved from the stored, unrendered workflow step template
- accept that owner only for a verified kind-9 event signed by the
relay's current NIP-11 `self` key, with unique canonical workflow
metadata and an explicit workflow mention for the receiving agent
- route the verified owner through the existing author and in-flight
mode policies in both normal and setup listeners
- refresh relay identity after reconnects, retaining the last verified
key on transient fetch errors while treating a successful response
without `self` as definitive removal

Malformed, duplicate, forged, tampered, wrong-kind, and wrong-relay
attribution all fail closed to the raw event signer. `respond-to=nobody`
remains absolute. Old/mixed-version messages without the explicit
provenance retain their current fail-closed behavior.

## Trust boundary

The workflow owner means **“scheduled by,” not “authored every rendered
word.”** Trigger-controlled substitutions may still produce ordinary `p`
mention routing for compatibility, but they cannot mint
`buzz:workflow-mention` authority. Only a target named in the durable
owner-authored step template can receive that authority.

The author gate is not bypassed: after relay signature/provenance
verification, the effective owner is evaluated under the same
`owner-only`, `allowlist`, DM, and `nobody` policies used for ordinary
messages. Owner control commands continue to use the raw event signer.

## Why this PR

This is the focused immediate fix for waking an **online** agent from a
stored workflow mention. Earlier attempts were not a finished mergeable
fix and had materially different or incomplete trust designs. Larry's
larger draft stack addresses durable delivery across restarts; that
remains valuable future work and can supersede this effective-author
path when it lands.

## Validation

At exact clean commit `fe5b55619fe44176343eefb4cb7fe180df45a7d8`:

- `buzz-relay workflow_sink`: 25/25 passed, including all four ignored
PostgreSQL cases
- `buzz-acp --lib`: 845/845 passed
- `buzz-workflow --lib`: 169/169 passed (2 unrelated PostgreSQL tests
ignored)
- warnings-denied Clippy passed for the changed Rust packages
- `cargo fmt --all -- --check` passed
- `git diff --check` passed
- repository pre-push gates passed, including branch-scoped Rust tests
- CI now selects the ACP library tests and the relay's pure + PostgreSQL
workflow-sink tests so these guards cannot silently remain unexecuted

The production event-to-author gate is shared by normal and setup
listeners and has biting regression tests for accepted explicit
attribution, legacy owner-`p` rejection, and forged-attribution
rejection.

## Exact-head local relay + ACP proof

Following the release-binary/local-relay shape in `TESTING.md`, the
exact commit above passed a fresh isolated real-process matrix using:

- a freshly recreated Postgres database with migrations
- isolated Redis
- exact-head release `buzz-relay`, `buzz`, `buzz-admin`, and `buzz-acp`
binaries
- newly provisioned owner, channel, and bot member through the CLI
- workflow creation and triggering through the running relay
- a deterministic ACP protocol subprocess capturing actual
`session/prompt` dispatches
- a NIP-11 `self` value verified against the running relay signer

Cases:

1. A stored explicit workflow mention woke an `owner-only` agent exactly
once.
2. A workflow message without an agent mention did not wake it.
3. A non-relay signer forging every workflow authority tag did not wake
it.
4. Trigger-controlled `{{trigger.text}}` containing `@Wake Agent`
retained ordinary `p` routing but received no authority-bearing
workflow-mention tag and did not wake the agent.
5. `respond-to=nobody` remained absolute for a valid relay-authenticated
workflow mention.

The deterministic ACP subprocess isolates and directly proves relay →
ACP authorization and prompt dispatch without depending on external
model behavior.

## Deployment and residual risk

Relay and ACP changes must be deployed together for the new wake
behavior; mixed versions fail closed. Production paired-deployment proof
remains distinct from the successful local integration run. Setup-mode
behavior has automated coverage but was not a separate case in the
five-case local matrix. Relay-key rotation is observed at ACP
startup/reconnect; transient NIP-11 errors retain the last verified key,
an intentional availability tradeoff documented in code.

---------

Signed-off-by: Pinky <5f5ab050ec58ae208332edd544ebf705221e24c1b86d82a6ca07038a7a8f6ac9@buzz.block.builderlab.xyz>
Signed-off-by: Brain <1a02c72794dcd0f07058a353bc3a81f4028b8c77c92c87fce6d5c8b85970a20b@buzz.block.builderlab.xyz>
Signed-off-by: Wes <wesbillman@users.noreply.github.com>
Co-authored-by: Pinky <5f5ab050ec58ae208332edd544ebf705221e24c1b86d82a6ca07038a7a8f6ac9@buzz.block.builderlab.xyz>
Co-authored-by: LioLionel <62820906+LioLionel@users.noreply.github.com>
Co-authored-by: Brain <1a02c72794dcd0f07058a353bc3a81f4028b8c77c92c87fce6d5c8b85970a20b@buzz.block.builderlab.xyz>
Co-authored-by: Carl <9d00794d3df50972eb8b615511783cab12a77a8fd5dd5edd58073ec73b54bd8b@buzz.block.builderlab.xyz>

* fix(relay): reject a frame on its own acknowledgement channel (#6961)

Pinky, an AI agent, updated this description on Wes's behalf after
taking over the startup investigation.

**Category:** fix

**User Impact:** An EVENT refused by WebSocket admission or handler
saturation receives a correlated `OK(event_id, false, reason)` instead
of an uncorrelated NOTICE, so the client can settle that refusal without
waiting for its publish timeout. Rate-limited refusals also arm client
backoff. This fixes a protocol failure mechanism; it does not establish
that every startup send will succeed or that the reported Desktop
startup incident is fully resolved.

**Problem:** Startup opens several live subscriptions and publishes at
once, and the relay's WebSocket admission gate is a fixed 5-second
window (`ws_admission_budget` = `human_ws_events_per_sec * 5`). If that
shared per-principal quota is exhausted, `enforce_ws_admission`
previously rejected an EVENT with a bare `["NOTICE", reason]`. Quota
pressure is a possible trigger, not proof of the original incident's
complete cause.

A NOTICE carries no event id. Both clients settle a pending publish
*only* from an `OK` keyed by event id (desktop `pendingEvents`, mobile
`_pendingEvents`), so nothing settled — and `handle_text_message`
returns early, so no `OK` ever followed either. The send **could not
fail**; it could only time out at `PUBLISH_TIMEOUT_MS` = 25s. That
explains how this rejection mechanism can produce a roughly 25-second
timeout; attributing the original report to it still requires the actual
startup/send workflow.

The handler-semaphore saturation path had the identical defect, and that
one needs no quota burst to fire.

**Solution:** NIP-01 gives each request type its own acknowledgement
channel, and a rejection is only actionable on the same one. Reject a
REQ with `CLOSED`, an EVENT with `OK(id, false, reason)`, and fall back
to `NOTICE` only where no per-request correlation exists. COUNT refusals
now also use `CLOSED(query_id, reason)` per NIP-45, covering both quota
admission and handler saturation (added in
`cd12c93804b87a24b61075dfd171dc471a0a527f`).

Reason strings are unchanged, so the `rate-limited:` prefix and `retry
in {N}s` hint that existing client gates parse keep working (desktop
`parseRateLimitHint`, mobile `RelayRateLimitGate`, buzz-acp
`set_rate_limit_gate`). Only the frame *type* changes, so
`docs/multi-tenant-relay.md` L7 stays satisfied.

Two notes on how this landed, both worth a reviewer's attention:

1. **A survived mutation became a design change.**
`send_admission_result` originally took a `RejectionTarget` parameter,
and reverting the *second* call site (the per-minute message quota)
survived the whole suite — with Redis unreachable the first quota check
short-circuits, so that line is unreachable in test. Rather than test
around it, the parameter is gone: the target is derived from the frame,
so no call site can name the wrong channel.

2. **The relay fix would have caused a client regression on its own.**
Gate arming lived only in the NOTICE branch. Once rejections arrive as
`OK:false`, `handleOk` failed the send without ever backing off — the
client would retry straight into the same quota. Desktop and Mobile now
arm on a `rate-limited:` OK rejection. ACP was subsequently fixed in
`3b06dd32493596ec650f20abf8805791c50fdc24`: it arms the gate and
re-parks only the refused observer frame, preserving other in-flight
frames. Desktop gets `activateRateLimitIfSignalled` as the single owner
of that prefix test, called from both `handleOk` and the NOTICE branch.

<details>
<summary>File changes</summary>

**crates/buzz-relay/src/rejection.rs** (new)
Owns the admission-rejection concern: `RejectionTarget`,
`rejection_target_for`, `request_rejection_message`,
`send_admission_result`, and `enforce_ws_admission`, moved out of
`connection.rs`. Six tests, two of which drive the real
`enforce_ws_admission` against a real `AppState`.

**crates/buzz-relay/src/connection.rs**
Fix the EVENT handler-semaphore rejection to correlate to the event id;
delegate admission to the new module. Add two tests that drive the real
`handle_text_message` with every handler permit held. Down from 1319 to
1116 lines.

**crates/buzz-relay/src/state.rs**
Widen the existing `test_state` helper to `pub(crate)` so the rejection
tests reuse it rather than adding a ninth copy of `AppState`
construction.

**desktop/src/shared/api/relayRateLimitGate.ts**
Add `activateRateLimitIfSignalled` — one owner for the `rate-limited:`
prefix test, since three inbound frame types now carry it.

**desktop/src/shared/api/relayClientSession.ts**
Arm the gate on a rate-limited OK rejection; route the NOTICE branch
through the same helper. Net zero lines, which keeps this
already-oversized file within the differential ratchet.

**desktop/src/shared/api/relayClientPublishRejection.test.mjs** (new)
Four tests against the real `RelayClient`: a rate-limited OK settles the
pending publish and arms the gate; an ordinary rejection does not arm
it; an accepted OK still resolves.

**mobile/lib/shared/relay/relay_session.dart**
Arm the gate in `_handleOk` for a rate-limited rejection.

**mobile/test/shared/relay/relay_session_test.dart**
Two tests driving the real `publish` + `debugHandleMessage` path.

</details>

<details>
<summary>Validation</summary>

**Mutation-tested — 5 mutations, all now killed.** Each production call
site was reverted to the defective behaviour to confirm a test fails.
This caught two false-negative tests:

| # | Mutation | Result |
|---|----------|--------|
| 1 | `rejection_target_for`: EVENT → `Connection` | 4 tests fail |
| 2 | EVENT handler-semaphore call site → bare NOTICE | **survived at
first** |
| 3 | per-minute quota call site → `Connection` | **survived**; fixed by
removing the parameter |
| 4 | desktop `handleOk` gate arming removed | 1 test fails |
| 5 | mobile `_handleOk` gate arming removed | 1 test fails |

Mutation 2 is the lesson: my first saturation test called
`request_rejection_message` directly, so reverting the real call site
inside the `match` arm left it green. It now drives
`handle_text_message` itself and dies on that mutation.

- `cargo test -p buzz-relay` — 928 passed, 1 failed:
`api::mesh_demo::tests::demo_join_forwarded_arm_round_trips_echo`,
**pre-existing**, reproduced with all changes stashed at `4dd4d73de`.
- `cd desktop && npm test` — 5721 passed, 0 failed (full suite).
- `cd mobile && flutter test` — 1876 passed, 0 failed (full suite).
- `just fmt-check`, `just clippy`, `just desktop-check`, `just
mobile-check`, `just file-size-check` — clean. Desktop's 5 biome
warnings are pre-existing (reproduced with changes stashed).
- All 9 pre-push lanes green, including `rust-tests` and
`desktop-tauri-checks`.

**Not verified:** not reproduced end-to-end against a live relay under a
forced quota burst. The causal chain is source-proven and
mutation-proven at the frame level; the ~25s attribution follows from
`PUBLISH_TIMEOUT_MS` but is not directly measured. A packaged-build
click-through would close that gap.

</details>

Related work: #6957 bounds Desktop HTTP event submission, but safe
retained-operation recovery after exhausted/ambiguous outcomes remains
unfinished. #6998 is the separately reviewable Desktop
readiness/duplicate-subscription slice. Neither is claimed to complete
native before/after startup-send validation.

Diagnosis note: `RESEARCH/DESKTOP_STARTUP_SEND_STALL_2026_08_27.md`
(Brain's workspace).

## Current review disposition (2026-08-28)

The [review on
`cd12c938`](https://github.com/block/buzz/pull/6961#pullrequestreview-5052902510)
identified ACP's missing rate-limited-OK handling. Commit
`3b06dd32493596ec650f20abf8805791c50fdc24` fixes gate arming, re-parking
the specifically refused observer frame, and the stale NOTICE comment.
Two regressions drive the real frame dispatcher. See [the implementation
and validation
response](https://github.com/block/buzz/pull/6961#issuecomment-5455032054).

The Mobile generation-check inline thread is resolved: its `async
publish` returns a failed Future when superseded; it does not throw
synchronously at invocation. No further production change was indicated
by that comment.

The validation counts above describe the original slice, not a new
rerun. At `3b06dd324`, the current GitHub check rollup has successful
completed test/build checks (non-applicable jobs skipped). The
security-review comment still requires review for the current base/head
range; do not read a green authorization job as a completed security
review. Approval and merge remain human decisions.

---------

Signed-off-by: Brain <1a02c72794dcd0f07058a353bc3a81f4028b8c77c92c87fce6d5c8b85970a20b@buzz.block.builderlab.xyz>
Signed-off-by: Wes <wesbillman@users.noreply.github.com>
Co-authored-by: Brain <1a02c72794dcd0f07058a353bc3a81f4028b8c77c92c87fce6d5c8b85970a20b@buzz.block.builderlab.xyz>
Co-authored-by: Carl <9d00794d3df50972eb8b615511783cab12a77a8fd5dd5edd58073ec73b54bd8b@buzz.block.builderlab.xyz>

* feat(desktop): add protected-build Bestie experiment (#6902)

## Summary

Introduces a protected-build boundary for the default-off Bestie
experiment without adding any Bestie product surface.

- Official OSS builds select an empty protected-feature module and emit
no Bestie/Chief metadata or implementation content.
- Protected internal builds select a separate module graph containing
the Bestie experiment definition.
- Within an internal build, Bestie remains disabled until the user opts
in under Settings → Experiments.
- The production build runs an artifact matrix and fails if OSS output
contains protected content or internal output lacks the Bestie manifest.

## Build contract

| Build variant | User opt-in | Result |
| --- | --- | --- |
| Official OSS | Any/forged | Bestie absent from the compiled artifact |
| Protected internal | Off | Bestie available but disabled |
| Protected internal | On | Bestie enabled |

The companion protected-release change is squareup/buzz-releases#91. It
sets `VITE_BUZZ_BESTIE=1`, requires that exact value, forwards it into
the signed macOS build, and asserts the contract in release validation.

## Why this is separate

This gives later Bestie PRs one build-selected import seam. Protected
implementations must be reachable only from the internal module so they
never enter the official OSS module graph.

## Non-goals

- No Bestie persona or provisioning
- No sidebar, app-chrome, or message-toolbar UI
- No entitlement or secrecy claim: the source is public; this boundary
controls official Block artifacts

## Verification

- Exact commit `523cf49ced03cba9be43836a54d6aa5d6923cc82`
- Full `just ci`: 5,673 Desktop tests, 2,773 Tauri tests, 1,860 mobile
tests, Rust/Tauri/web/mobile static checks and builds
- OSS production artifact: scanner confirms no `Bestie`, `Chief of
Staff`, or `builtin:bestie` content
- Internal production artifact: scanner confirms the protected Bestie
manifest is emitted
- Both build orders verified; `dist` retains the requested variant for
Vite/Tauri packaging

---------

Signed-off-by: Arjun Mahanti <arjun@squareup.com>
Signed-off-by: Fizz <fizz@buzz.local>
Signed-off-by: Fizz <dae5f6af70b8695a8b83c8deae555f63be41630ec2b8cd493e41a439c9527dd8@buzz.block.builderlab.xyz>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Fizz <fizz@buzz.local>
Co-authored-by: Fizz <dae5f6af70b8695a8b83c8deae555f63be41630ec2b8cd493e41a439c9527dd8@buzz.block.builderlab.xyz>

* add public descriptions to agent personas (#7126)

**Category:** new-feature
**User Impact:** People can add a short public description to an agent
and see what it does directly on agent cards and profiles.

**Problem:** Agent cards previously showed only a model label, so people
had to open an agent and inspect its instructions to understand its
purpose. Public metadata also needed one trustworthy lifecycle across
local edits, relay catalogs, profiles, and portable snapshots.

**Solution:** Add an optional owner-authored description with a
280-character visible-text policy, publish it as profile `about`, and
prefer it on agent cards while retaining the model fallback. Description
metadata is excluded from the spawn-content hash, remains
definition-owned, and is validated independently at every untrusted or
persistence boundary.

<details>
<summary>File changes</summary>

**desktop/src-tauri/src/commands/agent_config_tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.

**desktop/src-tauri/src/commands/agent_discovery/relay_directory.rs**
Updates relay-directory profile test publication for the expanded
profile contract.

**desktop/src-tauri/src/commands/agent_models_tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.

**desktop/src-tauri/src/commands/agent_models_update.rs**
Preserves the effective `about` value when instance edits republish a
complete profile event.

**desktop/src-tauri/src/commands/agents.rs**
Carries the effective authored description into initial managed-agent
profile publication.

**desktop/src-tauri/src/commands/agents_profile.rs**
Adds `about` to profile reconciliation and keeps description, name, and
avatar synchronized against relay state.

**desktop/src-tauri/src/commands/agents_tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.

**desktop/src-tauri/src/commands/personas/card.rs**
Materializes the definition-owned description before minting a portable
agent card snapshot.

**desktop/src-tauri/src/commands/personas/create.rs**
Normalizes and validates raw authored descriptions before persona
persistence.

**desktop/src-tauri/src/commands/personas/delete_cascade_tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.

**desktop/src-tauri/src/commands/personas/inbound.rs**
Validates descriptions at inbound relay ingress and applies accepted
values to local definitions.


**desktop/src-tauri/src/commands/personas/inbound/catalog_reconcile_tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.

**desktop/src-tauri/src/commands/personas/inbound/inbound_tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.

**desktop/src-tauri/src/commands/personas/mod.rs**
Centralizes raw-byte validation followed by trim/empty normalization for
description writes.

**desktop/src-tauri/src/commands/personas/pending.rs**
Revalidates descriptions before preparing public persona publications.

**desktop/src-tauri/src/commands/personas/sharing.rs**
Carries the optional public description through this managed-agent
compatibility path.

**desktop/src-tauri/src/commands/personas/snapshot.rs**
Materializes definition-owned descriptions into portable instance
snapshots without creating a second persisted authority.

**desktop/src-tauri/src/commands/personas/snapshot/fidelity_tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.

**desktop/src-tauri/src/commands/personas/snapshot/import.rs**
Restores snapshot descriptions onto imported definitions while keeping
linked instance copies absent.

**desktop/src-tauri/src/commands/personas/snapshot/tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.

**desktop/src-tauri/src/commands/personas/update.rs**
Persists persona description edits, republishes linked profiles, and
preserves legacy avatars during complete kind:0 replacements.


**desktop/src-tauri/src/commands/personas/update/name_propagation_tests.rs**
Proves description-only profile sync does not write instance state or
clear a legacy avatar.

**desktop/src-tauri/src/commands/team_snapshot.rs**
Round-trips member descriptions through team snapshots and imported
definitions.

**desktop/src-tauri/src/commands/team_snapshot/tests.rs**
Covers team member description export and import fidelity.

**desktop/src-tauri/src/commands/teams/adopt/apply.rs**
Starts adopted team catalog members without synthesizing an unauthored
description.

**desktop/src-tauri/src/commands/teams/adopt/tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.

**desktop/src-tauri/src/commands/teams/pending/tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.

**desktop/src-tauri/src/commands/teams/sharing/tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.

**desktop/src-tauri/src/egress_guard_tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.

**desktop/src-tauri/src/event_sync_team_catalog_tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.

**desktop/src-tauri/src/managed_agents/agent_description.rs**
Defines the canonical Rust description resolution used by profile
publication and reconciliation.

**desktop/src-tauri/src/managed_agents/agent_events.rs**
Updates managed-agent record construction for the optional public
description field.

**desktop/src-tauri/src/managed_agents/agent_snapshot.rs**
Includes descriptions as snapshot profile `about` metadata and validates
them at decode ingress.

**desktop/src-tauri/src/managed_agents/agent_snapshot_envelope.rs**
Updates managed-agent record construction for the optional public
description field.

**desktop/src-tauri/src/managed_agents/agent_snapshot_tests.rs**
Covers snapshot description export and rejection of unsafe or overlong
imported metadata.

**desktop/src-tauri/src/managed_agents/config_bridge/reader_tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.

**desktop/src-tauri/src/managed_agents/definition_validation.rs**
Adds the shared 280-character visible-text policy for public
descriptions.

**desktop/src-tauri/src/managed_agents/discovery/tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.

**desktop/src-tauri/src/managed_agents/effective_config/tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.

**desktop/src-tauri/src/managed_agents/global_config/tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.

**desktop/src-tauri/src/managed_agents/mod.rs**
Exports the description resolution and validation helpers to
managed-agent consumers.

**desktop/src-tauri/src/managed_agents/nest/render_tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.

**desktop/src-tauri/src/managed_agents/parallelism.rs**
Updates managed-agent fixtures for the optional description field
without changing runtime configuration behavior.

**desktop/src-tauri/src/managed_agents/persona_events.rs**
Adds description to persona event content while deliberately excluding
it from the spawn-relevant content hash.

**desktop/src-tauri/src/managed_agents/persona_events/tests.rs**
Pins description event round-tripping and proves description-only edits
do not change the restart hash.

**desktop/src-tauri/src/managed_agents/personas.rs**
Initializes built-in persona records without authored descriptions for
backward-compatible defaults.

**desktop/src-tauri/src/managed_agents/personas/tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.

**desktop/src-tauri/src/managed_agents/readiness.rs**
Updates managed-agent fixtures for the optional description field
without changing runtime configuration behavior.

**desktop/src-tauri/src/managed_agents/restore.rs**
Includes the effective description in launch-time profile
reconciliation.

**desktop/src-tauri/src/managed_agents/runtime/test_fixtures.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.

**desktop/src-tauri/src/managed_agents/runtime/tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.

**desktop/src-tauri/src/managed_agents/spawn_snapshot/tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.

**desktop/src-tauri/src/managed_agents/team_catalog/tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.

**desktop/src-tauri/src/managed_agents/team_snapshot.rs**
Updates managed-agent record construction for the optional public
description field.

**desktop/src-tauri/src/managed_agents/teams_tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.

**desktop/src-tauri/src/managed_agents/types.rs**
Adds optional description metadata to persona and managed-agent records
and their compatibility projections.

**desktop/src-tauri/src/managed_agents/types/requests.rs**
Accepts optional descriptions on persona create and update IPC requests.

**desktop/src-tauri/src/managed_agents/types/tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.

**desktop/src-tauri/src/migration_avatar_tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.

**desktop/src-tauri/src/persona_catalog.rs**
Parses and validates descriptions at the untrusted community-catalog
boundary.

**desktop/src-tauri/src/persona_catalog_tests.rs**
Covers valid catalog descriptions plus rejection of malformed,
invisible, and overlong values.

**desktop/src-tauri/src/relay.rs**
Publishes and queries kind:0 `about` so relay profiles preserve authored
descriptions.

**desktop/src-tauri/src/relay/tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.

**desktop/src/features/agents/AGENTS.md**
Documents description ownership, validation, snapshot, hashing, and
display invariants for future changes.

**desktop/src/features/agents/lib/agentDescription.test.mjs**
Pins Unicode counting, paste clamping, trimming, and empty
authored-description behavior.

**desktop/src/features/agents/lib/agentDescription.ts**
Provides shared display resolution, Unicode-scalar counting, and paste
clamping for descriptions.

**desktop/src/features/agents/lib/personaCatalogRelay.ts**
Maps validated catalog descriptions into catalog persona projections.

**desktop/src/features/agents/ui/AgentDefinitionDialog.tsx**
Adds the description draft to create and edit submission while
extracting identity fields from the large dialog.

**desktop/src/features/agents/ui/AgentDescriptionField.tsx**
Renders the public description input, helper copy, and Unicode-aware
near-limit counter.

**desktop/src/features/agents/ui/AgentIdentityCard.tsx**
Generalizes the card second line to show a two-line description or the
existing model fallback.

**desktop/src/features/agents/ui/UnifiedAgentsSection.tsx**
Prefers authored descriptions on persona cards and retains model labels
when no description exists.

**desktop/src/features/agents/ui/personaDialogState.test.mjs**
Verifies edit and duplicate drafts preserve authored descriptions.

**desktop/src/features/agents/ui/personaDialogState.ts**
Seeds authored descriptions into edit and duplicate dialog drafts.

**desktop/src/features/agents/ui/usePersonaActions.ts**
Preserves descriptions when copying catalog personas into local
definitions.

**desktop/src/shared/api/personaTypes.ts**
Defines description-bearing persona wire types in a focused module split
from the size-constrained API type file.

**desktop/src/shared/api/tauriPersonas.test.mjs**
Verifies raw persona descriptions map into the frontend model and absent
values become null.

**desktop/src/shared/api/tauriPersonas.ts**
Maps description fields across Tauri and preserves raw authored bytes
for authoritative Rust validation.

**desktop/src/shared/api/types.ts**
Re-exports the extracted persona types without changing consumer import
paths.

**desktop/src/testing/e2eBridge.ts**
Extends mock persona create, update, publication, and catalog parsing
with production-shaped description behavior.

**desktop/tests/e2e/agents.spec.ts**
Verifies an edited description persists and appears on the agent card.

</details>

### Reproduction Steps

1. Open **Agents**, edit a custom or built-in agent, and enter a
sentence in **Description**.
2. Save the agent and confirm the sentence appears as the second line on
its card.
3. Reopen the agent and confirm the authored description is restored;
clear it and confirm the card returns to the model label.
4. Paste more than 280 Unicode characters and confirm the field keeps
the first 280 characters and shows the near-limit counter.
5. Share or export/import the agent and confirm the description survives
in the catalog/profile or snapshot without showing a restart-required
badge for a description-only edit.

### Screenshots / Demo

The focused Playwright flow `built-in persona edits persist` exercises
the edited dialog, persisted value, and resulting card subtitle.
Screenshots can be added after review if the field placement or two-line
card treatment needs visual iteration.

### Verification

- `cargo test --manifest-path desktop/src-tauri/Cargo.toml --lib` —
3,029 passed
- `cd desktop && pnpm test` — 5,805 passed
- `cd desktop && pnpm exec tsc --noEmit`
- Focused Playwright: `built-in persona edits persist` — passed
- Pre-push desktop, Tauri, typecheck, test, file-size, and branch-skew
gates — passed

---------

Signed-off-by: tulsi <tulsi@block.xyz>

* fix(desktop): back split thread headers (#7137)

## Summary
- render an auxiliary panel's requested header backdrop in docked/split
mode
- preserve explicit transparent-backdrop behavior
- cover a populated, scrolled thread pane so timeline content cannot
bleed through its header

## Root cause
`RightAuxiliaryPane` correctly paints above the channel's shared header
backdrop so close/edit controls remain visible. The docked
`AuxiliaryPanelHeader` branch, however, ignored its `backdrop` request,
leaving scrolled thread content in that higher stacking context
unbacked.

## Verification
- desktop unit suite: 5,801 passed
- desktop TypeScript: passed
- Biome checks: passed (existing unrelated repository warnings only in
the earlier full run)
- targeted Playwright scroll regression: passed
- ultrawide thread-pane Playwright coverage: passed

Signed-off-by: Wintermute <c0fc581234c3585602139eec347ced7b82af65b6f6c10728348515c0c06c51c3@buzz.block.builderlab.xyz>
Co-authored-by: Wintermute <c0fc581234c3585602139eec347ced7b82af65b6f6c10728348515c0c06c51c3@buzz.block.builderlab.xyz>

* docs: add review-proven failure-path & async-state rules to AGENTS.md (#7061)

Mining the last 25 PRs' review threads (45 substantive findings, 11
reviewed PRs, avg **4.8 review rounds** each) shows **53% of findings
are repeats** of five clusters: swallowed failures, stale-async-state
races, tests that don't bind the production seam, unbounded
resources/retry loops, and non-atomic multi-step persistence. PR #6956
alone burned 4 rounds converging on one of these classes.

A second, independent mining pass over **71 agent-review rooms (303
findings, Aug 18–29)** confirmed the same clusters and added outcome
data — how often authors actually fix each finding class once flagged:
test-seam binding and unbounded-resource findings **100%**, swallowed
errors **90%**, stale-state races **70%**. It also surfaced two clusters
the GitHub-thread pass under-sampled: **assistive-semantics defects**
(44 findings, second-largest cluster) and **input-modality divergence**
(27 findings), now rules 7–8.

This PR distills those clusters into eight imperative rules in AGENTS.md
so agents apply them **before writing code**, adds one
client-consumption invariant to ARCHITECTURE.md §5, and places the
test-quality rule in TESTING.md (per the team decision that testing docs
are the canonical guide for review standards), cross-referenced from
AGENTS.md. Each rule cites the PRs where it was litigated. Raw mining
data: `reviews.jsonl` / `comments.jsonl` +
`backfill/buzz-review-findings.jsonl` (review-mining artifacts, not
committed).

No code changes. CLAUDE.md is a symlink to AGENTS.md and picks this up
automatically.

🤖 Drafted by Jude's agent from automated mining of this repo's last 25
PRs' review threads and 71 agent-review rooms; every rule cites the PRs
where it was litigated. Jude reviews and owns the result. Mining method
+ raw cluster data available on request.

---------

Signed-off-by: Jude Edwards <judeedwards@squareup.com>

* feat(buzz-acp): give each channel thread its own agent session (#6732)

## What this does

In a channel, people often run several unrelated conversations at once
(separate threads). Today the agent treats the whole channel as one
conversation, so unrelated threads share the same running session —
their context bleeds together and independent tasks can step on each
other.

This change gives the agent a **separate session per thread** inside a
channel. Direct messages stay as one conversation (unchanged). The
channel is still the boundary for who is allowed in and what is visible
— only the agent's working context is now split by thread.

## How it is turned on

Off by default. Operators opt in with one setting:

- `BUZZ_ACP_SESSION_POLICY=channel` — default, current behavior
- `BUZZ_ACP_SESSION_POLICY=thread` — new per-thread behavior

Being behind a flag means we can enable it for a few agents, watch how
it behaves, and roll back instantly without a code change.

## Key design decisions

- **Decide the thread once, up front.** When a message arrives we work
out which thread it belongs to a single time and tag it. Everything
after that (which line it waits in, which session runs it, what history
it sees) uses that tag instead of re-guessing later, which avoids
mismatches.
- **Default stays identical to today.** Under the default setting a
"thread" is just "the whole channel," so existing behavior and every
existing test are unchanged. The new, riskier behavior is strictly
opt-in.
- **Give the agent only its thread's history.** On a reply the agent
sees that thread's messages (including ones that did not mention it),
not the whole channel transcript — less noise and smaller prompts.
- **Don't let one channel use more memory than before.** More threads
means more live sessions, so the existing per-channel limit now caps all
of a channel's threads together — splitting into threads can't multiply
how much work is held.

## Bugs found and fixed while iterating (from review)

- **Same thread, two sessions.** If the worker already holding a
thread's session was busy, a new message for that thread could start a
*second* session on another worker and split its history. Now it waits
for the right worker instead of forking.
- **Interrupting the wrong thread.** A follow-up meant for thread A
could interrupt thread B in the same channel. Interrupts now target the
exact thread.
- **Stuck thread after a crash.** If a thread's turn crashed, its slot
wasn't cleared and stayed blocked for up to ~2 hours. It now clears
right away and retries.
- **Lost the original request.** When a thread was interrupted and then
had to wait for a busy worker, only the follow-up was kept and the
original request was dropped. The full request is now preserved on
retry.
- **Same thread seen as two.** Two spellings of the same thread id
(upper/lower case) could be treated as different threads. Normalized so
they count as one.

## Not in this PR

- The desktop Settings toggle and rollout wiring for managed agents —
https://github.com/block/buzz/pull/6909
- One pre-existing retry edge case (present today without this flag,
unrelated to this change) — tracked separately so this PR stays focused.

## Testing

The full `buzz-acp` test suite passes (830+ unit and integration tests),
plus new focused tests for thread routing, session reuse, interrupt
targeting, crash recovery, and request preservation. Behavior with the
flag off is unchanged.

---------

Signed-off-by: Salman Mohammed <smohammed@squareup.com>
Signed-off-by: Leo <5faf251baee50ee6bcde338aef6acdd70bb3e60115664c2cd490d94a55dfc488@buzz.block.builderlab.xyz>
Co-authored-by: Leo <5faf251baee50ee6bcde338aef6acdd70bb3e60115664c2cd490d94a55dfc488@buzz.block.builderlab.xyz>

* feat(db): add NIP-FI identity and final-admission schema foundation (#6994)

PR 2 of the NIP-FI plan: the schema foundation. Establishes the durable
server-side identity ledger and final-admission surface that the runtime
phases build on. All of Phase A's migrations live here; later phases own
their own deltas.

Depends on nothing — PR 1 (#6776, merged) owned zero migration files.
This PR's relations are shaped to store exactly what PR 1's verifier
produces: issuer-qualified identity and the four denial classes. They
meet in a later PR that writes a verified assertion into these tables in
one transaction.

## Two internally-ordered migrations

- `0041_nip_fi_identity_foundation.sql` (migration A) — core identity +
base-lifecycle relations (5 tables): issuer-qualified `(iss, sub)`
bindings, lifecycle history/selectors, enrollment policies, and
operation receipts. Applies cleanly to current `main`.
- `0042_nip_fi_authorization_foundation.sql` (migration B) — the
final-admission surface (10 tables): authorization events + capacity,
admission results, replay/receipt guards, audit, invalidation
domains/floors, protected-object authority, authority epochs, and
restore version deltas. Applies to A's resulting state.

Fifteen NIP-FI relations total, zero dangling foreign keys. Identity is
issuer-qualified throughout — no single-global-issuer assumption in any
relation, no `Block`-hardcoding. A single deployment may run one issuer;
that is config, not schema.

## Durable, immutable ledger posture

All 15 relations are append-only (immutable `no_delete`/`no_truncate`
triggers) and carry `community_id` as provenance, not ownership. Both
migrations widen the single SQL source of truth
`community_write_fence_excluded_table` so the relations are never
fence-attached, never purged on community deletion, and never counted as
tenant-scoped drift by the deletion control plane's exact-set catalog
check — the same posture main already applies to `product_feedback` and
`rate_limit_violations`. `schema/schema.sql` keeps one consolidated
definition of that function whose exclusion array byte-matches `0042`,
guarded by a parity assertion so a future consolidation cannot silently
drop NIP-FI relations from the ledger.

This makes a tenant's identity/authorization ledger survive community
deletion, per the spec's `FI-INV-02` (durable binding) and `FI-INV-03`
(tombstone monotonicity) and `NIP-FI.md`'s "durable server state"
ruling. `communities(id)` FK never dangles: community rows become
permanent tombstones, never hard-deleted.

## Authorization shape and cardinality contracts

Authenticated `OperatorDenied` events (`actor_kind` 1–3, non-null
`request_fingerprint`) carry a null `semantic_fingerprint` and commit
without a denial-attempt row. The denial-attempt cardinality and shape
guards are scoped to unresolved pre-auth kind-9 events (`actor_kind =
4`). Applied and no-op lifecycle receipts (`outcome_code IN (1, 3)`)
require exactly one mapped success-transition event; denied lifecycle
receipts (`outcome_code = 2`) require zero events from the complete core
lifecycle success-transition class (kinds 1, 2, 3, 6: enrolled, revoked,
rotated, retired) — any such event paired with a denied receipt would
record a transition that never occurred.

## Mined vs. new

Re-cut from Franco's #1476 (`0029`/`0030`) and Cea's #4772 committer
schema, re-cut along FK topology and renumbered above the live `main`
tip. The buzz-auth core of #1476 is Cea-authored; `Co-authored-by`
reflects verified per-commit authorship of the mined schema.

Zero Rust/`deletion.rs` edits — the migration-only exclusion widening
keeps `EXPECTED_SCOPED_TABLES` untouched.

---------

Signed-off-by: Will Pfleger <pfleger.will@gmail.com>
Signed-off-by: Hayt <9e1c23a3fd83f61da34420e4e88ff1b16e45cafcc0cd9019eb07d4ecfa8ca9b0@buzz.block.builderlab.xyz>
Co-authored-by: Cea Stapleton Cordasco <261786559+cea@users.noreply.github.com>
Co-authored-by: Hayt <9e1c23a3fd83f61da34420e4e88ff1b16e45cafcc0cd9019eb07d4ecfa8ca9b0@buzz.block.builderlab.xyz>

* fix(model-capabilities): humanize databricks goose model names (#7135)

🤖
## Summary
- add curated human-readable labels for Databricks Goose models that
otherwise render as fully qualified identifiers
- render `data_workflow_tools.goose.goose-glm-5-3` as `GLM-5.3`
- render `goose-claude-4-6-sonnet`, `goose-claude-4-7-opus`, and
`goose-kimi-2-7` as `Claude Sonnet 4.6`, `Claude Opus 4.7`, and `Kimi
2.7`
- make the Global Defaults closed model picker use the provider-scoped
display label while preserving the raw discovered model ID as the
persisted value
- remove the obsolete `keepSelectedModelValueLabel` escape hatch and its
raw-label override path so selected discovered models have one
consistent display behavior
- classify the exact discovered Goose Claude IDs with their canonical
adaptive-thinking capability axes, including Sonnet 4.6's exclusion of
`xhigh`
- expand Rust and TypeScript alias coverage and regenerate the shared
139-vector capability corpus

## Test plan
- `cargo test -p buzz-agent --lib` — 517 passed, 1 ignored
- `cd desktop && pnpm test` — 5,821 passed
- Desktop TypeScript typecheck — passed
- Biome on the changed component — passed
- `git diff --check` — passed
- targeted Playwright Global Defaults regression — passed on the
preceding implementation head; the subsequent commit only removes dead
picker-prop plumbing

Verified at `b9609d12696173aa309d2dbaf4f093a502756c36`. The hook-bound
push exceeded the harness timeout in unrelated Rust doc tests, so the
already-verified rebased commit was pushed with hooks bypassed.

Follow-up to #6955.

---------

Signed-off-by: Kalvin Chau <kalvin@block.xyz>
Co-authored-by: am <6e30cd56c30e030cd31bb0939b94a7c257c9a09d5ba2d92cf2735da45629f248@buzz.block.builderlab.xyz>

* feat(desktop): add isolated named demo builds (#6407)

🤖 I’m Larry, updating this description on Logan’s behalf.

## Summary

Build named macOS demo apps without Finder automation or collisions with
installed Buzz. `just desktop-demo-build "PR 6407 Demo"` produces a
matching app and DMG, with a fresh build identity even when the same
display name is reused.

- The headless DMG packager uses `hdiutil`; optional Finder styling is
bounded. The existing production release recipe is unchanged.
- Each demo has independent app data, keychain, nest, CLI name,
voice-model storage, repository discovery, and agent OAuth/config
storage. Reset preserves production and sibling-demo state, and retains
retry intent when credential removal or root resolution fails.
- Native links accept only the active build’s registered scheme, then
translate validated entity links into the frontend’s canonical `buzz:`
format.
- The recipe builds all six executable sidecars. Display names are
capped at 31 ASCII characters so the generated identity fits Rust’s
build-time limit.

**Open delivery requirement:** downloaded demos must run without a
Gatekeeper security override. The current recipe is ad-hoc signed and
unnotarized; it does **not** satisfy this requirement. Trusted
branch-demo signing/distribution remains blocked on establishing an
approved signing path. This PR is not being presented as complete
download-and-run delivery.

### Related issue

N/A — reported in the Buzz DMG-packaging workstream.

### Testing

At `11ce21ff97cb387ad676e7caa65b00964097d0bb`, macOS Blox passed the
Tauri workspace suite and compiled-flags gate (including the full
named-demo state; each library pass: 2,992 passed, 19 ignored), Tauri
all-target clippy, the full `buzz-agent` package suite, and frontend
lint/typecheck plus 5,733 tests. Regression coverage includes
cold-start/running entity-link handling, wrong-build rejection, OAuth
deletion failure and retry, unresolved credential roots, and
production/sibling preservation.

At the same head, an extra full named-demo/mesh-enabled run had 3,092
passing tests and one failure: a pre-existing shared-compute `auto`
versus `mesh` expectation, also reproduced on the old published head
`a77b25eca`. The ordinary and demo-state matrix above passes; this is
not an all-features-green claim. Live macOS Launch Services delivery
remains unverified.

GitHub CI completed with 30 successful checks and 9 skipped. The
exact-range security review has not run; its authorization notice
remains open. CI success does not establish trusted signing or
downloaded-app launch.

Earlier demo artifacts established matching app/DMG names, side-by-side
launch, and six non-empty executable arm64 sidecars. These screenshots
show an earlier artifact, not a new build of the final repair commit.
Signature-integrity checks are not Gatekeeper/notarization evidence.

<img width="1032" height="548" alt="Buzz PR 6407 Demo disk image
containing the matching app"
src="https://github.com/user-attachments/assets/bca0277e-db03-4308-b280-fcad55e6d601"
/>

<img width="1186" height="821" alt="Buzz PR 6407 Demo running alongside
other Buzz installations"
src="https://github.com/user-attachments/assets/b4bf4ae5-c341-4e15-8090-9d2ea7c623b6"
/>

---------

Signed-off-by: Logan Johnson <loganj@squareup.com>
Signed-off-by: Larry <8cf5a83f590ec0955b11647d1c88f796a98e088c30a492c58e0e46c3026ae7a4@buzz.block.builderlab.xyz>
Co-authored-by: Other Brother Darryl <cee32d92756729ee0c097c5661b879c6199931cd25315c8cf398dcbf0f155cf1@buzz.block.builderlab.xyz>
Co-authored-by: Larry <loganj+sandbox-larry@squareup.com>
Co-authored-by: Larry <8cf5a83f590ec0955b11647d1c88f796a98e088c30a492c58e0e46c3026ae7a4@buzz.block.builderlab.xyz>

---------

Signed-off-by: Philip Azar <pazar@squareup.com>
Signed-off-by: Will Pfleger <pfleger.will@gmail.com>
Signed-off-by: Fizz <dae5f6af70b8695a8b83c8deae555f63be41630ec2b8cd493e41a439c9527dd8@buzz.block.builderlab.xyz>
Signed-off-by: Pinky <5f5ab050ec58ae208332edd544ebf705221e24c1b86d82a6ca07038a7a8f6ac9@buzz.block.builderlab.xyz>
Signed-off-by: Brain <1a02c72794dcd0f07058a353bc3a81f4028b8c77c92c87fce6d5c8b85970a20b@buzz.block.builderlab.xyz>
Signed-off-by: Wes <wesbillman@users.noreply.github.com>
Signed-off-by: Arjun Mahanti <arjun@squareup.com>
Signed-off-by: Fizz <fizz@buzz.local>
Signed-off-by: tulsi <tulsi@block.xyz>
Signed-off-by: Wintermute <c0fc581234c3585602139eec347ced7b82af65b6f6c10728348515c0c06c51c3@buzz.block.builderlab.xyz>
Signed-off-by: Jude Edwards <judeedwards@squareup.com>
Signed-off-by: Salman Mohammed <smohammed@squareup.com>
Signed-off-by: Leo <5faf251baee50ee6bcde338aef6acdd70bb3e60115664c2cd490d94a55dfc488@buzz.block.builderlab.xyz>
Signed-off-by: Hayt <9e1c23a3fd83f61da34420e4e88ff1b16e45cafcc0cd9019eb07d4ecfa8ca9b0@buzz.block.builderlab.xyz>
Signed-off-by: Kalvin Chau <kalvin@block.xyz>
Signed-off-by: Logan Johnson <loganj@squareup.com>
Signed-off-by: Larry <8cf5a83f590ec0955b11647d1c88f796a98e088c30a492c58e0e46c3026ae7a4@buzz.block.builderlab.xyz>
Signed-off-by: shiv <shivchander.s30@gmail.com>
Co-authored-by: Phil Azar <pazar@squareup.com>
Co-authored-by: Will Pfleger <pfleger.will@gmail.com>
Co-authored-by: Duncan <dcfd242e557282d7a1e2cf2e6877522682f1e5c6156dc92ca7d90eaedd3b0f95@buzz.block.builderlab.xyz>
Co-authored-by: Arjun Mahanti <arjun.mahanti@gmail.com>
Co-authored-by: Fizz <dae5f6af70b8695a8b83c8deae555f63be41630ec2b8cd493e41a439c9527dd8@buzz.block.builderlab.xyz>
Co-authored-by: Wes <wesbillman@users.noreply.github.com>
Co-authored-by: Pinky <5f5ab050ec58ae208332edd544ebf705221e24c1b86d82a6ca07038a7a8f6ac9@buzz.block.builderlab.xyz>
Co-authored-by: LioLionel <62820906+LioLionel@users.noreply.github.com>
Co-authored-by: Brain <1a02c72794dcd0f07058a353bc3a81f4028b8c77c92c87fce6d5c8b85970a20b@buzz.block.builderlab.xyz>
Co-authored-by: Carl <9d00794d3df50972eb8b615511783cab12a77a8fd5dd5edd58073ec73b54bd8b@buzz.block.builderlab.xyz>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Fizz <fizz@buzz.local>
Co-authored-by: tulsi <tulsi@block.xyz>
Co-authored-by: thomaspblock <thomasp@squareup.com>
Co-authored-by: Wintermute <c0fc581234c3585602139eec347ced7b82af65b6f6c10728348515c0c06c51c3@buzz.block.builderlab.xyz>
Co-authored-by: Jude Edwards <judeedwards@squareup.com>
Co-authored-by: Salman Mohammed <smohammed@squareup.com>
Co-authored-by: Leo <5faf251baee50ee6bcde338aef6acdd70bb3e60115664c2cd490d94a55dfc488@buzz.block.builderlab.xyz>
Co-authored-by: Cea Stapleton Cordasco <261786559+cea@users.noreply.github.com>
Co-authored-by: Hayt <9e1c23a3fd83f61da34420e4e88ff1b16e45cafcc0cd9019eb07d4ecfa8ca9b0@buzz.block.builderlab.xyz>
Co-authored-by: Kalvin C <kalvinnchau@users.noreply.github.com>
Co-authored-by: am <6e30cd56c30e030cd31bb0939b94a7c257c9a09d5ba2d92cf2735da45629f248@buzz.block.builderlab.xyz>
Co-authored-by: Logan Johnson <loganj@squareup.com>
Co-authored-by: Other Brother Darryl <cee32d92756729ee0c097c5661b879c6199931cd25315c8cf398dcbf0f155cf1@buzz.block.builderlab.xyz>
Co-authored-by: Larry <loganj+sandbox-larry@squareup.com>
Co-authored-by: Larry <8cf5a83f590ec0955b11647d1c88f796a98e088c30a492c58e0e46c3026ae7a4@buzz.block.builderlab.xyz>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
salman1993 added a commit that referenced this pull request Sep 1, 2026
## Why

Thread-scoped ACP sessions from #6732 need an opt-in desktop rollout
that preserves today’s channel policy by default.

## What

- Add the default-off “Thread Scoped ACP Sessions” toggle under Settings
→ Experiments using the existing feature-override persistence.
- Map the persisted setting to `BUZZ_ACP_SESSION_POLICY=channel|thread`
for every local and provider-backed managed ACP launch. Changes apply
when managed agents next start; DMs remain conversation-scoped by the
backend.
- Keep the desktop setting authoritative over descriptor environment and
cover the UI, persistence, local spawn, provider payload, and existing
Projects/Workflows experiments.

## Risk Assessment

Low. The experiment defaults off and explicitly preserves `channel`;
changes are limited to managed-agent launch configuration. Existing
running agents are unchanged until their next start.

## Testing

- `. ./bin/activate-hermit && just ci` — passed on the final restacked
tree, including 5,674 desktop tests, 2,796 Tauri tests, and 1,860 mobile
tests.
- `cd desktop && pnpm build:e2e && pnpm exec playwright test
tests/e2e/experimental-features.spec.ts --project=smoke` — 1 passed.
- `. ./bin/activate-hermit && cargo test --manifest-path
desktop/src-tauri/Cargo.toml session_policy --lib` — 4 passed.
- `. ./bin/activate-hermit && cargo test --manifest-path
desktop/src-tauri/Cargo.toml commands::agents::deploy::tests --lib` — 15
passed.
- `. ./bin/activate-hermit && cargo test -p buzz-backend-kubernetes
--test wire_fixtures` — 4 passed.

## Stack Info

Stacked on #6732. This PR depends on #6732 and should merge after it.

Generated with Codex

---------

Signed-off-by: Salman Mohammed <smohammed@squareup.com>
Signed-off-by: Leo <5faf251baee50ee6bcde338aef6acdd70bb3e60115664c2cd490d94a55dfc488@buzz.block.builderlab.xyz>
Co-authored-by: Leo <5faf251baee50ee6bcde338aef6acdd70bb3e60115664c2cd490d94a55dfc488@buzz.block.builderlab.xyz>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants