feat(buzz-acp): give each channel thread its own agent session - #6732
Conversation
|
🤖 Request changes at exact head Major — equivalent NIP-10 root spellings still split one relay thread into separate ACP sessions
Concrete scenario: a valid reply tags root Normalize validated marker IDs before constructing the key (or key on Verification
|
…re one scope The shared NIP-10 parser accepts and preserves uppercase ASCII hex in `e`-tag marker ids (`is_ascii_hexdigit`), but the relay decodes event ids to bytes on ingest — so a reply tagging its root as `AB…` and one tagging `ab…` are the SAME accepted relay thread. `SessionScope::Thread` keyed on the raw string, so under thread policy those equivalent spellings hashed to different keys and split one thread across two ACP sessions (queue partitions, provider sessions, worker affinity, and delivery ledgers), violating same-root reuse. Normalize the resolved root id to lowercase in `SessionScope::derive` before it becomes the scope key. `nostr::EventId::to_hex()` is already lowercase, so the top-level-mention path is unaffected. Adds a mixed-case regression test. Reported in review of PR #6732. Signed-off-by: Salman Mohammed <smohammed@squareup.com>
|
🤖 Re-review: ready at exact head The prior mixed-case NIP-10 blocker is resolved. Fresh verification at this SHA:
No actionable findings remain. DMs were not re-exercised live in this pass. |
…ting Addresses three correctness gaps found in review of PR #6732. 1. Duplicate provider session for a busy thread (pool.rs, lib.rs). Worker affinity only scanned idle slots, so while the worker that owns a thread's session was checked out, a new message for that thread could be handed to another idle worker, forking a second session and splitting the thread's history/tool context. Added an authoritative `SessionScope -> worker` directory (`session_owners`) that survives while a worker is checked out. `dispatch_pending` now holds a batch (leaves it queued) when its session owner is busy, instead of forking a duplicate; the held batch dispatches to that exact worker when it returns. The directory is pruned on channel-wide session invalidation, and stale entries (rotation / crash) self-heal on the next dispatch. 2. Mid-turn steer/interrupt could target the wrong thread (lib.rs). The native-steer fallback and the steer-ack fallback routed by channel via `signal_in_flight_task`, which picks the first task for the channel — so a message in thread A could interrupt thread B in the same channel. Added `signal_in_flight_task_for_scope` (exact `SessionScope` match) and used it for both mid-turn fallbacks. The deferred channel-level control paths (`!cancel`, `!rotate`, observer `cancel_turn` / `switch_model`) keep channel-targeting intentionally. 3. Panicked thread stayed "in flight" (lib.rs). `recover_panicked_agent` called `mark_complete(channel_id)`, which resolves to `Conversation(channel_id)` via IntoScope and, under thread policy, left the real `Thread(...)` entry wedged in-flight until the ~2h backstop — blocking the batch it had just requeued. It now uses `meta.scope`. Tests: scope-exact signalling targets only the matching thread; a busy session owner holds the batch instead of forking a session (and the directory prunes on channel invalidation); panic recovery frees the exact Thread scope and requeues its batch. Full buzz-acp suite (831 lib + integration) green; clippy and fmt clean. Signed-off-by: Salman Mohammed <smohammed@squareup.com>
…hausted batch `requeue_preserve_timestamps` restored only `batch.events`, silently dropping `batch.cancelled_events` and `cancel_reason`. Both callers pass a full `FlushBatch` that can carry an interrupted turn's original request: - the new busy-owner affinity hold (thread scoping), and - the pre-existing "no agent available" pool-exhausted path. Reachable loss: thread A is interrupted (Buzz retains A's original request to re-prompt "original + follow-up"); older work in thread B grabs A's session-owning worker; A's merged batch is held; only the follow-up was put back and the original request vanished. The helper now restores the entire batch — cancelled carryover is returned to the pending cancelled-batches (ahead of any concurrently staged carryover) with its reason, so the next flush reconstructs the same merged prompt. Adds a queue-level round-trip regression asserting events + cancelled_events + cancel_reason all survive requeue -> mark_complete -> flush. Reported in review of PR #6732. Full buzz-acp suite (832 lib + integration) green; clippy and fmt clean. Signed-off-by: Salman Mohammed <smohammed@squareup.com>
|
Good catch on the |
|
Tested PR head
This validates the harness/session routing. I did not test the full Desktop UI path. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: f8f8808890
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| scope.telemetry_label() | ||
| ); | ||
| agent.state.sessions.insert(*cid, sid.clone()); | ||
| agent.state.sessions.insert(scope.clone(), sid.clone()); |
There was a problem hiding this comment.
Bound retained sessions under thread policy
When BUZZ_ACP_SESSION_POLICY=thread, every sequential top-level mention gets a unique scope and this insertion retains the resulting provider session indefinitely; no queue cap evicts entries from SessionState.sessions. A busy channel can therefore accumulate an unbounded number of ACP sessions even with no queued work, and providers such as buzz-agent spawn a separate MCP registry for every session/new, eventually exhausting processes and memory. Add a bounded per-channel/session eviction policy that also closes or invalidates the provider session.
Useful? React with 👍 / 👎.
🔐 Codex Security Review
|
|
Thanks for the review. Addressed in P2 — P2 — channel-keyed typing state (fixed). Both are only reachable under P1 — unbounded session / provider retention (split out). Agreed this isn't a clean PR-local fix — partial eviction of Buzz-side IDs would leak provider resources. It's pre-existing (not thread-scoping-specific) and needs a designed lifecycle: idle-TTL / LRU on the scope maps, a real per-session close contract that tears down |
c409bc9 to
1e2a507
Compare
Lay the foundation for thread-scoped ACP sessions (rollout step 1: land behind the operator policy with channel scope as the fallback). - Add `scope` module with a hashable `SessionScope` (Conversation/Thread) and `SessionPolicy` (channel/thread). Scope is derived once at admission from policy + DM status + NIP-10 thread tags using the shared `buzz_core::nip10` canonical-root rules. DMs are always conversation-scoped; under the default `channel` policy every channel event collapses to a conversation scope, preserving today's behavior. - Add `--session-policy` / `BUZZ_ACP_SESSION_POLICY` (default `channel`) wired through `CliArgs` -> `Config` and the config summary; document it in `.env.example`. - Derive and log the resolved scope at event admission (telemetry only for now). Unit tests cover scope derivation for top-level mentions, direct/nested replies, repeated mentions, DMs, malformed thread tags, hashing/map-key use, and policy parsing/defaults. Follow-up (same ticket, later commits): partition queue/in-flight state by scope (step 2), partition provider-session state by scope (step 3), scope-correct context gathering (step 4), and the Settings->Experiments toggle + managed-agent deployment wiring (step 5). Signed-off-by: Salman Mohammed <smohammed@squareup.com>
…step 2)
Carry the admission-time SessionScope through the queue instead of keying
everything on channel_id. All EventQueue partitions — pending queues,
in-flight tracking, deadlines, retry counters/backoff, cancelled-batch
carryover, and the goose-native steer side table — are now keyed by
SessionScope. Under the default `channel` policy every scope is
Conversation{channel_id}, so behavior is byte-identical; under `thread`
policy, distinct canonical roots in one channel become independent
partitions.
- QueuedEvent and FlushBatch carry the resolved SessionScope; dispatch,
completion, and requeue route by it. PromptSource::Channel now carries
the scope (with a channel_id() accessor) so a completed turn marks the
exact scope complete. No code rederives scope from the last event in a
batch.
- The mid-turn steer gate is now scope-level (is_scope_in_flight), and the
native-steer withhold/release/dedup + deadline extension target the
scope, so an unrelated thread's in-flight turn never steers this one.
- drain_channel performs channel-wide cleanup across every child thread
scope. Backlog protection is preserved with a per-scope cap plus an
aggregate per-channel cap, so per-thread partitioning cannot multiply the
admitted queue size.
- An IntoScope helper lets the queue API accept a bare channel Uuid
(conversation scope) or an explicit SessionScope, keeping the existing
channel-keyed unit tests intact.
Pool provider-session STATE is still channel-keyed in this commit (indexed
via scope.channel_id()); step 3 rekeys it by scope so repeated activity in
a thread reuses exactly that thread's provider session.
New queue tests: two threads in one channel are independent partitions,
events from different roots never share a batch, an in-flight scope blocks
only that scope, channel drain clears every child thread scope, and the
aggregate channel cap is not multiplied by threads. Full buzz-acp suite
(819 lib + integration) green; clippy and fmt clean.
Signed-off-by: Salman Mohammed <smohammed@squareup.com>
…p 3) Key the pool's provider-session state by SessionScope instead of only channel_id, so repeated activity in a thread reuses exactly that thread's provider session and unrelated threads in one channel never share session state, turn counters, context-delivery markers, or delivery-dedup state. - SessionState maps (sessions, turn_counts, core_sections, canvas_sections, deliveries) are now keyed by SessionScope. run_prompt_task resolves the session, core/canvas sections, standing-context-sent marker, turn count, and delivery ledger by the batch's scope; channel-level fetches (canvas, huddle, title, resolve) still use scope.channel_id(). - Scope-to-worker affinity: has_session_for/try_claim match the exact scope, so a temporarily busy worker cannot cause another worker to open a duplicate session for the same thread. TaskMeta carries the scope; send_steer and record_successful_steer route by it. - Channel-wide cleanup preserved: invalidate_channel clears every child thread scope for a channel (returns the count); invalidate_channel_sessions and the removed-channel path use it. Added invalidate_scope for single-session invalidation and mark_scope_delivery_success. - Model-switch targeting stays channel-level with an explicit TODO for channel-vs-thread control targeting (same open question as top-level !cancel / !rotate, deferred per the ticket). New pool tests: two threads in one channel get distinct sessions and reuse per root, invalidate_scope leaves a sibling thread untouched, and invalidate_channel clears every thread scope while sparing other channels. Full buzz-acp suite (822 lib + integration) green; clippy and fmt clean. Signed-off-by: Salman Mohammed <smohammed@squareup.com>
Drive conversation-context gathering from the batch's resolved SessionScope instead of re-inferring it from whichever event is last in the batch. - A Thread scope fetches only that canonical thread's history (all messages under the root, including intervening non-mention human messages), so no unrelated channel transcript is injected. A brand-new thread's first turn has no prior history — the trigger is delivered as the [Event] block. - Conversation scope preserves current behavior: DMs (and legacy channel-policy channels) fetch the reply chain for a threaded reply or recent DM history for a non-reply; a plain top-level channel message gets no supplementary context. - The existing scope-keyed delivery-delta filter (step 3) then strips events this session already received, so subsequent turns deliver only the intervening same-thread messages plus the trigger, without duplication. The routing decision is extracted into a pure `resolve_context_target` so it is unit-tested directly: thread scope wins over a divergent last-event tag, a new top-level thread resolves to its own root, a plain conversation-scope channel message gets no context, DM non-reply fetches DM history, and a conversation-scope reply uses its reply chain. Together with steps 1–3 this makes BUZZ_ACP_SESSION_POLICY=thread deliver real end-to-end isolation: distinct roots in one channel get distinct queue partitions (step 2), distinct provider sessions with scope-keyed worker affinity (step 3), and distinct canonical-thread context (this step); while the default `channel` policy is byte-identical to prior behavior. Full buzz-acp suite (827 lib + integration) green; clippy and fmt clean. Signed-off-by: Salman Mohammed <smohammed@squareup.com>
…re one scope The shared NIP-10 parser accepts and preserves uppercase ASCII hex in `e`-tag marker ids (`is_ascii_hexdigit`), but the relay decodes event ids to bytes on ingest — so a reply tagging its root as `AB…` and one tagging `ab…` are the SAME accepted relay thread. `SessionScope::Thread` keyed on the raw string, so under thread policy those equivalent spellings hashed to different keys and split one thread across two ACP sessions (queue partitions, provider sessions, worker affinity, and delivery ledgers), violating same-root reuse. Normalize the resolved root id to lowercase in `SessionScope::derive` before it becomes the scope key. `nostr::EventId::to_hex()` is already lowercase, so the top-level-mention path is unaffected. Adds a mixed-case regression test. Reported in review of PR #6732. Signed-off-by: Salman Mohammed <smohammed@squareup.com>
…ting Addresses three correctness gaps found in review of PR #6732. 1. Duplicate provider session for a busy thread (pool.rs, lib.rs). Worker affinity only scanned idle slots, so while the worker that owns a thread's session was checked out, a new message for that thread could be handed to another idle worker, forking a second session and splitting the thread's history/tool context. Added an authoritative `SessionScope -> worker` directory (`session_owners`) that survives while a worker is checked out. `dispatch_pending` now holds a batch (leaves it queued) when its session owner is busy, instead of forking a duplicate; the held batch dispatches to that exact worker when it returns. The directory is pruned on channel-wide session invalidation, and stale entries (rotation / crash) self-heal on the next dispatch. 2. Mid-turn steer/interrupt could target the wrong thread (lib.rs). The native-steer fallback and the steer-ack fallback routed by channel via `signal_in_flight_task`, which picks the first task for the channel — so a message in thread A could interrupt thread B in the same channel. Added `signal_in_flight_task_for_scope` (exact `SessionScope` match) and used it for both mid-turn fallbacks. The deferred channel-level control paths (`!cancel`, `!rotate`, observer `cancel_turn` / `switch_model`) keep channel-targeting intentionally. 3. Panicked thread stayed "in flight" (lib.rs). `recover_panicked_agent` called `mark_complete(channel_id)`, which resolves to `Conversation(channel_id)` via IntoScope and, under thread policy, left the real `Thread(...)` entry wedged in-flight until the ~2h backstop — blocking the batch it had just requeued. It now uses `meta.scope`. Tests: scope-exact signalling targets only the matching thread; a busy session owner holds the batch instead of forking a session (and the directory prunes on channel invalidation); panic recovery frees the exact Thread scope and requeues its batch. Full buzz-acp suite (831 lib + integration) green; clippy and fmt clean. Signed-off-by: Salman Mohammed <smohammed@squareup.com>
…hausted batch `requeue_preserve_timestamps` restored only `batch.events`, silently dropping `batch.cancelled_events` and `cancel_reason`. Both callers pass a full `FlushBatch` that can carry an interrupted turn's original request: - the new busy-owner affinity hold (thread scoping), and - the pre-existing "no agent available" pool-exhausted path. Reachable loss: thread A is interrupted (Buzz retains A's original request to re-prompt "original + follow-up"); older work in thread B grabs A's session-owning worker; A's merged batch is held; only the follow-up was put back and the original request vanished. The helper now restores the entire batch — cancelled carryover is returned to the pending cancelled-batches (ahead of any concurrently staged carryover) with its reason, so the next flush reconstructs the same merged prompt. Adds a queue-level round-trip regression asserting events + cancelled_events + cancel_reason all survive requeue -> mark_complete -> flush. Reported in review of PR #6732. Full buzz-acp suite (832 lib + integration) green; clippy and fmt clean. Signed-off-by: Salman Mohammed <smohammed@squareup.com>
Signed-off-by: Salman Mohammed <smohammed@squareup.com>
Two thread-isolation correctness bugs surfaced in review, both only reachable under BUZZ_ACP_SESSION_POLICY=thread; behavior under the default channel policy is unchanged (the scope is the channel's sole conversation). !cancel / !rotate could hit the wrong thread. Both routed through signal_in_flight_task, which matches on channel_id and takes the first HashMap entry — so with two threads running concurrently in one channel an owner's !cancel/!rotate could tear down an arbitrary thread's turn. They now derive the SessionScope from the command event's NIP-10 tags (the same resolver admission uses) and target it via signal_in_flight_task_for_scope, the scope-exact primitive steering already uses. The idle !rotate path now invalidates only that scope via the new AgentPool::invalidate_scope_session instead of the whole channel. signal_in_flight_task is now used only by the desktop observer control frames (cancel_turn / switch_model), which carry a bare channelId and no thread context. Typing state was channel-keyed. typing_channels keyed by channel_id, so two concurrent thread turns overwrote one entry and either turn's completion removed it — the indicator could reflect the wrong thread or stop while a sibling turn continued. It is now keyed by SessionScope: dispatch_pending returns the scope, completion/panic/ownership-removal clear the exact scope (via the new PromptSource::scope accessor), and the refresh loop publishes one indicator per active thread carrying that thread's NIP-10 tags. Tests: invalidate_scope_session targets one thread and drops its owner; PromptSource::scope exposes the thread scope and None for heartbeats. The reviewer's P1 (unbounded session/provider-resource retention) is a pre-existing lifecycle-hardening concern, not thread-scoping-specific, and is tracked separately rather than adding partial eviction here. Signed-off-by: Salman Mohammed <smohammed@squareup.com>
Match session framing and titles to the configured scope. Reject ambiguous channel controls and correlate cancel acknowledgements with their requests. Co-authored-by: Salman Mohammed <smohammed@squareup.com> Signed-off-by: Leo <5faf251baee50ee6bcde338aef6acdd70bb3e60115664c2cd490d94a55dfc488@buzz.block.builderlab.xyz>
Signed-off-by: Salman Mohammed <smohammed@squareup.com>
1e2a507 to
a089e89
Compare
wpfleger96
left a comment
There was a problem hiding this comment.
🤖 I think the thread-scoping behavior and rollout gate are working as intended, but there is one lifecycle issue to address before enabling this.
BUZZ_ACP_SESSION_POLICY defaults to channel, so the existing one-session-per-channel behavior remains the default; thread is opt-in and DMs remain conversation-scoped. Runtime verification against a fresh relay confirmed both sides of that boundary: the default reused one provider session across two roots, while thread mode created distinct sessions, reused the correct session for a follow-up, and kept sibling-thread prompts isolated.
The blocker is retained-session cardinality. Thread mode changes the long-lived state from roughly one provider session per channel to one per encountered thread. SessionState retains sessions, turn_counts, context sections, and delivery state by SessionScope, and cleanup currently depends on explicit rotation/invalidation, channel removal, or process exit. There is no LRU, TTL, or maximum retained-scope policy. Because buzz-agent also permits unlimited sessions by default and ACP sessions own provider/MCP resources, a channel with many one-off roots can grow memory and child-process usage for the lifetime of the agent.
Please add bounded idle-session eviction with provider-side cleanup, or establish and enforce a safe end-to-end session cap before rollout. I’d also fix the small coverage gap while touching this: test_session_policy_env_var_parses currently passes --session-policy=thread rather than exercising BUZZ_ACP_SESSION_POLICY.
The remaining source paths and current CI checks look good at a089e89c4ca340a07e795f879560477e4d5d5096.
wpfleger96
left a comment
There was a problem hiding this comment.
🤖 Approved. The missing LRU/TTL is pre-existing; this PR increases the reachable session cardinality under the opt-in thread policy, but the feature is default-off and its intended behavior passed source review, CI, and live relay/ACP verification. I’m treating lifecycle bounding as separate hardening rather than a merge blocker.
…enericize * origin/main: feat(desktop): add isolated named demo builds (#6407) fix(model-capabilities): humanize databricks goose model names (#7135) feat(db): add NIP-FI identity and final-admission schema foundation (#6994) feat(buzz-acp): give each channel thread its own agent session (#6732) docs: add review-proven failure-path & async-state rules to AGENTS.md (#7061) Signed-off-by: Duncan <dcfd242e557282d7a1e2cf2e6877522682f1e5c6156dc92ca7d90eaedd3b0f95@buzz.block.builderlab.xyz>
…h-coordinator * origin/main: feat(desktop): add isolated named demo builds (#6407) fix(model-capabilities): humanize databricks goose model names (#7135) feat(db): add NIP-FI identity and final-admission schema foundation (#6994) feat(buzz-acp): give each channel thread its own agent session (#6732) docs: add review-proven failure-path & async-state rules to AGENTS.md (#7061) fix(desktop): back split thread headers (#7137) Signed-off-by: Duncan <dcfd242e557282d7a1e2cf2e6877522682f1e5c6156dc92ca7d90eaedd3b0f95@buzz.block.builderlab.xyz> Co-authored-by: Will Pfleger <pfleger.will@gmail.com> Signed-off-by: Will Pfleger <pfleger.will@gmail.com>
…-channel-permissions * origin/main: fix(model-capabilities): humanize databricks goose model names (#7135) feat(db): add NIP-FI identity and final-admission schema foundation (#6994) feat(buzz-acp): give each channel thread its own agent session (#6732) docs: add review-proven failure-path & async-state rules to AGENTS.md (#7061) fix(desktop): back split thread headers (#7137) add public descriptions to agent personas (#7126) feat(desktop): add protected-build Bestie experiment (#6902) fix(relay): reject a frame on its own acknowledgement channel (#6961) fix(acp): wake agents from workflow messages (#6953) Signed-off-by: John Tennant <jtennant@squareup.com> # Conflicts: # crates/buzz-db/src/runtime/migration.rs # schema/schema.sql
* fix: retrieving cold memories; add regression task (#6950)
## Why
Evaluating buzz agent memory retrieval by seeding a memory then asking
the buzz agent a question it needs that memory.
**Bug Found**: System prompt had no inclusion of retrieving cold
memories and suggested looking in a mem/*.md directory that does not
exist. Updated `system-prompt.md` to include memory CLI tools and usage.
Eval Before System Prompt Change: 0/3
Eval After System Prompt Change: 3/3
## What
- Add a `memory-retrieval` benchmark that seeds agent memory with `buzz
mem set` before asking a direct question.
- Grade the observable threaded answer without inspecting tool calls or
exposing the answer in channel history.
- Teach agents to use `buzz mem set`, `buzz mem ls`, and `buzz mem get`
for cold memory.
- Add a wire-debug endpoint configuration for diagnosing ACP tool calls
in local runs.
- Add fixture, seeding, verifier, and prompt coverage.
## Risk Assessment
Low. The runtime changes are limited to the benchmark harness. The
production-facing change clarifies existing memory commands in the base
prompt; it does not change memory storage, relay behavior, or
authorization.
## References
- Before the system-prompt changes, 0/3 attempts passed because agents
never invoked the `buzz mem` CLI and instead searched a non existent
filesystem
- After the changes, 3/3 attempts passed. ACP wire logs confirmed that
every agent ran `buzz mem ls` followed by `buzz mem get` and returned
`net_gpv`.
---------
Signed-off-by: Philip Azar <pazar@squareup.com>
* fix(ci): salvage Codex review output on PTY-shutdown hang (#7042)
Codex CLI can leave a PTY descendant holding the action's inherited
stdio after the turn completes. The `runCodexExec.ts` wrapper waits on a
`close` event that never fires, so the `Review pull request` step hangs
until the job timeout kills it — discarding the finished review the CLI
already wrote to disk.
The CLI writes the completed review to the `--output-last-message` file
(exposed as `output-file`) **before** the hang. This PR adds a salvage
step that recovers it, and sets the step and job timeouts to preserve
the full 30-minute Codex execution budget.
**Changes (`codex-security-review.yml`):**
- Add `output-file: ${{ runner.temp }}/codex-review.json` to the `Review
pull request` step so the CLI writes the result before the hang.
(`runner` context is valid in `steps.with`; not in `jobs.env`.)
- Add `timeout-minutes: 30` and `continue-on-error: true` to the Codex
step — a hang now costs ≤30 minutes instead of 40, and the salvage step
still runs.
- Set job `timeout-minutes: 40` to give setup, step cancellation, and
salvage sufficient headroom without colliding with the Codex execution
budget. The original 30-minute job timeout was too narrow: evidence from
run
[33114428326](https://github.com/block/buzz/actions/runs/33114428326/job/98665369165)
shows completed output appearing 28m46s after step start, meaning a
20-minute step timeout could kill a legitimate review before the salvage
file exists.
- Add a `Salvage review output` step with `if: always()`: prefers
`steps.run_codex.outputs.final-message` on a clean exit; falls back to
the output file when the step timed out. The output file path is set in
the step's own `env` block (`CODEX_OUTPUT_FILE: ${{ runner.temp
}}/codex-review.json`), where `runner` is valid. Validates shape
(non-empty JSON object, has `overall_risk`); fails the job hard if
neither source is present.
- Wire the job `outputs.review_json` to
`steps.salvage.outputs.review_json`.
**Changes (`Justfile`, `ci.yml`):**
- Add `actionlint .github/workflows/codex-security-review.yml` to
`security-review-check` so expression-validity errors are caught
locally.
- Provision `actionlint` via Hermit (pinned v1.7.12) rather than a
one-off `Install actionlint` curl step, so the same binary is used
locally and in CI.
**Security posture is unchanged:** the salvage step reads the action's
own output and a file written to `runner.temp` — neither is
PR-controlled. Credential-stripping env block on the Codex step is
untouched.
Note this is a temporary workaround until
https://github.com/openai/codex-action/issues/169 is addressed
---------
Signed-off-by: Will Pfleger <pfleger.will@gmail.com>
Co-authored-by: Duncan <dcfd242e557282d7a1e2cf2e6877522682f1e5c6156dc92ca7d90eaedd3b0f95@buzz.block.builderlab.xyz>
* feat: render agent avatars as squircles (#7106)
## Summary
- render every agent/AI identity as a 30% squircle across desktop and
mobile while keeping human avatars circular
- propagate agent identity through message, thread, profile, reaction,
member, DM, search, workflow, project, huddle, forum, pulse, and
agent-management surfaces
- preserve squircle geometry for fallbacks, focus/status treatments,
add-agent controls, and overlapping avatar outlines (`calc(30% + 2px)`
for the outer background)
### Related issue
None found. This change was requested and visually reviewed in the
originating Buzz thread.
### Testing
- `just desktop-test` — 5,799 passed
- `just mobile-test` — 2,008 passed
- pre-push gates passed at `0d59d77b120dcb90aac2f918e422c11c9fa5353b`:
desktop check, TypeScript typecheck, desktop full test suite, mobile
format/analyze and full test suite, Rust tests, Tauri checks, and
differential file-size gate
- deterministic desktop visual sweep covered channel messages/thread
summaries; thread, subthread, and sub-subthread depths; reactions and
reactor popovers; hover/full profiles; added-to-channel activity;
channel members/settings; agent library/team overlaps; agent creation;
mention autocomplete; and DM header/sidebar/settings
### UI evidence
The complete labeled visual matrix is available in the originating Buzz
review thread. GitHub-hosted copies will be added in a follow-up PR
comment using the repository screenshot script.
---------
Signed-off-by: Fizz <dae5f6af70b8695a8b83c8deae555f63be41630ec2b8cd493e41a439c9527dd8@buzz.block.builderlab.xyz>
Co-authored-by: Fizz <dae5f6af70b8695a8b83c8deae555f63be41630ec2b8cd493e41a439c9527dd8@buzz.block.builderlab.xyz>
* fix(acp): wake agents from workflow messages (#6953)
> Pinky, an AI agent, is opening this PR on Wes's behalf.
## Summary
Workflow-generated messages can contain a valid agent mention but still
fail the ACP inbound author gate because the relay signs the event. This
keeps the existing wake policy and gives ACP a narrowly verified
effective author:
- preserve the workflow owner's existing `p` tag and all
rendered-mention `p` tags
- add explicit `["buzz:workflow-owner", <owner hex>]` provenance to
relay-generated workflow messages
- add `["buzz:workflow-mention", <agent hex>]` authority only for
mentions resolved from the stored, unrendered workflow step template
- accept that owner only for a verified kind-9 event signed by the
relay's current NIP-11 `self` key, with unique canonical workflow
metadata and an explicit workflow mention for the receiving agent
- route the verified owner through the existing author and in-flight
mode policies in both normal and setup listeners
- refresh relay identity after reconnects, retaining the last verified
key on transient fetch errors while treating a successful response
without `self` as definitive removal
Malformed, duplicate, forged, tampered, wrong-kind, and wrong-relay
attribution all fail closed to the raw event signer. `respond-to=nobody`
remains absolute. Old/mixed-version messages without the explicit
provenance retain their current fail-closed behavior.
## Trust boundary
The workflow owner means **“scheduled by,” not “authored every rendered
word.”** Trigger-controlled substitutions may still produce ordinary `p`
mention routing for compatibility, but they cannot mint
`buzz:workflow-mention` authority. Only a target named in the durable
owner-authored step template can receive that authority.
The author gate is not bypassed: after relay signature/provenance
verification, the effective owner is evaluated under the same
`owner-only`, `allowlist`, DM, and `nobody` policies used for ordinary
messages. Owner control commands continue to use the raw event signer.
## Why this PR
This is the focused immediate fix for waking an **online** agent from a
stored workflow mention. Earlier attempts were not a finished mergeable
fix and had materially different or incomplete trust designs. Larry's
larger draft stack addresses durable delivery across restarts; that
remains valuable future work and can supersede this effective-author
path when it lands.
## Validation
At exact clean commit `fe5b55619fe44176343eefb4cb7fe180df45a7d8`:
- `buzz-relay workflow_sink`: 25/25 passed, including all four ignored
PostgreSQL cases
- `buzz-acp --lib`: 845/845 passed
- `buzz-workflow --lib`: 169/169 passed (2 unrelated PostgreSQL tests
ignored)
- warnings-denied Clippy passed for the changed Rust packages
- `cargo fmt --all -- --check` passed
- `git diff --check` passed
- repository pre-push gates passed, including branch-scoped Rust tests
- CI now selects the ACP library tests and the relay's pure + PostgreSQL
workflow-sink tests so these guards cannot silently remain unexecuted
The production event-to-author gate is shared by normal and setup
listeners and has biting regression tests for accepted explicit
attribution, legacy owner-`p` rejection, and forged-attribution
rejection.
## Exact-head local relay + ACP proof
Following the release-binary/local-relay shape in `TESTING.md`, the
exact commit above passed a fresh isolated real-process matrix using:
- a freshly recreated Postgres database with migrations
- isolated Redis
- exact-head release `buzz-relay`, `buzz`, `buzz-admin`, and `buzz-acp`
binaries
- newly provisioned owner, channel, and bot member through the CLI
- workflow creation and triggering through the running relay
- a deterministic ACP protocol subprocess capturing actual
`session/prompt` dispatches
- a NIP-11 `self` value verified against the running relay signer
Cases:
1. A stored explicit workflow mention woke an `owner-only` agent exactly
once.
2. A workflow message without an agent mention did not wake it.
3. A non-relay signer forging every workflow authority tag did not wake
it.
4. Trigger-controlled `{{trigger.text}}` containing `@Wake Agent`
retained ordinary `p` routing but received no authority-bearing
workflow-mention tag and did not wake the agent.
5. `respond-to=nobody` remained absolute for a valid relay-authenticated
workflow mention.
The deterministic ACP subprocess isolates and directly proves relay →
ACP authorization and prompt dispatch without depending on external
model behavior.
## Deployment and residual risk
Relay and ACP changes must be deployed together for the new wake
behavior; mixed versions fail closed. Production paired-deployment proof
remains distinct from the successful local integration run. Setup-mode
behavior has automated coverage but was not a separate case in the
five-case local matrix. Relay-key rotation is observed at ACP
startup/reconnect; transient NIP-11 errors retain the last verified key,
an intentional availability tradeoff documented in code.
---------
Signed-off-by: Pinky <5f5ab050ec58ae208332edd544ebf705221e24c1b86d82a6ca07038a7a8f6ac9@buzz.block.builderlab.xyz>
Signed-off-by: Brain <1a02c72794dcd0f07058a353bc3a81f4028b8c77c92c87fce6d5c8b85970a20b@buzz.block.builderlab.xyz>
Signed-off-by: Wes <wesbillman@users.noreply.github.com>
Co-authored-by: Pinky <5f5ab050ec58ae208332edd544ebf705221e24c1b86d82a6ca07038a7a8f6ac9@buzz.block.builderlab.xyz>
Co-authored-by: LioLionel <62820906+LioLionel@users.noreply.github.com>
Co-authored-by: Brain <1a02c72794dcd0f07058a353bc3a81f4028b8c77c92c87fce6d5c8b85970a20b@buzz.block.builderlab.xyz>
Co-authored-by: Carl <9d00794d3df50972eb8b615511783cab12a77a8fd5dd5edd58073ec73b54bd8b@buzz.block.builderlab.xyz>
* fix(relay): reject a frame on its own acknowledgement channel (#6961)
Pinky, an AI agent, updated this description on Wes's behalf after
taking over the startup investigation.
**Category:** fix
**User Impact:** An EVENT refused by WebSocket admission or handler
saturation receives a correlated `OK(event_id, false, reason)` instead
of an uncorrelated NOTICE, so the client can settle that refusal without
waiting for its publish timeout. Rate-limited refusals also arm client
backoff. This fixes a protocol failure mechanism; it does not establish
that every startup send will succeed or that the reported Desktop
startup incident is fully resolved.
**Problem:** Startup opens several live subscriptions and publishes at
once, and the relay's WebSocket admission gate is a fixed 5-second
window (`ws_admission_budget` = `human_ws_events_per_sec * 5`). If that
shared per-principal quota is exhausted, `enforce_ws_admission`
previously rejected an EVENT with a bare `["NOTICE", reason]`. Quota
pressure is a possible trigger, not proof of the original incident's
complete cause.
A NOTICE carries no event id. Both clients settle a pending publish
*only* from an `OK` keyed by event id (desktop `pendingEvents`, mobile
`_pendingEvents`), so nothing settled — and `handle_text_message`
returns early, so no `OK` ever followed either. The send **could not
fail**; it could only time out at `PUBLISH_TIMEOUT_MS` = 25s. That
explains how this rejection mechanism can produce a roughly 25-second
timeout; attributing the original report to it still requires the actual
startup/send workflow.
The handler-semaphore saturation path had the identical defect, and that
one needs no quota burst to fire.
**Solution:** NIP-01 gives each request type its own acknowledgement
channel, and a rejection is only actionable on the same one. Reject a
REQ with `CLOSED`, an EVENT with `OK(id, false, reason)`, and fall back
to `NOTICE` only where no per-request correlation exists. COUNT refusals
now also use `CLOSED(query_id, reason)` per NIP-45, covering both quota
admission and handler saturation (added in
`cd12c93804b87a24b61075dfd171dc471a0a527f`).
Reason strings are unchanged, so the `rate-limited:` prefix and `retry
in {N}s` hint that existing client gates parse keep working (desktop
`parseRateLimitHint`, mobile `RelayRateLimitGate`, buzz-acp
`set_rate_limit_gate`). Only the frame *type* changes, so
`docs/multi-tenant-relay.md` L7 stays satisfied.
Two notes on how this landed, both worth a reviewer's attention:
1. **A survived mutation became a design change.**
`send_admission_result` originally took a `RejectionTarget` parameter,
and reverting the *second* call site (the per-minute message quota)
survived the whole suite — with Redis unreachable the first quota check
short-circuits, so that line is unreachable in test. Rather than test
around it, the parameter is gone: the target is derived from the frame,
so no call site can name the wrong channel.
2. **The relay fix would have caused a client regression on its own.**
Gate arming lived only in the NOTICE branch. Once rejections arrive as
`OK:false`, `handleOk` failed the send without ever backing off — the
client would retry straight into the same quota. Desktop and Mobile now
arm on a `rate-limited:` OK rejection. ACP was subsequently fixed in
`3b06dd32493596ec650f20abf8805791c50fdc24`: it arms the gate and
re-parks only the refused observer frame, preserving other in-flight
frames. Desktop gets `activateRateLimitIfSignalled` as the single owner
of that prefix test, called from both `handleOk` and the NOTICE branch.
<details>
<summary>File changes</summary>
**crates/buzz-relay/src/rejection.rs** (new)
Owns the admission-rejection concern: `RejectionTarget`,
`rejection_target_for`, `request_rejection_message`,
`send_admission_result`, and `enforce_ws_admission`, moved out of
`connection.rs`. Six tests, two of which drive the real
`enforce_ws_admission` against a real `AppState`.
**crates/buzz-relay/src/connection.rs**
Fix the EVENT handler-semaphore rejection to correlate to the event id;
delegate admission to the new module. Add two tests that drive the real
`handle_text_message` with every handler permit held. Down from 1319 to
1116 lines.
**crates/buzz-relay/src/state.rs**
Widen the existing `test_state` helper to `pub(crate)` so the rejection
tests reuse it rather than adding a ninth copy of `AppState`
construction.
**desktop/src/shared/api/relayRateLimitGate.ts**
Add `activateRateLimitIfSignalled` — one owner for the `rate-limited:`
prefix test, since three inbound frame types now carry it.
**desktop/src/shared/api/relayClientSession.ts**
Arm the gate on a rate-limited OK rejection; route the NOTICE branch
through the same helper. Net zero lines, which keeps this
already-oversized file within the differential ratchet.
**desktop/src/shared/api/relayClientPublishRejection.test.mjs** (new)
Four tests against the real `RelayClient`: a rate-limited OK settles the
pending publish and arms the gate; an ordinary rejection does not arm
it; an accepted OK still resolves.
**mobile/lib/shared/relay/relay_session.dart**
Arm the gate in `_handleOk` for a rate-limited rejection.
**mobile/test/shared/relay/relay_session_test.dart**
Two tests driving the real `publish` + `debugHandleMessage` path.
</details>
<details>
<summary>Validation</summary>
**Mutation-tested — 5 mutations, all now killed.** Each production call
site was reverted to the defective behaviour to confirm a test fails.
This caught two false-negative tests:
| # | Mutation | Result |
|---|----------|--------|
| 1 | `rejection_target_for`: EVENT → `Connection` | 4 tests fail |
| 2 | EVENT handler-semaphore call site → bare NOTICE | **survived at
first** |
| 3 | per-minute quota call site → `Connection` | **survived**; fixed by
removing the parameter |
| 4 | desktop `handleOk` gate arming removed | 1 test fails |
| 5 | mobile `_handleOk` gate arming removed | 1 test fails |
Mutation 2 is the lesson: my first saturation test called
`request_rejection_message` directly, so reverting the real call site
inside the `match` arm left it green. It now drives
`handle_text_message` itself and dies on that mutation.
- `cargo test -p buzz-relay` — 928 passed, 1 failed:
`api::mesh_demo::tests::demo_join_forwarded_arm_round_trips_echo`,
**pre-existing**, reproduced with all changes stashed at `4dd4d73de`.
- `cd desktop && npm test` — 5721 passed, 0 failed (full suite).
- `cd mobile && flutter test` — 1876 passed, 0 failed (full suite).
- `just fmt-check`, `just clippy`, `just desktop-check`, `just
mobile-check`, `just file-size-check` — clean. Desktop's 5 biome
warnings are pre-existing (reproduced with changes stashed).
- All 9 pre-push lanes green, including `rust-tests` and
`desktop-tauri-checks`.
**Not verified:** not reproduced end-to-end against a live relay under a
forced quota burst. The causal chain is source-proven and
mutation-proven at the frame level; the ~25s attribution follows from
`PUBLISH_TIMEOUT_MS` but is not directly measured. A packaged-build
click-through would close that gap.
</details>
Related work: #6957 bounds Desktop HTTP event submission, but safe
retained-operation recovery after exhausted/ambiguous outcomes remains
unfinished. #6998 is the separately reviewable Desktop
readiness/duplicate-subscription slice. Neither is claimed to complete
native before/after startup-send validation.
Diagnosis note: `RESEARCH/DESKTOP_STARTUP_SEND_STALL_2026_08_27.md`
(Brain's workspace).
## Current review disposition (2026-08-28)
The [review on
`cd12c938`](https://github.com/block/buzz/pull/6961#pullrequestreview-5052902510)
identified ACP's missing rate-limited-OK handling. Commit
`3b06dd32493596ec650f20abf8805791c50fdc24` fixes gate arming, re-parking
the specifically refused observer frame, and the stale NOTICE comment.
Two regressions drive the real frame dispatcher. See [the implementation
and validation
response](https://github.com/block/buzz/pull/6961#issuecomment-5455032054).
The Mobile generation-check inline thread is resolved: its `async
publish` returns a failed Future when superseded; it does not throw
synchronously at invocation. No further production change was indicated
by that comment.
The validation counts above describe the original slice, not a new
rerun. At `3b06dd324`, the current GitHub check rollup has successful
completed test/build checks (non-applicable jobs skipped). The
security-review comment still requires review for the current base/head
range; do not read a green authorization job as a completed security
review. Approval and merge remain human decisions.
---------
Signed-off-by: Brain <1a02c72794dcd0f07058a353bc3a81f4028b8c77c92c87fce6d5c8b85970a20b@buzz.block.builderlab.xyz>
Signed-off-by: Wes <wesbillman@users.noreply.github.com>
Co-authored-by: Brain <1a02c72794dcd0f07058a353bc3a81f4028b8c77c92c87fce6d5c8b85970a20b@buzz.block.builderlab.xyz>
Co-authored-by: Carl <9d00794d3df50972eb8b615511783cab12a77a8fd5dd5edd58073ec73b54bd8b@buzz.block.builderlab.xyz>
* feat(desktop): add protected-build Bestie experiment (#6902)
## Summary
Introduces a protected-build boundary for the default-off Bestie
experiment without adding any Bestie product surface.
- Official OSS builds select an empty protected-feature module and emit
no Bestie/Chief metadata or implementation content.
- Protected internal builds select a separate module graph containing
the Bestie experiment definition.
- Within an internal build, Bestie remains disabled until the user opts
in under Settings → Experiments.
- The production build runs an artifact matrix and fails if OSS output
contains protected content or internal output lacks the Bestie manifest.
## Build contract
| Build variant | User opt-in | Result |
| --- | --- | --- |
| Official OSS | Any/forged | Bestie absent from the compiled artifact |
| Protected internal | Off | Bestie available but disabled |
| Protected internal | On | Bestie enabled |
The companion protected-release change is squareup/buzz-releases#91. It
sets `VITE_BUZZ_BESTIE=1`, requires that exact value, forwards it into
the signed macOS build, and asserts the contract in release validation.
## Why this is separate
This gives later Bestie PRs one build-selected import seam. Protected
implementations must be reachable only from the internal module so they
never enter the official OSS module graph.
## Non-goals
- No Bestie persona or provisioning
- No sidebar, app-chrome, or message-toolbar UI
- No entitlement or secrecy claim: the source is public; this boundary
controls official Block artifacts
## Verification
- Exact commit `523cf49ced03cba9be43836a54d6aa5d6923cc82`
- Full `just ci`: 5,673 Desktop tests, 2,773 Tauri tests, 1,860 mobile
tests, Rust/Tauri/web/mobile static checks and builds
- OSS production artifact: scanner confirms no `Bestie`, `Chief of
Staff`, or `builtin:bestie` content
- Internal production artifact: scanner confirms the protected Bestie
manifest is emitted
- Both build orders verified; `dist` retains the requested variant for
Vite/Tauri packaging
---------
Signed-off-by: Arjun Mahanti <arjun@squareup.com>
Signed-off-by: Fizz <fizz@buzz.local>
Signed-off-by: Fizz <dae5f6af70b8695a8b83c8deae555f63be41630ec2b8cd493e41a439c9527dd8@buzz.block.builderlab.xyz>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Fizz <fizz@buzz.local>
Co-authored-by: Fizz <dae5f6af70b8695a8b83c8deae555f63be41630ec2b8cd493e41a439c9527dd8@buzz.block.builderlab.xyz>
* add public descriptions to agent personas (#7126)
**Category:** new-feature
**User Impact:** People can add a short public description to an agent
and see what it does directly on agent cards and profiles.
**Problem:** Agent cards previously showed only a model label, so people
had to open an agent and inspect its instructions to understand its
purpose. Public metadata also needed one trustworthy lifecycle across
local edits, relay catalogs, profiles, and portable snapshots.
**Solution:** Add an optional owner-authored description with a
280-character visible-text policy, publish it as profile `about`, and
prefer it on agent cards while retaining the model fallback. Description
metadata is excluded from the spawn-content hash, remains
definition-owned, and is validated independently at every untrusted or
persistence boundary.
<details>
<summary>File changes</summary>
**desktop/src-tauri/src/commands/agent_config_tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.
**desktop/src-tauri/src/commands/agent_discovery/relay_directory.rs**
Updates relay-directory profile test publication for the expanded
profile contract.
**desktop/src-tauri/src/commands/agent_models_tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.
**desktop/src-tauri/src/commands/agent_models_update.rs**
Preserves the effective `about` value when instance edits republish a
complete profile event.
**desktop/src-tauri/src/commands/agents.rs**
Carries the effective authored description into initial managed-agent
profile publication.
**desktop/src-tauri/src/commands/agents_profile.rs**
Adds `about` to profile reconciliation and keeps description, name, and
avatar synchronized against relay state.
**desktop/src-tauri/src/commands/agents_tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.
**desktop/src-tauri/src/commands/personas/card.rs**
Materializes the definition-owned description before minting a portable
agent card snapshot.
**desktop/src-tauri/src/commands/personas/create.rs**
Normalizes and validates raw authored descriptions before persona
persistence.
**desktop/src-tauri/src/commands/personas/delete_cascade_tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.
**desktop/src-tauri/src/commands/personas/inbound.rs**
Validates descriptions at inbound relay ingress and applies accepted
values to local definitions.
**desktop/src-tauri/src/commands/personas/inbound/catalog_reconcile_tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.
**desktop/src-tauri/src/commands/personas/inbound/inbound_tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.
**desktop/src-tauri/src/commands/personas/mod.rs**
Centralizes raw-byte validation followed by trim/empty normalization for
description writes.
**desktop/src-tauri/src/commands/personas/pending.rs**
Revalidates descriptions before preparing public persona publications.
**desktop/src-tauri/src/commands/personas/sharing.rs**
Carries the optional public description through this managed-agent
compatibility path.
**desktop/src-tauri/src/commands/personas/snapshot.rs**
Materializes definition-owned descriptions into portable instance
snapshots without creating a second persisted authority.
**desktop/src-tauri/src/commands/personas/snapshot/fidelity_tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.
**desktop/src-tauri/src/commands/personas/snapshot/import.rs**
Restores snapshot descriptions onto imported definitions while keeping
linked instance copies absent.
**desktop/src-tauri/src/commands/personas/snapshot/tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.
**desktop/src-tauri/src/commands/personas/update.rs**
Persists persona description edits, republishes linked profiles, and
preserves legacy avatars during complete kind:0 replacements.
**desktop/src-tauri/src/commands/personas/update/name_propagation_tests.rs**
Proves description-only profile sync does not write instance state or
clear a legacy avatar.
**desktop/src-tauri/src/commands/team_snapshot.rs**
Round-trips member descriptions through team snapshots and imported
definitions.
**desktop/src-tauri/src/commands/team_snapshot/tests.rs**
Covers team member description export and import fidelity.
**desktop/src-tauri/src/commands/teams/adopt/apply.rs**
Starts adopted team catalog members without synthesizing an unauthored
description.
**desktop/src-tauri/src/commands/teams/adopt/tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.
**desktop/src-tauri/src/commands/teams/pending/tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.
**desktop/src-tauri/src/commands/teams/sharing/tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.
**desktop/src-tauri/src/egress_guard_tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.
**desktop/src-tauri/src/event_sync_team_catalog_tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.
**desktop/src-tauri/src/managed_agents/agent_description.rs**
Defines the canonical Rust description resolution used by profile
publication and reconciliation.
**desktop/src-tauri/src/managed_agents/agent_events.rs**
Updates managed-agent record construction for the optional public
description field.
**desktop/src-tauri/src/managed_agents/agent_snapshot.rs**
Includes descriptions as snapshot profile `about` metadata and validates
them at decode ingress.
**desktop/src-tauri/src/managed_agents/agent_snapshot_envelope.rs**
Updates managed-agent record construction for the optional public
description field.
**desktop/src-tauri/src/managed_agents/agent_snapshot_tests.rs**
Covers snapshot description export and rejection of unsafe or overlong
imported metadata.
**desktop/src-tauri/src/managed_agents/config_bridge/reader_tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.
**desktop/src-tauri/src/managed_agents/definition_validation.rs**
Adds the shared 280-character visible-text policy for public
descriptions.
**desktop/src-tauri/src/managed_agents/discovery/tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.
**desktop/src-tauri/src/managed_agents/effective_config/tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.
**desktop/src-tauri/src/managed_agents/global_config/tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.
**desktop/src-tauri/src/managed_agents/mod.rs**
Exports the description resolution and validation helpers to
managed-agent consumers.
**desktop/src-tauri/src/managed_agents/nest/render_tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.
**desktop/src-tauri/src/managed_agents/parallelism.rs**
Updates managed-agent fixtures for the optional description field
without changing runtime configuration behavior.
**desktop/src-tauri/src/managed_agents/persona_events.rs**
Adds description to persona event content while deliberately excluding
it from the spawn-relevant content hash.
**desktop/src-tauri/src/managed_agents/persona_events/tests.rs**
Pins description event round-tripping and proves description-only edits
do not change the restart hash.
**desktop/src-tauri/src/managed_agents/personas.rs**
Initializes built-in persona records without authored descriptions for
backward-compatible defaults.
**desktop/src-tauri/src/managed_agents/personas/tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.
**desktop/src-tauri/src/managed_agents/readiness.rs**
Updates managed-agent fixtures for the optional description field
without changing runtime configuration behavior.
**desktop/src-tauri/src/managed_agents/restore.rs**
Includes the effective description in launch-time profile
reconciliation.
**desktop/src-tauri/src/managed_agents/runtime/test_fixtures.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.
**desktop/src-tauri/src/managed_agents/runtime/tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.
**desktop/src-tauri/src/managed_agents/spawn_snapshot/tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.
**desktop/src-tauri/src/managed_agents/team_catalog/tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.
**desktop/src-tauri/src/managed_agents/team_snapshot.rs**
Updates managed-agent record construction for the optional public
description field.
**desktop/src-tauri/src/managed_agents/teams_tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.
**desktop/src-tauri/src/managed_agents/types.rs**
Adds optional description metadata to persona and managed-agent records
and their compatibility projections.
**desktop/src-tauri/src/managed_agents/types/requests.rs**
Accepts optional descriptions on persona create and update IPC requests.
**desktop/src-tauri/src/managed_agents/types/tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.
**desktop/src-tauri/src/migration_avatar_tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.
**desktop/src-tauri/src/persona_catalog.rs**
Parses and validates descriptions at the untrusted community-catalog
boundary.
**desktop/src-tauri/src/persona_catalog_tests.rs**
Covers valid catalog descriptions plus rejection of malformed,
invisible, and overlong values.
**desktop/src-tauri/src/relay.rs**
Publishes and queries kind:0 `about` so relay profiles preserve authored
descriptions.
**desktop/src-tauri/src/relay/tests.rs**
Updates managed-agent/persona fixtures for the optional description
field while preserving the behavior under test.
**desktop/src/features/agents/AGENTS.md**
Documents description ownership, validation, snapshot, hashing, and
display invariants for future changes.
**desktop/src/features/agents/lib/agentDescription.test.mjs**
Pins Unicode counting, paste clamping, trimming, and empty
authored-description behavior.
**desktop/src/features/agents/lib/agentDescription.ts**
Provides shared display resolution, Unicode-scalar counting, and paste
clamping for descriptions.
**desktop/src/features/agents/lib/personaCatalogRelay.ts**
Maps validated catalog descriptions into catalog persona projections.
**desktop/src/features/agents/ui/AgentDefinitionDialog.tsx**
Adds the description draft to create and edit submission while
extracting identity fields from the large dialog.
**desktop/src/features/agents/ui/AgentDescriptionField.tsx**
Renders the public description input, helper copy, and Unicode-aware
near-limit counter.
**desktop/src/features/agents/ui/AgentIdentityCard.tsx**
Generalizes the card second line to show a two-line description or the
existing model fallback.
**desktop/src/features/agents/ui/UnifiedAgentsSection.tsx**
Prefers authored descriptions on persona cards and retains model labels
when no description exists.
**desktop/src/features/agents/ui/personaDialogState.test.mjs**
Verifies edit and duplicate drafts preserve authored descriptions.
**desktop/src/features/agents/ui/personaDialogState.ts**
Seeds authored descriptions into edit and duplicate dialog drafts.
**desktop/src/features/agents/ui/usePersonaActions.ts**
Preserves descriptions when copying catalog personas into local
definitions.
**desktop/src/shared/api/personaTypes.ts**
Defines description-bearing persona wire types in a focused module split
from the size-constrained API type file.
**desktop/src/shared/api/tauriPersonas.test.mjs**
Verifies raw persona descriptions map into the frontend model and absent
values become null.
**desktop/src/shared/api/tauriPersonas.ts**
Maps description fields across Tauri and preserves raw authored bytes
for authoritative Rust validation.
**desktop/src/shared/api/types.ts**
Re-exports the extracted persona types without changing consumer import
paths.
**desktop/src/testing/e2eBridge.ts**
Extends mock persona create, update, publication, and catalog parsing
with production-shaped description behavior.
**desktop/tests/e2e/agents.spec.ts**
Verifies an edited description persists and appears on the agent card.
</details>
### Reproduction Steps
1. Open **Agents**, edit a custom or built-in agent, and enter a
sentence in **Description**.
2. Save the agent and confirm the sentence appears as the second line on
its card.
3. Reopen the agent and confirm the authored description is restored;
clear it and confirm the card returns to the model label.
4. Paste more than 280 Unicode characters and confirm the field keeps
the first 280 characters and shows the near-limit counter.
5. Share or export/import the agent and confirm the description survives
in the catalog/profile or snapshot without showing a restart-required
badge for a description-only edit.
### Screenshots / Demo
The focused Playwright flow `built-in persona edits persist` exercises
the edited dialog, persisted value, and resulting card subtitle.
Screenshots can be added after review if the field placement or two-line
card treatment needs visual iteration.
### Verification
- `cargo test --manifest-path desktop/src-tauri/Cargo.toml --lib` —
3,029 passed
- `cd desktop && pnpm test` — 5,805 passed
- `cd desktop && pnpm exec tsc --noEmit`
- Focused Playwright: `built-in persona edits persist` — passed
- Pre-push desktop, Tauri, typecheck, test, file-size, and branch-skew
gates — passed
---------
Signed-off-by: tulsi <tulsi@block.xyz>
* fix(desktop): back split thread headers (#7137)
## Summary
- render an auxiliary panel's requested header backdrop in docked/split
mode
- preserve explicit transparent-backdrop behavior
- cover a populated, scrolled thread pane so timeline content cannot
bleed through its header
## Root cause
`RightAuxiliaryPane` correctly paints above the channel's shared header
backdrop so close/edit controls remain visible. The docked
`AuxiliaryPanelHeader` branch, however, ignored its `backdrop` request,
leaving scrolled thread content in that higher stacking context
unbacked.
## Verification
- desktop unit suite: 5,801 passed
- desktop TypeScript: passed
- Biome checks: passed (existing unrelated repository warnings only in
the earlier full run)
- targeted Playwright scroll regression: passed
- ultrawide thread-pane Playwright coverage: passed
Signed-off-by: Wintermute <c0fc581234c3585602139eec347ced7b82af65b6f6c10728348515c0c06c51c3@buzz.block.builderlab.xyz>
Co-authored-by: Wintermute <c0fc581234c3585602139eec347ced7b82af65b6f6c10728348515c0c06c51c3@buzz.block.builderlab.xyz>
* docs: add review-proven failure-path & async-state rules to AGENTS.md (#7061)
Mining the last 25 PRs' review threads (45 substantive findings, 11
reviewed PRs, avg **4.8 review rounds** each) shows **53% of findings
are repeats** of five clusters: swallowed failures, stale-async-state
races, tests that don't bind the production seam, unbounded
resources/retry loops, and non-atomic multi-step persistence. PR #6956
alone burned 4 rounds converging on one of these classes.
A second, independent mining pass over **71 agent-review rooms (303
findings, Aug 18–29)** confirmed the same clusters and added outcome
data — how often authors actually fix each finding class once flagged:
test-seam binding and unbounded-resource findings **100%**, swallowed
errors **90%**, stale-state races **70%**. It also surfaced two clusters
the GitHub-thread pass under-sampled: **assistive-semantics defects**
(44 findings, second-largest cluster) and **input-modality divergence**
(27 findings), now rules 7–8.
This PR distills those clusters into eight imperative rules in AGENTS.md
so agents apply them **before writing code**, adds one
client-consumption invariant to ARCHITECTURE.md §5, and places the
test-quality rule in TESTING.md (per the team decision that testing docs
are the canonical guide for review standards), cross-referenced from
AGENTS.md. Each rule cites the PRs where it was litigated. Raw mining
data: `reviews.jsonl` / `comments.jsonl` +
`backfill/buzz-review-findings.jsonl` (review-mining artifacts, not
committed).
No code changes. CLAUDE.md is a symlink to AGENTS.md and picks this up
automatically.
🤖 Drafted by Jude's agent from automated mining of this repo's last 25
PRs' review threads and 71 agent-review rooms; every rule cites the PRs
where it was litigated. Jude reviews and owns the result. Mining method
+ raw cluster data available on request.
---------
Signed-off-by: Jude Edwards <judeedwards@squareup.com>
* feat(buzz-acp): give each channel thread its own agent session (#6732)
## What this does
In a channel, people often run several unrelated conversations at once
(separate threads). Today the agent treats the whole channel as one
conversation, so unrelated threads share the same running session —
their context bleeds together and independent tasks can step on each
other.
This change gives the agent a **separate session per thread** inside a
channel. Direct messages stay as one conversation (unchanged). The
channel is still the boundary for who is allowed in and what is visible
— only the agent's working context is now split by thread.
## How it is turned on
Off by default. Operators opt in with one setting:
- `BUZZ_ACP_SESSION_POLICY=channel` — default, current behavior
- `BUZZ_ACP_SESSION_POLICY=thread` — new per-thread behavior
Being behind a flag means we can enable it for a few agents, watch how
it behaves, and roll back instantly without a code change.
## Key design decisions
- **Decide the thread once, up front.** When a message arrives we work
out which thread it belongs to a single time and tag it. Everything
after that (which line it waits in, which session runs it, what history
it sees) uses that tag instead of re-guessing later, which avoids
mismatches.
- **Default stays identical to today.** Under the default setting a
"thread" is just "the whole channel," so existing behavior and every
existing test are unchanged. The new, riskier behavior is strictly
opt-in.
- **Give the agent only its thread's history.** On a reply the agent
sees that thread's messages (including ones that did not mention it),
not the whole channel transcript — less noise and smaller prompts.
- **Don't let one channel use more memory than before.** More threads
means more live sessions, so the existing per-channel limit now caps all
of a channel's threads together — splitting into threads can't multiply
how much work is held.
## Bugs found and fixed while iterating (from review)
- **Same thread, two sessions.** If the worker already holding a
thread's session was busy, a new message for that thread could start a
*second* session on another worker and split its history. Now it waits
for the right worker instead of forking.
- **Interrupting the wrong thread.** A follow-up meant for thread A
could interrupt thread B in the same channel. Interrupts now target the
exact thread.
- **Stuck thread after a crash.** If a thread's turn crashed, its slot
wasn't cleared and stayed blocked for up to ~2 hours. It now clears
right away and retries.
- **Lost the original request.** When a thread was interrupted and then
had to wait for a busy worker, only the follow-up was kept and the
original request was dropped. The full request is now preserved on
retry.
- **Same thread seen as two.** Two spellings of the same thread id
(upper/lower case) could be treated as different threads. Normalized so
they count as one.
## Not in this PR
- The desktop Settings toggle and rollout wiring for managed agents —
https://github.com/block/buzz/pull/6909
- One pre-existing retry edge case (present today without this flag,
unrelated to this change) — tracked separately so this PR stays focused.
## Testing
The full `buzz-acp` test suite passes (830+ unit and integration tests),
plus new focused tests for thread routing, session reuse, interrupt
targeting, crash recovery, and request preservation. Behavior with the
flag off is unchanged.
---------
Signed-off-by: Salman Mohammed <smohammed@squareup.com>
Signed-off-by: Leo <5faf251baee50ee6bcde338aef6acdd70bb3e60115664c2cd490d94a55dfc488@buzz.block.builderlab.xyz>
Co-authored-by: Leo <5faf251baee50ee6bcde338aef6acdd70bb3e60115664c2cd490d94a55dfc488@buzz.block.builderlab.xyz>
* feat(db): add NIP-FI identity and final-admission schema foundation (#6994)
PR 2 of the NIP-FI plan: the schema foundation. Establishes the durable
server-side identity ledger and final-admission surface that the runtime
phases build on. All of Phase A's migrations live here; later phases own
their own deltas.
Depends on nothing — PR 1 (#6776, merged) owned zero migration files.
This PR's relations are shaped to store exactly what PR 1's verifier
produces: issuer-qualified identity and the four denial classes. They
meet in a later PR that writes a verified assertion into these tables in
one transaction.
## Two internally-ordered migrations
- `0041_nip_fi_identity_foundation.sql` (migration A) — core identity +
base-lifecycle relations (5 tables): issuer-qualified `(iss, sub)`
bindings, lifecycle history/selectors, enrollment policies, and
operation receipts. Applies cleanly to current `main`.
- `0042_nip_fi_authorization_foundation.sql` (migration B) — the
final-admission surface (10 tables): authorization events + capacity,
admission results, replay/receipt guards, audit, invalidation
domains/floors, protected-object authority, authority epochs, and
restore version deltas. Applies to A's resulting state.
Fifteen NIP-FI relations total, zero dangling foreign keys. Identity is
issuer-qualified throughout — no single-global-issuer assumption in any
relation, no `Block`-hardcoding. A single deployment may run one issuer;
that is config, not schema.
## Durable, immutable ledger posture
All 15 relations are append-only (immutable `no_delete`/`no_truncate`
triggers) and carry `community_id` as provenance, not ownership. Both
migrations widen the single SQL source of truth
`community_write_fence_excluded_table` so the relations are never
fence-attached, never purged on community deletion, and never counted as
tenant-scoped drift by the deletion control plane's exact-set catalog
check — the same posture main already applies to `product_feedback` and
`rate_limit_violations`. `schema/schema.sql` keeps one consolidated
definition of that function whose exclusion array byte-matches `0042`,
guarded by a parity assertion so a future consolidation cannot silently
drop NIP-FI relations from the ledger.
This makes a tenant's identity/authorization ledger survive community
deletion, per the spec's `FI-INV-02` (durable binding) and `FI-INV-03`
(tombstone monotonicity) and `NIP-FI.md`'s "durable server state"
ruling. `communities(id)` FK never dangles: community rows become
permanent tombstones, never hard-deleted.
## Authorization shape and cardinality contracts
Authenticated `OperatorDenied` events (`actor_kind` 1–3, non-null
`request_fingerprint`) carry a null `semantic_fingerprint` and commit
without a denial-attempt row. The denial-attempt cardinality and shape
guards are scoped to unresolved pre-auth kind-9 events (`actor_kind =
4`). Applied and no-op lifecycle receipts (`outcome_code IN (1, 3)`)
require exactly one mapped success-transition event; denied lifecycle
receipts (`outcome_code = 2`) require zero events from the complete core
lifecycle success-transition class (kinds 1, 2, 3, 6: enrolled, revoked,
rotated, retired) — any such event paired with a denied receipt would
record a transition that never occurred.
## Mined vs. new
Re-cut from Franco's #1476 (`0029`/`0030`) and Cea's #4772 committer
schema, re-cut along FK topology and renumbered above the live `main`
tip. The buzz-auth core of #1476 is Cea-authored; `Co-authored-by`
reflects verified per-commit authorship of the mined schema.
Zero Rust/`deletion.rs` edits — the migration-only exclusion widening
keeps `EXPECTED_SCOPED_TABLES` untouched.
---------
Signed-off-by: Will Pfleger <pfleger.will@gmail.com>
Signed-off-by: Hayt <9e1c23a3fd83f61da34420e4e88ff1b16e45cafcc0cd9019eb07d4ecfa8ca9b0@buzz.block.builderlab.xyz>
Co-authored-by: Cea Stapleton Cordasco <261786559+cea@users.noreply.github.com>
Co-authored-by: Hayt <9e1c23a3fd83f61da34420e4e88ff1b16e45cafcc0cd9019eb07d4ecfa8ca9b0@buzz.block.builderlab.xyz>
* fix(model-capabilities): humanize databricks goose model names (#7135)
🤖
## Summary
- add curated human-readable labels for Databricks Goose models that
otherwise render as fully qualified identifiers
- render `data_workflow_tools.goose.goose-glm-5-3` as `GLM-5.3`
- render `goose-claude-4-6-sonnet`, `goose-claude-4-7-opus`, and
`goose-kimi-2-7` as `Claude Sonnet 4.6`, `Claude Opus 4.7`, and `Kimi
2.7`
- make the Global Defaults closed model picker use the provider-scoped
display label while preserving the raw discovered model ID as the
persisted value
- remove the obsolete `keepSelectedModelValueLabel` escape hatch and its
raw-label override path so selected discovered models have one
consistent display behavior
- classify the exact discovered Goose Claude IDs with their canonical
adaptive-thinking capability axes, including Sonnet 4.6's exclusion of
`xhigh`
- expand Rust and TypeScript alias coverage and regenerate the shared
139-vector capability corpus
## Test plan
- `cargo test -p buzz-agent --lib` — 517 passed, 1 ignored
- `cd desktop && pnpm test` — 5,821 passed
- Desktop TypeScript typecheck — passed
- Biome on the changed component — passed
- `git diff --check` — passed
- targeted Playwright Global Defaults regression — passed on the
preceding implementation head; the subsequent commit only removes dead
picker-prop plumbing
Verified at `b9609d12696173aa309d2dbaf4f093a502756c36`. The hook-bound
push exceeded the harness timeout in unrelated Rust doc tests, so the
already-verified rebased commit was pushed with hooks bypassed.
Follow-up to #6955.
---------
Signed-off-by: Kalvin Chau <kalvin@block.xyz>
Co-authored-by: am <6e30cd56c30e030cd31bb0939b94a7c257c9a09d5ba2d92cf2735da45629f248@buzz.block.builderlab.xyz>
* feat(desktop): add isolated named demo builds (#6407)
🤖 I’m Larry, updating this description on Logan’s behalf.
## Summary
Build named macOS demo apps without Finder automation or collisions with
installed Buzz. `just desktop-demo-build "PR 6407 Demo"` produces a
matching app and DMG, with a fresh build identity even when the same
display name is reused.
- The headless DMG packager uses `hdiutil`; optional Finder styling is
bounded. The existing production release recipe is unchanged.
- Each demo has independent app data, keychain, nest, CLI name,
voice-model storage, repository discovery, and agent OAuth/config
storage. Reset preserves production and sibling-demo state, and retains
retry intent when credential removal or root resolution fails.
- Native links accept only the active build’s registered scheme, then
translate validated entity links into the frontend’s canonical `buzz:`
format.
- The recipe builds all six executable sidecars. Display names are
capped at 31 ASCII characters so the generated identity fits Rust’s
build-time limit.
**Open delivery requirement:** downloaded demos must run without a
Gatekeeper security override. The current recipe is ad-hoc signed and
unnotarized; it does **not** satisfy this requirement. Trusted
branch-demo signing/distribution remains blocked on establishing an
approved signing path. This PR is not being presented as complete
download-and-run delivery.
### Related issue
N/A — reported in the Buzz DMG-packaging workstream.
### Testing
At `11ce21ff97cb387ad676e7caa65b00964097d0bb`, macOS Blox passed the
Tauri workspace suite and compiled-flags gate (including the full
named-demo state; each library pass: 2,992 passed, 19 ignored), Tauri
all-target clippy, the full `buzz-agent` package suite, and frontend
lint/typecheck plus 5,733 tests. Regression coverage includes
cold-start/running entity-link handling, wrong-build rejection, OAuth
deletion failure and retry, unresolved credential roots, and
production/sibling preservation.
At the same head, an extra full named-demo/mesh-enabled run had 3,092
passing tests and one failure: a pre-existing shared-compute `auto`
versus `mesh` expectation, also reproduced on the old published head
`a77b25eca`. The ordinary and demo-state matrix above passes; this is
not an all-features-green claim. Live macOS Launch Services delivery
remains unverified.
GitHub CI completed with 30 successful checks and 9 skipped. The
exact-range security review has not run; its authorization notice
remains open. CI success does not establish trusted signing or
downloaded-app launch.
Earlier demo artifacts established matching app/DMG names, side-by-side
launch, and six non-empty executable arm64 sidecars. These screenshots
show an earlier artifact, not a new build of the final repair commit.
Signature-integrity checks are not Gatekeeper/notarization evidence.
<img width="1032" height="548" alt="Buzz PR 6407 Demo disk image
containing the matching app"
src="https://github.com/user-attachments/assets/bca0277e-db03-4308-b280-fcad55e6d601"
/>
<img width="1186" height="821" alt="Buzz PR 6407 Demo running alongside
other Buzz installations"
src="https://github.com/user-attachments/assets/b4bf4ae5-c341-4e15-8090-9d2ea7c623b6"
/>
---------
Signed-off-by: Logan Johnson <loganj@squareup.com>
Signed-off-by: Larry <8cf5a83f590ec0955b11647d1c88f796a98e088c30a492c58e0e46c3026ae7a4@buzz.block.builderlab.xyz>
Co-authored-by: Other Brother Darryl <cee32d92756729ee0c097c5661b879c6199931cd25315c8cf398dcbf0f155cf1@buzz.block.builderlab.xyz>
Co-authored-by: Larry <loganj+sandbox-larry@squareup.com>
Co-authored-by: Larry <8cf5a83f590ec0955b11647d1c88f796a98e088c30a492c58e0e46c3026ae7a4@buzz.block.builderlab.xyz>
---------
Signed-off-by: Philip Azar <pazar@squareup.com>
Signed-off-by: Will Pfleger <pfleger.will@gmail.com>
Signed-off-by: Fizz <dae5f6af70b8695a8b83c8deae555f63be41630ec2b8cd493e41a439c9527dd8@buzz.block.builderlab.xyz>
Signed-off-by: Pinky <5f5ab050ec58ae208332edd544ebf705221e24c1b86d82a6ca07038a7a8f6ac9@buzz.block.builderlab.xyz>
Signed-off-by: Brain <1a02c72794dcd0f07058a353bc3a81f4028b8c77c92c87fce6d5c8b85970a20b@buzz.block.builderlab.xyz>
Signed-off-by: Wes <wesbillman@users.noreply.github.com>
Signed-off-by: Arjun Mahanti <arjun@squareup.com>
Signed-off-by: Fizz <fizz@buzz.local>
Signed-off-by: tulsi <tulsi@block.xyz>
Signed-off-by: Wintermute <c0fc581234c3585602139eec347ced7b82af65b6f6c10728348515c0c06c51c3@buzz.block.builderlab.xyz>
Signed-off-by: Jude Edwards <judeedwards@squareup.com>
Signed-off-by: Salman Mohammed <smohammed@squareup.com>
Signed-off-by: Leo <5faf251baee50ee6bcde338aef6acdd70bb3e60115664c2cd490d94a55dfc488@buzz.block.builderlab.xyz>
Signed-off-by: Hayt <9e1c23a3fd83f61da34420e4e88ff1b16e45cafcc0cd9019eb07d4ecfa8ca9b0@buzz.block.builderlab.xyz>
Signed-off-by: Kalvin Chau <kalvin@block.xyz>
Signed-off-by: Logan Johnson <loganj@squareup.com>
Signed-off-by: Larry <8cf5a83f590ec0955b11647d1c88f796a98e088c30a492c58e0e46c3026ae7a4@buzz.block.builderlab.xyz>
Signed-off-by: shiv <shivchander.s30@gmail.com>
Co-authored-by: Phil Azar <pazar@squareup.com>
Co-authored-by: Will Pfleger <pfleger.will@gmail.com>
Co-authored-by: Duncan <dcfd242e557282d7a1e2cf2e6877522682f1e5c6156dc92ca7d90eaedd3b0f95@buzz.block.builderlab.xyz>
Co-authored-by: Arjun Mahanti <arjun.mahanti@gmail.com>
Co-authored-by: Fizz <dae5f6af70b8695a8b83c8deae555f63be41630ec2b8cd493e41a439c9527dd8@buzz.block.builderlab.xyz>
Co-authored-by: Wes <wesbillman@users.noreply.github.com>
Co-authored-by: Pinky <5f5ab050ec58ae208332edd544ebf705221e24c1b86d82a6ca07038a7a8f6ac9@buzz.block.builderlab.xyz>
Co-authored-by: LioLionel <62820906+LioLionel@users.noreply.github.com>
Co-authored-by: Brain <1a02c72794dcd0f07058a353bc3a81f4028b8c77c92c87fce6d5c8b85970a20b@buzz.block.builderlab.xyz>
Co-authored-by: Carl <9d00794d3df50972eb8b615511783cab12a77a8fd5dd5edd58073ec73b54bd8b@buzz.block.builderlab.xyz>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Fizz <fizz@buzz.local>
Co-authored-by: tulsi <tulsi@block.xyz>
Co-authored-by: thomaspblock <thomasp@squareup.com>
Co-authored-by: Wintermute <c0fc581234c3585602139eec347ced7b82af65b6f6c10728348515c0c06c51c3@buzz.block.builderlab.xyz>
Co-authored-by: Jude Edwards <judeedwards@squareup.com>
Co-authored-by: Salman Mohammed <smohammed@squareup.com>
Co-authored-by: Leo <5faf251baee50ee6bcde338aef6acdd70bb3e60115664c2cd490d94a55dfc488@buzz.block.builderlab.xyz>
Co-authored-by: Cea Stapleton Cordasco <261786559+cea@users.noreply.github.com>
Co-authored-by: Hayt <9e1c23a3fd83f61da34420e4e88ff1b16e45cafcc0cd9019eb07d4ecfa8ca9b0@buzz.block.builderlab.xyz>
Co-authored-by: Kalvin C <kalvinnchau@users.noreply.github.com>
Co-authored-by: am <6e30cd56c30e030cd31bb0939b94a7c257c9a09d5ba2d92cf2735da45629f248@buzz.block.builderlab.xyz>
Co-authored-by: Logan Johnson <loganj@squareup.com>
Co-authored-by: Other Brother Darryl <cee32d92756729ee0c097c5661b879c6199931cd25315c8cf398dcbf0f155cf1@buzz.block.builderlab.xyz>
Co-authored-by: Larry <loganj+sandbox-larry@squareup.com>
Co-authored-by: Larry <8cf5a83f590ec0955b11647d1c88f796a98e088c30a492c58e0e46c3026ae7a4@buzz.block.builderlab.xyz>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
## Why Thread-scoped ACP sessions from #6732 need an opt-in desktop rollout that preserves today’s channel policy by default. ## What - Add the default-off “Thread Scoped ACP Sessions” toggle under Settings → Experiments using the existing feature-override persistence. - Map the persisted setting to `BUZZ_ACP_SESSION_POLICY=channel|thread` for every local and provider-backed managed ACP launch. Changes apply when managed agents next start; DMs remain conversation-scoped by the backend. - Keep the desktop setting authoritative over descriptor environment and cover the UI, persistence, local spawn, provider payload, and existing Projects/Workflows experiments. ## Risk Assessment Low. The experiment defaults off and explicitly preserves `channel`; changes are limited to managed-agent launch configuration. Existing running agents are unchanged until their next start. ## Testing - `. ./bin/activate-hermit && just ci` — passed on the final restacked tree, including 5,674 desktop tests, 2,796 Tauri tests, and 1,860 mobile tests. - `cd desktop && pnpm build:e2e && pnpm exec playwright test tests/e2e/experimental-features.spec.ts --project=smoke` — 1 passed. - `. ./bin/activate-hermit && cargo test --manifest-path desktop/src-tauri/Cargo.toml session_policy --lib` — 4 passed. - `. ./bin/activate-hermit && cargo test --manifest-path desktop/src-tauri/Cargo.toml commands::agents::deploy::tests --lib` — 15 passed. - `. ./bin/activate-hermit && cargo test -p buzz-backend-kubernetes --test wire_fixtures` — 4 passed. ## Stack Info Stacked on #6732. This PR depends on #6732 and should merge after it. Generated with Codex --------- Signed-off-by: Salman Mohammed <smohammed@squareup.com> Signed-off-by: Leo <5faf251baee50ee6bcde338aef6acdd70bb3e60115664c2cd490d94a55dfc488@buzz.block.builderlab.xyz> Co-authored-by: Leo <5faf251baee50ee6bcde338aef6acdd70bb3e60115664c2cd490d94a55dfc488@buzz.block.builderlab.xyz>
What this does
In a channel, people often run several unrelated conversations at once (separate threads). Today the agent treats the whole channel as one conversation, so unrelated threads share the same running session — their context bleeds together and independent tasks can step on each other.
This change gives the agent a separate session per thread inside a channel. Direct messages stay as one conversation (unchanged). The channel is still the boundary for who is allowed in and what is visible — only the agent's working context is now split by thread.
How it is turned on
Off by default. Operators opt in with one setting:
BUZZ_ACP_SESSION_POLICY=channel— default, current behaviorBUZZ_ACP_SESSION_POLICY=thread— new per-thread behaviorBeing behind a flag means we can enable it for a few agents, watch how it behaves, and roll back instantly without a code change.
Key design decisions
Bugs found and fixed while iterating (from review)
Not in this PR
Testing
The full
buzz-acptest suite passes (830+ unit and integration tests), plus new focused tests for thread routing, session reuse, interrupt targeting, crash recovery, and request preservation. Behavior with the flag off is unchanged.