Sync fork with upstream block/buzz (398 commits) - #1
Draft
marketing-shibata50 wants to merge 398 commits into
Draft
marketing-shibata50 wants to merge 398 commits into
marketing-shibata50 wants to merge 398 commits into
Conversation
## What - morph the composer plus button into the attachment menu, camera, and photo surfaces - add ordered multi-select with inline recent photos and system picker fallback - add native iOS attachment/photo popovers and align the Android camera treatment ## Stack - follows block#3312 ## Validation - `just mobile-check` - `flutter test test/features/channels/compose_bar_test.dart` - full mobile pre-push suite --------- Signed-off-by: kenny lopez <klopez4212@gmail.com>
## Why Allow operators to install wrapper binaries and override the relay entrypoint without maintaining a duplicated Deployment outside the OSS chart. `extraManifests` can create independent resources but cannot extend the chart-managed relay Pod. ## What - Add opt-in init-container, volume, volume-mount, command, and args extension points - Preserve image defaults when extensions are empty and compose generic init containers with the MinIO readiness gate - Document the distinction from `extraManifests`, add schema coverage, and release chart 0.1.7 ## Risk Assessment Low — all new values are opt-in, and default rendered manifests are unchanged apart from version-derived metadata. Merge publishes a new chart version without modifying existing installations. ## References - [OpenTelemetry Collector Pod extensions](https://github.com/open-telemetry/opentelemetry-helm-charts/blob/main/charts/opentelemetry-collector/templates/_pod.tpl) alongside [extraManifests](https://github.com/open-telemetry/opentelemetry-helm-charts/blob/main/charts/opentelemetry-collector/templates/extraManifests.yaml) - [Argo CD extraObjects](https://github.com/argoproj/argo-helm/blob/main/charts/argo-cd/templates/extra-manifests.yaml) alongside component-scoped Pod extension hooks - `helm unittest` 0.8.2: 43/43 tests passed - Helm lint, schema validation, fixture renders, and chart packaging passed - Oracle review found no functional issues; its literal no-`tpl` regression test recommendation is included Generated with Amp --------- Signed-off-by: David Grochowski <dgrochowski@squareup.com> Co-authored-by: Amp <amp@ampcode.com>
**Category:** fix **User Impact:** Composer block formatting now applies to the intended line or selection without collapsing multiline content. **Problem:** Block formatting from a Shift+Enter line could convert the entire draft, selected visual lines could collapse into one list item, and code conversion could lose line breaks. **Solution:** Scope caret formatting to its hard-break-delimited line and normalize explicit selections for the destination block type while preserving neighboring content and visual line boundaries. <details> <summary>File changes</summary> **desktop/src/features/messages/lib/selectionBlockFormatting.ts** Scopes collapsed-caret block actions to the active visual line and normalizes multiline selections for lists and code blocks. **desktop/src/features/messages/lib/selectionBlockFormatting.test.mjs** Adds unit coverage for caret-line isolation across line positions and selection directions. **desktop/src/features/messages/ui/FormattingToolbar.tsx** Routes list, quote, and code-block actions through the selection-aware formatting transaction. **desktop/tests/e2e/composer-selection-formatting.spec.ts** Covers caret-only formatting, multiline list conversion, list-to-code conversion, preserved hard breaks, Markdown output, and backward selections. </details> ## Reproduction steps 1. In the desktop composer, enter several lines using Shift+Enter and place the caret on one line. 2. Apply a bullet list, ordered list, quote, or code block; only the caret line should change. 3. Select several Shift+Enter lines and apply a list; each visual line should become its own item. 4. Select several list items and apply Code block; they should become one multiline code block while unselected neighbors remain intact. 5. Select several Shift+Enter lines and apply Code block; each line break should remain visible. ## Screenshots/Demos <img width="508" height="222" alt="Screen Recording 2026-07-27 at 5 29 19 PM" src="https://github.com/user-attachments/assets/35640dea-0cfb-44f1-9b0b-a993c69cb55f" /> Expected multiline code-block result: https://buzz.block.builderlab.xyz/media/d2e2668093af3b67d896a32e9799daccd236da9fc9e24ec56ddb4ebf7d01dd96.png --------- Signed-off-by: Taylor Ho <taylorkmho@gmail.com> Co-authored-by: npub1223z34hd7vtwc6qj4s7flsxkj644nlre2nthu7lrrmkumhu3xddsrx9r6w <52a228d6edf316ec6812ac3c9fc0d696ab59fc7954d77e7be31eedcddf91335b@buzz.block.builderlab.xyz>
…ock#3253) ## Summary The desktop client renders a persistent user status (NIP-38 kind:30315, `d:general`) as the status line on profiles, but the CLI had no way to set it — only ephemeral presence (`set-presence`, kind:20001). Integrations that want a scriptable, durable status line (for example a now-playing music bridge that shows the current TIDAL track on a profile) had no entry point. ## Screenshots <img width="1455" height="960" alt="1" src="https://github.com/user-attachments/assets/f1669ec6-212b-4f6e-ad53-07df9aacffc9" /> <img width="1455" height="960" alt="2" src="https://github.com/user-attachments/assets/5bf70f47-e5b5-4eb0-a426-b5f1ef90d2ec" /> This adds: ```bash buzz users set-status --text "Working on the relay" --emoji "🔧" buzz users set-status --text "" --emoji "🎶" # intentional emoji-only status buzz users set-status --clear # removes the status ``` - Signs and submits the replaceable kind:30315 event via the HTTP bridge (no WS needed — unlike presence, user status is a stored event). - Uses the `d:general` coordinate the desktop client already reads for the profile status line, and the same `emoji` tag shape `SetStatusDialog` publishes. - Event construction lives in `buzz_sdk::build_user_status()`, keyed off `buzz_core::kind::KIND_USER_STATUS`, so the CLI command is a thin sign/submit wrapper. Text and emoji are trimmed; a blank emoji is omitted rather than emitted as an empty tag. - Clearing is the explicit `--clear` flag, mutually exclusive with `--text`/`--emoji`. It publishes an empty-content event carrying only `d:general`, which the desktop treats as no status. `--text ""` with an `--emoji` is an emoji-only status, not a clear. --------- Signed-off-by: Kagan Yaldizkaya <kagan@squareup.com> Signed-off-by: Will Pfleger <pfleger.will@gmail.com> Co-authored-by: Will Pfleger <pfleger.will@gmail.com>
The codex adapter version gate accepted any `major >= 1`, so a 1.x `codex-acp` older than the version that fixes outbound relay access for `buzz` CLI subprocesses classified as `Available` and was never offered a reinstall. Only the 0.16.x `@zed-industries/codex-acp` adapter — which fails `--version` outright — was caught. `probe_codex_acp_version` now returns the full `(major, minor, patch)` triple and `codex_adapter_availability` compares it against a new `MIN_CODEX_ACP_VERSION` floor of `1.1.7`, the current npm latest. An adapter below the floor classifies as `AdapterOutdated`, which routes it through the existing uninstall-then-install reinstall plan. The parse requires exactly three numeric dot-separated components. Partial versions (`1.2`) and prerelease tags (`1.2.0-rc1`) return `None` and therefore classify as `AdapterOutdated` — a version Buzz cannot compare against the floor fails closed, offering a reinstall rather than running an adapter of unknown vintage. Both the floor's bump policy and the strict-parse behavior are stated in doc comments rather than left implicit. Supersedes [block#3097](block#3097) by @Bharathchinneni, whose semver floor and behavior tests this carries. That PR could not land as written: the two `probe_codex_acp_major_version` compatibility wrappers it kept had no non-test callers, which is a hard `clippy -D warnings` failure. The wrappers are deleted here and their call sites collapsed onto `probe_codex_acp_version`. Signed-off-by: Will Pfleger <pfleger.will@gmail.com>
## Why The Inbox surface was briefly renamed to **Activity** during block#2045 and picked up a bell icon to match. The name was reverted to **Inbox** before merge, but the icon was not. A bell says "notification tray." Inbox is a destination — a focused, conversation-oriented place to catch up on work relevant to you, including drafts and reminders that have nothing to do with notifications. The glyph should say that. ## What changed - Swap the sidebar entry from Lucide `Bell` to Lucide `Inbox`. - Assert the icon in `inbox-refactor-screenshots.spec.ts`. Nothing pinned it before, which is exactly how it drifted through a rename. This also brings desktop back in line with mobile, which already uses `LucideIcons.inbox300` / `inbox500` for the same destination. ## Deliberately unchanged The bell on **reminder** rows in the list pane (`InboxListPane.tsx`, reminders → bell, drafts → file) stays. A bell is the right glyph for a reminder; that one was never about the surface's identity. ## Verification - The new assertion is a real guard, not a no-op: with `Bell` restored the test fails with `Expected: 1, Received: 0` on `svg.lucide-inbox`. Confirmed before committing. - `biome` and `tsc` clean. - Playwright smoke: `inbox-refactor-screenshots` 4 passed; `smoke`, `navigation`, `channels`, `sidebar-more-unread-overlap`, `home-collapsed-top-chrome`, `workspace-rail` — 107 passed, 1 skipped. - Screenshot below is the regenerated `02-current-controls` shot from the spec. Signed-off-by: Clay Delk <clay.delk@gmail.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
## What - add the shared desktop-style arc spinner for mobile - replace app loading indicators with the shared component - preserve a static pose when reduced motion is enabled ## Stack - follows block#3313 ## Validation - `just mobile-check` - focused spinner and pairing widget tests --------- Signed-off-by: kenny lopez <klopez4212@gmail.com>
Extracts the dense inline DCO paragraph from the "Before You Open a PR" section into a dedicated `### Sign Your Commits` subsection. ## What changed - Adds a `### Sign Your Commits` heading directly below the Conventional Commits paragraph - Leads with the command (`git commit -s`) in a code block - Follows with a plain-English explainer of what the sign-off does - Adds linkable `#### Fix unsigned commits already pushed` and `#### Auto-setup for future commits` subheadings - Removes the old inline paragraph (content preserved, structure only changed) ## Why The existing guidance was buried mid-paragraph; contributors may not find it until CI blocks them. This makes the requirement and its fix immediately visible and actionable. ## Notes Docs-only change, no code modified. Signed-off-by: Cameron Hotchkies <chotchkies@block.xyz> Co-authored-by: npub1ep9tf72jk6xgwamqj5m2j0xvqvwm9vdu3zxlz7cesxg53x52tkkqf6pa42 <c84ab4f952b68c8777609536a93ccc031db2b1bc888df17b198191489a8a5dac@buzz.block.builderlab.xyz>
## Summary Drafts were showing up in the Home Inbox **All** view, mixed in with messages and reminders (reported in `#buzz-bugs`). Drafts are private composer state, not inbox activity — they now appear only under the dedicated **Drafts** filter. ## Changes - **`inboxListRows.ts`** — drop the `draft` row variant from `buildInboxListRows`; the mixed view builds only `inbox` + `reminder` rows. - **`InboxListPane.tsx`** — remove the draft branch of the All-view render path; `PersonalItemRow` now renders reminders only. - **`useHomePersonalInbox.ts`** — stop enabling draft selection (and its root-status relay probing) for the mixed view; draft selection is scoped to the Drafts filter. - Drafts filter behavior is unchanged: the filter badge count, `DraftsPanel` list, and `DraftDetailPane` all still work. ## Testing - `pnpm test` (desktop unit suite): 3697 passed, 0 failed. - `pnpm exec biome check src/features/home tests`: clean. - Updated `inboxListRows.test.mjs` for the two-variant row model. - Updated the e2e test (`channels.spec.ts`) to assert All never lists drafts and that the draft is still reachable under the Drafts filter. - Added `drafts-all-fix-screenshots.spec.ts` capturing both states (screenshots below). ### All view — draft is gone, messages/reminders unaffected  ### Drafts filter — the draft is still listed and editable  Signed-off-by: Thomas Petersen <thomasp@squareup.com>
## Summary - add custom-agent catalog sharing and hide built-ins from discovery - let owners publish later catalog updates from Share or while saving edits — the save always persists locally, and the publish reports `published` or `queued` (flushed automatically once the relay is reachable again) - preserve agent type, model, and runtime across snapshot import/export - simplify agent and team entry points and tighten catalog layout - migrate the legacy global retention queue into the owner's active scope so pending catalog publishes survive the upgrade - keep the agent list and edits usable in recovery mode by degrading to unshared projections when scope resolution or the retention DB fails - track catalog provenance on copied personas, so adding an already-added foreign agent resolves to the existing copy instead of creating a duplicate - scope inbound persona events to the community relay they arrived on - page the catalog read past the relay's 1,000-row query clamp - unify the share dialog's memory-level choice into a single "What's included" selector that drives both DM-send and copy-link delivery (all six combinations preserved), group the delivery rows above the option rows, and label the catalog toggle "Not shared" / "Shared" - keep emoji avatars on catalog entries — they persist as inline percent-encoded SVG, which the catalog projection's http(s)-only URL guard used to drop, so a shared agent showed initials instead of its avatar - drop the "Active in communities" card from Agents settings, superseded by the per-channel runtime controls in the members sidebar ## Screenshots ### Agent actions  ### Team avatar stack  ### Catalog sharing  ### Publish while editing  ### Publish from Share  ### Catalog details  --------- Signed-off-by: kenny lopez <klopez4212@gmail.com> Signed-off-by: Will Pfleger <pfleger.will@gmail.com> Co-authored-by: Will Pfleger <pfleger.will@gmail.com>
Search migrated to Postgres FTS (commit f8bbe6e). The Typesense container was removed from compose.yml and the Helm chart, but the cleanup missed two template/config files: - `deploy/compose/.env.example`: `TYPESENSE_API_KEY` and `TYPESENSE_PORT` are dead — no typesense service exists in compose.yml and the relay binary no longer reads `TYPESENSE_API_KEY`. The `CHANGE_ME_RANDOM_API_KEY` placeholder was never consumed, so removing it also unbreaks the sed loop in the blog draft (one fewer no-op secret to generate). - `benchmarks/harbor-buzz-orchestra/scripts/benchmark.py`: generates a typesense_api_key in state and writes `TYPESENSE_API_KEY` to the .env file it creates. - *Editing this file caused the https://github.com/block/buzz/blob/main/.github/workflows/benchmark-harbor.yml linter ci checks to run, which seemingly haven't run before, so I needed fix the lint issues to pass this.* --------- Signed-off-by: Kalvin Chau <kalvin@block.xyz> Co-authored-by: npub1c4alndp82zyt9veaklm5d965quss79vlhk9awv7qu5erwhmf42qqlvc25c <c57bf9b4275088b2b33db7f746975407210f159fbd8bd733c0e532375f69aa80@buzz.block.builderlab.xyz>
Registering a custom ACP harness works today, but only from Settings → Agents. Anyone whose first touchpoint is "New agent" has no way to discover the custom path — the dropdown just lists the baked-in presets plus whatever was registered earlier. This adds an inline "Add custom harness…" entry to the harness dropdown in all three agent surfaces: create, edit-definition (`AgentDefinitionDialog`), and instance edit (`AgentInstanceEditDialog`). The entry is a sentinel option (`ADD_CUSTOM_HARNESS_VALUE`, NUL-prefixed so it can never collide with a real harness id — backend ids match `[a-z0-9_][a-z0-9_-]*`), mirroring the `CUSTOM_ENTRY_ID` trick already used in `HarnessCatalogDialog`. Picking it never writes into form state; it opens `AddCustomHarnessDialog`, a thin modal wrapper hosting the existing `CustomHarnessForm` in `chromeless` mode. `CustomHarnessForm`'s `onSaved` now carries the saved `definition.id` (the form may rewrite it); the two existing call sites ignore the argument, so their behavior is unchanged. Selection after save is deferred rather than immediate. `usePendingHarnessSelection` holds the saved id until the runtime catalog actually publishes it via discovery, then selects it exactly once — so the dialog never selects an id it cannot render, and back-to-back registrations resolve correctly. The wait is scoped to the owning dialog's `open` state: both host dialogs stay mounted when closed, so an unpublished id is dropped on close rather than selecting into reset form state when discovery later catches up. Selection is routed through each dialog's normal dropdown change handler, so provider/model reset (and command pinning in the instance dialog) behave identically to a hand-picked harness. Dismissing the modal leaves the previous selection untouched. `AgentInstanceEditDialog`'s existing "Custom command" option is a different feature (ad-hoc command override vs. a registered reusable harness) and is untouched. Coverage is 16 unit tests in `addCustomHarness.test.mjs` (real React mount, following the existing `.test.mjs` pattern) plus 4 Playwright specs in `inline-custom-harness.spec.ts` covering all three surfaces end-to-end. Both suites were mutation-verified: treating the sentinel as a real selection, selecting before the catalog publishes, never clearing the pending id, ignoring the dialog's open state, and reversing latest-save-wins each turn the unit tests red; reverting the two dialog diffs turns all four e2e specs red. The `check-file-sizes.mjs` overrides for the two dialogs are ratcheted to their exact new counts (1048 and 1229) — verified tight in both directions, N passes and N−1 fails, so no headroom is introduced. --------- Signed-off-by: Will Pfleger <pfleger.will@gmail.com>
…ck#3144) A Buzz install with a scheduled goose recipe fires each cron entry once per `goose acp` child instead of once, because every child unconditionally starts its own cron scheduler over the shared `~/.local/share/goose/schedule.json`. With a pool of N children per harness and multiple harnesses, one scheduled recipe fans out to N × harness_count executions — each running under the managed agent's identity rather than the operator's, and racing the operator's own standalone goose over the same schedule file. This injects `GOOSE_ACP_SCHEDULER_DISABLED=true` into every child spawned by `AcpClient::spawn`, so a managed agent never owns the operator's cron schedule. ## Placement The `cmd.env` call is set last — after the `extra_env` operator-wins loop and after the `CODEX_CONFIG` merge — deliberately with no escape hatch. Managed children not running the operator's schedule is a correctness invariant rather than an operator-tunable default, so the injection must beat both a conflicting persona `extra_env` entry and any value inherited from the parent process. It is injected for all agents, not just goose. Agent builds that don't recognize the variable ignore it. ## Sequencing The goose-side flag that reads this variable and skips scheduler startup lands separately (repo TBD). Until it does, this change is a forward-compatible no-op: it sets an environment variable nothing currently reads. Merging it first means no coordinated release is needed — the fix takes effect as soon as the goose side ships. Related: aaif-goose/goose#10738 Signed-off-by: Will Pfleger <pfleger.will@gmail.com>
## Summary - paint the community rail across the full app height instead of exposing the parent background through external margins - preserve the existing community-button alignment and balanced horizontal gutters by moving vertical spacing inside the rail - update the rail geometry coverage to require full-height paint ownership ## Root cause PR block#2972 aligned the rail box with the inset content by adding top and bottom margins to the `bg-sidebar` element. Margins are outside the painted box, so flat light and dark themes exposed a differently colored app background above and below the rail. ## Validation - pre-push `desktop-check` - pre-push desktop unit suite: 3,751 passed - `git diff --check` Local Playwright/E2E was not run; CI owns the full browser matrix. Signed-off-by: Wes <wesbillman@users.noreply.github.com> Co-authored-by: Carl <c7ebe626f000404285d3686e1dc74cc07cc60a9754a150041ba132e14bd3e2ec@buzz.block.builderlab.xyz>
…res (block#3396) ## Problem The prerequisites table lists language toolchains (Rust, Node, pnpm, Flutter, Docker, `just`) but no system libraries. Hermit pins the former and not the latter, so following the setup section exactly on Linux still leaves `just ci` unable to run: it fails partway through its first dependency, `just check`, at `desktop-tauri-clippy`. ``` The system library `gdk-pixbuf-2.0` required by crate `gdk-pixbuf-sys` was not found. The file `gdk-pixbuf-2.0.pc` needs to be installed and the PKG_CONFIG_PATH environment variable must contain its parent directory. ``` The desktop crates link against GTK and WebKitGTK. CI installs those packages explicitly, so it never sees this — which is exactly why the gap is invisible from the maintainer side. Since `check` runs first in the `ci` chain, the failure also masks everything after it (`test-unit`, `desktop-test`, `web-build`, `mobile-test` never run), which makes it read as a broken repo rather than a missing dependency. ## Change Adds a `#### Linux: Tauri system libraries` subsection under Prerequisites with: - The apt list copied from `.github/workflows/ci.yml`, so a local run matches CI rather than drifting from it - A pointer to [Tauri's prerequisites](https://tauri.app/start/prerequisites/) for non-Debian distributions - A note that server-side contributors can skip it — `just fmt-check`, `just clippy`, `just test-unit`, and `just test` need no GTK Docs only. No TOC entry needed, since the TOC lists `##` headings and this is a `####` subsection. ## How I hit it Running `just ci` before pushing block#3372, on Ubuntu under WSL2 with the Hermit toolchain active and all Docker services healthy. Everything the guide asks for was in place. The four `check` steps before `desktop-tauri-clippy` (`fmt-check`, `clippy`, `desktop-check`, `desktop-tauri-fmt-check`) passed, which is what makes the failure point specific rather than a general build problem. ## Closest existing work None found. I searched open and closed issues and PRs for `gdk-pixbuf`, `libgtk`, `webkit2gtk`, `system dependencies`, `prerequisites`, `just ci`, and `linux setup`. The Linux/GTK issues that exist (block#2604, block#2643, block#2982, block#2811, block#2562) are all runtime bugs in shipped builds, not setup-path failures. ## Verification The package list is transcribed from `.github/workflows/ci.yml:152-163`; the same list appears in `release.yml` and `linux-canary.yml`. I have not installed the packages on my machine, so I can confirm the failure and the source of the fix but not that the list is exhaustive on a clean box — worth a second pair of eyes from anyone who has done a fresh Linux setup recently. Signed-off-by: Kyler Cao <kcao@gssmail.com>
…lock#2004) ## Summary Fixes 4 flaky DM expansion E2E tests in Desktop Smoke shard 1 that were failing non-deterministically on CI (also reproducing on `main` at run `29526844596`). **Failing tests:** - `channels.spec.ts:652` — creates the DM before preparing a persona mention - `channels.spec.ts:760` — routes an agent mention from an existing DM to the expanded conversation - `channels.spec.ts:815` — routes a relay-agent mention from an existing DM to the expanded conversation - `channels.spec.ts:940` — drops an expanded DM after the first message fails ## Root Cause Race condition: under fast CI execution, mock command completions (create_managed_agent, open_dm) can resolve in non-deterministic order, causing assertions to observe stale or mid-transition state. ## Fix - **:652** — Move the `new-message-recipient-popover` hidden assertion after `chat-title` settles (both names present), so it runs post-transition rather than mid-transition. - **:760, :940** — Add `createManagedAgentDelayMs: 100` to ensure persona provisioning doesn't collapse into the same tick as the expanded-DM open/start sequence. - **:815** — Add `openDmDelayMs: 100` so the two open_dm calls resolve in deterministic order. ## Validation All 4 tests pass with `--repeat-each=3` (12/12 green) locally. Biome lint clean. ## Scope Test-only change: 12 insertions, 1 deletion in `desktop/tests/e2e/channels.spec.ts`. --- Investigated by Ferret, reviewed by Grumplestiltzkin. Signed-off-by: Cameron Hotchkies <chotchkies@block.xyz> Co-authored-by: Goose <opensource@block.xyz>
…block#3271) On some Linux GPU/driver/compositor combinations, WebKitGTK's dmabuf renderer aborts the web process during startup, so Buzz comes up with no window at all and the user has no way to fix it. Setting `WEBKIT_DISABLE_DMABUF_RENDERER=1` avoids the abort by falling back to the shared-memory buffer path. WebKit reads each of its rendering variables exactly once per process, so the choice has to be made before anything initializes — there is no runtime toggle and no second chance later in the same process. This decides up front from two cheap preflight signals rather than reacting to a crash: - **NVIDIA GPU** — any DRM device under `/sys/class/drm` reporting PCI vendor `0x10de`, the driver family behind most upstream reports. - **AppImage** — the `APPIMAGE` environment variable. linuxdeploy's AppRun hook pins `GDK_BACKEND=x11`, and the dmabuf renderer buys nothing on that XWayland path. Either signal disables the dmabuf renderer. Neither signal leaves the environment untouched. ## Escape hatches `--safe-rendering` forces the safest configuration for one launch — `WEBKIT_DISABLE_DMABUF_RENDERER` plus `WEBKIT_DISABLE_COMPOSITING_MODE` — for a machine neither signal recognises. Any user assignment of a variable this module may set stands the heuristic down **wholesale**. Presence is the test, not truthiness, so `VAR=0` and `VAR=` both count: a user asking for the dmabuf renderer *on* gets it, even on a machine the heuristic would have opted out. `--safe-rendering` against such an assignment is refused with a diagnostic naming both the assignment and the key to unset, and exits non-zero — the flag and the environment are two incompatible answers to one question, and neither is guessed. ## Placement `webkit_rendering::apply()` runs at the top of `fn main()`, before `buzz_lib::run()`. That is the only point where the process is still single threaded with no GTK object alive, which is what makes `std::env::set_var` sound; the module doc and the call site both say so. The whole module is `#[cfg(target_os = "linux")]` — macOS and Windows compile none of it. The decision is a pure function of argv, an injected environment lookup, and an injected DRM root, so all of it is unit-testable without mutating the process environment. Closes block#2338. Upstream: [tauri#9394](tauri-apps/tauri#9394). Same approach and same variable as [clash-verge-rev](https://github.com/clash-verge-rev/clash-verge-rev/blob/main/src-tauri/src/utils/linux/workarounds.rs) `workarounds.rs` and [screenpipe](https://github.com/screenpipe/screenpipe/blob/main/apps/screenpipe-app-tauri/src-tauri/src/linux_webkit_env.rs) `linux_webkit_env.rs`. Signed-off-by: Will Pfleger <pfleger.will@gmail.com>
…lock#3007) Mid-turn steering was reachable only through goose's `_goose/unstable/session/steer`, which requires an `expectedRunId` sourced from `_meta.goose.activeRunId`. claude-agent-acp and codex-acp never emit a run id, so every mid-turn mention to those harnesses bailed at the run-id guard before writing a byte and degraded to cancel + merge, destroying in-flight tool calls. Both adapters ship `_session/steering` (params `{sessionId, prompt}`, result `{outcome}`) and advertise it as `_meta.steering.supported` on the `initialize` response. This adds it as a second steer transport selected at write time, reusing the existing withhold/release, ack routing, and cancel+merge fallback machinery unchanged. ## Transport selection | `active_run_id` | `steering_supported` | Transport | |---|---|---| | `Some(run_id)` | any | `_goose/unstable/session/steer` + `expectedRunId` (unchanged) | | `None` | `true` | `_session/steering` with `{sessionId, prompt}` | | `None` | `false` | ack `ExpectedRunIdMissing`, write nothing (unchanged) | goose keeps priority when both are present — `expectedRunId` is strictly more precise about *which* run is being steered. ## Two load-bearing safety properties **The advertised capability is the only gate — never error-code probing.** codex-acp's `extMethod` answers unrecognized extension methods with a bare `{}`, which is a JSON-RPC *success* rather than `-32601`. Buzz maps a steer success to `queue.remove_event`, so probing an unknown method would silently delete the user's message with no error, no fallback, and no log line. **An `outcome` must be positively recognized.** Only `injected` and `startedNewTurn` count as delivery. Anything else — codex's `failed`, an unknown value, or a missing `outcome` entirely — is `SteerError::OutcomeRejected`, which releases the withheld event and fires the cancel+merge fallback. This makes the silent-loss path above unreachable even if an adapter mis-advertises. `startedNewTurn` acks `Success`, because the message really was delivered and must not be redelivered, but deliberately does **not** renew the read loop's hard deadline: the turn Buzz was awaiting had already settled, and renewing would extend the clock on a finished turn. ## Notes for reviewers - `SteerError::OutcomeRejected` needs no new arm in the `PoolEvent::SteerAck` match — the existing catch-all `Ok(SteerAck::Err(_)) => (true, false, true)` already gives release + fallback, and the two `AgentError` arms above it match that variant specifically, so they do not shadow it. - Comments that described the old goose-only "try-and-tolerate" `-32601` behavior are corrected; that assumption was never valid for codex-acp. - No CI job runs `buzz-acp` tests. The full package suite was run locally: **617 passing, 0 failing**. Signed-off-by: Will Pfleger <pfleger.will@gmail.com> Co-authored-by: npub1mn7jgtj4w2pd0g0zeuhxsa6jy6p0rewxz4kujt98my82ahfmp72sxjexk7 <dcfd242e557282d7a1e2cf2e6877522682f1e5c6156dc92ca7d90eaedd3b0f95@buzz.block.builderlab.xyz>
## Why Publish chart 0.1.7 after the feature PR merged from a fork and therefore intentionally skipped the internal-branch auto-tag job. ## What - Trigger the `chart-release/0.1.7` release lane - Update the quickstart example to reference chart 0.1.7 ## Risk Assessment Low — the chart implementation is already merged and tested; this PR creates its immutable release tag and OCI artifact. ## References - Chart implementation: block#3322 - `helm unittest` 0.8.2: 43/43 tests passed - Local pre-push checks passed Generated with Amp Signed-off-by: David Grochowski <dgrochowski@squareup.com> Co-authored-by: Amp <amp@ampcode.com>
Three main-branch runs today had shards killed at exactly 20m17s
("exceeded the maximum execution time of 20m0s"); the killed shard was
actively passing tests seconds before the cap. Shard runtime has grown
to the limit. 30 matches the other desktop jobs in the same workflow.
Signed-off-by: Will Pfleger <pfleger.will@gmail.com>
## Summary - replace the whole-tree file-size gate with a stateless differential ratchet - allow inherited files over 1,000 lines to hold or shrink, but never grow - delete the 44-entry numeric override ledger and run the same policy across Desktop, Web, and Mobile CI - fail closed when the local base cannot be resolved and cover policy, Git status parsing, and base resolution in unit tests This removes the shared mutable policy state that caused unrelated PRs to fail after neighboring merges. It does **not** by itself prevent two stale green PRs from becoming invalid when combined; that requires merge queue or up-to-date branch enforcement. ### Related issue None found. This follows the design discussion in the linked Buzz channel. ### Testing - `node --test scripts/check-file-sizes-core.test.mjs` (6/6) - Desktop, Web, and Mobile ratchet entrypoints - `just desktop-check` - `just web-check` - Mobile analysis - `git diff --check` The repository pre-push suite also exposed an unrelated existing Mobile widget failure in `ChannelDetailPage keeps follow mode off while a tall newest message stays visible`; it reproduces in isolation and this branch does not touch Mobile widget behavior. Signed-off-by: Wes <wesbillman@users.noreply.github.com> Co-authored-by: Carl <c7ebe626f000404285d3686e1dc74cc07cc60a9754a150041ba132e14bd3e2ec@buzz.block.builderlab.xyz>
## Summary - reconcile anchored-scroll state when passive layout changes put a thread at its physical floor - route thread composer-padding growth and shrink through the same hook-owned settlement path - preserve pinned thread targets while clearing stale new-message state ## Root cause Thread bottom state was updated primarily by native `scroll` events. Deferred replies, viewport changes, and composer-overlay padding can finish changing geometry after the user's last scroll—or after the initial open pin—without another scroll event. The thread could visibly reach the floor while `isAtBottom` and `newMessageCount` remained stale, leaving the “N new messages” pill visible. ## Verification - `pnpm check` - `pnpm typecheck` - `pnpm test` — 3,768 passed - push hook: branch-skew, Desktop check, and Desktop full unit suite passed Signed-off-by: Wes <wesbillman@users.noreply.github.com>
Fixes block#2787 - Added `claude-opus-5` to `config.rs` model classification and adaptive effort helpers. - Updated fixture test configurations to cover `claude-opus-5`. - Verified with `cargo test` and JS unit tests. Signed-off-by: Apurva Shaw <apurvashaw@Apurvas-MacBook-Air.local> Co-authored-by: Apurva Shaw <apurvashaw@Apurvas-MacBook-Air.local>
## Summary - drop the `subs` DashMap guard before mutating subscription indexes - snapshot fan-out candidate vectors so index guards are dropped before looking up `subs` - add concurrent fan-out/replacement regression coverage ## Why `fan_out_scoped` previously held an index guard while `push_match` acquired `subs`, while CLOSE and same-ID replacement held `subs` while removing from an index. The reverse ordering made an AB/BA deadlock reachable and could synchronously park all Tokio workers. ## Validation - `rustup run 1.95.0 cargo test -p buzz-relay` — 769 library tests passed, 33 ignored; 11 binary tests passed; doc tests passed - push hooks with pinned Rust 1.95 — branch-skew, repository Rust suites, and desktop Tauri suite passed - `git diff --check` ## Residual risk Fan-out now clones bounded candidate vectors before matching. This adds allocation/copy cost proportional to the indexed candidate set, in exchange for eliminating nested DashMap guards. This fixes the concrete lock cycle but does not prove every observed production wedge had this cause. --------- Signed-off-by: npub12gtutshhh76rx0jx697f32f9tffd4hhp3hx58fp4x6u4uemkm7sqf8f757 <5217c5c2f7bfb4333e46d17c98a9255a52dadee18dcd43a43536b95e6776dfa0@buzz.block.builderlab.xyz> Signed-off-by: npub1jh9wn95s0472h86ahapupaf7m6kx4v9sx2n0atj2hltcfer8k06s5n3pyf <95cae996907d7cab9f5dbf43c0f53edeac6ab0b032a6feae4abfd784e467b3f5@buzz.block.builderlab.xyz> Co-authored-by: npub12gtutshhh76rx0jx697f32f9tffd4hhp3hx58fp4x6u4uemkm7sqf8f757 <5217c5c2f7bfb4333e46d17c98a9255a52dadee18dcd43a43536b95e6776dfa0@buzz.block.builderlab.xyz> Co-authored-by: npub1jh9wn95s0472h86ahapupaf7m6kx4v9sx2n0atj2hltcfer8k06s5n3pyf <95cae996907d7cab9f5dbf43c0f53edeac6ab0b032a6feae4abfd784e467b3f5@buzz.block.builderlab.xyz>
…figured MCP startup (block#3420) ## Summary - add a generic per-runtime env-defaults table, `config::default_agent_env()`, mirroring the existing `default_agent_args()` / `codex_network_env()` precedent, and merge it once in `AcpClient::spawn` with the established precedence: **runtime defaults < persona `extra_env` < inherited parent env** - first (and only) row: Buzz-owned Hermes processes get `HERMES_ACP_SKIP_CONFIGURED_MCP=1`, so Hermes does not preload unrelated profile-configured MCP servers before answering ACP `initialize` (fixes the 10s model-discovery timeout in block#3355 — Buzz supplies session MCP servers explicitly through `session/new`, per Hermes's documented host-integration contract for this variable) - normalize Windows `.cmd`/`.bat` shims alongside `.exe` in `normalize_agent_command_identity` (npm installs resolve to those wrappers) - switch the `extra_env` parent-presence check from `var()` to `var_os()` so non-UTF-8 parent values are honored Replaces the runtime-specific approach in block#3386: same behavior, but the mechanism is generic runtime spawn metadata in `config.rs` rather than a Hermes/ACP special case in `acp.rs`, and the seam covers every launch path (Desktop spawn, `buzz-acp models`, CLI) because they all funnel through `AcpClient::spawn`. ~15 lines of production code. Fixes block#3355 ## Testing - `cargo test -p buzz-acp` — **639 passed, 0 failed** (full package, includes the new `default_agent_env_recognizes_hermes_identities` unit test and `spawn_applies_runtime_env_defaults_with_extra_env_precedence` integration test covering default injection, extra_env override, and non-Hermes exclusion) - `cargo fmt --all -- --check`, `cargo clippy -p buzz-acp --all-targets -- -D warnings` — clean - live-local with real Hermes v0.19.0 (`hermes-acp`): `buzz-acp models` returned **13 models / currentModelId in 2.6–3.0s** (was a 10.0s timeout on the first cold run without isolation); a wrapper probe confirmed the child received `HERMES_ACP_SKIP_CONFIGURED_MCP=1` by default and `0` when the parent env set it explicitly (operator wins) - lefthook pre-push suite green: rust-tests, desktop-check, desktop-test, desktop-tauri-test, mobile-test, branch-skew No UI changes; subprocess environment behavior only. Signed-off-by: npub1qyvc0c5kl4gqv2fd97fsk46tu378sqgy35vc83rvgfwne90sel7s0ed67d <011987e296fd5006292d2f930b574be47c7801048d1983c46c425d3c95f0cffd@buzz.block.builderlab.xyz> Signed-off-by: Tyler Longwell <tlongwell@block.xyz> Co-authored-by: npub1qyvc0c5kl4gqv2fd97fsk46tu378sqgy35vc83rvgfwne90sel7s0ed67d <011987e296fd5006292d2f930b574be47c7801048d1983c46c425d3c95f0cffd@buzz.block.builderlab.xyz> Co-authored-by: Tyler Longwell <tlongwell@block.xyz> Co-authored-by: mr-r0b0t.eth <adam.manning@pro-serveinc.com>
## Summary - align mobile message metadata and enlarge attachment-menu content - smooth keyboard-to-camera/photo transitions and initialize the iOS photo grid at the intended scale - fix horizontal gallery loading, edge overflow, and end spacing ## Why The attachment surfaces were reacting to keyboard and compact-menu geometry during presentation, while gallery clipping and image lifecycle behavior caused misalignment and occasional blank previews. ## Testing - `just mobile-check` - `flutter test` (881 passed, 1 skipped) - native `RunnerTests` (17 passed) - verified standalone Release build on a physical iPhone --------- Signed-off-by: kenny lopez <klopez4212@gmail.com>
…y/TLS passthrough) (block#3463) > On 8 tasks matched by name across the two runs, cost fell $8.36 → $1.77 (4.71×) and wall-clock 12,423 s → 1,085 s (11.45×). ## Summary Two independent, self-contained fixes to `buzz-agent`/`buzz-acp`, split out of the benchmark branch so they can land while the harness work continues: 1. **Request and surface Anthropic prompt caching.** buzz never sent a `cache_control` breakpoint, so on the Databricks Anthropic route `cache_read_input_tokens` was **structurally always 0** and the ~10× cache-read discount was never claimed. This teaches `anthropic_body()` to mark the cacheable prefix, and plumbs the cache split end-to-end so accounting can price it. 2. **Pass proxy + TLS-trust env into MCP tool subprocesses**, so agent tools on a proxy-only host stop reporting a live network as offline. ## Why the caching gap matters The Anthropic Messages API does **not** cache unless the request carries a `cache_control` breakpoint, and the Databricks AI Gateway — a third-party proxy in front of the model, in the same category as Bedrock/Vertex — does **not** auto-cache (only the first-party Anthropic API and Claude-on-AWS do zero-config caching). So every request was billed cold. Measured live against the Databricks gateway (`databricks-claude-opus-5`, 2026-07-28), the same call with and without a single `cache_control` marker: | Run | `input_tokens` | `cache_creation` | `cache_read` | latency | |---|---|---|---|---| | No `cache_control`, two byte-identical calls | 121,625 | 0 | **0** | ~9.3 s | | With one marker — cold (write) | 4 | 121,625 | 0 | 9.3 s | | With one marker — warm (read) | 4 | 0 | **121,625** | **4.5 s** | One marker moved 121,625 tokens from full-price input to a 0.1× cache read and roughly halved latency (a clean, isolated ~2.07× prefill speedup on this single-threaded microbenchmark). The gateway honours `cache_control`; buzz simply never sent it. At fleet scale this was a real budget item. Across matched Terminal-Bench solo sweeps (89 tasks, `-n 20`, before the fix), the two OpenAI-route models independently landed at ~86–87% cache reads — the expected shape for an agentic loop, where system + tools + append-only history repeat every turn — while the Anthropic route returned a hard 0% on every receipt: | Condition | Route | Input tokens | Cache reads | Cost | Cost if uncached | Discount | |---|---|---|---|---|---|---| | luna (`gpt-5-6`) | OpenAI | 20,320,818 | **17.7M (87.0%)** | $6.96 | $22.87 | **3.28×** | | sol (`gpt-5-6`) | OpenAI | 22,312,290 | **19.2M (85.9%)** | $37.07 | $123.35 | **3.33×** | | opus (`claude-opus-5`) | Anthropic | 12,459,822 | **0 (0.0%)** | $81.31 | $81.31 | **1.00×** | Applying luna's measured 87% read rate to the opus token counts at list prices (`input $5/M`, `cached_input $0.5/M`, `output $25/M`) puts the opus run at **~$32.53 vs the $81.31 actually paid — a ~60% overspend on those 49 trials (~$89 on a full sweep)**. That is an upper bound (it prices every cached token at the 0.1× read rate and ignores the 1.25× write premium), and the opus discount is structurally smaller than luna/sol's because opus emits ~3.5× more uncacheable output per trial, which sets a floor on what caching can recover. There is also a plausible **second-order effect**: Databricks appears to meter its per-minute rate limit on *uncached* input tokens, so the missing cache also cost rate-limit headroom — the opus endpoint lost 63% of its trials to fatal 429s while running alone at one-third of a GPT endpoint's raw throughput. This is a hypothesis, not a proven mechanism (the only zero-cache condition is also the only Anthropic endpoint), but it is the reading that explains the throttling with one rule instead of two. ## Post-fix results (provisional — first trials of an in-flight re-run) On 8 tasks matched by name across the two runs, cost fell **$8.36 → $1.77 (4.71×)** and wall-clock **12,423 s → 1,085 s (11.45×)**. | Metric | before (`4a955a858`) | after (`3bef1f6a`) | |---|---|---| | Cache reads as % of input | **0.0%** | **78.7%** (still climbing toward the ~86% steady state) | | `cost_usd_no_cache_discount / cost_usd` | **1.00×** | **2.18×** (tracking the projected ~2.5×) | | Trials with a fatal 429 (same `-n 20`) | **63%** | **15–19%** | To be clear about attribution: **~2× of that is the clean prefill saving from caching itself**; the rest is second-order — cached requests burn far less rate-limit budget, so they stall less and redo less destroyed work. The 11.45× is a system-level result specific to this throttled workspace, not a caching benchmark. Quality held (7/8 solved in each run). A controlled low-`-n` A/B (neither arm hitting a 429), which the `BUZZ_AGENT_PROMPT_CACHING` opt-out exists to enable, is still owed before this becomes a published claim. ## What changed ### 1. Request caching (`llm.rs`, `config.rs`) `anthropic_body()` emits ephemeral `cache_control` breakpoints, gated by `BUZZ_AGENT_PROMPT_CACHING` (**default on**, `=0` to opt out): - **Static prefix** — marker on the `system` block. Prefix order is `tools → system → messages`, so this single marker caches **tools + system** together. Byte-identical on every turn of a run, and survives a context handoff (system/tools come from cfg/mcp, not `self.history`). - **Rolling tail + leapfrog** — marker on the last block of the last **two** messages. The append-only history re-reads the prior turn's prefix from cache; marking two messages (not one) keeps consecutive breakpoints inside Anthropic's **20-block lookback window** even as tool parallelism rises, avoiding a silent full-price miss. An empty system prompt stays a bare string (Anthropic rejects empty text blocks), and below-threshold prefixes are silently not cached, so the flag is safe on by default. ### 2. Surface the cache split end-to-end — the plumbing (`types.rs`, `llm.rs`, `agent.rs`, `lib.rs`, `usage.rs`, `acp.rs`) This is the part that makes gaps like the one above **visible** instead of silent. A consumer that prices all of `input_tokens` at the full rate can't tell a route that's caching from one that isn't — the total looks right either way. So: - `LlmResponse` gains `cached_input_tokens` (a **subset** of `input_tokens`, never an addition); `parse_anthropic` / `parse_openai` / `parse_responses` each populate it. - A `usage_first()` helper reads the cache count wherever a provider hides it — flat `cache_read_input_tokens` (Anthropic), `prompt_tokens_details.cached_tokens` (OpenAI chat), `input_tokens_details.cached_tokens` (Responses) — taking the **first present value, never a sum**. Reading only flat keys is exactly why the OpenAI route's nested `cached_tokens` had *also* been going unclaimed: `prompt_tokens` is already inclusive, so the total looked correct while the discount silently went unreported. - The per-turn/per-session accumulators and the goose `usage_update` payload now carry `accumulatedCachedInputTokens`; `buzz-acp` deserializes it (`serde` default `0` for goose, which doesn't send it) and logs `cached=<n>`. ### 3. Fix a Databricks MLflow-route double-count (`llm.rs`) The Databricks MLflow route reports the flat Anthropic-spelled `cache_read_input_tokens` *alongside* an already-inclusive `prompt_tokens`, so the old code summed them and nearly doubled the count — inflating both the context-budget gate and cost. `openai_chat_input_tokens()` now reads `prompt_tokens` alone. Verified on a live `databricks-glm-5-2` response where `prompt_tokens + completion == total` proves inclusivity. (Anthropic's native route genuinely *excludes* the cache fields and is still summed — the two never collide, because `claude*` models route to the Anthropic path.) ### 4. Proxy + TLS-trust passthrough into MCP tools (`mcp.rs`) — independent fix `buzz-agent` `env_clear()`s each MCP child, and the allowlist carried no proxy/TLS vars. On a proxy-only host that doesn't degrade the tools, it **blinds** them: apt, curl, pip, git connect directly, the egress firewall resets the socket, and the agent reports "Connection reset by peer" — indistinguishable from a genuinely offline task. Adds both spellings of `HTTP(S)_PROXY`/`NO_PROXY`/`ALL_PROXY` (curl/git read lowercase; Go/Python read uppercase; libcurl ignores uppercase `HTTP_PROXY`) plus `SSL_CERT_FILE`/`SSL_CERT_DIR` for TLS-terminating proxies that present their own CA. ## Testing - `cargo fmt --all -- --check`, `cargo clippy -p buzz-agent -p buzz-acp --all-targets -- -D warnings` — clean. - `cargo test -p buzz-agent -p buzz-acp` — **all green** (632 + 299 lib tests plus integration suites, 0 failures). New tests cover: the three breakpoints and the disabled/empty-system/single-message edge cases; nested-vs-flat cache parsing for all three routes; the Databricks inclusive-`prompt_tokens` fix; wire deserialization of `accumulatedCachedInputTokens`; and the proxy/TLS passthrough allowlist. - Pre-push lefthook suite green (branch-skew, rust-tests, test, desktop-check/test/tauri). ## Relationship to the benchmark branch These are the non-`benchmarks/` changes from `benchmark/harness-accounting-and-solo`, lifted onto a clean base off `main` so they can merge independently. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Signed-off-by: Atish Patel <atish@squareup.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
## Summary - Present channel and thread typing status in a composer-matched container. - Animate the strip so the message list moves smoothly as typing begins and ends. - Increase typing-label contrast and avatar/padding for readability. ## Pixel 10 snapshot  ## Validation - `flutter test test/features/channels/channel_detail_page_test.dart` - `flutter analyze` Signed-off-by: kenny lopez <klopez4212@gmail.com>
## Summary - Simplify the community invite dialog around link sharing. - Add matching expiry and use-limit dropdowns, with sensible preset use caps. - Cover the default unlimited and selected-limit invite payloads. ## Validation - `pnpm -C desktop run build:e2e` - `pnpm -C desktop exec playwright test tests/e2e/invite-link-copy.spec.ts tests/e2e/invites-settings-screenshots.spec.ts --project=smoke` Signed-off-by: kenny lopez <klopez4212@gmail.com>
…wire (block#3538) ## Summary Databricks v2 chooses the gateway wire format — OpenAI Responses, Anthropic Messages, or MLflow chat — purely from substrings in the endpoint name. There is no family field on the endpoint to key off, so the substring set *is* the routing contract. The matcher only recognised `gpt-5`/`gpt5` and `claude`, which makes correct billing depend on every Claude endpoint happening to be named with the literal string "claude". ## Why this matters Getting a Claude model onto the Anthropic Messages route is exactly what lets buzz attach the `cache_control` breakpoint (the fix in block#3463). If a Claude endpoint's catalog name omits "claude" — an alias, a bare `opus-5`, a `goose-opus-5` — it silently falls through to the MLflow (OpenAI-wire) path, where Anthropic prompt caching is **structurally impossible**. The result is the same failure block#3463 fixed: 0% cache reads, the full ~10x read discount lost, and no error — a naming convention quietly holding up a billing-correctness invariant. ## What changed `databricks_v2_route_for_model` (`crates/buzz-agent/src/llm.rs`) now matches broader, case-insensitive marker sets: - **Claude → Anthropic Messages:** `claude`, `opus`, `sonnet`, `haiku`, `mythos`, `fable` — the Claude family names and release code names, so a Claude endpoint reaches the cache-capable route regardless of how it's named. - **GPT → OpenAI Responses:** the `gpt` family (now `gpt` on its own, not just `gpt-5`) plus the GPT-5 launch code names `sol`, `luna`, `terra`. OpenAI markers are evaluated first, preserving the prior `gpt-5`-first precedence for any name that could carry both. Names matching neither set still fall through to the MLflow chat route. ## Testing - `cargo fmt`, `cargo clippy -p buzz-agent --all-targets -- -D warnings` — clean. - `cargo test -p buzz-agent` — all green (299 lib + integration suites, 0 failures). The `databricks_v2_routes_by_model_family` test was expanded to cover each new marker, the GPT-5 code names, case-insensitivity, and the unchanged MLflow fallback (including `gemini`). ## Relationship to block#3463 block#3463 taught the Anthropic path to request caching; this makes sure Claude models actually land on that path. Follow-up still open: surfacing `cache_creation_input_tokens` end-to-end so a persistent `reads == 0 && writes == 0` reveals a disabled cache regardless of which wire a model takes — happy to do that next. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Signed-off-by: Atish Patel <atish@squareup.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
**Category:** fix **User Impact:** Custom emoji with valid 64-character names can now be used as reactions without errors. **Problem:** Buzz accepted 64-character custom emoji names during registration, but rejected them as reactions after the required surrounding colons made the payload 66 characters. Validation also differed between desktop, SDK, relay, and storage boundaries. <img width="554" height="47" alt="image" src="https://github.com/user-attachments/assets/4013452f-210e-4dd3-9003-f45ff3b28dc8" /> **Solution:** Keep the product limit at 64 ASCII characters for custom emoji names, enforce it consistently when emoji sets are registered, and allow only valid matching custom reaction payloads up to 66 characters. Widen the reaction projection to preserve the wrapped payload while retaining the existing 64-character limit for ordinary reactions. <details> <summary>File changes</summary> **crates/buzz-sdk/src/builders.rs** Defines the shared custom emoji boundaries and covers accepted 64-character and rejected 65-character shortcodes. **crates/buzz-relay/src/handlers/ingest.rs** Validates emoji-set shortcodes and permits 66-character reactions only when they are valid colon-wrapped custom emoji with a matching tag. **crates/buzz-db/src/event.rs** Adds storage regression coverage for maximum-length custom emoji reactions. **crates/buzz-db/src/migration.rs** Verifies the reaction column migration is applied correctly. **desktop/src/shared/api/customEmoji.ts** Enforces the existing 64-character shortcode maximum during desktop normalization and registration/import. **desktop/src/shared/api/customEmoji.test.mjs** Covers the desktop shortcode boundary. **migrations/0027_long_reaction_payloads.sql** Widens stored reaction payloads to 66 characters for the two required surrounding colons. **schema/schema.sql** Keeps the desired schema aligned with the migration. </details> ## Reproduction Steps 1. Register or import a custom emoji whose ASCII shortcode is exactly 64 characters. 2. Select that emoji as a reaction to a message. 3. Confirm the reaction publishes, persists, and renders without an error. 4. Attempt to register a 65-character shortcode and confirm it is rejected. 5. Publish an ordinary or malformed reaction over 64 characters and confirm the relay rejects it. ## Verification - `cargo test -p buzz-sdk`: 243 passed - `cargo test -p buzz-db`: 94 passed, 152 Postgres-required tests ignored - `pnpm test` in `desktop`: 3,859 passed - `cargo test -p buzz-relay`: 795 passed, 9 existing Postgres-unavailable failures, 35 ignored; new reaction boundary tests pass directly - `cargo fmt --all -- --check` - `git diff --check` Originating Buzz channel: `f2ec9671-d78e-4cde-894c-9f4c458c7f1f` --------- Signed-off-by: Taylor Ho <taylorkmho@gmail.com> Co-authored-by: npub1223z34hd7vtwc6qj4s7flsxkj644nlre2nthu7lrrmkumhu3xddsrx9r6w <52a228d6edf316ec6812ac3c9fc0d696ab59fc7954d77e7be31eedcddf91335b@buzz.block.builderlab.xyz>
… 'Attach file' (block#2381) (block#4304) Fixes block#2381. ## What was broken The message composer's paperclip accepts generic attachments — images, videos, PDFs, archives, and any other supported file — but its tooltip and accessible name still read **"Attach image"**. Sighted users might reasonably believe the control is image-only, and screen-reader users get an incomplete description of what the button does. ## The fix Rename the accessible name and tooltip text on the generic composer paperclip in `MessageComposerToolbar.tsx`: - `aria-label` — `"Attach image"` → `"Attach file"` - `<TooltipContent>` — `"Attach image"` → `"Attach file"` Plus update the 12 affected Desktop e2e selectors across five spec files to reference the new accessible name: - `desktop/tests/e2e/file-attachment.spec.ts` (2 selectors) - `desktop/tests/e2e/spoiler.spec.ts` (2) - `desktop/tests/e2e/composer-image-draw.spec.ts` (2) - `desktop/tests/e2e/image-attachment-gallery.spec.ts` (4) - `desktop/tests/e2e/video-attachment.spec.ts` (2) ## Scope (per the issue) The feedback screenshot dialog (`desktop/src/features/settings/ui/SendFeedbackDialog.tsx`) is **unchanged** — that dialog itself is image-only, so its "Attach image" wording is accurate. This PR only touches the generic composer control. ## Test plan - All **105** unit tests in `desktop/src/features/messages/ui/*.test.mjs` pass locally. - Verified no remaining `"Attach image"` string outside the intentionally preserved feedback dialog: ```sh grep -rn '"Attach image"' desktop/ # → only hits in SendFeedbackDialog.tsx ``` - The six e2e specs are only exercised in CI; the selector updates are mechanical and verified by grep to reference the new a11y name. ## Blast radius - **Files touched**: `MessageComposerToolbar.tsx` (two strings); five e2e spec files (12 selector updates). - **User-facing behaviour**: one tooltip + one screen-reader name change; no functional or visual changes otherwise. - **No API or state change.** ## Out of scope - The feedback dialog's "Attach image" wording — kept per the issue's own "Scope" guidance. - Any i18n plumbing — Buzz Desktop doesn't currently localize these strings. Signed-off-by: Sarthak Singh <sarthak.singh@juspay.in> Signed-off-by: Ravneet Arora <rarora@squareup.com> Co-authored-by: Ravneet Arora <rarora@squareup.com> Co-authored-by: Cursor <cursoragent@cursor.com>
**Category:** improvement **User Impact:** Message usernames are now bolder, making it easier to distinguish who said what at a glance. **Problem:** Usernames and surrounding message metadata had too little visual separation, which made message headers slower to scan. **Solution:** Increase the shared message-author label from semibold to bold while preserving its existing size, spacing, and interaction behavior. <details> <summary>File changes</summary> **desktop/src/features/messages/ui/MessageHeader.tsx** Raises the shared message-author font weight so standard and system message usernames gain consistent visual contrast. </details> ## Reproduction steps 1. Open a channel containing messages from multiple people or agents. 2. Compare each message username with its timestamp and message body. 3. Confirm the username renders in bold while the surrounding typography and layout remain unchanged. ## Screenshots | Before | After | | --- | --- | |  |  | Signed-off-by: Taylor Ho <taylorkmho@gmail.com>
**Category:** fix **User Impact:** Expanded thread panels now stay fully visible within the desktop channel area instead of being cut off. **Problem:** The resize handler clamped the thread panel against the full window width, even though the panel renders inside a narrower channel surface. On a 1720px window, this allowed a 1160px requested width where only 1111px could render, leaving persisted and visible geometry out of sync. **Solution:** Clamp resizing against the measured channel-surface width so the stored width matches what the layout can render while preserving the minimum 300px main pane. <details> <summary>File changes</summary> **desktop/src/features/channels/ui/ChannelScreen.tsx** Passes the measured channel-surface width into the thread-panel sizing hook. **desktop/src/shared/hooks/useThreadPanelWidth.ts** Clamps drag-resize updates against the available channel width instead of the full viewport. **desktop/tests/e2e/threadpane-ultrawide.spec.ts** Adds a 1720px regression proving the requested and rendered panel widths match, while retaining the ultrawide expansion case. </details> ### Reproduction steps 1. Open a channel thread in the desktop app at a 1720×900 window size. 2. Drag the thread panel's left resize handle toward the left edge to expand it as far as possible. 3. Confirm the panel remains fully bounded inside the channel surface and the main channel pane remains at least 300px wide. 4. Reload the channel and confirm the persisted expanded width renders without clipping. ### Testing - `pnpm --dir desktop build:e2e` - `pnpm --dir desktop exec playwright test tests/e2e/threadpane-ultrawide.spec.ts` — 2 passed - Push hooks: `desktop-check` and `desktop-test` passed - `git diff --check origin/main..HEAD` ### Screenshot  ### Related issue None found. Signed-off-by: Taylor Ho <taylorkmho@gmail.com>
**Category:** improvement **User Impact:** Selected communities now use a clear offset outline without tinting or covering their icon. **Problem:** The selected community state replaced the icon surface with an accent fill, obscuring image icons and changing the tile's content treatment. Hover also changed the fill, text color, shape, and opacity, making navigation states visually jumpy. **Solution:** Preserve each community tile's neutral surface and content while using a primary CSS outline for selection and a lighter outline for hover. The transparent outline offset leaves the space around image edges unpainted, and adjusted spacing prevents neighboring outlines from colliding. <img width="200" height="152" alt="Screen Recording 2026-08-05 at 3 23 32 PM" src="https://github.com/user-attachments/assets/5c25b1c0-4be8-41c4-8f1d-ad0010310c92" /> <details> <summary>File changes</summary> **desktop/src/features/sidebar/ui/CommunityRail.tsx** Replaces selected and hover fills with offset outlines, keeps icon presentation stable across states, and adjusts rail and tooltip spacing for the new outline geometry. **desktop/tests/e2e/community-rail.spec.ts** Covers the shared active/inactive surface, radius, text color, opacity, and outline behavior, including hover invariants. </details> ## Reproduction steps 1. Run the desktop app with two or more communities. 2. Give the active community an image icon. 3. Confirm the active icon keeps its original image and receives a 2px primary outline with a transparent 2px gap. 4. Hover another community and confirm only a lighter outline appears; its fill, text color, opacity, and corner radius remain unchanged. 5. Switch communities and confirm the outline follows the active community. ## Screenshots **Full desktop context**  --------- Signed-off-by: Taylor Ho <taylorkmho@gmail.com>
…ZZ_DRAIN_JITTER_MS) (block#4542) ## Problem On SIGTERM the relay sends every live WebSocket a **1012 Service Restart** close frame via `ConnectionManager::drain_all()` — all in the same instant (`main.rs` shutdown task → `state.rs::drain_all`). On a pod holding thousands of sessions, that makes every client reconnect simultaneously: the thundering-herd reconnect behind the DB pool-timeout bursts observed on each rolling deploy. Client-side jitter can't fix this — the desktop client *resets* its backoff to base on a 1012 and reconnects with only ±25% jitter (`relayClientSession.ts`), so the spread has to come from the server. ## Change Add `BUZZ_DRAIN_JITTER_MS` (default `0` = unchanged behavior). The two paths are kept **deliberately separate** so the default is byte-for-byte the previously shipped shutdown: - **Jitter off (`0`/unset, the default):** the original synchronous, all-at-once `drain_all()` runs unchanged — queue the 1012 on each connection's control channel, cancel, return. No new machinery on the default path. - **Jitter on (`> 0`):** a separate async `drain_all_jittered(jitter_ms)` spreads each connection's restart close over an independent uniform delay in **`[1, jitter_ms]`**. Each delayed close travels a dedicated `RestartClose` channel; the writer flushes the 1012 frame and **acknowledges the flush over a oneshot**, so drain waits for confirmed delivery (up to `RESTART_CLOSE_ACK_TIMEOUT` = 5s) rather than assuming it, falling back to cancellation if the channel is full/closed or the ack times out. The drain future is **owned and awaited** by the shutdown task, and the 30s hard-drain backstop is aborted only after a clean drain — so a clean roll exits `0`. The two methods can be unified and the old one dropped later once the jittered path is proven for all cases. - **`config.rs`** — `drain_jitter_ms`: non-negative parse, clamped to `MAX_DRAIN_JITTER_MS` = **20s** (leaving 10s of the 30s budget for flush). Junk fails loudly at startup; **empty/whitespace-only is treated as unset (jitter off)** so a `BUZZ_DRAIN_JITTER_MS=""` kill switch does not crashloop the relay (matches the sibling env vars in this file). - **`state.rs`** — `drain_all()` (unchanged synchronous default) + `drain_all_jittered()` (jittered + flush-ack). Both set the sticky `draining` flag before the first await. A registration that lands mid-shutdown always self-signals via the **immediate** control-frame + cancel path — jitter smears already-established sockets, not late arrivals. - **`main.rs`** — shutdown task dispatches: `drain_jitter_ms == 0` → `drain_all()`, else `drain_all_jittered(...).await`. ## Safety - **Default off is the currently-committed path.** With jitter unset/0 the shutdown runs the original synchronous `drain_all()` — no restart channel, no ack wait. Safe to deploy dark and dial up. - **Shutdown-boundary race preserved.** Sticky flag set before any await; a late registration self-signals its close with no jitter. - **Owned + backstopped.** The jittered drain future is awaited; the 30s hard-drain `process::exit(1)` remains the ceiling. `MAX_DRAIN_JITTER_MS` (20s) + `RESTART_CLOSE_ACK_TIMEOUT` (5s) = 25s, inside the 30s budget; 5s pre-sleep + 25s = 30s against `terminationGracePeriodSeconds: 60`. ## Known behavior to note (not a blocker, flagged from review) On a **successful** flush the jittered path deliberately does not cancel the connection token — teardown then depends on the client echoing our Close, or on process exit. Compliant clients echo; a silent client rides to the 30s hard exit. The default (jitter-off) path cancels deterministically as before. ## Tests - `config::tests::drain_jitter_defaults_off_and_rejects_junk` — default off, `20000`, clamp `60000`→`20000`, explicit `0`, junk `"soon"` fails, **empty `""` and whitespace-only treated as off**. - `state::tests::drain_all_is_immediate` — default path queues frame + cancels synchronously. - `state::tests::drain_all_sends_restart_close_and_cancels_every_conn`, `drain_all_full_control_buffer_still_cancels`, `register_after_drain_self_signals_restart_close_and_cancel`. - `state::tests::drain_all_jittered_defers_close_until_within_jitter_window` (paused time). - `state::tests::drain_all_jittered_waits_for_writer_acknowledgement_without_cancelling`. - `state::tests::drain_all_jittered_cancels_when_restart_channel_is_full_or_closed`. - `state::tests::drain_all_jittered_cancels_when_flush_ack_times_out` (paused time — the 5s ack-timeout fallback). Validation at `46c690940`: `cargo fmt -p buzz-relay --check`, `cargo clippy -p buzz-relay --all-targets -- -D warnings`, and the drain/config unit suite all clean. Local live SIGTERM test with a real relay process + 200 NIP-42-authenticated sockets — see the PR comment for the before/after distribution and exit codes. ## Rollout Ship with default `0`, then set `BUZZ_DRAIN_JITTER_MS` (e.g. 10000–20000) on bb-block first, watch the roll-window pool-timeout metric, then bb-public. `""` is a safe kill switch. Complements the preStop `sleep` (stops routing before close). --------- Signed-off-by: npub1srl70fhzyu3fsnahl06vw2czvqc2w3ds37hyzvjnk8ve8f03ngcqg9le2w <80ffe7a6e22722984fb7fbf4c72b026030a745b08fae413253b1d993a5f19a30@buzz.block.builderlab.xyz> Signed-off-by: npub128x7j3pwgm4vs8yra3c42fcgcwcvh94g3luwzkqa376du2q6l0esqcrwch <51cde9442e46eac81c83ec71552708c3b0cb96a88ff8e1581d8fb4de281afbf3@buzz.block.builderlab.xyz> Signed-off-by: Brad Seiler <seiler@squareup.com> Co-authored-by: npub1srl70fhzyu3fsnahl06vw2czvqc2w3ds37hyzvjnk8ve8f03ngcqg9le2w <80ffe7a6e22722984fb7fbf4c72b026030a745b08fae413253b1d993a5f19a30@buzz.block.builderlab.xyz> Co-authored-by: npub128x7j3pwgm4vs8yra3c42fcgcwcvh94g3luwzkqa376du2q6l0esqcrwch <51cde9442e46eac81c83ec71552708c3b0cb96a88ff8e1581d8fb4de281afbf3@buzz.block.builderlab.xyz>
### What changed? Inbox detail now gives the current user's messages the same ownership-gated Edit action as channel view. Editing reuses the existing composer and mutation flow, preserves attachment metadata, and refreshes structural overlays so the edited content appears immediately. Foreign authors' messages remain non-editable, including grouped Inbox conversations whose selected event is not the representative item. | Own Inbox message exposes **Edit message**. | Saving the edit updates the Inbox detail immediately. | | --- | --- | |  |  | ### Why? Inbox rows did not pass an edit handler into the shared message action bar, so a user's own messages could be edited from channel view but not from Inbox detail. ### How is it tested? Desktop checks, unit tests, and the full local CI gate passed. The focused Inbox Playwright regression passed 3 consecutive runs and covers current-user edit/save, foreign and archived-channel denial, and attachment preservation when a just-sent reply is edited before its relay echo arrives. Added tests: - [`inbox-edit.spec.ts`](https://github.com/block/buzz/blob/inbox-message-edit-action/desktop/tests/e2e/inbox-edit.spec.ts) - [`inboxViewHelpers.test.mjs`](https://github.com/block/buzz/blob/inbox-message-edit-action/desktop/src/features/home/lib/inboxViewHelpers.test.mjs) *🤖 This PR was authored [with an agent](buzz://message?channel=7f2d7e02-f4d5-4fb0-a426-0ca60ed3a1c3&id=c09ee04d18399b90296c3f932d22ab0377fa05f7e690ec7b08c36483ee633fbb).* --------- Signed-off-by: Tom Brow <tomb@block.xyz> Signed-off-by: npub1ft62tztwwm2x9xamk25smmuaj4sfckdkldksruf2x2jwqalffkrq0g7arr <4af4a5896e76d4629bbbb2a90def9d95609c59b6fb6d01f12a32a4e077e94d86@sprout-oss.stage.blox.sqprod.co> Co-authored-by: npub1ft62tztwwm2x9xamk25smmuaj4sfckdkldksruf2x2jwqalffkrq0g7arr <4af4a5896e76d4629bbbb2a90def9d95609c59b6fb6d01f12a32a4e077e94d86@sprout-oss.stage.blox.sqprod.co>
…lock#4633) ## Summary - Keep mobile thread reply badges current by merging relay recounts with replies observed locally. - Retain replies in the local channel store while continuing to filter them from the main timeline. - Match the thread summary behavior already used on desktop, including the reply count, latest reply time, and participant avatars. ## Why On mobile, the "N replies" badge under a channel message can stall at a stale count or remain missing after a reply arrives. This makes the badge unreliable and can cause people to miss replies. The badge has two inputs: best-effort recounts from the relay and replies the client sees arrive. Mobile previously let any positive relay recount override the local view, while also discarding replies from its local message store. A delayed or lost recount, or a reply received after the recount, could therefore leave the badge behind. This change combines both inputs by using the higher reply count, the later last-reply time, and a merged participant list. Relay timestamps have one-second precision, so equal timestamps do not prove that a recount included a locally observed reply. Comparing counts preserves that reply instead of trusting recency alone. Desktop already uses this merge behavior. ## Validation At commit `4e3356636f5ad62e8f07910af305c532186c6c08` with a clean worktree: - `flutter test` for mobile: 1105 passed, 1 skipped - `flutter analyze` for mobile: no issues found - Reverting the merge so a positive relay recount shadows local replies fails 4 of the new tests, including the same-second and reply-after-recount cases. Restoring the store-level reply drop fails both new provider tests. Added tests: - [`timeline_message_test.dart`](https://github.com/block/buzz/tree/main/mobile/test/features/channels/timeline_message_test.dart), covering relay-only recounts, a reply newer than the recount, a reply in the same second as the recount, a lost recount, a zero recount, nested replies at the root and at the reply they answer, a deleted reply, and participant merging and capping. - [`channel_messages_provider_test.dart`](https://github.com/block/buzz/tree/main/mobile/test/features/channels/channel_messages_provider_test.dart), covering a live reply reaching the store while staying out of the main timeline, and a reply newer than the relay recount raising the badge. --------- Signed-off-by: Tom Brow <tomb@block.xyz> Co-authored-by: npub12uu53ml9upy7ww9apmtv6vm0u8xlcldx7znsjvwgsr7uvy5g0kssw943ca <573948efe5e049e738bd0ed6cd336fe1cdfc7da6f0a70931c880fdc612887da1@buzz.block.builderlab.xyz>
This change enables a Tauri content security policy that limits executable content to the packaged application and does not allow inline scripts. Relay, media, asset, and Tauri IPC schemes remain available for desktop compatibility. The policy contains the impact of a future renderer injection; it does not itself remove an injection bug. ## Testing - `git diff --check origin/main...codex/security-desktop-csp` - Rebased onto `origin/main` at `5c98932` - Full CI pending Originating Buzz thread: `buzz://message?channel=3928fe05-df61-4b5d-b9c7-d623b9b10ea1&id=3c6c02312f763fbe0d2bfc33a6c1a362f91d0354f3d18b039cf7a0558c1439d1` --------- Signed-off-by: Jordan Mecom <jm@squareup.com> Signed-off-by: Eli Foster <efoster@squareup.com> Co-authored-by: Eli Foster <efoster@squareup.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…opes (block#4917) ## Problem Observer telemetry is the noisiest client of the relay: the old pacer (167ms spacing + 90/min rolling cap) let a busy session bill up to 6 events/second against the owner's message quota, and the rolling cap silently *dropped* frames once exceeded. Ruling from the rate-limiting investigation thread (channel `826fc99b-1472-40e7-a529-6b9db8943b8c`): pace at 1/s, always emit, minimal PR. **Review round 1 (Max, Sami)** found the first cut wrong in three ways — tick burst (all pending frames per tick), startup burst (`interval` fires at t=0), and per-channel quota arithmetic. All fixed and mutation-verified in round 1. **Review round 2 (Sami, Max)** found two more against the round-1 head: 1. **Drain-rate collapse (Sami, blocker):** front-run-only packing meant a frame held ONE event whenever channels interleaved — measured 275 B/s vs 63.5 KB/s, so an ordinary 2-channel session fell minutes behind with zero drops and no warning. Silent unbounded latency. 2. **Coalescer byte-cap bypass (Max):** chunks pending in `ObserverChunkCoalescer` were unbounded and outside the 4 MiB cap — 500 distinct-`messageId` 50KB chunks retained ~48MB with `pending_bytes == 0` and zero drops. **Review round 3 (Max)** found the drop accounting undercounted merged chunks: a coalescer entry that merged N same-`messageId` chunks counted as **1** in `dropped_events` when evicted (50 merged 1KB chunks evicted → counter read 1, 49 generated events unaccounted). Fixed: accounting is now denominated in **source (generated) observer events** end to end — each pending entry tracks how many chunks it absorbed, eviction charges that count, and the count survives flush into the publish FIFO. **Review round 4 (Sami, Max)** found three more against the round-3 head: 1. **Coalescer byte undercount (both, independently):** a pending merged entry retains its first chunk's text **twice** until flush — once inside the serialized event skeleton and once in the extracted text accumulator — but was charged only `serialized_len`, so true retention overshot the 4 MiB cap ~2× (measured 8.3 MB). Fixed: `push_pending` charges `serialized_len(&event) + text.len()`. 2. **Cap regressions asserted the accumulator against itself,** which is how the undercount hid. All three cap tests now assert on independently **walked** retained bytes (`serialized_len` per FIFO entry + `serialized_len + text.len()` per coalescer entry), with a secondary `accumulator >= walked` sanity check. Reverting the fix makes them fail at exactly 8,328,386 / 8,328,272 bytes. 3. **FIFO-arm source accounting was implemented but untested (Sami M13; Max reproduced at `cc9333b7c` with 102/151):** the round-3 regression only evicted a merged entry while still in the coalescer. New test forces a merged entry (50×1KB, `source_events=50`) through flush into the publish FIFO, then evicts it from there — mutating the FIFO eviction to `dropped += 1` fails with the reviewers' exact numbers (102 vs 151). **Review round 5 (Sami 9/9/9, Max 9/9/9)** — production judged merge-safe by both; remaining items are tests only, all landed at `63d821620`: 1. **The walker instrument was itself unverified (Sami M17–M20; Max independently confirmed the `return 0` mutant survives):** every cap test asks `walked_retained_bytes()` only for `<= CAP`, so a blinded walker passes everything — and paired with a reverted `push_pending` fix the two mutations cancel, hiding exactly the 8.3 MB overshoot it exists to detect. New two-sided pin: the walker must SEE the first chunk's text twice, and must agree with the accumulator EXACTLY while both stores are non-empty. Kills M17, M18, M19, M20. 2. **Two pre-existing snapshot-clone siblings (Sami D5b/D5c; byte-identical at merge-base `7334ad1e1` — not this PR's regression, but the PR made the class visible):** aliasing the inner turns map leaks a post-save turn into the snapshot; aliasing the inner tombstones map leaks a post-save terminal that blocks a legitimate post-restore resurrection. Two isolation tests with in-test controls — all three inner-map clones in `saveActiveAgentTurnsForCommunity` are now pinned. ## Change **Harness (`crates/buzz-acp`)** - **Global pacer: AT MOST ONE relay frame per second**, regardless of channel count or backlog size. `interval_at(now + 1s)` restores the no-startup-burst property; `MissedTickBehavior::Skip` is now pinned by a paused-time test (a stalled tick arm fires one catch-up frame, not one per missed deadline). At 1 frame/s telemetry spends ≤60/min of the shared 120/min quota; `OBSERVER_PUBLISH_TICK` documents the tradeoff as the knob. - **`ObserverPublishQueue` with gather-packing:** events wait as byte-accounted events (FIFO). `next_frame()` packs the front event's channel **gathered queue-wide in FIFO order** — frames never mix channels, and each channel's events keep their FIFO order, but cross-channel frame order MAY differ from arrival order. That is what keeps the drain rate in **bytes per slot** (one ~64KB frame/s) instead of front-run-length events per slot. **Null-channel events (`agent_panic`-class) are packing barriers** nothing gathers across, so causally-global events keep exact order against every channel. - **One byte cap over BOTH stores:** the event FIFO and the coalescer's pending chunk buffer count against the 4 MiB budget together; eviction is oldest-first across both (queue front, then coalescer front — structural age order) with accounting (warn + counter). A high-cardinality chunk flood is bounded exactly like a plain event flood. Coalescer entries are charged their **true** retention (`serialized_len + text.len()` — the first chunk's text lives in both the serialized skeleton and the extracted accumulator until flush). - **Shutdown is not a burst bypass:** paced one-frame-per-tick drain until empty. **Desktop** - `unwrapObserverBatch` expands envelopes on the live relay path and archive-ingest seam (round 1, unchanged). - **`activeAgentTurnsStore` watermark re-keyed per (agent, channel)** with a dedicated null-channel bucket: the per-agent `(timestamp, seq)` gate would silently skip a delayed channel's frames as stale under gather-packing's intentional cross-channel reorder. Safe because every turn-mutating path is channel-scoped by the event's own `channelId` (endTurn's null-turnId fallback matches `turn.channelId`; resurrectTurn keys on `event.channelId`), so per-channel serialization preserves each guard the per-agent gate provided. The tombstone-cap justification is rewritten for the new keying (worst case for an evicted tombstone is a ghost badge the prune reaps — bounded cosmetic staleness, not corruption). Community-switch save/restore deep-clones the nested map. Other per-agent maps stay agent-keyed: the clock offset is a running minimum (order-insensitive); turns/tombstones mutate only through channel-scoped paths. ## Version skew — old desktop + new harness Gather-packing *intentionally* emits cross-channel-reordered frames. An **old desktop** (per-agent watermark) against a **new harness** will silently skip a delayed channel's turn-state events as stale — working badges on that channel can go stale/missing until its next fresh event. Transcript and archive are unaffected (the transcript store sorts + rebuilds on out-of-order arrival; the archive is per-channel by construction). Ship desktop and harness together; skew degrades badges only, not data at rest. ## Throughput ceiling — "lossless" is qualified Sustained lossless rate is what fits in one ~64KB frame per second, now genuinely in bytes under interleaving: | event payload | events per frame | sustained ceiling | |---|---|---| | 100 B | 250 | 250 ev/s | | 500 B | 99 | 99 ev/s | | 2 KB | 30 | 30 ev/s | | 10 KB | 6 | 6 ev/s | With C channels producing concurrently, publish slots round-robin between them: per-channel drain is ~64KB/C per second and the 4 MiB burst budget (~64s single-channel) shortens accordingly. Beyond budget, oldest-first drops **with accounting** — visible, designed loss. **Accounting semantics:** `dropped_events` counts SOURCE (generated) observer events, not retained entries — evicting a coalesced entry that merged N chunks charges N. On the published side, a merged entry ships all N sources' text in ONE event, so the reconciliation invariant is `ingested == dropped_events + Σ source_events over published events` (for unmerged events, source_events = 1). ## Verification At `63d821620d3513505e8766ac691a8002f9d4a96f` (this head; `git rev-parse HEAD` matched in the same shell as every run), rustc 1.95.0: - `cargo test -p buzz-acp`: **689 lib + 9 integration, 0 failed** — regressions: interleaved 2-channel drain packs into ≤4 frames not 200 slots; null-channel barrier; queue-wide gather with within-channel FIFO; distinct-key 50KB chunk flood bounded by the cap with event-level accounting (published + dropped == ingested, survivors newest); paused-time `MissedTickBehavior::Skip` pin (verified to fail under `Burst`: 3 frames vs 1); merged-key eviction accounts every absorbed source chunk in BOTH arms — coalescer-side (Max's round-3 probe) and post-flush FIFO-side (Sami M13 / Max's round-4 probe: fails 102 vs 151 under `+= 1`). All three cap tests assert on independently walked retained bytes, not the accumulator (verified to fail without the `+text.len()` fix: 8,328,386 / 8,328,272 vs 4 MiB); the walker itself is pinned two-sided against the accumulator (all four blinding mutants M17–M20 verified to fail it, including the walker+fix cancellation pair). - `cargo clippy -p buzz-acp --all-targets -- -D warnings` clean, `cargo fmt --check` clean - Desktop: `tsc --noEmit` clean; node tests **4366 passed, 0 failed** — snapshot-clone family fully pinned: watermark aliasing (round 4), turns aliasing and tombstone aliasing (round 5, pre-existing gaps; each mutant verified to fail exactly its target test with an in-test control). Prior rounds: cross-channel reorder processed, cross-channel-delayed null-turnId `turn_error` evicts only its own channel's turn, null-bucket replay idempotency, same-channel stale/duplicate still skipped, watermark survives community-switch save/restore - All pre-push hooks green at the pushed commit (branch-skew, desktop-check, desktop-test, rust-tests, desktop-tauri-checks) Part of the rate-limiting fix stack; independent of `eva/rate-limit-fixes` by design (separate minimal PR per Tyler's ruling). --------- Signed-off-by: Eva <011987e296fd5006292d2f930b574be47c7801048d1983c46c425d3c95f0cffd@buzz.block.builderlab.xyz> Signed-off-by: Sami <f4a42a97e594b77bdbd8ee35191c8b28a94a4cb871d96f32921558275421fb68@buzz.block.builderlab.xyz> Co-authored-by: Eva <011987e296fd5006292d2f930b574be47c7801048d1983c46c425d3c95f0cffd@buzz.block.builderlab.xyz> Co-authored-by: Sami <f4a42a97e594b77bdbd8ee35191c8b28a94a4cb871d96f32921558275421fb68@buzz.block.builderlab.xyz>
**Category:** fix **User Impact:** Pull requests can once again pass the Desktop smoke test suite. **Problem:** The inbox attachment-edit smoke test still looked for the composer's former “Attach image” label after the shared action was renamed to “Attach file,” causing shard 3 and the aggregate Desktop CI job to fail on every PR. **Solution:** Update the stale accessible-name selector to match the current composer control while preserving the test's media-tag coverage. <details> <summary>File changes</summary> **desktop/tests/e2e/inbox-edit.spec.ts** Updates the attachment button selector to use the current accessible label so the existing attachment-edit regression test reaches the behavior it is meant to verify. </details> ## Reproduction steps 1. Build the Desktop E2E application with `pnpm -C desktop build:e2e`. 2. Run `cd desktop && pnpm exec playwright test --project=smoke tests/e2e/inbox-edit.spec.ts -g "editing an immediate attachment reply preserves its media tags"`. 3. Confirm the test locates the “Attach file” control and passes. Signed-off-by: Taylor Ho <taylorkmho@gmail.com>
## Problem Managed agents in internal Buzz builds should answer only their owner. Previously, an agent could keep a broader access setting and respond to other people, which did not match the access policy for internal builds. This PR makes owner-only access effective for every managed agent in internal builds and makes that restriction clear in the Desktop UI. Open source builds remain configurable. ## Changes - Enforce owner-only access when any managed agent starts or is deployed from an internal build. - Show the agent access control as locked to **Only me** in Desktop, with an explanation of why it cannot be changed. - Keep Welcome teammates working under the same rule without triggering unnecessary restarts. - Leave open source build behavior unchanged. This changes effective runtime access without rewriting stored or relay-advertised settings. The companion [block#4064](block#4064) explains the restriction in-thread when someone without access mentions an agent. The enforcement will remain inactive in shipped builds until [squareup/buzz-releases#74](squareup/buzz-releases#74) marks internal releases during the build. ## Screenshots | Before | After | | --- | --- | |  |  | ## Tests Added coverage for: - Runtime enforcement for [locally run agents](https://github.com/block/buzz/blob/e0165f52b52741a74184c9899e2b51eeec40c939/desktop/src-tauri/src/managed_agents/runtime/tests.rs#L196) and [deployed agents](https://github.com/block/buzz/blob/e0165f52b52741a74184c9899e2b51eeec40c939/desktop/src-tauri/src/commands/agents_tests.rs#L510). - The [current-build deployment path](https://github.com/block/buzz/blob/e0165f52b52741a74184c9899e2b51eeec40c939/desktop/src-tauri/src/commands/agents_tests.rs#L455), [invalid stored access](https://github.com/block/buzz/blob/e0165f52b52741a74184c9899e2b51eeec40c939/desktop/src-tauri/src/managed_agents/access_policy.rs#L98), and the [local startup guard](https://github.com/block/buzz/blob/e0165f52b52741a74184c9899e2b51eeec40c939/desktop/src-tauri/src/managed_agents/env_vars/tests.rs#L149). - Consistent enforcement across [both agent backends](https://github.com/block/buzz/blob/e0165f52b52741a74184c9899e2b51eeec40c939/desktop/src-tauri/src/managed_agents/access_policy.rs#L112). - Welcome teammates created as [locally run](https://github.com/block/buzz/blob/e0165f52b52741a74184c9899e2b51eeec40c939/desktop/src/features/onboarding/welcomeGuide.test.mjs#L384) or [deployed](https://github.com/block/buzz/blob/e0165f52b52741a74184c9899e2b51eeec40c939/desktop/src/features/onboarding/welcomeGuide.test.mjs#L393) agents, including [access-only](https://github.com/block/buzz/blob/e0165f52b52741a74184c9899e2b51eeec40c939/desktop/src/features/onboarding/welcomeKickoff.test.mjs#L202) and [runtime-related](https://github.com/block/buzz/blob/e0165f52b52741a74184c9899e2b51eeec40c939/desktop/src/features/onboarding/welcomeKickoff.test.mjs#L225) restart behavior. The full Desktop Rust and JavaScript suites, type checks, formatting, clippy, and file-size checks passed. Playwright E2E was not run. --- Originated from Buzz channel [buzz-agent-control](buzz://channel?id=cf5dada7-e26a-4887-ae41-b3bd5f42d3b2). Supersedes block#2537. --------- Signed-off-by: Tom Brow <tomb@block.xyz> Signed-off-by: npub1tquskdu6yc4h8l7xxtceculxw600grekeq0xg2ukqfrwl7vrzg3quz3gmp <58390b379a262b73ffc632f19c73e6769ef40f36c81e642b960246eff9831222@buzz.block.builderlab.xyz> Co-authored-by: npub1tquskdu6yc4h8l7xxtceculxw600grekeq0xg2ukqfrwl7vrzg3quz3gmp <58390b379a262b73ffc632f19c73e6769ef40f36c81e642b960246eff9831222@buzz.block.builderlab.xyz> Co-authored-by: Amp <amp@ampcode.com>
## Summary - virtualize the unfiltered channel member roster instead of eagerly mounting every member card - retain the existing member search/add flow and archived-member behavior - cover a 500-member roster, bounded mounted rows, and scrolling to the final member in E2E ## Cause The members sidebar rendered every active member card at once. On large channels this mounted hundreds or thousands of avatars, profile/presence consumers, menus, and DOM rows, blocking the renderer even though fetching the roster itself is fast. ## Testing - `pnpm typecheck` - `pnpm exec biome check src/features/channels/ui/MembersSidebar.tsx tests/e2e/channels.spec.ts` - `pnpm build:e2e` - `pnpm exec playwright test tests/e2e/channels.spec.ts --grep 'members sidebar (virtualizes large channel rosters|can invite relay-authorized agents|can invite and remove managed agents|collapses same-persona managed agents)'` (4 passed) - pre-push: `desktop-check`, full `desktop-test` (4,371 passed), branch-skew Implemented by Carl on Wes's behalf. Signed-off-by: Wes <wesbillman@users.noreply.github.com> Co-authored-by: Carl <c7ebe626f000404285d3686e1dc74cc07cc60a9754a150041ba132e14bd3e2ec@buzz.block.builderlab.xyz>
…ny — with real nodes (block#3862) ## Summary CI now proves the full Buzz shared-compute join story end to end: a member can discover another member's served model **through the Buzz relay alone** and run inference over the mesh, while a non-member gets nothing — the relay rejects its auth, and the mesh refuses to route for it even holding a leaked endpoint address. This is deliberately different from mesh-llm's own CI smokes (which bootstrap two nodes with a hand-carried invite token / mdns): here the **relay is the control plane**, exactly like the desktop app: 1. **Membership** — identities A and B are added via `buzz-admin` (kind:13534 NIP-43 roster); C is not. 2. **Advertise** — each member publishes a client-signed kind:30003 discovery note carrying its MeshLLM owner binding and (for the serve node) `serveTargets[].endpointAddr`, covered by an endpoint-binding signature — the exact payload shape the desktop coordinator publishes. 3. **Trust** — the serve node derives its admission allowlist from the relay (statuses ∩ roster) and requires the **exact expected {A, B} owner-id set** before starting with `TrustPolicy::Allowlist`. 4. **Join** — the client verifies owner + endpoint bindings and membership, then dials the relay-discovered endpoint (the desktop join-watcher's `dial_endpoint_addr` step). No out-of-band token. 5. **Infer** — a chat completion against the client's local OpenAI endpoint routes over QUIC to the serve node's model (CPU, SmolLM2-135M, ~105MB). 6. **Deny (differential)** — the stranger's NIP-42 auth must fail with the relay's own membership rejection (`restricted: not a relay member` — successful auth or any unrelated connect error fails the run), and dialing the leaked endpoint must not produce a routed inference — **while the trusted client re-proves inference immediately afterwards**, so a dead serve node can't masquerade as an admission denial. ## What's in the PR - `crates/buzz-relay/examples/mesh_relay_lifecycle_smoke.rs` — the harness. One process per node (mesh-llm keeps process-global state under `~/.mesh-llm`), orchestrator + serve/client/stranger roles, byte-identical binding payloads to `desktop/src-tauri/src/mesh_llm/identity.rs` (called out with keep-in-sync comments). Child stdout is pumped through a reader thread so every wait has a hard deadline; timed-out children are killed; exit statuses are checked. - `scripts/ci-mesh-lifecycle-smoke.sh` — provisions a membership-gated relay (throwaway owner + signing identities via `buzz-admin generate-key`), runs the harness, cleans up. Fails fast if :3000 is already occupied (a stale open relay would mask gating). - `scripts/start-relay-for-tests.sh` — gains opt-in NIP-43 membership env passthrough (`BUZZ_REQUIRE_RELAY_MEMBERSHIP` + `RELAY_OWNER_PUBKEY` + `BUZZ_RELAY_PRIVATE_KEY`). Default behavior unchanged. - `.github/workflows/mesh-lifecycle.yml` — separate, path-filtered, non-required workflow (mesh paths, the harness's dependency crates, `Cargo.lock`, dispatch), pinned to `ubuntu-24.04`. Caches the mesh native runtime + HF model keyed on the lockfile hash, so a mesh pin bump rolls the runtime cache. Uploads relay + harness logs on failure. ## Scope This is an **independent protocol harness**: it speaks the same wire protocol and payload shapes as the desktop but re-implements the binding/verification logic (the desktop crate is outside the workspace). Regressions inside the desktop's own discovery filtering are the desktop unit tests' job; what this smoke proves is that the relay + mesh-llm SDK + admission stack support the lifecycle end to end. ## Relationship to mesh-llm's CI Follows the shape mesh-llm's own CI proved stable (tiny CPU model, one runner, multiple real mesh-llm processes over real QUIC — cf. their `ci-two-node-client-serving-smoke.sh`), but swaps the token bootstrap for the relay-driven lifecycle, which is the part only Buzz can test. ## Validation Green on GitHub Actions (ubuntu-24.04) across three runs, including after rebases onto the mesh v0.74 upgrade (block#3467) and latest main: ``` PASS 1/6: relay-derived allowlist is exactly {A, B} PASS 2/6: serve member ready + advertised model: jc-builds/SmolLM2-135M-Instruct-Q4_K_M-GGUF:Q4_K_M PASS 3/6: client member discovered + joined via relay PASS 4/6: inference routed over the mesh: "PONG" PASS 5/6: relay rejected the stranger's NIP-42 auth (membership gate) PASS 6/6: stranger denied (gossip visible, inference rejected: 503 all tunnels failed) while trusted inference still routes PASS: full relay-driven mesh lifecycle verified ``` Also validated locally on macOS. `cargo fmt --all --check` and `cargo clippy -p buzz-relay --all-targets -- -D warnings` pass. ## Notes - The harness follows the repo's mesh `[dev-dependencies]` pin automatically, so it doubles as a canary for future mesh upgrades (it already caught the v0.73.1 → v0.74.0 bump during development). - The stranger "deny" accepts either shape mesh-llm exhibits: no model visibility at all, or gossip visibility with inference refused — mesh-llm applies the receiving node's owner policy after the gossip handshake, so admission gates *routing*, not gossip. The differential trusted-inference re-check (PASS 6/6) is what makes that a real denial rather than a dead server. - Model-visibility windows are tunable via `MESH_CLIENT_WINDOW_SECS` / `MESH_STRANGER_WINDOW_SECS` if shared runners prove slow — pin a longer window in the workflow env rather than re-running the job. --------- Signed-off-by: Michael Neale <michael.neale@gmail.com>
## Summary - require the macOS process to be running from an actual `.app` bundle before initializing `UNUserNotificationCenter` - keep the existing bundle-identifier requirement - cover packaged, case-insensitive `.app`, raw `target/debug`, and extensionless paths ## Why PR block#4799 guarded native notification initialization with `NSBundle.mainBundle.bundleIdentifier != nil`. Tauri embeds a bundle identifier in raw development executables, so `tauri dev` passed that guard and `UNUserNotificationCenter.current()` raised an uncaught `NSInternalInconsistencyException` because LaunchServices had no bundle proxy. ## Validation - focused macOS notification tests: 6 passed - direct raw debug executable no longer raises the notification-center exception - pre-commit formatting hook passed - pre-push package checks passed on pushed commit `f29a6664d2a863e7b8aa527f6149fd00b183e4de` The first push attempt hit an unrelated timing-test failure in `relay_admission::tests::concurrent_429_extends_the_window_for_parked_waiters`; its focused rerun passed, and the complete pre-push package suite passed on the next push. Signed-off-by: Wes <wesbillman@users.noreply.github.com> Co-authored-by: Carl <c7ebe626f000404285d3686e1dc74cc07cc60a9754a150041ba132e14bd3e2ec@buzz.block.builderlab.xyz>
…the authenticated socket (block#4990) ## Problem Users on v0.5.5 report "Can't reach the relay" toggling with brief "connected" flashes (field reports; also the macOS confirmation in block#4908). block#4737 closed the stuck-reconnect gaps; this is the opposite failure: the client redials fine, but then kills its own healthy socket. Mechanism (all on `main`): 1. AUTH succeeds → session emits `connected` (`relayClientSession.ts:583`), then awaits `replayLiveSubscriptions()`. 2. Paged channel backfill issues history REQs (`relayReconnectReplay.ts`, page limit 500). 3. A `CLOSED rate-limited:` on a **history** REQ arms the rate-limit gate but still rejects the history promise (`relayClosedRecovery.ts:38-51`). 4. The rejection escapes `replayLiveSubscriptions()` → `resetConnection()` tears down the authenticated socket. 5. Reconnect → AUTH OK → replay rate-limited again → loop. Each iteration re-spends the rate-limit budget, so the loop is self-sustaining. ## Fix Contain backfill failures inside the replay. Each subscription's paged backfill now retries behind the rate-limit gate up to `PAGE_REPLAY_MAX_ATTEMPTS` (3), then degrades to live-only **for this connection**. Socket health no longer depends on backfill success. Nothing is lost: the replay cursor (`lastSeenCreatedAt`) only advances on delivered events, so the next reconnect replays the same missed window. ## Red/green proof - Commit 1 (Pinky): e2e injecting `CLOSED rate-limited:` into the mid-replay history REQ — **red on main** (expected 1 reconnect dial, observed 2; connected-flash then teardown). - Commit 2 (this fix): same test **green unchanged** — one dial, state stays `connected` through the rate-limit hint plus the next backoff window. Why existing coverage missed it: the prior rate-limit e2e pre-armed the gate *before* replay (replay politely waits), and the CLOSED-injection test targeted a *live* subscription (which has its own retry path). Nobody injected back-pressure from the history REQ itself. ## Verification - `pnpm test`: 4374/4374 pass. - `playwright test tests/e2e/relay-reconnect.spec.ts`: 14/14 pass, including the new spec. - `tsc --noEmit` clean; Biome clean on touched files (pre-existing warnings on main in `personaCatalogRelay.test.mjs` / `terminal.css` untouched). ## Not addressed here (follow-ups from the same field reports) - AUTH terminal latch is too aggressive for relay-internal `error:` rejections (3 strikes during a relay bad window → stuck until click/relaunch; block#4908). - Server-side: `relay.drainJitterMs` (block#4542) defaults to 0 — enabling it on the hosted relay removes the deploy thundering herd that triggers these rate-limit storms. --------- Signed-off-by: Pinky <44b8e82baa6e0e254e0208d68f335c283c94e7b78dd1fa10d5a49d3f13dd0435@buzz.block.builderlab.xyz> Signed-off-by: Brain <21994759fc7a6fa6b965551d35cfd7897d262f2495467f2d78694ddcfa6a5c7e@buzz.block.builderlab.xyz> Co-authored-by: Pinky <44b8e82baa6e0e254e0208d68f335c283c94e7b78dd1fa10d5a49d3f13dd0435@buzz.block.builderlab.xyz> Co-authored-by: Brain <21994759fc7a6fa6b965551d35cfd7897d262f2495467f2d78694ddcfa6a5c7e@buzz.block.builderlab.xyz>
## Summary - add a persistent vertical pill beside the active community - keep the selected state visually distinct from unread dots and mention badges - preserve the existing `aria-current` selection semantics ## Screenshot  ## Test plan - `pnpm exec biome check src/features/sidebar/ui/CommunityRail.tsx tests/e2e/community-rail.spec.ts` - `pnpm test` (4,387 passed) - `pnpm build:e2e && pnpm exec playwright test tests/e2e/community-rail.spec.ts --project=smoke` (20 passed) - pre-push hooks: desktop check and 4,387 desktop tests passed on `c1e80c66d12f73c4eb5c03a19e932439b14caf2d` Signed-off-by: Wes <wesbillman@users.noreply.github.com> Co-authored-by: Carl <c7ebe626f000404285d3686e1dc74cc07cc60a9754a150041ba132e14bd3e2ec@buzz.block.builderlab.xyz>
## Summary - add stable three-step guidance to desktop mobile pairing - move code confirmation inline and show animated completion states - preserve pairing reset behavior and reduced-motion support ## Test plan - `pnpm --dir desktop check` - `pnpm --dir desktop exec tsc --noEmit` - `pnpm --dir desktop exec playwright test tests/e2e/mobile-pairing-qr.spec.ts --project=smoke` - pre-push desktop suite: 4,387 tests passed --------- Signed-off-by: kenny lopez <klopez4212@gmail.com> Co-authored-by: Fizz <50a12680c76f1a52c0b7af8dbb17e02c583227c290fb93b9a3defb456114223f@buzz.block.builderlab.xyz>
## Why The focus/split E2E test could capture the thread root before its programmatic middle-thread scroll had settled, then incorrectly report a scroll-restoration failure. ## What - Poll until the requested middle-thread scroll position is applied - Require the captured anchor to intersect the thread viewport and differ from the root - Preserve the existing focus-to-split-to-focus viewport assertions ## Risk Assessment Low — test-only synchronization change with no production behavior changes. ## References - Original failure: https://github.com/block/buzz/actions/runs/30231271427/job/89870533541 - Buzz thread: buzz://message?channel=12dd513d-45fd-48ff-80ac-8596d2fcc9d3&id=87ce6024b4bf74bfac2fa75d9f7bbbcc8f8fe2df460afe534152c495929f51ba - Reproduced confidence: 20 consecutive targeted passes, full spec pass, `just desktop-ci`, and `just ci` Generated with Codex Signed-off-by: npub1x4hk035p3p9q39a3fcrd2fe30lpkrhr5dwe0cqzzjphxyyh8m0gsq4vqap <356f67c681884a0897b14e06d527317fc361dc746bb2fc0042906e6212e7dbd1@buzz.block.builderlab.xyz> Co-authored-by: npub1x4hk035p3p9q39a3fcrd2fe30lpkrhr5dwe0cqzzjphxyyh8m0gsq4vqap <356f67c681884a0897b14e06d527317fc361dc746bb2fc0042906e6212e7dbd1@buzz.block.builderlab.xyz> Co-authored-by: Wes <wesbillman@users.noreply.github.com>
## Summary - replace ambiguous avatar play controls with centered Start and Restart pills - preserve avatar clipping while smoothly morphing actions into the running status dot - use accessible warning contrast and real restart behavior without a duplicate status badge ## Validation - `just ci` - focused Playwright coverage for morphing, shared geometry, and light/dark contrast Signed-off-by: kenny lopez <klopez4212@gmail.com>
…ges (block#4959) ## Problem When `buzz-agent` exhausts retries on a stalled LLM call, the error message reads: ``` transport: error sending request for url (...) (cumulative 721s, 3 attempts) ``` That text is reqwest's generic pre-response failure string — identical whether the cause is a TLS abort, a reset connection, or a `read_timeout` fire. An operator reading the log cannot tell whether something broke or whether the LLM generation legitimately took longer than the configured timeout. ## Root cause (probe-confirmed) A live probe against `goose-claude-fable-5` with a 900s client timeout completed in **370s** — well past the default `BUZZ_AGENT_LLM_TIMEOUT_SECS=240`. Extended-thinking models emit zero bytes on non-streaming calls until generation is complete, so reqwest's `read_timeout` fires on byte-silence regardless of whether the server is healthy. The 46× exact-721s stall signatures in production logs (3 × 240s + backoff) are deterministic self-inflicted timeouts, not network faults. ## Fix ### Pure classifier over `{is_connect, llm_timeout, phase}` A new `timeout_message(is_connect: bool, llm_timeout: Duration, phase: TimeoutPhase)` pure function produces factual messages with the configured duration value embedded verbatim. Two thin wrappers (`classify_transport_error`, `classify_body_read_error`) extract the reqwest flags and delegate. The duration reaches the classifiers through a new `read_timeout: Duration` parameter on `post()` and `openrouter_post()`; callers pass `cfg.llm_timeout`. ### Messages emitted | Case | Message | |---|---| | Connect-phase timeout (`is_connect && is_timeout`) | `connect timeout: no connection established within 10s` | | Transport read-timeout | `read timeout: no response bytes received within 240s (consider raising BUZZ_AGENT_LLM_TIMEOUT_SECS)` | | Body-read timeout | `read timeout: no further response bytes received within 240s (consider raising BUZZ_AGENT_LLM_TIMEOUT_SECS)` | | Non-timeout | `transport: {reqwest text}` / `body read: {reqwest text}` (unchanged) | `LLM_CONNECT_TIMEOUT` is now a named `const` (was inline `from_secs(10)`). **Out of scope by explicit decision:** streaming support, changes to `MAX_RETRIES` or backoff. ## Files changed - `crates/buzz-agent/src/llm.rs` — `timeout_message` pure fn + `TimeoutPhase` enum + `LLM_CONNECT_TIMEOUT` const; two classifier wrappers updated; `post()` and `openrouter_post()` gain `read_timeout` param; tests replaced. ## Tests `cargo test -p buzz-agent`: **397 passed, 0 failed** at `294ce5897`. **Pure-function tests (no network):** - `timeout_message_connect_true_shows_connect_timeout` — `is_connect=true` → connect-flavored text with 10s value; both phases checked - `timeout_message_transport_phase_shows_read_timeout_and_duration` — transport phase includes 240s and config knob - `timeout_message_body_read_phase_says_no_further_bytes_and_duration` — body phase says "no further", shows 300s - `timeout_message_duration_is_not_hardcoded` — 600s supplied → 600s in output, not 240s **Loopback reqwest integration tests:** - `classify_transport_error_read_timeout_is_loopback_verified` — TCP connect succeeds, server sends no bytes; verifies reqwest sets `is_timeout && !is_connect` and message contains 50ms value - `classify_transport_error_non_timeout_preserves_reqwest_text` — controlled accept-then-close on an owned loopback listener → non-timeout error; asserts exact `transport: {err}` output equality - `classify_body_read_error_timeout_says_no_further_bytes` — loopback server sends headers + 4 bytes of a declared-1024-byte body, then holds; verifies `is_timeout`, "no further", 100ms value, config knob No test performs egress beyond loopback (`127.0.0.1`). The TEST-NET-3 dial is deleted. --------- Signed-off-by: Will Pfleger <pfleger.will@gmail.com> Co-authored-by: Duncan <dcfd242e557282d7a1e2cf2e6877522682f1e5c6156dc92ca7d90eaedd3b0f95@buzz.block.builderlab.xyz>
…k#4845) **Category:** new-feature **User Impact:** People who lose a desktop identity can securely restore it from a signed-in Buzz phone without creating a replacement identity. **Problem:** A fresh or identity-lost desktop could not recover its existing full Buzz identity from an already-authorized phone. **Solution:** Add a SAS-confirmed reverse NIP-AB transfer, durable desktop import, a dedicated mobile recovery entry point, and clearer desktop recovery dialogs with tested loading, drag-and-drop, and failure states. https://github.com/user-attachments/assets/e9215c9c-80d0-462f-9161-0fa184ca2f74 <details> <summary>File changes</summary> **crates/buzz-core/src/pairing/session.rs** Adds the reverse encrypted payload and source-completion state transitions used for phone-to-desktop recovery. **desktop/src-tauri/src/commands/identity.rs** Exposes the existing guarded identity commit path for recovery imports. **desktop/src-tauri/src/commands/pairing.rs** Adds recovery-mode pairing, durable nsec import, start serialization, stale-task protection, and explicit rejection of unsupported recovery payloads. **desktop/src-tauri/src/lib.rs** Registers the recovery pairing command. **desktop/src/app/App.tsx** Refreshes the recovered identity before continuing onboarding. **desktop/src/features/onboarding/machineOnboarding.ts** Adds recovery transitions to the onboarding state machine. **desktop/src/features/onboarding/ui/BackupPasswordTimeline.tsx** Adds the visual backup-to-password-to-unlock progression. **desktop/src/features/onboarding/ui/IdentityRecoveryPairing.tsx** Implements QR generation, copy fallback, SAS confirmation, cancellation, expiry, and completion UI. **desktop/src/features/onboarding/ui/MachineOnboardingFlow.tsx** Connects private-key, phone, and backup recovery paths to the onboarding flow. **desktop/src/features/onboarding/ui/NostrKeyImportForm.tsx** Polishes recovery dialogs, backup drag-and-drop, loading stability, and security copy. **desktop/src/shared/api/tauri.ts** Keeps the existing pairing API surface focused on standard desktop-to-mobile pairing. **desktop/src/shared/api/tauriPairing.ts** Adds the recovery pairing invoke without growing the ratcheted shared API file. **desktop/src/testing/e2eBridge.ts** Mocks recovery pairing commands and lifecycle events for browser tests. **desktop/tests/e2e/identity-lost.spec.ts** Covers lost-identity entry, QR/copy recovery, SAS, cancellation, expiry, success, errors, backup import, drag-and-drop, and screenshots. **desktop/tests/e2e/onboarding.spec.ts** Verifies recovered identities continue through harness setup without replacement-key side effects. **mobile/lib/features/pairing/pairing_page.dart** Adds recovery-only scanning and explicit identity-handoff warnings. **mobile/lib/features/pairing/pairing_provider.dart** Recognizes recovery codes, returns the signed-in nsec after mutual SAS approval, and waits for desktop completion. **mobile/lib/features/settings/settings_page.dart** Accepts the recovery route builder at the app composition boundary to preserve feature isolation. **mobile/lib/features/settings/settings_page/connection_section.dart** Adds the signed-in “Send identity to desktop” settings action. **mobile/test/features/pairing/pairing_page_test.dart** Covers recovery-only validation and handoff messaging. **mobile/test/features/pairing/pairing_provider_test.dart** Covers reverse payload encryption, confirmation ordering, success, failure, timeout, and cleanup. </details> ## Reproduction steps 1. Launch Buzz Desktop with identity-lost state and choose **Recover from your phone**. 2. Confirm the QR and persistent **Copy pairing code** fallback appear without layout shift. 3. On a signed-in phone, open **Settings → Send identity to desktop**, scan or paste the recovery code, and compare the six-digit SAS on both devices. 4. Confirm on both sides and verify Desktop restores the identity and continues to harness setup. 5. Repeat from identity-lost state with **Recover from a backup file**; verify picker and drag-and-drop both advance to password entry and restore the encrypted backup. 6. Exercise cancellation, mismatched/unsupported codes, expired sessions, and an invalid backup; verify each returns actionable, non-stuck UI. ## Screenshots ### Desktop phone recovery — complete flow | Recovery entry | Pairing QR | Code match | Receiving identity | |---|---|---|---| |  |  |  |  | ### iOS Simulator — complete handoff flow | Settings entry | Recovery scanner | Manual recovery code | Code confirmation | |---|---|---|---| |  |  |  |  | ### Encrypted backup recovery — adjusted file flow | File picker | Drag-and-drop target | Password step | |---|---|---| |  |  |  | ## Verification - `cargo test -p buzz-core pairing` — 71 passed - `just mobile-test` — 1,169 passed - `pnpm build:e2e && pnpm exec playwright test identity-lost.spec.ts --project=smoke` — 15 passed - Full pre-push gates — desktop checks, desktop unit tests, Rust tests, Tauri checks, and mobile tests passed --------- Signed-off-by: Taylor Ho <taylorkmho@gmail.com> Co-authored-by: npub1223z34hd7vtwc6qj4s7flsxkj644nlre2nthu7lrrmkumhu3xddsrx9r6w <52a228d6edf316ec6812ac3c9fc0d696ab59fc7954d77e7be31eedcddf91335b@buzz.block.builderlab.xyz> Co-authored-by: Carl <acda9e433d19dcd0e6b6840f7f4b98f3a56f1fab98049d444c087019e6d36560@buzz.block.builderlab.xyz>
The local pre-push gate ran biome (`desktop-check`) and node:test (`desktop-test`) for desktop changes but never `tsc`, so TypeScript errors surface no earlier than CI's `desktop-core` job (`just desktop-build` = `tsc && vite build`). A branch with type errors passes every local hook today. This adds a `desktop-typecheck` pre-push command running `just desktop-typecheck` (`tsc --noEmit`) with the same glob/exclude as `desktop-check`, and updates the hook documentation in `AGENTS.md`. CI is unchanged — it already typechecks via `desktop-build`. Signed-off-by: Will Pfleger <pfleger.will@gmail.com>
…age boot (block#5086) Fixes the bug where running a dev build with stale localStorage would publish outdated channel sections, sort preferences, starred channels, and muted channels to the relay, clobbering the DMG installation's live state. ## Root cause All four sidebar-preference sync managers (`channelSectionsSync`, `channelSortSync`, `channelStarsSync`, `channelMutesSync`) collapsed five distinct fetch outcomes — no event, timeout, error, auth-race empty result, decrypt/parse failure — into a single `null`. Each hook's boot effect treated `null` as "no remote exists" and seed-published whatever was in localStorage, stamped at `max(now, lastRemoteCreatedAt+1)` with `lastRemoteCreatedAt` reset to 0 on every boot. A dev build with stale localStorage therefore re-signed old state as newer, and the DMG's live subscription applied it. ## Two guards **1. Tri-state fetch result** (`found | absent | failed`) — decrypt failure on an existing event reports `failed` and records `event.created_at`, so seed-publish is blocked even when the payload is unreadable. **2. Persisted head watermark** (`sidebarSyncWatermark.ts`) — keyed `{blobType, pubkey, normalizedRelayUrl}`, written to localStorage on every observed remote event (before decrypt on all paths: initial fetch, live subscription, `fetchOwnBlobBeforePublish`), hydrated at construction. Any session that has ever seen a remote blob skips seed-publish on the next boot even when the fetch returns empty. Relay URLs are normalised via `shared/lib/normalizeRelayUrl` (also used by profile storage) so the same relay written two ways never produces two keys. **Bootstrap owns the seed.** Each manager exposes `bootstrap(localStore)` that fetches, records the raw head, and delegates the decision to the single `runBootstrap` policy: hold on `failed` or `absent + prior watermark`, seed on genuine first-sync (`absent + zero watermark + non-empty local`), `apply-remote` when a blob was found. Hooks only act on `apply-remote`; they cannot publish during bootstrap. First-time sync is unchanged: successful EOSE with no event, zero watermark, and non-empty local state still seeds. ## LWW baseline preservation `fetchOwnBlobBeforePublish` for sections/sort snapshots the watermark before `recordRemoteHead` advances it, then compares the fetched event against the snapshot — advancing first would make `remote.createdAt > lastRemoteCreatedAt` always false and silently kill the whole-blob LWW merge. Stars/mutes merge per-entry via `mergeStores`, so no snapshot is needed there. ## Relay lifecycle All four hooks require a defined `relayUrl` (plumbed from `communitiesHook.activeCommunity?.relayUrl` in `AppShell.tsx`); while it is undefined no manager is constructed and no boot/live/reconnect effect binds. All effects depend on `[pubkey, relayUrl]`, so community switches tear down and rebind. `destroy()` cancels pending publishes without flushing — flushing would race community switching and could publish relay A's state to relay B via the shared `relayClient` singleton. Pending debounce-window edits are intentionally dropped: stars/mutes entries survive via per-entry merge on the next publish; a dropped sections/sort edit is lost because bootstrap whole-blob-replaces from remote on return. Known trade-off: a first boot with the relay unreachable holds (never seeds) until the user's next explicit edit — preferred over risking a stale seed-publish. ## Files - `sidebarSyncWatermark.ts` — watermark persistence + `runBootstrap` policy (tri-state `FetchResult`, `readWatermark`, `advanceWatermark`) - `shared/lib/normalizeRelayUrl.ts` — relay-URL normalisation shared by watermark keys and profile storage - `channelSectionsSync.ts`, `channelSortSync.ts`, `channelStarsSync.ts`, `channelMutesSync.ts` — tri-state fetch, pre-decrypt `recordRemoteHead` on all paths, sections/sort watermark snapshot for LWW, `bootstrap()`, cancel-without-flush `destroy()` - `useChannelSections.ts`, `useChannelSortPreference.ts`, `useChannelStars.ts`, `useChannelMutes.ts` — act on `bootstrap()` results, gate on `relayUrl`, `[pubkey, relayUrl]` deps on all effects - `AppShell.tsx` — passes `activeCommunity?.relayUrl` to `useChannelMutes` and `useChannelStars` - `sidebarSyncTestHelpers.mjs` — shared fake-window/localStorage/Tauri mocks for the four manager suites - Test suites — mutation-sensitive coverage: `failed→hold`, `absent+watermark→hold`, first-sync seeds, undecryptable head recorded on all paths, relay-A/B watermark isolation, watermark restart round-trip, sections/sort LWW baseline --------- Signed-off-by: Will Pfleger <pfleger.will@gmail.com> Co-authored-by: Duncan <dcfd242e557282d7a1e2cf2e6877522682f1e5c6156dc92ca7d90eaedd3b0f95@buzz.block.builderlab.xyz>
Owners and admins of a Buzz community get a desktop notification the
first time a new key joins their community. Requested by Tyler in
buzz-development ("we already have this [roster] — can we alert owners
and admins when a new key joins for the first time?"); design and
verification thread: channel `community-members-visibility`.
## Why the shape is what it is
- **kind:13534 membership snapshot is the alerting signal, not the
kind:8000 delta.** 8000 is leaky on two independent axes: its fan-out is
pod-local (no Redis hop — being fixed separately in block#4887), and
`buzz-admin add-member` publishes no 8000 at all by documented design.
The 13534 snapshot is the only signal covering every production join
path with cross-pod delivery (completeness audit: every
membership-insertion path enumerated at base `8342dfcc5`, all emit
13534).
- **This adds Desktop's first live 13534 subscription** — deliberate
line item. The existing read (`relayMembers.ts`) is a one-shot fetch;
without a live subscription no snapshot ever arrives passively and
nothing could fire.
- **8000 is subscribed only as a latency accelerator** and it *refetches
the authoritative snapshot* rather than alerting from its own payload,
so one ledger governs both signals and they cannot double-alert.
- **Persisted per-community/per-viewer ledger, written before the
notification fires.** Snapshot publication is eventual (60s reconciler
repairs failed best-effort publishes) and a reconciler-republished
snapshot is indistinguishable from a fresh one — only a durable record
answers "is this new". Also what makes reconnect replay (`since - 5s`
skew; `since === undefined` full-backlog edge) safe.
- **First snapshot per community seeds silently** (no notification storm
for existing members), and `seeded` is an explicit persisted bit — not
inferred from ledger non-emptiness, which would swallow the first
genuine join in a community whose only member is the viewer.
- Mounted in `useAppShellDesktopNotifications` (owns the
notifications-enabled precondition; `AppShell.tsx` is at the file-size
ratchet ceiling — net growth zero).
5 files, +3103 (production +752, tests +2,351), desktop-only. No relay
changes.
## Verification
**Current reviewed tip: `1854c4a5` — review-blessed code at
`d992ed295ead8c8423f81a752f4ad614718d85c6`** (clean tree, HEAD checked
in the same shell as each gate; history is `0e791f2d3` → merge of main
`2034e693a` → `fdeda44f0` → `5d0d2b4c3` → `a20a7d8cb` → merge of main
`0cfe4832` → `d992ed29` → `1854c4a5`, all fast-forward, no rebase or
force). Independently gated by Eva, Wren, and Sami: typecheck rc=0,
`pnpm check` rc=0 (pre-existing 1 warning / 2 infos), full Desktop unit
package **4431/4431**; push hooks pass. Wren's adversarial verdict at
`d992ed29`: APPROVE — minimalness 9, elegance 9, correctness 9, all four
cancellation seams plus 1b re-derived independently. `1854c4a5` is
assertions and comments only — no production behaviour change, so the
test count is unchanged.
**Remediation commits (review thread `community-members-visibility`):**
- `fdeda44f0` — authorization read from the signed snapshot being
reconciled (a demoting/removing snapshot fails closed before it can
disclose the joins it carries); >3 joins collapse to one summary;
8000-triggered refetches coalesce on a 500ms trailing window.
- `5d0d2b4c3` — join alerts coalesce **across** snapshots, not just
within one: a live burst arrives as several growing rosters, so delivery
defers onto a 1.5s trailing quiet window while ledger persistence and
dedupe stay synchronous per snapshot. Max measured 10 banners from 50
real joins before this; the same shape now produces one.
- `a20a7d8cb` — cancellation covers flushes already in flight, not just
queued timers: a generation token (bumped only by `clearPending`) is
rechecked after the profile lookup and before every send, so
demotion/removal/unmount/community-switch landing mid-flush suppresses
delivery; the notification title is captured with the batch rather than
read at send time. Concurrent-flush semantics pinned: a newer authorized
batch neither cancels nor is cancelled by an in-flight flush.
- `d992ed29` — the stale-authorized-frame disclosure, independently
reproduced at `0cfe4832` (held-open refetch released after a newer
demoting frame: `notifications=1`, body naming the joiner, where 0 is
required). Three fixes in one shape: every callback acts on a
per-effect-run session object (community id, viewer, ledger, ordering
state) instead of ambient current values, closing the community-switch
window; a `created_at` fence plus a fail-closed revocation latch, as one
mechanism, because the relay can publish two snapshots in the same
second so neither `<` nor `<=` alone is safe — the invariant is
“revocation wins”, not “newest wins”; and a 5s clamp on the 1.5s
trailing window so a sustained drip cannot defer delivery without bound.
Red-first: the four new arms fail at `0cfe4832` (25/29) and pass after
(29/29).
- `1854c4a5` — the privacy arm now asserts the persisted ledger is
unchanged across the delayed frame's release, not only the notification
count. Mutation-checked: moving the revoked check after the ledger
advance keeps notifications at 0 and passes the old assertion, and is
killed by the new one. Assertions and comments only.
**Mutation testing:** 9/9 mounted-hook mutants killed at `a20a7d8cb`,
each with a control row before and after — role/enabled gates,
reconnect, 8000 authority, failed-write handling and ref ordering,
community re-key/read, and query invalidation. The reducer/storage fix
separately killed 6/6 mutants with 15/0 controls; the foundational
ledger suite killed 9/9. At `d992ed29`: spelling the fence `<=` kills 5
arms; moving the empty-roster guard after the fence advance kills
exactly the fence-advance arm and nothing else (28/29). One
qualification stated rather than buried — moving the fence advance
itself up to the comparison SURVIVES the whole suite. That is an
equivalent mutant, not a coverage gap: the empty-roster guard returns
before the comparison, and authorization rejection latches `revoked` so
a later frame having moved the fence is unobservable. The scope is
written into the test's docstring. At `1854c4a5`: the
revoked-check-after-ledger-advance mutant is killed by the new ledger
assertion (and by the 1b arm).
**Scale/storage correction in `f6e5a3c57`:** the original 5,000-key cap
could evict members still present in a 5,001+ roster, causing them to
re-alert on every snapshot; read-time truncation reopened the same loop
after reload; and a raw quota exception could reject before notification
dispatch. The fix retains every on-roster key, caps only departed keys,
removes read-time truncation, and uses the app's quota-aware writer.
**Final ordering correction in `0e791f2d3`:** a failed post-recovery
write now skips notification and leaves the in-memory ledger unchanged,
so the next snapshot retries and delivers only after persistence
succeeds.
**Live-local matrix vs a real relay, executed at exact unchanged
`d75cc6cd9` and transferred to the current tip:** a 4,800-sequence
differential found zero old/new reducer divergences below the cap while
exercising the positive alert path; its negative control diverged as
required at 5,100 members (old re-alerts 100; new re-alerts 0). The
final hook change affects only the newly tested failed-write branch;
successful writes follow the same alert path exercised live. The live
communities were sub-cap and persisted successfully, so the matrix
remains applicable without a redundant rerun.
- Invite claim: owner and admin each exactly one notification; plain
member zero; 1.5s quiet window held (8000+13534 deduped); both open
clients live-refreshed the roster. Screenshot receipts SHA-256-pinned
and independently replicated.
- **CLI `buzz-admin add-member` (13534-only path):** DB counts moved
8000 `9→9`, 13534 `15→16` — zero accelerator events, exactly one alert
per manager. Proves snapshot-diff alone alerts.
- Plain member: zero notifications **and** zero
`buzz-community-join-seen.v1:*` localStorage keys before/after the join
(gate sits before the ledger).
- Staggered reload + replay dedupe: no alerts from startup
refetch/replay; republished already-seen snapshot produced zero through
a 2s quiet window.
- Community switch: independent per-community seed state; effect
re-keys; one alert per community, quiet window held at exactly two.
**Live re-verification at `d992ed29` is in progress** (Max; the
after-fix matrix leads with the delayed-refetch demotion arm, A→B switch
ledger isolation, the 5s sustained-drip timing, and packaged-app click
routing behind the positive/NIP-43 controls); earlier receipts at
`a20a7d8cb` cover the instrumented storm and cap-boundary re-drive;
earlier live receipts at `fdeda44f0` — privacy matrix
(demote/remove/promote), summary click-through — transfer where the diff
left those paths untouched.
## Known and accepted
- **8000 cross-pod fan-out is broken relay-side** — fixed in block#4887
(separate lane, not a blocker here): on a multi-pod relay the
accelerator only fires on the claim-handling pod; 13534 still covers
everyone, just not instantly.
- **Late-not-lost semantics.** A live frame missed during a
reload/socket gap is recovered by the next snapshot, reconnect refetch,
or remount backfill (`limit: 1`) diffed against the persisted ledger.
One live-run observation of an admin missing an immediate post-reload
fresh join is attributed to harness rate limiting; the recovery paths
above bound the damage to lateness, never duplicates.
- **Remote promotion activates on reload, not on the next snapshot**
(measured by Sami at `fdeda44f0`): the subscriptions are mounted from
the cached membership lookup, so a viewer promoted to admin by someone
else starts receiving join alerts only after a reload, community switch,
or local membership mutation refreshes that cache. Fails safe
(under-notify). Ruled accepted for v1 by Eva; the fix direction
(subscribing before authorization) is a deliberate design change
deferred to a follow-up if product wants instant activation.
- **Cross-user live-delivery staleness reproduced at the PR's own base**
(`2034e693a`, clean relay): a persisted send can fail to appear in an
already-open recipient timeline. Detached from this PR by a pinned-base
discriminator (identical failure with zero PR code) and tracked
separately in issue `6e2bda3092fa`; current main passes 4/4.
- **A stale demoting frame latches a genuine admin until reload or
community switch** (reverse ordering of the stale-frame privacy race,
`d992ed29`): if a snapshot that does not list the viewer as a manager
arrives out of order, the fail-closed revocation latch trips even though
the viewer is still an admin. The invalidation the latch fires refetches
the membership lookup, which correctly returns admin, so `active` stays
true, the effect deps do not change, and the session stays latched.
Fails safe (under-notify, never over-disclose) and consistent with the
promotion-on-reload semantics above. Ruled accepted for v1 by Eva;
self-clearing the latch would cost a third piece of timing state. Pinned
as documented behaviour in `useCommunityJoinAlerts.test.mjs` — and the
suppressed join is re-announced rather than lost, because a latched
session never records it in the ledger.
- **Lifetime-first-only semantics:** ever-seen ledger means
remove→re-add does not re-alert. Flagged for product ruling; one-line
change if re-adds should ping.
---------
Signed-off-by: Sami <f4a42a97e594b77bdbd8ee35191c8b28a94a4cb871d96f32921558275421fb68@buzz.block.builderlab.xyz>
Signed-off-by: Eva <011987e296fd5006292d2f930b574be47c7801048d1983c46c425d3c95f0cffd@buzz.block.builderlab.xyz>
Co-authored-by: Sami <f4a42a97e594b77bdbd8ee35191c8b28a94a4cb871d96f32921558275421fb68@buzz.block.builderlab.xyz>
Co-authored-by: Eva <011987e296fd5006292d2f930b574be47c7801048d1983c46c425d3c95f0cffd@buzz.block.builderlab.xyz>
…ock#4978) **Category:** fix **User Impact:** Users can navigate back while an identity key is being created, while Next remains visible and unavailable until creation finishes. **Problem:** The key-creation hold hid both navigation actions, leaving users without an escape route or a clear indication of what would happen next. **Solution:** Keep the onboarding footer mounted throughout creation, leave Back enabled, and gate Next on the completed identity state. <details> <summary>File changes</summary> **desktop/src/features/onboarding/ui/BackupStep.tsx** Keeps the onboarding navigation footer visible during key creation, with Back available and Next disabled until the identity is ready. **desktop/tests/e2e/onboarding-backup.spec.ts** Covers the loading and completed navigation states so the intended behavior cannot quietly crawl back out of the pit. </details> ## Reproduction steps 1. Start desktop onboarding and choose to create a new identity. 2. Submit the profile step and observe the key-creation screen. 3. Confirm Back is enabled while Next is visible but disabled. 4. Wait for key creation to finish and confirm Next becomes enabled. ## Screenshots | Before | After | | --- | --- | | Navigation actions are hidden during key creation. | Back remains enabled while Next stays visible and disabled. | |  |  | Signed-off-by: Taylor Ho <taylorkmho@gmail.com>
**Category:** fix **User Impact:** Agent cards and catalog listings now show the avatar belonging to the identity they represent. **Problem:** Running agent cards could show a stale definition avatar instead of the concrete agent profile, while adding another publisher's catalog entry could let local edits repaint that publisher's listing. This made agent identity look inconsistent across My Agents and the Agent Catalog. **Solution:** Treat the concrete agent pubkey profile as authoritative for running-card avatars, with the linked definition as fallback. Keep relay publications authoritative for foreign catalog presentation while using local copies only for linkage and selection state. | before | after | |--|--| | <img width="874" height="592" alt="Screenshot 2026-08-06 at 3 48 43 PM" src="https://github.com/user-attachments/assets/2cc6c9f7-ea50-413c-9c7b-4d34bd8b4ec7" /> | <img width="884" height="597" alt="Screenshot 2026-08-06 at 3 48 40 PM" src="https://github.com/user-attachments/assets/b14a865c-65c4-458f-9c30-d1a557c877d7" /> | | agent-set avatar not showing | agent-set avatar is showing | ## Changes <details> <summary>File changes</summary> **desktop/src/features/agents/lib/agentCardAvatar.ts** Adds the explicit avatar precedence rule for running agent cards and blocks avatar-dependent actions until the authoritative profile query settles. **desktop/src/features/agents/lib/agentCardAvatar.test.mjs** Covers profile precedence, definition fallback, blank avatar handling, and the profile-loading transition for linked-agent actions. **desktop/src/features/agents/lib/personaCatalogRelay.ts** Keeps publisher-provided catalog identity and behavior fields authoritative after a local copy is added. **desktop/src/features/agents/lib/personaCatalogRelay.test.mjs** Verifies local copies contribute linkage and selection without overriding publisher presentation. **desktop/src/features/agents/ui/UnifiedAgentsSection.tsx** Uses the concrete agent profile avatar before the linked definition avatar on running-agent cards. </details> ## Reproduction Steps ### Running agent card uses the agent profile avatar Use two visibly different, publicly reachable image URLs: **A** for the saved definition and **B** for the running agent profile. 1. In **Settings → Experiments**, enable **Agent-managed profiles**. This prevents Desktop from restoring the definition avatar over an agent's own relay-profile changes. 2. In **Agents**, create an agent with image **A** as its avatar and start it. 3. In a channel containing that agent, ask it to update its own Buzz profile avatar to image **B**. The exact CLI operation under the agent identity is `buzz users set-profile --avatar <image-B-url>`. 4. After the agent confirms the update, reopen **Agents → My Agents** (or reload the page so its kind:0 profile is fetched again). 5. Verify the running agent card shows image **B**, not definition image **A**. Open **⋯ → Share** and verify the share flow also uses image **B**. Before this fix, the My Agents card and share flow preferred image **A** whenever the linked definition had an avatar. ### Catalog listing remains publisher-authoritative This scenario requires a second Buzz identity so the entry is foreign to the account under test. 1. As the publisher identity, create an agent definition with a distinctive name, avatar, and instructions, then use **Share → Share to catalog**. 2. As the test identity, open **Agents → Discover agents**, find that publication, and add it. 3. In **My Agents**, open the added copy's **⋯ → Edit**, change its name, avatar, and instructions, and save. 4. Return to **Discover agents** and find the same publisher entry. 5. Verify it remains selected/added but still shows the publisher's original name, avatar, and instructions—not the test identity's local edits. ## Validation - `pnpm test` — 4,376 passed - `pnpm typecheck` — passed - `pnpm check` — passed with existing non-error notices --------- Signed-off-by: Taylor Ho <taylorkmho@gmail.com>
This change requires a valid signed Blossom authorization request and current relay membership for every media GET and HEAD request. It removes the unauthenticated compatibility path and updates desktop reads to send the required authorization. This blocks anonymous retrieval and access after relay-membership revocation. It does not yet bind a blob to its originating channel, so someone removed from a private channel can still read a known blob while remaining a relay member. That channel-ACL follow-up remains required before closing the full finding. ## Testing - `git diff --check origin/main...codex/security-media-read-auth` - Rebased onto `origin/main` at `5c98932` - Full CI pending Originating Buzz thread: `buzz://message?channel=3928fe05-df61-4b5d-b9c7-d623b9b10ea1&id=3c6c02312f763fbe0d2bfc33a6c1a362f91d0354f3d18b039cf7a0558c1439d1` --------- Signed-off-by: Jordan Mecom <jm@squareup.com> Signed-off-by: Alex Rosenzweig <arosenzweig@squareup.com> Signed-off-by: Eli Foster <efoster@squareup.com> Co-authored-by: Eli Foster <efoster@squareup.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…block#5133) ## What Relay-only carve-out of the ingest half of block#4999: generic EVENT ingest now accepts kind:30179 (NIP-PMA private managed-agent config). One file, `crates/buzz-relay/src/handlers/ingest.rs`, 16 insertions / 15 deletions; **two semantic lines**, byte-identical to the ingest hunk of block#4999 at `6f486e88`: 1. `required_scope_for_kind`: 30179 requires `Scope::UsersWrite` — same arm as its public sibling 30177 and the other owner-authored NIP-AP kinds. 2. `is_global_only_kind`: 30179 is owner-global, keyed `(pubkey, kind, d-tag)`; a stray `h` tag must not channel-scope it. The rest is import reflow plus replacing the guard test with a positive one (`private_managed_agent_kind_is_owner_scoped_global_user_data`: asserts UsersWrite scope, global-only, no h-channel scope). ## Why the guard test can be retired The removed test (`private_managed_agent_kind_remains_rejected_until_atomic_ingest_exists`) pinned a stated precondition: *"must not enter generic EVENT ingest before privacy and aggregate CAS deploy."* Both halves are resolved: - **Privacy** — the author-only read gates for 30179 shipped to main with block#4593: `AUTHOR_ONLY_KINDS` membership, `req.rs` pre-filter + result gates, `count.rs`, `event.rs` fanout, and the bridge pre-filter (`bridge.rs:999-1000` returns `restricted: author-only kinds require authors=[self]` / 403). Only the author can read the event back. - **Aggregate CAS** — block#4999 settled generation as **advisory**: the `g` tag is shape-validated, never relay-enforced. Last-write-wins per coordinate is the contract of record (see the kind:30179 contract blurb in block#4999), so no CAS mechanism is pending on the relay side. ## Why this is inert to existing relays and clients - No production desktop code on main authors kind:30179 — the codec (`private_managed_agent.rs`) has zero non-test callers. This PR accepts a kind nobody can produce yet. - Content is opaque NIP-44 ciphertext to the relay; the relay never decrypts it. - Reads remain author-only via the already-shipped gates above. - Storage is the standard parameterized-replaceable path already exercised by kinds 30175–30178. No schema, config, or migration changes. ## Testing - Full `buzz-relay` package suite at this commit: 859 passed, 1 failed — `api::mesh_demo::tests::demo_join_forwarded_arm_round_trips_echo` (504 vs 200), which **reproduces identically on clean main `769ac70b`** with this change stashed; pre-existing/environmental, not introduced here. - New positive ingest test passes. - Pre-push hooks green (branch-skew, rust-tests, desktop-tauri-checks). ## Relationship to block#4999 block#4999 (relay-primary agent config, desktop half) stays DO-NOT-MERGE pending live relay receipts + real CI; once this lands and deploys, its live test simplifies to plain `desktop-standalone` against the real relay, and block#4999 rebases to drop its now-duplicate ingest hunk (identical bytes → trivial rebase). Originating thread: buzz://message?channel=06f13ed3-0557-4ac2-922c-1545dd00bf97&id=2a43b3b4933a2ea78b77088619251c061355f9b7b6dc29ea0d702193f2344149 ## Brownfield FTS note (review findings, operator-ruled non-blocking for this PR) Max and Sami independently identified that the FTS privacy skip-set is regime-dependent: migration 0008 installs the positive allowlist (`kind IN (0, 9, 40002, 45001, 45003)`) **only on an empty events table**; an already-populated database keeps the 0001/0005 negative skip-list (wrapped by 0014 to add 30350), which omits 30179 — so on such an installation this PR admits 30179 rows whose NIP-44 ciphertext gets indexed by `to_tsvector`. Sami measured both regimes against real Postgres (brownfield: 30179 INDEXED; fresh: NULL) and demonstrated the existing drift test only exercises the fresh regime. `schema/schema.sql:222`'s canonical literal is also the negative list and omits 30179. Migration dates put any relay deployed with data before 0008 landed (2026-07-13) in the brownfield class. **Scope of exposure (Sami's trace):** not a content leak — `event_visible_to_reader` / `is_author_only_event` gates hold on both search surfaces (`req.rs:725`, `bridge.rs:1770`), so foreign readers receive nothing. Lost is the storage-level NULL-tsv backstop plus FTS page budget burned on post-filtered hits. **Operator ruling (Tyler, events `1472e5b6`, `cbd368ed`):** ship this PR without an exclusion migration. Safety argument that makes this sound rather than merely accepted: main has **zero non-test 30179 writers** until block#4999's desktop half deploys — no 30179 rows can exist, so nothing can be indexed in any regime while this PR is the only half live. **Additional review characterizations (Sami, non-blocking, on the record):** - *Behavioral delta enumerated:* routing triple (`required_scope_for_kind` / `is_global_only_kind` / `requires_h_channel_scope`) compared for all 65,536 kinds at base `769ac70b` vs head `77eeba6e` — exactly one row differs (30179). No other kind or client changes behavior. - *"SQL visibility before LIMIT" (NIP-PMA step 2):* no `AUTHOR_ONLY_KINDS` pushdown clause exists in `buzz-db` (only `SHARED_GATED_KINDS` has one). Author-only kinds are protected by the pre-filter (`author_only_filters_authorized`) plus post-filter omission; mixed-kind filters can burn candidate-page budget on discarded rows. Pre-existing and identical for 30300/30350 — not introduced here; noted so the NIP's step-2 checkbox is not read as fully ticked. - *Envelope validation gap:* 30179 is the only parameterized-replaceable kind at ingest with no per-kind envelope validator (codec grammar checks run in the desktop writer, not the relay). Generic limits only (256 KiB, ±15 min, pubkey==identity, d-tag bound). Self-inflicted footgun bounded to the author's own coordinate — candidate companion to the exclusion migration in the block#4999 rebase, deliberately not added here. **Bound follow-up (required before/with the block#4999 desktop half):** a 0014-shape additive migration (`pg_get_expr` capture + `CASE WHEN kind = 30179 THEN NULL ELSE (<existing>) END` wrap), add 30179 to the `schema/schema.sql:221` literal, and a brownfield-regime variant of the FTS drift test, per Sami's finding. Deploy-time spot check if ever wanted: `SELECT pg_get_expr(d.adbin, d.adrelid) FROM pg_attrdef d JOIN pg_attribute a ON a.attrelid = d.adrelid AND a.attnum = d.adnum WHERE d.adrelid = 'events'::regclass AND a.attname = 'search_tsv';` Signed-off-by: Tyler Longwell <tlongwell@block.xyz> Co-authored-by: Eva <011987e296fd5006292d2f930b574be47c7801048d1983c46c425d3c95f0cffd@buzz.block.builderlab.xyz> Co-authored-by: Tyler Longwell <tlongwell@block.xyz>
…lock#5136) ## Problem The harness posts each trial's task via `buzz messages send`, relying on `@<orchestrator-id>` name resolution. Task text is untrusted payload: when it contains @-tokens of its own, the CLI's mention resolver tries to resolve them as channel members, fails, and refuses to send — killing the trial with `RuntimeLaunchError` before the agent ever saw the task. Live occurrence: TB 2.1's `large-scale-text-editing` task embeds Vim macros (`:%normal! @a`). In the tb21-solo-1 run the trial died at launch: ``` RuntimeLaunchError: buzz messages send ... exited 1: {"error":"user_error","message":"mention '@A' does not match a current channel member; retry with --mention <pubkey>"} ``` Any TB task whose statement contains @-syntax is silently zeroed this way. ## Fix Pass the orchestrator's pubkey as an explicit `--mention` when posting the task. The CLI demotes unresolved @-tokens in the text to presentation-only when any explicit identity is supplied, so delivery still targets exactly the orchestrator and every @-token in the task statement becomes inert. The harness already holds the orchestrator's `AgentCredential` (it writes that pubkey into the worker roster tables), so no persistence is needed — fresh key per trial, fresh `--mention` per trial. Verified both halves against a live relay: a fenced `@a` without `--mention` still hard-fails (the resolver is not markdown-aware); the same content with an explicit `--mention` sends clean with `mention_pubkeys` containing only the target. ## Testing - `benchmarks/harbor-buzz-orchestra`: full pytest suite — 35 passed (34 baseline + new `test_send_mentions_by_pubkey_so_task_text_stays_inert`), ruff clean. Run against `origin/main` 769ac70 with exactly this patch applied. - `testbed`: full pytest suite — 23 passed, 1 skipped; ruff clean. ## Acceptance A task statement containing arbitrary @-tokens (Vim registers, emails, decorators) launches and delivers to the orchestrator instead of dying in `_send`. Originating Buzz thread: `buzz://message?channel=c3252dd2-0142-4e01-88c7-a2183c3960a5&id=74a65a0990fd2197882b66b5ea2707169d4a3dbd2020d1610c45150fb99f140b` Signed-off-by: Eva <011987e296fd5006292d2f930b574be47c7801048d1983c46c425d3c95f0cffd@buzz.block.builderlab.xyz> Co-authored-by: Eva <011987e296fd5006292d2f930b574be47c7801048d1983c46c425d3c95f0cffd@buzz.block.builderlab.xyz>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Syncs this fork's
mainwith the upstream block/buzzmainbranch. The fork was 398 commits behind upstream with no local-only commits, so this is a clean fast-forward — the branch is set to upstreammainatf53bbd1(fix(bench): mention the orchestrator by pubkey when posting the task, block#5136).Related issue
N/A — routine fork sync, none found.
Testing
No code authored in this PR; it only brings in already-reviewed upstream commits. Verified locally that
origin/mainis strictly behindupstream/main(0 commits ahead, 398 behind), so merging introduces no conflicts or divergence.🤖 Generated with Claude Code
https://claude.ai/code/session_01MFmJunfBUnwYLjiz5GgPgq
Generated by Claude Code