Repository navigation
fix: normalize x-api-key before the JWT/API-key auth selector - #954
Conversation
A real Anthropic SDK client constructed the documented way
(Anthropic(api_key=...)) sends its credential on x-api-key, never
Authorization (confirmed from the installed SDK's own source:
_api_key_auth returns {"X-Api-Key": api_key}). anthropic.APIKeyNormalizer
already existed to rewrite that header to Authorization: Bearer, but was
wired only at the mux leaf (mux.Handle("/v1/messages",
anthropic.APIKeyNormalizer(...))), inside authSelectorMiddleware, which
wraps the entire mux and inspects Authorization before any mux-level
handler runs. Every x-api-key-only request therefore had no Authorization
header at the point auth.Selector chose a path, fell through to the JWT
handler unconditionally, and 401'd with "missing bearer" regardless of key
validity.
Reproduced live against the deployed gateway with the real anthropic
Python SDK (0.122.0): the default api_key= construction always 401s with
"missing bearer"; the same call with auth_token= (forcing an
Authorization header directly) reaches the real API-key path instead.
Fix: apply anthropic.APIKeyNormalizer around the selector in
authSelectorMiddleware, so x-api-key is normalized before auth.Selector
ever inspects the header, for every /v1/* route. No-op wherever
Authorization is already set.
Also reshapes every /v1/messages error this surface owns (its own
validation refusals, and delegated-chain errors it forwards) into
Anthropic's real error envelope ({"type":"error","error":{"type",
"message"}}) instead of the OpenAI-shaped one it used unconditionally
before. Verified live against api.anthropic.com and the installed SDK's
_exceptions.py: the SDK's exception class is driven by HTTP status, not
body shape, so this did not previously crash a client, but err.type read
values outside Anthropic's documented enum and err.request_id was always
empty. Reuses the existing provider-blind sanitizer verbatim (through a
small local http.ResponseWriter capture) rather than duplicating its
regexes, so no sanitization logic changes.
Pre-dispatch middleware errors (the selector's own 401, budgetGate's 429)
and the global CompatHeaders OpenAI-style header injection on this
surface are left as documented gaps: fixing those touches shared
internal/middleware and internal/errors code used by every other
surface, out of this pass's scope.
Adds docs/anthropic-sdk-integration.md: the exact base URL, auth header,
copy-pasteable Python/TypeScript snippets, an OpenCode-style config, the
model aliases that actually work (hive-default, hive-fast, hive-auto,
live from GET /catalog/models), and every known limitation found during
this pass, each one verified rather than assumed.
Buglog entry (carried in the PR body per repo convention, to be appended
to main via a follow-up buglog-only PR):
{"date":"2026-08-17","error_message":"real Anthropic SDK client 401s with 'missing bearer' regardless of API key validity when pointed at Hive with only base_url and api_key set","root_cause":"authSelectorMiddleware inspects only the Authorization header and runs outside the mux; anthropic.APIKeyNormalizer (x-api-key -> Authorization: Bearer) was wired only at the mux leaf for /v1/messages, so it never ran before the selector's routing decision","fix":"apply anthropic.APIKeyNormalizer around the selector itself in authSelectorMiddleware (apps/edge-api/cmd/server/main.go), before auth.Selector inspects Authorization","tags":["anthropic","auth","edge-api","x-api-key","routing"]}
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
|
Bugbot is not enabled for your account, so this pull request was not reviewed. Enable Bugbot in the Cursor dashboard to get automatic reviews on future PRs. |
|
Warning Review limit reached
Next review available in: 29 minutes Limit details: You’ve used all 1 included review currently available under your plan. You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository. How can I continue?After more reviews become available, a review can be triggered using the To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews. How do review limits work?CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability. For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window. Please refer docs for additional details. Review details⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Pro Plus Run ID: 📒 Files selected for processing (7)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Adversarial review summary (dynamic selector, D-038)This PR touches an auth-routing boundary ( CodeRabbit CLI — ran ( ecc:code-review — SKIPPED. Structural capability gap: this task was
Plain adversarial pass (ticket-first) and manual security pass — run
Verdict: no security or correctness findings beyond the one CodeRabbit |
) ## Summary This PR is the output of live end-to-end verification of PR #954 (the x-api-key auth fix): a real Anthropic Python SDK client, a fresh API key minted through the actual developer console (`POST /api/v1/accounts/current/api-keys`, the same call the console UI button makes), run against the deployed gateway. PR #954 itself is confirmed working: a bogus `x-api-key` now correctly answers `401 authentication_error` (not the old `missing bearer`), and a valid key reaches the model on `hive-default` and `hive-auto` for non-streaming completions and a full tool-use round trip. That same live pass found a second, pre-existing bug that #954 did not touch and did not introduce: **every streamed response's first content block omitted the `index` key entirely.** `StreamEvent.Index` was a plain `int` tagged `json:"index,omitempty"`. Go's encoder drops a zero value under `omitempty`, and index 0 is the common case (the first, and usually only, content block in a plain-text or single-tool-call reply), so the wire event carried no `"index"` key at all. The real Anthropic Python SDK's own streaming accumulator (`anthropic/lib/streaming/_messages.py: accumulate_event`) indexes `current_snapshot.content[event.index]` and crashed with `TypeError: list indices must be integers or slices, not NoneType` on literally the first event of any streamed call — reproduced live 2026-08-18 against `api-hive.scubed.co` with `anthropic==0.122.0`. No existing test caught this. The existing assertions decode into `map[string]interface{}` and read `blockStarts[0]["index"].(float64)` via the comma-ok pattern, which silently yields the same zero-value `0` on a **missing** key as it does on a **present** key holding `0`. That is the exact "camouflaged test" shape: green whether the bug is present or not. ## Fix - `apps/edge-api/internal/anthropic/types.go`: `StreamEvent.Index` is now `*int` with `omitempty` on the pointer rather than the value, so `message_start`/`message_delta` (which never set `Index`, and which the real Anthropic protocol never puts an `index` field on) still carry no `"index"` key at all, while every `content_block_*` event gets an explicit, always-present index. - `apps/edge-api/internal/anthropic/stream.go`: every `content_block_start`, `content_block_delta`, and `content_block_stop` call site now passes `indexPtr(n)` explicitly, including `n=0`. - New regression test `TestSSETranslator_FirstBlockIndexIsPresentAndZero` checks key **presence** via `json.RawMessage` rather than a lenient map decode, so it fails the way the existing tests should have and now cannot silently pass the way they did before. ## Docs correction `docs/anthropic-sdk-integration.md` (written by a previous pass without ever exercising a real credential) is corrected against what this pass actually ran: - Every example switched from `model="hive-fast"` to `model="hive-default"`. **`hive-fast`'s live route currently 404s on both `/v1/messages` and `/v1/chat/completions`**: it maps to Groq's `llama-3.1-8b-instant`, which Groq now answers with `model_not_found` ("does not exist or you do not have access to it"). Confirmed live 2026-08-18 on both surfaces. Filed as issue (see below), not fixed in this PR: it is a DB-driven routing config change (`supabase/migrations/20260801_14_route_groq_fast_cheapest_model.sql` set this mapping), not a code change, and out of this PR's scope. - Documents that `hive-fast`'s failure response leaks LiteLLM's internal fallback-group bookkeeping (`Fallbacks=[{'hive-fast': ['hive-fast']}, ...]`, retry counts) verbatim into the customer-visible error message. No literal provider name appears, so the existing provider-name sanitizer doesn't catch it, but it is still internal routing implementation detail that should never reach a caller. - Documents a separate live finding, unrelated to #954: a console user who owns more than one workspace (a real fixture account used for this verification did) can have their session's default "current" account be one with **no** `tenant_billing_accounts` mapping. A key minted from that default state looks completely normal (`200`, a real `hk_...` secret) but answers every `/v1/messages` request with `403 permission_error` / `account_not_provisioned`, with no warning anywhere in the mint flow. The console does support selecting a different workspace via an `hive_account_id` cookie (forwarded as `X-Hive-Account-ID`), but nothing in the UI surfaces which account is "current" before a key is minted. ## Two issues filed from this verification pass, not fixed here 1. `hive-fast` routes to a deprecated/unavailable Groq model, and the resulting error leaks internal routing bookkeeping. Needs a DB routing update (out of this PR's scope) plus a look at whether the provider-blind sanitizer should also scrub LiteLLM fallback-group text, not just provider names. 2. A multi-workspace console user can mint a dead-on-arrival API key from their default account with no warning. Needs either a provisioning backfill for orphaned accounts, a mint-time check that surfaces `account_not_provisioned` before the key is created, or a UI affordance showing which workspace is "current." PR #962 (README refresh, open, not yet merged) also uses `model="hive-fast"` in both its OpenAI-SDK and Anthropic-SDK examples; flagging there directly since it's about to publish the same broken example. ## Test plan - [x] `TestSSETranslator_FirstBlockIndexIsPresentAndZero` — RED before the fix (confirmed: all three content_block_* events showed `"index"` absent from the wire JSON), GREEN after - [x] Full existing `apps/edge-api/internal/anthropic/...` suite passes unchanged (`go test ./apps/edge-api/internal/anthropic/... -count=1 -v`) - [x] Full `apps/edge-api/...` suite green (`go test ./apps/edge-api/... -count=1 -short`) - [x] `go build ./apps/edge-api/...` and `go vet ./apps/edge-api/...` clean - [x] Live re-verification pending this PR's deploy: real SDK `client.messages.stream()` against `hive-default` on the deployed box, full event sequence checked in order, will be posted as a PR comment once `deploy-demo-box.yml` for the merge commit succeeds. ## Boundary honored No billing, reservation, ledger, or metering code touched. No API key, token, or credential value printed anywhere in this PR, its tests, or its commit message.
The Web E2E (full stack) gate was intermittently red on the console
spend-alerts and profile-completion specs, both timing out on
`getByRole('heading', ...)` / `page.fill` because the console's generic
error boundary ("Something went wrong on this page") had tripped instead
of the real page. Both failures passed on Playwright's retry, which is
what the suite's flake-rate gate exists to catch rather than paper over.
The Next.js server log for the failing run showed the actual thrown
error: "could not load your workspace" (control-plane's writeInternal
message for accounts.EnsureViewerContext failures), with a digest
matching the error boundary's displayed reference number.
Root cause: web-console's Server Components call getViewer() unmemoized
per component (layout.tsx and each page.tsx independently), so a
brand-new user's very first page load can fire two concurrent
GET /api/v1/viewer requests. Both see zero active memberships and both
call provisionDefaultWorkspace, which derives an identical slug from the
same viewer's display name and inserts it with no idempotency guard.
The loser hits accounts.slug's unique constraint, and CreateAccount
propagated that raw pgx error unhandled, surfacing as an opaque 500 on
whichever request lost the race. E2E mints a fresh test user per run, so
its first console page load reliably lands in this window; the failure
is timing-dependent, hence flaky rather than a hard break.
Fixes:
- accounts.ErrSlugTaken sentinel, returned by CreateAccount on a unique
violation on slug (mirrors the existing CreateMembership /
ErrAlreadyMember pattern).
- provisionDefaultWorkspace now tells the two situations a slug
collision can mean apart: if the current viewer already holds an
active membership after the collision, a concurrent request for this
same viewer won the race and there is nothing left to do (recovered,
no error). Otherwise this is a genuine collision between two
different viewers whose display names happen to produce the same
slug (buildSlug has no per-user uniqueifier), and it retries with a
de-duplicated slug instead of failing that viewer's first sign-in.
Two new unit tests cover both branches directly against the accounts
service: TestEnsureDefaultAccount_SlugCollisionFromConcurrentSelf_Recovers
and TestEnsureDefaultAccount_SlugCollisionFromDifferentViewer_RetriesWithSuffix.
CI history check: this is a genuine, intermittent product bug, not a
selector-drift or fixture problem, and not caused by PR #954 or #955
(both merged while this check was still pending, per the standing
instruction to review that). The last full pass of this job on main was
1dc67ee (2026-08-17 16:56); #955's own pre-merge run hit this exact
race. The separate stale-selector failure on the OWUI e2e harness
("New chat" button vs `<a>`) is unrelated and already in flight on #951;
not touched here.
Buglog entry (to be appended to .wolf/buglog.jsonl on a buglog-only
branch cut from main, per the no-buglog-on-feature-branch rule):
{"error_message": "could not load your workspace (500) on GET /api/v1/viewer, surfacing as the web console's generic error boundary on the spend-alerts and billing-settings pages", "root_cause": "provisionDefaultWorkspace has no idempotency guard: two concurrent Server Component calls to getViewer() for the same brand-new user both attempt to create a personal workspace with the same deterministic slug, and CreateAccount let the resulting pg unique-violation on accounts.slug escape as a raw 500 instead of recovering the race", "fix": "translate the unique violation into accounts.ErrSlugTaken and have provisionDefaultWorkspace re-check the viewer's own memberships on collision: recover silently if a concurrent request for the same viewer already won, otherwise retry once with a de-duplicated slug for the genuine different-viewer collision case", "tags": ["accounts", "web-e2e", "race-condition", "ci-gate", "flaky-test"]}
#963) ## Summary Unblocks the required `Web E2E (full stack)` check on `main`, which was intermittently red and hard-blocking every queued PR (including #960, a live cross-tenant data exposure fix). - Fixes a genuine, intermittent backend race condition, not a test-selector or fixture drift. - `provisionDefaultWorkspace` (default-workspace auto-provisioning on first sign-in) had no idempotency guard. Next.js Server Components call `getViewer()` unmemoized per component (`layout.tsx` and each `page.tsx` independently), so a brand-new user's first console page load can fire two concurrent `GET /api/v1/viewer` requests. Both see zero active memberships and both try to create a personal workspace with the same deterministic slug (`buildSlug` derives it purely from the viewer's own display name). The loser hits `accounts.slug`'s unique constraint, and `CreateAccount` let that raw pg error escape unhandled as an opaque 500 ("could not load your workspace"), tripping the console's generic error boundary. - E2E mints a fresh test user per run, so its very first page load reliably lands in this race window on a real, unlucky-timing basis, hence flaky (fails once, passes on Playwright's retry) rather than a hard, permanent break. This is exactly what the suite's flake-rate gate ("A test that only passes on a retry is not a passing test") is there to catch. ## Root cause investigation - Walked `Web E2E (full stack)` job history on `main`: last clean pass was `1dc67ee6` (2026-08-17 16:56). PR #955's own pre-merge CI run hit this exact race on two specs (`console-spend-alerts.spec.ts`, `profile-completion.spec.ts`), both reported "2 flaky" (failed once, passed on retry) by the job's flake gate, which is what failed the job, not a hard 0-passed failure. - Downloaded that run's `nextjs-log-web-e2e` artifact: the underlying server error was `Error: could not load your workspace` with digests matching the two `error-context.md` snapshots' "quote reference" numbers exactly (`270132166`, `3259196144`). Traced that message to `writeInternal(w, r, "could not load your workspace", err)` in `accounts/http.go`'s `handleGetViewer`, then to `EnsureViewerContext` → `provisionDefaultWorkspace` → `CreateAccount`'s unhandled `INSERT ... accounts (slug ...)` unique-violation. - This is unrelated to the separate, already-known-stale `OWUI nightly e2e` selector issue (`button` named "New chat" vs the `<a>` OWUI actually renders) that's in flight on #951 — different harness, different spec, not touched here. - Neither #954 nor #955 introduced this: #955 is a test-only Go change and could not plausibly affect the web console; #954 touches auth-header normalization, unrelated to workspace provisioning. Both merged while this check was still pending (a known process gap, flagged separately), but neither is the cause. - This gate has **not** been silently ignored: the two prior real (non-skipped, non-cancelled) job completions on `main` before this PR were both green (`1dc67ee6` success, `ef06d693` success). PR #955 merged with its own pre-merge run showing this failure, which is the actual process gap worth naming; it did not merge past a run that had already gone green. ## Fix - New sentinel `accounts.ErrSlugTaken`, returned by `CreateAccount` on a unique violation on `slug` (mirrors the existing `CreateMembership` → `ErrAlreadyMember` translation already in the same file). - `provisionDefaultWorkspace` now distinguishes the two situations a slug collision can mean: 1. A concurrent request for the **same** viewer already won the provisioning race — re-checking that viewer's own memberships finds the winner's row, and this call recovers silently (no error; the caller re-lists memberships right after, same as the success path). 2. A genuine collision between two **different** viewers whose display names happen to produce the same slug (`buildSlug` has no per-user uniqueifier) — this viewer still holds no membership after the collision, so it retries with a de-duplicated slug (bounded, 3 attempts) instead of failing that viewer's first sign-in outright. ## Test plan - [x] `go build ./apps/control-plane/...` — clean. - [x] `go vet ./apps/control-plane/...` — clean. - [x] `gofmt -l apps/control-plane/internal/accounts/` — no unformatted files. - [x] `go test ./apps/control-plane/internal/accounts/... -count=1` — all pass, including two new regression tests: - `TestEnsureDefaultAccount_SlugCollisionFromConcurrentSelf_Recovers` - `TestEnsureDefaultAccount_SlugCollisionFromDifferentViewer_RetriesWithSuffix` - [ ] `Web E2E (full stack)` on this PR — will confirm the spend-alerts and profile-completion specs no longer flake once CI runs. ## Buglog entry Per `.wolf/` policy, not appended to `.wolf/buglog.jsonl` on this branch. To be carried into a separate buglog-only PR off `main` once this merges: ```json {"error_message": "could not load your workspace (500) on GET /api/v1/viewer, surfacing as the web console's generic error boundary on the spend-alerts and billing-settings pages", "root_cause": "provisionDefaultWorkspace has no idempotency guard: two concurrent Server Component calls to getViewer() for the same brand-new user both attempt to create a personal workspace with the same deterministic slug, and CreateAccount let the resulting pg unique-violation on accounts.slug escape as a raw 500 instead of recovering the race", "fix": "translate the unique violation into accounts.ErrSlugTaken and have provisionDefaultWorkspace re-check the viewer's own memberships on collision: recover silently if a concurrent request for the same viewer already won, otherwise retry once with a de-duplicated slug for the genuine different-viewer collision case", "tags": ["accounts", "web-e2e", "race-condition", "ci-gate", "flaky-test"]} ``` <!-- greptile_comment --> <details open><summary><h3>Greptile Summary</h3></summary> The PR makes first-login workspace provisioning atomic and resilient to both same-viewer races and cross-viewer slug collisions. - Adds a per-viewer PostgreSQL advisory transaction lock and lock-scoped membership re-check. - Atomically inserts the account, owner membership, and profile. - Translates slug uniqueness violations into a sentinel and retries cross-viewer collisions with a deduplicated slug. - Adds regression coverage and updates repository test doubles for the expanded interface. </details> <details open><summary><h3>Confidence Score: 5/5</h3></summary> The PR appears safe to merge, with no actionable changed-code defect identified. The transaction serializes provisioning by viewer, re-checks the same active-membership invariant used by the service, atomically persists all required workspace rows, and correctly handles the current slug uniqueness constraint. </details> <details open><summary><h3>Important Files Changed</h3></summary> | Filename | Overview | |----------|----------| | apps/control-plane/internal/accounts/repository.go | Adds atomic, transaction-scoped workspace provisioning with a per-viewer advisory lock, active-membership re-check, and precise slug-conflict translation. | | apps/control-plane/internal/accounts/service.go | Routes default provisioning through the atomic repository method and retries genuine cross-viewer slug collisions with bounded randomized suffixes. | | apps/control-plane/internal/accounts/service_test.go | Adds focused regression tests for concurrent same-viewer recovery and different-viewer slug deduplication. | | apps/control-plane/internal/accounts/types.go | Introduces the ErrSlugTaken sentinel used to distinguish retryable slug collisions. | | apps/control-plane/internal/accounting/http_test.go | Updates the account repository test double for the new atomic provisioning interface. | | apps/control-plane/internal/budgets/http_test.go | Updates the account repository test double for the new atomic provisioning interface. | | apps/control-plane/internal/ledger/service_test.go | Updates the account repository test double for the new atomic provisioning interface. | | apps/control-plane/internal/profiles/service_test.go | Updates the profile-aware repository test double for atomic provisioning. | | apps/control-plane/internal/usage/service_test.go | Updates the account repository test double for the new atomic provisioning interface. | </details> <details><summary><h3>Sequence Diagram</h3></summary> ```mermaid sequenceDiagram participant A as Viewer request A participant B as Viewer request B participant S as Accounts service participant DB as PostgreSQL A->>S: EnsureViewerContext B->>S: EnsureViewerContext A->>DB: Begin + lock(viewer) B->>DB: Begin + wait for lock(viewer) A->>DB: Re-check membership A->>DB: Insert account, membership, profile A->>DB: Commit and release lock B->>DB: Acquire lock and re-check DB-->>B: Existing active membership B-->>S: "wonElsewhere = true" S-->>B: Re-list and use winner workspace ``` </details> <sub>Reviews (1): Last reviewed commit: ["fix: close the workspace-provisioning ra..."](e08b57f) | [Re-trigger Greptile](https://app.greptile.com/api/retrigger?id=54270438)</sub> **Context used:** - Knowledge Base — [Control-plane identity and access](https://app.greptile.com/scubed/-/custom-context/knowledge-base/sakibsadmanshajib/hive/-/docs/control-plane-identity-access.md) <!-- /greptile_comment -->
…pplication (#951) Owner directive, 2026-08-17: "stop using iframes and keep it native inside OpenWebUI." `https://chat-hive.scubed.co/agents` rendered exactly one `<iframe>`, `loading="eager"`, pointed at `/agent-workspace/tasks`, which booted `apps/agent-console` (a second whole Next.js application) inside the page. That is why the agent surface still looked exactly as it did before the shell landed, why it was slow, and why a task could never become part of the conversation. The frame is gone. Scope is bounded deliberately: visual and functional parity with what the frame showed, natively, which is the task composer and the task list. No tool call cards, no progress panel, no transcript. All three are specified in `spec-2026-08-17-agent-run-surface` and blocked on the event relay (its step S6), and a tool card with nothing to render is worse than today's status row. ## Two premises in the brief that were already stale Both checked in the tree at `1dc67ee69`, both reported before building anything. 1. **`apps/agent-console/middleware.ts` does not send `X-Frame-Options: DENY` or `frame-ancestors 'none'`.** It sends `SAMEORIGIN` and `frame-ancestors 'self'` (`middleware.ts:80-84`), changed in #938 with a comment naming the shell as the reason. Nothing was working around anything: the frame was permitted because one Caddy listener serves both applications on one origin, so `'self'` is a real origin check that matches. 2. **The frame's `src` carried no credential.** It was `embed=1` and `theme=<light|dark>`, both re-validated and rebuilt rather than reflected (`middleware.ts:25-31`, `:93-100`). What the credential path actually was: `apps/agent-console` holds its own first-party Supabase session in cookies on the shared origin, and its browser code calls `/v1/agent/*` directly with `session.access_token`. That token comes from a direct Supabase sign-in, so it carries the `tenant_id` claim the `custom_access_token_hook` mints, and `JWTMiddleware` accepts it with no fallback. The frame existed because a second application had a second, better token. ## The problem, and why the recommended fix needed one correction The chat frontend holds Open WebUI's own session token, which edge-api has never heard of, plus a Supabase access token minted through Supabase's **OAuth-server** grant, which does not run `custom_access_token_hook` and therefore carries **no `tenant_id`**. `JWTMiddleware` rejects that (`middleware.go:123-129`). That is D-023 working, pinned by `inert_token_test.go`, and it is not relaxed here. Worse for any browser-side fix: **that OAuth token is not in the browser at all.** Open WebUI keeps it server side and resolves it per request through `get_system_oauth_token`, refreshing it when expired (`utils/middleware.py:3012-3045`). Handing it to page JavaScript would be a downgrade, not a fix. So the recommendation to reuse the mechanism `hive_jwt_forward.py` plus `owui_unwrap.go` already run in production is adopted. The correction: **`__metadata.upstream_auth` cannot serve this surface**, because three of the four agent-task calls have no JSON body to hide a token in. | Call | Method | Body | |---|---|---| | `GET /v1/agent/tasks` | GET | none | | `POST /v1/agent/tasks` | POST | JSON | | `GET /v1/agent/tasks/{id}` | GET | none | | `POST /v1/agent/tasks/{id}/cancel` | POST | none | (`apps/edge-api/internal/agenttask/handler.go:38-72`.) ## The design, three layers ### 1. edge-api: a second carrier on the same boundary `X-Hive-Upstream-Auth` carries the per-user token for requests with no body to carry it in. **This is not a new auth boundary.** The trust decision is unchanged and unchanged in strength: the header is honoured only when `Authorization` is exactly the shim key, which is the identical gate the body carrier already sits behind. What changes is the carrier, not who is allowed to use it. Properties, each with a test: * Honoured only under the shim key. Under a real API key, a real JWT, or no credential at all, it is ignored. * **Stripped from the forwarded request on every branch**, including the ignored one and including when the middleware is disabled. A header a client controls must never be readable by a handler, a log or an audit sink, because a handler that can read it is one that could later be taught to trust it. * Same 8 KiB cap as the body carrier, shared through one normaliser so the two cannot drift into accepting different shapes of the same credential. * Fails closed when present but unusable, rather than forwarding with the shim key still on `Authorization`. `requiresPerUserAuth` grows the agent-task paths. It previously covered `/v1/chat/completions` only, which meant a shim-key request to `/v1/agent/tasks` with no user token would pass through and **bind to the shim's own principal**, listing and cancelling the shim account's tasks instead of the user's. Nothing sent that request; this change is what keeps the new proxy from being one bug away from sending it. It is a tightening, not a widening: a silent mis-attribution becomes a 401. One latent defect fixed while in the same function: a present-but-empty `__metadata.upstream_auth` previously returned `unwrapOK` with an empty token, which skipped the fail-closed arm and forwarded to `/v1/chat/completions` with the shim key intact. It now takes the missing-carrier arm, so it 401s and logs like every other missing token. Not touched, deliberately: the tenant fallback is still gated on `IsOWUIUnwrapped` and is not widened by one line; `Selector` is unchanged; `inert_token_test.go` passes unmodified. ### 2. The chat container: a server-side proxy Only the **frontend** of `hive-open-webui` is built from the fork; the Python backend stays upstream's pinned image on purpose (Dockerfile header, D-036). So this arrives as a copied module plus one asserted splice, the same shape as `hive_model_picker` and `hive_rag_env_config`. The frontend builds with `adapter-static`, so there is no SvelteKit server to put it in instead. `deploy/docker/owui-patches/hive_agent_proxy.py`, mounted at `/api/v1/hive/agent`: * Every route depends on `get_verified_user`, and **no route reads a user id, a tenant id, an email or a token from the request**, so a caller cannot influence which principal edge-api resolves. * Presents the shim key on `Authorization` and the user's token on the carrier header. Neither is ever returned to the browser, logged, or included in an error. * Four named operations with the task id validated as a UUID before it is interpolated. Deliberately not a general-purpose proxy, which is the shape #770 removed from this image. * The create body is **rebuilt from two named fields** rather than passed through, so nothing a caller invents (a `__metadata` block, for one) reaches edge-api. * Adds no configuration: `OPENAI_API_BASE_URL` and `OPENAI_API_KEY` are already set on this container, so the shim key gains no new home. ### 3. The frontend `src/lib/hive/agentTasks.ts` and `AgentTasks.svelte`, plus the `/agents` route. Ported from the console, including two decisions that are easy to mistake for accidents and are not: the `unknown` status sentinel (a status this build cannot name still gets a row, because dropping it made a submitted task vanish) and the two engine sentinels (a deployment with no runtime reads "Blocked" in the warning register, never "Failed" in red next to a real failure). ## What happened to `/agent-workspace`, stated plainly It still exists and it is still served. Nothing in this PR removes it, and the PR should not be read as removing it. Three separate things get confused under that name, so all three, precisely: * **The `agent-hive.scubed.co` subdomain is already gone**, and was before this PR. It is NXDOMAIN and decommissioned (D-026, verified 2026-07-28). There is no separate agent host and has not been one for some time. * **`/agent-workspace` is a path on the chat host**, proxied to the `agent-console` container by `deploy/docker/Caddyfile.owui:161-168`. That rule is untouched here, so the path still resolves and still serves `apps/agent-console`. * **Nothing in the product points at it any more.** The nav row goes to `/agents`, and `/agents` no longer loads it. A user reaches it only by typing the URL. So after this PR the old surface is unreachable through the interface but still reachable by URL. Making it actually stop being a thing is the retirement described in the next section, and it is deliberately not in this PR: the console is the only surface that has ever launched a task against the live engine, so it is the control if the native path misbehaves on the box. Recommendation is to retire it once this has run there for a day. ## On the composer: the extraction was taken, not the fallback Answering the directive's point 1 plainly. `MessageInput.svelte` cannot be consumed wholesale: 2222 lines, 29 exported props, several required and chat-specific (`history`, `selectedModels`, `createMessagePair`, `stopResponse`, `taskIds`, `messageQueue`), and its container's classes are computed from chat stores. Cloning its CSS was banned, correctly. So the **presentational shell was extracted**, which was the preferred option: * `src/lib/hive/ComposerShell.svelte` is `#message-input-container` lifted verbatim: the rounded surface, the border, the hover and focus-within treatment, the inner padding. It reads no store; the two values that used to compute its classes arrive as props. * `src/lib/hive/ComposerSendButton.svelte` is `#send-message-button` lifted verbatim, classes and glyph unchanged. * `MessageInput.svelte` now renders both. **No behavioural edit was needed**, which was the stated condition for taking this option: the diff there is two tag swaps, one component swap and one import. What is deliberately **not** shared is the row of controls inside the container. Chat's carries attach, tools, skills, web search, voice and the model picker; the agent composer's carries the pack toggle. Sharing that row would mean sharing chat state with a surface that has none of it, which is the coupling the extraction exists to avoid. Per the directive: the heading block and the helper paragraph are deleted. The pack choice is a segmented **Knowledge work** / **Coding** toggle inside the composer. The one line describing the selected mode survives as a single muted line under the composer, because it is genuinely useful and it does not turn the composer back into a form. `Enter` sends and `Shift+Enter` is a newline, which is the chat composer's own behaviour. The two packs and their semantics are unchanged, and so is the default (`coding-pack`): this is a presentation change, not a capability change. Chat's own composer is proved unchanged two ways, per the condition set: the vitest suite is green, and before and after screenshots of the **chat** composer are posted below alongside the agent one. ## Does `apps/agent-console` still need to exist? Honest assessment, not acted on here. **No route in it is reachable from the product any more.** Nothing frames it, and the only two links that ever pointed at it are gone: the nav row goes to `/agents`, and `/agents` no longer loads `/agent-workspace/tasks`. It is now a signed-out URL that a user reaches only by typing it. What retiring it would take, in order: 1. Delete the `@agentConsole` matcher and its `reverse_proxy` from `deploy/docker/Caddyfile.owui` (lines 145 to 168), which is what makes `/agent-workspace` resolve at all. 2. Remove the `agent-console` service from `deploy/docker/docker-compose.yml` and its `Dockerfile.agent-console`, plus the `NEXT_PUBLIC_EDGE_API_BASE_URL` and `EDGE_API_INTERNAL_BASE_URL` wiring that exists only for it. 3. Remove the `agent-console-unit` job from `.github/workflows/ci.yml`, which is a required check, so branch protection needs updating in the same change or the merge gate waits forever on a job that no longer runs. 4. Delete `apps/agent-console/`. Two things worth keeping first, and neither is code: its `proof/` harness, and the accessibility reasoning in `task-console.tsx` (it refuses `--color-ink-3` for body text on a measured contrast argument). The second is already carried across into `AgentTasks.svelte`'s comments and CSS. One caveat against deleting it immediately: it is the only surface that has ever launched a task against the live engine, so it is the control if the native path misbehaves on the box. Recommendation is to retire it in a follow-up once this has run on the demo box for a day, not in this PR. ## Testing * `go test ./apps/edge-api/... -count=1 -short` in the toolchain container: green, whole module, including nine new cases for the carrier and the fail-closed paths. * `python3 scripts/test_owui_agent_proxy.py`, wired into `make test-scripts` (a required CI check). Pure standard library, no framework, matching the other `scripts/test_owui_*.py` self-checks. It is mutation tested: swapping the two credentials onto the wrong headers, deleting the missing-token guard, and forwarding the whole request body each make it fail. * `npm run test:frontend` for the fork, covering the decoder, the status mapping and the four calls. **A gap named rather than hidden.** `ci.yml` has vitest lanes for `apps/desktop`, `apps/web-console` and `apps/agent-console`, and none for `vendor/open-webui`, so every test ever written under `src/lib/hive` was coverage on paper only. This adds `npm run test:frontend -- --run` to the image's frontend build stage, which means the fork's tests gate the nightly build and the demo deploy. That is not the same as gating a pull request and is not claimed as such; a PR-time lane for this tree is a follow-up. ## Buglog entry ```json {"id":"owui-agent-shim-principal-gap","date":"2026-08-17","title":"OWUI shim-key requests to /v1/agent/tasks would have bound to the shim's own principal","error_message":"none observed; latent","root_cause":"requiresPerUserAuth in apps/edge-api/internal/auth/owui_unwrap.go returned true only for /v1/chat/completions, so a shim-key request to any other path fell through with the shim key still on Authorization and resolved as the shim account. Correct for the paths that existed then (embeddings and text-to-speech authenticate as the shim by design), wrong the moment a second per-user OWUI path appeared. A present-but-empty __metadata.upstream_auth had the same effect on the chat path itself, because it returned unwrapOK with an empty token and skipped the fail-closed arm.","fix":"Extended requiresPerUserAuth to /v1/agent/tasks and its subtree, and made an empty upstream_auth report as a missing carrier so it takes the 401 arm and the warn log.","tags":["auth","edge-api","owui","fail-closed","tenancy"]} ``` ## Visual proof: not posted yet, and exactly what exists today Stated precisely, because an earlier version of this section said screenshots were "committed under `docs/proof/` and posted in a comment". Neither was true. Nothing is committed under `docs/proof/` in this diff, no image is on this PR, and `owui-nightly.yml` has no step that commits or comments; it only uploads a transient CI artifact. That sentence described intent before the run finished and read as completed work. It will not be restated until an image is actually visible on this page. **What has been verified on a real stack with a real signed-in session**, run 32066161156, the OWUI end-to-end job on this branch: - `the agent workspace opens inside the shell and frames nothing` passed. That is the load-bearing one: an `iframe` count of zero on `/agents`, the composer present and accepting typed input, both toggle options clickable, the list painting exactly the rows the API returned matched against that response's own text, and the call to the proxy answering something other than 401, which is the single failure mode this design has. - `proof capture: the agent surface, natively, in both palettes` passed, so the images were produced. - Every chat spec passed on that same build: send and stream, multi-turn, model switch, tenant model visibility, sign-out. That is evidence the composer extraction did not break chat, and it does not depend on anyone reading a screenshot. **Why no image was attached at first, and a finding worth its own paragraph.** The captures were being written into `playwright-report-owui/proof`, inside the HTML reporter's own output folder. That reporter clears its folder before writing the report, so **every capture was produced and then deleted**, and the result looked exactly like a capture step that had never run. Nothing failed, nothing warned, and the test that wrote them passed. Confirmed by downloading the artifact and finding no `proof/` directory rather than by inferring it. Anyone adding a capture step to a Playwright suite in this repository will walk into the same trap, which is why it is recorded here rather than only fixed: proof must never be written inside a reporter's output directory. The captures now go to a sibling directory with its own upload step. This is the same silent-absence shape that let the end-to-end gate rot unnoticed, where an artifact that is missing and an artifact that was never asked for are indistinguishable after the fact. **Two attempts to re-capture since then both failed outside this branch**, and neither is counted as anything: - Run 32112158077: the sign-in helper's consent hop exceeded its 30 second budget. The container log shows the token exchange completing 28 seconds after its redirect and the OAuth session being stored, so the login worked and was simply late. - Runs 32113599893 and 32114089597: the seeder's first write returned `POST /tenants -> 522` in both. Without fixture credentials the OWUI project matches zero spec files, so nothing in this PR ran at all. **Correction to an earlier version of this section**, which called that 522 an external Supabase problem. It is very probably ours. The same job log carries `(ECHECKOUTTIMEOUT) unable to check out connection from the pool after 15000ms in Session mode`, which is the documented 15-client session-mode pooler ceiling being hit, and that single cause explains both symptoms: a PostgREST request that cannot get a database connection hangs until Cloudflare gives up with a 522, while a direct `pg` client gets the checkout timeout by name. Blaming the provider was the comfortable read and the evidence does not support it. I also want to retract a piece of reasoning rather than quietly drop it: I probed `/rest/v1/` unauthenticated, got a 401, and treated that as evidence the origin had recovered. An unauthenticated request never reaches the connection pool, so that probe could not have detected the condition that matters and proved nothing. Neither failure touches a file in this diff, and a run that never reached the suite is an absence of information rather than a result. The captures are outstanding for that reason and for no other. **How they will be posted.** Through `scripts/post-pr-visual-proof.sh`, which uploads to a release asset. A raw link pinned to this branch 404s the moment the branch is deleted, and this repo squash-merges and deletes branches, so that mechanism produces proof which expires exactly when the PR it proves gets merged. ## A third thing, and the budget I did not raise Investigating why the proof runs kept failing turned up a cause that is not a test problem. The consent and sign-in pages are web-console's, on a different origin, and web-console is served by a development build, so it compiles each page the first time anyone requests it. On run 32112158077 the first two login attempts produced no page within 30 seconds and the third completed the whole exchange in 21 seconds once warm. That is a route compiling on demand, not a slow protocol. The harness now waits for the consent app to be serving before it starts timing the login. **The 30 second login budget is deliberately unchanged.** Raising it was the obvious move and the wrong one: it would have buried a cold compile inside a per-login timeout, and every future slow login would then look exactly like this one, which is how the next real regression gets absorbed instead of noticed. The new wait is a separate budget for a separate thing, the app becoming reachable at all. It also adds no new failure mode, because the assertion it precedes already requires a DOM that exists only on the consent origin. Why that app is served in dev mode is a product question, it is deliberate per D-022, and it is filed as #967 rather than decided here. It is a plausible cause of the slow sign-in being reported in real use. ## Two things a reviewer should know **The OWUI end-to-end harness was red before this branch and is fixed here.** It still is on `main`, at that same readiness check, which is independent evidence this fix is real and not a symptom of something else on this branch. It is also a different defect from the provisioning race PR #963 is separately fixing, so the two should not be read as one problem. Its sign-in readiness signal waited for a `button` named "New chat", and Open WebUI renders both New Chat controls as anchors carrying an `aria-label`, so that query could never match. The sign-in it guards actually succeeds; the downloaded failure screenshot shows a working signed-in chat page. A sibling file already matched either role for exactly this reason. That fix is why any of the verification above exists, and it is separate from this feature. **Two nav tests from #938 went red the first time they ever ran**, which was in this branch after that fix. Neither is a regression here and neither was weakened. Both assumed a sidebar that starts expanded while this fixture's starts collapsed, so a text assertion resolved to the icon-only rail and read "", and the collapsed-rail test waited for a control that only exists while the sidebar is open. Each test now drives the sidebar into the state it asserts about. --- ## Rebase onto current main, 2026-08-22, and what this closes Rebased onto `main` at `c30882491`. The base this branch was validated against, `1dc67ee69`, predates the migration of the entire database off hosted Supabase onto the self-hosted instance on the demo box, the move of the public auth origin to a same-origin `/auth/v1` route on the console, and the switch to admin-provisioned-only signup. One conflict, in `Makefile`, where `main` had added seven self-checks to `test-scripts` and this branch adds one; both are kept. ### This is the fix for the top demo blocker, issue #540 Re-confirmed live on the box on 2026-08-22: a user signs into `chat-hive.scubed.co` normally, clicks **Agents**, and is presented with an email and password form captioned that the workspace is separate from chat. Proved there by difference rather than guessed: chat's OAuth handshake mints an Open WebUI token and never writes the `sb-...-auth-token` cookie that the embedded `apps/agent-console` reads, and injecting the same session's Supabase cookies on the `chat-hive` origin makes the real workspace render immediately. That is the second-application problem this branch removes. The credential the embedded application was looking for is no longer needed by anything, because the surface is native and authenticates with the Open WebUI session the user already has: the panel calls `/api/v1/hive/agent/*` on the chat origin, and `deploy/docker/owui-patches/hive_agent_proxy.py` brokers it server side through `get_system_oauth_token`, which is the one place the user's Supabase token is reachable. #540's own analysis named that as the only possible broker. ### Does a second credential prompt remain reachable **Yes, by one route, and this branch does not close it.** `/agent-workspace/*` is still proxied to `apps/agent-console` by `deploy/docker/Caddyfile.owui`, and that application still renders its own email and password form, so a typed or bookmarked `https://chat-hive.scubed.co/agent-workspace` still reaches one. This branch's only edit to that file is a comment; it changes no route. No route inside the chat interface leads there any more. `vendor/open-webui/src/lib/hive/nav.ts` points the sidebar entry at `/agents`, and the only two remaining mentions of the old path anywhere in the front end are historical comments. So the demo path is fixed and the residual is a URL nobody is shown. Answering that path 404 was considered and rejected here rather than skipped: the Tauri desktop app targets `/agent-workspace` as its console base path (`apps/desktop/src/settings.ts`, `apps/desktop/src-tauri/src/settings.rs`), and `apps/web-console/tests/e2e/_probe/agent-workspace-flows.spec.ts` is a coverage ledger over that surface. Retiring `apps/agent-console`, or moving the desktop app onto the native surface first, is a separate decision with a far larger blast radius than this pull request. ### One coupling worth knowing before this is demoed The proxy reads the user's Supabase token through `get_system_oauth_token`, which goes through `OAuthManager.get_oauth_token`, which refreshes five minutes before expiry and **deletes the OAuth session outright when that refresh fails**. That is issue #782, still unfixed on `main` and fixed by PR #787. Until #787 lands, this agent surface stops working roughly 55 minutes after sign-in for the same reason chat does, and the panel will correctly report "Your Hive sign-in could not be confirmed. Sign in again and retry." A demo that runs longer than that from a single sign-in needs #787 merged first. ### Review findings addressed on the rebased head Two threads, both fixed rather than argued, plus one defect found while verifying them. A list refresh overlapping a create or a cancel replaced the whole task array with its older answer, dropping the row the user had just submitted or reverting one they had just cancelled, and when that stale answer held no in-flight task the poll loop stopped too and the screen did not recover without a reload. A refresh now captures a mutation counter before its request goes out and discards its own answer if a create or a cancel landed while the request was open. The third entry in the identity smell tuple in `scripts/test_owui_agent_proxy.py` could not fail: for `headers.get('Authorization')` the accessor templates expanded to `request.query_params.get('headers.get('Authorization')')`, a string no Python source can contain, so that iteration reported a pass over the header read it was named after. The header read now has its own direct test against the whole `request.headers` attribute, and it fails on purpose when the string is planted on a line that never executes. Found while running the above: `agentTasks.test.ts`, 203 lines of assertions, was running in no job at all. The module imported `$lib/constants`, which reaches `$app/environment`, and the only runner that covers this front end runs plain vitest with no alias resolution, so the file could not be loaded. It would have turned that required check red on whichever of #951 and #952 merged second. The API base is now a parameter with a production default and the component passes the dev-aware value, so a built bundle and `npm run dev` both behave exactly as before, and the tests load with no configuration: 16 of them, one of which fails when `encodeURIComponent` is dropped from the cancel path. ### Verification on the rebased head - 31 vendored front end tests pass across three files, 16 of them in `agentTasks.test.ts`, which had never executed before this rebase. - `make test-scripts` green, including `test_owui_agent_proxy.py` and the seven self-checks `main` added. - `go test ./apps/edge-api/internal/auth/... -short` green, which covers the `owui_unwrap.go` change against `main`'s newer `x-api-key` normalisation (#954). - `Caddyfile.owui` validates against the pinned Caddy image, using `main`'s new "Every Caddyfile adapts" CI step. ### Visual proof Posted on this pull request as release assets, with the full method, the artefacts of the stub, and the `/agent-workspace` residual in `docs/proof/agents-native-no-iframe-2026-08-22/README.md`. Measured in the page rather than asserted: zero password inputs, zero iframes, and no "Sign in" text on `/agents`. The capture uses the real bundle from `docker build --target frontend` on this branch with a stubbed backend on a loopback origin, and says so on every artefact. A live capture is not available before merge for two reasons that are not properties of this branch: the panel needs the `owui-patches` router this branch adds, which exists only in a rebuilt image, so bundle interception against the deployed origin would 404; and the development box's `.env` still points `SUPABASE_URL` at the hosted Supabase project deleted in the migration, which now returns `ENOTFOUND`, so no session can be minted. No password was set, reset or rotated to work around that. ### Buglog entry To be appended to `.wolf/buglog.jsonl` on `main` in a separate buglog-only pull request after this merges, per the openwolf protocol. ```json {"id":"owui-agents-second-credential-prompt","date":"2026-08-22","error_message":"Clicking Agents in the chat sidebar presented its own email and password form captioned that the workspace is separate from chat, one click after a successful chat sign-in","root_cause":"The Agents route embedded apps/agent-console in an iframe, and that application authenticates from a Supabase SSR cookie on the chat origin which chat's OAuth handshake never writes: the handshake mints an Open WebUI token only, and the user's Supabase token is reachable only server-side inside the chat container as the stored OAuth token","fix":"Removed the iframe and rendered the task surface natively in the chat application, authenticating with the Open WebUI session the user already holds and brokering the Supabase token server-side through a FastAPI router added by owui-patches/hive_agent_proxy.py","tags":["owui","auth","agents","iframe","session","demo-blocker","issue-540"]} {"id":"agent-tasks-stale-poll-overwrites-mutation","date":"2026-08-22","error_message":"A newly created agent task row disappeared, or a cancelled row reverted, and polling sometimes stopped entirely until the page was reloaded","root_cause":"refresh assigned the fetched list over the whole tasks array, so a poll already in flight when a create or cancel completed landed afterwards and overwrote it, and schedulePoll then decided from that stale array and stopped when it held no in-flight task","fix":"A mutation counter captured before the request goes out; refresh discards its own answer when a create or cancel landed while the request was open","tags":["svelte","race","polling","agents"]} {"id":"agenttasks-unit-tests-never-ran","date":"2026-08-22","error_message":"agentTasks.test.ts reported no failures because it was never loaded by any job","root_cause":"The module imported $lib/constants, which reaches $app/environment, and scripts/test-owui-hive-frontend.sh runs plain vitest over copied files with no SvelteKit alias resolution, so the test file could not be loaded at all","fix":"The API base is a parameter with a production default and the component passes the dev-aware value, so the module imports nothing and the tests load with no configuration","tags":["tests","unfailable-check","vitest","sveltekit"]} {"id":"owui-agent-proxy-smell-check-unfailable","date":"2026-08-22","error_message":"The identity smell loop in test_owui_agent_proxy.py reported a pass over a header read it was named after","root_cause":"The tuple entry headers.get('Authorization') was expanded through accessor templates into request.query_params.get('headers.get('Authorization')'), a string no Python source can contain, so that iteration asserted nothing","fix":"Separated the header smell into a direct substring test against the whole request.headers attribute, demonstrated failing on purpose","tags":["tests","unfailable-check","security"]} ``` <!-- greptile_comment --> <details open><summary><h3>Greptile Summary</h3></summary> The PR replaces the iframe-hosted agent workspace with a native Open WebUI task surface and brokers per-user agent API credentials through a constrained server-side proxy. - Adds a guarded upstream-auth header carrier and fail-closed agent-task authentication. - Adds native task composition, listing, cancellation, polling, and shared composer presentation. - Updates frontend, proxy, authentication, and end-to-end verification coverage. </details> <details open><summary><h3>Confidence Score: 5/5</h3></summary> The PR appears safe to merge. No blocking failure remains; the mutation-generation guard prevents stale refreshes from overwriting completed creates or cancellations, and ID filtering prevents the create/refresh duplicate-row race. </details> <details open><summary><h3>Important Files Changed</h3></summary> | Filename | Overview | |----------|----------| | apps/edge-api/internal/auth/owui_unwrap.go | Adds a stripped, shim-gated header carrier for per-user agent credentials and extends fail-closed authentication to agent-task paths. | | deploy/docker/owui-patches/hive_agent_proxy.py | Adds the authenticated server-side broker that resolves each user’s OAuth token and forwards only the four supported task operations. | | vendor/open-webui/src/lib/hive/AgentTasks.svelte | Implements native task composition, polling, cancellation, stale-refresh suppression, and create-row deduplication; both previously reported races are fixed. | | vendor/open-webui/src/lib/hive/agentTasks.ts | Defines the typed agent-task client, decoding, status handling, and constrained proxy calls used by the native surface. | | vendor/open-webui/src/routes/(app)/agents/+page.svelte | Replaces the framed agent application with the native AgentTasks component. | | vendor/open-webui/src/lib/components/chat/MessageInput.svelte | Adopts extracted composer shell and send-button components without changing the chat composer’s behavior. | | apps/web-console/e2e/phase-19/owui/09-agent-workspace-nav.spec.ts | Verifies native rendering, absence of iframes, proxy authentication behavior, navigation, task rows, and visual-proof capture. | </details> <details><summary><h3>Sequence Diagram</h3></summary> ```mermaid sequenceDiagram participant U as Signed-in user participant UI as Native OWUI agent surface participant P as OWUI agent proxy participant A as Edge API auth middleware participant T as Agent-task API U->>UI: Open /agents UI->>P: "GET/POST /api/v1/hive/agent/*" P->>P: Resolve server-side OAuth token P->>A: Shim authorization + upstream-auth carrier A->>A: Validate shim gate, strip carrier, unwrap user token A->>T: Request authorized as signed-in user T-->>A: Task response A-->>P: Task response P-->>UI: Sanitized response UI-->>U: Native task list and composer ``` </details> <sub>Reviews (4): Last reviewed commit: ["fix: deduplicate the created task row ag..."](f3df569) | [Re-trigger Greptile](https://app.greptile.com/api/retrigger?id=54086925)</sub> **Context used:** - Knowledge Base — [Control-plane marketplace-backed agent task workflow](https://app.greptile.com/scubed/-/custom-context/knowledge-base/sakibsadmanshajib/hive/-/docs/control-plane-agent-workflows.md) - Knowledge Base — [Edge API security boundary](https://app.greptile.com/scubed/-/custom-context/knowledge-base/sakibsadmanshajib/hive/-/docs/edge-api-security-boundary.md) <!-- /greptile_comment --> --- ## Buglog entry, third (added 2026-08-23) Alongside the two above. To be appended to `.wolf/buglog.jsonl` on `main` in the same buglog-only pull request after this merges, per the openwolf protocol. Not appended on this branch: a branch that appends a line conflicts with every other branch that appended one, and an unmergeable pull request gets no CI run at all (issue #873). ```json {"id":"owui-unwrap-carrier-strip-unobservable","date":"2026-08-23","title":"Header-carrier strip asserted on a path no test could observe","error_message":"TestOWUIUnwrap_HeaderCarrierPresentButBlank_StrippedAndRejected asserted only the Rejected half of its own name; the Stripped half was unobservable because next never runs on a 401","root_cause":"The strip was applied to a clone of the request, so the only observation point was the downstream handler. Every rejection branch answers without calling next, so nothing on those branches could be checked, and an outer middleware still holding the pre-clone pointer would also have kept seeing a live per-user token. Moving Header.Del into the forwarding branches alone would have left the whole test file green.","fix":"Strip the carrier from the inbound request in place instead of from a clone, making the invariant one fact rather than one fact per branch, and assert in both rejection tests that the header is gone from the request the middleware was handed. Safe because net/http never re-reads request headers after the handler returns and the header is ours alone.","verification":"Green confirmed on unmodified code first. The production strip was then narrowed to fire only for a usable carrier, which is the exact regression described; all four blank sub-cases and the over-long case went red naming the new assertion, and the pass-through case went red too. Restoring the unconditional strip returned ./apps/edge-api/... to green.","tags":["auth","edge-api","owui-unwrap","test-quality","guard-cannot-fail","security","pr-951"]} ``` ## Visual proof: now posted, correcting the section above The section headed "Visual proof: not posted yet" is out of date and its own condition has been met. Images are now visible on this page, posted as permanent release assets through `scripts/post-pr-visual-proof.sh`, from a live signed-in capture against a complete local stack rather than a bundle or a CI artifact. The capture was taken after rebasing onto `main` at `82375d07f`, where #952 and #956 have landed, so it exercises the real post-merge state. Measured in the live DOM at both hops: zero password inputs, zero iframes, zero child browsing contexts, zero anchors to `/agent-workspace`. `GET /api/v1/agent/tasks` answered 200 through the running proxy chain, so the authenticated data path is exercised and not only the rendering. `/agent-workspace` still answers 307 to `/agent-workspace/tasks`, verified after the rebase and deliberately unchanged. Method, substrate and credential handling: `docs/proof/agents-native-live-2026-08-23/README.md`. --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…#1274, #1260, #1261, #1259) (#1296) Closes #1274. Closes #1260. Closes #1261. Closes #1259. ## Are these one root cause or four? Two families, not one, and not four independent fixes. **Family A, Go zero-value JSON encoding dropping values the Anthropic contract requires to be present (#1274 and #1260).** Both are the gateway emitting a valid Go zero value where Anthropic requires an explicit empty one, and in both cases the Go-side assertion a normal test would write cannot see the defect at all: `len(content) == 0` is true for a nil slice and an empty one alike, and a `Text` field set to `""` looks correct right up until `encoding/json` runs. The mechanism differs (`omitempty` on a string in one, a nil slice with no tag in the other), so it is one class with two instances rather than a single line to change. **Family B, an auth model that only ever covered POST /v1/messages (#1261 and #1259).** Every other route on the Anthropic surface was left to inherit an auth path built for a different dialect. `count_tokens` is the only route here that does not delegate to the chat chain, so it never gained the API-key authority that chain provides, and `GET /v1/models` was never given the leaf `APIKeyNormalizer` that `/v1/messages` has carried since #954. The two share a cause but not a line of code. #1259 additionally carries a response-shape defect that belongs to neither family: the route answered an Anthropic client with the OpenAI list envelope and the OpenAI error envelope. Fixed here in the same pass, since fixing only the auth half would have moved the failure one step later. ## Ground truth API shapes were verified against the live specification on 2026-08-28, not from model memory, per the repository rule. - `https://docs.claude.com/en/docs/build-with-claude/streaming` shows `content_block_start` as `{"type":"text","text":""}` for a text block and `{"type":"tool_use","id":...,"name":...,"input":{}}` for a tool_use block, in every one of its five worked examples. - `https://docs.claude.com/en/api/models-list` shows the list envelope as `data` (entries of `type`/`id`/`display_name`/`created_at`) plus `has_more`, `first_id` and `last_id`, the last two typed "string or null". - The SDK's own accumulator (`anthropic/lib/streaming/_messages.py`, `accumulate_event`) was read directly to confirm the failure mechanism: it does `content.text += event.delta.text` for a `text_delta`, which is the `None += str` crash, and for an `input_json_delta` it **assigns** `content.input` from its own `jiter` buffer rather than reading the block-start value. That is why the missing `text` was fatal and the missing `input` was not, and it is why this PR treats the two differently in its claims. `openwolf bug search anthropic` and `openwolf bug search omitempty` both returned no prior entries. ## What changed | Issue | Change | | --- | --- | | #1274 | `StreamContentBlock.Text` is now `*string` so the empty string serializes, following the precedent `StreamEvent.Index` already set in this same struct for the identical reason. `Input json.RawMessage` added and set to `{}` on a tool_use block start. | | #1260 | `FromOAIResponse` seeds `Content` with an empty non-nil slice at construction and appends into it, which covers the early `len(resp.Choices) == 0` return as well as the normal path. | | #1261 | `anthropic.Deps.AuthorizeAPIKey` added; `count_tokens` accepts a JWT session principal or an API-key principal. Nil leaves it session-only and fail-closed. The refusal is the authorizer's own verdict run through the shared `WriteAuthFailure` status mapping and then reshaped into the Anthropic envelope, so no status logic is duplicated. | | #1259 | `GET /v1/models` is registered through a named `modelsHandler` applying `anthropic.APIKeyNormalizer` at the leaf and `anthropic.ModelsCompat` for shape. | ### Why the leaf normalizer on /v1/models is not redundant `authSelectorMiddleware` already wraps every `/v1/*` route in `APIKeyNormalizer` (#954), but only when JWT auth is wired: `handler = authSelectorMiddleware(jwtMW, handler)` sits behind `if jwtMW != nil`. On a deployment where Supabase JWT config is absent, edge-api logs "JWT auth wiring skipped" and mounts no selector at all, so nothing normalizes `x-api-key` and `handleModels` reads an empty `Authorization` header. `/v1/messages` kept working there purely because it carries its own leaf wrapper. That asymmetry is exactly the reported symptom in #1259 (same key, same connection, `/v1/messages` fine and `/v1/models` 401) and it is why the fix is a leaf wrapper rather than a change to the selector. ### Provider-blind invariant The `ModelsCompat` translation reads `id`, `created` and `name` and writes `type`, `id`, `display_name` and `created_at`. `owned_by` and every other OpenAI field are dropped rather than passed through, so this strictly reduces what reaches an Anthropic client. `TestModelsCompat_AnthropicClientGetsTheAnthropicListShape` asserts `owned_by` is absent from the body. Nothing on any path touched here emits a provider name, an upstream id or a `system_fingerprint`; the existing `TestHandleModelsDoesNotLeakProviderNames` and `TestSSETranslator_MessageID_NeverLeaksUpstreamID` guards are unchanged and still pass. ## Mutation check Every added test was proved load-bearing by reverting the fix it guards and observing it fail for the right reason, then restoring and observing it pass. Full transcript, six mutations, all red, then green on restore: | # | Mutation applied | Test | Result | | --- | --- | --- | --- | | 1 | `StreamContentBlock.Text` back to a plain `string`, call site back to `Text: ""` | `TestSSETranslator_TextBlockStartCarriesExplicitEmptyText` | FAIL: `text content_block_start omits the "text" key entirely ... map[type:text]` | | 2 | Delete `Input: json.RawMessage("{}")` from the tool_use block start | `TestSSETranslator_ToolUseBlockStartCarriesExplicitEmptyInput` | FAIL: `tool_use content_block_start omits the "input" key: map[id:call_1 name:get_weather type:tool_use]` | | 3 | Delete the `Content: []ResponseBlock{}` seed | `TestFromOAIResponse_EmptyCompletionSerializesContentAsArray` | FAIL on both subtests: `got {... "content":null ...}` | | 4 | `if h.deps.AuthorizeAPIKey == nil` forced to `if true`, i.e. count_tokens back to session-only | `TestHandler_CountTokens_AcceptsAPIKeyPrincipal`, `..._RejectedAPIKeyKeepsTheAuthorizersOwnRefusal` | FAIL: `want 200 got 401 body={"type":"error","error":{"type":"authentication_error","message":"missing user"}}`, and `error.message: want the authorizer's own refusal got missing user` | | 5 | `if !IsAnthropicClient(r)` forced to `if true`, i.e. never re-shape | `TestModelsCompat_*` (3 tests) | FAIL: OpenAI body passed through, `first_id`/`last_id` nil, `envelope: want top-level type=error got <nil>` | | 6 | `modelsHandler` unwrapped to `return handleModels(client, authorizer)` | `TestModelsHandlerServesAnAnthropicSDKClient` | FAIL: `x-api-key caller: want 200 got 401: {"error":{"message":"You didn't provide an API key...` | After restoring all six, `go test ./apps/edge-api/internal/anthropic/... ./apps/edge-api/cmd/server/...` is `ok` for both packages, and the wider `go test ./apps/edge-api/... -count=1 -short` passes with no failures. `gofmt -l` reports none of the touched files. Four of the added tests are deliberately non-regression guards rather than defect guards, and do not go red under any mutation above. The count said two on first posting and was corrected in review; the table is the record anyone reads later, so it should say what is actually true. - `TestModelsCompat_OpenAIClientIsUntouched` and `TestModelsHandlerKeepsTheOpenAIShapeForOpenAIClients` exist to fail if a later change starts re-shaping unconditionally and empties Open WebUI's model picker. Both compare byte-identical bodies, so the inverse mutation of always reshaping does turn them red. - `TestHandler_CountTokens_WithoutAnAPIKeyAuthorityFailsClosed` stays green under mutation 4, since a handler wired without an API-key authority writes 401 either way. It pins the fail-closed default and goes red if someone makes a nil authority permissive. - `TestHandler_CountTokens_SessionPrincipalDoesNotConsultTheAPIKeyAuthority` also stays green under mutation 4, since the session branch returns before the nil check is reached. It goes red if the two principal checks are ever reordered. `TestIsAnthropicClient` is a unit test of a new pure function that none of the six mutations touch, so it is neither a defect guard nor a non-regression guard for anything in the table. ## Review round two Three findings from review, all fixed in `25b2820`, all in `apps/edge-api/internal/anthropic`. **Refusal retry headers were being dropped.** `apierrors.WriteAuthFailure` is the shared source of truth precisely so a retryable 429 is never collapsed into a non-retryable refusal, and it delivers the retryable part through headers rather than the body: `retry-after` plus the four `x-ratelimit-*` values on a real 429, and `retry-after` on both the degraded-limiter branch and the `upstream_unavailable` branch. Recording the refusal in a `headerlessRecorder` and reshaping only its body threw all of that away. `count_tokens` lost it latently, and `GET /v1/models` lost it live through `authorizeAliasRequest`, which also left the two client shapes disagreeing about retry metadata for the same refusal on the same route. `headerlessRecorder.reshapeInto` now carries the headers over and reshapes in one call, so a call site cannot take one half of a refusal without the other. `Content-Type` and `Content-Length` are deliberately not carried: the reshaped body is a different envelope of a different length, and a stale `Content-Length` would truncate it on the wire. Everything that is forwarded is retry metadata and carries no provider identity, so the provider-blind invariant is unchanged. **`GET /v1/models` served two representations for one URL without declaring it.** Nothing on the route sets `Cache-Control` and the route needs a credential, so no correct cache stores it today and no live exploit exists. The declaration is what keeps a later edge cache in front of Caddy, or an intermediary keying on URL alone, from handing an Anthropic-shaped body to Open WebUI and emptying its model picker, which is the regression `TestModelsCompat_OpenAIClientIsUntouched` exists to prevent. `Vary: anthropic-version, x-api-key` is now set on both branches, before either writes. **`Finish` could emit `message_delta` before any `message_start`.** `Translate` calls `Finish` unconditionally and `Finish` did not check `t.started`, so an upstream stream that yielded no parseable chunk at all produced a terminal pair with nothing before it, and the SDK accumulator raises an unexpected-event-order error when a `message_delta` arrives while its snapshot is still nil. Review suggested a follow-up issue for this rather than widening the PR; the fix turned out to be three lines in a file this PR already changes, which is smaller than the issue describing it, so it is taken here. `FeedLine` and `Finish` now share one `ensureStarted`, which emits `message_start` exactly once per stream whichever of them reaches it first. Four more mutations, each applied alone, each restored after: | # | Mutation applied | Test | Result | | --- | --- | --- | --- | | 7 | `reshapeInto`'s header loop forced to skip every key | `TestHandler_CountTokens_RefusalCarriesTheAuthorizersRetryHeaders`, `TestHandler_CountTokens_UpstreamUnavailableCarriesRetryAfter`, `TestModelsCompat_RefusalCarriesTheRetryHeaders` | FAIL: `retry-after: want 30 got ""`, `retry-after: want 5 got ""`, and the three `x-ratelimit-*` assertions on both routes | | 8 | `Content-Length` removed from the skip list, so the delegated body's length is forwarded | `TestModelsCompat_RefusalCarriesTheRetryHeaders` | FAIL: `content-length: the delegated body length must not describe the reshaped one, got "4096"` | | 9 | The `Vary` line deleted | `TestModelsCompat_DeclaresVary` | FAIL on both subtests: `vary: want both request headers that select the representation, got ""` | | 10 | `ensureStarted` removed from `Finish` | `TestSSETranslator_EmptyStreamStillOpensTheMessage` | FAIL: `first event: want message_start got message_delta` | `TestSSETranslator_FinishAfterAStartedStreamDoesNotRepeatMessageStart` is a fifth non-regression guard: it stays green under all ten mutations and exists so that opening the message from `Finish` can never add a second `message_start` to a stream that already carried one. ### Adversarial review of the round-two fixes The round-two commit went back through the review pipeline before this was called done. Two findings, both fixed in `f0e4c56`. **The `Vary` declaration used `Set` where it should use `Add`.** Nothing in edge-api declares a `Vary` on this route today (`grep` finds exactly one writer, the new line itself), so overwriting was harmless in the current tree and wrong the moment an outer middleware declares one of its own, a CORS layer setting `Vary: Origin` being the obvious case. The wrapper is now additive. **The 2xx branch of `ModelsCompat` dropped every header the delegated handler set.** That path re-encodes the body rather than reshaping it, which is exactly why the carry-over the refusal path had just gained was easy to leave out of it. A header set alongside a success is as much part of that response as the body is, and this route is where a 2xx rate-limit budget would surface. The header copy is now its own method on the recorder, `copyHeadersTo`, and both branches call it. | # | Mutation applied | Test | Result | | --- | --- | --- | --- | | 11 | `Vary` back to `Set` | `TestModelsCompat_VaryDoesNotClobberAnExistingDeclaration` | FAIL: `vary: want "Origin" preserved, got "anthropic-version, x-api-key"` | | 12 | `rec.copyHeadersTo(w)` deleted from the success branch | `TestModelsCompat_SuccessCarriesTheDelegatedHeaders` | FAIL: `x-ratelimit-remaining-requests: want 42 got ""` | ### On `display_name` and the provider-blind edge Review confirmed independently that the `ModelsCompat` translation is provider-blind (it reads `id`, `created` and `name` only, so `owned_by` and `description` cannot ride along) and flagged the remaining edge: `display_name` is `model_aliases.display_name`, free text, and one seeded row already reads `Openrouter Auto (Task Aware)`. That row is `visibility='internal'` today, so the `visibility IN ('public', 'preview')` predicate keeps it off both list paths, but the migration that added it describes a later flip to public as a one-line follow-up. Nothing here closes #1284, and this PR never claimed to: the OpenAI branch is unchanged and still ships the `hive-stt` and `hive-tts` descriptions that #1284 reports. The `display_name` edge is recorded as a comment on #1284 rather than a new issue, since it is the same route and the same family, and the guard it asks for (assert no listed alias's `display_name` or `summary` matches a known provider name) belongs where the catalog is listed, not in this translation. ## Effect on #1278 This should flip all four `xfail` markers in `packages/sdk-tests/python/tests/test_anthropic_messages.py` to real assertions. Not touched here, since that branch is not mine to edit. - `test_streaming_event_sequence_integrity` (#1274) - `test_tool_choice_none_forbids_tool_use` (#1260) - `test_count_tokens` (#1261) - `test_models_list` (#1259) One caveat on the last one. `client.models.list()` will now authenticate and return the Anthropic list envelope, but the SDK's typed page model may expect fields this gateway has no source for. If that assertion still fails after this merges, it is a narrower follow-up about specific fields, not the auth and envelope defects #1259 reported. ## Scope deliberately not taken #1260 also observes that a zero-content streaming turn emits `message_start`, `message_delta`, `message_stop` with no `content_block_start`/`stop` pair. That is left alone: the streaming path already emits `"content":[]` in `message_start`, so an SDK folding the stream back into a message gets an empty array rather than null, and Anthropic's own specification does not require a content block on a turn that produced no content. The reported crash is in the non-streaming builder, which is what this changes. The neighbouring case where `message_start` never fires at all is a different defect, was not what #1260 reported, and is now fixed here rather than deferred. See Review round two above. ## Test plan - [x] `go test ./apps/edge-api/... -count=1 -short` through the toolchain container, green, re-run after the review-round-two fixes - [x] `gofmt -l apps/edge-api` clean for every touched file - [x] Twelve-way mutation check, all three tables above - [ ] CI required checks green - [ ] `packages/sdk-tests` Anthropic conformance suite re-run against a live stack once #1278 merges, to confirm the four xfails flip ## Buglog entry ```json {"id":"bug-2026-08-28-anthropic-sdk-wire-conformance","date":"2026-08-28","title":"Anthropic surface broke the real SDK four ways: content_block_start dropped text, empty completions serialized content null, count_tokens 401d API keys, /v1/models rejected x-api-key and answered OpenAI-shaped","error_message":"TypeError: unsupported operand type(s) for +=: 'NoneType' and 'str' in anthropic/lib/streaming/_messages.py accumulate_event; TypeError: 'NoneType' object is not iterable on msg.content; anthropic.AuthenticationError 401 {\"type\":\"error\",\"error\":{\"type\":\"authentication_error\",\"message\":\"missing user\"}} on count_tokens; anthropic.AuthenticationError 401 invalid_api_key on models.list()","root_cause":"Two families. (1) Go zero-value JSON encoding dropped values the Anthropic wire contract requires to be explicitly present: omitempty on StreamContentBlock.Text dropped the required \"text\":\"\" from every text content_block_start, and a nil []ResponseBlock in FromOAIResponse marshaled to null instead of []. Neither is visible to a Go-side assertion, since len() cannot distinguish nil from empty and a \"\" field looks correct until encoding/json runs. (2) The Anthropic surface's auth model only ever covered POST /v1/messages: count_tokens is the only route that does not delegate to the chat chain so it never gained an API-key authority and recognised session principals only, and GET /v1/models never got the leaf APIKeyNormalizer /v1/messages carries, which matters because the global one in authSelectorMiddleware only exists when jwtMW is non-nil.","fix":"StreamContentBlock.Text became *string and gained Input json.RawMessage set to {} on tool_use starts, matching the live streaming spec. FromOAIResponse seeds Content with an empty non-nil slice at construction so both exits are covered. anthropic.Deps.AuthorizeAPIKey added and count_tokens accepts either principal, fail-closed when nil, refusals reshaped from the shared WriteAuthFailure mapping. GET /v1/models registered through modelsHandler applying APIKeyNormalizer at the leaf plus a ModelsCompat wrapper that answers Anthropic-shaped callers with the Anthropic list and error envelopes while leaving OpenAI callers byte-identical.","tags":["anthropic","sdk-conformance","edge-api","json-encoding","omitempty","auth","streaming","provider-blind"],"issues":[1274,1260,1261,1259],"pr":"fix/anthropic-stream-content-block-start"} ```
Summary
A real Anthropic SDK client constructed the documented way (
Anthropic(api_key=...)) sends its credential onx-api-key, neverAuthorization. Confirmed directly from the installed SDK's own source:_api_key_authreturns{"X-Api-Key": api_key};Authorizationonly appears when the caller passesauth_token=instead.apps/edge-api/internal/anthropic/handler.goalready shipsAPIKeyNormalizer, unit-tested in isolation, which rewritesx-api-keytoAuthorization: Bearer. The bug was entirely in where it was wired:main.goapplied it only at the mux leaf (mux.Handle("/v1/messages", anthropic.APIKeyNormalizer(anthropicHandler))), butauthSelectorMiddleware(which decides API-key vs JWT path by inspectingAuthorization) wraps the entire mux and runs first. Everyx-api-key-only request had noAuthorizationheader at the point the selector decided, fell through to the JWT path unconditionally, and 401'd with"missing bearer"regardless of key validity.Reproduced live against the deployed gateway (
https://api-hive.scubed.co) with the realanthropicPython SDK (0.122.0): the defaultapi_key=construction always 401s withmissing bearer; the same call withauth_token=(forcing anAuthorizationheader directly) reaches the real API-key path instead. This is the empirical confirmation that the owner's literal ask — point a real SDK at Hive by changing onlybase_urland the key — could not have worked on any key before this fix.Fix
authSelectorMiddleware(apps/edge-api/cmd/server/main.go): applyanthropic.APIKeyNormalizeraround the selector itself, sox-api-keyis normalized beforeauth.SelectorinspectsAuthorization, for every/v1/*route. No-op whereverAuthorizationis already set.apps/edge-api/internal/anthropic/errors.go: reshapes every/v1/messageserror this surface owns (its own validation refusals, and delegated-chain errors it forwards) into Anthropic's real error envelope ({"type":"error","error":{"type","message"}}) instead of the OpenAI-shaped one used unconditionally before. Verified live againstapi.anthropic.comand the installed SDK's_exceptions.py: the SDK's exception class is driven by HTTP status, not body shape, so this never crashed a client, buterr.typeread values outside Anthropic's documented enum anderr.request_idwas always empty. Reuses the existing provider-blind sanitizer verbatim (through a small localhttp.ResponseWritercapture) rather than duplicating its regexes; no sanitization logic changed.docs/anthropic-sdk-integration.md: exact base URL, auth header, copy-pasteable Python/TypeScript snippets, an OpenCode-style config, the model aliases that actually answer/v1/messages(hive-default,hive-fast,hive-auto, live fromGET /catalog/models), and every known limitation found during this pass, each one verified rather than assumed.Documented, not fixed here (filed as follow-up, out of this pass's scope)
budgetGate's 429) still use the generic envelope: fixing those touches sharedinternal/middleware/internal/errorscode used by every other surface, a materially larger blast radius, deliberately deferred.middleware.CompatHeaders()(global, unconditional) stampsx-request-id/openai-version/openai-processing-mson every response including/v1/messages— confirmed live. Harmless to the SDK (unknown headers ignored) but means.request_idstays empty and leaks OpenAI-family headers onto an Anthropic-shaped endpoint.top_k, extended thinking (thinking/redacted_thinkingblocks), prompt caching (cache_control),container/server tools,service_tier,inference_geo, and stop-sequence identification in the response are not implemented. Neither of Hive's two configured providers (OpenRouter, Groq) supports Anthropic's native versions of these, so these are provider-capability gaps, not oversights.Test plan
TestAuthSelectorMiddlewareAcceptsXAPIKeyHeader— RED before the fix (confirmed), GREEN afterTestAuthSelectorMiddlewareStillRoutesJWTBearerToJWTPath— regression guard, a non-hk_ session bearer still routes to the JWT path onlyTestHandler_ValidationErrorUsesAnthropicEnvelope,TestHandler_DownstreamErrorUsesAnthropicEnvelope,TestHandler_UpstreamErrorUsesAnthropicEnvelope— new envelope shape assertionsapps/edge-api/internal/anthropic/...suite passes unchanged (message/status assertions matched verbatim; only the envelope changed)apps/edge-api/...suite green (go test ./apps/edge-api/... -count=1 -short)go vet ./apps/edge-api/...cleanBuglog entry
{"date":"2026-08-17","error_message":"real Anthropic SDK client 401s with 'missing bearer' regardless of API key validity when pointed at Hive with only base_url and api_key set","root_cause":"authSelectorMiddleware inspects only the Authorization header and runs outside the mux; anthropic.APIKeyNormalizer (x-api-key -> Authorization: Bearer) was wired only at the mux leaf for /v1/messages, so it never ran before the selector's routing decision","fix":"apply anthropic.APIKeyNormalizer around the selector itself in authSelectorMiddleware (apps/edge-api/cmd/server/main.go), before auth.Selector inspects Authorization","tags":["anthropic","auth","edge-api","x-api-key","routing"]}Boundary honored
No billing, reservation, ledger, or metering code touched (a separate agent owns that concurrently). No API key, token, or credential value printed anywhere in this PR or its tests (all test keys are obvious literals:
hk_test_key,hk_bogus_test_key_...).