Skip to content

fix: normalize x-api-key before the JWT/API-key auth selector - #954

Merged
sakibsadmanshajib merged 1 commit into
mainfrom
fix/anthropic-x-api-key-auth-routing
Aug 17, 2026
Merged

sakibsadmanshajib merged 1 commit into
mainfrom
fix/anthropic-x-api-key-auth-routing

Conversation

@sakibsadmanshajib

Copy link
Copy Markdown
Owner

Summary

A real Anthropic SDK client constructed the documented way (Anthropic(api_key=...)) sends its credential on x-api-key, never Authorization. Confirmed directly from the installed SDK's own source: _api_key_auth returns {"X-Api-Key": api_key}; Authorization only appears when the caller passes auth_token= instead.

apps/edge-api/internal/anthropic/handler.go already ships APIKeyNormalizer, unit-tested in isolation, which rewrites x-api-key to Authorization: Bearer. The bug was entirely in where it was wired: main.go applied it only at the mux leaf (mux.Handle("/v1/messages", anthropic.APIKeyNormalizer(anthropicHandler))), but authSelectorMiddleware (which decides API-key vs JWT path by inspecting Authorization) wraps the entire mux and runs first. Every x-api-key-only request had no Authorization header at the point the selector decided, fell through to the JWT path unconditionally, and 401'd with "missing bearer" regardless of key validity.

Reproduced live against the deployed gateway (https://api-hive.scubed.co) with the real anthropic Python SDK (0.122.0): the default api_key= construction always 401s with missing bearer; the same call with auth_token= (forcing an Authorization header directly) reaches the real API-key path instead. This is the empirical confirmation that the owner's literal ask — point a real SDK at Hive by changing only base_url and the key — could not have worked on any key before this fix.

Fix

  • authSelectorMiddleware (apps/edge-api/cmd/server/main.go): apply anthropic.APIKeyNormalizer around the selector itself, so x-api-key is normalized before auth.Selector inspects Authorization, for every /v1/* route. No-op wherever Authorization is already set.
  • New apps/edge-api/internal/anthropic/errors.go: reshapes every /v1/messages error this surface owns (its own validation refusals, and delegated-chain errors it forwards) into Anthropic's real error envelope ({"type":"error","error":{"type","message"}}) instead of the OpenAI-shaped one used unconditionally before. Verified live against api.anthropic.com and the installed SDK's _exceptions.py: the SDK's exception class is driven by HTTP status, not body shape, so this never crashed a client, but err.type read values outside Anthropic's documented enum and err.request_id was always empty. Reuses the existing provider-blind sanitizer verbatim (through a small local http.ResponseWriter capture) rather than duplicating its regexes; no sanitization logic changed.
  • docs/anthropic-sdk-integration.md: exact base URL, auth header, copy-pasteable Python/TypeScript snippets, an OpenCode-style config, the model aliases that actually answer /v1/messages (hive-default, hive-fast, hive-auto, live from GET /catalog/models), and every known limitation found during this pass, each one verified rather than assumed.

Documented, not fixed here (filed as follow-up, out of this pass's scope)

  • Errors emitted before the request reaches the anthropic package at all (the selector's own 401, budgetGate's 429) still use the generic envelope: fixing those touches shared internal/middleware/internal/errors code used by every other surface, a materially larger blast radius, deliberately deferred.
  • middleware.CompatHeaders() (global, unconditional) stamps x-request-id/openai-version/openai-processing-ms on every response including /v1/messages — confirmed live. Harmless to the SDK (unknown headers ignored) but means .request_id stays empty and leaks OpenAI-family headers onto an Anthropic-shaped endpoint.
  • top_k, extended thinking (thinking/redacted_thinking blocks), prompt caching (cache_control), container/server tools, service_tier, inference_geo, and stop-sequence identification in the response are not implemented. Neither of Hive's two configured providers (OpenRouter, Groq) supports Anthropic's native versions of these, so these are provider-capability gaps, not oversights.

Test plan

  • TestAuthSelectorMiddlewareAcceptsXAPIKeyHeader — RED before the fix (confirmed), GREEN after
  • TestAuthSelectorMiddlewareStillRoutesJWTBearerToJWTPath — regression guard, a non-hk_ session bearer still routes to the JWT path only
  • TestHandler_ValidationErrorUsesAnthropicEnvelope, TestHandler_DownstreamErrorUsesAnthropicEnvelope, TestHandler_UpstreamErrorUsesAnthropicEnvelope — new envelope shape assertions
  • Full existing apps/edge-api/internal/anthropic/... suite passes unchanged (message/status assertions matched verbatim; only the envelope changed)
  • Full apps/edge-api/... suite green (go test ./apps/edge-api/... -count=1 -short)
  • go vet ./apps/edge-api/... clean
  • Live reproduction of the bug and the intended fix behavior against the deployed box, with the real Anthropic Python SDK, both pre- and post-fix code paths (not deployed yet; deploy is the orchestrator's step)

Buglog entry

{"date":"2026-08-17","error_message":"real Anthropic SDK client 401s with 'missing bearer' regardless of API key validity when pointed at Hive with only base_url and api_key set","root_cause":"authSelectorMiddleware inspects only the Authorization header and runs outside the mux; anthropic.APIKeyNormalizer (x-api-key -> Authorization: Bearer) was wired only at the mux leaf for /v1/messages, so it never ran before the selector's routing decision","fix":"apply anthropic.APIKeyNormalizer around the selector itself in authSelectorMiddleware (apps/edge-api/cmd/server/main.go), before auth.Selector inspects Authorization","tags":["anthropic","auth","edge-api","x-api-key","routing"]}

Boundary honored

No billing, reservation, ledger, or metering code touched (a separate agent owns that concurrently). No API key, token, or credential value printed anywhere in this PR or its tests (all test keys are obvious literals: hk_test_key, hk_bogus_test_key_...).

A real Anthropic SDK client constructed the documented way
(Anthropic(api_key=...)) sends its credential on x-api-key, never
Authorization (confirmed from the installed SDK's own source:
_api_key_auth returns {"X-Api-Key": api_key}). anthropic.APIKeyNormalizer
already existed to rewrite that header to Authorization: Bearer, but was
wired only at the mux leaf (mux.Handle("/v1/messages",
anthropic.APIKeyNormalizer(...))), inside authSelectorMiddleware, which
wraps the entire mux and inspects Authorization before any mux-level
handler runs. Every x-api-key-only request therefore had no Authorization
header at the point auth.Selector chose a path, fell through to the JWT
handler unconditionally, and 401'd with "missing bearer" regardless of key
validity.

Reproduced live against the deployed gateway with the real anthropic
Python SDK (0.122.0): the default api_key= construction always 401s with
"missing bearer"; the same call with auth_token= (forcing an
Authorization header directly) reaches the real API-key path instead.

Fix: apply anthropic.APIKeyNormalizer around the selector in
authSelectorMiddleware, so x-api-key is normalized before auth.Selector
ever inspects the header, for every /v1/* route. No-op wherever
Authorization is already set.

Also reshapes every /v1/messages error this surface owns (its own
validation refusals, and delegated-chain errors it forwards) into
Anthropic's real error envelope ({"type":"error","error":{"type",
"message"}}) instead of the OpenAI-shaped one it used unconditionally
before. Verified live against api.anthropic.com and the installed SDK's
_exceptions.py: the SDK's exception class is driven by HTTP status, not
body shape, so this did not previously crash a client, but err.type read
values outside Anthropic's documented enum and err.request_id was always
empty. Reuses the existing provider-blind sanitizer verbatim (through a
small local http.ResponseWriter capture) rather than duplicating its
regexes, so no sanitization logic changes.

Pre-dispatch middleware errors (the selector's own 401, budgetGate's 429)
and the global CompatHeaders OpenAI-style header injection on this
surface are left as documented gaps: fixing those touches shared
internal/middleware and internal/errors code used by every other
surface, out of this pass's scope.

Adds docs/anthropic-sdk-integration.md: the exact base URL, auth header,
copy-pasteable Python/TypeScript snippets, an OpenCode-style config, the
model aliases that actually work (hive-default, hive-fast, hive-auto,
live from GET /catalog/models), and every known limitation found during
this pass, each one verified rather than assumed.

Buglog entry (carried in the PR body per repo convention, to be appended
to main via a follow-up buglog-only PR):
{"date":"2026-08-17","error_message":"real Anthropic SDK client 401s with 'missing bearer' regardless of API key validity when pointed at Hive with only base_url and api_key set","root_cause":"authSelectorMiddleware inspects only the Authorization header and runs outside the mux; anthropic.APIKeyNormalizer (x-api-key -> Authorization: Bearer) was wired only at the mux leaf for /v1/messages, so it never ran before the selector's routing decision","fix":"apply anthropic.APIKeyNormalizer around the selector itself in authSelectorMiddleware (apps/edge-api/cmd/server/main.go), before auth.Selector inspects Authorization","tags":["anthropic","auth","edge-api","x-api-key","routing"]}

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

@cursor

cursor Bot commented Aug 17, 2026

Copy link
Copy Markdown

Bugbot is not enabled for your account, so this pull request was not reviewed.

Enable Bugbot in the Cursor dashboard to get automatic reviews on future PRs.

@coderabbitai

coderabbitai Bot commented Aug 17, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@sakibsadmanshajib, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 29 minutes

Limit details: You’ve used all 1 included review currently available under your plan.

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 7f613d3e-c80f-44ca-b0e0-d9e3b6922e1e

📥 Commits

Reviewing files that changed from the base of the PR and between 1dc67ee and 8acd48d.

📒 Files selected for processing (7)
  • apps/edge-api/cmd/server/main.go
  • apps/edge-api/cmd/server/main_test.go
  • apps/edge-api/internal/anthropic/errors.go
  • apps/edge-api/internal/anthropic/handler.go
  • apps/edge-api/internal/anthropic/handler_test.go
  • apps/edge-api/internal/anthropic/translate_writer.go
  • docs/anthropic-sdk-integration.md

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Comment thread apps/edge-api/internal/anthropic/errors.go
@sakibsadmanshajib

Copy link
Copy Markdown
Owner Author

Adversarial review summary (dynamic selector, D-038)

This PR touches an auth-routing boundary (authSelectorMiddleware) and an
input-parsing/error-serialization boundary (/v1/messages error envelope),
so the mandatory security review applies.

CodeRabbit CLI — ran (coderabbit review --base main --agent). 1 finding
(major, errors.go: suggested removing the Anthropic-specific envelope in
favor of the shared OpenAI-shaped one). Addressed with a reasoned rebuttal
inline on the finding and resolved: reverting would reintroduce the exact
defect this PR fixes, confirmed live against api.anthropic.com and the
installed Anthropic SDK source that the two envelopes are mutually exclusive
by construction, and the change is scoped to only the errors this surface
owns, reusing the shared sanitizer verbatim rather than duplicating it. See
the resolved thread on errors.go.

ecc:code-review — SKIPPED. Structural capability gap: this task was
dispatched to a subagent whose toolset is Read/Write/Edit/Bash only, with no
Skill-invocation tool available to run a plugin skill. Reporting this
precisely per the Gate Compliance rule rather than attempting a workaround.

/codex:adversarial-review — SKIPPED. Attempted (codex review --base main); the CLI reported ERROR: You've hit your usage limit... try again at Sep 10th, 2026. A genuinely unavailable stream, not a clean pass — recorded
as SKIPPED per the pipeline's fallback rule, not treated as passing review.

Plain adversarial pass (ticket-first) and manual security pass — run
directly, in place of the two tool-gated streams above, since the ticket
itself (owner: point a real Anthropic SDK at Hive by changing only base_url
and the key) is squarely a correctness-and-security question this pass could
still answer without those tools:

  • Auth-bypass check: anthropic.APIKeyNormalizer only rewrites x-api-key to
    Authorization: Bearer when Authorization is not already set
    (normalizeAPIKeyHeader), so an existing credential can never be
    overridden, and auth.Selector's own anti-smuggling check (any Authorization
    value containing a comma or whitespace is refused as a possible
    credential-smuggling attempt, selector.go:33-40) still runs downstream of
    the normalizer unchanged, so a multi-value or malformed x-api-key still
    gets refused rather than routed. No new bypass.
  • Broadening the normalizer from the /v1/messages mux leaf to every /v1/*
    route (via authSelectorMiddleware) grants no new capability to an invalid
    credential; it only lets an already-valid hk_ key authenticate via an
    additional, standard header name on more routes, matching
    APIKeyNormalizer's own pre-existing doc comment ("normalising Anthropic
    x-api-key credentials... before dispatching", never scoped to one route in
    its own text).
  • Information disclosure check: reshapeToAnthropicError extracts only
    .error.message/.error.code from an already-sanitized body (either the
    delegated chain's own provider-blind refusal, or the output of calling the
    existing apierr.WriteProviderBlindUpstreamError sanitizer verbatim into a
    local, in-memory http.ResponseWriter capture) and never introduces a new
    source of raw upstream text. The error.type field is remapped from a
    fixed, hardcoded status-to-enum table (anthropicErrorType), with no
    attacker-controlled input reaching it. No new leak.
  • Pre-existing, unchanged precedence question (not introduced by this PR,
    noted for completeness): if a caller sends both x-api-key and
    Authorization, the existing Authorization always wins and x-api-key is
    silently ignored. This was true before this PR (already the behavior at the
    old mux-leaf wiring) and is unchanged by moving where it fires.
  • No credential, token, or key value appears anywhere in the diff; all test
    literals are obviously fake (hk_test_key, hk_bogus_test_key_...,
    hk_live_test_key).

Verdict: no security or correctness findings beyond the one CodeRabbit
finding above, which is resolved.

@sakibsadmanshajib
sakibsadmanshajib merged commit 8cb172d into main Aug 17, 2026
18 of 19 checks passed
@github-actions
github-actions Bot deleted the fix/anthropic-x-api-key-auth-routing branch August 17, 2026 20:31
sakibsadmanshajib added a commit that referenced this pull request Aug 18, 2026
)

## Summary

This PR is the output of live end-to-end verification of PR #954 (the
x-api-key auth fix): a real Anthropic Python SDK client, a fresh API key
minted through the actual developer console
(`POST /api/v1/accounts/current/api-keys`, the same call the console UI
button makes), run against the deployed gateway.

PR #954 itself is confirmed working: a bogus `x-api-key` now correctly
answers `401 authentication_error` (not the old `missing bearer`), and a
valid key reaches the model on `hive-default` and `hive-auto` for
non-streaming completions and a full tool-use round trip.

That same live pass found a second, pre-existing bug that #954 did not
touch and did not introduce: **every streamed response's first content
block omitted the `index` key entirely.** `StreamEvent.Index` was a
plain
`int` tagged `json:"index,omitempty"`. Go's encoder drops a zero value
under
`omitempty`, and index 0 is the common case (the first, and usually
only,
content block in a plain-text or single-tool-call reply), so the wire
event
carried no `"index"` key at all. The real Anthropic Python SDK's own
streaming accumulator (`anthropic/lib/streaming/_messages.py:
accumulate_event`) indexes `current_snapshot.content[event.index]` and
crashed with `TypeError: list indices must be integers or slices, not
NoneType` on literally the first event of any streamed call — reproduced
live 2026-08-18 against `api-hive.scubed.co` with `anthropic==0.122.0`.

No existing test caught this. The existing assertions decode into
`map[string]interface{}` and read `blockStarts[0]["index"].(float64)`
via
the comma-ok pattern, which silently yields the same zero-value `0` on a
**missing** key as it does on a **present** key holding `0`. That is the
exact "camouflaged test" shape: green whether the bug is present or not.

## Fix

- `apps/edge-api/internal/anthropic/types.go`: `StreamEvent.Index` is
now
  `*int` with `omitempty` on the pointer rather than the value, so
`message_start`/`message_delta` (which never set `Index`, and which the
  real Anthropic protocol never puts an `index` field on) still carry no
  `"index"` key at all, while every `content_block_*` event gets an
  explicit, always-present index.
- `apps/edge-api/internal/anthropic/stream.go`: every
`content_block_start`,
  `content_block_delta`, and `content_block_stop` call site now passes
  `indexPtr(n)` explicitly, including `n=0`.
- New regression test
  `TestSSETranslator_FirstBlockIndexIsPresentAndZero` checks key
**presence** via `json.RawMessage` rather than a lenient map decode, so
it
fails the way the existing tests should have and now cannot silently
pass
  the way they did before.

## Docs correction

`docs/anthropic-sdk-integration.md` (written by a previous pass without
ever
exercising a real credential) is corrected against what this pass
actually
ran:

- Every example switched from `model="hive-fast"` to
`model="hive-default"`.
  **`hive-fast`'s live route currently 404s on both `/v1/messages` and
`/v1/chat/completions`**: it maps to Groq's `llama-3.1-8b-instant`,
which
  Groq now answers with `model_not_found` ("does not exist or you do not
have access to it"). Confirmed live 2026-08-18 on both surfaces. Filed
as
issue (see below), not fixed in this PR: it is a DB-driven routing
config
change
(`supabase/migrations/20260801_14_route_groq_fast_cheapest_model.sql`
  set this mapping), not a code change, and out of this PR's scope.
- Documents that `hive-fast`'s failure response leaks LiteLLM's internal
fallback-group bookkeeping (`Fallbacks=[{'hive-fast': ['hive-fast']},
...]`,
retry counts) verbatim into the customer-visible error message. No
literal
  provider name appears, so the existing provider-name sanitizer doesn't
  catch it, but it is still internal routing implementation detail that
  should never reach a caller.
- Documents a separate live finding, unrelated to #954: a console user
who
  owns more than one workspace (a real fixture account used for this
verification did) can have their session's default "current" account be
one with **no** `tenant_billing_accounts` mapping. A key minted from
that
default state looks completely normal (`200`, a real `hk_...` secret)
but
  answers every `/v1/messages` request with `403 permission_error` /
`account_not_provisioned`, with no warning anywhere in the mint flow.
The
  console does support selecting a different workspace via an
`hive_account_id` cookie (forwarded as `X-Hive-Account-ID`), but nothing
  in the UI surfaces which account is "current" before a key is minted.

## Two issues filed from this verification pass, not fixed here

1. `hive-fast` routes to a deprecated/unavailable Groq model, and the
resulting error leaks internal routing bookkeeping. Needs a DB routing
update (out of this PR's scope) plus a look at whether the
provider-blind
   sanitizer should also scrub LiteLLM fallback-group text, not just
   provider names.
2. A multi-workspace console user can mint a dead-on-arrival API key
from
   their default account with no warning. Needs either a provisioning
   backfill for orphaned accounts, a mint-time check that surfaces
`account_not_provisioned` before the key is created, or a UI affordance
   showing which workspace is "current."

PR #962 (README refresh, open, not yet merged) also uses
`model="hive-fast"`
in both its OpenAI-SDK and Anthropic-SDK examples; flagging there
directly
since it's about to publish the same broken example.

## Test plan

- [x] `TestSSETranslator_FirstBlockIndexIsPresentAndZero` — RED before
the
      fix (confirmed: all three content_block_* events showed `"index"`
      absent from the wire JSON), GREEN after
- [x] Full existing `apps/edge-api/internal/anthropic/...` suite passes
unchanged (`go test ./apps/edge-api/internal/anthropic/... -count=1 -v`)
- [x] Full `apps/edge-api/...` suite green
      (`go test ./apps/edge-api/... -count=1 -short`)
- [x] `go build ./apps/edge-api/...` and `go vet ./apps/edge-api/...`
clean
- [x] Live re-verification pending this PR's deploy: real SDK
`client.messages.stream()` against `hive-default` on the deployed box,
full event sequence checked in order, will be posted as a PR comment
      once `deploy-demo-box.yml` for the merge commit succeeds.

## Boundary honored

No billing, reservation, ledger, or metering code touched. No API key,
token, or credential value printed anywhere in this PR, its tests, or
its
commit message.
sakibsadmanshajib added a commit that referenced this pull request Aug 22, 2026
The Web E2E (full stack) gate was intermittently red on the console
spend-alerts and profile-completion specs, both timing out on
`getByRole('heading', ...)` / `page.fill` because the console's generic
error boundary ("Something went wrong on this page") had tripped instead
of the real page. Both failures passed on Playwright's retry, which is
what the suite's flake-rate gate exists to catch rather than paper over.

The Next.js server log for the failing run showed the actual thrown
error: "could not load your workspace" (control-plane's writeInternal
message for accounts.EnsureViewerContext failures), with a digest
matching the error boundary's displayed reference number.

Root cause: web-console's Server Components call getViewer() unmemoized
per component (layout.tsx and each page.tsx independently), so a
brand-new user's very first page load can fire two concurrent
GET /api/v1/viewer requests. Both see zero active memberships and both
call provisionDefaultWorkspace, which derives an identical slug from the
same viewer's display name and inserts it with no idempotency guard.
The loser hits accounts.slug's unique constraint, and CreateAccount
propagated that raw pgx error unhandled, surfacing as an opaque 500 on
whichever request lost the race. E2E mints a fresh test user per run, so
its first console page load reliably lands in this window; the failure
is timing-dependent, hence flaky rather than a hard break.

Fixes:
- accounts.ErrSlugTaken sentinel, returned by CreateAccount on a unique
  violation on slug (mirrors the existing CreateMembership /
  ErrAlreadyMember pattern).
- provisionDefaultWorkspace now tells the two situations a slug
  collision can mean apart: if the current viewer already holds an
  active membership after the collision, a concurrent request for this
  same viewer won the race and there is nothing left to do (recovered,
  no error). Otherwise this is a genuine collision between two
  different viewers whose display names happen to produce the same
  slug (buildSlug has no per-user uniqueifier), and it retries with a
  de-duplicated slug instead of failing that viewer's first sign-in.

Two new unit tests cover both branches directly against the accounts
service: TestEnsureDefaultAccount_SlugCollisionFromConcurrentSelf_Recovers
and TestEnsureDefaultAccount_SlugCollisionFromDifferentViewer_RetriesWithSuffix.

CI history check: this is a genuine, intermittent product bug, not a
selector-drift or fixture problem, and not caused by PR #954 or #955
(both merged while this check was still pending, per the standing
instruction to review that). The last full pass of this job on main was
1dc67ee (2026-08-17 16:56); #955's own pre-merge run hit this exact
race. The separate stale-selector failure on the OWUI e2e harness
("New chat" button vs `<a>`) is unrelated and already in flight on #951;
not touched here.

Buglog entry (to be appended to .wolf/buglog.jsonl on a buglog-only
branch cut from main, per the no-buglog-on-feature-branch rule):
{"error_message": "could not load your workspace (500) on GET /api/v1/viewer, surfacing as the web console's generic error boundary on the spend-alerts and billing-settings pages", "root_cause": "provisionDefaultWorkspace has no idempotency guard: two concurrent Server Component calls to getViewer() for the same brand-new user both attempt to create a personal workspace with the same deterministic slug, and CreateAccount let the resulting pg unique-violation on accounts.slug escape as a raw 500 instead of recovering the race", "fix": "translate the unique violation into accounts.ErrSlugTaken and have provisionDefaultWorkspace re-check the viewer's own memberships on collision: recover silently if a concurrent request for the same viewer already won, otherwise retry once with a de-duplicated slug for the genuine different-viewer collision case", "tags": ["accounts", "web-e2e", "race-condition", "ci-gate", "flaky-test"]}
sakibsadmanshajib added a commit that referenced this pull request Aug 22, 2026
#963)

## Summary

Unblocks the required `Web E2E (full stack)` check on `main`, which was
intermittently red and hard-blocking every queued PR (including #960, a
live cross-tenant data exposure fix).

- Fixes a genuine, intermittent backend race condition, not a
test-selector or fixture drift.
- `provisionDefaultWorkspace` (default-workspace auto-provisioning on
first sign-in) had no idempotency guard. Next.js Server Components call
`getViewer()` unmemoized per component (`layout.tsx` and each `page.tsx`
independently), so a brand-new user's first console page load can fire
two concurrent `GET /api/v1/viewer` requests. Both see zero active
memberships and both try to create a personal workspace with the same
deterministic slug (`buildSlug` derives it purely from the viewer's own
display name). The loser hits `accounts.slug`'s unique constraint, and
`CreateAccount` let that raw pg error escape unhandled as an opaque 500
("could not load your workspace"), tripping the console's generic error
boundary.
- E2E mints a fresh test user per run, so its very first page load
reliably lands in this race window on a real, unlucky-timing basis,
hence flaky (fails once, passes on Playwright's retry) rather than a
hard, permanent break. This is exactly what the suite's flake-rate gate
("A test that only passes on a retry is not a passing test") is there to
catch.

## Root cause investigation

- Walked `Web E2E (full stack)` job history on `main`: last clean pass
was `1dc67ee6` (2026-08-17 16:56). PR #955's own pre-merge CI run hit
this exact race on two specs (`console-spend-alerts.spec.ts`,
`profile-completion.spec.ts`), both reported "2 flaky" (failed once,
passed on retry) by the job's flake gate, which is what failed the job,
not a hard 0-passed failure.
- Downloaded that run's `nextjs-log-web-e2e` artifact: the underlying
server error was `Error: could not load your workspace` with digests
matching the two `error-context.md` snapshots' "quote reference" numbers
exactly (`270132166`, `3259196144`). Traced that message to
`writeInternal(w, r, "could not load your workspace", err)` in
`accounts/http.go`'s `handleGetViewer`, then to `EnsureViewerContext` →
`provisionDefaultWorkspace` → `CreateAccount`'s unhandled `INSERT ...
accounts (slug ...)` unique-violation.
- This is unrelated to the separate, already-known-stale `OWUI nightly
e2e` selector issue (`button` named "New chat" vs the `<a>` OWUI
actually renders) that's in flight on #951 — different harness,
different spec, not touched here.
- Neither #954 nor #955 introduced this: #955 is a test-only Go change
and could not plausibly affect the web console; #954 touches auth-header
normalization, unrelated to workspace provisioning. Both merged while
this check was still pending (a known process gap, flagged separately),
but neither is the cause.
- This gate has **not** been silently ignored: the two prior real
(non-skipped, non-cancelled) job completions on `main` before this PR
were both green (`1dc67ee6` success, `ef06d693` success). PR #955 merged
with its own pre-merge run showing this failure, which is the actual
process gap worth naming; it did not merge past a run that had already
gone green.

## Fix

- New sentinel `accounts.ErrSlugTaken`, returned by `CreateAccount` on a
unique violation on `slug` (mirrors the existing `CreateMembership` →
`ErrAlreadyMember` translation already in the same file).
- `provisionDefaultWorkspace` now distinguishes the two situations a
slug collision can mean:
1. A concurrent request for the **same** viewer already won the
provisioning race — re-checking that viewer's own memberships finds the
winner's row, and this call recovers silently (no error; the caller
re-lists memberships right after, same as the success path).
2. A genuine collision between two **different** viewers whose display
names happen to produce the same slug (`buildSlug` has no per-user
uniqueifier) — this viewer still holds no membership after the
collision, so it retries with a de-duplicated slug (bounded, 3 attempts)
instead of failing that viewer's first sign-in outright.

## Test plan

- [x] `go build ./apps/control-plane/...` — clean.
- [x] `go vet ./apps/control-plane/...` — clean.
- [x] `gofmt -l apps/control-plane/internal/accounts/` — no unformatted
files.
- [x] `go test ./apps/control-plane/internal/accounts/... -count=1` —
all pass, including two new regression tests:
  - `TestEnsureDefaultAccount_SlugCollisionFromConcurrentSelf_Recovers`
-
`TestEnsureDefaultAccount_SlugCollisionFromDifferentViewer_RetriesWithSuffix`
- [ ] `Web E2E (full stack)` on this PR — will confirm the spend-alerts
and profile-completion specs no longer flake once CI runs.

## Buglog entry

Per `.wolf/` policy, not appended to `.wolf/buglog.jsonl` on this
branch. To be carried into a separate buglog-only PR off `main` once
this merges:

```json
{"error_message": "could not load your workspace (500) on GET /api/v1/viewer, surfacing as the web console's generic error boundary on the spend-alerts and billing-settings pages", "root_cause": "provisionDefaultWorkspace has no idempotency guard: two concurrent Server Component calls to getViewer() for the same brand-new user both attempt to create a personal workspace with the same deterministic slug, and CreateAccount let the resulting pg unique-violation on accounts.slug escape as a raw 500 instead of recovering the race", "fix": "translate the unique violation into accounts.ErrSlugTaken and have provisionDefaultWorkspace re-check the viewer's own memberships on collision: recover silently if a concurrent request for the same viewer already won, otherwise retry once with a de-duplicated slug for the genuine different-viewer collision case", "tags": ["accounts", "web-e2e", "race-condition", "ci-gate", "flaky-test"]}
```


<!-- greptile_comment -->

<details open><summary><h3>Greptile Summary</h3></summary>

The PR makes first-login workspace provisioning atomic and resilient to
both same-viewer races and cross-viewer slug collisions.
- Adds a per-viewer PostgreSQL advisory transaction lock and lock-scoped
membership re-check.
- Atomically inserts the account, owner membership, and profile.
- Translates slug uniqueness violations into a sentinel and retries
cross-viewer collisions with a deduplicated slug.
- Adds regression coverage and updates repository test doubles for the
expanded interface.
</details>


<details open><summary><h3>Confidence Score: 5/5</h3></summary>

The PR appears safe to merge, with no actionable changed-code defect
identified.

The transaction serializes provisioning by viewer, re-checks the same
active-membership invariant used by the service, atomically persists all
required workspace rows, and correctly handles the current slug
uniqueness constraint.
</details>


<details open><summary><h3>Important Files Changed</h3></summary>




| Filename | Overview |
|----------|----------|
| apps/control-plane/internal/accounts/repository.go | Adds atomic,
transaction-scoped workspace provisioning with a per-viewer advisory
lock, active-membership re-check, and precise slug-conflict translation.
|
| apps/control-plane/internal/accounts/service.go | Routes default
provisioning through the atomic repository method and retries genuine
cross-viewer slug collisions with bounded randomized suffixes. |
| apps/control-plane/internal/accounts/service_test.go | Adds focused
regression tests for concurrent same-viewer recovery and
different-viewer slug deduplication. |
| apps/control-plane/internal/accounts/types.go | Introduces the
ErrSlugTaken sentinel used to distinguish retryable slug collisions. |
| apps/control-plane/internal/accounting/http_test.go | Updates the
account repository test double for the new atomic provisioning
interface. |
| apps/control-plane/internal/budgets/http_test.go | Updates the account
repository test double for the new atomic provisioning interface. |
| apps/control-plane/internal/ledger/service_test.go | Updates the
account repository test double for the new atomic provisioning
interface. |
| apps/control-plane/internal/profiles/service_test.go | Updates the
profile-aware repository test double for atomic provisioning. |
| apps/control-plane/internal/usage/service_test.go | Updates the
account repository test double for the new atomic provisioning
interface. |

</details>


<details><summary><h3>Sequence Diagram</h3></summary>

```mermaid
sequenceDiagram
  participant A as Viewer request A
  participant B as Viewer request B
  participant S as Accounts service
  participant DB as PostgreSQL
  A->>S: EnsureViewerContext
  B->>S: EnsureViewerContext
  A->>DB: Begin + lock(viewer)
  B->>DB: Begin + wait for lock(viewer)
  A->>DB: Re-check membership
  A->>DB: Insert account, membership, profile
  A->>DB: Commit and release lock
  B->>DB: Acquire lock and re-check
  DB-->>B: Existing active membership
  B-->>S: "wonElsewhere = true"
  S-->>B: Re-list and use winner workspace
```
</details>

<sub>Reviews (1): Last reviewed commit: ["fix: close the
workspace-provisioning
ra..."](e08b57f)
| [Re-trigger
Greptile](https://app.greptile.com/api/retrigger?id=54270438)</sub>

**Context used:**

- Knowledge Base — [Control-plane identity and
access](https://app.greptile.com/scubed/-/custom-context/knowledge-base/sakibsadmanshajib/hive/-/docs/control-plane-identity-access.md)

<!-- /greptile_comment -->
sakibsadmanshajib added a commit that referenced this pull request Aug 23, 2026
…pplication (#951)

Owner directive, 2026-08-17: "stop using iframes and keep it native
inside OpenWebUI."

`https://chat-hive.scubed.co/agents` rendered exactly one `<iframe>`,
`loading="eager"`, pointed at `/agent-workspace/tasks`, which booted
`apps/agent-console` (a second whole Next.js application) inside the
page. That is why the agent surface still looked exactly as it did
before the shell landed, why it was slow, and why a task could never
become part of the conversation. The frame is gone.

Scope is bounded deliberately: visual and functional parity with what
the frame showed, natively, which is the task composer and the task
list. No tool call cards, no progress panel, no transcript. All three
are specified in `spec-2026-08-17-agent-run-surface` and blocked on the
event relay (its step S6), and a tool card with nothing to render is
worse than today's status row.

## Two premises in the brief that were already stale

Both checked in the tree at `1dc67ee69`, both reported before building
anything.

1. **`apps/agent-console/middleware.ts` does not send `X-Frame-Options:
DENY` or `frame-ancestors 'none'`.** It sends `SAMEORIGIN` and
`frame-ancestors 'self'` (`middleware.ts:80-84`), changed in #938 with a
comment naming the shell as the reason. Nothing was working around
anything: the frame was permitted because one Caddy listener serves both
applications on one origin, so `'self'` is a real origin check that
matches.
2. **The frame's `src` carried no credential.** It was `embed=1` and
`theme=<light|dark>`, both re-validated and rebuilt rather than
reflected (`middleware.ts:25-31`, `:93-100`).

What the credential path actually was: `apps/agent-console` holds its
own first-party Supabase session in cookies on the shared origin, and
its browser code calls `/v1/agent/*` directly with
`session.access_token`. That token comes from a direct Supabase sign-in,
so it carries the `tenant_id` claim the `custom_access_token_hook`
mints, and `JWTMiddleware` accepts it with no fallback. The frame
existed because a second application had a second, better token.

## The problem, and why the recommended fix needed one correction

The chat frontend holds Open WebUI's own session token, which edge-api
has never heard of, plus a Supabase access token minted through
Supabase's **OAuth-server** grant, which does not run
`custom_access_token_hook` and therefore carries **no `tenant_id`**.
`JWTMiddleware` rejects that (`middleware.go:123-129`). That is D-023
working, pinned by `inert_token_test.go`, and it is not relaxed here.

Worse for any browser-side fix: **that OAuth token is not in the browser
at all.** Open WebUI keeps it server side and resolves it per request
through `get_system_oauth_token`, refreshing it when expired
(`utils/middleware.py:3012-3045`). Handing it to page JavaScript would
be a downgrade, not a fix.

So the recommendation to reuse the mechanism `hive_jwt_forward.py` plus
`owui_unwrap.go` already run in production is adopted. The correction:
**`__metadata.upstream_auth` cannot serve this surface**, because three
of the four agent-task calls have no JSON body to hide a token in.

| Call | Method | Body |
|---|---|---|
| `GET /v1/agent/tasks` | GET | none |
| `POST /v1/agent/tasks` | POST | JSON |
| `GET /v1/agent/tasks/{id}` | GET | none |
| `POST /v1/agent/tasks/{id}/cancel` | POST | none |

(`apps/edge-api/internal/agenttask/handler.go:38-72`.)

## The design, three layers

### 1. edge-api: a second carrier on the same boundary

`X-Hive-Upstream-Auth` carries the per-user token for requests with no
body to carry it in. **This is not a new auth boundary.** The trust
decision is unchanged and unchanged in strength: the header is honoured
only when `Authorization` is exactly the shim key, which is the
identical gate the body carrier already sits behind. What changes is the
carrier, not who is allowed to use it.

Properties, each with a test:

* Honoured only under the shim key. Under a real API key, a real JWT, or
no credential at all, it is ignored.
* **Stripped from the forwarded request on every branch**, including the
ignored one and including when the middleware is disabled. A header a
client controls must never be readable by a handler, a log or an audit
sink, because a handler that can read it is one that could later be
taught to trust it.
* Same 8 KiB cap as the body carrier, shared through one normaliser so
the two cannot drift into accepting different shapes of the same
credential.
* Fails closed when present but unusable, rather than forwarding with
the shim key still on `Authorization`.

`requiresPerUserAuth` grows the agent-task paths. It previously covered
`/v1/chat/completions` only, which meant a shim-key request to
`/v1/agent/tasks` with no user token would pass through and **bind to
the shim's own principal**, listing and cancelling the shim account's
tasks instead of the user's. Nothing sent that request; this change is
what keeps the new proxy from being one bug away from sending it. It is
a tightening, not a widening: a silent mis-attribution becomes a 401.

One latent defect fixed while in the same function: a present-but-empty
`__metadata.upstream_auth` previously returned `unwrapOK` with an empty
token, which skipped the fail-closed arm and forwarded to
`/v1/chat/completions` with the shim key intact. It now takes the
missing-carrier arm, so it 401s and logs like every other missing token.

Not touched, deliberately: the tenant fallback is still gated on
`IsOWUIUnwrapped` and is not widened by one line; `Selector` is
unchanged; `inert_token_test.go` passes unmodified.

### 2. The chat container: a server-side proxy

Only the **frontend** of `hive-open-webui` is built from the fork; the
Python backend stays upstream's pinned image on purpose (Dockerfile
header, D-036). So this arrives as a copied module plus one asserted
splice, the same shape as `hive_model_picker` and `hive_rag_env_config`.
The frontend builds with `adapter-static`, so there is no SvelteKit
server to put it in instead.

`deploy/docker/owui-patches/hive_agent_proxy.py`, mounted at
`/api/v1/hive/agent`:

* Every route depends on `get_verified_user`, and **no route reads a
user id, a tenant id, an email or a token from the request**, so a
caller cannot influence which principal edge-api resolves.
* Presents the shim key on `Authorization` and the user's token on the
carrier header. Neither is ever returned to the browser, logged, or
included in an error.
* Four named operations with the task id validated as a UUID before it
is interpolated. Deliberately not a general-purpose proxy, which is the
shape #770 removed from this image.
* The create body is **rebuilt from two named fields** rather than
passed through, so nothing a caller invents (a `__metadata` block, for
one) reaches edge-api.
* Adds no configuration: `OPENAI_API_BASE_URL` and `OPENAI_API_KEY` are
already set on this container, so the shim key gains no new home.

### 3. The frontend

`src/lib/hive/agentTasks.ts` and `AgentTasks.svelte`, plus the `/agents`
route. Ported from the console, including two decisions that are easy to
mistake for accidents and are not: the `unknown` status sentinel (a
status this build cannot name still gets a row, because dropping it made
a submitted task vanish) and the two engine sentinels (a deployment with
no runtime reads "Blocked" in the warning register, never "Failed" in
red next to a real failure).

## What happened to `/agent-workspace`, stated plainly

It still exists and it is still served. Nothing in this PR removes it,
and the PR should not be read as removing it.

Three separate things get confused under that name, so all three,
precisely:

* **The `agent-hive.scubed.co` subdomain is already gone**, and was
before this PR. It is NXDOMAIN and decommissioned (D-026, verified
2026-07-28). There is no separate agent host and has not been one for
some time.
* **`/agent-workspace` is a path on the chat host**, proxied to the
`agent-console` container by `deploy/docker/Caddyfile.owui:161-168`.
That rule is untouched here, so the path still resolves and still serves
`apps/agent-console`.
* **Nothing in the product points at it any more.** The nav row goes to
`/agents`, and `/agents` no longer loads it. A user reaches it only by
typing the URL.

So after this PR the old surface is unreachable through the interface
but still reachable by URL. Making it actually stop being a thing is the
retirement described in the next section, and it is deliberately not in
this PR: the console is the only surface that has ever launched a task
against the live engine, so it is the control if the native path
misbehaves on the box. Recommendation is to retire it once this has run
there for a day.

## On the composer: the extraction was taken, not the fallback

Answering the directive's point 1 plainly.

`MessageInput.svelte` cannot be consumed wholesale: 2222 lines, 29
exported props, several required and chat-specific (`history`,
`selectedModels`, `createMessagePair`, `stopResponse`, `taskIds`,
`messageQueue`), and its container's classes are computed from chat
stores. Cloning its CSS was banned, correctly.

So the **presentational shell was extracted**, which was the preferred
option:

* `src/lib/hive/ComposerShell.svelte` is `#message-input-container`
lifted verbatim: the rounded surface, the border, the hover and
focus-within treatment, the inner padding. It reads no store; the two
values that used to compute its classes arrive as props.
* `src/lib/hive/ComposerSendButton.svelte` is `#send-message-button`
lifted verbatim, classes and glyph unchanged.
* `MessageInput.svelte` now renders both. **No behavioural edit was
needed**, which was the stated condition for taking this option: the
diff there is two tag swaps, one component swap and one import.

What is deliberately **not** shared is the row of controls inside the
container. Chat's carries attach, tools, skills, web search, voice and
the model picker; the agent composer's carries the pack toggle. Sharing
that row would mean sharing chat state with a surface that has none of
it, which is the coupling the extraction exists to avoid.

Per the directive: the heading block and the helper paragraph are
deleted. The pack choice is a segmented **Knowledge work** / **Coding**
toggle inside the composer. The one line describing the selected mode
survives as a single muted line under the composer, because it is
genuinely useful and it does not turn the composer back into a form.
`Enter` sends and `Shift+Enter` is a newline, which is the chat
composer's own behaviour. The two packs and their semantics are
unchanged, and so is the default (`coding-pack`): this is a presentation
change, not a capability change.

Chat's own composer is proved unchanged two ways, per the condition set:
the vitest suite is green, and before and after screenshots of the
**chat** composer are posted below alongside the agent one.

## Does `apps/agent-console` still need to exist?

Honest assessment, not acted on here.

**No route in it is reachable from the product any more.** Nothing
frames it, and the only two links that ever pointed at it are gone: the
nav row goes to `/agents`, and `/agents` no longer loads
`/agent-workspace/tasks`. It is now a signed-out URL that a user reaches
only by typing it.

What retiring it would take, in order:

1. Delete the `@agentConsole` matcher and its `reverse_proxy` from
`deploy/docker/Caddyfile.owui` (lines 145 to 168), which is what makes
`/agent-workspace` resolve at all.
2. Remove the `agent-console` service from
`deploy/docker/docker-compose.yml` and its `Dockerfile.agent-console`,
plus the `NEXT_PUBLIC_EDGE_API_BASE_URL` and
`EDGE_API_INTERNAL_BASE_URL` wiring that exists only for it.
3. Remove the `agent-console-unit` job from `.github/workflows/ci.yml`,
which is a required check, so branch protection needs updating in the
same change or the merge gate waits forever on a job that no longer
runs.
4. Delete `apps/agent-console/`.

Two things worth keeping first, and neither is code: its `proof/`
harness, and the accessibility reasoning in `task-console.tsx` (it
refuses `--color-ink-3` for body text on a measured contrast argument).
The second is already carried across into `AgentTasks.svelte`'s comments
and CSS.

One caveat against deleting it immediately: it is the only surface that
has ever launched a task against the live engine, so it is the control
if the native path misbehaves on the box. Recommendation is to retire it
in a follow-up once this has run on the demo box for a day, not in this
PR.

## Testing

* `go test ./apps/edge-api/... -count=1 -short` in the toolchain
container: green, whole module, including nine new cases for the carrier
and the fail-closed paths.
* `python3 scripts/test_owui_agent_proxy.py`, wired into `make
test-scripts` (a required CI check). Pure standard library, no
framework, matching the other `scripts/test_owui_*.py` self-checks. It
is mutation tested: swapping the two credentials onto the wrong headers,
deleting the missing-token guard, and forwarding the whole request body
each make it fail.
* `npm run test:frontend` for the fork, covering the decoder, the status
mapping and the four calls.

**A gap named rather than hidden.** `ci.yml` has vitest lanes for
`apps/desktop`, `apps/web-console` and `apps/agent-console`, and none
for `vendor/open-webui`, so every test ever written under `src/lib/hive`
was coverage on paper only. This adds `npm run test:frontend -- --run`
to the image's frontend build stage, which means the fork's tests gate
the nightly build and the demo deploy. That is not the same as gating a
pull request and is not claimed as such; a PR-time lane for this tree is
a follow-up.

## Buglog entry

```json
{"id":"owui-agent-shim-principal-gap","date":"2026-08-17","title":"OWUI shim-key requests to /v1/agent/tasks would have bound to the shim's own principal","error_message":"none observed; latent","root_cause":"requiresPerUserAuth in apps/edge-api/internal/auth/owui_unwrap.go returned true only for /v1/chat/completions, so a shim-key request to any other path fell through with the shim key still on Authorization and resolved as the shim account. Correct for the paths that existed then (embeddings and text-to-speech authenticate as the shim by design), wrong the moment a second per-user OWUI path appeared. A present-but-empty __metadata.upstream_auth had the same effect on the chat path itself, because it returned unwrapOK with an empty token and skipped the fail-closed arm.","fix":"Extended requiresPerUserAuth to /v1/agent/tasks and its subtree, and made an empty upstream_auth report as a missing carrier so it takes the 401 arm and the warn log.","tags":["auth","edge-api","owui","fail-closed","tenancy"]}
```

## Visual proof: not posted yet, and exactly what exists today

Stated precisely, because an earlier version of this section said
screenshots were "committed under `docs/proof/` and posted in a
comment". Neither was true. Nothing is committed under `docs/proof/` in
this diff, no image is on this PR, and `owui-nightly.yml` has no step
that commits or comments; it only uploads a transient CI artifact. That
sentence described intent before the run finished and read as completed
work. It will not be restated until an image is actually visible on this
page.

**What has been verified on a real stack with a real signed-in
session**, run 32066161156, the OWUI end-to-end job on this branch:

- `the agent workspace opens inside the shell and frames nothing`
passed. That is the load-bearing one: an `iframe` count of zero on
`/agents`, the composer present and accepting typed input, both toggle
options clickable, the list painting exactly the rows the API returned
matched against that response's own text, and the call to the proxy
answering something other than 401, which is the single failure mode
this design has.
- `proof capture: the agent surface, natively, in both palettes` passed,
so the images were produced.
- Every chat spec passed on that same build: send and stream,
multi-turn, model switch, tenant model visibility, sign-out. That is
evidence the composer extraction did not break chat, and it does not
depend on anyone reading a screenshot.

**Why no image was attached at first, and a finding worth its own
paragraph.** The captures were being written into
`playwright-report-owui/proof`, inside the HTML reporter's own output
folder. That reporter clears its folder before writing the report, so
**every capture was produced and then deleted**, and the result looked
exactly like a capture step that had never run. Nothing failed, nothing
warned, and the test that wrote them passed. Confirmed by downloading
the artifact and finding no `proof/` directory rather than by inferring
it.

Anyone adding a capture step to a Playwright suite in this repository
will walk into the same trap, which is why it is recorded here rather
than only fixed: proof must never be written inside a reporter's output
directory. The captures now go to a sibling directory with its own
upload step. This is the same silent-absence shape that let the
end-to-end gate rot unnoticed, where an artifact that is missing and an
artifact that was never asked for are indistinguishable after the fact.

**Two attempts to re-capture since then both failed outside this
branch**, and neither is counted as anything:

- Run 32112158077: the sign-in helper's consent hop exceeded its 30
second budget. The container log shows the token exchange completing 28
seconds after its redirect and the OAuth session being stored, so the
login worked and was simply late.
- Runs 32113599893 and 32114089597: the seeder's first write returned
`POST /tenants -> 522` in both. Without fixture credentials the OWUI
project matches zero spec files, so nothing in this PR ran at all.

**Correction to an earlier version of this section**, which called that
522 an external Supabase problem. It is very probably ours. The same job
log carries `(ECHECKOUTTIMEOUT) unable to check out connection from the
pool after 15000ms in Session mode`, which is the documented 15-client
session-mode pooler ceiling being hit, and that single cause explains
both symptoms: a PostgREST request that cannot get a database connection
hangs until Cloudflare gives up with a 522, while a direct `pg` client
gets the checkout timeout by name. Blaming the provider was the
comfortable read and the evidence does not support it.

I also want to retract a piece of reasoning rather than quietly drop it:
I probed `/rest/v1/` unauthenticated, got a 401, and treated that as
evidence the origin had recovered. An unauthenticated request never
reaches the connection pool, so that probe could not have detected the
condition that matters and proved nothing.

Neither failure touches a file in this diff, and a run that never
reached the suite is an absence of information rather than a result. The
captures are outstanding for that reason and for no other.

**How they will be posted.** Through `scripts/post-pr-visual-proof.sh`,
which uploads to a release asset. A raw link pinned to this branch 404s
the moment the branch is deleted, and this repo squash-merges and
deletes branches, so that mechanism produces proof which expires exactly
when the PR it proves gets merged.

## A third thing, and the budget I did not raise

Investigating why the proof runs kept failing turned up a cause that is
not a test problem. The consent and sign-in pages are web-console's, on
a different origin, and web-console is served by a development build, so
it compiles each page the first time anyone requests it. On run
32112158077 the first two login attempts produced no page within 30
seconds and the third completed the whole exchange in 21 seconds once
warm. That is a route compiling on demand, not a slow protocol.

The harness now waits for the consent app to be serving before it starts
timing the login. **The 30 second login budget is deliberately
unchanged.** Raising it was the obvious move and the wrong one: it would
have buried a cold compile inside a per-login timeout, and every future
slow login would then look exactly like this one, which is how the next
real regression gets absorbed instead of noticed. The new wait is a
separate budget for a separate thing, the app becoming reachable at all.
It also adds no new failure mode, because the assertion it precedes
already requires a DOM that exists only on the consent origin.

Why that app is served in dev mode is a product question, it is
deliberate per D-022, and it is filed as #967 rather than decided here.
It is a plausible cause of the slow sign-in being reported in real use.

## Two things a reviewer should know

**The OWUI end-to-end harness was red before this branch and is fixed
here.** It still is on `main`, at that same readiness check, which is
independent evidence this fix is real and not a symptom of something
else on this branch. It is also a different defect from the provisioning
race PR #963 is separately fixing, so the two should not be read as one
problem. Its sign-in readiness signal waited for a `button` named "New
chat", and Open WebUI renders both New Chat controls as anchors carrying
an `aria-label`, so that query could never match. The sign-in it guards
actually succeeds; the downloaded failure screenshot shows a working
signed-in chat page. A sibling file already matched either role for
exactly this reason. That fix is why any of the verification above
exists, and it is separate from this feature.

**Two nav tests from #938 went red the first time they ever ran**, which
was in this branch after that fix. Neither is a regression here and
neither was weakened. Both assumed a sidebar that starts expanded while
this fixture's starts collapsed, so a text assertion resolved to the
icon-only rail and read "", and the collapsed-rail test waited for a
control that only exists while the sidebar is open. Each test now drives
the sidebar into the state it asserts about.










---

## Rebase onto current main, 2026-08-22, and what this closes

Rebased onto `main` at `c30882491`. The base this branch was validated
against, `1dc67ee69`, predates the migration of the entire database off
hosted Supabase onto the self-hosted instance on the demo box, the move
of the public auth origin to a same-origin `/auth/v1` route on the
console, and the switch to admin-provisioned-only signup. One conflict,
in `Makefile`, where `main` had added seven self-checks to
`test-scripts` and this branch adds one; both are kept.

### This is the fix for the top demo blocker, issue #540

Re-confirmed live on the box on 2026-08-22: a user signs into
`chat-hive.scubed.co` normally, clicks **Agents**, and is presented with
an email and password form captioned that the workspace is separate from
chat. Proved there by difference rather than guessed: chat's OAuth
handshake mints an Open WebUI token and never writes the
`sb-...-auth-token` cookie that the embedded `apps/agent-console` reads,
and injecting the same session's Supabase cookies on the `chat-hive`
origin makes the real workspace render immediately.

That is the second-application problem this branch removes. The
credential the embedded application was looking for is no longer needed
by anything, because the surface is native and authenticates with the
Open WebUI session the user already has: the panel calls
`/api/v1/hive/agent/*` on the chat origin, and
`deploy/docker/owui-patches/hive_agent_proxy.py` brokers it server side
through `get_system_oauth_token`, which is the one place the user's
Supabase token is reachable. #540's own analysis named that as the only
possible broker.

### Does a second credential prompt remain reachable

**Yes, by one route, and this branch does not close it.**
`/agent-workspace/*` is still proxied to `apps/agent-console` by
`deploy/docker/Caddyfile.owui`, and that application still renders its
own email and password form, so a typed or bookmarked
`https://chat-hive.scubed.co/agent-workspace` still reaches one. This
branch's only edit to that file is a comment; it changes no route.

No route inside the chat interface leads there any more.
`vendor/open-webui/src/lib/hive/nav.ts` points the sidebar entry at
`/agents`, and the only two remaining mentions of the old path anywhere
in the front end are historical comments. So the demo path is fixed and
the residual is a URL nobody is shown.

Answering that path 404 was considered and rejected here rather than
skipped: the Tauri desktop app targets `/agent-workspace` as its console
base path (`apps/desktop/src/settings.ts`,
`apps/desktop/src-tauri/src/settings.rs`), and
`apps/web-console/tests/e2e/_probe/agent-workspace-flows.spec.ts` is a
coverage ledger over that surface. Retiring `apps/agent-console`, or
moving the desktop app onto the native surface first, is a separate
decision with a far larger blast radius than this pull request.

### One coupling worth knowing before this is demoed

The proxy reads the user's Supabase token through
`get_system_oauth_token`, which goes through
`OAuthManager.get_oauth_token`, which refreshes five minutes before
expiry and **deletes the OAuth session outright when that refresh
fails**. That is issue #782, still unfixed on `main` and fixed by PR
#787. Until #787 lands, this agent surface stops working roughly 55
minutes after sign-in for the same reason chat does, and the panel will
correctly report "Your Hive sign-in could not be confirmed. Sign in
again and retry." A demo that runs longer than that from a single
sign-in needs #787 merged first.

### Review findings addressed on the rebased head

Two threads, both fixed rather than argued, plus one defect found while
verifying them.

A list refresh overlapping a create or a cancel replaced the whole task
array with its older answer, dropping the row the user had just
submitted or reverting one they had just cancelled, and when that stale
answer held no in-flight task the poll loop stopped too and the screen
did not recover without a reload. A refresh now captures a mutation
counter before its request goes out and discards its own answer if a
create or a cancel landed while the request was open.

The third entry in the identity smell tuple in
`scripts/test_owui_agent_proxy.py` could not fail: for
`headers.get('Authorization')` the accessor templates expanded to
`request.query_params.get('headers.get('Authorization')')`, a string no
Python source can contain, so that iteration reported a pass over the
header read it was named after. The header read now has its own direct
test against the whole `request.headers` attribute, and it fails on
purpose when the string is planted on a line that never executes.

Found while running the above: `agentTasks.test.ts`, 203 lines of
assertions, was running in no job at all. The module imported
`$lib/constants`, which reaches `$app/environment`, and the only runner
that covers this front end runs plain vitest with no alias resolution,
so the file could not be loaded. It would have turned that required
check red on whichever of #951 and #952 merged second. The API base is
now a parameter with a production default and the component passes the
dev-aware value, so a built bundle and `npm run dev` both behave exactly
as before, and the tests load with no configuration: 16 of them, one of
which fails when `encodeURIComponent` is dropped from the cancel path.

### Verification on the rebased head

- 31 vendored front end tests pass across three files, 16 of them in
`agentTasks.test.ts`, which had never executed before this rebase.
- `make test-scripts` green, including `test_owui_agent_proxy.py` and
the seven self-checks `main` added.
- `go test ./apps/edge-api/internal/auth/... -short` green, which covers
the `owui_unwrap.go` change against `main`'s newer `x-api-key`
normalisation (#954).
- `Caddyfile.owui` validates against the pinned Caddy image, using
`main`'s new "Every Caddyfile adapts" CI step.

### Visual proof

Posted on this pull request as release assets, with the full method, the
artefacts of the stub, and the `/agent-workspace` residual in
`docs/proof/agents-native-no-iframe-2026-08-22/README.md`. Measured in
the page rather than asserted: zero password inputs, zero iframes, and
no "Sign in" text on `/agents`.

The capture uses the real bundle from `docker build --target frontend`
on this branch with a stubbed backend on a loopback origin, and says so
on every artefact. A live capture is not available before merge for two
reasons that are not properties of this branch: the panel needs the
`owui-patches` router this branch adds, which exists only in a rebuilt
image, so bundle interception against the deployed origin would 404; and
the development box's `.env` still points `SUPABASE_URL` at the hosted
Supabase project deleted in the migration, which now returns
`ENOTFOUND`, so no session can be minted. No password was set, reset or
rotated to work around that.

### Buglog entry

To be appended to `.wolf/buglog.jsonl` on `main` in a separate
buglog-only pull request after this merges, per the openwolf protocol.

```json
{"id":"owui-agents-second-credential-prompt","date":"2026-08-22","error_message":"Clicking Agents in the chat sidebar presented its own email and password form captioned that the workspace is separate from chat, one click after a successful chat sign-in","root_cause":"The Agents route embedded apps/agent-console in an iframe, and that application authenticates from a Supabase SSR cookie on the chat origin which chat's OAuth handshake never writes: the handshake mints an Open WebUI token only, and the user's Supabase token is reachable only server-side inside the chat container as the stored OAuth token","fix":"Removed the iframe and rendered the task surface natively in the chat application, authenticating with the Open WebUI session the user already holds and brokering the Supabase token server-side through a FastAPI router added by owui-patches/hive_agent_proxy.py","tags":["owui","auth","agents","iframe","session","demo-blocker","issue-540"]}
{"id":"agent-tasks-stale-poll-overwrites-mutation","date":"2026-08-22","error_message":"A newly created agent task row disappeared, or a cancelled row reverted, and polling sometimes stopped entirely until the page was reloaded","root_cause":"refresh assigned the fetched list over the whole tasks array, so a poll already in flight when a create or cancel completed landed afterwards and overwrote it, and schedulePoll then decided from that stale array and stopped when it held no in-flight task","fix":"A mutation counter captured before the request goes out; refresh discards its own answer when a create or cancel landed while the request was open","tags":["svelte","race","polling","agents"]}
{"id":"agenttasks-unit-tests-never-ran","date":"2026-08-22","error_message":"agentTasks.test.ts reported no failures because it was never loaded by any job","root_cause":"The module imported $lib/constants, which reaches $app/environment, and scripts/test-owui-hive-frontend.sh runs plain vitest over copied files with no SvelteKit alias resolution, so the test file could not be loaded at all","fix":"The API base is a parameter with a production default and the component passes the dev-aware value, so the module imports nothing and the tests load with no configuration","tags":["tests","unfailable-check","vitest","sveltekit"]}
{"id":"owui-agent-proxy-smell-check-unfailable","date":"2026-08-22","error_message":"The identity smell loop in test_owui_agent_proxy.py reported a pass over a header read it was named after","root_cause":"The tuple entry headers.get('Authorization') was expanded through accessor templates into request.query_params.get('headers.get('Authorization')'), a string no Python source can contain, so that iteration asserted nothing","fix":"Separated the header smell into a direct substring test against the whole request.headers attribute, demonstrated failing on purpose","tags":["tests","unfailable-check","security"]}
```




<!-- greptile_comment -->

<details open><summary><h3>Greptile Summary</h3></summary>

The PR replaces the iframe-hosted agent workspace with a native Open
WebUI task surface and brokers per-user agent API credentials through a
constrained server-side proxy.
- Adds a guarded upstream-auth header carrier and fail-closed agent-task
authentication.
- Adds native task composition, listing, cancellation, polling, and
shared composer presentation.
- Updates frontend, proxy, authentication, and end-to-end verification
coverage.
</details>


<details open><summary><h3>Confidence Score: 5/5</h3></summary>

The PR appears safe to merge.

No blocking failure remains; the mutation-generation guard prevents
stale refreshes from overwriting completed creates or cancellations, and
ID filtering prevents the create/refresh duplicate-row race.
</details>


<details open><summary><h3>Important Files Changed</h3></summary>




| Filename | Overview |
|----------|----------|
| apps/edge-api/internal/auth/owui_unwrap.go | Adds a stripped,
shim-gated header carrier for per-user agent credentials and extends
fail-closed authentication to agent-task paths. |
| deploy/docker/owui-patches/hive_agent_proxy.py | Adds the
authenticated server-side broker that resolves each user’s OAuth token
and forwards only the four supported task operations. |
| vendor/open-webui/src/lib/hive/AgentTasks.svelte | Implements native
task composition, polling, cancellation, stale-refresh suppression, and
create-row deduplication; both previously reported races are fixed. |
| vendor/open-webui/src/lib/hive/agentTasks.ts | Defines the typed
agent-task client, decoding, status handling, and constrained proxy
calls used by the native surface. |
| vendor/open-webui/src/routes/(app)/agents/+page.svelte | Replaces the
framed agent application with the native AgentTasks component. |
| vendor/open-webui/src/lib/components/chat/MessageInput.svelte | Adopts
extracted composer shell and send-button components without changing the
chat composer’s behavior. |
| apps/web-console/e2e/phase-19/owui/09-agent-workspace-nav.spec.ts |
Verifies native rendering, absence of iframes, proxy authentication
behavior, navigation, task rows, and visual-proof capture. |

</details>


<details><summary><h3>Sequence Diagram</h3></summary>

```mermaid
sequenceDiagram
  participant U as Signed-in user
  participant UI as Native OWUI agent surface
  participant P as OWUI agent proxy
  participant A as Edge API auth middleware
  participant T as Agent-task API
  U->>UI: Open /agents
  UI->>P: "GET/POST /api/v1/hive/agent/*"
  P->>P: Resolve server-side OAuth token
  P->>A: Shim authorization + upstream-auth carrier
  A->>A: Validate shim gate, strip carrier, unwrap user token
  A->>T: Request authorized as signed-in user
  T-->>A: Task response
  A-->>P: Task response
  P-->>UI: Sanitized response
  UI-->>U: Native task list and composer
```
</details>

<sub>Reviews (4): Last reviewed commit: ["fix: deduplicate the created
task row
ag..."](f3df569)
| [Re-trigger
Greptile](https://app.greptile.com/api/retrigger?id=54086925)</sub>

**Context used:**

- Knowledge Base — [Control-plane marketplace-backed agent task
workflow](https://app.greptile.com/scubed/-/custom-context/knowledge-base/sakibsadmanshajib/hive/-/docs/control-plane-agent-workflows.md)
- Knowledge Base — [Edge API security
boundary](https://app.greptile.com/scubed/-/custom-context/knowledge-base/sakibsadmanshajib/hive/-/docs/edge-api-security-boundary.md)

<!-- /greptile_comment -->

---

## Buglog entry, third (added 2026-08-23)

Alongside the two above. To be appended to `.wolf/buglog.jsonl` on
`main` in
the same buglog-only pull request after this merges, per the openwolf
protocol. Not appended on this branch: a branch that appends a line
conflicts
with every other branch that appended one, and an unmergeable pull
request
gets no CI run at all (issue #873).

```json
{"id":"owui-unwrap-carrier-strip-unobservable","date":"2026-08-23","title":"Header-carrier strip asserted on a path no test could observe","error_message":"TestOWUIUnwrap_HeaderCarrierPresentButBlank_StrippedAndRejected asserted only the Rejected half of its own name; the Stripped half was unobservable because next never runs on a 401","root_cause":"The strip was applied to a clone of the request, so the only observation point was the downstream handler. Every rejection branch answers without calling next, so nothing on those branches could be checked, and an outer middleware still holding the pre-clone pointer would also have kept seeing a live per-user token. Moving Header.Del into the forwarding branches alone would have left the whole test file green.","fix":"Strip the carrier from the inbound request in place instead of from a clone, making the invariant one fact rather than one fact per branch, and assert in both rejection tests that the header is gone from the request the middleware was handed. Safe because net/http never re-reads request headers after the handler returns and the header is ours alone.","verification":"Green confirmed on unmodified code first. The production strip was then narrowed to fire only for a usable carrier, which is the exact regression described; all four blank sub-cases and the over-long case went red naming the new assertion, and the pass-through case went red too. Restoring the unconditional strip returned ./apps/edge-api/... to green.","tags":["auth","edge-api","owui-unwrap","test-quality","guard-cannot-fail","security","pr-951"]}
```

## Visual proof: now posted, correcting the section above

The section headed "Visual proof: not posted yet" is out of date and its
own
condition has been met. Images are now visible on this page, posted as
permanent release assets through `scripts/post-pr-visual-proof.sh`, from
a
live signed-in capture against a complete local stack rather than a
bundle or
a CI artifact.

The capture was taken after rebasing onto `main` at `82375d07f`, where
#952
and #956 have landed, so it exercises the real post-merge state.
Measured in
the live DOM at both hops: zero password inputs, zero iframes, zero
child
browsing contexts, zero anchors to `/agent-workspace`. `GET
/api/v1/agent/tasks` answered 200 through the running proxy chain, so
the
authenticated data path is exercised and not only the rendering.
`/agent-workspace` still answers 307 to `/agent-workspace/tasks`,
verified
after the rebase and deliberately unchanged.

Method, substrate and credential handling:
`docs/proof/agents-native-live-2026-08-23/README.md`.

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
sakibsadmanshajib added a commit that referenced this pull request Aug 29, 2026
…#1274, #1260, #1261, #1259) (#1296)

Closes #1274. Closes #1260. Closes #1261. Closes #1259.

## Are these one root cause or four?

Two families, not one, and not four independent fixes.

**Family A, Go zero-value JSON encoding dropping values the Anthropic
contract requires to be present (#1274 and #1260).** Both are the
gateway emitting a valid Go zero value where Anthropic requires an
explicit empty one, and in both cases the Go-side assertion a normal
test would write cannot see the defect at all: `len(content) == 0` is
true for a nil slice and an empty one alike, and a `Text` field set to
`""` looks correct right up until `encoding/json` runs. The mechanism
differs (`omitempty` on a string in one, a nil slice with no tag in the
other), so it is one class with two instances rather than a single line
to change.

**Family B, an auth model that only ever covered POST /v1/messages
(#1261 and #1259).** Every other route on the Anthropic surface was left
to inherit an auth path built for a different dialect. `count_tokens` is
the only route here that does not delegate to the chat chain, so it
never gained the API-key authority that chain provides, and `GET
/v1/models` was never given the leaf `APIKeyNormalizer` that
`/v1/messages` has carried since #954. The two share a cause but not a
line of code.

#1259 additionally carries a response-shape defect that belongs to
neither family: the route answered an Anthropic client with the OpenAI
list envelope and the OpenAI error envelope. Fixed here in the same
pass, since fixing only the auth half would have moved the failure one
step later.

## Ground truth

API shapes were verified against the live specification on 2026-08-28,
not from model memory, per the repository rule.

- `https://docs.claude.com/en/docs/build-with-claude/streaming` shows
`content_block_start` as `{"type":"text","text":""}` for a text block
and `{"type":"tool_use","id":...,"name":...,"input":{}}` for a tool_use
block, in every one of its five worked examples.
- `https://docs.claude.com/en/api/models-list` shows the list envelope
as `data` (entries of `type`/`id`/`display_name`/`created_at`) plus
`has_more`, `first_id` and `last_id`, the last two typed "string or
null".
- The SDK's own accumulator (`anthropic/lib/streaming/_messages.py`,
`accumulate_event`) was read directly to confirm the failure mechanism:
it does `content.text += event.delta.text` for a `text_delta`, which is
the `None += str` crash, and for an `input_json_delta` it **assigns**
`content.input` from its own `jiter` buffer rather than reading the
block-start value. That is why the missing `text` was fatal and the
missing `input` was not, and it is why this PR treats the two
differently in its claims.

`openwolf bug search anthropic` and `openwolf bug search omitempty` both
returned no prior entries.

## What changed

| Issue | Change |
| --- | --- |
| #1274 | `StreamContentBlock.Text` is now `*string` so the empty string
serializes, following the precedent `StreamEvent.Index` already set in
this same struct for the identical reason. `Input json.RawMessage` added
and set to `{}` on a tool_use block start. |
| #1260 | `FromOAIResponse` seeds `Content` with an empty non-nil slice
at construction and appends into it, which covers the early
`len(resp.Choices) == 0` return as well as the normal path. |
| #1261 | `anthropic.Deps.AuthorizeAPIKey` added; `count_tokens` accepts
a JWT session principal or an API-key principal. Nil leaves it
session-only and fail-closed. The refusal is the authorizer's own
verdict run through the shared `WriteAuthFailure` status mapping and
then reshaped into the Anthropic envelope, so no status logic is
duplicated. |
| #1259 | `GET /v1/models` is registered through a named `modelsHandler`
applying `anthropic.APIKeyNormalizer` at the leaf and
`anthropic.ModelsCompat` for shape. |

### Why the leaf normalizer on /v1/models is not redundant

`authSelectorMiddleware` already wraps every `/v1/*` route in
`APIKeyNormalizer` (#954), but only when JWT auth is wired: `handler =
authSelectorMiddleware(jwtMW, handler)` sits behind `if jwtMW != nil`.
On a deployment where Supabase JWT config is absent, edge-api logs "JWT
auth wiring skipped" and mounts no selector at all, so nothing
normalizes `x-api-key` and `handleModels` reads an empty `Authorization`
header. `/v1/messages` kept working there purely because it carries its
own leaf wrapper. That asymmetry is exactly the reported symptom in
#1259 (same key, same connection, `/v1/messages` fine and `/v1/models`
401) and it is why the fix is a leaf wrapper rather than a change to the
selector.

### Provider-blind invariant

The `ModelsCompat` translation reads `id`, `created` and `name` and
writes `type`, `id`, `display_name` and `created_at`. `owned_by` and
every other OpenAI field are dropped rather than passed through, so this
strictly reduces what reaches an Anthropic client.
`TestModelsCompat_AnthropicClientGetsTheAnthropicListShape` asserts
`owned_by` is absent from the body. Nothing on any path touched here
emits a provider name, an upstream id or a `system_fingerprint`; the
existing `TestHandleModelsDoesNotLeakProviderNames` and
`TestSSETranslator_MessageID_NeverLeaksUpstreamID` guards are unchanged
and still pass.

## Mutation check

Every added test was proved load-bearing by reverting the fix it guards
and observing it fail for the right reason, then restoring and observing
it pass. Full transcript, six mutations, all red, then green on restore:

| # | Mutation applied | Test | Result |
| --- | --- | --- | --- |
| 1 | `StreamContentBlock.Text` back to a plain `string`, call site back
to `Text: ""` |
`TestSSETranslator_TextBlockStartCarriesExplicitEmptyText` | FAIL: `text
content_block_start omits the "text" key entirely ... map[type:text]` |
| 2 | Delete `Input: json.RawMessage("{}")` from the tool_use block
start | `TestSSETranslator_ToolUseBlockStartCarriesExplicitEmptyInput` |
FAIL: `tool_use content_block_start omits the "input" key: map[id:call_1
name:get_weather type:tool_use]` |
| 3 | Delete the `Content: []ResponseBlock{}` seed |
`TestFromOAIResponse_EmptyCompletionSerializesContentAsArray` | FAIL on
both subtests: `got {... "content":null ...}` |
| 4 | `if h.deps.AuthorizeAPIKey == nil` forced to `if true`, i.e.
count_tokens back to session-only |
`TestHandler_CountTokens_AcceptsAPIKeyPrincipal`,
`..._RejectedAPIKeyKeepsTheAuthorizersOwnRefusal` | FAIL: `want 200 got
401
body={"type":"error","error":{"type":"authentication_error","message":"missing
user"}}`, and `error.message: want the authorizer's own refusal got
missing user` |
| 5 | `if !IsAnthropicClient(r)` forced to `if true`, i.e. never
re-shape | `TestModelsCompat_*` (3 tests) | FAIL: OpenAI body passed
through, `first_id`/`last_id` nil, `envelope: want top-level type=error
got <nil>` |
| 6 | `modelsHandler` unwrapped to `return handleModels(client,
authorizer)` | `TestModelsHandlerServesAnAnthropicSDKClient` | FAIL:
`x-api-key caller: want 200 got 401: {"error":{"message":"You didn't
provide an API key...` |

After restoring all six, `go test ./apps/edge-api/internal/anthropic/...
./apps/edge-api/cmd/server/...` is `ok` for both packages, and the wider
`go test ./apps/edge-api/... -count=1 -short` passes with no failures.
`gofmt -l` reports none of the touched files.

Four of the added tests are deliberately non-regression guards rather
than defect guards, and do not go red under any mutation above. The
count said two on first posting and was corrected in review; the table
is the record anyone reads later, so it should say what is actually
true.

- `TestModelsCompat_OpenAIClientIsUntouched` and
`TestModelsHandlerKeepsTheOpenAIShapeForOpenAIClients` exist to fail if
a later change starts re-shaping unconditionally and empties Open
WebUI's model picker. Both compare byte-identical bodies, so the inverse
mutation of always reshaping does turn them red.
- `TestHandler_CountTokens_WithoutAnAPIKeyAuthorityFailsClosed` stays
green under mutation 4, since a handler wired without an API-key
authority writes 401 either way. It pins the fail-closed default and
goes red if someone makes a nil authority permissive.
-
`TestHandler_CountTokens_SessionPrincipalDoesNotConsultTheAPIKeyAuthority`
also stays green under mutation 4, since the session branch returns
before the nil check is reached. It goes red if the two principal checks
are ever reordered.

`TestIsAnthropicClient` is a unit test of a new pure function that none
of the six mutations touch, so it is neither a defect guard nor a
non-regression guard for anything in the table.

## Review round two

Three findings from review, all fixed in `25b2820`, all in
`apps/edge-api/internal/anthropic`.

**Refusal retry headers were being dropped.**
`apierrors.WriteAuthFailure` is the shared source of truth precisely so
a retryable 429 is never collapsed into a non-retryable refusal, and it
delivers the retryable part through headers rather than the body:
`retry-after` plus the four `x-ratelimit-*` values on a real 429, and
`retry-after` on both the degraded-limiter branch and the
`upstream_unavailable` branch. Recording the refusal in a
`headerlessRecorder` and reshaping only its body threw all of that away.
`count_tokens` lost it latently, and `GET /v1/models` lost it live
through `authorizeAliasRequest`, which also left the two client shapes
disagreeing about retry metadata for the same refusal on the same route.
`headerlessRecorder.reshapeInto` now carries the headers over and
reshapes in one call, so a call site cannot take one half of a refusal
without the other. `Content-Type` and `Content-Length` are deliberately
not carried: the reshaped body is a different envelope of a different
length, and a stale `Content-Length` would truncate it on the wire.
Everything that is forwarded is retry metadata and carries no provider
identity, so the provider-blind invariant is unchanged.

**`GET /v1/models` served two representations for one URL without
declaring it.** Nothing on the route sets `Cache-Control` and the route
needs a credential, so no correct cache stores it today and no live
exploit exists. The declaration is what keeps a later edge cache in
front of Caddy, or an intermediary keying on URL alone, from handing an
Anthropic-shaped body to Open WebUI and emptying its model picker, which
is the regression `TestModelsCompat_OpenAIClientIsUntouched` exists to
prevent. `Vary: anthropic-version, x-api-key` is now set on both
branches, before either writes.

**`Finish` could emit `message_delta` before any `message_start`.**
`Translate` calls `Finish` unconditionally and `Finish` did not check
`t.started`, so an upstream stream that yielded no parseable chunk at
all produced a terminal pair with nothing before it, and the SDK
accumulator raises an unexpected-event-order error when a
`message_delta` arrives while its snapshot is still nil. Review
suggested a follow-up issue for this rather than widening the PR; the
fix turned out to be three lines in a file this PR already changes,
which is smaller than the issue describing it, so it is taken here.
`FeedLine` and `Finish` now share one `ensureStarted`, which emits
`message_start` exactly once per stream whichever of them reaches it
first.

Four more mutations, each applied alone, each restored after:

| # | Mutation applied | Test | Result |
| --- | --- | --- | --- |
| 7 | `reshapeInto`'s header loop forced to skip every key |
`TestHandler_CountTokens_RefusalCarriesTheAuthorizersRetryHeaders`,
`TestHandler_CountTokens_UpstreamUnavailableCarriesRetryAfter`,
`TestModelsCompat_RefusalCarriesTheRetryHeaders` | FAIL: `retry-after:
want 30 got ""`, `retry-after: want 5 got ""`, and the three
`x-ratelimit-*` assertions on both routes |
| 8 | `Content-Length` removed from the skip list, so the delegated
body's length is forwarded |
`TestModelsCompat_RefusalCarriesTheRetryHeaders` | FAIL:
`content-length: the delegated body length must not describe the
reshaped one, got "4096"` |
| 9 | The `Vary` line deleted | `TestModelsCompat_DeclaresVary` | FAIL
on both subtests: `vary: want both request headers that select the
representation, got ""` |
| 10 | `ensureStarted` removed from `Finish` |
`TestSSETranslator_EmptyStreamStillOpensTheMessage` | FAIL: `first
event: want message_start got message_delta` |

`TestSSETranslator_FinishAfterAStartedStreamDoesNotRepeatMessageStart`
is a fifth non-regression guard: it stays green under all ten mutations
and exists so that opening the message from `Finish` can never add a
second `message_start` to a stream that already carried one.

### Adversarial review of the round-two fixes

The round-two commit went back through the review pipeline before this
was called done. Two findings, both fixed in `f0e4c56`.

**The `Vary` declaration used `Set` where it should use `Add`.** Nothing
in edge-api declares a `Vary` on this route today (`grep` finds exactly
one writer, the new line itself), so overwriting was harmless in the
current tree and wrong the moment an outer middleware declares one of
its own, a CORS layer setting `Vary: Origin` being the obvious case. The
wrapper is now additive.

**The 2xx branch of `ModelsCompat` dropped every header the delegated
handler set.** That path re-encodes the body rather than reshaping it,
which is exactly why the carry-over the refusal path had just gained was
easy to leave out of it. A header set alongside a success is as much
part of that response as the body is, and this route is where a 2xx
rate-limit budget would surface. The header copy is now its own method
on the recorder, `copyHeadersTo`, and both branches call it.

| # | Mutation applied | Test | Result |
| --- | --- | --- | --- |
| 11 | `Vary` back to `Set` |
`TestModelsCompat_VaryDoesNotClobberAnExistingDeclaration` | FAIL:
`vary: want "Origin" preserved, got "anthropic-version, x-api-key"` |
| 12 | `rec.copyHeadersTo(w)` deleted from the success branch |
`TestModelsCompat_SuccessCarriesTheDelegatedHeaders` | FAIL:
`x-ratelimit-remaining-requests: want 42 got ""` |

### On `display_name` and the provider-blind edge

Review confirmed independently that the `ModelsCompat` translation is
provider-blind (it reads `id`, `created` and `name` only, so `owned_by`
and `description` cannot ride along) and flagged the remaining edge:
`display_name` is `model_aliases.display_name`, free text, and one
seeded row already reads `Openrouter Auto (Task Aware)`. That row is
`visibility='internal'` today, so the `visibility IN ('public',
'preview')` predicate keeps it off both list paths, but the migration
that added it describes a later flip to public as a one-line follow-up.

Nothing here closes #1284, and this PR never claimed to: the OpenAI
branch is unchanged and still ships the `hive-stt` and `hive-tts`
descriptions that #1284 reports. The `display_name` edge is recorded as
a comment on #1284 rather than a new issue, since it is the same route
and the same family, and the guard it asks for (assert no listed alias's
`display_name` or `summary` matches a known provider name) belongs where
the catalog is listed, not in this translation.

## Effect on #1278

This should flip all four `xfail` markers in
`packages/sdk-tests/python/tests/test_anthropic_messages.py` to real
assertions. Not touched here, since that branch is not mine to edit.

- `test_streaming_event_sequence_integrity` (#1274)
- `test_tool_choice_none_forbids_tool_use` (#1260)
- `test_count_tokens` (#1261)
- `test_models_list` (#1259)

One caveat on the last one. `client.models.list()` will now authenticate
and return the Anthropic list envelope, but the SDK's typed page model
may expect fields this gateway has no source for. If that assertion
still fails after this merges, it is a narrower follow-up about specific
fields, not the auth and envelope defects #1259 reported.

## Scope deliberately not taken

#1260 also observes that a zero-content streaming turn emits
`message_start`, `message_delta`, `message_stop` with no
`content_block_start`/`stop` pair. That is left alone: the streaming
path already emits `"content":[]` in `message_start`, so an SDK folding
the stream back into a message gets an empty array rather than null, and
Anthropic's own specification does not require a content block on a turn
that produced no content. The reported crash is in the non-streaming
builder, which is what this changes.

The neighbouring case where `message_start` never fires at all is a
different defect, was not what #1260 reported, and is now fixed here
rather than deferred. See Review round two above.

## Test plan

- [x] `go test ./apps/edge-api/... -count=1 -short` through the
toolchain container, green, re-run after the review-round-two fixes
- [x] `gofmt -l apps/edge-api` clean for every touched file
- [x] Twelve-way mutation check, all three tables above
- [ ] CI required checks green
- [ ] `packages/sdk-tests` Anthropic conformance suite re-run against a
live stack once #1278 merges, to confirm the four xfails flip

## Buglog entry

```json
{"id":"bug-2026-08-28-anthropic-sdk-wire-conformance","date":"2026-08-28","title":"Anthropic surface broke the real SDK four ways: content_block_start dropped text, empty completions serialized content null, count_tokens 401d API keys, /v1/models rejected x-api-key and answered OpenAI-shaped","error_message":"TypeError: unsupported operand type(s) for +=: 'NoneType' and 'str' in anthropic/lib/streaming/_messages.py accumulate_event; TypeError: 'NoneType' object is not iterable on msg.content; anthropic.AuthenticationError 401 {\"type\":\"error\",\"error\":{\"type\":\"authentication_error\",\"message\":\"missing user\"}} on count_tokens; anthropic.AuthenticationError 401 invalid_api_key on models.list()","root_cause":"Two families. (1) Go zero-value JSON encoding dropped values the Anthropic wire contract requires to be explicitly present: omitempty on StreamContentBlock.Text dropped the required \"text\":\"\" from every text content_block_start, and a nil []ResponseBlock in FromOAIResponse marshaled to null instead of []. Neither is visible to a Go-side assertion, since len() cannot distinguish nil from empty and a \"\" field looks correct until encoding/json runs. (2) The Anthropic surface's auth model only ever covered POST /v1/messages: count_tokens is the only route that does not delegate to the chat chain so it never gained an API-key authority and recognised session principals only, and GET /v1/models never got the leaf APIKeyNormalizer /v1/messages carries, which matters because the global one in authSelectorMiddleware only exists when jwtMW is non-nil.","fix":"StreamContentBlock.Text became *string and gained Input json.RawMessage set to {} on tool_use starts, matching the live streaming spec. FromOAIResponse seeds Content with an empty non-nil slice at construction so both exits are covered. anthropic.Deps.AuthorizeAPIKey added and count_tokens accepts either principal, fail-closed when nil, refusals reshaped from the shared WriteAuthFailure mapping. GET /v1/models registered through modelsHandler applying APIKeyNormalizer at the leaf plus a ModelsCompat wrapper that answers Anthropic-shaped callers with the Anthropic list and error envelopes while leaving OpenAI callers byte-identical.","tags":["anthropic","sdk-conformance","edge-api","json-encoding","omitempty","auth","streaming","provider-blind"],"issues":[1274,1260,1261,1259],"pr":"fix/anthropic-stream-content-block-start"}
```
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant