Repository navigation
fix: give supabase-storage the S3 protocol credentials its clients sign with (#1282) - #1368
Conversation
…gn with POST /v1/files returned 500 on the demo box for every upload, and with it every Batches input object and every artifact write, because the self-hosted Supabase Storage service was never given S3_PROTOCOL_ACCESS_KEY_ID or S3_PROTOCOL_ACCESS_KEY_SECRET. Without that pair Storage cannot recompute an AWS SigV4 signature at all, so it refuses every signed request with 403 AccessDenied and the body "Missing S3 Protocol Access Key ID or Secret Key Environment variables", while the container reports healthy, the buckets exist and the client-side configuration looks complete. The edge log line added for issue #1255 is what surfaced the real reason. The credentials come from the same S3_ACCESS_KEY and S3_SECRET_KEY the consumers already read. A separate pair would let the signing half and the verifying half drift apart with no boot error on either side, and the only symptom would be a 403 from a service that looks configured. S3_PROTOCOL_PREFIX is the second half of the fix. Caddyfile.supabase strips /storage/v1 with handle_path before Storage sees the request, and Storage recomputes the canonical URI as prefix plus the path it received, so an empty prefix fails every request with SignatureDoesNotMatch instead, which reads as a wrong key rather than a wrong path. The recipe .env.example and docker-compose.yml carried, pointing both S3 variables at the service_role key, never worked and could not have: a service_role JWT is not an S3 credential in single tenant mode, and using one value for both halves discloses it, since SigV4 sends the access key id in the clear inside the Authorization header. Both files now say to generate a distinct pair. Two guards in scripts/test_selfhost_supabase_seam.py, which CI already runs through make test-scripts: one fails if Storage stops being handed the same credentials the consumers sign with, the other if the prefix, the endpoint and the gateway route stop agreeing. Both were confirmed to go red against the unfixed compose file. Refs #1282, #1324
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
|
Warning Review limit reachedNext included review available in 58 minutes. View limit detailsLimit details: You’ve used the included review currently available. You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository. Review configuration: ⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Pro Plus Run ID: 📒 Files selected for processing (6)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
Correction to the box state note in the description: the temporary |
…ng for it
Review finding. A fresh EnterpriseEdge install still landed on a broken
upload after the compose fix, because scripts/install.sh prompted for
S3_ACCESS_KEY and S3_SECRET_KEY as required free text and labelled the
endpoint with a hosted project URL. That installer starts the enterprise
profile, which runs supabase-storage in the stack, so all three values
describe that in-stack service and none of them exist on a hosted project.
The prompt is what steered operators to the service_role key, which Storage
does not accept as an S3 credential.
The pair is now generated with openssl the way the installer already
generates its other secrets, and the endpoint defaults to the gateway form
Caddy serves. Nothing is left for the operator to look up.
Also tightens the new seam guard: it matched a bare prefix, so
${S3_ACCESS_KEY_ID} would have satisfied an assertion about
${S3_ACCESS_KEY}, which is precisely the drift it exists to catch. Confirmed
red against that substitution.
Refs #1282
Pre-merge reviewTwo streams. One ran, one refused. Antigravity (gemini-3.1-pro-high, high effort): RANSix points put to it: the shared-variable choice, the blast radius of the Finding 1, fresh EnterpriseEdge install still broken. Accepted, fixed in ea8451e. One part of the finding was wrong on the mechanism and worth recording so nobody re-derives it: it claimed Finding 2, the Finding 3, the guards asserted text rather than the property. Partly accepted, fixed in ea8451e. The variable check matched a bare prefix, so Findings 4, 5 and 6: nothing wrong. It independently confirmed the security reading (the service_role JWT was travelling in the clear in the SigV4 CodeRabbit CLI: SKIPPED, not a passRecorded as SKIPPED rather than clean. Nothing from this stream was read, so its absence is not evidence of anything. |
TestHandleChat_SuppressedChunkUsageNotAccounted asserted that the response body contains no "999", meaning the bogus usage numbers on the frame that must be suppressed. Every frame this handler emits carries a generated ragchat-<uuid> id, and a hex uuid contains that substring often enough to turn the guard red with nothing wrong in the code under test. It did exactly that on this branch, on ragchat-39990550-3e90-48db-b254-c54af332a90d, in CI run 33238715016, on a change that touches no Go code at all. The check now names the fields rather than the bare number. The marker assertion beside it is unchanged and still covers the same frame, so the guard loses no coverage: the suppressed chunk carries both. A flake inside a regression guard is worse than no guard, because the next red on it gets assumed to be the flake.
|
CI note. The Fixed in d23392c rather than re-run and ignored: the check now names the usage fields instead of the bare number, and the |
…easuring auth order Reading the assertions rather than the test names, which is what the previous pass should have done: three of the five live-conformance failures were not pre-existing failures at all. All three reported "Expect test to fail", the error vitest raises when an it.fails test PASSES. files.test.ts and batches.test.ts both carried an issue #1324 marker for the same environment defect: supabase-storage never received S3_PROTOCOL_ACCESS_KEY_ID and S3_PROTOCOL_ACCESS_KEY_SECRET, so it refused every SigV4 request with 403 while reporting healthy. PR #1368 supplied them. Both markers come off, exactly as their own comments said they should. The third marker in batches.test.ts stays, because that test still fails, but its comment no longer blames #1324. The sibling test performs the same upload and passes, so the remaining blocker is downstream of the upload. The likely candidate, recorded as a lead rather than a diagnosis, is the provider batch capability gap in CLAUDE.md Known Issue 4. usage-accounting.test.ts carried an issue #1317 marker for a terminal usage chunk that never arrived. #1317 is closed, PR #1334 relays the frame, and the first live run afterwards reported the unexpected pass. Marker off, assertions unchanged, so a documented gap becomes a plain regression guard. error-shape.test.ts asserted 404 for an unknown endpoint while sending no Authorization header, so it measured auth ordering instead. Everything under /v1/ passes the auth selector before the mux, which is what the real OpenAI API does as well, so an unauthenticated request is answered as unauthenticated whatever path it names. Adding the header its sibling test already sends makes it measure its own subject. Verified against the deployed gateway: GET /v1/nonexistent with an hk_ bearer returns 404 with code unknown_endpoint. The key need not be valid, because the unsupported-endpoint middleware sits outside the mux and so runs ahead of any per-route authorizer. Issue #1377 is unaffected and stays open: it is about /v1/audio/voices, a route deliberately registered without an authorizer that this same ordering makes unreachable.
…e true (#1381, #1382, #1324) (#1383) Follow-up to #1371, which merged before it could be measured live. Refs #1381, refs #1382, refs #1324, refs #1377. Two jobs here, and they pull in opposite directions on purpose: 1. Two markers this repo had are restored, because those tests genuinely still fail. 2. Three markers this repo had are removed, because those tests genuinely now pass. Same file family, opposite treatment, one rule: **a marker states what is true today, and it comes off the moment it stops being true in either direction.** An `it.fails` on a passing test is not a harmless leftover, it is a red badge on good news, and it teaches the next reader to skim the suite instead of reading it. ## Restored markers: audio, on #1381 The voice half of #1318 is fixed and proven live. The call now gets past the voice check and the upstream refuses the next thing: ``` litellm.exceptions.BadRequestError: litellm.BadRequestError: GroqException - response_format must be one of [wav] ``` The OpenAI SDK omits `response_format`, the OpenAI default is mp3, and the route accepts only wav. LiteLLM classifies that as a BadRequestError and answers 500 anyway, so the 400 relabelling from #1371 never gets the chance to help. Full write-up in #1381. The transcription-only test stays a plain `it`, so a defect there cannot hide behind the round-trip marker. ## Images: keeps `it`, drops the retries The empty 200 is gone. The box logged the guard firing: ``` provider_blind_upstream_error alias="hive-auto" status=502 raw_message="{\"created\":1787988601,...,\"data\":[],...}" ``` `maxRetries: 0` on that one call. A 5xx is retryable by definition and this SDK retries twice, but a route whose capability flag is a carried-forward legacy flag can never succeed on a retry. On run 33240963131 the first attempt refused correctly in about four seconds and the test still hit its 60 second ceiling, because the later attempts hung upstream. With one attempt it passes. What remains open is the catalog decision, tracked in #1382. ## Retired markers: files, batches, usage accounting These three reported `Expect test to fail`, which is what vitest raises when an `it.fails` test PASSES. My first pass called them pre-existing failures in untouched areas, by reading the test names instead of the assertions. That was wrong, and the phrase was doing load-bearing work it had not earned. - **`files.test.ts`** and **`batches.test.ts` (unsupported endpoint value)** both carried an issue #1324 marker for the same environment defect: `supabase-storage` never received `S3_PROTOCOL_ACCESS_KEY_ID` and `S3_PROTOCOL_ACCESS_KEY_SECRET`, so it refused every SigV4 request with 403 while reporting healthy. PR #1368 supplied them. Both markers come off, exactly as their own comments said they would. - **`usage-accounting.test.ts`** carried an issue #1317 marker for a terminal usage chunk that never arrived. #1317 is closed, PR #1334 relays the frame, and the live run reported the unexpected pass. Marker off, assertions untouched: a documented gap becomes a plain regression guard. The third marker in `batches.test.ts` (submit, retrieve, cancel) **stays**, because that test does still fail. Its comment no longer blames #1324 though: the sibling test does the same upload and passes, so the blocker is downstream of the upload. The comment now names the likely candidate as a lead rather than a diagnosis (the provider batch capability gap, CLAUDE.md Known Issue 4) and says outright that nobody has seen the actual error yet, because an `it.fails` marker swallows it. ## error-shape: not a marker, a test that measured the wrong thing It asserted 404 for an unknown endpoint while sending no Authorization header, so what it actually measured was auth ordering. Everything under `/v1/` passes the auth selector before the mux, so an unauthenticated request is answered as unauthenticated whatever path it names, which is also what the real OpenAI API does. Adding the header its sibling test in the same file already sends makes it measure its own subject. Verified against the deployed gateway before changing it: ``` GET /v1/nonexistent -H "Authorization: Bearer hk_definitely_not_a_real_key" 404 {"error":{"message":"Unknown endpoint: GET /v1/nonexistent","type":"invalid_request_error","param":null,"code":"unknown_endpoint"}} ``` The key does not need to be valid: the unsupported-endpoint middleware sits outside the mux and therefore ahead of any per-route authorizer. No marker was added, because there is no defect here to track. **#1377 is unaffected and stays open.** It is about `/v1/audio/voices`, a route deliberately registered without an authorizer that this same ordering makes unreachable, which is a real defect and a different one. ## The one genuine red left: sampling-params `honors temperature and top_p without erroring` timed out at the suite-wide 60 second budget. Not a marker problem and not something this PR touches: it is latency on the pooled free alias, whose sibling `honors max_tokens on the pooled alias` took 17.9 seconds in the same run. One sample, so no issue filed yet; if it repeats it wants either its own issue or a per-test timeout, and I would rather it stay red once than be marked on a single observation. ## Verification Re-labelled to run the live suite from this branch. Previous run 33241302607 is the before-picture. Local `node --experimental-strip-types --check` passes on all six touched files; the always-on `sdk-tests-js` job runs the same suite against a throwaway stack.
) (#1390) Fixes #1324. Fixes #1380. Diagnoses #1370 and leaves it open with the cause recorded. ## Was restoring this coverage worth it? It caught a real wire-format bug on its first sighted run. That is the short answer to the only question that matters here, and it is evidence rather than principle. A suite that had been blind for four months was given a real object store to talk to. The first thing it found was that **every signed S3 request this product makes was going out in absolute form** (`PUT http://host/s3/bucket/key HTTP/1.1`), which RFC 9112 reserves for requests **to a proxy**. An origin server expects `PUT /s3/bucket/key`. The cause was `URL.Opaque`, set for the SigV4 signer and never cleared, which `net/url.RequestURI` returns verbatim with the scheme glued back on. It was invisible on the deployed box for one reason only: every S3 request there passes through Caddy, which normalizes the target before proxying it on, so Supabase Storage never saw the malformed form. Reached directly, Storage refuses it with a 500 that names neither the path nor the credential. Nothing in four months of green CI could have found this, because CI was pointed at a hostname that stopped resolving in August. Details, the wire capture, and the fix are in the pinned comment on this pull request. ## What was actually wrong, and it is one thing rather than three The `S3_ENDPOINT` repository secret was last written on `2026-04-21T23:15:49Z`, four months before the August self-hosted Supabase cutover, and still names the Supabase Cloud project this repository left. That project is deleted, so the hostname does not resolve. Confirmed with `gh secret list`, along with the rest of that April-vintage family: `S3_ACCESS_KEY`, `S3_SECRET_KEY`, `S3_REGION`, `S3_BUCKET_FILES`, `S3_BUCKET_IMAGES`, `SUPABASE_URL`, `SUPABASE_DB_HOST`. **#1324.** `live-integration` reads those six at job level (`ci.yml:1221`), so every `POST /v1/files` died at the socket. The Files and Batches conformance tests were carried as `it.fails` markers, which means the suite reported green for four months while covering nothing. **#1380 names the wrong job, and this is worth stating plainly rather than quietly fixing.** `ci.yml:2026` and `ci.yml:2565` are `web-e2e` and `interaction-coverage`, both Playwright. Neither runs `sdk-tests-js`; those run in `live-integration` (`ci.yml:1612`), which was reading the dead secret. So #1380's conclusion, that CI cannot fail on the storage path, is correct, but its cause is #1324's secret and not the discard port. I read the comment at `ci.yml:2016` before touching either line, as asked, and its author's three reasons all survive scrutiny, so the stub stays and the comment now names the job that does exercise storage. **#1370 is a third victim of the same secret family.** The nightly does not report `skipped` at run level: every scheduled run since 2026-08-09 reports `failure` (runs 33197858823, 33097409194, 32939092305, 32817812570, 32698750004, 32623212748, 32557183055, 32455134986, 32340347866). What is skipped is the work. Job log 98939498320 shows the cause: ``` socket.gaierror: [Errno -2] Name or service not known urllib.error.URLError: <urlopen error [Errno -2] Name or service not known> File "scripts/seed-owui-e2e-user.py", line 709, in main ``` The seeder resolves `secrets.SUPABASE_URL`. It dies at step 5, so steps 6 through 15 including the entire Playwright suite are skipped, and the only red left is step 16 failing on `browserType.launch: Executable doesn't exist`, because the step that installs browsers was one of the skipped ones. A reader who opens the run sees a Playwright error and no mention of a dead hostname. That is what made a week of failures read as flake. ## The decision, and why it is not a new secret **Can a GitHub-hosted runner reach the box's storage endpoint? No, and it must not.** `deploy/docker/Caddyfile.supabase` puts `/storage/v1` on the internal listener only. The public listener is a separate port carrying `/auth/v1` minus its admin and self-service routes, and the file states the constraint directly: > Consequence for whoever does the cutover: publish or tunnel port 8080 ONLY. Ports 80 and 443 are the in-network surface and must never be exposed. So repointing the secret is blocked by a deliberate security posture, not by an unfinished chore. Even with that posture relaxed, CI would then be reading and writing the live demo box's buckets, which is the arrangement the `web-e2e` comment already records as a defect that was fixed. **MinIO was considered and rejected, and not on the grounds you might expect.** The standing no-MinIO line lives in `docker-compose.enterprise.yml:312` and in CLAUDE.md as a statement about what a deployment persists objects into. On a literal reading it does not reach a container that lives for six minutes and holds nothing, and I do not think it does. MinIO is rejected on a different ground: the defect that actually took production down, #1282, was a `supabase-storage` server-side credential defect. MinIO has no `S3_PROTOCOL_*` concept at all, so a MinIO fixture is structurally incapable of catching that or its next relative, and picking it would rebuild a weaker version of the blindness this PR exists to remove. **What this does instead:** `live-integration` stands up the real `supabase/storage-api`, at the same immutable digest `docker-compose.enterprise.yml` pins, on the throwaway Postgres the job already boots. That database already has everything Storage needs: `deploy/supabase/init/00-extensions.sql` creates the `storage` schema, the Supabase roles, and the default privileges in `storage`, and its own comments already name this consumer. Nothing external is dialed, no secret is involved, and a Dependabot-triggered run behaves identically to every other run. Verified before writing a line of workflow YAML: the image boots against a bare `pgvector/pgvector:pg17` with only `00-extensions.sql` applied, runs its own migrations into `storage`, and serves a full signed round trip on `/s3/hive-files/<key>` with no PostgREST and no Caddy. **Deliberate gap, stated rather than buried.** Production reaches Storage through the gateway, which strips `/storage/v1`, so Storage is given `S3_PROTOCOL_PREFIX=/storage/v1` to rebuild the canonical URI. The fixture has no gateway, so it runs with an empty prefix and the endpoint ends in `/s3`. That prefix agreement is therefore not exercised live here. It is not unchecked: `scripts/test_selfhost_supabase_seam.py` asserts the prefix, the endpoint and the Caddy route all agree, and CI runs it through `make test-scripts`. Adding Caddy to the fixture would add a component and its failure modes for a seam that already has a guard. ## No secret needs creating, rotating or repointing Stated explicitly because the obvious reading of #1324 is "fix the secret". It is not. The three workflows that read `secrets.S3_*` now read none of them: - `live-integration` takes its values from the fixture. - `agent-visual-proof` takes the same secret-free stub the Playwright jobs use, because it touches no object either. Its `github.actor != 'dependabot[bot]'` guard is byte identical after the change; `git diff` shows no line touching it. - `owui-nightly` keeps its six, deliberately, and that needs the explanation below. Nothing else in this repository reads the six `S3_*` secrets after this change. ### Why `owui-nightly` keeps them, and what the lint taught me My first attempt at that file added a preflight step resolving `secrets.SUPABASE_URL`, failing with a message naming the dead project. `tools/lint-no-dead-supabase-secrets.mjs` rejected it, correctly, and it is worth restating why rather than just complying. That guard pins the exact per-secret reference counts for the file rather than exempting the file, so a freshly copy-pasted `secrets.SUPABASE_*` line cannot ride in on an exemption. My step took `SUPABASE_URL` from 3 to 4. Improving an error message is not a reason to widen an exemption documented as one that must only ever shrink. The diagnostic therefore moved into the existing `Reconstruct SUPABASE_DB_URL` step, which already reads `SUPABASE_DB_HOST`, and resolves that instead. Same deleted project, same answer, zero new references. **The `PENDING` map is unchanged and the lint passes on the audited counts exactly as they stand:** ``` pending owui-nightly.yml: 13 audited reference(s) to the retired project's secrets, tracked in #1055 ok: 16 workflow file(s) scanned, 1 pending, no unexempted reference to the retired project's secrets ``` I checked whether this PR could take that file to zero and retire the entry, which is what #1055 actually wants. It can in principle, by the route the linter itself names: `scripts/ci-throwaway-db.sh` plus `scripts/ci-supabase-stack.sh` plus this PR's new `scripts/ci-object-store.sh`, the arrangement `web-e2e` already has. I am not doing it here, for reasons I would rather state than have discovered: - It replaces every credential path in a job whose suite has not passed since 2026-08-18, so a configuration fix is unlikely to be the last thing wrong with it. That is #1370's own fix, not a storage PR's. - It is unverifiable from this branch in any way I would trust. The job runs 45 minutes and builds Open WebUI from source, and I would be asserting it works on one dispatch of a suite with unknown other breakage. This repository has a standing lesson about exactly that shape of sincere false green. - Its S3 half specifically cannot take the stub the other jobs take. `web-e2e`, `interaction-coverage` and `agent-visual-proof` can be stubbed at the discard port because they touch no object. This one boots Open WebUI, whose knowledge and document surfaces do write objects, so a stub would convert a blocked suite into a suite that runs and fails on storage. It needs the real fixture, which needs the throwaway Postgres, which is the #1370 change. The block now carries that reasoning inline, so the next person to open it does not rediscover any of it. ## Proving the signal can actually fail Every restored or new assertion was verified red before being kept. **The fixture's own round-trip check.** Booted with `S3_PROTOCOL_ACCESS_KEY_ID` and `S3_PROTOCOL_ACCESS_KEY_SECRET` removed, it exits non-zero with, verbatim, the error PR #1368 read off the production box: ``` ::error::the throwaway object store answered its health check but could not complete a signed PUT and GET. That is the shape of issue #1282: Storage reports healthy and refuses every signed request. aws: [ERROR]: An error occurred (AccessDenied) when calling the PutObject operation: Missing S3 Protocol Access Key ID or Secret Key Environment variables ``` With the credentials present, same database, same command: `round trip OK` and seven `KEY=value` lines. The ordering guard was also verified red, by running it against a Postgres with no `storage` schema. **The workflow rot guard**, `apps/web-console/tests/unit/ci-live-integration-storage.test.ts`, run against the pre-change `ci.yml`: ``` × reads no S3 repository secret × stands up a real object store fixture × keeps the S3 values out of the job-level env block Tests 3 failed | 1 passed (4) ``` and green against the post-change file. All four of its assertions matter: a returning secret means the job dials a dead host again, a missing fixture step means it dials nothing, and a job-level `S3_*` entry silently overrides what the step writes to `$GITHUB_ENV`, which is the exact trap that got `SUPABASE_*` and `HIVE_API_KEY` removed from this job rather than reassigned. **The digest drift guard**, added to `scripts/test_selfhost_supabase_seam.py`, verified red by bumping the fixture pin one patch version: ``` AssertionError: the CI object-store fixture and the deployed stack pin different Storage images. compose: supabase/storage-api:v1.11.13@sha256:1e85dad... fixture: supabase/storage-api:v1.11.12@sha256:1e85dad... ``` Without it, the claim that CI tests what the box runs is a comment rather than a property. **End to end in CI.** The `it.fails` markers came off, which is itself a two-way check: `it.fails` reports an unexpected pass as a failure, so a premature removal and a premature marker both go red. The deliberate break and its revert are recorded below once both runs land. The break was made in the code path rather than in the fixture, so what is proven is that a real storage defect turns the check red: `apps/edge-api/cmd/server/main.go` handed the files handler a bucket that does not exist, which is the shape of the family #1282 belongs to, where the request reaches a live server and is refused. **Run 33243735217, commit `1e160189b`, the deliberate break. `Live integration (SDK tests + smoke)` (job 99077797492): failure.** The fixture itself came up clean on the hosted runner, which was the one thing that could not be proven locally: ``` bucket hive-files: created bucket hive-images: created round trip OK ``` and then the restored coverage did exactly what it exists to do: ``` ❯ tests/files/files.test.ts (2 tests | 1 failed) × Files > uploads, lists, retrieves metadata, downloads content, then deletes a file → 500 Failed to store file ✓ Files > rejects an invalid purpose with a structured 4xx error ❯ tests/batches/batches.test.ts (2 tests | 1 failed) × Batches > rejects an unsupported batch endpoint value with a structured 4xx error → 500 Failed to store file ``` Both failures name the storage path. Before this PR the same break would have changed nothing, because the endpoint was already dead and both assertions were carried as expected failures. That run was later cancelled by the `ci-${{ github.ref }}` concurrency group when the revert was pushed. Its `Live integration` job had already concluded, so the evidence above is from a completed job and not from a cancelled one. It also failed `Repo policy lints (tenant + audit)` on the dead-secret linter, which is the finding described in the section above and is fixed in `4411d7af9`. **Run 33244166031, commit `b8b8ae75e`, break reverted plus the lint fix.** Result recorded here once it lands. **Run 33246016113, commit `9994dba07`.** The storage half is green: ``` == sdk-tests-js (exit 0) ✓ tests/files/files.test.ts (2 tests) ✓ tests/batches/batches.test.ts (2 tests) Test Files 22 passed (22) Tests 58 passed | 1 skipped (59) ``` | Run | Commit | Expectation | Result | | --- | --- | --- | --- | | 33243735217 | `1e160189b` | `Live integration` red, naming the storage path | failure on `500 Failed to store file` | | 33246016113 | `9994dba07` | Files and Batches green | both suites pass, 58/58 in JS | Between those two the restored coverage did its job and caught a real defect, which is why the middle runs are not a clean green: see the pinned comment on this pull request. `packages/storage` was emitting an absolute-form request target on every S3 request, which Caddy normalizes on the box and Supabase Storage refuses outright. That is fixed here, verified red then green, then end to end against a real Storage container. `Live integration` is still red overall, on a strict `xfail` for issue **#1274** in the Python suite that now passes. It is pre-existing, present in run 33243735217 before any of this branch's storage work, unrelated to storage, and deliberately left alone: clearing a strict `xfail` on one green sample is how an intermittent defect gets recorded as fixed. ## Required contexts: ran, or satisfied by a skip Filled in from the PR's own checks once they resolve, per the standing hazard that a `skipped` required check satisfies branch protection and a non-required check's absence is invisible on the page. ## Also in this PR, at the coordinator's request `.claude/skills/worktree-compose-stack.md` advertised "the current dead-Supabase-host blocker" in its frontmatter description, which is the line the skill router matches on, and named a specific Supabase Cloud hostname in `.env` as the thing stopping a local stack, citing #1254. Both halves are now false, and the way they are false is expensive: #1254 is closed, and `.env` no longer contains that hostname. Verified directly, values masked: `SUPABASE_URL`, `SUPABASE_DB_URL`, `S3_ENDPOINT` and `NEXT_PUBLIC_SUPABASE_URL` are all present with a zero-length value. An agent reading the old text greps for the hostname, finds nothing, and concludes either that the skill is wrong or that a local stack should now work. Neither is true. The blocker changed shape rather than going away, and the two shapes produce different errors: - Then: `control-plane` came up and warned `database unreachable after 6 attempt(s)`, naming no cause. - Now: `control-plane` never reaches the database check. `loadStorageConfigFromEnv` requires the S3 names to be non-empty and `log.Fatalf`s, so it exits at boot with `storage unavailable: missing S3_ENDPOINT, S3_ACCESS_KEY, S3_SECRET_KEY, S3_REGION`. The section now leads with "do not grep `.env` for a Supabase hostname to confirm this", states both shapes, and points at `scripts/ci-supabase-stack.sh` plus the new `scripts/ci-object-store.sh` as the sandbox-workable alternative. It belongs in this PR because it is the same root cause at a different site: one deleted Supabase Cloud project, still referenced by stale configuration in more than one place. The CI-side instance is the secret this PR removes. ## Not in scope, deliberately - **Fully fixing #1370.** See the section above on why, and what it would take. What this PR adds there is a resolution check inside a step that already reads the host, failing in seconds with a message naming the stale secret family and the issue, instead of spending 45 minutes producing a misleading missing-browser cascade. The issue stays open with the cause recorded. - **The accumulating `owui-e2e-shim` API keys** noted in #1370. Separate hygiene finding. - **`/v1/batches` success path.** Still not exercisable with the current provider mix, CLAUDE.md Known Issues item 4, unchanged. The submit, retrieve and cancel path is what the marker was hiding and is what comes back. ## Effect on the demo None. `deploy-demo-box.yml` is untouched, and no file this PR changes is read by it. ## Plan `ObsidianVault/hive/plan-2026-08-29-ci-object-storage-blindness.md`, with the requirements and acceptance criteria folded in. ## Issue #1274's stale marker, cleared with two-sample evidence `Live integration` was red on one thing that was not storage: a `strict=True` `xfail` in `packages/sdk-tests/python/tests/test_anthropic_messages.py` that now passes, rendering as `XPASS(strict)`. That is the same defect family as the three `it.fails` markers #1383 cleared this morning, and a marker states what is true today. **What it tracked and what fixed it,** so the closure is traceable rather than "it passes now": #1274 was `content_block_start` omitting an empty `text` field, because `json:"text,omitempty"` drops a Go string's empty value as readily for the required-but-empty case as for unset. The Anthropic SDK seeds its accumulator snapshot from that field and then does `content.text += delta.text`, so a missing key made that a `None += str` TypeError on every streamed text response. **PR #1296, commit `0b4e5ffed`,** fixed it by making `Text` a pointer so the required-but-empty value reaches the wire. The comment on `StreamContentBlock` in `apps/edge-api/internal/anthropic/types.go` records that, and issue #1274 is closed. **Not cleared on one sample.** It reported `XPASS(strict)` on five independent runs across five commits, two of them after rebasing onto #1410's classifier fix, with no run in which it failed legitimately: | Run | Commit | | --- | --- | | 33244935750 | `f8dddd47e` | | 33246016113 | `9994dba07` | | 33247050365 | `3060119a5` | | 33247571957 | `11d7e2eae` (post-rebase) | | 33248088001 | `eac8729b4` (post-rebase) | `gh run rerun --failed` refuses on this branch ("its workflow file may be broken", the known behaviour when a run's workflow file has changed), so the second sample comes from separate runs on separate commits rather than a re-run of one. That is the stronger form of the evidence, not a substitute for it. The `run-live-integration` label stays on. Removing it would make a required check skip, and a skipped required check satisfies branch protection: that would trade a loud failure for a quiet absence on the exact lane this pull request exists to restore. ## Correction posted to #1380 That issue attributes the blindness to the `http://127.0.0.1:9` stub. The conclusion is right and the cause is not: `ci.yml:2026` and `ci.yml:2565` are `web-e2e` and `interaction-coverage`, both Playwright, and neither runs `sdk-tests-js`. The SDK suites run in `live-integration`, which was reading the dead secret. Corrected on the issue itself rather than only here, so the next reader does not inherit the wrong model: #1380 (comment) ## `ci.yml` regions this branch touches Three, all narrow, listed so a rebase order against #1410 and #1417 can be worked out: 1. The `live-integration` job env block, where six `secrets.S3_*` entries are removed and replaced by a comment. 2. One new step in `live-integration` after the throwaway-database step, one new teardown step, one new failure-dump step, and one sentence added to the existing `Pre-verify secrets reached runner env` comment. 3. Comment-only additions inside the `web-e2e` and `interaction-coverage` S3 stub blocks. No value changes in either. It does not touch the `Repo policy lints` region that #1417 edits, and it does not touch the upstream-refusal classifier that #1410 changed. Rebased onto `0d089744a` cleanly with no conflicts. ## Buglog entry ```json {"id":"bug-2026-08-29-ci-object-storage-dead-secret","date":"2026-08-29","title":"CI object storage endpoint pointed at a deleted Supabase Cloud project, so the file upload path could not fail","error_message":"file upload to object storage failed bucket=hive-files error=\"Put \\\"https://<project>.supabase.co/storage/v1/s3/hive-files/...\\\": dial tcp: lookup <project>.supabase.co: no such host\"","root_cause":"The S3_ENDPOINT repository secret was last written 2026-04-21, four months before the August self-hosted Supabase cutover, and still named the deleted Supabase Cloud project. The live-integration job read it at job level, so every POST /v1/files in CI died at DNS and the Files and Batches conformance assertions were carried as it.fails markers rather than as coverage. The same stale secret family also killed the OWUI nightly at its seeding step via SUPABASE_URL (#1370), which then skipped the entire Playwright suite and left a misleading missing-browser cascade as the only visible red. Issue #1380 attributed the blindness to the http://127.0.0.1:9 stub in ci.yml, but that stub is in two Playwright jobs that run no SDK suite.","fix":"Removed every secrets.S3_* reference from live-integration and agent-visual-proof. live-integration now boots its own throwaway supabase/storage-api at the digest docker-compose.enterprise.yml pins, on the throwaway Postgres it already stands up, whose 00-extensions.sql already creates the storage schema and roles. The fixture asserts a signed PUT and GET before printing its values. The it.fails markers came off. Guards added: a unit test asserting the job keeps the fixture and no S3 secret, and a seam-test assertion that the fixture and compose pin the same image.","tags":["ci","storage","s3","supabase","dark-detector","stale-secret","sigv4"],"issues":[1324,1380,1370,1282]} ``` ```json {"id":"bug-2026-08-29-s3-absolute-form-request-target","date":"2026-08-29","title":"S3 client sent absolute-form request targets, which Supabase Storage refuses with an unnamed 500","error_message":"s3 PUT /s3/hive-files/<key> failed with status 500: <Code>InternalError</Code><Message>Internal Server Error</Message>; storage side: {\"code\":\"ERR_INVALID_URL\",\"input\":\"http://localhost:8080http://172.17.0.1:5000/s3/hive-files/<key>\"} TypeError: Invalid URL at SignatureV4.constructCanonicalRequest","root_cause":"packages/storage/signing.go SignHTTP sets req.URL.Opaque to \"//host/path\" so the AWS SigV4 signer receives an un-normalized path, and never cleared it. net/url.RequestURI returns Opaque verbatim and re-attaches the scheme when it starts with \"//\", so every signed S3 request went on the wire in absolute form (PUT http://host/s3/bucket/key HTTP/1.1). RFC 9112 reserves absolute form for a request to a proxy. The deployed box never saw it because every S3 request passes through Caddy, which normalizes the target before proxying. Reached directly, Supabase Storage builds its canonical URI as new URL(\"http://localhost:8080\" + prefix + request.url) and throws on the doubled origin.","fix":"Clear req.URL.Opaque after signer.SignHTTP returns. The signature is unaffected: the signer derives its canonical URI from Opaque while set, stripping the leading //host, and that string is byte identical to the EscapedPath the wire form falls back to once Opaque is empty. presignHTTP still keeps Opaque, because url.URL.String renders a bare-path Opaque as http:/s3/... and the //host form is what makes a presigned URL absolute. Four tests added, each verified red first.","tags":["storage","s3","sigv4","http","rfc9112","proxy-masked","dark-detector"],"issues":[1324,1380,1282]} ```
## Summary This is the batched buglog follow-up for the pull requests merged to `main` on 2026-08-29. Its diff is `.wolf/buglog.jsonl` and nothing else. Per `.claude/rules/openwolf.md`, every fixed bug, error, failed test or failed build must be logged, but the line may never be appended on a fix branch. `merge=union` in `.gitattributes` resolves concurrent appends locally and is ignored by GitHub's server side merge, so two branches that both appended land in hard conflict there. An unmergeable pull request gets no `refs/pull/N/merge`, no `pull_request` run and therefore zero checks, and the required status gate then blocks the merge for a reason the page never states (issue #873). Each fix accordingly carried its entry in its own pull request body, and this pull request copies them onto `main` in one batch, which the protocol explicitly prefers over one pull request per entry. ## Scope examined Fifty nine pull requests merged to `main` on 2026-08-29. Forty eight of them carried at least one entry, for eighty two entries in total. Thirty two of those were already on `main` and are skipped, leaving fifty appended here from thirty four pull requests. The largest block of skips comes from #1342, the equivalent batch for the 2026-08-28 merges, which merged earlier the same day and already landed thirty six entries covering #1257, #1268, #1276, #1277, #1287, #1292, #1293, #1294, #1296, #1301, #1303, #1305, #1313, #1335 and #1337. ## What landed Fifty entries appended, one JSON object per line, append only. The 232 pre-existing lines are byte identical to `origin/main` (verified by hashing the first 232 lines of the result against the base file). Every line in the resulting file parses as JSON and carries `error_message`, `root_cause`, `fix` and `tags`. | Source | Entries | |---|---| | #1083 | 2 | | #1277 | 1 | | #1278 | 1 | | #1298 | 1 | | #1334 | 1 | | #1336 | 3 | | #1343 | 1 | | #1346 | 1 | | #1351 | 1 | | #1365 | 2 | | #1368 | 1 | | #1369 | 1 | | #1371 | 3 | | #1375 | 3 | | #1376 | 1 | | #1378 | 1 | | #1379 | 2 | | #1388 | 5 | | #1389 | 3 | | #1390 | 2 | | #1393 | 1 | | #1394 | 1 | | #1410 | 1 | | #1417 | 1 | | #1421 | 1 | | #1423 | 1 | | #1424 | 1 | | #1426 | 1 | | #1429 | 1 | | #1431 | 1 | | #1433 | 1 | | #1434 | 1 | | #1436 | 1 | | #1439 | 1 | Entries are copied verbatim from their source pull request bodies. Nothing was rewritten, no field was invented, and no field was added. No JSON needed repair: all eighty two extracted entries parsed on the first attempt and all four required fields were present on every one. ## Merged pull requests that carried no entry Eleven of the fifty nine. Recorded here because the gap is itself the useful signal. | Pull request | Title | Assessment | |---|---|---| | #1013 | chore(deps): bump the go-minor-patch group across 1 directory with 4 updates | Dependabot bump, no defect fixed, no entry expected | | #1015 | chore(deps): bump the go-minor-patch group across 1 directory with 6 updates | Dependabot bump, no entry expected | | #1016 | chore(deps): bump golang from 1.26-alpine to 1.27-alpine in /deploy/docker | Dependabot bump, no entry expected | | #1218 | chore(deps): bump postcss from 8.5.19 to 8.5.26 in /apps/desktop | Dependabot bump, no entry expected | | #1219 | chore(deps): bump golang.org/x/crypto from 0.41.0 to 0.52.0 in /apps/control-plane | Dependabot bump, no entry expected | | #1342 | chore: batch buglog entries for the 2026-08-28 merges | The previous batch pull request itself, correctly carries no entry of its own | | #1364 | chore: remove four dead skills and record the patterns that cost time | Protocol gap. The body records patterns that cost time, which is the shape of a buglog entry, but none was written as one | | #1383 | test: retire stale expected-failure markers, restore the ones that are true (#1381, #1382, #1324) | Protocol gap. Stale `it.fails` markers reading as red is a real defect that was fixed here and should have carried an entry | | #1384 | docs: correct D-047, hive-auto reverted to variable pricing (D-059) | Decision ledger correction, arguably a documentation defect, no entry written | | #1387 | chore(deps): bump next from 15.5.23 to 16.3.3 in /apps/agent-console | Dependabot bump, no entry expected | | #1398 | docs: rescue the 2026-08-25 parity captures and add the 2026-08-29 QA matrix evidence | Documentation and evidence rescue, no entry written | Six of the eleven are Dependabot bumps and one is the previous batch, so the genuine protocol gaps are #1364, #1383, #1384 and #1398. Of those, #1383 is the one worth a follow-up: it fixed a real defect class (a stale expected-failure marker reads as a red "Expect test to fail" and gets dismissed as pre-existing) and left no record. ## Entries skipped as already present Thirty two. Thirty of them matched an entry already on `main` on `error_message`, `id` or `fix`. Two more from #1278 are semantic duplicates that an exact match would have missed, and were skipped after reading the landed entries they duplicate: - #1278's `streaming content_block_start omits text field` entry is covered by the consolidated `bug-2026-08-28-anthropic-sdk-wire-conformance` entry landed from #1296, whose root cause names the same `omitempty` on `StreamContentBlock.Text`. - #1278's `GET /v1/models leaked an upstream provider name` entry is covered by `BUG-1284`, landed from #1300, which names the same `public.model_aliases.summary` publication path. #1278's third entry, on `top_k` forwarding producing a 400, is not covered anywhere on `main` and is appended here. #1342 recorded #1278 as fully "merged into #1296", which was accurate for two of its three entries. ## Note on entry quality One appended entry is thin: #1277's parity re-score record carries `error_message` of `n/a` and a root cause of "console had no privacy/data-policy surface at all". It is a parity gap record rather than a defect record. It is included exactly as written rather than embellished, per the protocol's preference for the author's own words. ## Test plan - [x] Branch cut fresh from `origin/main`, diff is `.wolf/buglog.jsonl` and nothing else - [x] First 232 lines byte identical to the base file (md5 match) - [x] All 282 resulting lines parse as JSON and carry `error_message`, `root_cause`, `fix` and `tags` - [x] No `.wolf/` telemetry (`anatomy.md`, `memory.md`, `token-ledger.json`, `hooks/_session.json`, `buglog.json`) in the commit - [ ] The six required checks report green via the inert path allowlist in `.github/workflows/ci.yml` --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Fixes #1282.
Cause, read from the running box rather than guessed
The error log added for issue #1255 made this one command away. On the demo box:
That message is Storage's own, from
createServerSignaturein itssignature-v4plugin. In single tenant mode Storage verifies an S3 request by recomputing the AWS SigV4 signature fromS3_PROTOCOL_ACCESS_KEY_IDandS3_PROTOCOL_ACCESS_KEY_SECRET, and with neither set it cannot verify anything, so it refuses every signed request. Thesupabase-storageservice indocker-compose.enterprise.ymlhas never had either variable, confirmed against the running container's environment.This was configuration, not code.
edge-apiandcontrol-planewere correct throughout: correct endpoint, correct region, correct bucket names, buckets present as rows instorage.buckets, container healthy. Nothing on the client side could have been changed to make it work.The prior guidance made it worse rather than random.
.env.exampleanddocker-compose.ymlboth said to setS3_ACCESS_KEYandS3_SECRET_KEYto theENTERPRISE_SERVICE_ROLE_KEY, and the box followed it: both values were byte identical service_role JWTs. A service_role JWT is not an S3 credential in single tenant mode, and identical halves are their own problem, since SigV4 sends the access key id in the clear inside the Authorization header, so every request would have disclosed the service_role key to anything on the path.The change
supabase-storagenow receives:S3_PROTOCOL_ACCESS_KEY_IDandS3_PROTOCOL_ACCESS_KEY_SECRET, from the sameS3_ACCESS_KEYandS3_SECRET_KEYvariables the consumers sign with. Deliberately the same two: a separate pair lets the signing half and the verifying half drift with no boot error on either side, and the only symptom is a 403 from a service that looks configured.S3_PROTOCOL_PREFIX: /storage/v1.Caddyfile.supabasestrips that prefix withhandle_pathbefore Storage sees the request, and Storage rebuilds the canonical URI ass3ProtocolPrefix + request.url, so leaving it empty fails every request with SignatureDoesNotMatch instead, which reads as a wrong key rather than a wrong path..env.exampleand thedocker-compose.ymlcomment now say to generate a distinct pair, and say why.Guards
Two new checks in
scripts/test_selfhost_supabase_seam.py, which CI already runs throughmake test-scripts:S3_ENDPOINTthe operator is told to use, and the gateway'shandle_pathroute all agree.Both were confirmed to go red against the unfixed compose file before being kept.
Evidence against the deployed box
Same API key, same file, same public endpoint, before and after.
Before:
After:
Round trip, not just the write:
The two dependent surfaces
Batches: the storage half is fixed, one unrelated failure remains. A batch submission now gets all the way past every storage dependent stage: the input file uploads, ownership resolves,
validateJSONLdownloads and parses the object out ofhive-files, and the credit reservation is taken. It then fails incontrol-planeat route selection, identically for every alias offered by the deployment:That is alias independent, has nothing to do with object storage, and matches
CLAUDE.mdKnown Issues item 4: no provider in the current mix has a batch API. It is pre-existing and out of scope here. Worth its own issue, and worth noting that the live conformance suite will keep reporting Batches red for this second reason after this merges.RAG document upload never depended on object storage.
apps/edge-api/internal/ragforwards parsed content tocontrol-plane, which chunks, embeds and writes to pgvector. NoUploadcall anywhere in either package, and the chat front end stores its own uploads on local disk, with no S3 variables in the Open WebUI container at all. ProbingPOST /v1/rag/documentsagainst production returns 403ACCESS_DENIEDfromfeaturegate, which is the tenant feature flag for the probe account and not a storage failure. So the claim that this 500 also broke the RAG upload beat does not hold, and no tenant flag was changed to check.What was done to the running box, exactly
~/hive/.env:S3_ACCESS_KEYandS3_SECRET_KEYreplaced with a freshly generated distinct pair, replacing the two identical service_role JWTs. Previous file backed up on the box.deploy/docker/docker-compose.fix1282.ymlcarrying the same three variables this PR adds, so the fix could be proven ahead of the merge without touching a tracked file, which would have broken the deploy'sgit pull --ff-only. It is passed only to the manual command below and is deleted after verification; the deploy workflow's compose flags never reference it.docker compose ... up -d --no-deps --no-build supabase-storage edge-api control-plane. Three containers recreated and nothing else.supabase-storageto receive the new variables,edge-apiandcontrol-planebecause their S3 credentials changed. All three came back healthy within twenty seconds.Once this merges, the deploy recreates those services from the tracked file with the same values from the box's
.env, so the state converges rather than depending on the override.What this does not close
#1324 stays open, and this is not its fix. That issue is about CI, where the
S3_ENDPOINTrepository secret still names the deleted Supabase Cloud project. Production never pointed there: the box has pointed athttp://caddy-supabase/storage/v1/s3since the cutover, and its failure was the missing server side credentials, not a dead hostname. Two separate causes producing the same blind 500 in two places.I cannot close the CI half from here, and it is not a secret rotation either: the box's Storage is not reachable from a GitHub hosted runner, so there is no correct value for the current secret. The real repair is for the throwaway CI stack to run its own
supabase-storagecontainer with these same three variables. Related, and part of why nobody noticed:ci.ymlpoints the local stack'sS3_ENDPOINTathttp://127.0.0.1:9with a stub key, so the always onsdk-tests-jslane cannot exercise an upload either, and reports green while doing so.Also noticed and left alone:
control-planelogsauditarchive: cold storage configured bucket=hive-audit-cold, and no such bucket exists.supabase-initcreateshive-filesandhive-imagesonly. Nothing fails today because that cron only moves records older than ninety hot days, but it will fail the first time it runs. Separate issue, not folded in here.Buglog entry
{"date":"2026-08-29","title":"Supabase Storage refused every S3 request because the server had no S3 protocol credentials","error_message":"s3 PUT /storage/v1/s3/hive-files/... failed with status 403: AccessDenied: Missing S3 Protocol Access Key ID or Secret Key Environment variables","root_cause":"The self-hosted supabase-storage service was never given S3_PROTOCOL_ACCESS_KEY_ID, S3_PROTOCOL_ACCESS_KEY_SECRET or S3_PROTOCOL_PREFIX. In single tenant mode Storage recomputes the SigV4 signature from that pair, so with neither set it refuses every signed request while staying healthy with the buckets present. The documented recipe compounded it by telling operators to point both client variables at the service_role key, which is not an S3 credential and discloses itself through the access key id that SigV4 sends in the clear.","fix":"Hand supabase-storage S3_PROTOCOL_ACCESS_KEY_ID and S3_PROTOCOL_ACCESS_KEY_SECRET from the same S3_ACCESS_KEY and S3_SECRET_KEY the consumers sign with, plus S3_PROTOCOL_PREFIX=/storage/v1 to match the prefix Caddy strips. Corrected the credential guidance in .env.example and docker-compose.yml, and added two guards to scripts/test_selfhost_supabase_seam.py.","tags":["storage","supabase","s3","sigv4","compose","config","issue-1282"]}