Repository navigation
fix: stop reporting a gateway cooldown as an upstream refusal (issue #1374) - #1410
Conversation
The live-integration job's "Name the real cause when the upstream refused" step exists so a red check states its own cause (issue #1088). On 2026-08-29 it stated the wrong one. Run 33243396287 printed UPSTREAM REFUSAL, not a performance regression: the provider rate limited us (429). over a job whose actual failures were `500 Failed to store file` in two JS suites and one strict XPASS on issue #1274. Nothing upstream refused anything. The free pool was healthy: the job's own "Probe each free pool member" step reported 3 healthy and 1 unhealthy in that run and in 33244166031, and direct probes of the other two providers answered 200, with Groq reporting 955 of its 1000 daily requests remaining. Three log shapes were being read as upstream refusals when none is one. LiteLLM's `No deployments available for selected model ... cooldown_list=[...]` is the router declining to dispatch to deployments it has already benched. No upstream call is made at all, and in that run it carried alias `deepseek-v4-flash`, so a paid alias's gateway cooldown was reported as a refusal of the free pool. LiteLLM's `Cooldown Deployments=[...]` table quotes the exception that caused each cooldown, including `'status_code': '429'` and RateLimitError verbatim, and reprints it on every subsequent dispatch for as long as the entry lives. One past event re-reports itself indefinitely. `No fallback model group found` is emitted on any failed dispatch to a group with no configured fallback, whatever the underlying error. In the same artifact it appears on a deliberate over-length prompt, on deepseek not supporting image input, and on a Groq TTS suite asserting an invalid response_format is refused. All three are expected behaviour under test. The dead-member branch now requires the NotFoundError shape together with the free pool's own group name, which is the case its message actually describes and the shape recorded in apps/edge-api/internal/inference/retry_test.go. Corroboration that the classification was noise rather than a spent allowance: issue #1374 records seven consecutive failures that morning and only two of them classified. A spent daily allowance classifies every time; an incidental hit inside a rolling 2000-line window does not. The logic moves from a bash heredoc into scripts/classify-upstream-refusal.py because YAML can carry no regression test, and every branch, both false positives and the genuine dead-member case are now pinned by fixtures copied from the run's own output. The old reporting `grep ... | head -5` also exited 2 on SIGPIPE under `set -o pipefail`, the exact trap its neighbouring comment warned about, visible in that run as `grep: write error: Broken pipe`. Reading the file into memory removes it. report-free-pool-health.py gains the same distinction. A member answering 429 is rate limited, not retired, and the two remedies are opposites: replacing the row of a rate-limited member deletes one that is coming back. It also surfaces the provider's quota window and retry hint, pulled from the full error text rather than truncated away at 300 characters, so the next reader can tell a transient per-minute limit from a spent daily allowance without opening a log artifact. That was exactly the question this failure raised and could not answer. Both self-checks join `make test-scripts`, which is already a required check. report-free-pool-health.py has shipped a --selfcheck since issue #1064 that no workflow ever invoked, which is the same coverage-on-paper shape the Makefile comment already blames for issue #717. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WyEwUxZCArdn1ZUDkTvuQ1
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
|
Warning Review limit reachedNext included review available in 26 minutes. View limit detailsLimit details: You’ve used the included review currently available. You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository. Review configuration: ⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Pro Plus Run ID: 📒 Files selected for processing (4)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Addresses the Antigravity review on PR #1410, plus three findings from a self-review of the same diff. Discounting LiteLLM's cooldown bookkeeping for every branch traded the false positive being fixed for a false negative on the two conditions no rerun can clear. The cooldown table quotes the exception that caused each entry, so when the original refusal has scrolled out of the --tail=2000 window the table is the only surviving evidence of it. A 429 quoted there is stale by construction, because LiteLLM reprints the table on every dispatch for as long as the entry lives, and that staleness is the whole defect. A daily token budget or a 402 quoted there is still true. The spend branches now read the unfiltered lines and only the transient rate-limit branch reads the filtered ones, with self-check cases for a TPD and a 402 quoted inside a cooldown entry. Verified this does not reopen the original false positive: run 33243396287's compose-logs artifact still classifies as no refusal, because its cooldown entries quote a RateLimitError and a NotFoundError and never a daily budget or a 402. An unclassified run now says what it discounted and shows the closest matching lines it did see. Silence over a log visibly full of 429s is the failure mode that would have replaced the one being fixed, and the evidence block also covers the branches that were deliberately narrowed, so a NotFoundError in a model group other than the free pool still reaches the reader instead of being dropped without trace. The exit code of report-free-pool-health.py is computed from the payload again rather than by scanning rendered output for an `::error::` prefix. The behaviour was identical either way, but an exit code that depends on message wording breaks the next time somebody rewords a message, and this one decides whether a free-tier outage stops the run. _window_hint bounds its input before scanning. WINDOW_HINT has no nested quantifier so it cannot backtrack catastrophically, but the error string is provider-controlled text of unbounded length and the scan is linear per start position, so the bound is cheaper than trusting a remote service to keep its messages short. The cross-step contract with GITHUB_ENV now has a test. The downstream "File or update tracking issue on failure" step reads UPSTREAM_FAILURE_CLASS from there, and if that write breaks the step files an issue with no cause in it, which is the exact failure mode issue #1088 was opened about. The test pins that the value is written, is newline-terminated, stays on one line so it cannot corrupt later entries, and is not written at all when there is no refusal. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WyEwUxZCArdn1ZUDkTvuQ1
Adversarial review, two streamsStream 1: CodeRabbit CLI — RAN, no findingsNot skipped and not rate limited on this run, so this is a real clean pass rather than an absence. Stream 2: Antigravity (
|
|
Correction to my review comment above, because the distinction matters and I blurred it. There are two CodeRabbit surfaces on this PR and they did not have the same outcome:
So the honest tally for stage 6 is: CodeRabbit CLI RAN clean, CodeRabbit bot SKIPPED (rate limited), Antigravity RAN with four findings, two of which I accepted and fixed in c7719c5. Two independent streams did run, so the review requirement is met, but nobody should read the green |
|
Filed the deferred items from this investigation, so they do not live only in the PR body:
None of these is fixed here. This PR stays scoped to the classifier. |
…1417) Fixes #1409. ## What was wrong `repo-policy-lints` is a required check declared with `if: always()`, and every step in it carried `if: needs.changes.outputs.run != 'false'`. On a docs-only pull request the job therefore reported success having executed nothing. One step in that job guards a docs path. `tools/lint-no-token-in-proof-captures.mjs` resolves its scan root to `docs/proof/` and reads nothing else, and `docs/*` is on the inert-path allow-list that computes the docs-only verdict. So the only automated guard standing in front of a credential committed into a proof capture switched itself off on the exact pull-request shape that adds proof captures. PR #1398 committed seventy-one captures plus a README. Its `Repo policy lints (tenant + audit)` check reported pass in six seconds, and it was merged on that green. Running the linter by hand on merged `main` afterwards reported `ok (217 files under docs/proof/ scanned)`, so nothing had leaked. Safe by luck, not by gate. ## Answers to the three questions on the issue ### 1. Where the verdict is computed, and what it short circuits `.github/workflows/ci.yml`, the `changes` job, step `Decide whether the real suite must run`. It lists the pull request's files through the GitHub API and classifies each one. An executable extension (`*.js`, `*.mjs`, `*.cjs`, `*.jsx`, `*.ts`, `*.mts`, `*.cts`, `*.tsx`, `*.sh`, `*.py`) is denied first and forces `run=true`. Then `*.md`, `LICENSE`, `NOTICE` and anything under `docs/`, `.wolf/`, `.claude/`, `.vscode/`, `.cursor/` are inert. Anything unmatched forces `run=true`. Jobs skipped wholesale when the verdict is docs-only (`if: needs.changes.outputs.run == 'true'`): `rust-tests`, `desktop-tests`, `agent-console-unit`, `docker-build`, `markitdown-sidecar`, `live-integration`, `web-e2e`, `interaction-coverage`, `egress-firewall-e2e`. Every one of those builds or exercises code, so none of them has the "the thing I guard is a docs path" character. Jobs always scheduled but gated per step: `go-tests`, `web-unit`, `repo-policy-lints`. I checked the rest of `repo-policy-lints` for the same hole, since the concern is not this one linter. Its other guards read workflows, Dockerfiles, shell scripts, Go sources, `package.json`, or `.claude/hooks/*.js`. Every one of those inputs already forces `run=true`, either through the extension-deny arm (`.js`, `.mjs`, `.sh`, `.py`) or through the unmatched-path arm (`.github/workflows/*.yml`, `deploy/docker/Dockerfile.*`, `Makefile`). `lint:proof-tokens` is the single guard in the tree whose subject lives under an allow-listed prefix, so it is the only one with this defect today. Nothing else needs naming, and nothing else was left unfixed. ### 2. Which fix, and why Three candidates: 1. **Exempt `docs/proof/**` from the docs-only classification.** Rejected. It would run the entire heavy suite, `docker-build` and `web-e2e` included, over a pull request of screenshots, to gain one sub-second filesystem scan. It also pays that price for the image half of a capture, which the scanner cannot read at all. 2. **Run `lint:proof-tokens` unconditionally.** Chosen. 3. **Add a third `changes` output that is true when a `docs/proof/**` path changed, and gate the step on it.** Rejected. It copies the scanner's scope into the workflow, and the allow-list comment a few lines above already states the governing rule, that a second copy of a list is a second thing to forget. It is also strictly weaker than option 2: the scanner walks the committed tree rather than the diff, so an unconditional run also covers a capture that landed earlier and was never scanned, which a diff-scoped gate cannot. Cheapest measured against the failure mode, which is the instruction on the issue. The failure mode removed is a credential published to a public repository with no other automated backstop at all, since GitHub secret scanning, push protection and GitGuardian do not inspect a release asset and inspect no pixels anywhere. PR #578 is the precedent, four live invitation tokens, and that pull request was very likely docs-only itself. Every other step in the job keeps its gate, so a docs-only pull request still skips the tenancy, audit, deploy and workflow guards that genuinely have nothing to do with it. ### 3. The other direction Yes, and this pull request does not close it. The scanner reads `docs/proof/` by design and by stated rule, so a capture log committed anywhere else is unscanned no matter which checks run. This fix should not be read as more complete than it is: it fixes when the scanner runs, not where it looks. Filed separately as #1420, with the candidate designs and the same red-first acceptance bar. ## What this costs **On an ordinary code pull request, nothing.** `run=true` there, so checkout, `setup-node`, `npm ci` and every lint in the job already ran on every such pull request. The four steps simply no longer evaluate a condition that was always true. No step was added, no step made slower, and the job's work is identical. **On a docs-only pull request, nineteen seconds instead of six.** Measured, not estimated, from the green demonstration job below (`10:09:26Z` to `10:09:45Z`): checkout 8s, `setup-node` 2s, `npm ci --ignore-scripts` 3s over the two root dev dependencies (`pg` and `yaml`), and the scan of all 213 files under 1s. Every other step in the job still skips, and all of them together still resolve inside that same second. The `.wolf/buglog.jsonl`-only pull requests described in `.claude/rules/openwolf.md` are on that docs-only path and now pay the same thirteen extra seconds. That is a deliberate trade and worth stating plainly. ## Acceptance, red first The bar on the issue is not "the check runs". It is a deliberate synthetic credential in a docs-only pull request turning the check RED. That demonstration cannot be run on this pull request. The fix edits `.github/workflows/ci.yml`, so this branch can never be classified docs-only, and `on.pull_request.branches` is `[main]`, so a pull request based on this branch triggers no CI at all. The demonstration therefore ran on two throwaway branches: a base branch carrying this exact fix plus a temporary entry in that branches filter, and a head branch whose only change was one planted capture file. Neither the temporary filter entry nor the planted credential appears in this pull request. The planted value was `token=` followed by sixty-two characters of a cycled hex alphabet, the same synthetic constant the linter's own `MUST_CATCH` fixtures use. It was never a real credential, and it was redacted in the very next commit on the same branch. Demonstration pull request: #1418, base `ci/proof-lint-1409-base`, head `ci/proof-lint-1409-red`. Both branches are deleted once this pull request merges, so the run and job ids are recorded here rather than left on a branch, for the same reason visual proof moves to a release. **RED, run [33247080029](https://github.com/sakibsadmanshajib/hive/actions/runs/33247080029), job [99086306257](https://github.com/sakibsadmanshajib/hive/actions/runs/33247080029/job/99086306257)** `Detect changed paths` classified it docs-only, and every heavy job skipped accordingly: ``` verdict: SKIP — every changed path is docs-only web-e2e verdict: SKIP, no changed path can affect the booted browser flow ``` `Repo policy lints (tenant + audit)` then FAILED, on the one step that no longer skips: ``` Run npm run lint:proof-tokens lint-no-token-in-proof-captures: self-test ok (16 assertions) docs/proof/ci-gate-1409/capture-log.txt:11: looks like a real credential (query param, JWT, or API key) — step 3: navigated to http://web-console:3000/invitations/accept?token=0123456789abcdef… Process completed with exit code 1 ``` Job conclusions on that run: `Repo policy lints (tenant + audit)` failure; `Detect changed paths`, `Go tests` (all five modules) and `Web console` success; `rust-tests`, `desktop-tests`, `agent-console-unit`, `docker-build`, `markitdown-sidecar`, `live-integration`, `web-e2e`, `interaction-coverage`, `egress-firewall-e2e` all skipped. That skip list is the proof that the docs-only verdict really was in force while the linter ran. **GREEN, run [33247152597](https://github.com/sakibsadmanshajib/hive/actions/runs/33247152597), job [99086499814](https://github.com/sakibsadmanshajib/hive/actions/runs/33247152597/job/99086499814)** Same pull request, next commit, planted value replaced with `token=REDACTED_INVITE_TOKEN`. Still docs-only: ``` verdict: SKIP — every changed path is docs-only lint-no-token-in-proof-captures: self-test ok (16 assertions) lint-no-token-in-proof-captures: ok (213 files under docs/proof/ scanned) ``` `Repo policy lints (tenant + audit)`: success. So the check is falsifiable on the exact pull-request shape it previously could not fail on, which is the point of #797. Locally, before any push: clean tree `ok (212 files under docs/proof/ scanned)`, exit 0; same tree with the plant, exit 1 naming the file and line. ## Review - **CodeRabbit CLI**: ran, not rate limited. `Review complete, no findings`, one file reviewed. - **Antigravity** (`gemini-3.1-pro-high`, effort high): run with the diff pasted inline, since `agy` executes in its own scratch directory and cannot see this worktree. It quoted this diff's exact lines back, so it reviewed the right change. Two findings, both posted as inline comments and both answered there: ungate `npm ci` since the scanner imports only Node builtins (declined, three measured seconds against a silent invariant), and the unconditional scan blocking an unrelated docs pull request if the corpus already held a leak (accepted as intended behaviour, with the diff-scoped alternative recorded in #1420). ## Rebase note Rebased onto `0d089744a` (#1410, merged) before requesting merge. `.github/workflows/ci.yml` has other editors in flight: #1390 touches object storage, #1410 touched the upstream-refusal classifier. This diff touches only the `repo-policy-lints` job, lines 643 to 660 and 727 to 750, and the rebase was clean. ## Buglog entry ```json {"date":"2026-08-29","title":"lint:proof-tokens skipped itself on the only pull requests that add proof captures","error_message":"Repo policy lints (tenant + audit) reported pass in 6 seconds on PR #1398, which committed 71 proof captures. No step in the job executed.","root_cause":"The changes job in .github/workflows/ci.yml puts docs/ on the inert-path allow-list, so a pull request whose changed paths are all under docs/proof/ sets run=false. Every step in repo-policy-lints was gated on needs.changes.outputs.run != 'false', including npm run lint:proof-tokens, whose scanner reads docs/proof/ and nothing else. The gate therefore disabled itself for its own subject matter, and published a green that a reviewer trusts. Nothing else backstops the text half of a proof capture: GitHub secret scanning, push protection and GitGuardian do not read release assets and read no pixels.","fix":"Run npm run lint:proof-tokens unconditionally, along with the checkout, setup-node and npm ci steps it needs. Every other step in the job keeps its docs-only gate. Rejected carving docs/proof/** out of the classification (would run the whole heavy suite over a screenshot pull request) and a dedicated changes output (copies the scanner's scope into the workflow where it can drift, and cannot cover a capture that landed earlier). Demonstrated red before green on a throwaway docs-only pull request: run 33247080029 failed on a planted synthetic token, run 33247152597 passed once redacted.","tags":["ci","docs-only","proof-captures","credential-leak","unfalsifiable-green","issue-1409","issue-797"],"pr":1417} ``` Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Same defect family as the three markers cleared in #1383: a marker that outlived the bug it tracked and now renders as a red XPASS(strict) reading "Expect test to fail". Issue #1274 was content_block_start omitting an empty `text` field, because `json:"text,omitempty"` drops a Go string's empty value as readily for the required-but-empty case as for unset. The Anthropic SDK seeds its accumulator snapshot from that field and then does `content.text += delta.text`, so a missing key made that a `None += str` TypeError on every streamed text response. PR #1296 (commit 0b4e5ff) fixed it by making Text a pointer so the value reaches the wire, and the issue is closed. Not cleared on a single green sample. The test reported XPASS(strict) on five independent CI runs across five commits (33244935750, 33246016113, 33247050365, 33247571957, 33248088001), two of them after rebasing onto the classifier fix in #1410, with no run in which it failed legitimately.
) (#1390) Fixes #1324. Fixes #1380. Diagnoses #1370 and leaves it open with the cause recorded. ## Was restoring this coverage worth it? It caught a real wire-format bug on its first sighted run. That is the short answer to the only question that matters here, and it is evidence rather than principle. A suite that had been blind for four months was given a real object store to talk to. The first thing it found was that **every signed S3 request this product makes was going out in absolute form** (`PUT http://host/s3/bucket/key HTTP/1.1`), which RFC 9112 reserves for requests **to a proxy**. An origin server expects `PUT /s3/bucket/key`. The cause was `URL.Opaque`, set for the SigV4 signer and never cleared, which `net/url.RequestURI` returns verbatim with the scheme glued back on. It was invisible on the deployed box for one reason only: every S3 request there passes through Caddy, which normalizes the target before proxying it on, so Supabase Storage never saw the malformed form. Reached directly, Storage refuses it with a 500 that names neither the path nor the credential. Nothing in four months of green CI could have found this, because CI was pointed at a hostname that stopped resolving in August. Details, the wire capture, and the fix are in the pinned comment on this pull request. ## What was actually wrong, and it is one thing rather than three The `S3_ENDPOINT` repository secret was last written on `2026-04-21T23:15:49Z`, four months before the August self-hosted Supabase cutover, and still names the Supabase Cloud project this repository left. That project is deleted, so the hostname does not resolve. Confirmed with `gh secret list`, along with the rest of that April-vintage family: `S3_ACCESS_KEY`, `S3_SECRET_KEY`, `S3_REGION`, `S3_BUCKET_FILES`, `S3_BUCKET_IMAGES`, `SUPABASE_URL`, `SUPABASE_DB_HOST`. **#1324.** `live-integration` reads those six at job level (`ci.yml:1221`), so every `POST /v1/files` died at the socket. The Files and Batches conformance tests were carried as `it.fails` markers, which means the suite reported green for four months while covering nothing. **#1380 names the wrong job, and this is worth stating plainly rather than quietly fixing.** `ci.yml:2026` and `ci.yml:2565` are `web-e2e` and `interaction-coverage`, both Playwright. Neither runs `sdk-tests-js`; those run in `live-integration` (`ci.yml:1612`), which was reading the dead secret. So #1380's conclusion, that CI cannot fail on the storage path, is correct, but its cause is #1324's secret and not the discard port. I read the comment at `ci.yml:2016` before touching either line, as asked, and its author's three reasons all survive scrutiny, so the stub stays and the comment now names the job that does exercise storage. **#1370 is a third victim of the same secret family.** The nightly does not report `skipped` at run level: every scheduled run since 2026-08-09 reports `failure` (runs 33197858823, 33097409194, 32939092305, 32817812570, 32698750004, 32623212748, 32557183055, 32455134986, 32340347866). What is skipped is the work. Job log 98939498320 shows the cause: ``` socket.gaierror: [Errno -2] Name or service not known urllib.error.URLError: <urlopen error [Errno -2] Name or service not known> File "scripts/seed-owui-e2e-user.py", line 709, in main ``` The seeder resolves `secrets.SUPABASE_URL`. It dies at step 5, so steps 6 through 15 including the entire Playwright suite are skipped, and the only red left is step 16 failing on `browserType.launch: Executable doesn't exist`, because the step that installs browsers was one of the skipped ones. A reader who opens the run sees a Playwright error and no mention of a dead hostname. That is what made a week of failures read as flake. ## The decision, and why it is not a new secret **Can a GitHub-hosted runner reach the box's storage endpoint? No, and it must not.** `deploy/docker/Caddyfile.supabase` puts `/storage/v1` on the internal listener only. The public listener is a separate port carrying `/auth/v1` minus its admin and self-service routes, and the file states the constraint directly: > Consequence for whoever does the cutover: publish or tunnel port 8080 ONLY. Ports 80 and 443 are the in-network surface and must never be exposed. So repointing the secret is blocked by a deliberate security posture, not by an unfinished chore. Even with that posture relaxed, CI would then be reading and writing the live demo box's buckets, which is the arrangement the `web-e2e` comment already records as a defect that was fixed. **MinIO was considered and rejected, and not on the grounds you might expect.** The standing no-MinIO line lives in `docker-compose.enterprise.yml:312` and in CLAUDE.md as a statement about what a deployment persists objects into. On a literal reading it does not reach a container that lives for six minutes and holds nothing, and I do not think it does. MinIO is rejected on a different ground: the defect that actually took production down, #1282, was a `supabase-storage` server-side credential defect. MinIO has no `S3_PROTOCOL_*` concept at all, so a MinIO fixture is structurally incapable of catching that or its next relative, and picking it would rebuild a weaker version of the blindness this PR exists to remove. **What this does instead:** `live-integration` stands up the real `supabase/storage-api`, at the same immutable digest `docker-compose.enterprise.yml` pins, on the throwaway Postgres the job already boots. That database already has everything Storage needs: `deploy/supabase/init/00-extensions.sql` creates the `storage` schema, the Supabase roles, and the default privileges in `storage`, and its own comments already name this consumer. Nothing external is dialed, no secret is involved, and a Dependabot-triggered run behaves identically to every other run. Verified before writing a line of workflow YAML: the image boots against a bare `pgvector/pgvector:pg17` with only `00-extensions.sql` applied, runs its own migrations into `storage`, and serves a full signed round trip on `/s3/hive-files/<key>` with no PostgREST and no Caddy. **Deliberate gap, stated rather than buried.** Production reaches Storage through the gateway, which strips `/storage/v1`, so Storage is given `S3_PROTOCOL_PREFIX=/storage/v1` to rebuild the canonical URI. The fixture has no gateway, so it runs with an empty prefix and the endpoint ends in `/s3`. That prefix agreement is therefore not exercised live here. It is not unchecked: `scripts/test_selfhost_supabase_seam.py` asserts the prefix, the endpoint and the Caddy route all agree, and CI runs it through `make test-scripts`. Adding Caddy to the fixture would add a component and its failure modes for a seam that already has a guard. ## No secret needs creating, rotating or repointing Stated explicitly because the obvious reading of #1324 is "fix the secret". It is not. The three workflows that read `secrets.S3_*` now read none of them: - `live-integration` takes its values from the fixture. - `agent-visual-proof` takes the same secret-free stub the Playwright jobs use, because it touches no object either. Its `github.actor != 'dependabot[bot]'` guard is byte identical after the change; `git diff` shows no line touching it. - `owui-nightly` keeps its six, deliberately, and that needs the explanation below. Nothing else in this repository reads the six `S3_*` secrets after this change. ### Why `owui-nightly` keeps them, and what the lint taught me My first attempt at that file added a preflight step resolving `secrets.SUPABASE_URL`, failing with a message naming the dead project. `tools/lint-no-dead-supabase-secrets.mjs` rejected it, correctly, and it is worth restating why rather than just complying. That guard pins the exact per-secret reference counts for the file rather than exempting the file, so a freshly copy-pasted `secrets.SUPABASE_*` line cannot ride in on an exemption. My step took `SUPABASE_URL` from 3 to 4. Improving an error message is not a reason to widen an exemption documented as one that must only ever shrink. The diagnostic therefore moved into the existing `Reconstruct SUPABASE_DB_URL` step, which already reads `SUPABASE_DB_HOST`, and resolves that instead. Same deleted project, same answer, zero new references. **The `PENDING` map is unchanged and the lint passes on the audited counts exactly as they stand:** ``` pending owui-nightly.yml: 13 audited reference(s) to the retired project's secrets, tracked in #1055 ok: 16 workflow file(s) scanned, 1 pending, no unexempted reference to the retired project's secrets ``` I checked whether this PR could take that file to zero and retire the entry, which is what #1055 actually wants. It can in principle, by the route the linter itself names: `scripts/ci-throwaway-db.sh` plus `scripts/ci-supabase-stack.sh` plus this PR's new `scripts/ci-object-store.sh`, the arrangement `web-e2e` already has. I am not doing it here, for reasons I would rather state than have discovered: - It replaces every credential path in a job whose suite has not passed since 2026-08-18, so a configuration fix is unlikely to be the last thing wrong with it. That is #1370's own fix, not a storage PR's. - It is unverifiable from this branch in any way I would trust. The job runs 45 minutes and builds Open WebUI from source, and I would be asserting it works on one dispatch of a suite with unknown other breakage. This repository has a standing lesson about exactly that shape of sincere false green. - Its S3 half specifically cannot take the stub the other jobs take. `web-e2e`, `interaction-coverage` and `agent-visual-proof` can be stubbed at the discard port because they touch no object. This one boots Open WebUI, whose knowledge and document surfaces do write objects, so a stub would convert a blocked suite into a suite that runs and fails on storage. It needs the real fixture, which needs the throwaway Postgres, which is the #1370 change. The block now carries that reasoning inline, so the next person to open it does not rediscover any of it. ## Proving the signal can actually fail Every restored or new assertion was verified red before being kept. **The fixture's own round-trip check.** Booted with `S3_PROTOCOL_ACCESS_KEY_ID` and `S3_PROTOCOL_ACCESS_KEY_SECRET` removed, it exits non-zero with, verbatim, the error PR #1368 read off the production box: ``` ::error::the throwaway object store answered its health check but could not complete a signed PUT and GET. That is the shape of issue #1282: Storage reports healthy and refuses every signed request. aws: [ERROR]: An error occurred (AccessDenied) when calling the PutObject operation: Missing S3 Protocol Access Key ID or Secret Key Environment variables ``` With the credentials present, same database, same command: `round trip OK` and seven `KEY=value` lines. The ordering guard was also verified red, by running it against a Postgres with no `storage` schema. **The workflow rot guard**, `apps/web-console/tests/unit/ci-live-integration-storage.test.ts`, run against the pre-change `ci.yml`: ``` × reads no S3 repository secret × stands up a real object store fixture × keeps the S3 values out of the job-level env block Tests 3 failed | 1 passed (4) ``` and green against the post-change file. All four of its assertions matter: a returning secret means the job dials a dead host again, a missing fixture step means it dials nothing, and a job-level `S3_*` entry silently overrides what the step writes to `$GITHUB_ENV`, which is the exact trap that got `SUPABASE_*` and `HIVE_API_KEY` removed from this job rather than reassigned. **The digest drift guard**, added to `scripts/test_selfhost_supabase_seam.py`, verified red by bumping the fixture pin one patch version: ``` AssertionError: the CI object-store fixture and the deployed stack pin different Storage images. compose: supabase/storage-api:v1.11.13@sha256:1e85dad... fixture: supabase/storage-api:v1.11.12@sha256:1e85dad... ``` Without it, the claim that CI tests what the box runs is a comment rather than a property. **End to end in CI.** The `it.fails` markers came off, which is itself a two-way check: `it.fails` reports an unexpected pass as a failure, so a premature removal and a premature marker both go red. The deliberate break and its revert are recorded below once both runs land. The break was made in the code path rather than in the fixture, so what is proven is that a real storage defect turns the check red: `apps/edge-api/cmd/server/main.go` handed the files handler a bucket that does not exist, which is the shape of the family #1282 belongs to, where the request reaches a live server and is refused. **Run 33243735217, commit `1e160189b`, the deliberate break. `Live integration (SDK tests + smoke)` (job 99077797492): failure.** The fixture itself came up clean on the hosted runner, which was the one thing that could not be proven locally: ``` bucket hive-files: created bucket hive-images: created round trip OK ``` and then the restored coverage did exactly what it exists to do: ``` ❯ tests/files/files.test.ts (2 tests | 1 failed) × Files > uploads, lists, retrieves metadata, downloads content, then deletes a file → 500 Failed to store file ✓ Files > rejects an invalid purpose with a structured 4xx error ❯ tests/batches/batches.test.ts (2 tests | 1 failed) × Batches > rejects an unsupported batch endpoint value with a structured 4xx error → 500 Failed to store file ``` Both failures name the storage path. Before this PR the same break would have changed nothing, because the endpoint was already dead and both assertions were carried as expected failures. That run was later cancelled by the `ci-${{ github.ref }}` concurrency group when the revert was pushed. Its `Live integration` job had already concluded, so the evidence above is from a completed job and not from a cancelled one. It also failed `Repo policy lints (tenant + audit)` on the dead-secret linter, which is the finding described in the section above and is fixed in `4411d7af9`. **Run 33244166031, commit `b8b8ae75e`, break reverted plus the lint fix.** Result recorded here once it lands. **Run 33246016113, commit `9994dba07`.** The storage half is green: ``` == sdk-tests-js (exit 0) ✓ tests/files/files.test.ts (2 tests) ✓ tests/batches/batches.test.ts (2 tests) Test Files 22 passed (22) Tests 58 passed | 1 skipped (59) ``` | Run | Commit | Expectation | Result | | --- | --- | --- | --- | | 33243735217 | `1e160189b` | `Live integration` red, naming the storage path | failure on `500 Failed to store file` | | 33246016113 | `9994dba07` | Files and Batches green | both suites pass, 58/58 in JS | Between those two the restored coverage did its job and caught a real defect, which is why the middle runs are not a clean green: see the pinned comment on this pull request. `packages/storage` was emitting an absolute-form request target on every S3 request, which Caddy normalizes on the box and Supabase Storage refuses outright. That is fixed here, verified red then green, then end to end against a real Storage container. `Live integration` is still red overall, on a strict `xfail` for issue **#1274** in the Python suite that now passes. It is pre-existing, present in run 33243735217 before any of this branch's storage work, unrelated to storage, and deliberately left alone: clearing a strict `xfail` on one green sample is how an intermittent defect gets recorded as fixed. ## Required contexts: ran, or satisfied by a skip Filled in from the PR's own checks once they resolve, per the standing hazard that a `skipped` required check satisfies branch protection and a non-required check's absence is invisible on the page. ## Also in this PR, at the coordinator's request `.claude/skills/worktree-compose-stack.md` advertised "the current dead-Supabase-host blocker" in its frontmatter description, which is the line the skill router matches on, and named a specific Supabase Cloud hostname in `.env` as the thing stopping a local stack, citing #1254. Both halves are now false, and the way they are false is expensive: #1254 is closed, and `.env` no longer contains that hostname. Verified directly, values masked: `SUPABASE_URL`, `SUPABASE_DB_URL`, `S3_ENDPOINT` and `NEXT_PUBLIC_SUPABASE_URL` are all present with a zero-length value. An agent reading the old text greps for the hostname, finds nothing, and concludes either that the skill is wrong or that a local stack should now work. Neither is true. The blocker changed shape rather than going away, and the two shapes produce different errors: - Then: `control-plane` came up and warned `database unreachable after 6 attempt(s)`, naming no cause. - Now: `control-plane` never reaches the database check. `loadStorageConfigFromEnv` requires the S3 names to be non-empty and `log.Fatalf`s, so it exits at boot with `storage unavailable: missing S3_ENDPOINT, S3_ACCESS_KEY, S3_SECRET_KEY, S3_REGION`. The section now leads with "do not grep `.env` for a Supabase hostname to confirm this", states both shapes, and points at `scripts/ci-supabase-stack.sh` plus the new `scripts/ci-object-store.sh` as the sandbox-workable alternative. It belongs in this PR because it is the same root cause at a different site: one deleted Supabase Cloud project, still referenced by stale configuration in more than one place. The CI-side instance is the secret this PR removes. ## Not in scope, deliberately - **Fully fixing #1370.** See the section above on why, and what it would take. What this PR adds there is a resolution check inside a step that already reads the host, failing in seconds with a message naming the stale secret family and the issue, instead of spending 45 minutes producing a misleading missing-browser cascade. The issue stays open with the cause recorded. - **The accumulating `owui-e2e-shim` API keys** noted in #1370. Separate hygiene finding. - **`/v1/batches` success path.** Still not exercisable with the current provider mix, CLAUDE.md Known Issues item 4, unchanged. The submit, retrieve and cancel path is what the marker was hiding and is what comes back. ## Effect on the demo None. `deploy-demo-box.yml` is untouched, and no file this PR changes is read by it. ## Plan `ObsidianVault/hive/plan-2026-08-29-ci-object-storage-blindness.md`, with the requirements and acceptance criteria folded in. ## Issue #1274's stale marker, cleared with two-sample evidence `Live integration` was red on one thing that was not storage: a `strict=True` `xfail` in `packages/sdk-tests/python/tests/test_anthropic_messages.py` that now passes, rendering as `XPASS(strict)`. That is the same defect family as the three `it.fails` markers #1383 cleared this morning, and a marker states what is true today. **What it tracked and what fixed it,** so the closure is traceable rather than "it passes now": #1274 was `content_block_start` omitting an empty `text` field, because `json:"text,omitempty"` drops a Go string's empty value as readily for the required-but-empty case as for unset. The Anthropic SDK seeds its accumulator snapshot from that field and then does `content.text += delta.text`, so a missing key made that a `None += str` TypeError on every streamed text response. **PR #1296, commit `0b4e5ffed`,** fixed it by making `Text` a pointer so the required-but-empty value reaches the wire. The comment on `StreamContentBlock` in `apps/edge-api/internal/anthropic/types.go` records that, and issue #1274 is closed. **Not cleared on one sample.** It reported `XPASS(strict)` on five independent runs across five commits, two of them after rebasing onto #1410's classifier fix, with no run in which it failed legitimately: | Run | Commit | | --- | --- | | 33244935750 | `f8dddd47e` | | 33246016113 | `9994dba07` | | 33247050365 | `3060119a5` | | 33247571957 | `11d7e2eae` (post-rebase) | | 33248088001 | `eac8729b4` (post-rebase) | `gh run rerun --failed` refuses on this branch ("its workflow file may be broken", the known behaviour when a run's workflow file has changed), so the second sample comes from separate runs on separate commits rather than a re-run of one. That is the stronger form of the evidence, not a substitute for it. The `run-live-integration` label stays on. Removing it would make a required check skip, and a skipped required check satisfies branch protection: that would trade a loud failure for a quiet absence on the exact lane this pull request exists to restore. ## Correction posted to #1380 That issue attributes the blindness to the `http://127.0.0.1:9` stub. The conclusion is right and the cause is not: `ci.yml:2026` and `ci.yml:2565` are `web-e2e` and `interaction-coverage`, both Playwright, and neither runs `sdk-tests-js`. The SDK suites run in `live-integration`, which was reading the dead secret. Corrected on the issue itself rather than only here, so the next reader does not inherit the wrong model: #1380 (comment) ## `ci.yml` regions this branch touches Three, all narrow, listed so a rebase order against #1410 and #1417 can be worked out: 1. The `live-integration` job env block, where six `secrets.S3_*` entries are removed and replaced by a comment. 2. One new step in `live-integration` after the throwaway-database step, one new teardown step, one new failure-dump step, and one sentence added to the existing `Pre-verify secrets reached runner env` comment. 3. Comment-only additions inside the `web-e2e` and `interaction-coverage` S3 stub blocks. No value changes in either. It does not touch the `Repo policy lints` region that #1417 edits, and it does not touch the upstream-refusal classifier that #1410 changed. Rebased onto `0d089744a` cleanly with no conflicts. ## Buglog entry ```json {"id":"bug-2026-08-29-ci-object-storage-dead-secret","date":"2026-08-29","title":"CI object storage endpoint pointed at a deleted Supabase Cloud project, so the file upload path could not fail","error_message":"file upload to object storage failed bucket=hive-files error=\"Put \\\"https://<project>.supabase.co/storage/v1/s3/hive-files/...\\\": dial tcp: lookup <project>.supabase.co: no such host\"","root_cause":"The S3_ENDPOINT repository secret was last written 2026-04-21, four months before the August self-hosted Supabase cutover, and still named the deleted Supabase Cloud project. The live-integration job read it at job level, so every POST /v1/files in CI died at DNS and the Files and Batches conformance assertions were carried as it.fails markers rather than as coverage. The same stale secret family also killed the OWUI nightly at its seeding step via SUPABASE_URL (#1370), which then skipped the entire Playwright suite and left a misleading missing-browser cascade as the only visible red. Issue #1380 attributed the blindness to the http://127.0.0.1:9 stub in ci.yml, but that stub is in two Playwright jobs that run no SDK suite.","fix":"Removed every secrets.S3_* reference from live-integration and agent-visual-proof. live-integration now boots its own throwaway supabase/storage-api at the digest docker-compose.enterprise.yml pins, on the throwaway Postgres it already stands up, whose 00-extensions.sql already creates the storage schema and roles. The fixture asserts a signed PUT and GET before printing its values. The it.fails markers came off. Guards added: a unit test asserting the job keeps the fixture and no S3 secret, and a seam-test assertion that the fixture and compose pin the same image.","tags":["ci","storage","s3","supabase","dark-detector","stale-secret","sigv4"],"issues":[1324,1380,1370,1282]} ``` ```json {"id":"bug-2026-08-29-s3-absolute-form-request-target","date":"2026-08-29","title":"S3 client sent absolute-form request targets, which Supabase Storage refuses with an unnamed 500","error_message":"s3 PUT /s3/hive-files/<key> failed with status 500: <Code>InternalError</Code><Message>Internal Server Error</Message>; storage side: {\"code\":\"ERR_INVALID_URL\",\"input\":\"http://localhost:8080http://172.17.0.1:5000/s3/hive-files/<key>\"} TypeError: Invalid URL at SignatureV4.constructCanonicalRequest","root_cause":"packages/storage/signing.go SignHTTP sets req.URL.Opaque to \"//host/path\" so the AWS SigV4 signer receives an un-normalized path, and never cleared it. net/url.RequestURI returns Opaque verbatim and re-attaches the scheme when it starts with \"//\", so every signed S3 request went on the wire in absolute form (PUT http://host/s3/bucket/key HTTP/1.1). RFC 9112 reserves absolute form for a request to a proxy. The deployed box never saw it because every S3 request passes through Caddy, which normalizes the target before proxying. Reached directly, Supabase Storage builds its canonical URI as new URL(\"http://localhost:8080\" + prefix + request.url) and throws on the doubled origin.","fix":"Clear req.URL.Opaque after signer.SignHTTP returns. The signature is unaffected: the signer derives its canonical URI from Opaque while set, stripping the leading //host, and that string is byte identical to the EscapedPath the wire form falls back to once Opaque is empty. presignHTTP still keeps Opaque, because url.URL.String renders a bare-path Opaque as http:/s3/... and the //host form is what makes a presigned URL absolute. Four tests added, each verified red first.","tags":["storage","s3","sigv4","http","rfc9112","proxy-masked","dark-detector"],"issues":[1324,1380,1282]} ```
## Summary This is the batched buglog follow-up for the pull requests merged to `main` on 2026-08-29. Its diff is `.wolf/buglog.jsonl` and nothing else. Per `.claude/rules/openwolf.md`, every fixed bug, error, failed test or failed build must be logged, but the line may never be appended on a fix branch. `merge=union` in `.gitattributes` resolves concurrent appends locally and is ignored by GitHub's server side merge, so two branches that both appended land in hard conflict there. An unmergeable pull request gets no `refs/pull/N/merge`, no `pull_request` run and therefore zero checks, and the required status gate then blocks the merge for a reason the page never states (issue #873). Each fix accordingly carried its entry in its own pull request body, and this pull request copies them onto `main` in one batch, which the protocol explicitly prefers over one pull request per entry. ## Scope examined Fifty nine pull requests merged to `main` on 2026-08-29. Forty eight of them carried at least one entry, for eighty two entries in total. Thirty two of those were already on `main` and are skipped, leaving fifty appended here from thirty four pull requests. The largest block of skips comes from #1342, the equivalent batch for the 2026-08-28 merges, which merged earlier the same day and already landed thirty six entries covering #1257, #1268, #1276, #1277, #1287, #1292, #1293, #1294, #1296, #1301, #1303, #1305, #1313, #1335 and #1337. ## What landed Fifty entries appended, one JSON object per line, append only. The 232 pre-existing lines are byte identical to `origin/main` (verified by hashing the first 232 lines of the result against the base file). Every line in the resulting file parses as JSON and carries `error_message`, `root_cause`, `fix` and `tags`. | Source | Entries | |---|---| | #1083 | 2 | | #1277 | 1 | | #1278 | 1 | | #1298 | 1 | | #1334 | 1 | | #1336 | 3 | | #1343 | 1 | | #1346 | 1 | | #1351 | 1 | | #1365 | 2 | | #1368 | 1 | | #1369 | 1 | | #1371 | 3 | | #1375 | 3 | | #1376 | 1 | | #1378 | 1 | | #1379 | 2 | | #1388 | 5 | | #1389 | 3 | | #1390 | 2 | | #1393 | 1 | | #1394 | 1 | | #1410 | 1 | | #1417 | 1 | | #1421 | 1 | | #1423 | 1 | | #1424 | 1 | | #1426 | 1 | | #1429 | 1 | | #1431 | 1 | | #1433 | 1 | | #1434 | 1 | | #1436 | 1 | | #1439 | 1 | Entries are copied verbatim from their source pull request bodies. Nothing was rewritten, no field was invented, and no field was added. No JSON needed repair: all eighty two extracted entries parsed on the first attempt and all four required fields were present on every one. ## Merged pull requests that carried no entry Eleven of the fifty nine. Recorded here because the gap is itself the useful signal. | Pull request | Title | Assessment | |---|---|---| | #1013 | chore(deps): bump the go-minor-patch group across 1 directory with 4 updates | Dependabot bump, no defect fixed, no entry expected | | #1015 | chore(deps): bump the go-minor-patch group across 1 directory with 6 updates | Dependabot bump, no entry expected | | #1016 | chore(deps): bump golang from 1.26-alpine to 1.27-alpine in /deploy/docker | Dependabot bump, no entry expected | | #1218 | chore(deps): bump postcss from 8.5.19 to 8.5.26 in /apps/desktop | Dependabot bump, no entry expected | | #1219 | chore(deps): bump golang.org/x/crypto from 0.41.0 to 0.52.0 in /apps/control-plane | Dependabot bump, no entry expected | | #1342 | chore: batch buglog entries for the 2026-08-28 merges | The previous batch pull request itself, correctly carries no entry of its own | | #1364 | chore: remove four dead skills and record the patterns that cost time | Protocol gap. The body records patterns that cost time, which is the shape of a buglog entry, but none was written as one | | #1383 | test: retire stale expected-failure markers, restore the ones that are true (#1381, #1382, #1324) | Protocol gap. Stale `it.fails` markers reading as red is a real defect that was fixed here and should have carried an entry | | #1384 | docs: correct D-047, hive-auto reverted to variable pricing (D-059) | Decision ledger correction, arguably a documentation defect, no entry written | | #1387 | chore(deps): bump next from 15.5.23 to 16.3.3 in /apps/agent-console | Dependabot bump, no entry expected | | #1398 | docs: rescue the 2026-08-25 parity captures and add the 2026-08-29 QA matrix evidence | Documentation and evidence rescue, no entry written | Six of the eleven are Dependabot bumps and one is the previous batch, so the genuine protocol gaps are #1364, #1383, #1384 and #1398. Of those, #1383 is the one worth a follow-up: it fixed a real defect class (a stale expected-failure marker reads as a red "Expect test to fail" and gets dismissed as pre-existing) and left no record. ## Entries skipped as already present Thirty two. Thirty of them matched an entry already on `main` on `error_message`, `id` or `fix`. Two more from #1278 are semantic duplicates that an exact match would have missed, and were skipped after reading the landed entries they duplicate: - #1278's `streaming content_block_start omits text field` entry is covered by the consolidated `bug-2026-08-28-anthropic-sdk-wire-conformance` entry landed from #1296, whose root cause names the same `omitempty` on `StreamContentBlock.Text`. - #1278's `GET /v1/models leaked an upstream provider name` entry is covered by `BUG-1284`, landed from #1300, which names the same `public.model_aliases.summary` publication path. #1278's third entry, on `top_k` forwarding producing a 400, is not covered anywhere on `main` and is appended here. #1342 recorded #1278 as fully "merged into #1296", which was accurate for two of its three entries. ## Note on entry quality One appended entry is thin: #1277's parity re-score record carries `error_message` of `n/a` and a root cause of "console had no privacy/data-policy surface at all". It is a parity gap record rather than a defect record. It is included exactly as written rather than embellished, per the protocol's preference for the author's own words. ## Test plan - [x] Branch cut fresh from `origin/main`, diff is `.wolf/buglog.jsonl` and nothing else - [x] First 232 lines byte identical to the base file (md5 match) - [x] All 282 resulting lines parse as JSON and carry `error_message`, `root_cause`, `fix` and `tags` - [x] No `.wolf/` telemetry (`anatomy.md`, `memory.md`, `token-ledger.json`, `hooks/_session.json`, `buglog.json`) in the commit - [ ] The six required checks report green via the inert path allowlist in `.github/workflows/ci.yml` --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…(issue #1374) (#1471) Fixes #1374 ## What was actually wrong The live integration lane was red on main for most of 2026-08-29, and every report it produced named the wrong cause. Twelve runs on main were sampled between 00:59 and 20:53. The lane failed in eleven of them and passed in the last three, so the redness is **intermittent, not constant**. Five distinct causes were behind it: | Run | Live integration job | Real failure | | --- | --- | --- | | 33225137397 | 99027935246 | three stale strict XPASS markers (#1259, #1260, #1261) | | 33227399732 | 99034286025 | same three | | 33233439579 | 99050528402 | stale vitest `it.fails` (#1317) plus stale XPASS (#1274) | | 33237021259 | 99060035910 | same pair | | 33238912385 | 99065065399 | same pair (this is the run that filed #1374) | | 33242252822 | 99073896127 | `500 Failed to store file` on batches, files and tools (#1324) | | 33243396287 | 99080051438 | `500 Failed to store file` on batches and files | | 33248100043 | 99089550787 | `500 Failed to store file` on batches and files | | 33251802900 | 99099024922 | `AssertionError: expected 31 to be 5` on the `total_tokens` invariant | | 33256605935, 33261985601, 33274656635 | | green | Three of those five causes were fixed by other pull requests during the same day (#1383 retired the stale markers, #1390 gave CI a real object store, #1410 narrowed the refusal classifier). None of it was visible from the run page or the issue. The SDK suite step's only annotation was `Process completed with exit code 1`, and the "Name the real cause when the upstream refused" step wrote a 429 into the tracking issue. Issue #1374 was opened at 06:47 and accumulated eleven comments, six of which read > **Cause identified from the container logs:** the provider rate limited us (429). while the actual failures were the table above. The 429 was real and incidental every single time: one free pool member spends its Gemini free tier daily allowance most days, LiteLLM logs the refusal, the pool routes around it, and the "Probe each free pool member" step in the same job already prints > free pool members: 3 healthy, 1 unhealthy > RATE LIMITED MEMBER: openai/gemini-flash-latest > the pool routes around it meanwhile So the job already knew the 429 was benign, and then made it the headline anyway. A red required lane whose report is wrong hides every subsequent regression, which is what this issue actually is. ## The fix **`scripts/extract-sdk-failures.py`** (new). Reads the three suite logs the failing step already wrote and names the failing tests: - vitest: `FAIL` blocks, with the assertion or error message from the line under them. `Error: Expect test to fail` is unreadable without the test name in front of it, and that is the shape a stale marker takes. - pytest: `FAILED` and `ERROR` summary lines, with the reason recovered from the `FAILURES` section when pytest dropped it from the summary line (that is where the `[XPASS(strict)] issue #1274: ...` text with the issue number lives). - gradle: test failures, excluding the `> Task :test FAILED` line, which is a task and not a test. A suite that exits non-zero and names no test is reported as having died wholesale rather than failing an assertion. That is the one case where an upstream refusal really is the likely cause, and it is the case issue #1088 was written for. **Correction, from the review.** The first version of this pull request claimed that path was preserved, and it was not. The precedence rule below keyed on `SDK_FAILED_TESTS` being non-empty, and that variable is non-empty whenever any suite failed, the wholesale line included. A genuine daily-token-budget refusal was therefore demoted to a notice and filed as context, which is issue #1088's misreport restored inside the fix meant to remove it. The extractor now reports whether a real test name was parsed, as a separate `SDK_NAMED_TEST_FAILURES` flag written only in that case, and the precedence rule keys on the flag. The two states are asserted in both directions. Output goes four places: `::error::` annotations on the run page, a fenced job summary, and `SDK_FAILED_TESTS` plus `SDK_NAMED_TEST_FAILURES` in `$GITHUB_ENV` for the tracking issue step and the classifier. **`scripts/classify-upstream-refusal.py`** gains one precedence rule, keyed on `SDK_NAMED_TEST_FAILURES`. When a real test name was parsed out of a suite log, a refusal is written as `UPSTREAM_FAILURE_CONTEXT` and annotated as a notice, not as `UPSTREAM_FAILURE_CLASS` and an error. When no test was named, including when every suite died wholesale, the refusal is still the cause and is still reported as one. It still prints the refusal and its evidence lines in full, because occasionally the refusal is the cause; it just stops being the headline when something more specific is available. **The tracking issue step** leads with the failing tests, then the cause, then the refusal as context. Both scripts carry self-checks, and `scripts/extract-sdk-failures.py --selfcheck` joins `make test-scripts`, which is already a required check via repo-policy-lints. `scripts/fixtures/sdk-failures/run-33251802900-sdk-tests-js.log` is a committed capture of that run's real suite output, ANSI escapes and all, rather than an excerpt retyped into the script; the remaining fixtures are inline excerpts from the runs in the table, and the three that are synthetic say so in a comment and say why. ## Why this is loud, not merely green Nothing here changes what passes or fails. No assertion was weakened, no timeout raised, no retry added, no pass condition widened. The lane fails on exactly the same inputs it failed on before. What changed is that the failure now names itself. The wholesale case is deliberately kept loud too: a suite that dies without naming a test says so explicitly and points at the compose logs artifact, instead of being silently absent from the report. ## Verification Replayed both scripts against the real logs of run 33251802900, the run whose issue comment at 12:19 said "the provider rate limited us", using its own `compose-logs` artifact as the classifier input: ``` ::error::sdk-tests-js: tests/usage/usage-accounting.test.ts > Usage accounting > prompt_tokens_details.cached_tokens is present with a numeric value -- AssertionError: expected 31 to be 5 // Object.is equality ::error::sdk-tests-py: tests/test_anthropic_messages.py::test_streaming_event_sequence_integrity -- [XPASS(strict)] issue #1274: content_block_start omits an empty text field ... ::notice::an upstream refusal signature IS in the logs, but named SDK test failures exist, so this is context and not the cause. ``` `$GITHUB_ENV` after the two steps carries `SDK_FAILED_TESTS` (heredoc form, since the value is multi line) and `UPSTREAM_FAILURE_CONTEXT`, and no `UPSTREAM_FAILURE_CLASS`. That is the exact inverse of what was posted to the issue at 12:19. Also replayed the parsers against the raw job logs of runs 33238912385, 33243396287, 33248100043 and 33237021259. Each named the failures the table above lists, and named nothing on the passing suites. Self-checks: ``` python3 scripts/classify-upstream-refusal.py --selfcheck # ok python3 scripts/extract-sdk-failures.py --selfcheck # ok ``` ## Review round Ten inline findings, all confirmed against the head commit at the time. Every one is addressed here. | Severity | Finding | Disposition | | --- | --- | --- | | CRITICAL | The precedence rule fired on the wholesale-death sentinel, reversing issue #1088 in its own scenario | Extractor writes `SDK_NAMED_TEST_FAILURES` only when a real test name was parsed; the classifier keys on that. Both directions asserted, both observed red before the fix | | CRITICAL | Same defect at the extractor end: the two files disagreed on what "named" meant | Same fix, one contract, asserted from both sides | | HIGH | A pytest node id containing a space matched nothing, so a parameterized assertion failure read as a suite that died wholesale | Id taken lazily up to the reason separator; a spaced parameter id is in the self-check, with and without a reason | | HIGH | `_selfcheck` never called `report()`, so the `::error::` prefix, the `MAX_ANNOTATIONS` cap, `_write_summary` and the all-green early return could not go red | The self-check captures `report()`'s stdout, `$GITHUB_ENV` and `$GITHUB_STEP_SUMMARY`. Eleven mutations of this diff were run and every one turns a self-check red | | MEDIUM | The gradle fixture was fabricated and named a class that does not exist | Real package and class names, and a comment stating it is synthetic and why: no red `sdk-tests-java` run exists to capture. To be replaced by a capture the first time that suite goes red | | MEDIUM | The classify self-check missed the notice-versus-error branch and the sentinel case | Both covered, with the annotation level asserted on captured stdout, not only the env key | | MEDIUM | `\|\| true` on the extractor step silently reverted the lane | Replaced with a `::warning::` that says the cause below is a guess and not a named test | | LOW | The tracking issue body took the uncapped list | Capped at the same 20 as the annotations, with "and N more, see the run page" | | LOW | The pytest reason recovery picked up source lines and missed class-based tests | Prefers the `E ` line, skips source, and matches the rule header's bare function name | | LOW | The job summary rendered log-derived text as markdown | Written inside a four-backtick fence | Two review findings were confirmations rather than defects and are unchanged: `$GITHUB_ENV` injection is not possible through the random heredoc delimiter with its delimiter-line filter and single-line values, and nothing in this diff weakens an assertion, raises a timeout, adds a retry or widens a pass condition. ### Both directions of the CRITICAL, before and after Against the head commit under review, with the same real 429 in both cases: ``` DIRECTION 1, a real test was named: ::error::UPSTREAM REFUSAL ... UPSTREAM_FAILURE_CLASS DIRECTION 2, only the wholesale sentinel: ::notice::an upstream refusal ... UPSTREAM_FAILURE_CONTEXT ``` Exactly inverted. After the fix: ``` DIRECTION 1, a real test was named: ::notice::an upstream refusal ... UPSTREAM_FAILURE_CONTEXT DIRECTION 2, only the wholesale sentinel: ::error::UPSTREAM REFUSAL ... UPSTREAM_FAILURE_CLASS ``` `make test-scripts` passes, exit 0. ## Not fixed here, filed separately Run 33251802900's failure was a genuine product defect and is not a reporting problem: `total_tokens` came back as 31 while `prompt_tokens + completion_tokens` was 5, which breaks the OpenAI wire contract that the two are equal. The shape matches a thinking model whose reasoning tokens are counted in the total but not in `completion_tokens`, and it reaches billing, since `usage_clamp.go` only recomputes the total when `completion_tokens` is zero. It appeared in one of twelve sampled runs, which is one sample of an intermittent defect on a load balanced pool, so it needs its own investigation rather than being bundled into a CI reporting change. Filed as #1472. ## Buglog entry ```json {"id":"bug-2026-08-29-live-integration-misreports-cause","date":"2026-08-29","title":"Live integration reported an incidental provider 429 as the cause of every failure","error_message":"Cause identified from the container logs: the provider rate limited us (429)","root_cause":"The SDK suite step's only annotation was 'Process completed with exit code 1', so the tracking issue was composed entirely from the refusal classifier's verdict. One free pool member spends its Gemini free tier daily allowance most days, so a real 429 is in the container logs of nearly every run, and the classifier reported it as the cause whatever had actually failed. Issue #1374 collected eleven comments over five unrelated defects (stale strict XPASS markers #1259 #1260 #1261 #1274, a stale vitest it.fails #1317, the dead S3 endpoint #1324, and a broken total_tokens invariant), six of them headlining the 429.","fix":"Added scripts/extract-sdk-failures.py to name the failing tests from the three suite logs (vitest, pytest, gradle) as ::error:: annotations, a job summary and SDK_FAILED_TESTS in $GITHUB_ENV. Gave classify-upstream-refusal.py a precedence rule so a refusal alongside named test failures is written as UPSTREAM_FAILURE_CONTEXT and annotated as a notice, not as the cause. The tracking issue step leads with the failing tests. The precedence rule keys on a separate SDK_NAMED_TEST_FAILURES flag that the extractor writes only when a real test name was parsed, because the failure list is non-empty for a wholesale death too and keying on it demoted a genuine daily budget refusal, reversing issue #1088 inside the fix. Both scripts self-check under make test-scripts.","tags":["ci","live-integration","observability","free-pool","misreported-cause"],"files":[".github/workflows/ci.yml","scripts/extract-sdk-failures.py","scripts/classify-upstream-refusal.py","Makefile"],"related_issues":[1374,1088,1064,1324,1383,1390,1410]} ```
The finding is the opposite of the report
I was dispatched to fix a spent daily allowance on the
hive-freepool. There was no spent allowance. The pool was healthy the entire time, and the thing that said otherwise was our own diagnostic.Three
Live integration (SDK tests + smoke)runs failed this morning carrying the banner:That banner is wrong, and it is wrong in a way that cost a full investigation cycle. It is the same defect the step was built to prevent (issue #1088, a red check that misreports its own cause), wearing the opposite sign.
Evidence that the pool was fine
The job's own
Probe each free pool member and name any that is gonestep, in both failing runs (33243396287 and 33244166031):One member is rate limited, the Google one. Direct probes of the other two providers, one request each:
dots-studio/dots-3-note-preview:freecost: 0openai/gpt-oss-20bx-ratelimit-remaining-requests: 955of1000And the deployed box, one turn on
hive-free, run byqamatrixon a key created for the probe and revoked after:The live free tier serves. Forty-five of a thousand daily Groq requests spent is not an exhausted allowance.
What the job actually failed on
sdk-tests-jsexit 1 onError: 500 Failed to store fileintests/files/files.test.tsandtests/batches/batches.test.ts.sdk-tests-pyexit 1 on[XPASS(strict)] issue #1274.sdk-tests-javaexit 0. The PR branch under test in the second run is namedfix/ci-object-storage-blindness-1324. The failure is object storage, which is #1324 and #1390, not a provider.The three false positives, all from that run's own artifact
The classifier prints its own matching lines, which is how they were found.
1. LiteLLM's router cooldown. The line that set the class:
Wrong evidence twice over. The alias is
deepseek-v4-flash, which isHIVE_TOOLS_MODELand not the free pool, so a paid alias's problem was reported as the free tier's. AndNo deployments available ... cooldown_list=[...]is the gateway declining to dispatch to deployments it has already benched. No upstream call was made; nobody refused anything. It matched only because the regex carried the bare tokenstatus=429.2. LiteLLM's cooldown table. Found only by running the fixed script against the real
compose-logsartifact rather than against my own fixtures:It quotes the exception that caused the cooldown, verbatim, and LiteLLM reprints the whole table on every subsequent dispatch for as long as the entry lives. One past event re-reports itself indefinitely.
3.
No fallback model group found. LiteLLM emits this on any failed dispatch to a group with no configured fallback, whatever the underlying error. In the same artifact it appears on a deliberate over-length prompt (400), on deepseek not supporting image input (404), and on a Groq TTS suite asserting an invalidresponse_formatis refused (400). All three are expected behaviour under test. Had the 429 branch not fired first, the dead-member branch would have.The dead-member branch now requires the
NotFoundErrorshape andmodel_group=route-free-pool, which is the case its message describes and the exact shape already recorded inapps/edge-api/internal/inference/retry_test.gofrom run 32830060362.Corroboration that this was noise, not an allowance
Issue #1374 records seven consecutive failures on the morning of 2026-08-29 (06:47, 07:24, 07:40, 07:51, 08:13, 08:42, 09:11Z). Only two carry the 429 classification; four carry no cause at all. A spent daily allowance classifies every time. An incidental grep hit inside a rolling
--tail=2000window does not.What changed
scripts/classify-upstream-refusal.py(new). The classification moves out of the YAML heredoc because YAML can carry no regression test, and this logic had shipped false positives nothing could have caught. Follows the existingscripts/report-free-pool-health.pyshape: pure stdlib,--selfcheck, same credential redactor. Branch order is unchanged. Fixtures are copied from run 33243396287's own output, including both false positives and the genuine dead-member case, so the branch cannot be tightened into uselessness without a test going red..github/workflows/ci.yml. The step body shrinks to the redaction pipeline plus one script call. This also removes a second defect: the old reportinggrep ... | head -5exited 2 on SIGPIPE underset -o pipefail, the exact trap the neighbouring comment warns about for the classification greps, visible in that run asgrep: write error: Broken pipefollowed by##[error]Process completed with exit code 2. Reading the file into memory removes it by construction. I did not reproduce the SIGPIPE locally; a five-line fixture is too small to makegrepwrite afterheadexits. The evidence is the run's output.scripts/report-free-pool-health.py. A member answering 429 is rate limited, not retired, and the two remedies are opposites: the shipped advice was "replace the member row insupabase/migrations/", which for the Gemini member would have deleted one that is coming back. It now also surfaces the quota window and retry hint, extracted from the full error text rather than truncated away at 300 characters:That line is the one fact this whole investigation needed and could not get. The real log truncated before Google's
quotaId, so I still do not know whether the Gemini window is per-day or per-minute. Closing that gap is the point of the change, not a claim to have closed the question.Makefile. Both self-checks jointest-scripts, already a required check.report-free-pool-health.pyhas shipped a--selfchecksince issue #1064 that no workflow ever invoked, which is precisely the coverage-on-paper shape the Makefile comment already blames for issue #717.What this does not do
HIVE_TOOLS_MODELresolving todeepseek-v4-proburned 150 calls and 30.2M credits in one day on 2026-08-25. No model knob is touched..wolf/decisions.mdD-048, and that is deliberate. D-048 records the pool as four Hive-owned keys behind load-balanced failover: two free Groq text keys, an OpenRouter free key, a Gemini key. Migration20260824_02_free_pool_router.sqlships exactly that, member for member. The suspicion that20260823_21_groq_text_routes_to_openrouter_free.sqlhad collapsed the pool's diversity does not hold: that migration is dated the day before the pool existed and moved the legacy tier aliases, not the pool. D-048 matches shipped reality, and a correction that corrects nothing is worse than silence.The real redundancy behind
hive-free, for the recordThree distinct providers, four keys, and between three and four distinct allowances.
route-free-pool-groqandroute-free-pool-groq-2are the same model on the same provider differing only by key, and Groq enforces rate limits per organization, not per key. If both keys live in one Groq org they are one allowance wearing two hats. Nothing in the repo settles it. Filed below.One genuinely stale artefact found on the way:
deploy/litellm/config.yaml's "CONCENTRATION RISK" comment (lines 33 to 45) describes the pre-pool topology from 2026-08-23 and says five aliases resolve to one model at one provider.20260824_02superseded that the next day, and the checked-inconfig.yamlcontains noroute-free-poolentry at all: the pool reaches LiteLLM only through the job's "Sync LiteLLM model list from the DB catalog" step. That comment is what generated this investigation's false premise. Not touched here, because that file is hot and shared; filed instead.Filed, not built
deploy/litellm/config.yaml's concentration-risk comment is a day out of date and actively misleading.GROQ_API_KEYandGROQ_API_KEY_2are two Groq organizations or two keys on one. Settled by one probe from whoever holds both: comparex-ratelimit-remaining-requestsand see whether spending on one moves the other..envare only the first obstacle; the box's public listener deliberately refuses the admin API, sogenerate_link404s and minting must run against the internalcaddy-supabaselistener through an SSH tunnel with aHostheader rewrite. Coordinate with ci: give CI a real object store instead of a dead endpoint (#1324, #1380) #1390, which is already correcting.claude/skills/worktree-compose-stack.md.Testing
Verified RED before GREEN. The first run of the new self-check failed with the production string verbatim:
And against the real artifacts from the failing run, before and after:
Shared-file note
This touches
.github/workflows/ci.yml, which #1390 also edits. Confirmed with the coordinator that #1390 does not touch the classifier: its ci.yml diff has zero matches forupstream refus,UPSTREAM_FAILURE_CLASSandrate.?limit. Its hunks at@@ -1820,12 @@and@@ -1838,6 @@are adjacent to this region but do not overlap. Branched from currentorigin/main. If #1390 merges first I will rebase and re-read those hunks rather than trust a clean auto-merge.Buglog entry
{"id":"bug-2026-08-29-classifier-misattributes-gateway-cooldown","date":"2026-08-29","title":"CI upstream-refusal classifier reported a LiteLLM router cooldown on a paid alias as a whole-pool provider rate limit","error_message":"UPSTREAM REFUSAL, not a performance regression: the provider rate limited us (429). Check whether the window is per minute (transient, so a rerun may pass) or per day (an allowance that is spent).","root_cause":"The 'Name the real cause when the upstream refused' step in .github/workflows/ci.yml grepped a 2000-line window of the litellm and edge-api container logs for the bare token 'status=429' among others. Three log shapes matched that are not upstream refusals: LiteLLM's own 'No deployments available for selected model ... cooldown_list=[...]' router cooldown (no upstream call is made), its 'Cooldown Deployments=[...]' table (which quotes the causing exception verbatim and reprints it on every dispatch while the entry lives), and a bare 'No fallback model group found' (which LiteLLM emits on any failed dispatch to a group with no configured fallback, including three refusals the job deliberately tests for). The matched line also carried alias deepseek-v4-flash, so a paid alias's gateway cooldown was reported as a refusal of the hive-free pool. The free pool was healthy throughout: the job's own probe step reported 3 healthy 1 unhealthy in both runs, only the Gemini member was rate limited, and the actual job failures were '500 Failed to store file' and a strict XPASS on issue #1274. Only two of the seven consecutive failures on issue #1374 that morning classified at all, which is the signature of an incidental grep hit rather than a spent allowance.","fix":"Moved the classification from a bash heredoc into scripts/classify-upstream-refusal.py, which make test-scripts runs as a required check. Gateway cooldown lines in both shapes are dropped before any branch reads them. The dead-member branch now requires the NotFoundError shape together with model_group=route-free-pool. A classified refusal names the alias it saw. Fixtures are copied from run 33243396287's own output. This also removes a SIGPIPE exit 2 from the old 'grep ... | head -5' reporting pipeline under set -o pipefail. Separately, report-free-pool-health.py now distinguishes a rate-limited member from a retired one, whose remedies are opposites, and surfaces the provider's quota window and retry hint from the full error text instead of truncating them away at 300 characters.","tags":["ci","litellm","free-pool","signal-quality","false-positive","issue-1088","issue-1064","issue-1374"]}🤖 Generated with Claude Code
https://claude.ai/code/session_01WyEwUxZCArdn1ZUDkTvuQ1