ci: capture visual proof of the signed-in chat surface before merge (issue #1738) - #1739
Conversation
…issue #1738) Orchestrator rule 8 requires a screenshot taken against a running stack before any UI-touching change merges. For the chat surface there was no way to satisfy that before merge, and the reason was circular. A signed-in session is not decoration on this surface. hive_jwt_forward and hive_upstream_auth attach the signed-in user's OAuth access token to every outgoing completion, and edge-api's OWUIUnwrap middleware treats /v1/chat/completions as unconditionally requiring a per-user token, so with no session the shim key stands alone and every model answers 401. Signing in needs GoTrue's OAuth 2.1 authorization server, a registered client for Open WebUI and the console's own consent screen. ci-supabase-stack.sh pinned a GoTrue with no OAuth server at all, and owui-nightly.yml, the one job that boots Open WebUI, still points at the hosted Supabase project that was deleted during the move to the self-hosted data plane, together with the OAuth client that lived in its dashboard. The only stack that could render a signed-in chat was the demo box, which only runs already-merged code. This is the chat-surface sibling of agent-visual-proof.yml. It stands the whole chain up on a disposable hosted runner from the pull request's own merge ref, signs in through the real Continue with Hive journey, sends a real completion, and posts the screenshots to the permanent visual-proof-assets release. ci-supabase-stack.sh gains two opt-in flags. --oauth-server selects the same GoTrue pin the enterprise overlay already runs and enables the authorization server; --external-url pins the discovery document and the iss claim to one origin, which is what lets the browser, Open WebUI's container and the console agree on a single Supabase address. Both default to the previous behaviour, so existing callers are unaffected. Everything else is reused rather than rebuilt: register-owui-oauth-client.py for the client, seed-owui-e2e-user.py for the run-scoped tenant, user and shim key, docker-compose.agent-proof.yml for the JWKS front, the owui-setup Playwright project for the one implementation of the sign-in journey, and stamp.mjs for the capture footer. No shared account's password is set, reset or rotated anywhere in the job.
|
Warning Review limit reachedNext included review available in 27 minutes. View limit detailsLimit details: You’ve used the included review currently available. You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository. Review configuration: ⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Team Run ID: 📒 Files selected for processing (9)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Cowork visual proof, captured in CICaptured against a stack booted from launch-liveness-01-empty-consolelaunch-liveness-02-sandbox-launchedRun log (screenshot stamps carry no URL, and the log is redacted and linted by `lint:proof-tokens`) |
Run 33671163466 died fourteen seconds into the stack boot with compose reporting only "dependency failed to start: container hive-open-webui-1 is unhealthy", and the failure log dump that should have explained it returned nothing for that service. Two causes, both reproduced locally against the shipped image. The crash is the DSN scheme. compose hands SUPABASE_DB_URL to Open WebUI as PGVECTOR_DB_URL, SQLAlchemy resolves a scheme to a dialect plugin by name, and no dialect is registered under "postgres": the container exits with NoSuchModuleError before uvicorn binds. Both Go services accept either form because pgx takes the alias, so this is the one place three drivers share one string and the only correct spelling is postgresql://. The silence is the compose profile. Four services in docker-compose.yml carry no profiles key, so a compose call without one still resolves a valid project made of exactly those four and reports on it, which is why a dump that named open-webui explicitly returned edge-api, control-plane, litellm and markitdown. COMPOSE_PROFILES is now set at job level alongside COMPOSE_FILE. With the crash fixed, Open WebUI still takes 112 seconds to pass its first health check on a fresh database, mostly its own alembic chain, against a healthcheck allowing 30 seconds and five retries. Docker recovers on its own, but caddy-owui's depends_on does not wait that long, so the boot is now two waves with this job's own deadline and its own error message. The healthcheck is left alone deliberately: widening it changes how every deployment detects a dead chat container, which is a product decision and not this job's to make.
PR #1734 landed between this branch's first two runs and added assert_account_scope, which refuses a custom --account-slug from a run that configures no consumer: a custom slug means a real deployment, and revoking its keys while syncing nothing is the issue #560 outage. This job is neither case. Its account is a row in a database created minutes earlier by the job itself and destroyed with the runner, so the slug names nothing shared and there is no deployment anywhere holding a key a revocation could strand. Dropping both slug flags puts it on the reserved CI account, which is the arm the guard allows for a consumer-less run.
Cowork visual proof, captured in CICaptured against a stack booted from launch-liveness-01-empty-consolelaunch-liveness-02-sandbox-launchedRun log (screenshot stamps carry no URL, and the log is redacted and linted by `lint:proof-tokens`) |
Cowork visual proof, captured in CICaptured against a stack booted from launch-liveness-01-empty-consolelaunch-liveness-02-sandbox-launchedRun log (screenshot stamps carry no URL, and the log is redacted and linted by `lint:proof-tokens`) |
The bootstrap step of owui.setup.ts demanded that its throwaway account reach chat, and on a freshly created container that account cannot. utils/oauth.py's get_user_role returns DEFAULT_USER_ROLE early when user_count is zero, above the point where deploy/docker/owui-patches/tenant_role_from_db.py splices Hive's own membership lookup in, and issue #748 deleted the post-insert promotion upstream used to repair the first account with. DEFAULT_USER_ROLE on this deployment is pending, so the very first account a container ever sees gets the activation screen no matter what tenant membership it holds. That is the intended posture for an instance shared by every tenant: nothing self-promotes. Which outcome appears depends only on whether the container's volume is fresh, which this fixture neither controls nor tests. What it does test is unchanged and is asserted by both arms: a full OAuth round trip completed, and the instance now holds an account that is not the fixture user. The step now accepts either, so it stops failing on every freshly created container. Measured on run 33673391630, where the bootstrap login completed the whole redirect chain and then sat on Account Activation Pending for sixty seconds. This has not been caught before because owui-nightly.yml, the only other job that boots Open WebUI, has been pointed at a Supabase project that was deleted in August and has not reached this step since. Also excludes traces, videos and the HTML reporter's index.html from this workflow's failure artifact. A trace of this journey holds session cookies and the OAuth callback's code, no text linter can inspect either, and the repository is public. Same exclusion list, and the same reasoning, as owui-nightly.yml.
Cowork visual proof, captured in CICaptured against a stack booted from launch-liveness-01-empty-consolelaunch-liveness-02-sandbox-launchedRun log (screenshot stamps carry no URL, and the log is redacted and linted by `lint:proof-tokens`) |
…tener The check that the JWKS endpoint serves a usable key fired once, immediately after root.crt appeared on disk. That file existing means Caddy's local authority is written, which happens a beat before the listener will complete a handshake with the certificate it has just issued. Measured on run 33675164605: the curl ran 0.2 seconds after Caddy logged "certificate obtained successfully", came back empty, and the job failed claiming GoTrue had published no key at all. The two are not the same thing, and this assertion exists precisely to tell an empty key set from a populated one, so it must not also be able to fail on a listener that is a fraction of a second from ready. It now polls for up to a minute and reports the same failure with the same message when the key set really is empty. Affects agent-visual-proof.yml too, which runs the same probe.
Cowork visual proof, captured in CICaptured against a stack booted from launch-liveness-01-empty-consolelaunch-liveness-02-sandbox-launchedRun log (screenshot stamps carry no URL, and the log is redacted and linted by `lint:proof-tokens`) |
signInWithHive required Open WebUI's Continue with Hive button to be visible before it would proceed. This deployment sets OAUTH_AUTO_REDIRECT, so the landing page starts the authorize chain by itself, and whether a given load paints the button first or bounces before it can be clicked is a race. Losing it left the browser sitting on the console's own sign-in form while a thirty second assertion waited for a button on an origin the page had already left, which reads as a missing control rather than as a redirect that already happened. Measured on run 33676212820, where one run won that race on the bootstrap login and lost it on the fixture login thirty seconds later. The helper now clicks the button when it is there and accepts having been redirected when it is not. Nothing is weakened: the assertion that actually matters is the next one, which requires the browser to have reached the consent origin, and it is unchanged, so a page that neither offers the button nor redirects still fails there by name.
Cowork visual proof, captured in CICaptured against a stack booted from launch-liveness-01-empty-consolelaunch-liveness-02-sandbox-launchedRun log (screenshot stamps carry no URL, and the log is redacted and linted by `lint:proof-tokens`) |
Cowork visual proof, captured in CICaptured against a stack booted from launch-liveness-01-empty-consolelaunch-liveness-02-sandbox-launchedRun log (screenshot stamps carry no URL, and the log is redacted and linted by `lint:proof-tokens`) |
Capture log (screenshot stamps carry no URL, and the log is query-string stripped and linted by `lint:proof-tokens`) |
lint-no-sync-child-process-in-tests flagged the new proof spec, correctly. A synchronous child blocks the Playwright worker's event loop, so the 360 second timeout the spec sets would never have fired while the capture ran: the exact shape PR #838 found recorded as passed at 30194 ms under a 30000 ms timeout. Awaited spawn instead, with stdio inherited so the capture log still reaches the run log live, and the child's exit code surfaced as the test failure. The per-child timeout is gone with the sync call, which is the point: Playwright's own timeout is now the one that actually applies.
Cowork visual proof, captured in CICaptured against a stack booted from launch-liveness-01-empty-consolelaunch-liveness-02-sandbox-launchedRun log (screenshot stamps carry no URL, and the log is redacted and linted by `lint:proof-tokens`) |
Capture log (screenshot stamps carry no URL, and the log is query-string stripped and linted by `lint:proof-tokens`) |
…sue that owns it
The comment explaining the staged boot said the healthcheck allows a 30
second start period and five retries. Neither number is in effect. Nothing
in this repository defines a healthcheck for open-webui: docker-compose.yml
sets none, Dockerfile.open-webui sets none, docker-compose.enterprise.yml
sets none, so it is inherited whole from the upstream base image, which
carries a test and no other field.
$ docker image inspect hive-open-webui:v0.10.2-branded \
--format '{{json .Config.Healthcheck}}'
{"Test":["CMD-SHELL","curl --silent --fail http://localhost:${PORT:-8080}/health | jq -ne 'input.status == true' || exit 1"]}
With no Interval, Timeout, StartPeriod or Retries, Docker's defaults apply:
start period 0s, interval 30s, three retries. The first check therefore
fires 30 seconds into a container that is still migrating, and three
consecutive failures declare it unhealthy at about 90 seconds, roughly 22
seconds before it finishes booting. The wrong numbers described the same
defect but understated how little headroom there is.
The defect itself is now filed as issue #1757 rather than living only in
this comment and a pull request body that vanishes on merge. It affects
every cold boot against an unmigrated database, which is a fresh enterprise
install and a fresh local dev stack, not only this job.
Comment only. No step, condition or command changes.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WbVmp2Uh7FCgnqKB2TuBb5
The one red check was the known flaky agent-engine test, not this branch
That is issue #1427, open,
Same test, same line, same shape of error, on a checkout with nothing changed. The local failure names the socket outright, which is #1427's diagnosis in one line: the reap wins the race, tears the sandbox down, and the in-flight call finds no socket. CI's Not retried silently. The re-run is the push above, which is a comment-only change to the workflow correcting the healthcheck numbers, so every check re-ran on real content rather than on a bare Two loose ends from the earlier session, now closed
|
Cowork visual proof, captured in CICaptured against a stack booted from launch-liveness-01-empty-consolelaunch-liveness-02-sandbox-launchedRun log (screenshot stamps carry no URL, and the log is redacted and linted by `lint:proof-tokens`) |
Capture log (screenshot stamps carry no URL, and the log is query-string stripped and linted by `lint:proof-tokens`) |
Second bug-log reconciliation batch of the day. The first batch (#1743, merged 2026-09-02T19:44:04Z) appended 197 entries from 167 PRs, taking `.wolf/buglog.jsonl` on `main` to 511 lines. This batch sweeps every PR merged after that point which carried a `## Buglog entry` heading in its body, appending 14 entries from 7 PRs: - #1733 (1 entry) - #1735 (1 entry) - #1739 (6 entries) - #1740 (1 entry) - #1748 (1 entry) - #1749 (2 entries) - #1756 (2 entries) Checked and excluded: - #1727 carries no buglog entry. It is a docs/process PR (tracking-discipline rule), not a bug fix, and its body mentions `.wolf/buglog.jsonl` only in passing prose. - #1715, #1729, #1731 and #1734 merged before this batch's window and are already present in the first batch (#1743). Verified by id/error_message lookup against the 511 lines already on `main`. Every entry was extracted from its source PR body, parsed as JSON to confirm it is well-formed, and checked for the required `error_message`, `root_cause`, `fix` and `tags` fields (all present, none reconstructed). No duplicates were found against the existing 511 lines or within this batch, checked by both `id` and exact `error_message` match. Diff is exactly one file, 14 insertions, 0 deletions. The first 511 lines byte-match `main`'s current copy (verified with `diff` against `git show origin/main:.wolf/buglog.jsonl`). This PR was not opened on a fix or feature branch, per `.claude/rules/openwolf.md`: it is the dedicated buglog-only PR, branched directly from `main`, diffing only `.wolf/buglog.jsonl`. Refs #873 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…1758) (#1764) Closes #1758. ## The defect Both visual-proof workflows read their YAML from one ref and then check out a different tree to run it against, the target pull request's. The scripts the steps invoke by name carry command lines that live in the YAML, so taking those scripts from the target coupled every dispatch to whatever harness the target happened to have branched with. Run [33690589909](https://github.com/sakibsadmanshajib/hive/actions/runs/33690589909) exited 2 nine seconds in on `unknown argument: --oauth-server`, an argument main's YAML passed to a script only main carried, and PR #1730 had to merge main twice before it could be photographed. ## The fix One step per workflow, immediately after the target checkout, re-takes a named list of scripts from `github.sha`. That is the default branch's head on a `workflow_dispatch` and the pull request's own merge commit on a `pull_request` event, which satisfies both halves of the acceptance criteria with one expression: a dispatch at a stale pull request runs current harness, and a pull request that deliberately edits a harness script is still proven against its own version of it. ## Checkout shape, and why The issue offered two shapes: two checkouts into two paths, or a partial checkout of harness files over the target tree. This takes the second, because the three chain scripts resolve their data files from their own location: ``` scripts/ci-supabase-stack.sh repo_root -> deploy/supabase/init/00-extensions.sql scripts/ci-throwaway-db.sh repo_root -> supabase/migrations, .github/ci/test-db-bootstrap.sql scripts/apply-migrations.sh repo_root -> supabase/migrations, scripts/migration-baseline.conf ``` Run from a second directory, `repo_root` becomes that directory, and the migrations applied to the throwaway database would be main's rather than the target's. A pull request that adds a migration would then be proven against a schema that does not include it. Making the split work in a second directory needs a repo root override threaded through three scripts. Landing the harness at its normal paths instead needs nothing: every call site is unchanged, `working-directory: deploy/docker` steps keep their `../../scripts/...` relative paths, and the data files stay the target's because the workspace is still the target's tree. The overlay is not silent. The step prints `git diff --cached --stat` of exactly which harness files it replaced, and `git checkout` fails loudly if a listed path has been renamed on the harness ref. ## The file split, written down Stated in a comment above the step in both workflows, and enforced by the new lint. **Harness, taken from the workflow's own ref.** Anything the YAML invokes by name, plus anything reached from those with an argument interface. | chat-visual-proof.yml | agent-visual-proof.yml | | --- | --- | | `scripts/ci-supabase-stack.sh` | `scripts/ci-supabase-stack.sh` | | `scripts/ci-throwaway-db.sh` | `scripts/ci-throwaway-db.sh` | | `scripts/generate-enterprise-jwt-keys.py` | `scripts/generate-enterprise-jwt-keys.py` | | `scripts/register-owui-oauth-client.py` | `scripts/install-agent-engine-host.sh` | | `scripts/seed-owui-e2e-user.py` | `scripts/agent-engine-health-probe.sh` | | `scripts/redact-log-credentials.py` | `deploy/systemd-user` | | `scripts/post-pr-visual-proof.sh` | `scripts/redact-log-credentials.py` | The last two agent entries came out of review. `install-agent-engine-host.sh` resolves both through `REPO_DIR`, which the workflow sets to the workspace, so main's installer was installing the target's health probe and rendering the target's systemd unit templates. The template directory goes on the list as a directory rather than three files, because the installer interpolates the unit names. **Application, taken from the target.** Two entries deserve their reasons, because both look like harness: `apps/web-console/e2e/phase-19/` and `apps/agent-console/proof/harness/capture-live.mjs` are the capture drivers, and they select against the target's own DOM. A pull request that changes a selector and its driver together has to be proven with its own driver, never main's. The issue makes this point itself; the dispatch brief's parenthetical listing the chat capture driver as harness is the one place I have gone the other way, deliberately. `scripts/apply-migrations.sh` stays the target's. `ci-throwaway-db.sh` calls it with no arguments, so it has no interface with the YAML at all, and it validates `scripts/migration-baseline.conf` against `supabase/migrations`, both of which are the target's. Pinning the runner while leaving its two data inputs on the other side is the coupling this change exists to remove. Everything else follows from that: compose files, Dockerfiles, the Go services, the forked Open WebUI, the schema, and `tools/lint-no-token-in-proof-captures.mjs`. Two corrections to the issue's own lists. `scripts/ci-seed-api-key.sh` appears in `chat-visual-proof.yml` only inside a comment at line 544 and is never invoked, so it is not on the list. `scripts/generate-enterprise-jwt-keys.py` is reached from `ci-supabase-stack.sh` in both workflows, not just the agent one, so it is on both. ## The guard `tools/lint-visual-proof-harness-split.mjs`, wired as `npm run lint:proof-harness-split` and run in the same required check as its neighbours. The way this list rots is someone adding a step that calls a new script and not adding it to the list, which reintroduces the defect silently on workflows that do not run on most pull requests. The lint fails instead. It asserts every `scripts/...` path a workflow invokes is on that workflow's harness list or on a documented application-side exception list, and that every listed path exists. Comment lines and `paths:` trigger entries are not invocations and are skipped. It carries a MUST_CATCH and MUST_ALLOW self-test that runs as a preflight on every invocation, matching `lint-no-token-in-proof-captures.mjs`. It reads workflow YAML for what the steps invoke, and each listed script for the paths it reaches through its own repo-root variable, requiring both to be listed or declared. That second half is what turns the live seam under the exception, main's `ci-throwaway-db.sh` calling the target's `apply-migrations.sh`, from an invisible coincidence into a declared entry carrying its reason. The allowlist is a map from path to reason, so a target-side read cannot be added silently. One ceiling, stated in the source: the transitive scan keys on three repo-root variable names rather than doing dataflow. A wide scan for any repo-shaped substring was tried first and is unusable, matching container image names, URL paths and references inside Python docstrings, which would bury five real entries among twelve. Verified against real mutations rather than only the fixtures. All five go red: an unlisted invocation, a listed path that does not exist, an empty list, the step deleted outright, and a transitive call to an unlisted script from inside a listed one. ## Verification Two dispatches against a deliberately stale throwaway pull request, #1765, whose single commit reverts `scripts/ci-supabase-stack.sh` to its pre-#1739 content. That is what makes the control honest: `refs/pull/N/merge` is recomputed against current main, so a branch merely cut from an old commit picks the current harness back up and proves nothing. Reverting the file in the branch reproduces the tree shape #1758 describes deterministically. **Negative control, run [33701602008](https://github.com/sakibsadmanshajib/hive/actions/runs/33701602008).** The `pull_request` arm on main's unfixed YAML. Failed at `Stand up this run's own Supabase, with the OAuth server on` nine seconds in: ``` unknown argument: --oauth-server ##[error]Process completed with exit code 2. ``` Byte for byte the failure of run 33690589909 in the issue, so the fixture is genuinely stale. **The fix, run [33701642279](https://github.com/sakibsadmanshajib/hive/actions/runs/33701642279).** A `workflow_dispatch` of this branch's YAML at the same pull request. `gh workflow run ... --ref ci/1758-harness-from-main` runs the workflow file from this branch and sets `github.sha` to its head, so this arm is available before merge, and the dispatch path is the one that had to be proven. The harness step reported: ``` harness taken from 70fabce. Replaced, against the target's own copies: scripts/ci-supabase-stack.sh | 96 +++++++++++++++++++++++++++++++++++++++++--- 1 file changed, 90 insertions(+), 6 deletions(-) ``` One file replaced, the stale one, and the rest of the target's tree left alone. The run then went green end to end, not merely past the argument: the Supabase step succeeded, so did `Sign in and capture the proof`, `Refuse an empty capture` and `Post the captures on the pull request`. **agent-visual-proof.yml, run [33702459551](https://github.com/sakibsadmanshajib/hive/actions/runs/33702459551).** Dispatched the same way at the same pull request. Its harness step passed with the identical replacement, which is the part this change makes. Cancelled straight after, since the agent job's remaining twenty minutes exercise the sandbox rather than anything here. #1765 is closed and its branch deleted. Local, on the working tree: `npm run lint:proof-harness-split` passes on both workflows and its self-test; both workflow files and `ci.yml` parse; `lint:deploy-diagnosability`, `lint:compose-required-vars` and its self-check still pass. `lint:spec-wiring` cannot run on this box, it needs Playwright browsers in `apps/web-console`, and it is unaffected by an added step. ## Buglog entry ```json {"id": "1758-visual-proof-harness-from-target-tree", "date": "2026-09-03", "title": "Both visual-proof workflows ran main's YAML against the target pull request's harness scripts, so any pull request branched before a harness change could not be proven", "error_message": "chat-visual-proof.yml run 33690589909, workflow_dispatch against PR #1730, failed nine seconds in at 'Stand up this run's own Supabase, with the OAuth server on' with 'unknown argument: --oauth-server' and exit code 2, nowhere near the sabotage step it was dispatched to exercise", "root_cause": "Both proof workflows resolved refs/pull/N/merge and made it the only checkout in the job, so every step ran against the target pull request's tree. GitHub reads the workflow YAML from the default branch on a dispatch, so the command lines were current while the scripts those command lines invoked were whatever the target happened to carry. The --oauth-server arm of scripts/ci-supabase-stack.sh was added by #1739 and existed only on main. Seven harness scripts in chat-visual-proof.yml and five in agent-visual-proof.yml were exposed the same way, and the two workflows share three of them.", "fix": "Each workflow now re-takes a named list of scripts from github.sha immediately after the target checkout, which is the default branch's head on a workflow_dispatch and the pull request's own merge commit on a pull_request event. The overlay lands them at their normal paths so no call site changes and the scripts keep resolving their data files, supabase/migrations and scripts/migration-baseline.conf among them, from the target's tree. tools/lint-visual-proof-harness-split.mjs, wired as npm run lint:proof-harness-split in ci.yml, fails when a workflow invokes a scripts/ path that is on neither its harness list nor a documented application-side exception list.", "tags": ["ci", "github-actions", "visual-proof", "workflow", "checkout", "harness"]} ``` 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01WbVmp2Uh7FCgnqKB2TuBb5 --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>



























Closes #1738.
The chat-surface sibling of
.github/workflows/agent-visual-proof.yml. That workflow broke the proof-before-merge circularity for the Cowork agent surface; this one does the same for chat.Why it was circular
A signed-in session is not decoration on this surface, it is the only way the surface works.
deploy/docker/pipelines/hive_jwt_forward.pyanddeploy/docker/owui-patches/hive_upstream_auth.pyattach the signed-in user's OAuth access token to every outgoing completion, and edge-api'sOWUIUnwrapmiddleware treats/v1/chat/completionsas unconditionally requiring a per-user token. With no session, nothing is attached, the static shim key stands alone, and edge-api answers 401 for every model. There is no "photograph the chat surface without signing in" shortcut, and a stand-in identity provider would have produced a page that cannot hold a conversation. That is why this workflow performs the whole real sign-in rather than something cheaper.Signing in needs GoTrue's OAuth 2.1 authorization server, a registered OAuth client for Open WebUI, and the console's own
/oauth/consentscreen. None of it could be stood up per run:scripts/ci-supabase-stack.shpinnedsupabase/gotrue:v2.170.0, which has no OAuth server at all and answers 404 on every routescripts/register-owui-oauth-client.pyuses, andowui-nightly.yml, the one job that boots Open WebUI, points at the hosted Supabase project deleted during the move to the self-hosted data plane, together with the OAuth client that lived in its dashboard. The only stack that could render a signed-in chat was the demo box, which only runs already-merged code.Proof that it works
Run
33681434827posted three screenshots of a signed-in Hive chat on this pull request: the empty composer, the model picker, and a real streamed completion onHive Free. The images live on the permanentvisual-proof-assetsrelease, so they survive this branch being deleted on merge.The first green capture, run
33677630521, is also what exposed two defects in the job that made it: the run tenant could see no free alias, so the captured turn ran on a paid one, and it had a zero credit balance, so the published image carried an out-of-credits banner over the composer. Both are now provisioned and asserted, and the capture reads the composer's alias and refuses to send on one that does not match.Trigger, and why
workflow_dispatchaimed at a pull request number, pluspull_requeston the Hive-owned chat paths.Dispatch is the primary arm because only the dispatcher knows which pull request claims to change this surface. The path trigger exists so the job cannot rot:
agent-visual-proof.ymlrecords at length that it failed on every branch for four days with nothing saying so, and a proof job nobody runs is a proof job nobody trusts. The path list is deliberately the Hive-owned chat paths and notvendor/open-webui/**as a whole, because that tree is a vendored fork and a routine upstream sync touches thousands of files this job could attribute nothing to.Like the agent proof job, this one is not a required check. It announces a red run on a single reused tracking issue instead. Making it required would block unrelated pull requests and turn a slow provider into a merge stopper.
The dispatch arm itself is unverified, and cannot be verified before this merges.
workflow_dispatchonly offers a workflow that exists on the default branch, so neither the dispatch arm nor itssabotage: truenegative control can be fired from this branch. Every green run below reached the job through thepull_requestarm. Whoever merges this is accepting an outstanding verification, not a tested path: dispatch once against a pull request number after merge, and once withsabotage: true, before treating either as working.Measured runtime
11m07s end to end on run
33681434827, and 10m53s on33677630521before the LiteLLM sync step was added. This run's own Supabase 27s, the stack build and boot 7m36s (most of it compiling the forked Open WebUI frontend from source), harness dependencies and a browser 42s, the console build 25s, the whole sign-in journey 13s, the capture 7s, posting 9s. The job budget is 40 minutes: a cold Open WebUI boot has its own 420s deadline and the captured turn waits up to 180s for a provider to finish streaming, and a wedged one of either should fail inside the budget rather than sit until GitHub's six hour ceiling.What changed
.github/workflows/chat-visual-proof.yml(new). Throwaway Postgres, this run's own Supabase with the OAuth server on, registers the Open WebUI OAuth client, seeds a run-scoped tenant, user, membership and shim key, grants that tenant a free alias and a credit balance, boots the chat stack from the pull request's own merge ref, serves the console for the consent screen, signs in through the real journey, sends a real completion, posts the screenshots.scripts/ci-supabase-stack.sh. Two opt-in flags, both defaulting to the previous behaviour so existing callers are byte-for-byte unaffected.--oauth-serverselects the same GoTrue pindocker-compose.enterprise.ymlalready runs on the demo box and enables the authorization server.--external-urlpinsAPI_EXTERNAL_URLand theissclaim to one caller-chosen origin, which is what lets the browser, Open WebUI's container and the console agree on a single Supabase address. Also two robustness fixes that affectagent-visual-proof.ymltoo, both measured: the JWKS assertion no longer races its own TLS listener, and the GoTrue image pull is retried.apps/web-console/e2e/phase-19/sign-in-with-hive.ts. Stops requiring the "Continue with Hive" button, whichOAUTH_AUTO_REDIRECTcan bounce past before it is clickable. The assertion that matters, reaching the consent origin, is unchanged.apps/web-console/e2e/phase-19/owui/owui.setup.ts. The bootstrap login accepts the activation-pending screen. Since issue Open WebUI admin is derived from tenant OWNER, an active hazard for every legitimately provisioned tenant owner, not merely a latent one #748 removed the self-promotion, the first account a fresh container sees is deliberately pending, so demanding chat there failed on every freshly created container.apps/web-console/e2e/phase-19/owui/capture-chat-proof.mjs(new) and onepackage.jsonscript.What is reused rather than rebuilt
scripts/register-owui-oauth-client.py,scripts/seed-owui-e2e-user.py,scripts/ci-throwaway-db.shandscripts/apply-migrations.sh,deploy/docker/docker-compose.agent-proof.ymlverbatim for the JWKS-over-TLS front, theowui-setupPlaywright project for the one implementation of the sign-in journey,apps/agent-console/proof/harness/stamp.mjsfor the capture footer, andscripts/post-pr-visual-proof.shfor the upload.Safety
OWUI_E2E_RUN_KEYnamespaces every fixture address per run, so this run provisions its own identities in a GoTrue destroyed with the runner and shares no credential with any other run (docs/live-test-auth.md).stamp.mjs's footer never carries a URL, andnpm run lint:proof-tokensruns over the capture before anything is published. Every step that could republish it is gated on that lint passing.index.htmlare excluded from the failure artifact: a trace of this journey holds session cookies and the OAuth callback'scode, no text linter can inspect either, and this repository is public. One artifact from an earlier run of this branch was deleted for exactly that reason before this exclusion existed.agent-visual-proof.ymldocuments.Observation, not fixed here, now filed as #1757
Open WebUI takes about 112 seconds to pass its first health check against a fresh database. Nothing in this repository defines that healthcheck:
docker-compose.yml,Dockerfile.open-webuianddocker-compose.enterprise.ymlall set none, so it is inherited whole from the upstream base image, which carries a test and no other field. Docker's defaults therefore apply, start period 0s, interval 30s, three retries, and the container is declared unhealthy at about 90 seconds, roughly 22 seconds before it is actually up. Docker recovers on its own butcaddy-owui'sdepends_on: service_healthydoes not wait, so a colddocker compose upaborts with "dependency failed to start" naming no cause.This job works around it by staging the boot and waiting on its own deadline. Widening the healthcheck would change how every deployment detects a genuinely dead chat container, so it is a product decision and is left alone here. It is not left only in this description either: the defect is issue #1757, with the measurements, the
docker image inspectoutput they came from, and the options. It affects every cold boot against an unmigrated database, so a fresh enterprise install and a fresh local dev stack, not only CI.Buglog entry
{"date":"2026-09-02","title":"postgres:// DSN kills Open WebUI at boot","error_message":"sqlalchemy.exc.NoSuchModuleError: Can't load plugin: sqlalchemy.dialects:postgres","root_cause":"compose hands SUPABASE_DB_URL to open-webui as PGVECTOR_DB_URL. SQLAlchemy resolves a URL scheme to a dialect plugin by name and registers none under 'postgres', so the container exits before uvicorn binds. Both Go services accept the alias because pgx does, so the same string is valid for two of the three consumers and fatal for the third.","fix":"write the throwaway DSN as postgresql:// in .github/workflows/chat-visual-proof.yml","tags":["dsn","open-webui","sqlalchemy","ci","compose"]} {"date":"2026-09-02","title":"docker compose logs without a profile reports on the wrong project","error_message":"dependency failed to start: container hive-open-webui-1 is unhealthy, and a log dump naming open-webui returned only edge-api, control-plane, litellm and markitdown","root_cause":"four services in docker-compose.yml carry no profiles key, so a compose call with no profile still resolves a valid project made of exactly those four, succeeds, and reports on it. The one log that could explain the failure was silently absent.","fix":"set COMPOSE_PROFILES at job level alongside COMPOSE_FILE","tags":["docker-compose","profiles","ci","diagnostics"]} {"date":"2026-09-02","title":"JWKS assertion races the TLS listener it just started","error_message":"the JWKS endpoint served no usable key over TLS","root_cause":"scripts/ci-supabase-stack.sh probed once as soon as root.crt appeared. That file existing means Caddy's authority is written, which happens a beat before the listener completes a handshake with the certificate it just issued. The single-shot curl ran 0.2s after 'certificate obtained successfully' and came back empty, so the job blamed GoTrue for publishing no key.","fix":"poll the probe for up to a minute; the same message still fires when the key set really is empty","tags":["tls","jwks","race","ci","gotrue"]} {"date":"2026-09-02","title":"first Open WebUI account is permanently pending on a fresh container","error_message":"Account Activation Pending, on an account holding an ACTIVE tenant membership","root_cause":"utils/oauth.py get_user_role returns DEFAULT_USER_ROLE early when user_count is zero, above the point where owui-patches/tenant_role_from_db.py splices Hive's membership lookup in, and issue #748 deleted the post-insert promotion upstream used to repair the first account. DEFAULT_USER_ROLE here is pending, so the first account a container ever sees can never reach chat. owui.setup.ts demanded that it did.","fix":"accept either chat or the activation screen for the bootstrap login in owui.setup.ts","tags":["open-webui","oauth","roles","e2e","issue-748"]} {"date":"2026-09-02","title":"Continue with Hive helper races OAUTH_AUTO_REDIRECT","error_message":"expect(locator).toBeVisible() failed, getByRole('button', { name: /continue with hive/i }), while the page had already navigated to the console sign-in form","root_cause":"the deployment sets OAUTH_AUTO_REDIRECT, so Open WebUI's landing page starts the authorize chain by itself. signInWithHive required the button first, and whether a load paints it before the bounce is a race: one run won it on the bootstrap login and lost it on the fixture login thirty seconds later.","fix":"click the button when present, accept having been redirected when not; the assertion requiring the consent origin is unchanged","tags":["playwright","oauth","race","e2e"]} {"date":"2026-09-02","title":"restricted free alias silently downgrades a CI capture to a paid model","error_message":"none; the picker listed seven aliases, none free, and the captured turn ran on Deepseek V4 Flash","root_cause":"hive-free is visibility='restricted' since 20260831_01_restrict_free_pool_aliases_visibility.sql and the run's freshly minted tenant carried no tenant_model_visibility row, so Open WebUI fell back to the first alias it could see. Nothing failed and nothing said so; the only evidence was a model name in a screenshot.","fix":"grant the run tenant visible=true on hive-free the way scripts/ci-seed-api-key.sh does, and make the capture read the composer's alias and refuse to send on an unexpected one","tags":["catalog","visibility","billing","ci","owner-directive"]}Test plan
33681434827).hive-free, with the composer's alias read and asserted, and the credits banner reads a positive balance (run33681434827).agent visual proof, which sharesscripts/ci-supabase-stack.sh, still passes on this branch (runs33679174991,33680013891).33680014334reached a signed-in chat, got back "Invalid request for hive-free", went red, uploaded the failure images as an artifact and posted nothing on the pull request.sabotage: truedispatch. Not yet exercisable, and this is a GitHub constraint rather than a skip:workflow_dispatchonly offers a workflow that exists on the default branch, so neither the dispatch arm nor its sabotage input can be triggered until this merges. Everything above ran through thepull_requestarm. The sabotage arm should be run once after merge to confirm the negative control.