From f1980e9897e306897f7ed6e0aa63a890503f18ce Mon Sep 17 00:00:00 2001 From: Sakib Sadman Shajib Date: Sat, 29 Aug 2026 14:09:17 -0400 Subject: [PATCH 1/2] chore: batch buglog entries for the 2026-08-29 merges Append the fifty buglog entries carried in the bodies of the pull requests merged to main on 2026-08-29. The diff is .wolf/buglog.jsonl and nothing else, per the protocol in .claude/rules/openwolf.md. Entries are copied verbatim from their source pull request bodies. No line is rewritten and no field is invented. Entries already present on main, including the thirty six landed by the 2026-08-28 batch in #1342, are skipped. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01WyEwUxZCArdn1ZUDkTvuQ1 --- .wolf/buglog.jsonl | 50 ++++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 50 insertions(+) diff --git a/.wolf/buglog.jsonl b/.wolf/buglog.jsonl index c2cd407db..17495753e 100644 --- a/.wolf/buglog.jsonl +++ b/.wolf/buglog.jsonl @@ -230,3 +230,53 @@ {"id":"bug-2026-08-28-citation-hook-multiedit-blind","timestamp":"2026-08-29T03:18:49.755Z","related_bugs":[],"occurrences":1,"last_seen":"2026-08-29T03:18:49.755Z","date":"2026-08-28","title":"decision-citation-check hook exited 0 on every Claude Code MultiEdit","error_message":"No output, exit 0, on a MultiEdit payload citing a fabricated D-999; the settings matcher invoked the hook and it declined to look","root_cause":"The payload reader handled only Cursor's flat data.edits array. Claude Code nests a MultiEdit's edits under tool_input.edits, so ti.content, ti.new_string and the flat editsContent were all empty, content resolved to empty string, and the guard exited before auditing a single citation. The self-check could not see it because its claude payload shape was only ever tool_name Write with tool_input.content, so the MultiEdit shape was never exercised. Same defect as issue #1333 in secrets-scanner.js, inherited by copying the pattern.","fix":"Adopt the payload handling that fixed #1333 in secrets-scanner.js: read the edits array from either harness position, serialize any field that is not a string instead of coercing it, and audit every source rather than letting the first truthy one win. Run the decision-citation self-check cases over all three write shapes via the shared edit() helper, plus explicit payloads for content-and-edits together, a non-string new_string, and a fabrication in a later edit. Mutation tested: reverting the edits-array read turns nine cases red.","tags":["hooks","claude-code","payload-shape","fail-open","citation-check","pr-1313"],"related":["#1333"],"pr":1313} {"id":"bug-mtdt9x8l-387c1b","timestamp":"2026-08-29T03:18:49.988Z","related_bugs":[],"occurrences":1,"last_seen":"2026-08-29T03:18:49.988Z","date":"2026-08-29","tags":["auth","billing","provisioning","api-keys","console"],"error_message":"The console minted an API key (201, secret shown in the copy-it-now panel, key listed active) for an account with zero public.tenant_billing_accounts rows, and edge-api then rejected that key with 403 account_not_provisioned and an unactionable contact-support message","root_cause":"apikeys.Service.CreateKey and RotateKey generated and persisted a secret without ever resolving the account's billing tenant, even though the exact lookup (Repository.GetTenantIDByAccountID) already existed in the same package for ResolveSnapshot; the unprovisioned state is permanent for any user whose tenant is already billed by a different workspace, since tenant_billing_accounts is UNIQUE on both columns and signup.EnsureTenantBillingAccount refuses to guess a mapping","fix":"added Service.requireBillingTenant as a precondition on both mints, returning a new ErrAccountNotProvisioned sentinel mapped to 409 with an actionable message and the machine code account_not_provisioned; the console proxy route now maps that code to its own customer-facing wording (it never forwards upstream error text) and ApiKeyCreateForm renders a 409 body verbatim instead of replacing it with generic retry text; ResolveSnapshot deliberately still tolerates uuid.Nil so pre-existing keys keep failing closed at the boundary","verification":"four new Go tests plus a mutation check that forces the guard false and confirms all four go red; two new console unit tests; visual proof captured against a branch-built control-plane and console over a throwaway Supabase with all 112 migrations, showing the refusal, zero api_keys rows written, and a positive control where a mapped workspace still issues a key","pr":1335} {"id":"bug-mtdt9xel-580d70","timestamp":"2026-08-29T03:18:50.204Z","related_bugs":[],"occurrences":1,"last_seen":"2026-08-29T03:18:50.204Z","date":"2026-08-28","title":"Secrets scanner hook was blind to every MultiEdit and reported clean","error_message":"No output, exit 0, on a MultiEdit payload containing a hardcoded credential. Indistinguishable from a clean scan.","root_cause":"secrets-scanner.js resolved its edit array from the top-level data.edits field, which is Cursor's file-edit payload shape. Claude Code nests the array at tool_input.edits. On a MultiEdit all three content fallbacks (ti.content, ti.new_string, top-level data.edits) were empty, so content became the empty string and the early-return guard exited 0 with no output. The hooks.selfcheck.js harness fed both harness shapes but used Write for the Claude Code case, whose text lives in tool_input.content, so the nested-array path was never exercised; the harness was also not invoked by CI or any npm script, so even a covering case would not have run.","fix":"Consult both shapes, preferring tool_input.edits over data.edits. Add a claude-multiedit shape to hooks.selfcheck.js and run every secrets-scanner case under all three shapes. Wire hooks.selfcheck.js into the Repo policy lints required CI check so the harness actually runs.","tags":["hooks","secrets","claude-code","payload-shape","silent-failure","ci-coverage","issue-1333"],"pr":1337} +{"id":"sso-wave1-manifest-entry-outside-specs","date":"2026-08-29","title":"Playwright spec manifest entry landed outside the specs object","error_message":"scripts/verify-spec-collection.mjs would report UNDECLARED tests/e2e/sso-silent-consent.spec.ts runs but is not in the manifest","root_cause":"The manifest was rewritten with a JSON serializer that reformatted every array and appended the new spec as a key of the root object rather than of specs, and the guard reads document.specs only. The branch was never re-run against the guard after that edit because the pull request was already conflicting, and a conflicting pull request builds no refs/pull/N/merge, so GitHub created no pull_request run and the page showed only CodeQL, GitGuardian and CodeRabbit, which reads green at a glance.","fix":"Resolve the rebase conflict onto main's compact formatting and add the single line inside specs, then run npm run e2e:verify-collection to confirm 42 files collect into the pinned projects.","tags":["playwright","ci","manifest","rebase","web-console"]} +{"id":"sso-consent-redirect-guard-slash-loop","date":"2026-08-29","title":"OAuth consent landing redirect guard missed repeated slashes","error_message":"ERR_TOO_MANY_REDIRECTS between /oauth/consent and itself for an auto-approve target spelled //oauth/consent","root_cause":"isSafeRedirectTarget compared parsed.pathname with trailing slashes stripped and nothing else, so //oauth/consent and /oauth//consent matched neither the landing path nor its stripped form and were allowed. Next.js then 308-normalized the repeated slashes away and re-entered the landing, which asked GoTrue again and received the same redirect_url. The guard's comment reasoned about Next normalizing a trailing slash into the guarded path and stopped there, which is why it read as thorough and still missed the sibling case.","fix":"Collapse repeated slashes and decode percent escapes inside a try before comparing, and add a unit case that fails against the previous guard. Measured the real routing behaviour against a production console build rather than assuming it: encoded and case variants answer 404, repeated and trailing slashes answer 308.","tags":["oauth","redirect","auth","web-console","security-review"]} +{"date":"2026-08-28","error_message":"n/a","root_cause":"console had no privacy/data-policy surface at all (parity re-score, criterion weight 9, credit 0.25)","fix":"added /console/privacy grounded in verified backend state (UsageEventRow schema, routing.SelectionInput.AllowedProviders dead code path, provider-blind error handling), never rendering a claim or control that is not enforced; corrected mid-review to drop literal OpenRouter/Groq vendor names per the console-wide provider-blind convention, keeping the third-party-routing disclosure itself","tags":["web-console","privacy","compliance","parity"]} +{"timestamp":"2026-08-28T23:20:00Z","error_message":"Anthropic Messages request carrying top_k fails with 400 and the customer-facing message hive-free is not available.","root_cause":"edge-api forwards four non-OpenAI fields verbatim from the Anthropic surface (top_k, thinking, cache_control, session_id) because the OpenAI chat-completions shape has no equivalent. hive-free load balances across providers with different tolerances, so the request succeeded or hard-failed depending on which pool member answered, and the refusal reached the caller as model unavailability rather than as a refused field","fix":"dispatchWithRetry in apps/edge-api/internal/inference/retry.go now classifies a 400 that names one of those four fields, strips that single field, and retries. Bounded to one strip per field and to attempts that have a successor; a field the caller actually sent is never stripped. Known edge filed as issue #1323","tags":["anthropic","edge-api","free-pool","sdk-conformance"]} +{"date":"2026-08-29","title":"Pre-merge frontend gate compiled lib/hive only, so a Hive component added under lib/components was compiled and rendered by nothing","error_message":"make test-owui-frontend reported 16 files, 208 tests, 13/13 components compiled, exit 0 on a tree carrying a Svelte parse error, two transposed money figures and an inert settings tab","root_cause":"scripts/test-owui-hive-frontend.sh copies upstream components into its scratch tree as text fixtures and runs the svelte compile pass over lib/hive only, so a Hive authored component placed under lib/components/chat/Settings was read as text and never compiled, and no test rendered it; the image build that would have caught the parse error runs after merge in deploy-demo-box.yml","fix":"moved the component to vendor/open-webui/src/lib/hive/SettingsUsage.svelte where the compile guard already looks, taught the gate to install the lockfile pinned vite svelte plugin and write a vitest config so tests can import and server side render components, and re-ran the three mutations to confirm each one turns the suite red","tags":["ci","test-gate","open-webui","svelte","frontend","money-surface"]} +{"id":"bug-1329-anthropic-usage-zero","date":"2026-08-29","title":"/v1/messages reported usage input_tokens 0 and output_tokens 0 while billing charged correctly","error_message":"usage: {input_tokens: 0, output_tokens: 0} on every POST /v1/messages response","root_cause":"The post-finish chunk suppression in apps/edge-api/internal/inference/mint_id.go recognised a terminal usage frame as usage != nil AND len(choices) == 0. LiteLLM v1.98.0 sends that frame with one empty-delta choice, so the relay dropped the only usage frame on the stream. The Anthropic surface folds that same relay, so it reported zeros; settlement reads the accumulator, which sees the frame earlier, so billing stayed correct and hid the defect.","fix":"Key the suppression on whether a post-finish frame carries renderable payload rather than on the choice count, re-emit a forwarded terminal usage frame with an empty choices array, move the synthesized fallback usage frame ahead of the single [DONE] sentinel, and report input_tokens on the Anthropic message_delta.","tags":["anthropic","streaming","usage","litellm","sse","issue-1329"]} +{"date":"2026-08-29","area":"web-console","error_message":"Sign-up shows Something went wrong on our end. Reference AUTH-... after POST /auth/v1/signup returns a bare 404","root_cause":"The console shipped a sign-up route on deployments that refuse account creation at the gateway (Caddyfile.supabase 404) and at the GoTrue flag. The empty 404 body makes auth-js raise AuthUnknownError with no status, which the allow-list in lib/auth/auth-error.ts correctly withheld, so a deliberate policy rendered as an outage","fix":"Gate the sign-up route and its cross-links on NEXT_PUBLIC_DISABLE_SELF_SERVE_SIGNUP, fed from ENTERPRISE_DISABLE_SIGNUP so one variable drives GoTrue and the UI, and add toUserFacingSignUpMessage to report a 404, a 403, GoTrue signups-not-allowed copy, or the status-free AuthUnknownError shape as a policy refusal","tags":["web-console","auth","signup","issue-1328"]} +{"date":"2026-08-29","area":"web-console","error_message":"Rotate on /console/api-keys opens This page could not be found","root_cause":"api-key-list.tsx linked every active key to /console/api-keys/[id]/rotate, a route that was never built; the page header also promised rotation. The rotateApiKey client helper was dead and decoded the wrong response shape","fix":"Remove the link and the promise, name the working path (create a replacement, revoke the old) in the header, and delete the broken helper","tags":["web-console","api-keys","issue-1331"]} +{"date":"2026-08-29","area":"web-console","error_message":"Available credits renders 99,996,364,207 with no unit while the API keys table renders the same quantity as $0.000662","root_cause":"Two surfaces rendered the same credit quantity in different denominations, with no conversion shown on either","fix":"One CreditBalance component on the dashboard and the billing page, dollars as the headline with the credit figure and the conversion beneath, and a truncating formatUsdBalanceFromCredits so a balance never rounds up","tags":["web-console","billing","credits","issue-1332"]} +{"date":"2026-08-29","area":"web-console","error_message":"A negative credit balance displayed less debt than the account carried","root_cause":"formatUsdBalanceFromCredits rounded with Math.trunc, which rounds toward zero, so an available balance driven negative by reservations was flattered by one displayed unit","fix":"Round with Math.floor, identical for positive balances and conservative for negative ones, and state the direction in the docstring","tags":["web-console","billing","credits","issue-1332"]} +{"date":"2026-08-29","area":"chat","error_message":"The chat composer credit banner displayed more money than the account held, and its text disagreed with its own low-credit badge at the threshold","root_cause":"vendor/open-webui/src/lib/hive/credits.ts rendered the balance with formatUsdFromCredits, a port of the console PRICE formatter, which rounds to nearest; 9,996,364,207 credits rendered as $10.00 and 499,999,999 rendered as the $0.50 threshold while creditState called the same balance low","fix":"Port the console balance formatter as formatUsdBalanceFromCredits, rounding down in credit space, and move the two balance call sites onto it while today's spend keeps the price formatter","tags":["chat","open-webui","billing","credits","issue-1345"]} +{"id":"bug-2026-08-29-analytics-null-group-key","date":"2026-08-29","title":"Analytics summaries returned 500 when a grouped column was NULL","error_message":"usage: scan error summary row: cannot scan into dest[0] (col: group_key): cannot scan NULL into *string","root_cause":"GetErrorSummary, GetUsageSummary and GetSpendSummary scanned group_key into a non-nullable string. usage_events.api_key_id is nullable (NULL for chat traffic with no API key, for errors recorded before a key is resolved, and for keys deleted under ON DELETE SET NULL), so one unattributable row failed the scan for the entire result set and the console overview card rendered its error state on the first page after sign-in.","fix":"Scan group_key into a nullable destination in all three summaries and map NULL to an explicit unattributed bucket rather than dropping the row. The console Top API keys widget no longer labels that bucket Deleted key. Pinned by TestSummariesByAPIKey_NullKeyGroupsAsUnattributed, TestUnattributedGroupKeyMatchesTheConsole and a console unit case, each proven red by its own mutation.","tags":["control-plane","usage","analytics","postgres","null-scan","console-overview","web-console"],"files":["apps/control-plane/internal/usage/repository.go","apps/control-plane/internal/usage/repository_live_test.go","apps/control-plane/internal/usage/wire_contract_test.go","apps/web-console/lib/analytics/cache-metrics.ts","apps/web-console/lib/control-plane/contract.ts"]} +{"id":"bug-2026-08-29-agent-visual-proof-dependabot-secrets","date":"2026-08-29","title":"agent visual proof is permanently red on Dependabot pull requests because Dependabot events get no repository secrets","error_message":"HIVE_AGENT_ENGINE_LLM_API_KEY is required (value never printed)","root_cause":"A pull_request event raised by Dependabot is served from GitHub's separate Dependabot secrets store, so every secrets.* reference in agent-visual-proof.yml resolves to the empty string. The job's gate only excluded fork pull requests, and a Dependabot branch lives in the same repository, so it passed that check and then failed on the first secret read. Run 33234962802 shows the SIF correlation step succeeding and the launcher install step failing with LITELLM_MASTER_KEY, the whole S3 block and NEXT_PUBLIC_EDGE_API_BASE_URL blank in the same step log. Misdiagnosed twice as a trigger-path or artifact-correlation problem, which the step-level evidence rules out.","fix":"Skip the job for Dependabot pull requests alongside fork pull requests (github.actor != 'dependabot[bot]'), and route a bump that needs proof through the existing workflow_dispatch, which runs with real secrets.","tags":["ci","github-actions","dependabot","secrets","agent-visual-proof","false-red"]} +{"id":"bug-2026-08-29-sif-artifact-retention-clock","date":"2026-08-29","title":"agent visual proof gets less reliable the longer the sandbox image inputs stay unchanged","error_message":"no unexpired agent-engine-sif artifact was built from these image inputs","root_cause":"agent-engine-sif.yml builds only on a change to deploy/apptainer, apps/agent-engine/packs or vendor/openhands, while the uploaded artifact expires after 90 days regardless. The proof job refuses any artifact whose run did not build the same three trees, so a quarter without an image-input change expires the last match and every proof run hard-errors for a reason no pull request caused. A stable tree made CI less reliable, not more.","fix":"Add a monthly schedule to agent-engine-sif.yml so a matching unexpired artifact always exists, and name that schedule in the proof job's error message so a future hit is diagnosable.","tags":["ci","github-actions","artifact-retention","agent-visual-proof","latent"]} +{"date":"2026-08-29","title":"Supabase Storage refused every S3 request because the server had no S3 protocol credentials","error_message":"s3 PUT /storage/v1/s3/hive-files/... failed with status 403: AccessDenied: Missing S3 Protocol Access Key ID or Secret Key Environment variables","root_cause":"The self-hosted supabase-storage service was never given S3_PROTOCOL_ACCESS_KEY_ID, S3_PROTOCOL_ACCESS_KEY_SECRET or S3_PROTOCOL_PREFIX. In single tenant mode Storage recomputes the SigV4 signature from that pair, so with neither set it refuses every signed request while staying healthy with the buckets present. The documented recipe compounded it by telling operators to point both client variables at the service_role key, which is not an S3 credential and discloses itself through the access key id that SigV4 sends in the clear.","fix":"Hand supabase-storage S3_PROTOCOL_ACCESS_KEY_ID and S3_PROTOCOL_ACCESS_KEY_SECRET from the same S3_ACCESS_KEY and S3_SECRET_KEY the consumers sign with, plus S3_PROTOCOL_PREFIX=/storage/v1 to match the prefix Caddy strips. Corrected the credential guidance in .env.example and docker-compose.yml, and added two guards to scripts/test_selfhost_supabase_seam.py.","tags":["storage","supabase","s3","sigv4","compose","config","issue-1282"]} +{"id":"2026-08-29-demo-box-disk-90pct","date":"2026-08-29","area":"ops/deploy","error_message":"Demo box root filesystem at 90 percent (11 GB free of 98 GB) with deploy-demo-box.yml free to build into the remainder; run 33237021243 (job 99059755081) died at 05:55:11Z with 'Unhandled exception. System.IO.IOException: No space left on device' while the Actions runner was writing its own _diag worker log, reporting failure with an EMPTY failed-steps array because the runner process died rather than a step failing","root_cause":"The deploy job's two cleanup steps (docker image prune -f, docker builder prune -f --filter until=24h) both run only at the END of the job. `if: always()` covers a failed step but not the runner being killed, so a job that dies on ENOSPC or trips timeout-minutes mid build skips both. Separately, neither step targets what actually accumulated: tagged images that no container references and nothing in the repo names (three Playwright versions, superseded litellm, stale node and golang bases, 15 GB total). Dangling images were 0 and build cache was 0 B reclaimable, so both existing prunes were reclaiming approximately nothing.","fix":"Reclaimed 15 GB by removing 13 explicitly named unreferenced images with `docker rmi` (no prune, no --force, so docker refuses anything a container holds), taking the box from 12 GB to 27 GB free with all 22 running containers untouched, then a further 5.757 GB later that day with `docker builder prune -f` (no -a) once deploys had superseded those layers, ending at 25 GB free. Added scripts/check-deploy-disk.sh as the FIRST step of BOTH the migrate and deploy jobs, the only position a job timeout or ENOSPC death cannot skip: warn below 25 GB, hard fail below 15 GB (above the 11 GB at which the incident happened), printing docker system df and naming the unsafe reclaim moves. Floor overridable only via a workflow_dispatch input, never from a commit message. Guarded by scripts/test-deploy-disk-gate.sh in a required CI lane.","tags":["disk","docker","deploy-demo-box","prune","demo-box","timeout","ops"]} +{"id":"bug-mtdw1319-a41c02","timestamp":"2026-08-29T07:05:00.000Z","related_bugs":[],"occurrences":1,"last_seen":"2026-08-29T07:05:00.000Z","date":"2026-08-29","title":"images.generate answered 200 with an empty data array and billed the hold","error_message":"POST /v1/images/generations returned HTTP 200 with data: [], no image and no error, and the flat image reservation was finalized for it","root_cause":"Neither handleGeneration nor handleEdit checked the decoded ImageResponse for a usable payload before settling and writing 200; the only route with supports_image_generation is a text model carrying a legacy capability flag, so an imageless 2xx is its normal answer","fix":"Added hasImagePayload, counting entries that actually carry a url or b64_json rather than array length, and refused before settlement on both paths with the sanitized upstream error at 502 while releasing the hold; the upstream body handed to the sanitizer is bounded at 4 KiB so its operator log line cannot carry a multi-megabyte 2xx body","verification":"Six wire-level subtests over the three imageless shapes on both endpoints, asserting status, envelope, absence of a data key, hold released and not finalized; mutation with the guard disabled turns all six red","tags":["images","billing","fake-success","edge-api","issue-1319"],"pr":1371} +{"id":"bug-mtdw1318-7be519","timestamp":"2026-08-29T07:05:01.000Z","related_bugs":[],"occurrences":1,"last_seen":"2026-08-29T07:05:01.000Z","date":"2026-08-29","title":"audio.speech rejected every OpenAI voice name and answered 500","error_message":"POST /v1/audio/speech with the OpenAI default voice alloy failed with HTTP 500 and a bare internal error, making the endpoint uncallable by an unmodified OpenAI SDK","root_cause":"The handler rewrote only the model key and forwarded voice verbatim to groq/orpheus-v1-english, whose roster is six entirely different names; the upstream 400 then reached the caller through the provider-blind path as a sanitized 500, so a caller could not tell a bad parameter from a broken gateway and a retry could never succeed","fix":"Added resolveVoice, translating the eleven OpenAI stock names onto the six upstream voices and normalizing case and whitespace, with the resolved value written back into the dispatched body; a name in neither set is refused before route selection with a 400 invalid_request_error naming param voice, code invalid_value, and the supported roster, so no reservation is taken for a request that cannot succeed","verification":"Wire-level tests over all eleven stock names plus case and whitespace variants asserting the DISPATCHED body carries an accepted voice, an agreement test that every id advertised by GET /v1/audio/voices is accepted, and refusal tests for unknown, missing and blank voices; mutation disabling the guard and the rewrite turns sixteen subtests red","tags":["audio","tts","openai-compat","status-mapping","edge-api","issue-1318","issue-996","issue-1285"],"pr":1371} +{"id":"bug-mtdw1348-c93d44","timestamp":"2026-08-29T07:05:02.000Z","related_bugs":[],"occurrences":1,"last_seen":"2026-08-29T07:05:02.000Z","date":"2026-08-29","title":"Malformed chat requests were answered as a model availability problem","error_message":"An invalid message role and an oversized max_tokens both returned 400 with message hive-small is not available. and code upstream_error, pointing a customer debugging their own payload at model status","root_cause":"Two causes. Nothing validated the messages array, so a bad role was forwarded, refused upstream, and the LiteLLM fallback-group bookkeeping that came back was collapsed by sanitizeProviderBlindMessage into the is not available sentence regardless of status. Separately WriteProviderBlindUpstreamError labelled every non-429 non-503 upstream failure api_error with code upstream_error, including a 400 that means the caller request was refused. The pre-existing empty-messages check read len on a json.RawMessage, which is a byte length, so messages: [] passed it too","fix":"Added validateChatMessages refusing an unknown or missing role, an empty array and a non-object entry with invalid_request_error, the offending param and codes invalid_value, empty_array and invalid_type, matching the n refusal shape from PR 1305. Added providerBlindRequestShaped so an upstream 400 or 422 is relabelled invalid_request_error with code invalid_request and gets a request-shaped sentence instead of an availability verdict, while 404 and every 5xx keep theirs","verification":"Five wire-level subtests for the malformed shapes, a positive test that all six valid roles still pass, and two error-package tests pinning both sides of the status gate; mutations disabling the validator, cutting the role list, and forcing the status gate to false and to true each turn the matching test red, with the false mutation reproducing the reported body verbatim","tags":["error-mapping","chat-completions","validation","provider-blind","edge-api","issue-1348"],"pr":1371} +{"date":"2026-08-29","tags":["console","web","ui","dark-mode"],"title":"Analytics charts inverted in dark mode because their colours were hex literals","error_message":"Every analytics chart rendered a light panel with light-mode series colours against the dark console palette","root_cause":"The three chart components passed hex literals to recharts for the panel background and every series, and left the grid, axes, tooltip and legend at recharts light-tuned defaults. A colour handed to recharts as a prop is not reachable from a stylesheet, so the prefers-color-scheme override could not correct any of it.","fix":"Added a chart-theme module holding CSS variable references for the panel, grid, axes, tooltip, legend and every series, and pointed all three charts at it. A guard test reads the sources and fails on a quoted hex literal."} +{"date":"2026-08-29","tags":["console","web","ui","empty-state"],"title":"Console tables and charts reported an empty account while they were still loading","error_message":"Tables flashed 'No records yet.' during every fetch and empty charts rendered three bare axis frames with no explanation","root_cause":"DataTable had no pending state at all, so the empty branch was the only thing it could render before data arrived. ChartCard passed an empty series straight to recharts, which draws its axes regardless.","fix":"DataTable takes a loading prop and renders a pending row with aria-busy. ChartCard takes the rows it wraps and renders an empty state at the chart own height instead of rendering the chart."} +{"date":"2026-08-29","tags":["chat","owui","frontend","noise"],"title":"Chat front end reported two non-errors to the user and the console","error_message":"No token found in localStorage warned on every signed-out page load, and File not found. toasted whenever the file picker was cancelled","root_cause":"The socket connect handler treated an absent token as worth warning about when it is the normal anonymous state, and both composers treated a change event with an empty FileList as a missing file when it is also what pressing cancel produces.","fix":"Dropped the warning and kept the emit guard. Dropped the toast on an empty selection in both composers. A guard test pins all of it, plus the removal of the dead tools branches."} +{"id":"bug-1349-credit-precision-composer-overlap","date":"2026-08-29","title":"Balance rendered nine significant figures and the composer toolbar overlapped at 375px","error_message":"$0.000000858 shown as a credit balance in the chat banner and Settings > Usage; Cowork (x124-199) and Hive Auto (x145-272) drawn on top of each other at 375px","root_cause":"Two unrelated causes. (1) formatUsdBalanceFromCredits and the chat spend formatter both scaled decimal places to the value with Math.min(9, Math.max(2, 2 - magnitude)), a rule ported from the catalog PRICE formatter where nine decimals is correct, so a sub-cent balance rendered at the full width of one credit in chrome the customer cannot dismiss. The rule was hand-duplicated across two builds whose test suites cannot see each other, which is also how #1344 and #1345 happened. (2) The composer control row's left group was flex-1 min-w-0 while its children were shrink-0 and .hv-mode was flex-shrink: 0, so shrinking the group pushed its content out of the box and over the model chip.","fix":"Format balances and spend in whole cents, floored for balances and rounded for spend, with a sub-cent bound of \"< $0.01\" so a real non-zero figure still never renders as the empty-wallet string. Added tools/lint-credit-balance-formatter-parity.mjs, wired as a CI step, to fail the build when the two copies of formatUsdBalanceFromCredits diverge. Made the composer control row wrap, with the left group taking a full line below sm and the previous basis restored from sm up.","tags":["money","formatting","frontend","open-webui","web-console","layout","responsive","cross-build-duplication"]} +{"id":"bug-1372-hive-auto-flat-reservation-hold","date":"2026-08-29","title":"Hive Auto unusable below 2.00 USD: flat envelope-sized credit hold refused every request","error_message":"You exceeded your current quota, please check your plan and billing details. (429 insufficient_quota) shown while the account held 0.455 USD","root_cause":"hive-auto is priced upstream_actual, so the gateway takes an up-front credit hold instead of charging a catalog rate. reservation_estimate_credits was a flat 2000000000 (2.00 USD), the price of the largest request the variable-price bounds allow, and every request paid that authorization regardless of its actual size. control-plane's enforcePolicy refused any account whose balance was under it. Three silent enablers: the session chat path never ran EnforceVariablePriceBounds at all, so its hold covered a request with no size cap or completion ceiling; route-openrouter-auto-live carried no provider.max_price, so the rate half of the hold's coverage proof did not exist; and it carried no usage.include, so OpenRouter reported no cost and every settlement fell through to charging the full hold unconfirmed. The refusal text was OpenAI's canonical insufficient_quota sentence, which named a plan Hive does not sell.","fix":"Size the hold from the request: len(body) as a prompt-token upper bound plus the effective completion ceiling already written into the body, priced at the route's own rate ceiling through CreditsForUpstreamCost, capped by the catalog envelope and floored by the endpoint default. Run EnforceVariablePriceBounds on the session chat path. Add route-openrouter-auto-live to deploy/litellm/config.yaml with usage.include and provider.max_price. Split the refusal text by status and say what happened.","tags":["billing","reservation","upstream_actual","hive-auto","litellm-config","chat","error-message","money-path"],"files":["apps/edge-api/internal/inference/pricing.go","apps/edge-api/internal/chat/dispatch.go","apps/edge-api/internal/chat/billing.go","apps/edge-api/internal/inference/reservation_guard.go","deploy/litellm/config.yaml","apps/web-console/lib/quickstart-model.ts"],"issue":1372} +{"id":"bug-1367-console-mobile-nav","date":"2026-08-29","title":"Console had no navigation at all below the lg breakpoint, and three of its tables scrolled the whole document sideways at 375px","error_message":"At 375px on /console/api-keys the entire tabbable set was Documentation, Switch to EN, Switch to বাংলা, Sign out, Revoke: no route to any console section from any page. Separately, document.documentElement.scrollWidth was 803 on /console/api-keys and 818 on /console/catalog against a 375px viewport, and window.scrollX reached 428.","root_cause":"console-shell.tsx rendered the sidebar as `hidden lg:flex` with no small-screen replacement, so all thirteen nav links were display:none below 1024px. The overflow was two stacked faults in DataTable's wrapper: `overflow-hidden` clipped every column past the fold instead of scrolling it, and because the wrapper was not a containing block, Chromium propagated the over-wide table's layout overflow past that overflow:auto ancestor to the viewport, so the page scrolled sideways even though the scroller itself was the correct width.","fix":"Made the same aside element the drawer (fixed overlay plus scrim below lg, static grid column at lg and up) driven by a small client ConsoleFrame, keeping ConsoleShell a Server Component because it renders the async LocaleSwitcher. Closed state keeps `hidden` so the links stay out of the tab order. Changed DataTable's wrapper to `relative overflow-x-auto`.","tags":["console","responsive","accessibility","navigation","css-overflow","chromium","issue-1367"]} +{"id":"bug-1367-ledger-order","date":"2026-08-29","title":"Credit ledger returned a customer's money history in random order","error_message":"The billing overview's 'last five ledger events' rendered as 11, 16, 26, 16, 24 August, and the paginated Ledger tab was shuffled the same way.","root_cause":"ListEntriesWithCursor ordered by `id` and paginated on `id < cursor`. credit_ledger_entries.id is `gen_random_uuid()`, a v4 UUID, so it encodes no time order whatsoever. Both the overview and the Ledger tab read that one query, so a view-level sort would have fixed only half of it. An existing live test asserted descending id under the name 'default limit returns newest first', which encoded the bug as the requirement and kept it green.","fix":"ORDER BY created_at DESC, id DESC, with the keyset expressed as (created_at, id) < (SELECT created_at, id FROM credit_ledger_entries WHERE id = $cursor) so the opaque cursor stays the entry id and the API contract is unchanged. Corrected the existing assertion to compare created_at, and added three live Postgres tests using fixed UUIDs whose id order is the exact inverse of time order.","tags":["billing","ledger","postgres","pagination","keyset","control-plane","camouflaged-test"]} +{"id":"owui-skills-guard-on-page-not-layout","date":"2026-08-29","error_message":"/skills/create and /skills/edit rendered the full skill editor for a session without workspace.skills; the save then failed with 401 from POST /api/v1/skills/create","root_cause":"the permission guard was written in routes/(app)/skills/+page.svelte. A SvelteKit +page.svelte applies to its exact route only and is not inherited by nested routes, unlike +layout.svelte, so the two editor routes had no guard at all","fix":"moved the guard into routes/(app)/skills/+layout.svelte, which covers the index and both editor routes; skills-surface.test.ts now reads the layout rather than the page","tags":["owui","svelte","authz","sveltekit-routing"],"pr":1388} +{"id":"owui-workspace-components-need-workspace-chrome","date":"2026-08-29","error_message":"the skills index rendered underneath the sidebar with the New Skill button pinned to the window corner and the title overlapping the Hive wordmark","root_cause":"Skills.svelte and SkillEditor.svelte are upstream Workspace components authored for routes/(app)/workspace/+layout.svelte's container, which supplies both the md:max-w-[calc(100%-var(--sidebar-width))] constraint and the px-3 md:px-[18px] padding. Mounting them in a bare hv-panel-region flex region supplied neither","fix":"the /skills route layout reproduces that container minus the tab bar; a test pins all three class fragments so the chrome cannot silently go away again","tags":["owui","css","layout","fork"],"pr":1388} +{"id":"owui-group-access-grants-never-filtered","date":"2026-08-29","error_message":"a non-admin could attach an access grant naming any group id on the shared instance, including another tenant's, and have it stored","root_cause":"utils/access_control.filter_allowed_access_grants strips public grants and individual user grants for a non-admin and never inspects group grants at all. For skills that is cross-tenant prompt injection, because a skill body is appended verbatim to the chat request as a system message. Unreachable in practice only because GET /api/v1/groups filters to the caller's own memberships, which is secrecy of a UUID rather than an authorization check","fix":"owui-patches/apply_skill_group_grants_patch.py drops group grants naming a group the caller is not a member of, at all three skill write sites; the image build asserts all four markers landed. The general case across the other routers is filed as #1396","tags":["owui","security","authz","multi-tenancy","prompt-injection"],"pr":1388} +{"id":"agy-runs-outside-the-repo","date":"2026-08-29","error_message":"an Antigravity review returned detailed findings about a credit-formatting diff that does not exist on the branch under review","root_cause":"agy executes in ~/.gemini/antigravity-cli/scratch, not in the directory it is launched from. Measured by asking it to print pwd, git rev-parse --abbrev-ref HEAD and git diff --stat: the first returns the scratch path and both git commands fail with 'not a git repository'. Given no diff to read, it produced a plausible review of something else","fix":"paste the diff and the load-bearing surrounding code inline in the prompt rather than telling it to run git; discard any pass whose findings name files not in that paste","tags":["harness","review-streams","antigravity","false-green"],"pr":1388} +{"id":"shared-scratchpad-filename-collision-crossed-prs","date":"2026-08-29","error_message":"six review threads on PR #1083, another agent's pull request, were resolved by an agent that had never read them","root_cause":"agents in one session share a scratchpad directory. A helper written to the plain name resolve-threads.sh was overwritten by another agent's script of the same name, which targeted pull request 1083, and running it acted on their PR. Same failure class as the known parallel-agent report-path collision, but the blast radius is a GitHub mutation rather than a lost file","fix":"unresolved all six immediately to restore the state found, left a note on #1083 saying what happened and that the restore was blind rather than a judgement, and prefixed every later script with the worktree id","tags":["harness","parallel-agents","scratchpad","github"],"pr":1388} +{"id":"BUG-1198-upstream-actual-capture","date":"2026-08-29","title":"Zero-content capture replaced a reported upstream cost with the full authorization hold","error_message":"actual_credits = 100000000 where UpstreamActualSettlement had computed 560000 for the same generation","root_cause":"The zero-content branch in apps/edge-api/internal/inference/orchestrator.go overwrote actualCredits with capCaptureAtCeiling, whose contract is to return its credits argument untouched for an upstream_actual route; that argument is reservation.Held(). A 178x overcharge on hive-auto and openrouter-auto, the alias family with no catalog price to fall back on.","fix":"Guard the branch on !route.Pricing.IsUpstreamActual() at the call site rather than relying on the early return that caused it. UpstreamActualSettlement is already fail-closed for that route in all three of its outcomes.","tags":["money","billing","settlement","edge-api","overcharge","upstream-actual","issue-1198"]} +{"id":"BUG-1198-cached-prompt-double-charge","date":"2026-08-29","title":"A fully cached prompt was estimated back on top of its own cache tokens","error_message":"captureInputTokens returned a full-prompt estimate for a usage block whose fresh input count was a legitimate zero, and capCaptureAtCeiling then priced that estimate as fresh input alongside the same prompt's cache-read tokens","root_cause":"captureInputTokens tested freshInputTokens > 0 rather than whether the usage block reported any input quantity. NormalizeCacheUsage produces fresh = 0 for a 100 percent cache hit in the inclusive shape, so the real measurement was discarded and the prompt was billed twice, once at the full input rate and once at the cache rate.","fix":"Test hasUsage && freshInputTokens + cacheReadTokens + cacheWriteTokens > 0, which distinguishes nothing-reported from zero-reported, and pass the cache components from both call sites.","tags":["money","billing","settlement","edge-api","cache-tokens","double-charge","issue-1198"]} +{"id":"BUG-1198-hold-capture-overcharge","date":"2026-08-29","title":"Fail-closed hold capture billed the flat reservation hold when the caller set no max_tokens","error_message":"credit_reservations rows with consumed_credits = 100000000 and terminal_usage_confirmed = false, up to 355872x the median confirmed charge on the same alias","root_cause":"capCaptureAtCeiling in apps/edge-api/internal/inference/completion_ceiling.go short circuited on ceiling <= 0 and returned reservation.Held() unchanged, so both fail-closed capture branches (#1171 zero-content, #1215 missing usage block) charged the flat authorization floor as though it were a measurement. Reintroduced the exact defect issue #636 removed one layer down.","fix":"Apply the bound unconditionally, built from captureInputTokens plus the new captureCompletionTokens, capped by the caller ceiling where set, floored at one credit, still capped by the hold. Six tests that asserted the flat hold updated to the priced bound.","tags":["money","billing","settlement","edge-api","overcharge","issue-1198","D-034"]} +{"id":"bug-2026-08-29-ci-object-storage-dead-secret","date":"2026-08-29","title":"CI object storage endpoint pointed at a deleted Supabase Cloud project, so the file upload path could not fail","error_message":"file upload to object storage failed bucket=hive-files error=\"Put \\\"https://.supabase.co/storage/v1/s3/hive-files/...\\\": dial tcp: lookup .supabase.co: no such host\"","root_cause":"The S3_ENDPOINT repository secret was last written 2026-04-21, four months before the August self-hosted Supabase cutover, and still named the deleted Supabase Cloud project. The live-integration job read it at job level, so every POST /v1/files in CI died at DNS and the Files and Batches conformance assertions were carried as it.fails markers rather than as coverage. The same stale secret family also killed the OWUI nightly at its seeding step via SUPABASE_URL (#1370), which then skipped the entire Playwright suite and left a misleading missing-browser cascade as the only visible red. Issue #1380 attributed the blindness to the http://127.0.0.1:9 stub in ci.yml, but that stub is in two Playwright jobs that run no SDK suite.","fix":"Removed every secrets.S3_* reference from live-integration and agent-visual-proof. live-integration now boots its own throwaway supabase/storage-api at the digest docker-compose.enterprise.yml pins, on the throwaway Postgres it already stands up, whose 00-extensions.sql already creates the storage schema and roles. The fixture asserts a signed PUT and GET before printing its values. The it.fails markers came off. Guards added: a unit test asserting the job keeps the fixture and no S3 secret, and a seam-test assertion that the fixture and compose pin the same image.","tags":["ci","storage","s3","supabase","dark-detector","stale-secret","sigv4"],"issues":[1324,1380,1370,1282]} +{"id":"bug-2026-08-29-s3-absolute-form-request-target","date":"2026-08-29","title":"S3 client sent absolute-form request targets, which Supabase Storage refuses with an unnamed 500","error_message":"s3 PUT /s3/hive-files/ failed with status 500: InternalErrorInternal Server Error; storage side: {\"code\":\"ERR_INVALID_URL\",\"input\":\"http://localhost:8080http://172.17.0.1:5000/s3/hive-files/\"} TypeError: Invalid URL at SignatureV4.constructCanonicalRequest","root_cause":"packages/storage/signing.go SignHTTP sets req.URL.Opaque to \"//host/path\" so the AWS SigV4 signer receives an un-normalized path, and never cleared it. net/url.RequestURI returns Opaque verbatim and re-attaches the scheme when it starts with \"//\", so every signed S3 request went on the wire in absolute form (PUT http://host/s3/bucket/key HTTP/1.1). RFC 9112 reserves absolute form for a request to a proxy. The deployed box never saw it because every S3 request passes through Caddy, which normalizes the target before proxying. Reached directly, Supabase Storage builds its canonical URI as new URL(\"http://localhost:8080\" + prefix + request.url) and throws on the doubled origin.","fix":"Clear req.URL.Opaque after signer.SignHTTP returns. The signature is unaffected: the signer derives its canonical URI from Opaque while set, stripping the leading //host, and that string is byte identical to the EscapedPath the wire form falls back to once Opaque is empty. presignHTTP still keeps Opaque, because url.URL.String renders a bare-path Opaque as http:/s3/... and the //host form is what makes a presigned URL absolute. Four tests added, each verified red first.","tags":["storage","s3","sigv4","http","rfc9112","proxy-masked","dark-detector"],"issues":[1324,1380,1282]} +{"id":"bug-2026-08-29-checkout-rails-missing-profile","date":"2026-08-29","title":"Buy credits dead on the deployed console: checkout rails 500 on a missing account_profiles row, and a rails payload the console cannot decode","error_message":"GET /api/v1/accounts/current/checkout/rails -> 500 {\"error\":\"checkout temporarily unavailable\"} (console proxy reported {\"error\":\"Upstream service error\"})","root_cause":"Four call sites (payments.Service.GetCheckoutOptions, payments.Service.InitiateCheckout, and both through stubCountryAdapter into the demo stub) read the whole account profile purely to obtain CountryCode for AvailableRails, and each treated profiles.ErrNotFound as a server fault. country_code is already optional (the repository COALESCEs NULL to \"\") and AvailableRails(\"\") is defined, so a missing account_profiles row was a question the code could already answer. Issue #999 counted 14 such accounts live. Independently, RailOption never carried currency, label or enabled, which the console's decodeCheckoutRail requires, so the rail list reached the checkout modal empty and no payment method rendered even without the 500.","fix":"Added profiles.Service.CountryCode, which maps ErrNotFound to (\"\", nil) and propagates every other error, and narrowed payments.ProfileReader to take the bare country so the payment path can no longer fail on an absent row. stubCountryAdapter delegates to it. Extended RailOption with Label, Currency and Enabled and CheckoutOptions with CreditIncrement, MinCredits and MaxCredits, all built through one shared NewRailOption constructor; max_credits is the minimum ceiling across the rails on offer so the advertised bound is never looser than ValidatePurchaseAmount enforces.","tags":["money-path","checkout","payments","control-plane","web-console","profiles","500","wire-contract"],"issues":[1386,999,797],"files":["apps/control-plane/internal/profiles/service.go","apps/control-plane/internal/payments/service.go","apps/control-plane/internal/payments/stub/stub.go","apps/control-plane/cmd/server/main.go"]} +{"id":"BUG-2026-08-29-console-tenant-claim","date":"2026-08-29","title":"Console admin pages refused every real user because control-plane read a tenant field nothing writes","error_message":"GET /api/v1/admin/feature-gates -> 400 {\"error\":\"no tenant selected\"}; console renders \"Could not load feature gates\"","root_cause":"auth.Client resolved Viewer.TenantID only from raw_user_meta_data.selected_tenant_id. The only application writer of that field is POST /v1/tenants/switch, which the web console never calls: its workspace switcher sets an hive_account_id cookie the control-plane does not read, and signup writes tenant_users without it. Every account created through normal signup therefore carried uuid.Nil forever and platform.WorkspaceAdminGate refused it with 400. Only accounts seeded by scripts, which set the field through the admin API, could reach /console/feature-gates or /console/marketplace, which is why it was never caught. Meanwhile public.custom_access_token_hook had been stamping a correct, membership-validated tenant_id claim into every issued token the whole time, and control-plane discarded it.","fix":"auth.Client.LookupUser now reads the tenant_id claim from the bearer Supabase has just validated, sub-matched to the user Supabase resolved, preferring it over the metadata and keeping the metadata as a fallback for pre-hook tokens. The existing membershipCheck still runs on whichever candidate wins, so the claim cannot widen access and an errored check still fails closed.","tags":["auth","control-plane","web-console","featuregate","marketplace","tenant-scope","issue-543","issue-762"]} +{"id":"bug-2026-08-29-classifier-misattributes-gateway-cooldown","date":"2026-08-29","title":"CI upstream-refusal classifier reported a LiteLLM router cooldown on a paid alias as a whole-pool provider rate limit","error_message":"UPSTREAM REFUSAL, not a performance regression: the provider rate limited us (429). Check whether the window is per minute (transient, so a rerun may pass) or per day (an allowance that is spent).","root_cause":"The 'Name the real cause when the upstream refused' step in .github/workflows/ci.yml grepped a 2000-line window of the litellm and edge-api container logs for the bare token 'status=429' among others. Three log shapes matched that are not upstream refusals: LiteLLM's own 'No deployments available for selected model ... cooldown_list=[...]' router cooldown (no upstream call is made), its 'Cooldown Deployments=[...]' table (which quotes the causing exception verbatim and reprints it on every dispatch while the entry lives), and a bare 'No fallback model group found' (which LiteLLM emits on any failed dispatch to a group with no configured fallback, including three refusals the job deliberately tests for). The matched line also carried alias deepseek-v4-flash, so a paid alias's gateway cooldown was reported as a refusal of the hive-free pool. The free pool was healthy throughout: the job's own probe step reported 3 healthy 1 unhealthy in both runs, only the Gemini member was rate limited, and the actual job failures were '500 Failed to store file' and a strict XPASS on issue #1274. Only two of the seven consecutive failures on issue #1374 that morning classified at all, which is the signature of an incidental grep hit rather than a spent allowance.","fix":"Moved the classification from a bash heredoc into scripts/classify-upstream-refusal.py, which make test-scripts runs as a required check. Gateway cooldown lines in both shapes are dropped before any branch reads them. The dead-member branch now requires the NotFoundError shape together with model_group=route-free-pool. A classified refusal names the alias it saw. Fixtures are copied from run 33243396287's own output. This also removes a SIGPIPE exit 2 from the old 'grep ... | head -5' reporting pipeline under set -o pipefail. Separately, report-free-pool-health.py now distinguishes a rate-limited member from a retired one, whose remedies are opposites, and surfaces the provider's quota window and retry hint from the full error text instead of truncating them away at 300 characters.","tags":["ci","litellm","free-pool","signal-quality","false-positive","issue-1088","issue-1064","issue-1374"]} +{"date":"2026-08-29","title":"lint:proof-tokens skipped itself on the only pull requests that add proof captures","error_message":"Repo policy lints (tenant + audit) reported pass in 6 seconds on PR #1398, which committed 71 proof captures. No step in the job executed.","root_cause":"The changes job in .github/workflows/ci.yml puts docs/ on the inert-path allow-list, so a pull request whose changed paths are all under docs/proof/ sets run=false. Every step in repo-policy-lints was gated on needs.changes.outputs.run != 'false', including npm run lint:proof-tokens, whose scanner reads docs/proof/ and nothing else. The gate therefore disabled itself for its own subject matter, and published a green that a reviewer trusts. Nothing else backstops the text half of a proof capture: GitHub secret scanning, push protection and GitGuardian do not read release assets and read no pixels.","fix":"Run npm run lint:proof-tokens unconditionally, along with the checkout, setup-node and npm ci steps it needs. Every other step in the job keeps its docs-only gate. Rejected carving docs/proof/** out of the classification (would run the whole heavy suite over a screenshot pull request) and a dedicated changes output (copies the scanner's scope into the workflow where it can drift, and cannot cover a capture that landed earlier). Demonstrated red before green on a throwaway docs-only pull request: run 33247080029 failed on a planted synthetic token, run 33247152597 passed once redacted.","tags":["ci","docs-only","proof-captures","credential-leak","unfalsifiable-green","issue-1409","issue-797"],"pr":1417} +{"id":"BUG-1407","date":"2026-08-29","title":"Chat and console 502s during every container recreate, misattributed to per-request Docker DNS resolution","error_message":"dial tcp: lookup open-webui on 127.0.0.11:53: server misbehaving; dial tcp 172.18.0.x:8080: connect: connection refused","root_cause":"Caddy has no dial retry, so any request arriving while an upstream container is being recreated fails. The two error strings are the same window seen at two moments: while the old container's name is deregistered Docker's embedded DNS forwards the single-label name to the host resolver, which returns SERVFAIL and Go reports as 'server misbehaving'; once the new container has registered but is not listening the dial is refused. All 58 measured 502s across both origins fell inside a deploy run window and none against a healthy upstream.","fix":"lb_try_duration 30s and lb_try_interval 1s on the four reverse_proxy blocks in Caddyfile.owui and both in Caddyfile.console, bounded so a genuinely broken upstream still 502s. Guarded by scripts/test_caddy_upstream_retry.py.","tags":["caddy","docker-dns","deploy","502","demo-box","issue-1407"]} +{"id":"bug-2026-08-29-console-qa-cluster","date":"2026-08-29","title":"Unbounded API key nickname made the console key table unusable and unrepairable from any product surface","error_message":"API keys table renders 45,524 pixels wide; Revoke control of every key in the workspace is unreachable","root_cause":"The nickname field had no length bound at any layer (form, Next route, Go handler, Go service, Postgres column) and the table's name cell had no width constraint, so one 5000-character value sized the whole table. No rename or update route existed, so nothing in the product could shorten a stored value. The same shape recurred in three siblings: two CSV exports each hand-rolled their own cell escaping and neither neutralised a formula-leading cell, the analytics id-to-nickname join existed on one surface only, and an analytics grid item's default min-width of auto let one card size its track past the viewport.","fix":"Bound the nickname at 100 runes in a shared decodeMintBody used by both the create and rotate routes, and truncate the name cell with the full value on title so an already-stored row is repaired on render. Move CSV escaping into lib/csv.ts with formula neutralisation that exempts plain numbers. Move the api-key label join into lib/analytics/api-key-labels.ts and use it on the group tabs as well as the overview tile. Zero the analytics grid items' min-width.","tags":["console","validation","csv-injection","responsive","input-bounds","issue-1400","issue-1401","issue-1403","issue-1406","issue-1408"]} +{"date":"2026-08-29","title":"Demo box build cache grew unbounded, so the deploy disk gate refused three consecutive deploys","error_message":"demo box has 14G free on /var/lib/docker, below the 15G floor. Refusing to proceed.","root_cause":"Nothing reclaimed build cache in a way that reached the box. The deploy job's only cache prune carried --filter until=24h, and on a day of roughly 43 merges every cache record is younger than 24 hours, so the filter excluded exactly the records consuming the disk. Build cache reached 37.68GB with 12.42GB reclaimable while images were only 5 percent reclaimable.","fix":"Reclaim unused build cache inside scripts/check-deploy-disk.sh, below the warn line only, bounded by timeout 300, before the thresholds are compared. Inside the guard rather than as a preceding step so a failed reclaim cannot skip the check. Both the pre-reclaim and post-reclaim readings are reported so a non-cache disk problem still fails loudly.","tags":["deploy","disk","docker","build-cache","ci","issue-1419"]} +{"date":"2026-08-29","title":"Chat upload had no size or type limit: a 30 MB file stalled with no error and an .exe was accepted","error_message":"POST /api/v1/files/?process=true never returns for a 28.6 MB attachment; no progress, no timeout, no error. evil.exe uploads with 200.","root_cause":"Open WebUI enforces rag.file.max_size and rag.file.allowed_extensions in upload_file_handler, but both were unset. RAG_FILE_MAX_SIZE was in docker-compose.yml with an empty default and a comment claiming it fed the client-side guard, while the key was missing from hive_rag_env_config.py's RAG_CONFIG_ENV, so the persisted first-boot row outranked the environment and no value could ever reach a booted box. RAG_ALLOWED_FILE_EXTENSIONS was absent entirely.","fix":"Reconcile rag.file.max_size and rag.file.allowed_extensions through hive_rag_env_config.py, with an int coercion for the size and a list coercion for the allowlist (a persisted comma string turns upstream's not-in membership test into a substring test). Compose defaults 25 MB, matching RAG_MAX_UPLOAD_BYTES 26214400, and an allowlist derived from known_source_ext plus the loader's own document branches. A malformed size or an empty allowlist fails startup rather than booting unenforced.","tags":["open-webui","uploads","persistent-config","silent-failure","issue-1405","issue-722"]} +{"id":"1399-composer-smart-typography","date":"2026-08-29","title":"Chat composer rewrote typed characters and the rewritten text reached the model","error_message":"Typing `git push --force` into the chat composer sent `git push —force`; straight quotes became curly, `--` an em dash, `...` an ellipsis glyph, and `->`, `!=`, `<<`, `>>`, `2 * 3`, `^2` were also rewritten","root_cause":"vendor/open-webui/src/lib/components/common/RichTextInput.svelte registered @tiptap/extension-typography bare in the richText extension array. Its twenty-two entries are ProseMirror input rules, which rewrite the text buffer as a character is typed rather than its presentation, so the mutated string is what the editor serializes and what leaves the browser. Code inside a formed code block was exempt because input rules do not run there, which made the defect look narrower than it was and sent an earlier hunt to the marked renderer and to CSS ligatures, neither of which is the layer. A same-day QA pass could not reproduce it on the deployed box, most likely because richText is a per-user setting and the extension is registered only inside that branch.","fix":"Removed the import and the registration. The extension contributes nothing but input rules, so dropping it removes all twenty-two; the suggested configure() with six rules disabled would have left sixteen live, including -> and !=. Guarded by vendor/open-webui/src/lib/hive/composer-literal-input.test.ts, which types a corpus through a real ProseMirror input-rule pipeline with no DOM, asserts the serialized string, pins the harness against the measured corruption so it can go red, and keys on the imported module rather than the identifier so an aliased re-add is caught. Proven with a before/after wire capture on one pinned backend image differing only in the mounted frontend build.","tags":["chat","open-webui","composer","tiptap","prosemirror","frontend","issue-1399"]} +{"id":"bug-2026-08-29-owui-prompt-templates-unreachable","date":"2026-08-29","title":"Open WebUI's ten task and RAG prompt templates were unreachable on any booted deployment","error_message":"No error. Setting TITLE_GENERATION_PROMPT_TEMPLATE, RAG_TEMPLATE or any of the other eight in compose changed nothing on the demo box, silently.","root_cause":"All ten are Open WebUI persistent config. Config.seed_defaults only inserts keys that are absent, so the first boot seeded every one of them (confirmed on the demo box: rag.template holding upstream's full default text and nine *.prompt_template rows holding an empty string, all written at first boot) and the database outranked the environment from then on. The keys were not in deploy/docker/owui-patches/hive_rag_env_config.py's reconcile allowlist, which exists for exactly this failure. Compounding it, the surface that would otherwise edit them, Open WebUI's admin panel, is deleted from the fork and 404'd at the proxy, and every write verb under /api/v1/configs is denied, so the only remaining path was a hand-written SQLite UPDATE inside the owui-data volume on a live box.","fix":"Added the ten keys to RAG_CONFIG_ENV under upstream's own variable names, with a TEMPLATE_KEYS frozenset so a prompt's leading indentation and trailing newline are persisted unstripped while an all-whitespace value still counts as unset. Added the compose passthroughs with empty defaults on every profile. Proved delivery end to end by booting the pinned image twice against one volume and capturing the request body Open WebUI sent to the model.","tags":["open-webui","persisted-config","first-boot-wins","system-prompts","722","772","1405"]} +{"date":"2026-08-29","issue":491,"pr":null,"error_message":"Dark mode danger/warning/success text fails WCAG AA (danger badge 2.65:1)","root_cause":"globals.css dark media block redefined only the -soft semantic tokens; the three solid tokens (--color-success, --color-warning, --color-danger) fell through the cascade to their light-mode values, which are not readable against dark backgrounds","fix":"Added dark oklch(0.72 ...) overrides for the three solid tokens in globals.css, matching --color-accent's existing dark lightness; flipped button.tsx's danger variant label from hardcoded text-white to text-[var(--color-canvas)] since the lightened dark danger token drops white-on-danger to 2.72:1","tags":["web-console","theme","accessibility","wcag","css"]} +{"id":"BUG-1434-grafana-tiles-inert","date":"2026-08-29","title":"Analytics Grafana tiles gated on a GRAFANA_BASE_URL set nowhere, and wiring it would have leaked cross-tenant key ids","error_message":"Both Grafana tiles on /console/analytics render 'Not available on this deployment.' for every account; GRAFANA_BASE_URL is absent from every compose file, workflow, .env.example and from the running web-console container environment.","root_cause":"The tiles shipped with a runtime gate on an operator-provided variable that no deployment ever set, so the feature was inert from the first commit. The deeper defect is that completing the wiring was unsafe: Grafana runs with GF_AUTH_ANONYMOUS_ENABLED=true at the Viewer role bound on every interface, is fronted by no Caddyfile, and its rate-limit dashboard queries topk(10, sum by (key_id, tier) (...)) while the analytics page that renders the tiles carries no role guard, so every tenant member would have received an unauthenticated link naming other tenants' API key identifiers. Confirmed by capture: the pre-change component with GRAFANA_BASE_URL set renders two live Grafana links to a non-admin account.","fix":"Removed the two Grafana tiles and the nullable-href/disabled-state machinery that existed only to describe their absence, keeping the working first-party Request logs tile. Added a colocated regression test asserting the exact rendered link set, name-agnostic so a re-wiring under any variable name fails it, verified by mutation to go red on the pre-change component. Filed issue #1442 for hardening the Grafana instance itself, which this change deliberately does not touch.","tags":["console","analytics","grafana","observability","multi-tenant","information-disclosure","inert-code","dead-config"]} +{"id":"BUG-1428-upload-cap-divergence","date":"2026-08-29","title":"Chat upload cap and RAG ingest ceiling were two independently settable numbers, so the #1426 fix deployed inert","error_message":"Chat composer published a 100 MB attachment cap while edge-api and the markitdown sidecar enforced 26214400 bytes; a 30 MB attachment was accepted and stored by Open WebUI after a silent multi-minute upload","root_cause":"docker-compose.yml gave the open-webui service its own RAG_FILE_MAX_SIZE variable, in whole megabytes, settable independently of the RAG_MAX_UPLOAD_BYTES expression it passes edge-api and the sidecar in bytes. PR #1426 set that variable's compose default to 25, but the demo box's .env carried an explicit RAG_FILE_MAX_SIZE=100 and an explicit value beats a compose fallback, so the fix never reached the deployment it was written for.","fix":"Deleted RAG_FILE_MAX_SIZE as a settable knob. The open-webui service now takes the same ${RAG_MAX_UPLOAD_BYTES:-26214400} expression as the other two services, and hive_rag_env_config.derived_upload_cap floors it into whole megabytes, rounding down so the chat surface can never accept what the ingest path refuses. The container refuses to start if RAG_FILE_MAX_SIZE is present, and edge-api now fails the boot on a malformed ceiling instead of warning and falling back. A compose guard in scripts/test_owui_rag_env_config.py fails if a second knob is re-introduced.","tags":["config","docker-compose","open-webui","rag","uploads","silent-no-op","deploy-drift"]} +{"date":"2026-08-29","error_message":"OpenRouter 402 Payment Required (own account out of funds) forwarded verbatim as HTTP 402 to the Hive customer","root_cause":"WriteProviderBlindUpstreamError only sanitized the MESSAGE body for an upstream refusal; it never remapped the HTTP STATUS itself, so a provider funding refusal (about Hive's own OpenRouter balance) was indistinguishable, at the status-code level, from a caller quota refusal (about the customer's own Hive balance)","fix":"remap upstream 402 to 503/upstream_unavailable in WriteProviderBlindUpstreamError, reusing the existing 503/504 temporarily-unavailable message path; original status kept in the operator log","tags":["billing","provider-blind","accounting","reservations","issue-1411"]} From fbf70a99084f6fc2a4780f4db923f5d28942c4b7 Mon Sep 17 00:00:00 2001 From: Sakib Sadman Shajib Date: Sat, 29 Aug 2026 14:13:04 -0400 Subject: [PATCH 2/2] chore: fold in the three entries from the two late 2026-08-29 merges #1432 and #1448 merged after the first enumeration pass and both carried an entry. #1432 contributes two well formed JSON entries. #1448 wrote its entry as prose key and value lines rather than a JSON object, so it is reformatted here into one line with its content preserved verbatim and nothing added. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01WyEwUxZCArdn1ZUDkTvuQ1 --- .wolf/buglog.jsonl | 3 +++ 1 file changed, 3 insertions(+) diff --git a/.wolf/buglog.jsonl b/.wolf/buglog.jsonl index 17495753e..6a5b66d84 100644 --- a/.wolf/buglog.jsonl +++ b/.wolf/buglog.jsonl @@ -280,3 +280,6 @@ {"id":"BUG-1434-grafana-tiles-inert","date":"2026-08-29","title":"Analytics Grafana tiles gated on a GRAFANA_BASE_URL set nowhere, and wiring it would have leaked cross-tenant key ids","error_message":"Both Grafana tiles on /console/analytics render 'Not available on this deployment.' for every account; GRAFANA_BASE_URL is absent from every compose file, workflow, .env.example and from the running web-console container environment.","root_cause":"The tiles shipped with a runtime gate on an operator-provided variable that no deployment ever set, so the feature was inert from the first commit. The deeper defect is that completing the wiring was unsafe: Grafana runs with GF_AUTH_ANONYMOUS_ENABLED=true at the Viewer role bound on every interface, is fronted by no Caddyfile, and its rate-limit dashboard queries topk(10, sum by (key_id, tier) (...)) while the analytics page that renders the tiles carries no role guard, so every tenant member would have received an unauthenticated link naming other tenants' API key identifiers. Confirmed by capture: the pre-change component with GRAFANA_BASE_URL set renders two live Grafana links to a non-admin account.","fix":"Removed the two Grafana tiles and the nullable-href/disabled-state machinery that existed only to describe their absence, keeping the working first-party Request logs tile. Added a colocated regression test asserting the exact rendered link set, name-agnostic so a re-wiring under any variable name fails it, verified by mutation to go red on the pre-change component. Filed issue #1442 for hardening the Grafana instance itself, which this change deliberately does not touch.","tags":["console","analytics","grafana","observability","multi-tenant","information-disclosure","inert-code","dead-config"]} {"id":"BUG-1428-upload-cap-divergence","date":"2026-08-29","title":"Chat upload cap and RAG ingest ceiling were two independently settable numbers, so the #1426 fix deployed inert","error_message":"Chat composer published a 100 MB attachment cap while edge-api and the markitdown sidecar enforced 26214400 bytes; a 30 MB attachment was accepted and stored by Open WebUI after a silent multi-minute upload","root_cause":"docker-compose.yml gave the open-webui service its own RAG_FILE_MAX_SIZE variable, in whole megabytes, settable independently of the RAG_MAX_UPLOAD_BYTES expression it passes edge-api and the sidecar in bytes. PR #1426 set that variable's compose default to 25, but the demo box's .env carried an explicit RAG_FILE_MAX_SIZE=100 and an explicit value beats a compose fallback, so the fix never reached the deployment it was written for.","fix":"Deleted RAG_FILE_MAX_SIZE as a settable knob. The open-webui service now takes the same ${RAG_MAX_UPLOAD_BYTES:-26214400} expression as the other two services, and hive_rag_env_config.derived_upload_cap floors it into whole megabytes, rounding down so the chat surface can never accept what the ingest path refuses. The container refuses to start if RAG_FILE_MAX_SIZE is present, and edge-api now fails the boot on a malformed ceiling instead of warning and falling back. A compose guard in scripts/test_owui_rag_env_config.py fails if a second knob is re-introduced.","tags":["config","docker-compose","open-webui","rag","uploads","silent-no-op","deploy-drift"]} {"date":"2026-08-29","error_message":"OpenRouter 402 Payment Required (own account out of funds) forwarded verbatim as HTTP 402 to the Hive customer","root_cause":"WriteProviderBlindUpstreamError only sanitized the MESSAGE body for an upstream refusal; it never remapped the HTTP STATUS itself, so a provider funding refusal (about Hive's own OpenRouter balance) was indistinguishable, at the status-code level, from a caller quota refusal (about the customer's own Hive balance)","fix":"remap upstream 402 to 503/upstream_unavailable in WriteProviderBlindUpstreamError, reusing the existing 503/504 temporarily-unavailable message path; original status kept in the operator log","tags":["billing","provider-blind","accounting","reservations","issue-1411"]} +{"id":"bug-2026-08-29-consent-landing-client-roundtrip","date":"2026-08-29","title":"Cold sign-in spent 1984 ms rendering a consent panel only for the browser to discover there was no session","error_message":"GET /oauth/consent returned 200 with a rendered panel for a session-less visitor; the browser booted the panel, read the session, found none, and navigated to /auth/sign-in 1984 ms later","root_cause":"decideConsentLanding returned render-panel for !hasSession, deferring to the client panel a decision the server component had already made. The server reads the same @supabase/ssr cookie session the browser does, so the deferral bought no information and cost a full render plus bundle boot on the single hop every first-time sign-in passes through.","fix":"decideConsentLanding answers !hasSession with the sign-in redirect the page already knows how to perform. The missing-authorization_id case and the retried case still fall to the panel, each pinned by a test.","tags":["web-console","auth","oauth","performance","issue-967","issue-945"]} +{"id":"bug-2026-08-29-profile-gated-container-survives-deploy","date":"2026-08-29","title":"A next dev console kept running for four weeks after its compose service was gated behind a profile","error_message":"hive-web-console-1 up since 2026-07-31, restart unless-stopped, publishing 0.0.0.0:3000, serving /auth/sign-in in 5.56 s from a development build, across roughly thirty deploys","root_cause":"docker compose up --remove-orphans removes containers for services absent from the compose file, not containers for services present but outside the enabled profile set. PR #605's dev profile gate therefore stopped the service being started and never stopped it running, and no other deploy step evicted it.","fix":"deploy-demo-box.yml removes any running container in the compose project whose service is absent from `docker compose config --services` under the deploy's own flags, failing closed when that listing is empty. The already-running container was removed by hand.","tags":["deploy","docker-compose","security","issue-967"]} +{"id":"bug-console-silent-save-2026-08-29","date":"2026-08-29","title":"Profile settings save silently no-op'd when an unrelated required field was empty, and never confirmed success even when it worked.","error_message":"Editing Owner name on /console/settings/profile and clicking Save produced no visible success or error message and reverted to the original value on reload.","root_cause":"AccountProfileForm has five required inputs across three cards. The form had no noValidate attribute, so clicking Save while the pre-existing Country/State fields were empty (the default state for any newly-provisioned account, since /console/setup uses the same form) triggered native HTML5 constraint validation, which blocked the submit event before React's Server Action ever ran. Confirmed live with a full HAR capture against the deployed box: zero non-GET requests were sent for the save attempt. Separately, even a successful save rendered no confirmation at all, since redirect() inside the action just reloaded the page.","fix":"Added noValidate to the form so submission always reaches the Server Action, which already returns per-field errors the form already renders. The action now redirects to saved=1 on success, and the form renders a status confirmation, mirroring the existing /console/members?invited=1 pattern (issue #535). /console/setup (the form's other caller) was wired with the same prop.","verification":"New account-profile-form.test.tsx, 4 tests. Mutation tested: reverting the noValidate addition alone turns 2 of 4 tests red (the reproduction test and the per-field-error test), confirming the test suite actually catches this regression. npm run build passes clean.","tags":["console","web-console","forms","silent-failure","server-actions","ux"],"pr":1448}