Repository navigation
fix: hold requests across a container recreate instead of 502ing (#1407) - #1421
Conversation
Re-measuring the demo box shows the count in #1407 is right and the attribution is not. All 55 chat 502s and all 3 console 502s in the 24 hours to 2026-08-29T10:00Z fell inside a deploy-demo-box.yml run window, in nine bursts of 4 to 40 seconds. None happened against a healthy upstream. The two error strings are one defect seen at two moments of a container recreate. While the old container is gone and its name is deregistered, Docker's embedded DNS no longer holds the record and forwards the query to the host resolver in the container's resolv.conf, which refuses a single-label name with SERVFAIL; Go reports that as "server misbehaving". Confirmed directly: nslookup for an unknown container name against 127.0.0.11 returns SERVFAIL, not NXDOMAIN. Once the new container has registered but uvicorn is not listening yet, the dial is refused instead. So the fix suggested in the issue, an explicit resolvers 127.0.0.11, would point Caddy at the very resolver returning the SERVFAIL. Per-request resolution is not the defect either: resolving at dial time is what lets the proxy follow the new container's IP at all. What is missing is patience across the window. lb_try_duration 30s with lb_try_interval 1s is added to the four reverse_proxy blocks in Caddyfile.owui and both in Caddyfile.console. Caddy retries only while the upstream connection was never established, so a request already talking to an upstream is untouched and nothing already answered can be replayed. A genuinely broken upstream still returns a real 502 with the real error logged, 30 seconds later rather than immediately: a bounded delay, not a masked success. 30s stays under Cloudflare's 100s origin timeout so a held request never becomes a 524, and 1s rather than the 250ms default keeps a held request from issuing 120 lookups against a resolver that is already failing. Caddyfile.supabase and Caddyfile.artifacts measured zero 502s over the same window and are deliberately left alone. Caddyfile.agent-proof is a scratch harness, not a product surface. scripts/test_caddy_upstream_retry.py guards the recurrence rather than the fix. Caddyfile.owui has grown three reverse_proxy routes since it was written, each copied from the block above it, and a fourth added the same way from an older block would silently ship without the retry. The guard carries its own self-check so it is provably able to fail. This does not remove the outage. Every deploy that recreates open-webui still takes chat down for seconds. It makes that window invisible to most requests. A health-gated cutover is a separate and much larger change. Buglog entry: {"id":"BUG-1407","date":"2026-08-29","title":"Chat and console 502s during every container recreate, misattributed to per-request Docker DNS resolution","error_message":"dial tcp: lookup open-webui on 127.0.0.11:53: server misbehaving; dial tcp 172.18.0.x:8080: connect: connection refused","root_cause":"Caddy has no dial retry, so any request arriving while an upstream container is being recreated fails. The two error strings are the same window seen at two moments: while the old container's name is deregistered Docker's embedded DNS forwards the single-label name to the host resolver, which returns SERVFAIL and Go reports as 'server misbehaving'; once the new container has registered but is not listening the dial is refused. All 58 measured 502s across both origins fell inside a deploy run window and none against a healthy upstream.","fix":"lb_try_duration 30s and lb_try_interval 1s on the four reverse_proxy blocks in Caddyfile.owui and both in Caddyfile.console, bounded so a genuinely broken upstream still 502s. Guarded by scripts/test_caddy_upstream_retry.py.","tags":["caddy","docker-dns","deploy","502","demo-box","issue-1407"]} Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WyEwUxZCArdn1ZUDkTvuQ1
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
|
Warning Review limit reachedNext included review available in 19 minutes. View limit detailsLimit details: You’ve used the included review currently available. You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository. Review configuration: ⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Pro Plus Run ID: ⛔ Files ignored due to path filters (1)
📒 Files selected for processing (7)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Visual proofBrowser through a Caddy running the repo Caddyfile.owui, against an upstream name that does not resolve, which is exactly what a container recreate produces. First image is the file as it stands on main: HTTP 502 in 0.11s, the error page a visitor gets mid-deploy. Second is the same file with lb_try_duration 30s and lb_try_interval 1s, same unresolvable upstream at request time, the container appearing mid-hold: HTTP 200 in 4.17s. The upstream behind the second shot is a deliberately labelled throwaway file-server standing in for open-webui, not the Hive product, and its own page body says so. Full transcript, including the case where the upstream never returns and the request still gets a real 502 at the 30s bound, is at docs/proof/chat-502-recreate-retry-2026-08-29/capture.log. |
Three defects found on a self-review of the guard added in the previous commit, all of them ways it could report a pass it should not. Brace depth was counted on comment lines, so an unbalanced brace inside a comment within a reverse_proxy block would close the block early and hide every directive below it. Comment lines are now skipped before the depth arithmetic rather than only when reading directives. The reported line number was the block's closing brace rather than its opening one, which sent a reader to the wrong place in a file where these blocks are separated by long comments. The directive was matched with a prefix test, so a hypothetical reverse_proxy_something would have been treated as a proxy block. It is now matched as a whole token. The self-check grew from two synthetic blocks to three, covering a nested transport block, an environment placeholder, an unbalanced brace in a comment, and a guarded block after an unguarded one, so brace depth returning correctly to zero is now something the check actually proves. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WyEwUxZCArdn1ZUDkTvuQ1
Adversarial review found three real gaps in the guard's parser, all of them ways it could pass a file it should reject or reject one it should pass. An inline comment on a reverse_proxy's opening line made the block look like a bare one-liner, so the comment's last word was read as the upstream and the block body was parsed as top-level directives. Comments are now stripped from every line before it is examined, using Caddy's own rule that a hash opens a comment only where it begins a token. Durations were matched with a single-unit regex, so a perfectly valid lb_try_duration of 1m30s was reported as unparseable. Compound Go durations now parse, and the parser has its own table of cases including the shapes that must be rejected. The self-check exercised only the missing-duration branch, because the check short-circuits there, so the missing-interval branch was never entered by any test. The synthetic Caddyfile now carries a fourth block that has a duration and no interval, and the self-check asserts both failures by upstream and by cause. The same review claimed the directive forces Caddy to buffer every request body and can replay a non-idempotent POST. Both were measured against the real proxy rather than argued, and both are wrong: an A/B with a chunked body that pauses mid-stream shows the upstream receiving the request immediately in both variants, six seconds before the client finishes sending, with a request count of exactly one. Caddy retries after a successful connection only for requests matching retry_match, which defaults to GET and is not set here. The transcript is in section 8 of the proof log, along with the file-descriptor ceiling and the per-origin container isolation that bound the pile-up concern. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WyEwUxZCArdn1ZUDkTvuQ1
sakibsadmanshajib
left a comment
There was a problem hiding this comment.
Adversarial review, stream results
Two streams were run against this diff.
CodeRabbit: SKIPPED, not a pass. Both the CLI and the GitHub app returned Rate limit exceeded ("You've used all 3 included reviews currently available"), retried three times over eight minutes. The PR check named CodeRabbit reports pass with the detail Review rate limited, which is a check that cannot go red on this account state and must not be read as approval. No CodeRabbit finding exists for this diff, positive or negative.
Antigravity (gemini-3.1-pro-high, effort high): RAN. Four findings, posted inline below. Three of them are real and are fixed in 7575456; two are wrong on a mechanism that I measured rather than argued, and the measurements are in section 8 of the proof log.
Repo prose convention uses commas, colons or separate sentences between clauses rather than a dash. Three comment lines introduced by this branch used one. No behaviour change. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WyEwUxZCArdn1ZUDkTvuQ1
The proxy_blocks docstring still said only whole-line comments were stripped, which stopped being true when inline comment handling went in, and the synthetic Caddyfile's header still described three blocks after a fourth was added for the interval branch. Both now say what the code does, and the block list explains what each of the four is there to catch. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WyEwUxZCArdn1ZUDkTvuQ1
|
Baseline refresh, so the post-merge comparison has a current number to beat. Re-counted on the box while this PR was in review, about forty minutes after the original measurement: Up from 55 in the same rolling 24 hour window, with nine of those in the last hour alone, which is the deploy cadence continuing to do what the body of this PR describes. The defect is live and accumulating, not historical. After this merges and has been deployed for a day, re-run the same two commands. The 24 hour count should fall to the residue of recreate windows longer than the 30s bound, not to zero, since this change shortens the window rather than removing it. |
## Summary This is the batched buglog follow-up for the pull requests merged to `main` on 2026-08-29. Its diff is `.wolf/buglog.jsonl` and nothing else. Per `.claude/rules/openwolf.md`, every fixed bug, error, failed test or failed build must be logged, but the line may never be appended on a fix branch. `merge=union` in `.gitattributes` resolves concurrent appends locally and is ignored by GitHub's server side merge, so two branches that both appended land in hard conflict there. An unmergeable pull request gets no `refs/pull/N/merge`, no `pull_request` run and therefore zero checks, and the required status gate then blocks the merge for a reason the page never states (issue #873). Each fix accordingly carried its entry in its own pull request body, and this pull request copies them onto `main` in one batch, which the protocol explicitly prefers over one pull request per entry. ## Scope examined Fifty nine pull requests merged to `main` on 2026-08-29. Forty eight of them carried at least one entry, for eighty two entries in total. Thirty two of those were already on `main` and are skipped, leaving fifty appended here from thirty four pull requests. The largest block of skips comes from #1342, the equivalent batch for the 2026-08-28 merges, which merged earlier the same day and already landed thirty six entries covering #1257, #1268, #1276, #1277, #1287, #1292, #1293, #1294, #1296, #1301, #1303, #1305, #1313, #1335 and #1337. ## What landed Fifty entries appended, one JSON object per line, append only. The 232 pre-existing lines are byte identical to `origin/main` (verified by hashing the first 232 lines of the result against the base file). Every line in the resulting file parses as JSON and carries `error_message`, `root_cause`, `fix` and `tags`. | Source | Entries | |---|---| | #1083 | 2 | | #1277 | 1 | | #1278 | 1 | | #1298 | 1 | | #1334 | 1 | | #1336 | 3 | | #1343 | 1 | | #1346 | 1 | | #1351 | 1 | | #1365 | 2 | | #1368 | 1 | | #1369 | 1 | | #1371 | 3 | | #1375 | 3 | | #1376 | 1 | | #1378 | 1 | | #1379 | 2 | | #1388 | 5 | | #1389 | 3 | | #1390 | 2 | | #1393 | 1 | | #1394 | 1 | | #1410 | 1 | | #1417 | 1 | | #1421 | 1 | | #1423 | 1 | | #1424 | 1 | | #1426 | 1 | | #1429 | 1 | | #1431 | 1 | | #1433 | 1 | | #1434 | 1 | | #1436 | 1 | | #1439 | 1 | Entries are copied verbatim from their source pull request bodies. Nothing was rewritten, no field was invented, and no field was added. No JSON needed repair: all eighty two extracted entries parsed on the first attempt and all four required fields were present on every one. ## Merged pull requests that carried no entry Eleven of the fifty nine. Recorded here because the gap is itself the useful signal. | Pull request | Title | Assessment | |---|---|---| | #1013 | chore(deps): bump the go-minor-patch group across 1 directory with 4 updates | Dependabot bump, no defect fixed, no entry expected | | #1015 | chore(deps): bump the go-minor-patch group across 1 directory with 6 updates | Dependabot bump, no entry expected | | #1016 | chore(deps): bump golang from 1.26-alpine to 1.27-alpine in /deploy/docker | Dependabot bump, no entry expected | | #1218 | chore(deps): bump postcss from 8.5.19 to 8.5.26 in /apps/desktop | Dependabot bump, no entry expected | | #1219 | chore(deps): bump golang.org/x/crypto from 0.41.0 to 0.52.0 in /apps/control-plane | Dependabot bump, no entry expected | | #1342 | chore: batch buglog entries for the 2026-08-28 merges | The previous batch pull request itself, correctly carries no entry of its own | | #1364 | chore: remove four dead skills and record the patterns that cost time | Protocol gap. The body records patterns that cost time, which is the shape of a buglog entry, but none was written as one | | #1383 | test: retire stale expected-failure markers, restore the ones that are true (#1381, #1382, #1324) | Protocol gap. Stale `it.fails` markers reading as red is a real defect that was fixed here and should have carried an entry | | #1384 | docs: correct D-047, hive-auto reverted to variable pricing (D-059) | Decision ledger correction, arguably a documentation defect, no entry written | | #1387 | chore(deps): bump next from 15.5.23 to 16.3.3 in /apps/agent-console | Dependabot bump, no entry expected | | #1398 | docs: rescue the 2026-08-25 parity captures and add the 2026-08-29 QA matrix evidence | Documentation and evidence rescue, no entry written | Six of the eleven are Dependabot bumps and one is the previous batch, so the genuine protocol gaps are #1364, #1383, #1384 and #1398. Of those, #1383 is the one worth a follow-up: it fixed a real defect class (a stale expected-failure marker reads as a red "Expect test to fail" and gets dismissed as pre-existing) and left no record. ## Entries skipped as already present Thirty two. Thirty of them matched an entry already on `main` on `error_message`, `id` or `fix`. Two more from #1278 are semantic duplicates that an exact match would have missed, and were skipped after reading the landed entries they duplicate: - #1278's `streaming content_block_start omits text field` entry is covered by the consolidated `bug-2026-08-28-anthropic-sdk-wire-conformance` entry landed from #1296, whose root cause names the same `omitempty` on `StreamContentBlock.Text`. - #1278's `GET /v1/models leaked an upstream provider name` entry is covered by `BUG-1284`, landed from #1300, which names the same `public.model_aliases.summary` publication path. #1278's third entry, on `top_k` forwarding producing a 400, is not covered anywhere on `main` and is appended here. #1342 recorded #1278 as fully "merged into #1296", which was accurate for two of its three entries. ## Note on entry quality One appended entry is thin: #1277's parity re-score record carries `error_message` of `n/a` and a root cause of "console had no privacy/data-policy surface at all". It is a parity gap record rather than a defect record. It is included exactly as written rather than embellished, per the protocol's preference for the author's own words. ## Test plan - [x] Branch cut fresh from `origin/main`, diff is `.wolf/buglog.jsonl` and nothing else - [x] First 232 lines byte identical to the base file (md5 match) - [x] All 282 resulting lines parse as JSON and carry `error_message`, `root_cause`, `fix` and `tags` - [x] No `.wolf/` telemetry (`anatomy.md`, `memory.md`, `token-ledger.json`, `hooks/_session.json`, `buglog.json`) in the commit - [ ] The six required checks report green via the inert path allowlist in `.github/workflows/ci.yml` --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>


Fixes #1407.
The issue's count is right, its attribution is not
I re-measured on the box before building anything. Both halves of the issue's premise needed checking, and one of them fails.
The counts, over the 24 hours to 2026-08-29T10:00Z:
server misbehavingconnection refusedhive-caddy-owui-1(chat)hive-caddy-console-1hive-caddy-supabase-1hive-caddy-artifacts-1The issue says the eleven (now seventeen) DNS failures "happen against a healthy, running upstream". They do not. Timestamped to the second, all 55 chat 502s fall into nine bursts, and every burst is inside a
deploy-demo-box.ymlrun window:Zero 502s outside a deploy window. The chat container running at capture time brackets the last burst exactly:
created=2026-08-29T09:52:55Z started=2026-08-29T09:53:10Z.Mechanism
The two error strings are one defect seen at two moments of the same container recreate.
docker exec hive-caddy-owui-1 nslookup -type=a nosuchcontainer 127.0.0.11returns SERVFAIL, not NXDOMAIN. The container'sresolv.confrecordsExtServers: [host(127.0.0.53)], so Docker's embedded DNS forwards a name it does not hold to the host's systemd-resolved, which refuses a single-label name. Go's resolver maps SERVFAIL to "server misbehaving".So:
Why not the fix the issue suggests
An explicit
resolvers 127.0.0.11points Caddy at the very resolver returning the SERVFAIL. During the gap no resolver holds the record, because the container genuinely does not exist.Per-request resolution is not the defect either. Resolving at dial time is what lets this proxy follow the new container's IP across a recreate at all; a proxy that resolved once at startup would hold a dead IP after every deploy and fail permanently, which is strictly worse.
The change
lb_try_duration 30sandlb_try_interval 1son the fourreverse_proxyblocks inCaddyfile.owui(open-webui, agent-console, and the two edge-api routes) and both inCaddyfile.console.Caddyfile.supabaseandCaddyfile.artifactsmeasured zero 502s over the same window, so adding it there would be speculative.Caddyfile.agent-proofis a scratch proof harness, not a product surface.What it does to an in-flight request
Nothing. Caddy retries only while the connection to the upstream was never established, so a request already talking to an upstream is untouched and nothing already answered can be replayed. Both failure classes here are dial failures.
What it does to a request arriving during a recreate
It is held, retried once a second, and served as soon as the new container listens.
What it does to a genuinely broken upstream
It still returns a real 502, with the real error still logged verbatim, 30 seconds later rather than immediately. That is a bounded delay, not a masked success, and it is the point of the bound rather than an unbounded retry. The cost is real and accepted: while an upstream is crash-looping, every request waits the full window before erroring.
Why 30s and 1s specifically
30s covers the observed central mass of recreate windows (4s, 8s, 12s, 12s, 29s) whole and shortens the two long ones, while staying well under Cloudflare's 100s origin timeout, so a held request never turns into a 524 instead of the 502 it replaces. 1s rather than the 250ms default because at 250ms a single held request issues 120 lookups across the window against the same resolver that is already failing; recovery is detected at most one second later.
Proof
Red first, on the box, against a throwaway user-defined network holding no containers, so
open-webuiis exactly the unresolvable single-label name a recreate produces. The repo's own Caddyfile is mounted unchanged into the digest-pinnedcaddy:2-alpine.hive-open-webui-1upstream, healthycaddy validatereportsValid configurationfor both patched files.The full transcript, including the raw commands, is committed at
docs/proof/chat-502-recreate-retry-2026-08-29/capture.log, with the three harness scripts beside it so the A/B is re-runnable. Screenshots posted below as a separate comment.What is not proven here: the after-rate on the real box cannot be measured before this merges, since the Caddyfile only takes effect on the next deploy. The before-rate is the table at the top. Re-count with the section 1 commands over a comparable window once this has been deployed for a day.
What this does not fix: every deploy that recreates open-webui still takes chat down for 4 to 40 seconds, and there were more than forty deploys in the measured day. This makes that window invisible to most requests. A health-gated cutover is what removes it, and it is a separate and much larger change to
deploy-demo-box.ymland the compose model.Test
scripts/test_caddy_upstream_retry.py, wired intomake test-scripts, guards the recurrence rather than the fix.Caddyfile.owuihas grown threereverse_proxyroutes since it was written (agent console, agent API, featuregate probe), each copied from the block above it; a fourth added the same way from an older block would silently ship without the retry and serve 502s again on the next deploy. The guard also rejects a duration outside 5s to 60s, so it cannot be quietly defanged to zero or widened past the Cloudflare timeout.It carries its own self-check, so it is provably able to go red rather than only green. Against the pre-fix files it reports all six unguarded blocks:
It also rejects a duration outside 5s to 60s, so the guard cannot be quietly
defanged to zero or widened past the Cloudflare timeout, and it parses compound
Go durations so a legitimate
1m30sis not a false red.Buglog entry
To be appended to
.wolf/buglog.jsonlonmainin a separate buglog-only pull request after this merges, per.claude/rules/openwolf.md.{"id":"BUG-1407","date":"2026-08-29","title":"Chat and console 502s during every container recreate, misattributed to per-request Docker DNS resolution","error_message":"dial tcp: lookup open-webui on 127.0.0.11:53: server misbehaving; dial tcp 172.18.0.x:8080: connect: connection refused","root_cause":"Caddy has no dial retry, so any request arriving while an upstream container is being recreated fails. The two error strings are the same window seen at two moments: while the old container's name is deregistered Docker's embedded DNS forwards the single-label name to the host resolver, which returns SERVFAIL and Go reports as 'server misbehaving'; once the new container has registered but is not listening the dial is refused. All 58 measured 502s across both origins fell inside a deploy run window and none against a healthy upstream.","fix":"lb_try_duration 30s and lb_try_interval 1s on the four reverse_proxy blocks in Caddyfile.owui and both in Caddyfile.console, bounded so a genuinely broken upstream still 502s. Guarded by scripts/test_caddy_upstream_retry.py.","tags":["caddy","docker-dns","deploy","502","demo-box","issue-1407"]}🤖 Generated with Claude Code
https://claude.ai/code/session_01WyEwUxZCArdn1ZUDkTvuQ1