release #5541: iOS PWA streaming freeze + SSE reconnect reliability - #5586
Conversation
Two mobile reliability fixes on the crown-jewel chat streaming path: (1) content-visibility:auto on off-screen .msg-row under @media(pointer:coarse) so WKWebView skips layout/paint for off-screen rows during streaming (kills the long-chat freeze). The LIVE turn is kept content-visibility:visible via the STABLE #liveAssistantTurn id (covers ALL render modes — Compact Worklog, Transparent Stream, restored-live — not just the Transparent-Stream path that stamps data-live-assistant-turn), so a normal live turn on touch never blanks mid-stream and its height keeps growing so the new-message cue still fires. contain-intrinsic-size:auto 1px preserves flick-scroll momentum. Scoped to touch — desktop find-in-page untouched. Deliberately does NOT set overflow-anchor:none (inert on iOS WebKit; re-opens #4856/#5338). (2) SSE reconnect ladder extended 4->6 steps + a last-ditch _restoreSettledSession full-session poll (8s watchdog) after retries exhaust, so a response completed during an iOS Tailscale/VPN reconnect is recovered without an error banner. Co-authored-by: luperrypf <luperrypf@users.noreply.github.com>
|
| Filename | Overview |
|---|---|
| static/messages.js | Extends SSE retry ladder from 4 to 6 steps and adds a last-ditch _restoreSettledSession with an 8s watchdog; the watchdog covers only the second call, leaving the pre-existing probe-loop call unguarded |
| static/style.css | Adds content-visibility:auto + contain-intrinsic-size:auto 1px + contain:layout style to .msg-row under @media (pointer:coarse), with a correctly-targeted #liveAssistantTurn ID selector to keep the live turn fully rendered on all paths |
| tests/test_sse_error_multi_probe_reconnect.py | Correctly updated static assertions to reflect the expanded 6-step delay array and new last-ditch reconnect structure; brace-matching approach is robust against future growth |
| CHANGELOG.md | Release-managed CHANGELOG entry added for both the streaming-freeze fix and the SSE reconnect improvements; no issues |
Sequence Diagram
%%{init: {'theme': 'neutral'}}%%
sequenceDiagram
participant ES as EventSource
participant PR as _probeReconnect
participant SS as /api/stream/status
participant RS as _restoreSettledSession
participant UI as Composer UI
ES->>PR: onerror fires
PR->>UI: "Reconnecting… (1/6)"
loop For each of 6 retry delays
PR->>SS: GET /api/chat/stream/status
alt stream active
SS-->>PR: "{active:true}"
PR->>ES: reconnect EventSource
else replay available
SS-->>PR: "{replay_available:true}"
PR->>ES: reconnect with replay
else still not settled
SS-->>PR: "{}"
PR->>RS: _restoreSettledSession() [no watchdog]
Note over PR,RS: If RS hangs here on final probe, watchdog never starts
alt more retries remain
PR->>UI: "Reconnecting… (N/6)"
PR->>PR: schedule next probe
end
end
end
Note over PR: All 6 probes exhausted
PR->>UI: "Restoring session…"
PR->>RS: _restoreSettledSession() [8s watchdog]
alt RS resolves in time
RS-->>PR: true/false
PR->>PR: clearTimeout(watchdog)
else watchdog fires first (8s)
PR->>UI: _handleStreamError
end
%%{init: {'theme': 'base', 'themeVariables': {"darkMode": true, "background": "#0d1117", "primaryColor": "#21262d", "primaryTextColor": "#e6edf3", "primaryBorderColor": "#8b949e", "lineColor": "#8b949e", "textColor": "#e6edf3", "edgeLabelBackground": "#161b22", "actorBkg": "#21262d", "actorBorder": "#8b949e", "actorTextColor": "#e6edf3", "actorLineColor": "#8b949e", "signalColor": "#8b949e", "signalTextColor": "#e6edf3", "noteBkgColor": "#373320", "noteBorderColor": "#d4a72c", "noteTextColor": "#f0e6c0", "labelBoxBkgColor": "#21262d", "labelBoxBorderColor": "#8b949e", "labelTextColor": "#e6edf3", "loopTextColor": "#e6edf3", "activationBkgColor": "#30363d", "activationBorderColor": "#8b949e"}}}%%
sequenceDiagram
participant ES as EventSource
participant PR as _probeReconnect
participant SS as /api/stream/status
participant RS as _restoreSettledSession
participant UI as Composer UI
ES->>PR: onerror fires
PR->>UI: "Reconnecting… (1/6)"
loop For each of 6 retry delays
PR->>SS: GET /api/chat/stream/status
alt stream active
SS-->>PR: "{active:true}"
PR->>ES: reconnect EventSource
else replay available
SS-->>PR: "{replay_available:true}"
PR->>ES: reconnect with replay
else still not settled
SS-->>PR: "{}"
PR->>RS: _restoreSettledSession() [no watchdog]
Note over PR,RS: If RS hangs here on final probe, watchdog never starts
alt more retries remain
PR->>UI: "Reconnecting… (N/6)"
PR->>PR: schedule next probe
end
end
end
Note over PR: All 6 probes exhausted
PR->>UI: "Restoring session…"
PR->>RS: _restoreSettledSession() [8s watchdog]
alt RS resolves in time
RS-->>PR: true/false
PR->>PR: clearTimeout(watchdog)
else watchdog fires first (8s)
PR->>UI: _handleStreamError
end
Reviews (1): Last reviewed commit: "fix(#5541): iOS PWA streaming freeze + S..." | Re-trigger Greptile
| // Last-ditch: the stream may have finished while we were retrying. | ||
| // _restoreSettledSession polls the full session API (not just stream | ||
| // status) and can recover a completed response without an error banner. | ||
| // This is especially important on iOS where Tailscale reconnects can | ||
| // take longer than the retry window. | ||
| setComposerStatus('Restoring session…'); | ||
| let _restoreTimedOut=false; | ||
| const _restoreTimer=setTimeout(()=>{ | ||
| // If _restoreSettledSession hangs (flaky Tailscale), don't leave | ||
| // the UI stuck on "Restoring session…" forever. Fall through to | ||
| // _handleStreamError after 8s. | ||
| _restoreTimedOut=true; | ||
| if(!_terminalStateReached&&!_streamFinalized){ | ||
| if(_deferStreamErrorIfOffline()) return; | ||
| if(_deferStreamErrorIfPageHidden(source)) return; | ||
| _flushReasoningToAnchor(); | ||
| _scheduleAnchorRegistryCleanup(120000); | ||
| _handleStreamError(source); | ||
| } | ||
| },8000); | ||
| try{ | ||
| if(await _restoreSettledSession(source, {preserveVisibleOnShorterTerminalSnapshot:true})){ | ||
| if(_restoreTimedOut) return; // timer already fired _handleStreamError | ||
| clearTimeout(_restoreTimer); | ||
| return; | ||
| } | ||
| }catch(_){ | ||
| // _restoreSettledSession threw. If the timer already fired, | ||
| // _handleStreamError was called there; we return below. | ||
| // Otherwise the code below cancels the timer and calls it directly. | ||
| } | ||
| if(_restoreTimedOut) return; // timer already fired _handleStreamError | ||
| clearTimeout(_restoreTimer); | ||
| if(_terminalStateReached||_streamFinalized) return; | ||
| if(_deferStreamErrorIfOffline()) return; | ||
| if(_deferStreamErrorIfPageHidden(source)) return; | ||
| _flushReasoningToAnchor(); | ||
| _scheduleAnchorRegistryCleanup(120000); | ||
| _handleStreamError(source); |
There was a problem hiding this comment.
Watchdog gap: final probe calls
_restoreSettledSession twice, but only the second call is guarded
On the last probe (attempt=5, where nextDelay is undefined), the code first falls through to the existing _restoreSettledSession call at line 5788, which has no timeout. Only if that returns false does execution reach the new last-ditch block at line 5804 where the 8s _restoreTimer starts. If the line-5788 call hangs on a flaky Tailscale link — the exact scenario the watchdog is designed to handle — the timer never gets created, leaving the UI stuck on "Reconnecting… (6/6)" indefinitely instead of the intended 8s cap. The watchdog is well-designed for the second call but doesn't cover the prior call in the same probe.
Release: #5541 — iOS PWA streaming freeze + SSE reconnect reliability
Ships @luperrypf's mobile streaming reliability fix, rebuilt on current
master, with a gate-caught defect fixed and the full gate green. Nathan approved shipping on the green gate.The bug (#5541)
On iOS/Android PWA, a long streaming conversation could freeze — WKWebView re-lays-out the whole document on every
scrollHeightread as the response grows. Separately, a dropped SSE connection (iOS Tailscale/VPN reconnect) could exhaust the retry ladder and leave a permanently dead stream needing a manual refresh.The fix
@media (pointer: coarse)): off-screen.msg-rows getcontent-visibility:auto+contain-intrinsic-size:auto 1px+contain:layout style, so WKWebView skips their layout/paint during streaming. The live turn is kept fully rendered via the stable#liveAssistantTurnid (covers every render path — Compact Worklog, Transparent Stream, restored-live) plus thedata-live-assistant*forms, so it never blanks mid-stream and its height keeps growing so the new-message cue still fires. Scoped to touch — desktop rendering is pixel-identical. Deliberately does not setoverflow-anchor:none(inert on iOS WebKit; re-opens the bug: WebUI is unusable on Android - transcript jumps to top on every interaction (regression in v0.51.576) #4856/fix(chat): suppress browser overflow-anchor during JS scroll-anchor realign (mobile scroll jump-back) #5338 Android jump-to-top regression).[1500,3000,5000,8000]→+[12000,20000], then a last-ditch_restoreSettledSessionfull-session poll (8s-watchdog-guarded, with a "Restoring session…" affordance) recovers a response that completed while the connection was flapping — without an error banner.Gate-caught defect (fixed this round)
The first gate found the live-turn force-visible selector only covered
data-live-assistant-turn(stamped only by the Transparent Stream path) — so a normal live turn (#liveAssistantTurnwithout the attribute) on touch could becontent-visibility-skipped and blank out mid-stream. Fixed by adding#liveAssistantTurn(the stable id set on all render paths) to the force-visible selector. Verified live in-browser.Gate (green)
#liveAssistantTurnforce-visible confirmed live, desktopcontent-visibilityunaffected (fine-pointer), zero JS console errors.Attribution: original author @luperrypf (
Co-authored-bytrailer preserved).Closes #5541