Skip to content

perf(webui): stop repainting the live activity scene when nothing changed - #4

Open
alanjds wants to merge 13 commits into
masterfrom
claude/webui-streaming-fps
Open

alanjds wants to merge 13 commits into
masterfrom
claude/webui-streaming-fps

Conversation

@alanjds

@alanjds alanjds commented Aug 24, 2026

Copy link
Copy Markdown
Owner

While a turn streams, the whole page collapses to single-digit FPS on a large session — measured at 4 fps on a real instance (Firefox, ~2000-message transcript), with sustained ~40% CPU. Typing lags because keystrokes wait for a frame, not because the composer is slow.

Where the frames go

Chromium CDP sampling profiler against a synthetic SSE stream driving the real _wireSSE handlers, 2000-message transcript, virtualization off, ~64k chars accumulated in the live turn:

renderLiveAnchorActivityScene         265 calls / 10s (26.4/s)  4940ms  = 49% wall
_restorePinnedMessageScrollSnapshot   528 calls (2 per repaint) 2469ms  = 25%
_captureMessageScrollSnapshot + _captureMessageViewportAnchor   1448ms  = 14%
_renderAnchorSceneRowsIntoWorklog (the actual row build)         410ms  =  4%
renderMessages                        0 calls

renderMessages is not involved at all. The scene repaint dominates, and ~39% of wall clock is the scroll-snapshot capture/restore bracketing it — each reads scrollHeight / clientHeight / getBoundingClientRect immediately after a DOM teardown, forcing a synchronous layout of the entire transcript. Only ~4% is the row building the repaint exists to do.

The rate is also unbounded. _scheduleRender caps DOM updates at ~15fps precisely because they are expensive, but the scene repaint escapes that cap: one _doRender requests it twice (_renderLiveThinking → updateThinking → appendThinking, then _upsertAnchorProcessProse → _renderAnchorLiveScene), and every reasoning SSE event requests it again — 26.4/s against a 15fps budget, with roughly half producing byte-identical DOM.

The fix

Memoize the last painted scene at the shared chokepoint and skip the teardown/rebuild when a request would reproduce it. Both render modes are patched (compact_worklog and transparent_stream) since both rebuild from the same rows and pay the same reflow pair.

The comparison is structural over the rendering rows plus the globals the renderers read (mode, ids, _showThinking, worklog-open default, turn start) rather than a hand-picked field list, so no field can silently freeze a live row. Row/payload strings are shared by reference with the anchor registry, so the common case short-circuits on ===.

Fails closed. The skip is taken only when the exact turn element and row container we painted are still connected, still owned by this stream, and still hold the same child count. Any external mutation — removeThinking, a dedupe pass, a renderMessages rebuild, a session/stream/mode switch — invalidates the memo and the full repaint runs. The skipped call returns the same boolean the paint returned, because callers such as _upsertAnchorReasoning branch on it to decide whether to fall back to the legacy thinking card.

Measured

Synthetic SSE stream on the 2000-message session, compared at matched accumulated live-turn size:

live turn before after
22k chars 37.4 fps 41.6
45k chars 28.2 35.0
60k chars 21.0 31.0
80k chars 16.9 26.2
120k chars 13.4 15.9

What this does not fix

These numbers are virtualization OFF. With virtualize_transcript ON the same ramp measures 0.8–2.9 fps both before and after — no improvement. Off is the shipped default so this is still the common path, but the ON case has a separate bottleneck that this change does not touch.

The gain also decays as the live turn grows: both curves converge near 13 fps past ~140k chars, so a long enough thinking block still degrades.

Verification

  • Deterministic targeted sweep (the 155 test files that read the changed JS, plus the new file): 2462 passed, 13 failed — all 13 the pre-existing test_update_banner_fixes.py ordering pollution, identical on a clean tree. Zero non-baseline failures.
  • 9 new tests in tests/test_live_anchor_scene_repaint_memo.py, all passing.
  • ruff check clean, node --check static/ui.js clean.

A note on the full suite: running all of tests/ gives ~45–50 failures on both this branch and pristine master, and the sets differ in both directions (9 failures appear only on master). Those are all in test_workspace_git.py / test_webui_state_db_reconciliation.py — git operations and same-second timestamp granularity. The full suite is too flaky to serve as a regression signal at that scale; the targeted deterministic subset above is what this was validated against.

Helper access from the two renderers is typeof-guarded, matching how this function already reaches _moveLiveRunStatusToTurnEnd / scrollIfPinned / _restoreWorklogDetailDisclosureState, so the existing extract-and-eval node harnesses keep running the renderers with a partial dependency set.


Generated by Claude Code

claude and others added 13 commits August 24, 2026 16:23
…hanged

While a turn streams, the whole page collapses to single-digit FPS on a large
session. Profiling a live stream on a 2000-message transcript (Chromium CDP
sampling profiler, virtualization OFF, ~64k chars accumulated in the live turn)
attributes the frame budget as:

  renderLiveAnchorActivityScene   265 calls / 10s (26.4/s)  4940ms  = 49% wall
  _restorePinnedMessageScrollSnapshot  528 calls (2 per repaint)  2469ms = 25%
  _captureMessageScrollSnapshot + _captureMessageViewportAnchor  1448ms = 14%
  _renderAnchorSceneRowsIntoWorklog (the actual row build)        410ms  = 4%
  renderMessages                                                  0 calls

So the scene repaint dominates, and ~39% of wall clock is the scroll-snapshot
capture/restore bracketing it — each of those reads scrollHeight/clientHeight/
getBoundingClientRect straight after a DOM teardown, forcing a synchronous
layout of the ENTIRE transcript. Only ~4% is the row building the repaint
exists to do.

The repaint rate is also unbounded: _scheduleRender caps DOM updates at ~15fps
precisely because they are expensive, but the scene repaint escapes that cap.
One _doRender asks for it twice (_renderLiveThinking -> updateThinking ->
appendThinking, then _upsertAnchorProcessProse -> _renderAnchorLiveScene) and
every `reasoning` SSE event asks again -- 26.4/s against a 15fps budget, with
roughly half of those requests producing byte-identical DOM.

Fix at the shared chokepoint: memoize the last PAINTED scene and skip the
teardown/rebuild when a request would reproduce it. Both render modes are
patched (compact_worklog and transparent_stream) since both rebuild from the
same rows and pay the same reflow pair. The comparison is structural over the
rendering rows plus the globals the renderers read (mode, ids, _showThinking,
worklog-open default, turn start), not a hand-picked field list, so no field
can silently freeze a live row; row/payload strings are shared by reference
with the anchor registry so the common case short-circuits on ===.

Fail closed: the skip is taken only when the exact turn element and row
container we painted are still connected, still owned by this stream, and still
hold the same child count. Any external mutation (removeThinking, a dedupe
pass, a renderMessages rebuild, a session/stream/mode switch) invalidates the
memo and the full repaint runs. The skipped call returns the same boolean the
paint returned, because callers such as _upsertAnchorReasoning branch on it to
decide whether to fall back to the legacy thinking card.

Measured with a synthetic SSE stream (fake EventSource driving the real
_wireSSE handlers) on the 2000-message session, virtualization OFF, comparing
at matched accumulated live-turn size:

  live turn   22k chars   37.4 -> 41.6 fps
              45k chars   28.2 -> 35.0 fps
              60k chars   21.0 -> 31.0 fps
              80k chars   16.9 -> 26.2 fps
             120k chars   13.4 -> 15.9 fps

Helper access from the two renderers is typeof-guarded, matching how this
function already reaches _moveLiveRunStatusToTurnEnd / scrollIfPinned /
_restoreWorklogDetailDisclosureState, so the existing extract-and-eval node
harnesses keep running the renderers with a partial dependency set.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018y9pr7ppZgDuCZCT6x96kY
…x is provably identical; compression-anchor coverage (nesquena#6826 r3)
Removes the local re-import of api.streaming._session_payload_with_full_messages
inside test_in_tail_duplicate_guard_refuses_bounded_and_full_revision_accepted;
it is already imported at module level (line 17) and unused in the function body,
tripping F401 on CI's hosted lint. Mechanical maintainer fix per auto-fix policy.

Co-authored-by: webtecnica <webtecnica@users.noreply.github.com>
Bounded sidecar-anchored tail read for regenerate (closes nesquena#6826). Thanks @webtecnica — six rounds of adversarial re-gate to convergence.
…ail read (nesquena#7204, @webtecnica)

Adds the CHANGELOG entry for nesquena#7204 (closes nesquena#6826). Release-metadata only;
the code change ships via the merge of nesquena#7204 (contributor webtecnica).
Release exp-v0.52.264: fast regenerate via bounded sidecar-anchored tail read (nesquena#7204, @webtecnica)

alanjds commented Sep 4, 2026

Copy link
Copy Markdown
Owner Author

Consolidated long-conversation performance findings (input to the Svelte migration go/no-go)

Parking this here because issues are disabled on this fork and this PR carries the CDP profile it depends on. Another session asked for these findings to answer: "at what point does this stop being spike work and become 'let's plan the actual migration'? At the point that we got sure that our version can solve the long list issue without worse issues. Are we on that point already?"

Everything below is measured. Where something is unknown, I say so.

Short answer

Not yet — but the gap is one specific unmeasured thing, not general unease.

"The long list issue" was never one issue. It decomposed into five distinct bugs, four now fixed and measured. The fifth — why streaming collapses to ~1–3 fps with virtualize_transcript on — is unprofiled. Until someone profiles it, a rewrite cannot claim to solve it, because nobody currently knows what it is.

What it decomposed into

# Bug Scope Status
1 Reload ratchet — every same-session refresh refetched the whole transcript default path fixed, #5
2 Render-window ratchet — one completed turn undid the nesquena#6999 cap default path fixed, #5
3 Virtualized transcript re-rendered ~10×/sec forever while idle virtualization only fixed, #9
4 Reader in history drifted backward during streaming virtualization only fixed, #7
5 Streaming collapses the page to single-digit fps both paths partly fixed, this PR

Plus an adjacent correctness bug: submitEdit had no in-flight guard — seven concurrent /api/session/truncate calls from repeated clicks, with a keep_count race that could truncate a long session to a handful of messages (#3).

The numbers that bear on an architecture decision

The server is not the bottleneck. Synthetic 2000-message session (1.3 MB sidecar):

200 msgs 2000 msgs
Session.load() 0.6 ms 4.6 ms
Session.save() 4.7 ms 31.8 ms
GET /api/session ~330 ms ~330 ms

Flat. A migration justified by "the Python server can't keep up" would be justified by something that isn't happening.

Render cost tracks S.messages.length exactly (virtualization off — the shipped default):

loaded rows render DOM nodes
37 46 ms 960
537 546 ms 13,460
2000 2,883 ms 50,007

With virtualization on it is flat from 250 → 2000 rows (~120 ms, 1,555 nodes). The existing virtualization does work for the transcript DOM — windowing was never impossible in vanilla JS.

The streaming collapse is not in the transcript renderer — see the profile in this PR's description. renderMessages: zero calls. Half the frame budget is one function repainting the live activity scene ~26×/sec against a 15 fps budget it escapes, roughly half of those producing byte-identical DOM. Another 39% is scroll-snapshot capture/restore forcing full-transcript layout after a DOM teardown. Four percent is the work the repaint exists to do.

The open unknown, stated precisely

With virtualize_transcript on, the ramp measures 0.8–2.9 fps before and after this PR's fix. No improvement. On a real instance: 4 fps on, 10 fps off, ~40% sustained CPU.

Nobody has profiled the ON path — the profile above was taken with virtualization off. That is the largest unknown in the picture, and it is squarely in the territory a migration would be expected to fix.

This is the thing to measure before deciding. The harness exists (scratchpad/fps/ in the session that produced this PR: a fake EventSource driving the real _wireSSE handlers, plus CDP sampling). Pointing it at the ON path is hours, not days, and converts the go/no-go from a judgement call into a fact.

Methodological warning, offered seriously

Over this investigation I was wrong three times, each time reasoning from plausible architecture rather than measurement:

  1. Attributed the scroll drift to a follow-to-bottom path. Wrong — instrumentation showed nothing wrote scrollTop at all.
  2. Attributed it to stale anchor geometry. Wrong — fixing the geometry (34% → 6.6% error) changed the drift by exactly zero pixels (fix: calibrate virtual row heights from measurements, not flat constants #8). The real cause was postProcessRenderedMessages growing rows above the viewport, unabsorbed on desktop because .messages is overflow-anchor: none.
  3. Attributed composer lag to a quadratic <think> re-parse. Wrong as the primary cause — the page was at 4 fps and keystrokes were waiting for a frame.

Each correction came from instrumenting, not thinking harder. Relevance to the decision: "a clean rewrite will fix this" is exactly the kind of plausible-but-unmeasured claim that was wrong three times here. Two of these bugs (#3's loop, this PR's repaint) are architecture-independent — a Svelte version that reimplements the same live-scene repaint pattern, or brackets renders with full-transcript layout reads, inherits them.

What I would want true before calling it migration-planning

  1. The ON-path streaming collapse is profiled and named. Currently unknown.
  2. A target number exists. "Fast enough" needs a figure — e.g. ≥30 fps during streaming at 2000 messages / 80k live chars, since that is where the default path now sits after this PR.
  3. The spike is measured on the same harness. A spike judged by impression against a codebase judged by profiler is not a comparison.

1 and 3 are cheap. 2 is a product decision.

Where the detail lives

#5 ratchets (default path, largest win) · #9 re-render loop, both halves · #7 post-process anchor hold · this PR live-scene repaint memo · #8 row-height calibration, including its honest negative result · #3 submitEdit guard.

Each carries its own measurements, rejected alternatives, and what it explicitly does not fix.

🤖 Generated with Claude Code

https://claude.ai/code/session_018y9pr7ppZgDuCZCT6x96kY


Generated by Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants