Conversation
…ders
A conversation of a few thousand messages becomes progressively unusable:
every renderMessages() rebuilds the whole transcript, and renderMessages()
runs on every send, every SSE batch and every refresh. Measured on a
synthetic 2000-message session (desktop Chromium; a phone is several times
slower), with the shipped default of virtualize_transcript=false:
loaded rows render DOM nodes
37 46 ms 960
537 546 ms 13,460
2000 2,883 ms 50,007
Render cost tracks S.messages.length exactly, so the bug is whatever loads
the whole transcript into the browser. Two things did, and together they
formed a one-way ratchet that no later action undid.
1. _messageReloadLimitForSession() returned null — "no msg_limit", i.e. the
bare full-transcript request — in two cases: when the previous load was
not truncated, and when the reload hint exceeded the server's msg_limit
ceiling. The hint grows to the loaded row count, so paging back through a
long conversation crossed the 500-row ceiling and every subsequent
same-session refresh (tab focus, SSE catch-up, visibility change) refetched
and re-rendered all 2000 rows. The null was deliberate: a clamped window
would silently drop already-loaded older rows (nesquena#6152/nesquena#6154).
The window is now always bounded, and that invariant is upheld a different
way — rows older than the returned window that we already hold are retained
and re-prepended, so a bounded refresh still cannot lose history. The
re-join is skipped, letting the server window win, when the prior rows
can't be trusted to line up (the session shrank server-side via fork/undo/
truncate, or we don't hold enough rows to cover the gap): losing older rows
to a re-fetch is recoverable via "load earlier", splicing a stale prefix
onto a rewritten transcript is not. loadSession() clears S.messages and
_oldestIdx before the reload fetch, so the pre-clear offset is stashed
alongside the existing carry-forward snapshot and consumed with it.
2. nesquena#6999 capped the auto-expanding render window at 4x the default so it could
not grow to the whole loaded transcript, but the two stream-completion sites
in messages.js still expanded it to every loaded row. A single completed
turn silently undid the cap. All three sites now share one bounded helper in
ui.js so the bound cannot drift apart again. This restores the nesquena#6999
invariant (it does not by itself bound render cost — with virtualization off
the DOM is built for every loaded row regardless of the window).
Deliberately out of scope, per the "fix the class" rule: the render window is
also expanded to the full transcript by jump-to-session-start (ui.js),
jump-to-message when the target is outside the window (ui.js), and the outline
jump (outline.js), and outline.js still issues a bare no-msg_limit fetch.
Those are explicit user requests to reveal the whole transcript, addressed by
absolute index; windowing them needs a jump-to-index-with-context fetch, which
is a separate change.
Verified: reload requests always carry msg_limit; the render window stays at
the 200 cap across repeated turn completions instead of ratcheting to the
loaded row count; 2455 tests across the 155 files that read the changed JS
pass (the 13 failures in test_update_banner_fixes.py are pre-existing
test-ordering pollution — identical on a clean tree, and the file passes
standalone both before and after).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018y9pr7ppZgDuCZCT6x96kY
… transcript refetch (nesquena#6826)
…x is provably identical; compression-anchor coverage (nesquena#6826 r3)
…U) for bounded tail read (nesquena#6826 r3 round-3)
… data_version guard (nesquena#6826 r4)
…al WAL interleaving test (nesquena#6826 r5)
Removes the local re-import of api.streaming._session_payload_with_full_messages inside test_in_tail_duplicate_guard_refuses_bounded_and_full_revision_accepted; it is already imported at module level (line 17) and unused in the function body, tripping F401 on CI's hosted lint. Mechanical maintainer fix per auto-fix policy. Co-authored-by: webtecnica <webtecnica@users.noreply.github.com>
Bounded sidecar-anchored tail read for regenerate (closes nesquena#6826). Thanks @webtecnica — six rounds of adversarial re-gate to convergence.
…ail read (nesquena#7204, @webtecnica) Adds the CHANGELOG entry for nesquena#7204 (closes nesquena#6826). Release-metadata only; the code change ships via the merge of nesquena#7204 (contributor webtecnica).
Release exp-v0.52.264: fast regenerate via bounded sidecar-anchored tail read (nesquena#7204, @webtecnica)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Split 1 of 4 from a combined branch (
alanjds/hermes-webui#2). Affects the default path (virtualize_transcriptoff) — this is the one that removes the cliff.Problem
renderMessages()rebuilds the whole transcript and runs on every send, every SSE batch and every refresh. On a synthetic 2000-message session (desktop Chromium; a phone is several times slower), with the shipped defaultvirtualize_transcript=false:Cost tracks
S.messages.lengthexactly. The server is not involved —Session.load()is 4.6 ms andGET /api/session~330 ms flat at both 200 and 2000 messages. So the bug is whatever loads the whole transcript into the browser.Two things did, and together they formed a one-way ratchet nothing undid short of a page reload.
1. The bare full-transcript refetch
_messageReloadLimitForSession()returnednull— "send nomsg_limit" — when the previous load wasn't truncated, or when the reload hint exceeded the server ceiling. The hint grows to the loaded row count, so paging back crossed the 500-row ceiling and from then on every same-session refresh (tab focus, SSE catch-up, visibility change — which fire aggressively on mobile) refetched and re-rendered all 2000 rows.That
nullwas deliberate: a clamped window would silently drop already-loaded older rows (nesquena#6152/nesquena#6154). The window is now always bounded, and the invariant is upheld differently — rows older than the returned window that we already hold are retained and re-prepended.The re-join is skipped, letting the server window win, when the prior rows can't be trusted to line up: the session shrank server-side (fork / undo / truncate), or we don't hold enough rows to cover the gap. Losing older rows to a re-fetch is recoverable via "load earlier"; splicing a stale prefix onto a rewritten transcript is not.
loadSession()clearsS.messagesand_oldestIdxbefore the reload fetch, so the pre-clear offset is stashed alongside the existing carry-forward snapshot and consumed with it.2. The render-window ratchet
nesquena#6999 capped the auto-expanding render window at 4× the default, but the two stream-completion sites in
messages.jsstill expanded it to every loaded row — one completed turn undid the cap. All three sites now share one bounded helper.This restores the nesquena#6999 invariant; it does not by itself bound render cost (with virtualization off the DOM is built for every loaded row regardless). Part 1 is what removes the cliff.
Deliberately out of scope
The render window is also expanded to the full transcript by jump-to-session-start, jump-to-message when the target is outside the window, and the outline jump; and
outline.jsstill issues a bare no-msg_limitfetch. Those are explicit user requests to reveal the whole transcript, addressed by absolute index — windowing them needs a jump-to-index-with-context fetch.Verification
msg_limit, verified by intercepting the page's real/api/sessiontraffic.tests/run is too flaky to use as a signal — ~45–50 failures on both this branch and pristine master, differing in both directions, all intest_workspace_git.py/test_webui_state_db_reconciliation.py.test_webui_external_refresh_frontend.pypinned the exact source strings this changes; those assertions are updated to the new mechanism, plus two new tests (the retained-prefix re-join and its bail-outs; no expansion site still using the raw uncapped count).