fix(web-shell): batch transcript dispatch to avoid tab-return freeze - #7012
Conversation
Dispatching each buffered SSE event individually makes a tab-return burst O(events x blocks) on the main thread (per-dispatch block-array copy + freeze), freezing very long sessions for minutes. Coalesce the live stream into one dispatch per macrotask, cap the client's in-memory transcript window, and skip the dev-only block freeze in production.
|
Thanks for the PR! Template looks good ✓ Problem: Real, well-documented performance issue. The O(E×B) main-thread freeze on tab-return SSE burst is traced to specific code paths ( Direction: Aligned. Web Shell transcript performance is core to the product — long sessions becoming unusable after a tab switch is a genuine user-facing defect. Size: 174 production lines (140 in Approach: Focused and minimal. The batcher is local to the Moving on to code review. 🔍 中文说明感谢贡献! 模板完整 ✓ 问题:真实且有据可查的性能问题。标签页切回时 SSE 突发导致的 O(E×B) 主线程冻结,已追溯到具体代码路径(每次 dispatch 中 方向:对齐。Web Shell transcript 性能是产品核心体验——长会话在标签页切换后变得不可用是真实的用户级缺陷。 规模:174 行生产代码( 方案:聚焦且精简。批处理器局限在 进入代码审查 🔍 — Qwen Code · qwen3.7-max Reviewed at |
c2237a5 to
5866da8
Compare
|
Please do not rebase or force-push to an active PR as it invalidates existing review comments. Note for future reference, the bots always squash all changes into a single commit automatically as part of the integration. 中文请勿对活跃的 PR 执行 rebase 或 force-push,因为这会使已有的评审评论失效。另外,供日后参考:作为集成流程的一部分,机器人始终会自动将所有改动压缩(squash)为单个提交。 |
🖼️ web-shell visual previewRendered against a mock daemon (no real backend): the PR base vs this PR head Screenshots · before / afterFull-resolution recordings (.webm) are attached to the workflow run. — Qwen Code · web-shell visuals |
Code ReviewIndependent proposal (from title + "Why"): buffer transcript events in the live SSE loop, flush in a single The PR matches this exactly, and goes further in the right ways:
No critical blockers found. No AGENTS.md violations. TestingThis is a browser-side performance fix (main-thread freeze on tab-return from SSE burst). Terminal-based tmux testing cannot demonstrate the fix — the behavior change is dispatch coalescing, invisible in CLI output. CI and local test runs are the evidence. Local test results (PR code applied to worktree)DaemonSessionProvider tests (webui): Includes 3 new tests:
SDK tests: Typecheck: clean on both CI (GitHub Actions): all checks green — SummaryClean, focused performance fix. The implementation is correct, well-tested, and addresses every concern raised by the previous review round (ytahdn's observer-debug-guard finding, qwen-code-ci-bot's Criticals on unprotected dispatch and missing catch-block flush — all resolved). The design doc is unusually thorough and the test coverage pins the key invariants (coalescing, ordering, flush-not-drop on unmount, observer guard). 中文说明代码审查独立方案(仅根据标题和"为什么需要"):在实时 SSE 循环中缓冲 transcript 事件,以每个宏任务一次 PR 完全匹配此方案,并在正确的方向上更进一步:
未发现阻塞性问题。未违反 AGENTS.md 规范。 测试这是浏览器端的性能修复(标签页切回时 SSE 突发导致的主线程冻结)。基于终端的 tmux 测试无法展示此修复——行为变化是 dispatch 合并,在 CLI 输出中不可见。CI 和本地测试运行是证据。 本地测试结果(PR 代码应用到 worktree)
总结精简、聚焦的性能修复。实现正确,测试充分,解决了上一轮审查提出的所有问题(ytahdn 的 observer-debug-guard 发现,qwen-code-ci-bot 关于未保护 dispatch 和缺失 catch-block flush 的 Critical——全部已解决)。设计文档非常详尽,测试覆盖锁定了关键不变量(合并、排序、卸载时 flush 而非丢弃、observer 守卫)。 — Qwen Code · qwen3.7-max Reviewed at |
|
Confidence: 4/5 — clean performance fix, well-tested, every prior review concern resolved. This is a textbook example of a focused performance PR done right. The problem is real and well-traced (O(E×B) main-thread freeze on tab-return SSE burst), the fix is minimal (batcher local to the Going back to my independent proposal from Stage 2: buffer events, flush per macrotask, sync at control boundaries. The PR matches this exactly. I don't see a simpler path — the batcher is already about as small as it can be while being correct. The secondary improvements (production freeze skip, maxBlocks cap) are well-bounded and well-justified. Every concern from the prior review rounds has been addressed:
154 provider tests pass locally (including 3 new ones), 1356 SDK tests pass, typechecks clean, CI green. No blocking issues found. Approving. ✅ 中文说明信心度:4/5 — 干净的性能修复,测试充分,之前审查的所有问题均已解决。 这是一个聚焦的性能 PR 的典范。问题是真实且有据可查的(标签页切回时 SSE 突发的 O(E×B) 主线程冻结),修复精简(批处理器局限在 回到我在 Stage 2 的独立方案:缓冲事件、按宏任务刷新、在控制边界同步。PR 完全匹配此方案。我没有找到更简单的路径——批处理器在保持正确性的前提下已经尽可能精简了。附带改进(生产环境跳过冻结、maxBlocks 上限)有明确边界和充分理由。 之前审查轮次的所有问题均已解决:
本地 154 个 provider 测试通过(含 3 个新增),1356 个 SDK 测试通过,typecheck 无错误,CI 绿色。未发现阻塞性问题。 已批准。✅ — Qwen Code · qwen3.7-max Reviewed at |
qwen-code-ci-bot
left a comment
There was a problem hiding this comment.
LGTM, looks ready to ship. ✅
ReviewVerdict: the core fix is correct and does what it claims — I verified the asymptotics and every ordering control point against the code, and ran the affected suites locally at the PR head. No blockers. One small teardown change is worth making before merge, plus a portability hardening and some design-doc drift to clean up. Independently verified
Should fix before merge1. Teardown should flush, not drop — the design doc already says so. 2. First bare const FREEZE_TRANSCRIPT_BLOCKS =
typeof process !== 'undefined' && process.env.NODE_ENV !== 'production';App builds still fold this to Design doc / test-plan drift (minor)
Noted, no action needed
🤖 Generated with Claude Code — Claude Fable 5 |
… browser Address review feedback: teardown now flushes buffered transcript events instead of dropping them (the SSE client advances lastSeenEventId as events are yielded, so a dropped buffer would be skipped by a same-session incremental resume). Guard FREEZE_TRANSCRIPT_BLOCKS with typeof process so an unbundled browser consumer of the daemon/ui surface does not throw a ReferenceError. Add a dispatch-count assertion to the burst test and an unmount-flush regression test, and align the design doc (setTimeout-only flush, verification plan).
ytahdn
left a comment
There was a problem hiding this comment.
Requesting changes for one reproducible transcript-ordering regression in the new batching path. The existing provider suite passes, but a focused observer-burst regression test fails as described inline.
|
Overall assessment for Using a macrotask ( The remaining concern is architectural consistency: dispatch used to update the store synchronously, while this change creates two transcript states—the committed store and Minimum path forward: make the observer/debug guard account for queued events and add the focused burst regression test. Preferably, encapsulate reads that require the effective state (committed + pending), then re-audit every |
…ursts in one block Address ytahdn's PR QwenLM#7012 review: the batched-dispatch debug guard read the committed store's activeAssistantBlockId, which lags the pending buffer within a burst, so a debug event interleaved in an observer assistant burst was not filtered and split the block. Flush the buffer before the guard, scoped to observer-mode debug events (rare) so steady streaming keeps batching. Add a focused burst regression test, make the unmount-flush test deterministic with fake timers (it was timing-racy), and update the design doc.
ytahdn
left a comment
There was a problem hiding this comment.
Re-reviewed the complete PR at 930d08876. The previously reported observer-burst ordering regression is fixed and covered by a focused regression test.
Verified locally:
- DaemonSessionProvider: 154/154
- SDK daemon UI reducer: 269/269
- Web Shell: 50/50
- SDK and WebUI typecheck
The batching, reset, replay-complete, stream-end, teardown, and observer-debug ordering paths are consistent. No remaining blocking findings. Approving; the Ubuntu CI job is still running and should remain a merge requirement.
qwen-code-ci-bot
left a comment
There was a problem hiding this comment.
Not reviewed: chunk 1, chunk 2, chunk 3 — no agent reported covering these; nobody read them.
— qwen3.7-max via Qwen Code /review
qwen-code-ci-bot
left a comment
There was a problem hiding this comment.
Not reviewed: chunk 1, chunk 2, chunk 3 — no agent reported covering these; nobody read them.
— qwen3.7-max via Qwen Code /review
The catch block at the end of the connection loop skipped the post-loop flush, leaving buffered transcript events on a scheduled timer. The retriable path resumes via Last-Event-ID without resetting the store, and lastSeenEventId has already advanced past those events, so clearing the buffer would drop them on the incremental delta-resume. Flush instead. Also route the restored-prompt settle and replay_complete control dispatches through dispatchTranscriptNow so each is self-contained (flush + dispatch) rather than relying on an earlier flush by timing, and tighten the burst regression test from toContain(CHUNK_COUNT) to toEqual([CHUNK_COUNT]) so a regression emitting redundant per-event dispatches also fails. Addresses the ci-bot review.
qwen-code-ci-bot
left a comment
There was a problem hiding this comment.
Not reviewed: Agent 0: Issue fidelity & root-cause ownership — no prompt was built for it (agent-prompt --role 0 never ran).
Not reviewed: Agent 1a: Line-by-line correctness — no prompt was built for it (agent-prompt --role 1a never ran).
Not reviewed: Agent 2: Security — no prompt was built for it (agent-prompt --role 2 never ran).
Not reviewed: Agent 3: Code quality — no prompt was built for it (agent-prompt --role 3 never ran).
Not reviewed: Agent 4: Performance & efficiency — no prompt was built for it (agent-prompt --role 4 never ran).
Not reviewed: Agent 5: Test coverage — no prompt was built for it (agent-prompt --role 5 never ran).
Not reviewed: Agent 6a: Undirected audit — attacker mindset — no prompt was built for it (agent-prompt --role 6a never ran).
Not reviewed: Agent 6b: Undirected audit — 3 AM oncall mindset — no prompt was built for it (agent-prompt --role 6b never ran).
Not reviewed: Agent 6c: Undirected audit — six-months-later maintainer — no prompt was built for it (agent-prompt --role 6c never ran).
Not reviewed: Agent 1b: Removed-behavior audit — no prompt was built for it (agent-prompt --role 1b never ran).
Not reviewed: Agent 1c: Cross-file tracer — no prompt was built for it (agent-prompt --role 1c never ran).
Not reviewed: Agent 7: Build & test verification — no prompt was built for it (agent-prompt --role 7 never ran).
Not reviewed: reverse audit — no auditor ran (Step 5 builds its prompt with agent-prompt --role reverse-audit; none was recorded, so the pass that looks for what Step 3 missed was skipped).
Not reviewed: verification — the review posts findings, but no verifier ran (Step 4 builds its prompt with agent-prompt --role verify; none was recorded, so the findings were not verified).
— qwen3.7-max via Qwen Code /review
A reducer throw inside runTranscriptFlush escaped as an uncaught setTimeout error on the macrotask path and, via flushTranscriptSync, propagated out of the catch block (aborting lastSeenEventId bookkeeping, reconnect, auth branching, terminal cleanup, and pendingSessionLoad rejection) and out of the useEffect cleanup (leaving half-torn-down state). Wrap the dispatch in try/catch and log it with the batch size so the throw is surfaced without crashing the session or skipping teardown; one guard fixes all three paths. Also document the flush precondition on settleActivePromptFromTurnEvent, which dispatches assistant.done directly and previously carried that contract only as an inline comment at the call site. Addresses the ci-bot review.
… can act on it A role with no recorded prompt proves one thing: the brief never reached an agent. The roster check claimed more than that — "no prompt was built for it (`agent-prompt --role 0` never ran)" — and on #7012 it said that about all twelve dimensions of a review that had just posted two Criticals with line numbers. The agents were in the same comment the gate was calling empty. Both failures are real and neither is the other. An orchestrator that writes the launch by hand gets an agent that runs, reads the diff and finds things, having never seen the severity bar, the finding format or this project's rules — all of which live in the brief it was never given. That is worth blocking on. It is not "nobody looked", and a check may not report the reading it cannot see. Three changes, one shape: - The per-role text says the brief never reached an agent, and that the dimension was reviewed "if at all" from a prompt the run wrote for itself. It no longer speaks for the agent's existence. - Every role briefless collapses to one line. It is one failure — the run did not use the prompt builder — and saying it twelve times buries the fact that explains all twelve. - The public body drops the internal command. `agent-prompt --role 2` is not something a PR author can run; on #7012 fourteen lines of it were the whole CHANGES_REQUESTED while the findings sat inline below the fold. The call survives in check-coverage's stderr, where the orchestrator reads it, and the role number is already in each label. check-coverage no longer leads with a count: the collapsed line covers the whole roster, so "1 required brief" would undercount it by the size of the review. Behaviour is unchanged — the gate fires on exactly the same runs and still caps the verdict. Only the sentence changes, and only where it was overclaiming or talking to the wrong reader.
qwen-code-ci-bot
left a comment
There was a problem hiding this comment.
Reviewed. Suggestions are inline.
— qwen3.7-max via Qwen Code /review
Code review @
|
doudouOUC
left a comment
There was a problem hiding this comment.
Review: Approve ✅
Reviewed the full diff at 32927e7 against main in a clean worktree. This is a well-scoped main-thread performance fix — root cause, design, and the three fixes all hold up. All prior review threads (ci-bot + ytahdn) are resolved, and I re-verified each against the current commit.
Independently verified
packages/webui—DaemonSessionProvider.test.tsx: 154 passedpackages/sdk-typescript— daemon UI reducer: 269 passedpackages/web-shell—SplitView.test.tsx: 36 passedtypecheck(webui + sdk): cleaneslint(all changed files): clean
Correctness checks that hold
- Coalescing property: the macrotask (
setTimeout 0) flush is required — a microtask flush would fire between everyfor awaitevent and never coalesce. The burst test pins it withexpect(dispatchBatchSizes).toEqual([CHUNK_COUNT]), so a per-event regression fails rather than silently passing. - "Every store read in the live loop is flush-preceded": audited all
store.*sites. The two in-loopgetSnapshot()reads (debug guard,awaitingResync) are each preceded by aflushTranscriptSync(); all pre-loop dispatches (setup/replay) run on an empty buffer. - Debug guard: faithfully restores pre-batching semantics —
debugis filtered only when the committed store has an active assistant block, and the scoped flush commits prior-iteration chunks first. - No silent loss: buffered events are always either flushed or intentionally
clearPendingTranscriptEvents()'d beforestore.reset(). The catch path flushes (not clears) because the retriable resume advanceslastSeenEventId; unmount flushes for the same reason. - Throw containment:
runTranscriptFlushswallows a reducer throw, so neither the timer path nor the catch/unmountflushTranscriptSync()can cascade into error-recovery/teardown. settleActivePromptFromTurnEventcallers both satisfy the flush precondition (the in-loop caller is flushed just above; the replay-setup caller runs on an empty buffer).- passive-assistant-done timer (
setTimeout 80/3000) always fires after the batched flush (setTimeout 0) by timer ordering — even under background throttling — soassistant.donenever lands ahead of buffered content.
Non-blocking observations
- Fix B2 (dev-only freeze): skipping the freeze in production means a latent "external consumer mutates the COW-shared blocks array in place" bug would corrupt silently rather than throw; the dev/CI safety net remains. Documented, reasonable tradeoff.
- The reducer-throw
console.errorlogseventCountbut not payloads — limits post-mortem detail but avoids writing session content to logs; acceptable. dispatchTranscriptNow's controlstore.dispatchis unguarded, but this matches pre-PR behavior and sits inside the loop's try/catch — not a regression.
Nice work — the design doc's audit trail (rounds 1–8) makes the reasoning fully traceable. The only thing not covered by unit tests is a live large-session tab-away/return stress run, which the PR already flags as recommended final confirmation.
|
@qwen-code /triage |
ytahdn
left a comment
There was a problem hiding this comment.
增量 review 通过(930d088..32927e7,2 commits):
- catch 块的 flushTranscriptSync 正确补回了 try 块内被 throw 跳过的 post-loop flush,选择 flush 而非 clear 的理由(retriable path 不 reset store)注释到位
- restored-prompt / replay_complete 两条路径改用 dispatchTranscriptNow 保证了 buffered transcript 在 assistant.done 之前的正确顺序
- runTranscriptFlush 的 try/catch 堵住了 reducer throw 的三个逃逸路径(setTimeout uncaught、catch 块 error recovery、useEffect cleanup),一个 guard 修三处
- 测试从 toContain 收紧到 toEqual 能检测冗余 dispatch regression
qwen-code-ci-bot
left a comment
There was a problem hiding this comment.
LGTM, looks ready to ship. ✅
…its FIX Two findings from the bot review of the previous commit. The shellContextEnv suite isolated QWEN_CODE_SESSION_ID and QWEN_CODE_CLI but not QWEN_CODE_PROJECT_DIR — the third variable the CLI exports to every shell, and the one this suite's own per-session tests assign without cleanup. Reproduced: run the suite with it set, as any `npm test` from inside a qwen session does, and exactly the two `.toEqual()` exact-match tests fail on a key the test never set. Same isolation, same shape, and it retires the in-file leak too. The remediation channel covered blind agents and the Step 4/5 gaps and stopped there: missing briefs, rewritten launches, unread briefs and never-opened diffs still reached the body with no FIX line beside them. A body disclosure with no repair command is how #7012's orchestrator got to "the agents clearly did their job" — the whole reason the channel exists. Each category now pushes one remediation line (missing briefs point at `--roster`; the relaunch-shaped ones say relaunch with the same printed prompt), and a test pins the pair for a roster gap: the body says "brief never reached an agent" with no command in it, and the remediation names the roster call.
Two sentences, same class, both this branch's own thesis applied to itself. A chunk agent that ran on a hand-written prompt while its chunk was never built landed in the body as "no prompt was built for it (`agent-prompt` never ran for this chunk)" — an internal command on the author-facing surface, one line per chunk on a 3B replay of the #7012 shape. The label now says what happened in the author's register (ran on a prompt the run wrote itself; the brief never reached it); the rebuild command already rides the rewritten-launches remediation line on stderr. And the Step 4/5 `not-built` texts still said "no auditor ran" / "no verifier ran" — the one residual of the overclaim this branch exists to retire. `not-built` is decided before the transcripts are consulted: a run that skipped the builder and hand-wrote the launch leaves no brief on disk whose open could be looked for, so such an auditor is invisible to the check, and "no auditor ran" claims sight it does not have. Both texts now use the roster wording: what a missing record proves (no agent was launched with a prompt this skill builds), then what it costs ("ran, if at all, without the method its brief carries"). The Delivery docstring records why. Tests pin the new sentences positively and negatively; the register pin (no `agent-prompt`/`--chunk` in a body label) guards the first one.
…ne call (QwenLM#7033) * fix(review): name a rewritten launch as itself, and leave nothing to hand-assemble Dogfooded on a real 3A review of a live PR, and the run talked its way past the gate: compose-review printed: Verdict: Comment — an Approve was NOT available: a dimension nobody reviewed the run's next thought: "the compose-review flagged reverse audit as unreviewed (transcript visibility issue — the reverse audit did run substantively with two dry rounds). Let me proceed." the run then reported, and saved: Verdict: Approve The gap was right and its wording was wrong, and the wording is what let the run dismiss it. Two auditors HAD run — 16 and 23 tool calls each — and both HAD opened their brief. What had actually happened is that the orchestrator skipped `--findings` and hand-wrote their launches, keeping only the brief pointer, so no agent was launched with the prompt the CLI built. The gap said "no agent was launched with it that opened its brief", which is false as written, and "a transcript visibility issue" is what a reader concludes from a message that does not describe what happened. So the floor tells the four shapes apart instead of collapsing them into one boolean, and each says what happened and what to do: not-built — the step was skipped; `agent-prompt --role <r>` never ran not-launched — the prompt was built and nothing was launched with it rewritten — an agent ran and opened its brief, but no agent got the built prompt: the launch was written by hand instead of pasted brief-unread — an agent got the built prompt and never opened the brief `rewritten` is the one that just happened, and it is now un-dismissable: it concedes the agent ran and read its brief, and names the orchestrator's own edit as the defect. And the path that produced it is gone: `--findings` is now REQUIRED for a role that takes findings. There is no bare-block-plus-hand-assembly path left — the command refuses, and prints one block to paste. An early reverse-audit round with nothing confirmed yet passes an empty file, which the command renders as "Nothing is confirmed yet". SKILL: Step 4/5 say `--findings` is required. Step 6 gains the second half of the lesson — you may not overrule the line compose-review gives you; a cap you can explain is still a cap, and the fix is to make the step verifiable and re-run, not to keep the verdict you preferred. Step 8's report interpolates the verdict out of the composed JSON (`jq -r .event`) instead of typing it, because the terminal is prose and the archive is forever. * fix(review): call the CLI that is running, not whatever `qwen` PATH finds Reported from a real session: `npm run dev:daemon`, `/review 6998 --comment` in the web shell, and the run died on Missing required argument: chunk with a help screen for a command that has no `--role` at all. The daemon was running the checkout — the skill it loaded is the current one, and says `--role 0` — but the skill shells out to `qwen review agent-prompt …`, and `qwen` on that machine is `/usr/bin/qwen` → a v0.19.10 global install whose `agent-prompt` predates QwenLM#6892 entirely. The skill and the CLI it was talking to were different programs. The skill assumed `qwen` on PATH is the build running it. That holds for a single install and breaks for exactly the people most likely to run a dev daemon. It is also invisible when it breaks: the error names an argument, not a version. So the entry is passed down instead of rediscovered. `scripts/cli-entry.js` is the executable entry and the one thing that knows its own path, so it publishes it as `QWEN_CODE_CLI` (`||=`, so the relaunch into dist/cli.js keeps pointing callers back at the wrapper with the shebang, not at itself). `daemon-dev.js` sets it too — the dev daemon is started as `node scripts/dev.js` and never passes through the wrapper, which is why this bit there first. `getShellContextEnvVars` passes it to every shell subprocess, beside the session and project-dir vars that are already handed down for the same reason. The skill's 23 command sites now read `"${QWEN_CODE_CLI:-qwen}" review …`; the fallback keeps hosts that do not export it on the old behaviour. PATH was the other candidate and was rejected: prepending a shim dir means writing an executable at spawn time and overriding `PATH` in an env that `normalizePathEnvForWindows` has already normalised — a `Path`/`PATH` collision on Windows in exchange for saving one variable. The env var is isolated in `shellContextEnv.test.ts` the way the session id already is: the CLI now exports it to every shell it spawns, so `npm test` run from inside a qwen session inherits it, and the exact-equality assertion would have failed on a variable the test never set. Verified by running the suite with it set. * fix(review): point the dev daemon's CLI at the source it is running, not dist Verifying the previous commit on a real `npm run dev:daemon` caught it doing a smaller version of the bug it fixes. The daemon runs the TypeScript **source** through tsx; `cli-entry.js` runs `dist/cli.js`. Pointing QWEN_CODE_CLI there traded "the subprocess is a whole major version behind" for "the subprocess is however stale the last build was" — measured on the box that reported this, dist was **105 source files** behind the daemon. Same bug, smaller hat. `scripts/dev.js` is the entry that runs what the daemon itself runs, so the dev daemon points there. It gains a shebang and the exec bit, which is what lets a caller invoke it as `"${QWEN_CODE_CLI}" review …` without knowing it needs node — the same shape `cli-entry.js` already has for the published path. Verified end to end on a headless box: started the dev daemon, read /proc/<pid>/environ (QWEN_CODE_CLI=<repo>/scripts/dev.js, -rwxr-xr-x), and ran the command that started this whole thread. Before: `Missing required argument: chunk`. After: `agent-prompt: --role 0 needs a plan with prNumber and ownerRepo` — the role is understood, and the complaint is about the fixture, which is the correct answer. * fix(review): say what a missing brief proves, once, to the reader who can act on it A role with no recorded prompt proves one thing: the brief never reached an agent. The roster check claimed more than that — "no prompt was built for it (`agent-prompt --role 0` never ran)" — and on QwenLM#7012 it said that about all twelve dimensions of a review that had just posted two Criticals with line numbers. The agents were in the same comment the gate was calling empty. Both failures are real and neither is the other. An orchestrator that writes the launch by hand gets an agent that runs, reads the diff and finds things, having never seen the severity bar, the finding format or this project's rules — all of which live in the brief it was never given. That is worth blocking on. It is not "nobody looked", and a check may not report the reading it cannot see. Three changes, one shape: - The per-role text says the brief never reached an agent, and that the dimension was reviewed "if at all" from a prompt the run wrote for itself. It no longer speaks for the agent's existence. - Every role briefless collapses to one line. It is one failure — the run did not use the prompt builder — and saying it twelve times buries the fact that explains all twelve. - The public body drops the internal command. `agent-prompt --role 2` is not something a PR author can run; on QwenLM#7012 fourteen lines of it were the whole CHANGES_REQUESTED while the findings sat inline below the fold. The call survives in check-coverage's stderr, where the orchestrator reads it, and the role number is already in each label. check-coverage no longer leads with a count: the collapsed line covers the whole roster, so "1 required brief" would undercount it by the size of the review. Behaviour is unchanged — the gate fires on exactly the same runs and still caps the verdict. Only the sentence changes, and only where it was overclaiming or talking to the wrong reader. * fix(review): name the directory the missing briefs were missing from "The prompt builder never ran" and "the prompt builder ran against a different --plan" arrive at this check as the same thing — an absent file — and they are fixed differently. Nothing in the error told them apart. The record directory hangs off the plan path as given, so a relative --plan resolves against the caller's cwd, and the skill runs Steps 2-6 from inside the worktree it just created. Two cwds, one relative path, two directories. Proven locally: the same `--plan .qwen/tmp/p.json` from a repo root and from a worktree under it yields two record dirs. That is not a reason to resolve the path differently — resolving a relative path against the cwd is what a relative path means, and the mismatch mostly fails loudly, because the plan is not in the worktree either and the read errors. It is a reason to print where it looked. One line, on stderr, where the orchestrator reads it; the PR author gets no path to a temp directory. * feat(review): build the whole roster in one call, because compliance decays per call The launch prompts are already small — a role line, the brief pointer, the diff reads — and it did not save the run that stopped building them. Dogfooded on one PR, the same environment went from a clean review to "no prompt was built for any of twelve roles" over three reviews in a day. The per-agent form asks the orchestrator for ~30 build-then-launch round trips on a large review, and that is a compliance cost paid per agent, per review, forever; what decays under repetition eventually decayed. `agent-prompt --roster` builds every prompt the plan requires — chunk agents, dimension agents, invariants — in one call: one labelled block per agent, each recorded under the key `check-coverage` will look it up by. The list is `requiredAgents(plan)`, the same list the coverage gate reads, so what gets built is exactly what gets checked; a key the two derive differently is refused at build time rather than surfacing later as "brief never reached an agent" on a compliant run. The blocks are separated by lines that are visibly not prompt text, and a block copied lazily — separator included — still passes the add-only delivery check. That is load-bearing: if honest-but-sloppy copying read as a rewrite, the gate would punish exactly the behaviour this call exists to buy. The per-agent forms stay, for rebuilding a single prompt after Step 3D names a gap. Step 4/5 verify and reverse-audit are untouched: they are built per round, with the findings folded in. SKILL.md's Step 3A and 3B now ask for the roster once instead of one call per agent, and check-coverage's missing-brief error names the one-call fix first. * fix(review): close the review's five consistency gaps in the CLI-pinning story Review feedback on this PR found five places where the fix stopped short of its own thesis. All five, addressed: 1. Four copyable SKILL.md commands had missed the QWEN_CODE_CLI sweep — `pr-context` (lightweight mode), `cleanup` (cache hit), `capture-local --file` (file-path reviews), `agent-prompt --whole-diff` (Agent 8). On a skewed host those modes died exactly the way the motivating run did. All four now carry the prefix; the remaining bare mentions are prose. 2. check-coverage's own stderr recommended recovery with a bare `qwen` — the message is the interface the orchestrator acts on, and on a skewed host the recommended recovery reproduced the skew. All four recommendation sites now print the prefixed form, and the rebuild hint covers `--chunk <id>`, which a missing chunk agent needs and `--role` cannot express. 3. Ambient inheritance could silently re-point an entry at another session's CLI. A dev daemon started from inside another qwen session's shell — the usual dogfooding flow — inherited that session's QWEN_CODE_CLI through `??`/`||=` and called the OUTER build: the same skew, one level up, and silent. Every entry now stamps itself unconditionally; nested sessions each call their own build. The `||=` comment in cli-entry.js also claimed a relaunch hazard that does not exist (the relaunch child runs dist/cli.js and never re-executes the wrapper) — the comment now states the real reason. 4. The third dogfooding entry point was still unpinned: `npm run dev` and `npm start` published nothing, so a /review from a plain dev TUI fell back to PATH. `scripts/dev.js` now stamps the variable in the env it spawns with — which also covers the daemon, since daemon-dev launches serve through it, and the daemon's own deferring copy is gone (one writer, not two). `scripts/start.js` does the same and gains the shebang and exec bit that make it callable as the entry it now names. 5. `"${VAR:-fallback}"` is POSIX parameter expansion, which cmd.exe passes through literally and PowerShell rejects. The skill was already POSIX-bound (Step 0 pipes through `tee`); the requirement is now total, and SKILL.md says so where the variable is introduced: on Windows, run the review from git-bash. The unconditional stamp is pinned by a test that inherits a foreign QWEN_CODE_CLI and asserts the spawned child gets this checkout's dev.js; flipping the assignment back to `??` turns exactly that test red. * fix(review): finish the two-register split, and pin the last unpinned entry Round-2 review feedback: five more places where this PR's own rules were not yet applied to itself. The Agent 7 brief handed its subagent a bare `qwen`. Its two fenced command blocks (`build-test`, `test-efficacy`) are the one call site where a SUBAGENT shells out to the review CLI — reachable by neither the SKILL.md sweep nor the stderr hints. Its shell gets QWEN_CODE_CLI exactly as the orchestrator's does, so the standard prefix works verbatim; without it, an old PATH global likely lacks these subcommands entirely, wedging the agent between its mandate (no hand-run builds) and a command that does not exist. A test now rejects any line-initial bare `qwen review` in that brief. The Step 4/5 gap texts and the blind-agent line carried remediation commands into the posted body — the register §4 stripped from missingRoles, surviving in the sibling paths, and partly ADDED by this PR (the rewritten texts). Each gap is now two sentences for two readers: `gap` (author-facing, no internal commands, rendered under `Not reviewed:`) and `fix` (orchestrator-facing, printed by compose-review to stderr as `FIX:` lines, carried on the result as `remediation`). The four-shape precision is intact — it moved channels, not content — and tests pin both directions: the body may not contain `agent-prompt`/`--findings`, and the remediation must. Pinning start.js exposed a stdout contamination: check-build-status.js printed "Checking build status..." to stdout ahead of every child, and start.js is now an entry whose stdout callers consume — `review parse-args --stdin | tee` would write a plan file whose first line is not JSON. The checker's status lines go to stderr with its warnings; `./scripts/start.js --version` now emits the version alone. Also from review: the all-briefless hint no longer points at role labels the collapsed line does not carry, and start.js's stamp gets the same test dev.js has — inherit a foreign QWEN_CODE_CLI, assert the spawned child gets this checkout's entry. * fix(review): isolate the env var this PR exports, and give every gap its FIX Two findings from the bot review of the previous commit. The shellContextEnv suite isolated QWEN_CODE_SESSION_ID and QWEN_CODE_CLI but not QWEN_CODE_PROJECT_DIR — the third variable the CLI exports to every shell, and the one this suite's own per-session tests assign without cleanup. Reproduced: run the suite with it set, as any `npm test` from inside a qwen session does, and exactly the two `.toEqual()` exact-match tests fail on a key the test never set. Same isolation, same shape, and it retires the in-file leak too. The remediation channel covered blind agents and the Step 4/5 gaps and stopped there: missing briefs, rewritten launches, unread briefs and never-opened diffs still reached the body with no FIX line beside them. A body disclosure with no repair command is how QwenLM#7012's orchestrator got to "the agents clearly did their job" — the whole reason the channel exists. Each category now pushes one remediation line (missing briefs point at `--roster`; the relaunch-shaped ones say relaunch with the same printed prompt), and a test pins the pair for a roster gap: the body says "brief never reached an agent" with no command in it, and the remediation names the roster call. * fix(review): retire the last two overclaims the round-3 review found Two sentences, same class, both this branch's own thesis applied to itself. A chunk agent that ran on a hand-written prompt while its chunk was never built landed in the body as "no prompt was built for it (`agent-prompt` never ran for this chunk)" — an internal command on the author-facing surface, one line per chunk on a 3B replay of the QwenLM#7012 shape. The label now says what happened in the author's register (ran on a prompt the run wrote itself; the brief never reached it); the rebuild command already rides the rewritten-launches remediation line on stderr. And the Step 4/5 `not-built` texts still said "no auditor ran" / "no verifier ran" — the one residual of the overclaim this branch exists to retire. `not-built` is decided before the transcripts are consulted: a run that skipped the builder and hand-wrote the launch leaves no brief on disk whose open could be looked for, so such an auditor is invisible to the check, and "no auditor ran" claims sight it does not have. Both texts now use the roster wording: what a missing record proves (no agent was launched with a prompt this skill builds), then what it costs ("ran, if at all, without the method its brief carries"). The Delivery docstring records why. Tests pin the new sentences positively and negatively; the register pin (no `agent-prompt`/`--chunk` in a body label) guards the first one. * test(review): make the every-gap-has-a-FIX claim true, and pin the partial stderr shape Round-5 review caught a test whose title outran its body: "every coverage gap … has a FIX" exercised only the missing-roles path, so dropping the remediation push for unread briefs — or rewritten launches, or never-opened diffs — failed nothing. That is the exact disclosure-without-repair state the channel exists to prevent, asserted by a test that could not see it. The title now claims what the test covers, and a sibling test covers the rest: one plan, three defects — a chunk agent on a hand-written prompt, one that never opened its brief, one that never opened the diff — asserting each category's FIX line and that none of the three drags a command into the body. Between the blind-agent test, the missing-roles test and this one, every category that discloses is now asserted to repair; mutation-checked by deleting each push in turn, one red test each. Also from the review: the missing-briefs stderr had handler coverage only for the all-briefless collapse. The partial shape — one role missing, the rest briefed — reached stderr through no test, so a formatting regression there (a broken join, a lost --roster hint, a garbled Looked-in path) would ship unseen. A second handler test pins it: the per-role detail, the rebuild hints, and the record-dir line, with the collapse text asserted absent. * fix(review): close the round-5 findings — entry contracts, gap reach, repair loops A GPT-5 review pass filed twenty-eight findings against this branch. Nineteen were real and are fixed here; two were refuted with evidence (the scripts test suite IS in CI: `test:ci` runs `npm run test:scripts`); the rest are recorded follow-ups of documented floor designs. Entry contracts. The standalone package launches through a shim that carries the bundled Node and announces itself via QWEN_CODE_LAUNCHER_PATH — stamping cli-entry.js there handed subprocesses a `#!/usr/bin/env node` script on hosts that may have no system Node; the shim is now preferred, with a test. The variable also predates this branch with a second meaning: desktop tooling sets it to a vendored dist/cli.js — a module path, no shebang — which a POSIX shell would run as a shell script; getShellContextEnvVars now drops a shebang-less script (and only a script: a native binary needs none), restoring the bare `qwen` fallback for those hosts. And both dev launchers read a signal-killed child (`code === null`) as exit 0 — a killed gate command reported green; both now re-raise the signal, with close(null, 'SIGKILL') regressions. The production entry's stamp gets the test only the dev entries had. Gap reach. `not-launched` said the pass "did not run" — but a hand-written launch that never opened the brief lands in that shape too, so it now uses the certification language the other shapes got. The roster check judged only the FIRST transcript matching a built prompt, so a failed attempt masked the compliant relaunch that the remediation itself prescribes — all matches are consulted now. An agent flagged rewritten is no longer also flagged unopened (contradictory repairs for one agent), and the all-briefless collapse no longer coexists with one "none was built" line per chunk transcript. Repair loops. Every rebuild command the run prints is now executable as written — plan, selector, and `--rules` included, because a rebuild without the rules file writes a rules-free brief that every delivery check still passes; the verify variant stops inviting the empty findings file that is only legitimate for a reverse-audit round. check-coverage prints exact selectors beside the human labels. Idle agents and unread chunks get FIX lines too, and a handler test pins the boundary: every FIX on stderr, before the verdict, never in the JSON. SKILL.md Step 6 now says what FIX lines are for: one bounded repair round, recompose, then the cap stands. Roster integrity. The output is self-checking against the 30 000-character shell truncation the skill itself documents — numbered blocks, an end-of-roster line, and SKILL.md redirects it to a file read back paged. A PR-controlled filename can no longer forge a block boundary: control characters flatten to spaces in the label and the launch prompt, and a test pins the separator count. The jq interpolation in the report template is gone — the verdict line is copied from Step 6's output, not recomputed by a binary the host may not have. The findings read-error no longer advises omitting a flag another guard requires. * fix(review): filter by overwriting, not omitting — the spread carries what the record drops The shebang filter fixed the wrong layer. It omitted QWEN_CODE_CLI from the record getShellContextEnvVars returns — but every spawn site composes the child env as `{...process.env, ...vars}`, so a key omitted from the additive record arrives anyway, inherited through the spread. On exactly the hosts the filter was written for (desktop tooling setting the variable to a shebang-less vendored dist/cli.js), the value leaked through and every `"${QWEN_CODE_CLI:-qwen}"` in the skill died on exit 126 — where before this branch those hosts ran bare `qwen` and worked. The fix is the pattern this same function already documents for the agent/prompt IDs: write an EMPTY string, which overwrites the inherited value through the spread, and which the consumer's `:-` expansion treats exactly like unset. The test comment that justified omission — "an empty string would shadow the fallback" — was true only of the colon-less `${VAR-qwen}` form and is corrected where it stood, so the reasoning that produced the bug does not outlive it. The tests now assert on the channel the bug lived in: composing `{...process.env, ...getShellContextEnvVars()}` and reading the child env — for the shebang-less case, the unreadable-path case, and the pass-through case. Reverting the overwrite to an omission turns exactly the two filter tests red. Verified end-to-end: with the desktop shape in the parent env, a child shell resolves `"${QWEN_CODE_CLI:-qwen}"` to the PATH `qwen` again. Also from the same review: the two adjacent `missingReceipts` blocks in compose-review are one block now (disclosure and repair cannot drift apart), and the `Exact selectors:` line says a rebuild of an already-built role is idempotent, so the over-prescription cannot make an operator hesitate. * fix(review): reunite roleLabel with the doc comment the selectorOf insertion orphaned The insertion left roleLabel's one-line JSDoc stranded above selectorOf, stacked on top of the new function's own — a maintainer chasing a wrong-label bug would have edited the rebuild-flags function. Each doc sits on its function again. * fix(review): close the round-9 findings — convergence, injectivity, and the claims a record can carry Fourteen findings from a GPT-5 review of the previous head; twelve fixed here, one was already fixed in the commit the review missed, one re-recorded as the standing roster-design follow-up. Repair loops now converge. Coverage accumulated every historical failed transcript, so the relaunch its own FIX line prescribes ADDED a transcript while the failed one kept its flag — ok stayed false, the same FIX printed forever. A failed attempt is now superseded by a compliant attempt at the same target (same chunk served verbatim with the diff opened; same built prompt delivered to an agent that opened its brief), and a rewritten agent is not also told to relaunch the prompt that was the defect. One transcript, one credit. Pasting the whole roster output to a single agent produced one transcript that verbatim-contains every block, matched every requirement independently, and certified an N-agent fan-out with one reader (reproduced upstream: roster 8, agents 1, ok true). Requirements now claim distinct transcripts; the paste-all run fails with a sentence that names the mistake. Records claim only what they prove. "Its brief never reached an agent" said more than a missing record can see (the builder may have run against another --plan spelling); it now reads "no record shows its brief reaching an agent". The rewritten texts claimed the brief's method never arrived — but that shape is DETECTED by the brief being opened; they now state exactly that, and that the launch was not the built one. A zero-byte record (a torn write) no longer counts as built anywhere: one predicate serves the collapse, the roster loop and the chunk lookup. Entries the shell can actually run. The shebang filter now also requires the execute bit (a 0644 script passes the header check and dies on EACCES), and cli-entry consumes QWEN_CODE_LAUNCHER_PATH at stamp time — the serve/mcp fast path never reached the branch that deleted it, so a standalone daemon leaked the outer shim into every child, where a different checkout would republish it as its own entry. Inputs a PR cannot weaponize, commands an operator can run. The invariant brief interpolated the raw PR-controlled filename into the file the agent is told is the whole of its instructions — display sinks now flatten control characters and the functional read argument is JSON-quoted. Agent 7 no longer receives the review rules its own workflow forbids it (SKILL.md: deterministic commands, not code review). The verifier refuses an empty findings file — a vacuous pass that cleared the delivery floor while ruling on nothing — while the early reverse-audit round keeps it. FIX lines carry the run's real plan path instead of a `<plan>` placeholder that pastes as a shell redirection, the roster truncation hint names --file and --rules, and composed.json persists the exact verdictLine so the archived report copies rather than reconstructs it — event and cappedBy alone cannot express a presubmit downgrade. Every new behaviour is pinned: convergence, paste-all refusal, and the zero-byte collapse are mutation-checked (disabling each turns exactly its test red); the exec-bit, brief-injection, launcher-consumption and verdictLine contracts each carry a direct test. 1 236 tests across the affected suites. * fix(review): close the three paths the round-11 review found still open The Step 4/5 FIX lines still carried a literal `--plan <plan>`. Round 9 substituted the real path into compose-review's own remediation strings and check-coverage's hints, and left the one builder both Step 4/5 gaps flow through — `rebuildFix` — untouched: its output reached stderr through verificationGaps with the placeholder intact, and a literal `<plan>` pasted into a POSIX shell parses as input redirection, so the one repair round Step 6 prescribes could never run there. The push sites now substitute the plan path verificationGaps was handed, and the test that pins the fix text asserts no literal `<plan>` survives anywhere in the remediation. A lightweight cross-repo review can now be REQUIRED to run Agent 0. plan-diff takes `--pr <n> --repo <owner/repo>` — passed only after pr-context succeeds, so the pair's presence doubles as the context-availability signal — and writes the identity into the plan; the roster requires role 0 wherever the full identity is present, not only in worktree mode (fetch-pr always writes both fields, so PR-worktree behavior is unchanged). Half an identity is refused: a roster demanding an agent nobody can brief would wedge the run. SKILL.md's lightweight capture block carries the flags and the when-not-to-pass-them rule. And the path-inertness boundary is one function with a wider net: `inertPath` now flattens every control character (a terminal escape in a filename must not reach a terminal), the separator glyph, and the backtick — which could close the Markdown code span the path is rendered inside and let the tail of a PR-controlled filename run as markup in the brief the agent treats as authoritative. The roster label and launch-prompt sites that had their own narrower regexes now share it. The injection test's hostile filename gained a backtick and an ESC sequence, and asserts the rendered heading carries exactly the span's own backtick pair and no control bytes, while the JSON-quoted functional read argument still round-trips the raw path. Each fix is mutation-checked: reverting the substitution, re-gating the roster on worktree mode, and narrowing inertPath each turn exactly one test red. * fix(review): bind the receipt to what was delivered, and match what actually assigns Three review-integrity holes from the round-12 review, each with a reproduction, each fixed at the layer the reproduction named. The verify receipt could be satisfied by a partial delivery. The record was deliberately the findings-free launch block, so one key could serve every shard by the add-only rule — and that same rule let a caller build with a real findings file, launch the agent with only the recorded tail, and clear the gate while no verifier ever saw a finding. The record is now the EXACT printed prompt, findings folded in, keyed per findings-content digest (`verify--<sha>`, `reverse-audit--chunk-N--<sha>`); the delivery side collects the whole key family with the documented floor of one. Tail-only delivery matches nothing; each shard verifies against its own list; shard records no longer share a key, so none clobbers another. The injective roster matching was greedy, and greedy rejects valid assignments. With transcript T1 containing blocks A+B and T2 containing only A, first-come claiming took T1 for A and reported B missing — a compliant repair permanently capped by transcript filename order. The claim set is now a maximum bipartite matching (Kuhn's augmenting paths), seeded on the edges where the transcript also opened the requirement's brief and extended over all verbatim edges, so a requirement reports missing only when no injective completion exists at all. A rules-free rebuild could silently strip the brief. The launch prompt only points at the brief, so rebuilding a rules-bearing role without --rules left the recorded launch byte-identical while the project rules vanished from the one file the agent treats as authoritative — every delivery check kept passing. writeBrief now refuses the downgrade at the single choke point both build paths pass through, with the escape hatch named (delete the record dir to start over deliberately). All three are mutation-checked: regressing the record to findings-free, the matching to greedy, or disabling the downgrade guard each turns its own test red. 704 review tests green. * docs(review): let the docs and comments claim only what the new record design does The round-13 review caught the drift this branch's own thesis forbids: two SKILL.md sentences still described the findings-free record the previous commit retired — an orchestrator reasoning from them would conclude a findings-less delivery still matches, precisely the bypass that commit closed. Both now state the new contract: the record is the exact printed block, keyed per findings digest, and a launch that drops the list matches no record. And the matching comment claimed more than Kuhn guarantees: phase-2 augmentation can displace an opened match onto an unopened edge to enlarge the matching, so an unread flag describes the assignment, not an impossibility. The comment now says so, and why cardinality is the right thing to maximize. * docs(review): finish retiring the findings-free record from every sentence that described it Round 15 found the three survivors round 13 missed — all in code, not SKILL.md: the findingsSection docstring (all three of its clauses false since the digest-key commit), the findings field doc ('Printed, not recorded'), and the --findings --help text, which told an operator the exact opposite of what the command now does. Each now states the new contract: the findings are part of the recorded prompt, keyed per digest, and a launch that drops them matches no record. Also from the same review: the plan-path substitution uses a function replacer, so a path containing $& or $` cannot be misrendered as a replacement pattern. Practically unreachable for .qwen/tmp paths; closed because it costs four characters. * docs(review): the actually-last sentence describing the findings-free record Round 16 counted one survivor of the sweep the previous commit's title claimed complete: the acceptsFindings jsdoc in agent-briefs.ts, present-tense, whose '(see runAgentPrompt)' pointed at a function whose own comment says the opposite. It now states the digest-key contract like its siblings, and a whole-tree grep for present-tense descriptions of the retired design comes back empty. * test(review): pin the idle and missing-chunk FIX lines to the remediation channel Round-18 review: the two remediation pushes added for the every-gap-has-a-FIX rule had no test of their own — deleting either failed nothing, leaving a body disclosure whose repair could silently vanish, the exact state the channel exists to prevent. The idle-plan test now asserts the relaunch FIX; the blind-plan test, whose chunks nobody reads, now asserts the chunks-nobody-read FIX beside the blind one. Both mutation-checked: deleting each push turns exactly one test red. * fix(review): quote the plan path in every printed repair, and test the executable shebang-less shape Round-21 review, three items. The plan path is now single-quoted at all seven sites that print it into a repair command — a workspace path containing a space split the copy-pasted FIX at the space, exactly the operator moment the lines exist for; the earlier uniformity deferral ends here, uniformly. PlanDiffResult declares prNumber/ownerRepo so a refactor away from the conditional spread cannot silently drop the fields the roster's Agent-0 requirement reads. And the filter gains the test its primary target deserved: an EXECUTABLE shebang-less .js (the desktop vendored bundle shape) is rejected by the header read itself — the existing 0644 fixture never reached that branch, so a regression in the byte read would have passed every test. * fix(review): shell-quote the plan path properly — an apostrophe is not rarer than a space Round-22 review: the bare '…' wrap from the previous commit closed at an embedded apostrophe, so ~/Documents/John's Projects broke where it had worked unquoted — one breakage class traded for another instead of both closed. A shared shellQuotePath (the same '\'' dance as utils/standalone-update.ts) now serves all six repair-printing sites, and a test drives verificationGaps from a plan under an apostrophe directory, asserting the escaped form and rejecting the naive wrap. * fix(review): quote the --file selector, un-dead the spawn guard, test the half-identity Round-24/25 reviews, four small items. selectorOf now shell-quotes the --file path — the same copy-paste contract the --plan quoting just earned, on the one selector that carries a path. RULES_MARKER moves above writeBrief's JSDoc, which it had been silently stealing. The check-build-status test's reject guard was dead (execFile always delivers string stdout, so an ENOENT resolved and the empty-stdout assertion passed on a script that never ran) — it now rejects on spawn-level errors, which carry string codes, while non-zero exits still resolve. And the roster's ownerRepo guard gets the independent test it never had: a plan with prNumber but no ownerRepo requires no Agent 0, since the brief builder cannot serve half an identity.
- export the sidechannel API through the daemon barrel (selectUnrecognizedDiagnostics, UNRECOGNIZED_DIAGNOSTICS_LIMIT, DAEMON_UI_UNRECOGNIZED_DIAGNOSTIC_REASONS + types) and pin the reachability in daemon-public-surface.test.ts - restore the MAX_TEXT_BLOCK_LENGTH cap on sidechannel text, mirroring truncateText exactly (suffix fits within the cap) - ship the unrecognized reason subset as a runtime const array and route by membership, so a new reason cannot fall through to appendStatusBlock - copy the correlation fields createBase stamps (promptId, sourceRecordIds, branchRecordId, originatorClientId) onto sidechannel entries; drop the dead source/data switches - un-fuse the budget-history comment chain in scripts/build.js - update docs/developers/daemon-ui for the split routing - tests: full entry shape, text cap, block-path debugReason counterpart, and a webui malformed_payload interleave sibling so the #7012 flush-before-guard keeps a discriminating stimulus
…the routing predicate appendUnrecognizedDiagnostic left activeUserBlockId untouched while the replaced appendStatusBlock path reset it for every non-user block; a later mergeable user.text.delta with no promptId stamp (e.g. a peer client's $ <cmd> echo) then appended onto the earlier user block across the diagnostic, collapsing two user turns into one and skewing rewindTranscriptToUserTurn's kind==='user' turn indexing. Keep the reset (assistant/thought pointers stay untouched, the point of the sidechannel); witness test flip-verified red without the one-line reset. Also export isUnrecognizedDiagnosticReason from types.ts next to DAEMON_UI_UNRECOGNIZED_DIAGNOSTIC_REASONS and call it at all three routing-guard sites (reducer, provider flush condition, provider drop filter) so the #7012/#8823 guard pair classifies every debug event against one source instead of three hand-written copies.
…dechannel (QwenLM#9202) * fix(sdk): route unrecognized diagnostics onto a bounded transcript sidechannel Normalizer-classified unrecognized_event / unrecognized_session_update debug events no longer enter transcript blocks[]: they are mirrored onto a capped unrecognizedDiagnostics sidechannel instead. This stops them from finalizing a streaming assistant/thought block (which dropped a following assistant.usage frame) and from consuming the maxBlocks budget (which let repeated noise evict real conversation content). malformed_payload diagnostics and client-dispatched debug events keep their existing block semantics. * fix(sdk): align browser bundle budget * fix(sdk): close the sidechannel review round (QwenLM#8823) - export the sidechannel API through the daemon barrel (selectUnrecognizedDiagnostics, UNRECOGNIZED_DIAGNOSTICS_LIMIT, DAEMON_UI_UNRECOGNIZED_DIAGNOSTIC_REASONS + types) and pin the reachability in daemon-public-surface.test.ts - restore the MAX_TEXT_BLOCK_LENGTH cap on sidechannel text, mirroring truncateText exactly (suffix fits within the cap) - ship the unrecognized reason subset as a runtime const array and route by membership, so a new reason cannot fall through to appendStatusBlock - copy the correlation fields createBase stamps (promptId, sourceRecordIds, branchRecordId, originatorClientId) onto sidechannel entries; drop the dead source/data switches - un-fuse the budget-history comment chain in scripts/build.js - update docs/developers/daemon-ui for the split routing - tests: full entry shape, text cap, block-path debugReason counterpart, and a webui malformed_payload interleave sibling so the QwenLM#7012 flush-before-guard keeps a discriminating stimulus * fix(sdk): address round-2 sidechannel review for QwenLM#8823 - build.js: bump daemon browser bundle budget 191KB -> 192KB (195,591 bytes measured > 195,584 cap; build failed at head) - webui: narrow the observer-mode debug guard so unrecognized_* diagnostics reach the reducer sidechannel; only block-path debug events are dropped - webui: merge history-store unrecognizedDiagnostics in applyTranscriptHistory so paged-back sessions keep diagnostics - transcript: extract truncateTextAtLimit shared by the block and sidechannel truncation paths - transcript: reset unrecognizedDiagnostics on rewind alongside the sibling per-turn state resets - types: rename DaemonUnrecognizedDiagnostic.receivedAt to clientReceivedAt (matches the sibling block projection) - tests: reason-prefix conformance pin, rewind reset, narrowed guard, history pagination merge * fix(webui): avoid flushing sidechannel diagnostics * fix(sdk): preserve diagnostics across rewind * fix(webui): dedupe sidechannel history records * fix(webui): align the paging sidechannel test with the normalizer keys The paging test added in e6b40e5 failed deterministically (webui suite red, CI Test job red) for two reasons: 1. The fixtures stamped only _meta['qwen.session.recordId'], but the SDK normalizer's extractSourceRecordIds reads _meta.qwenTranscript.sourceRecordIds — no sidechannel entry ever carried sourceRecordIds, so the dedupe assertion could not pass and the new displayedRecordIds loop was never exercised by a passing test. Stamp BOTH keys, matching production replay frames (acp-bridge buildUpdateMeta) and the sibling dedupe test. 2. Cap arithmetic: LIMIT-1 live entries + 2 fresh history entries = LIMIT+1, so the newest-wins slice evicted record-old-1 which the test asserted present. Emit LIMIT-2 live events so the post-merge total lands exactly on the cap. Also correct the post-merge index assertions: history entries come first (old-1, old-2), then the deduped-once live overlap, then the first live mystery event. Suite 506/506, eslint + prettier clean. * fix(sdk): raise diagnostic sidechannel bundle budget * fix(sdk): raise the daemon browser bundle budget to 198KB and pin the diagnostics selector - The sidechannel routing + selector cost ~1037 B over the 197KB cap (bundle measured 201893 B), failing the browser-bundle size gate; bump MAX_DAEMON_BROWSER_BUNDLE_BYTES to 198 * 1024. - Fold the rebase-residue 190→191→192 KB ledger entries into the accurate 190→195→196→197→198 lineage so the next bump has one canonical history. - Add a behavioral pin for selectUnrecognizedDiagnostics: it must return the routed sidechannel itself (toBe), discriminating a `return []` or shallow-copy regression that the typeof-only surface test cannot see; flip-verified. * fix(sdk): reset the user pointer on sidechanneled diagnostics, share the routing predicate appendUnrecognizedDiagnostic left activeUserBlockId untouched while the replaced appendStatusBlock path reset it for every non-user block; a later mergeable user.text.delta with no promptId stamp (e.g. a peer client's $ <cmd> echo) then appended onto the earlier user block across the diagnostic, collapsing two user turns into one and skewing rewindTranscriptToUserTurn's kind==='user' turn indexing. Keep the reset (assistant/thought pointers stay untouched, the point of the sidechannel); witness test flip-verified red without the one-line reset. Also export isUnrecognizedDiagnosticReason from types.ts next to DAEMON_UI_UNRECOGNIZED_DIAGNOSTIC_REASONS and call it at all three routing-guard sites (reducer, provider flush condition, provider drop filter) so the QwenLM#7012/QwenLM#8823 guard pair classifies every debug event against one source instead of three hand-written copies. * fix(ci): prevent bite harness SIGPIPE --------- Co-authored-by: yiliang114 <yiliang114@users.noreply.github.com>


What this PR does
When a Web Shell tab is hidden and then restored, the SSE stream replays a burst of buffered transcript events. Previously each event was dispatched to the transcript store individually, and every dispatch copies and freezes the entire block array (O(blocks)). A burst of E events against a transcript of B blocks is therefore O(E×B) of synchronous main-thread work, which can freeze the tab for minutes or crash it on very long sessions.
This PR coalesces the live event stream into one dispatch per macrotask, so a burst collapses into a single O(B) reduction. It also caps the in-memory transcript window the client retains (the daemon stays the full source of truth), and skips the dev-only block freeze in production builds where it is pure overhead.
Why it's needed
Long-running Web Shell sessions became unusable after the tab was backgrounded: switching back triggered a multi-minute main-thread block (or a tab crash) while the buffered stream drained. The cost was quadratic in transcript size, so it got dramatically worse as sessions grew — matching the report of lag/crash with many sessions or one large session after switching the tab away and back.
Reviewer Test Plan
How to verify
Evidence (Before & After)
N/A — this is a main-thread performance fix; the effect (no multi-minute freeze on tab return) is not meaningfully capturable in screenshots. Verified via the unit burst test and the suites listed above. A live browser stress run on a large session is recommended as a final confirmation.
Tested on
Environment (optional)
Unit tests via vitest;
npm run build,npm run typecheck,npm run lint.Risk & Scope
maxBlocksonly caps the client's in-memory window; the daemon remains the authoritative full transcript.Linked Issues
Reported internally; no tracking issue.
中文说明
本 PR 做了什么
当 Web Shell 标签页被隐藏后再切回时,SSE 流会重放一批缓冲的 transcript 事件。此前每个事件都单独 dispatch 到 transcript store,而每次 dispatch 都会拷贝并冻结整个 block 数组(O(blocks))。因此 E 个事件 × B 个 block 就是 O(E×B) 的同步主线程开销,在很长的会话里会让标签页卡顿数分钟甚至崩溃。
本 PR 把实时事件流合并为每个宏任务一次 dispatch,于是一次突发突发折叠为一次 O(B) 归约。同时为客户端保留的内存 transcript 窗口设置上限(daemon 仍是完整真源),并在生产构建里跳过仅开发态使用的 block 冻结(它在生产环境纯属额外开销)。
为什么需要
长时间的 Web Shell 会话在标签页被切到后台后会变得不可用:切回来时,随着缓冲流被排空,会出现数分钟的主线程阻塞(或标签页崩溃)。该开销随 transcript 大小呈二次方增长,因此会话越大越严重——正好对应「会话很多或单个大会话、切走再切回就卡顿/崩溃」的反馈。
评审者测试计划
如何验证
证据(前后对比)
N/A——这是主线程性能修复,其效果(切回标签页不再卡顿数分钟)无法用截图有效表达。已通过单元突发测试和上述套件验证。建议再对大会话做一次真实浏览器压测作为最终确认。
测试环境
运行环境(可选)
vitest 单元测试;
npm run build、npm run typecheck、npm run lint。风险与范围
maxBlocks仅限制客户端的内存窗口;daemon 仍是完整 transcript 的权威来源。关联 Issue
内部反馈;无跟踪 issue。