fix(serve): detect stale SSE cursors across daemon restarts via epoch token; preserve turn attribution and surface compaction failures in replay - #7458
Conversation
…nd degraded-snapshot signaling (DAEMON-001/007/008)
|
@qwen-code /takeover |
|
Thanks for the PR! Template looks good ✓ Problem: observed bug with clear evidence. Issue #7457 documents three concrete defects in the daemon's event replay/reconnect chain. DAEMON-001 (stale cursor) is a provable logic bug — the numeric heuristic Direction: aligned. Daemon serve-mode reliability is core infrastructure, and these fixes harden the reconnect/replay path without changing the public contract. All changes are backward compatible in both directions (old client ↔ new daemon, new client ↔ old daemon). CHANGELOG: no direct reference to epoch tokens, but the daemon/SDK reliability series (#7386, #7400) is established prior art in this area. Size: not applicable — no core module paths ( Approach: the scope feels right. Three related fixes in one PR is reasonable here — they share the same code paths (eventBus, compactionEngine, bridge) and are part of an established batch series. Each fix is minimal: epoch is a Moving on to code review. 🔍 中文说明感谢贡献! 模板完整 ✓ 问题:已观测到的 bug,有明确证据。Issue #7457 记录了 daemon 事件重放/重连链路上的三个具体缺陷。DAEMON-001(过期游标)是可证明的逻辑缺陷——数字启发式 方向:对齐。Daemon serve 模式可靠性属于核心基础设施,这些修复加固了重连/重放路径而不改变公共契约。所有改动双向向后兼容(旧客户端 ↔ 新 daemon,新客户端 ↔ 旧 daemon)。CHANGELOG:无 epoch token 的直接引用,但 daemon/SDK 可靠性系列(#7386、#7400)是该方向的已有先例。 规模:不适用——未触及核心模块路径( 方案:范围合理。三个相关修复放在一个 PR 里在此处是合理的——它们共享相同的代码路径(eventBus、compactionEngine、bridge),且属于已确立的批次系列。每个修复都是最小化的:epoch 是一个 进入代码审查 🔍 — Qwen Code · qwen3.7-max Reviewed at |
🩺 serve daemon A/BBuilt the PR base vs this PR head ✅ No response changes against the PR base across 4 scenario(s). — Qwen Code · serve A/B |
Code ReviewIndependent proposal (before reading the diff): for the stale-cursor problem, I'd add a random token to EventBus on construction, advertise it via response headers and restore payloads, and force resync on mismatch. For attribution, carry latest stamps through compaction merges. For degradation, latch a flag on first compaction error and surface it in snapshots. Touch eventBus.ts, compactionEngine.ts, bridge.ts, bridgeTypes.ts, CLI serve routes, SDK transport/client. Comparison with the diff: the PR's approach matches my independent proposal almost exactly. No simpler path missed. No critical blockers found. A few observations:
Conventions: ESM ✓, no Real-Scenario TestingProtocol/replay-layer change — no TUI surface. Tested the daemon serve mode directly via REST/SSE endpoints. Server startupTest 1: SSE response advertises X-Qwen-Event-EpochTest 2: Matching epoch → normal resume (no resync frame)No Test 3: Stale epoch → forced resync with detail=epoch_mismatchEpoch mismatch deterministically forces resync — even though Test 4: Invalid epoch → degrades gracefullyNo resync frame — invalid token rejected, falls back to numeric heuristic. Server log: Test 5: Load response carries eventEpochTest 6: Server log for epoch mismatch resyncUnit testsAll changed test files pass:
中文说明代码审查独立方案(读 diff 之前):对过期游标问题,我会在 EventBus 构造时加一个随机 token,通过响应头和 restore 载荷下发,不匹配时强制 resync。对归属问题,在压缩合并时保留最新标记。对降级问题,在首次压缩失败时锁存标志并在快照中透出。涉及 eventBus.ts、compactionEngine.ts、bridge.ts、bridgeTypes.ts、CLI serve 路由、SDK transport/client。 与 diff 对比: PR 方案与我的独立方案几乎完全一致。没有遗漏更简路径。 未发现关键阻塞问题。几个观察:
规范:ESM ✓、无 真实场景测试协议/重放层改动——无 TUI 表面。通过 REST/SSE 端点直接测试 daemon serve 模式。
所有变更测试文件通过(共 1265 个测试)。 — Qwen Code · qwen3.7-max Reviewed at |
|
Confidence: 5/5 — clean across every stage; would merge without hesitation. The problem is real and provable: the numeric heuristic The implementation matches my independent proposal almost line-for-line. Each of the three fixes (epoch, attribution, degradation) is self-contained, adds no unnecessary abstraction, and carries comprehensive tests (1265 across the changed files). The real-scenario verification confirmed all three behaviors end-to-end: epoch header advertised on every SSE surface, mismatch forces resync with the Backward compatibility is clean in both directions — old clients never send the header, new clients talking to old daemons never learn an epoch. The WS transport's explicit "not applicable" comment is the right call. If I had to maintain this in six months, I'd thank the author: the code is well-structured, the JSDoc on public SDK fields is appropriate, and the inline comments explain why (DAEMON-001 references, backward-compat rationale) rather than narrating what. 中文说明置信度:5/5 —— 每个阶段都干净;毫不犹豫地合并。 问题是真实且可证明的:数字启发式 实现与我的独立方案几乎逐行一致。三个修复(epoch、归属、降级)各自独立,不添加不必要的抽象,并配有全面的测试(变更文件共 1265 个)。真实场景验证端到端确认了所有三个行为:epoch 头在每个 SSE 表面下发、不匹配强制 resync 并携带 双向向后兼容干净——旧客户端从不发送该头,新客户端连旧 daemon 学不到 epoch。WS transport 明确的"不适用"注释是正确的做法。 如果六个月后我要维护这段代码,我会感谢作者:代码结构良好,公共 SDK 字段的 JSDoc 恰当,内联注释解释为什么(DAEMON-001 引用、向后兼容理由)而非叙述做了什么。 — Qwen Code · qwen3.7-max Reviewed at |
qwen-code-ci-bot
left a comment
There was a problem hiding this comment.
LGTM, looks ready to ship. ✅
|
🤝 Takeover engaged: the autofix loop now manages this PR — it will address new review feedback and resolve base conflicts until the label is removed or the round cap is reached. Remove the 中文说明🤝 已接管:autofix 循环现在管理此 PR —— 将持续处理新的评审反馈与 base 冲突,直到移除标签或达到轮次上限。移除 |
|
🤖 Reviewed the latest feedback — no changes needed. Why, point by point: No action takenNo actionable feedback was found on PR #7458:
No code changes were made. 中文说明未采取任何操作PR #7458 上没有发现需要处理的反馈:
未进行任何代码更改。 Base-conflict check: no conflict with main. 🧠 Handled by Qwen Code · model/模型 |
qwen-code-ci-bot
left a comment
There was a problem hiding this comment.
Reviewed — no blockers. Suggestions are inline. 1 Suggestion-level finding(s) could not be anchored to a changed line and were dropped; nothing further to act on here.
— qwen3.7-max via Qwen Code /review
|
🤖 Addressed the latest review feedback (round 1/100). What changed, and what I pushed back on: Review feedback addressed[rc:3624409595]
|
qwen-code-ci-bot
left a comment
There was a problem hiding this comment.
No issues found. LGTM! ✅
— qwen3.7-max via Qwen Code /review
✅ Local runtime verification — LGTM, recommend mergeI built this PR's head into a real Setup: PR head DAEMON-001 — stale SSE cursor across a daemon restart (headline)I constructed the exact case the numeric heuristic misses: epoch‑1 yields an old cursor
SDK end-to-end (real DAEMON-007 — turn attribution survives compactionDrove the compiled DAEMON-008 — degraded snapshot is visibleDrove the compiled Unit suites — all 12 changed test files (PR worktree, real
|
| package | changed files | result |
|---|---|---|
@qwen-code/acp-bridge |
eventBus · compactionEngine · bridge | 544 passed |
@qwen-code/sdk-typescript |
RestSse · AcpHttp · DaemonClient · DaemonSessionClient | 423 passed |
@qwen-code/qwen-code (cli serve) |
transport · server · sse-last-event-id · multi-workspace · create-sub-session | 1180 passed |
2147 passed / 0 failed.
Verdict
All three fixes behave exactly as described; backward-compat holds in both directions (a legacy no-epoch client keeps today's numeric heuristic — Arm B; the epoch check never false-trips when it matches — Arm C); and the runtime observations have mutation teeth. LGTM — recommend merge.
Verified with a real daemon + mock OpenAI on an isolated HOME; screenshots are renders of the actual captured SSE frames / probe output.
🇨🇳 中文版
✅ 本地真实构建验证 —— LGTM,建议合并
我把本 PR head 构建成真实 dist,用真实的 qwen serve daemon(mock OpenAI 后端、隔离 HOME)在网络层端到端验证了三处修复,并做了变异(mutation)测试确认"咬合力",最后重跑了全部改动的测试文件。全部通过。
环境: PR head 60f682cd7 → 隔离 worktree 里全新 npm install → daemon 以 node packages/cli/dist/index.js serve 启动,使得 @qwen-code/acp-bridge 解析到 packages/acp-bridge/dist 里编译后的 PR 代码(而非过期 bundle)。SDK 直接跑源码。Node 22.22.2 + Bun 1.3.14。
DAEMON-001 —— daemon 重启后过期 SSE 游标(核心)
我构造了数字启发式会漏掉的场景:纪元 1 得到旧游标 Last-Event-ID=6;重启后新总线高水位涨到 14,于是 lastEventId >= nextId → 6 >= 14 → false——一个来自死纪元、却看起来像合法后缀续传的游标。随后在同一个重启后的 daemon 上,用完全相同的旧游标 Last-Event-ID: 6 以三种方式订阅 GET /session/:id/events:
| Arm | X-Qwen-Event-Epoch |
首个 SSE 帧 | 结果 |
|---|---|---|---|
| A · 修复(旧 epoch) | d2f56a15…5863d1 |
state_resync_required detail=epoch_mismatch |
✅ 全量重放 id 4→14 |
| B · 旧客户端(无该头) | — | session_update id=7 |
❌ 静默过期续传 —— 即 bug |
| C · 对照(当前 epoch) | dbad093b…0ff6 |
session_update id=7 |
✅ 正确后缀续传(不误触发) |
- daemon 同时在 stderr 打运维面包屑:
… reason=epoch_reset, detail=epoch_mismatch。 - 变异测试: 在编译后的总线里把 epoch 检查废掉(
epochMismatch=false)+ 重启,Arm A 退回0 个 resync 帧(从 id=7 续传)—— 复现了 PR 前的 bug,证明 epoch token 正是起作用的关键。 - 所有带游标的通道都下发: 202 prompt envelope、load/resume 响应体、REST SSE 响应头都携带同一个
eventEpoch,每次重启用randomUUID()重新生成、从不持久化;健康会话上replayDegraded正确缺省。/acpSSE 通道复用同一个parseEventEpochHeader与总线,由通过的 transport 单测覆盖。
SDK 端到端(从源码引入真实 RestSseTransport;Node 与 Bun 结果一致 PASS):subscribe({lastEventId, epoch: OLD}) 触发 resync(echo 路径 → 发送该头),subscribe({lastEventId}) 经 onEpoch 学到 daemon 当前 epoch(learn 路径 → 读响应头)。
DAEMON-007 —— 归属跨压缩存活
直接驱动编译后的 TurnBoundaryCompactionEngine(10/10):合并后的 agent_message_chunk/agent_thought_chunk 保留顶层 promptId/originatorClientId 和 data.sessionId;字段级合并保证 chunk 分别只带一个字段时两者都不丢;折叠的 tool_call 采用 latest-wins;对照组(源事件无归属)不产生多余字段。变异(去掉重新盖章)→ 4 项归属检查失败。
DAEMON-008 —— 降级快照可见
用会抛错的压缩引擎驱动编译后的 EventBus(7/7):首次 ingest 抛错时 onCompactionError 恰好触发一次,快照锁存 degraded:true,publish() 仍永不抛错;健康引擎从不降级(无误报)。变异(把锁存改成空操作)→ 2 项降级检查失败(且回调会在每次抛错时都触发,反证一次性锁存的作用)。
单测套件 —— 全部 12 个改动测试文件(PR worktree,真实 npm install)
| 包 | 改动文件 | 结果 |
|---|---|---|
@qwen-code/acp-bridge |
eventBus · compactionEngine · bridge | 544 passed |
@qwen-code/sdk-typescript |
RestSse · AcpHttp · DaemonClient · DaemonSessionClient | 423 passed |
@qwen-code/qwen-code(cli serve) |
transport · server · sse-last-event-id · multi-workspace · create-sub-session | 1180 passed |
2147 passed / 0 failed。
结论
三处修复行为与描述完全一致;双向向后兼容成立(旧的无 epoch 客户端保持现有数字启发式 —— Arm B;epoch 匹配时绝不误触发 —— Arm C);运行时观测均具备变异咬合力。LGTM —— 建议合并。
|
🤖 AutoFix could not start evaluation — it crashed or timed out before reading the feedback, so no fix was attempted. This PR is now marked terminal and future scans (including forced dispatch) will skip it. To recover: delete this bot's terminal AutoFix failed before producing a verified commit (the run crashed or timed out before it could explain why). Run log: https://github.com/QwenLM/qwen-code/actions/runs/29889165690 🧠 Handled by Qwen Code · model/模型 |
|
⏸️ Takeover paused: this PR reached its round cap (100/100). Comment 中文说明⏸️ 托管已暂停:本 PR 达到轮次上限(100/100)。评论 |
|
@qwen-code /resolve |
…outing Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>
|
Qwen Code resolved the merge conflicts and pushed the branch update. Merge Resolution Summary — PR #7458Root causeThe base branch ( Textual or semanticTextual. The two features touched adjacent lines of the same call expression but did not alter the same logic path. The resolution is a clean composition: the ternary dispatch from main wraps the subscribe call, and the Resolved block: const iterable = virtualKey
? await deps.virtualSubagentSessions?.subscribe(runtime, sessionId, {
signal: abort.signal,
lastEventId,
...(maxQueued !== undefined ? { maxQueued } : {}),
})
: runtime.bridge.subscribeEvents(sessionId, {
signal: abort.signal,
lastEventId,
...(eventEpoch !== undefined ? { epoch: eventEpoch } : {}),
...(maxQueued !== undefined ? { maxQueued } : {}),
...(snapshot ? { snapshot: true } : {}),
});
if (!iterable) {
res.status(404).json({ error: 'Subagent session not found', code: 'session_not_found', sessionId });
return;
}What is load-bearing
What you could not verify
中文说明合并解决摘要 — PR #7458根因基线分支 ( 文本冲突还是语义冲突文本冲突。 两个功能修改了同一调用表达式的相邻行,但没有修改同一逻辑路径。解决方案是干净地组合:main 的三元分发包裹 subscribe 调用,PR 的 关键约束
无法验证的内容
|
74faba9
qwen-code-ci-bot
left a comment
There was a problem hiding this comment.
Not reviewed: reverse audit — an auditor ran and opened its brief, but no agent was launched with the prompt the CLI built — the launch was written by hand, and what the agent was actually asked is not what this skill certifies.
— qwen3.7-max via Qwen Code /review
Await each session/load reply frame before opening the session stream so the GET cannot race conn.ownSession() into a 403; addresses the review Critical on the deg-0 arm.
|
🤖 Reviewed the latest feedback — no changes needed. Why, point by point: No action requiredAll review findings on PR #7458 have already been addressed in commits Finding triage
中文说明无需操作PR #7458 上的所有评审意见已在提交 意见分类
Base-conflict check: no conflict with main. 🧠 Handled by Qwen Code · model/模型 |
Code review — DAEMON-001 / 007 / 008 (epoch token, attribution, degraded replay)Reviewed the full diff (30 files, +1710/-30). This is a clean, well-scoped reliability PR. No blocking issues found — the analysis below is verification notes plus a few optional observations. What it doesThree independent daemon-replay hardenings, each gated so it's additive and bidirectionally backward-compatible:
Correctness — verified
Backward compatibilityGenuinely bidirectional: old client → new daemon keeps the numeric heuristic (no header sent); new client → old daemon never learns an epoch and falls back the same way. New fields/headers are optional and additive; Test coverage — excellentNegative and boundary cases are all present: invalid epoch header degrades to "not provided" and logs without aborting the stream, 64/65-char boundary, no-cursor ignores epoch, healthy path stays undegraded, older-daemon-omits-epoch, header-learned supersedes seeded, and the virtual-subagent path skips the throwing Security
Optional observations (non-blocking)
Nice work — the comment density and DAEMON-xxx traceability match the surrounding code, and the test matrix is thorough. |
qwen-code-ci-bot
left a comment
There was a problem hiding this comment.
Reviewed — no blockers. Suggestions are inline.
— qwen3.7-max via Qwen Code /review
Cross-origin SSE clients must send the epoch header through preflight and read it from the response, or stale-cursor detection (DAEMON-001) is silently disabled for every CORS client.
|
🤖 Reviewed the latest feedback — no changes needed. Why, point by point: Review feedback assessment — PR #7458No changes made this round. All feedback has been triaged below. Inline suggestion: CORS headers for
|
qwen-code-ci-bot
left a comment
There was a problem hiding this comment.
No issues found. LGTM! ✅
— qwen3.7-max via Qwen Code /review
…poch-compaction # Conflicts: # packages/sdk-typescript/scripts/build.js
|
🤖 Addressed the latest review feedback (round 2/100). What changed, and what I pushed back on: Review feedbackNo actionable feedback was present in this round: the feedback file contained no reviews, no inline comments, no issue-level comments, no failed checks, and no still-red checks. There were therefore no findings to triage or implement, and Action taken: merged
|
qwen-code-ci-bot
left a comment
There was a problem hiding this comment.
No issues found. LGTM! ✅
— qwen3.7-max via Qwen Code /review
🔬 Maintainer E2E Verification ReportBranch: 1. Build & Typecheck
2. Unit Tests (all three affected packages)
* The single timeout ( 3. E2E Real Daemon Tests (9/9 passed)Started a real
Key evidenceEpoch advertised on every cursor surface: Stale epoch deterministically forces resync (the core DAEMON-001 fix): Summary: 4. VerdictAll three DAEMON audit items (001 epoch-token restart detection, 007 compaction attribution, 008 degraded-snapshot signaling) are covered by the unit tests. The E2E run confirms the epoch token flows correctly through every client-facing surface (load, SSE header, prompt 202) and that a stale epoch deterministically forces a full resync — including across a real daemon restart. Backward compatibility is preserved (no epoch header → legacy numeric heuristic). Recommendation: ready to merge ✅ 中文版本🔬 维护者 E2E 验证报告分支: 1. 构建与类型检查
2. 单元测试(三个受影响包)
* 唯一超时的测试( 3. E2E 真实 Daemon 测试(9/9 通过)在端口 14170 启动真实
4. 结论三个 DAEMON 审计项(001 epoch-token 重启检测、007 压缩归属保留、008 降级快照信号)均有单测覆盖。E2E 运行确认 epoch token 正确流经所有客户端可见通道(load、SSE 头、prompt 202),且过期 epoch 确定性地强制全量 resync——包括真实 daemon 重启场景。向后兼容得到保持(不带 epoch 头 → 回落到数字启发式)。 建议:可以合入 ✅ |
✅ Local runtime verification (round 2, current head) — LGTM, recommend mergeRe-verified this PR against the current head Setup: PR head DAEMON-001 — stale SSE cursor across a daemon restart (headline)Drove the real
DAEMON-007 — turn attribution survives compactionDrove the compiled DAEMON-008 — degraded snapshot is visibleDrove the compiled Regression fixed — virtual-subagent REST SSE (my earlier High finding)
Changed-suite totals — fresh real build
2200 passed / 0 failed. ( VerdictAll three fixes behave exactly as described on the current head; backward-compat holds in both directions (Arm B legacy path unchanged, Arm C never false-trips); the headline has mutation teeth; and the previously-flagged virtual-subagent regression is fixed and test-locked. LGTM — recommend merge. Verified with the real daemon route + the compiled 🇨🇳 中文版✅ 本地真实构建验证(第 2 轮,当前 head)—— LGTM,建议合并针对当前 head 环境: PR head DAEMON-001 —— daemon 重启后过期 SSE 游标(核心)用真实
DAEMON-007 —— 归属跨压缩存活驱动编译后的 DAEMON-008 —— 降级快照可见用会抛错的压缩引擎驱动编译后的 回归已修 —— virtual-subagent REST SSE(我此前的 High 项)
改动测试套件总计 —— 全新真实构建
2200 通过 / 0 失败。( 结论三处修复在当前 head 的行为与描述完全一致;双向向后兼容成立(Arm B 旧路径不变、Arm C 绝不误触发);核心项具备变异咬合;此前标记的 virtual-subagent 回归已修复并被测试锁定。LGTM —— 建议合并。 |
Code Review —
|
… token; preserve turn attribution and surface compaction failures in replay (#7458) * fix(daemon): epoch-token restart detection, compaction attribution, and degraded-snapshot signaling (DAEMON-001/007/008) * fix(acp-bridge): field-level turn attribution merge and replayDegraded bridge test (#7458) * fix(serve): skip bus epoch lookup for virtual subagent SSE streams (#7458) The REST SSE route looked up the bus epoch for every session id, but virtual subagent sessions ride their own bus and their compound ids are not in the bridge's byId map, so the lookup threw and aborted the subscription — breaking subagent event streams. Skip the lookup for the virtual path and degrade a torn-down real session to a headerless stream (mirrors the /acp route). Also bumps the daemon browser SDK bundle budget (167KB -> 168KB) for the epoch fields and declares eventEpoch on DaemonSession so the create/attach path drops its inline type cast. * fix(serve): stamp eventEpoch on accepted continuations and surface replayDegraded in the SDK (#7458) Address three review suggestions: - POST /session/:id/continue now returns eventEpoch alongside lastEventId, mirroring the prompt 202 envelope so continuation-seeded SSE cursors detect daemon restarts (DAEMON-001) - DaemonSessionClient exposes replayDegraded from the load response so SDK consumers can prefer the full transcript over a degraded snapshot - add /acp dispatch-level regression test for the degraded-snapshot stderr breadcrumb (fires only when snapshot.degraded is set) * test(cli): fix load-reply race in the degraded-breadcrumb transport test Await each session/load reply frame before opening the session stream so the GET cannot race conn.ownSession() into a 403; addresses the review Critical on the deg-0 arm. * fix(serve): allow and expose X-Qwen-Event-Epoch in CORS headers Cross-origin SSE clients must send the epoch header through preflight and read it from the response, or stale-cursor detection (DAEMON-001) is silently disabled for every CORS client. --------- Co-authored-by: qwen-code-bot <qwen-code-bot@users.noreply.github.com> Co-authored-by: qwen-code-dev-bot <qwen-code-dev-bot@users.noreply.github.com> Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> Co-authored-by: Qwen Autofix <qwen-autofix[bot]@users.noreply.github.com>
) * fix(cli): correct queued message display style and ordering Mid-turn steer messages (user input queued while the model is responding) had two display bugs: 1. They rendered with notification styling (● icon) instead of user-input styling (> prefix) because accept() added them to UI history as MessageType.NOTIFICATION. 2. They appeared below the model's reply because accept() was only called in the finally block after the entire response stream completed, appending the user message after all model response items. Fix: use MessageType.USER with sentToModel: true for steer messages, and settle the steer input on the first stream event (after the user-content push lands but before model-response events are committed to UI history). Pass steer inputs through to recursive sendMessageStream calls so all takeSteerInput paths benefit from early settlement. Add a WeakSet guard to settleSteerInput for idempotency across recursive invocations. * test(core): add ordering test for early steer settlement Verify that accept() is called after the first stream event is pulled but before subsequent events reach the consumer, pinning the settle-before-content timing that ensures queued user messages render above the model's reply. * fix(cli): use sentToModel: false for steer messages, address review - Use sentToModel: false instead of true: steer messages are injected into an existing tool-result turn, not standalone user turns. sentToModel: true would make isRealUserTurn() count them as real turns, inflating the rewind turn index. - Remove unnecessary as HistoryItemWithoutId cast. - Add post-cleanup assertion in ordering test to verify the WeakSet guard prevents double-settlement. * fix(cli): align resumed mid-turn steer display with live session (#7381) Resume path now renders mid_turn_user_message as MessageType.USER with sentToModel: false, matching the live-session styling. Add a comment documenting the intentional sentToModel: false choice. * fix(cli): exclude steer messages from user-turn filters (#7381) Steer messages (sentToModel: false) were counted as real user turns by five downstream consumers that filter on type === 'user' without checking sentToModel, breaking cancel auto-restore, telemetry turn count, prompt recall, away-recap thresholds, and resume collapse boundaries. Add sentToModel !== false guards at each site. * test(cli): add coverage for sentToModel !== false guards (#7381) * test(cli): add coverage for sentToModel !== false guard in input-history filter (#7381) * test(cli): add coverage for sentToModel !== false guard in YOLO turn-count telemetry (#7381) * fix(cli): restore corrupted docs and classify steer items as synthetic (#7381) * fix(docs): restore corrupted autogenerated input names in GitHub Action docs (#7381) * fix(cli): deduplicate findLastUserItemIndex and add steerInput forwarding test (#7381) * fix(cli): keep code-block copy numbering continuous across steer items (#7381) * test(core): add Hook continuation steerInput forwarding test Verify that steerInput is forwarded through the Stop-hook continuation path and settled early on the first content event of the continuation turn, matching the existing Steer continuation coverage. * fix(cli): sync selection test fixtures with ink FrameCell/ReadonlyFrame types (#7381) * fix(core): align cron day wildcard semantics (#7464) Co-authored-by: destire-mio <248462155+destire-mio@users.noreply.github.com> * feat(core): keep completed background agents resident (#7426) * feat(core): keep background agents resident * fix(core): harden background continuation boundaries * docs(core): move per-spawn cleanup comment to subagentDispose The comment describing the per-spawn cleanup (which stays undefined on the fork-resume path) had drifted above the launchModel declaration, where it no longer applied and could mislead readers. Relocate it to the subagentDispose assignment in the non-fork branch it actually documents. * fix(core): close finishing window and release resident on error in background GOAL path - Non-worktree GOAL completion drained the message queue but never called registry.beginFinishing(), unlike the worktree path. A send_message racing the terminal transition could be accepted (status still running, finishingAgents empty) and then orphaned by complete(). Call beginFinishing() after the empty drain to reject the racing message instead. - The completion catch block never reset keepResident, so a throw from patchAgentMeta/registry.complete left the runtime resident but finalized as failed — a zombie that cleanupRuntime never reclaimed. Reset keepResident in the catch so the finally block disposes it. --------- Co-authored-by: Claude <noreply@anthropic.com> * ci(autofix): continue environment-specific fixes (#7444) * ci(autofix): continue environment-specific fixes * docs(autofix): align verification wording * docs(autofix): require bundle before integration tests * docs(autofix): scope surrogate verification rules * docs(autofix): require focused tests before integration checks * docs(autofix): clarify review verification guidance * fix(acp-bridge): close prompt-terminal follow-ups from the PR #7400 self-review (#7453) * fix(acp-bridge): close prompt-terminal follow-ups from PR #7400 self-review Keep a removed RUNNING prompt visible to the teardown flush via a removed flag so its terminal still publishes when the session closes before the agent cooperates; gate broadcastTurnError's session turn-state mutation to running prompts; propagate the typed PromptDeadlineExceededError from the pre-dispatch abort check; document the deadline FIFO-release overlap trade-off, the trailing prompt_cancelled after flush, and the result.then/finally ordering invariant; route the dedup log to the debug channel; drop the prompt-deadline re-export that pulled the bridge into a leaf module. Fixes #7451 * test(acp-bridge): cover promote-then-remove-then-settle duplicate completed guard (#7453) --------- Co-authored-by: Qwen Code Bot <qwen-code-bot@users.noreply.github.com> * fix(core): strip Qwen-internal daemon secrets from agent-spawned child env (#7256) * fix(core): strip Qwen-internal daemon secrets from agent-spawned child env Shell subprocesses (and the monitor tool and stdio MCP servers) inherited the full daemon process.env, including QWEN_SERVER_TOKEN (the serve-daemon bearer credential), so an agent-run command like printenv QWEN_SERVER_TOKEN could read an internal secret. Add a shared sanitizeChildEnv() that removes Qwen-internal daemon/server tokens (QWEN_SERVER_TOKEN, QWEN_DAEMON_TOKEN) before spawning, and apply it at the shell child_process + PTY paths, monitor.ts, and the mcp-client stdio transport. The denylist is deliberately narrow: it does NOT strip third-party credentials (GH_TOKEN, AWS_*, NPM_TOKEN, ...) that real shell workflows legitimately inherit -- only Qwen-internal secrets. Exported from the package root so the desktop denylists can consolidate onto it later. Fixes #6601. * test(core): cover daemon-secret stripping on monitor and mcp-client spawn sites * test(core): replace process.env instead of mutating in shell sanitization tests The file restores process.env by reference in afterEach, so in-place key mutations leaked into later tests. Use the replacement pattern already used by setupConflictingPathEnv. * docs(core): align JSDoc @param names with actual function signatures (#7492) Fix 6 instances where JSDoc @param tags had drifted from their corresponding function signatures — parameters were renamed, removed, or undocumented over time but the doc blocks were not updated. Closes #7446 * feat(serve): support forced MCP reconnects (#7488) * feat(serve): support forced MCP reconnects * test(serve): cover forced MCP reconnect options --------- Co-authored-by: 克竟 <dingbingzhi.dbz@alibaba-inc.com> * fix(cli): insert newline on Shift+Enter and stop streaming thinking-block flicker (#7397) * fix(cli): re-push Kitty keyboard flags onto the alternate screen in VP mode In VP mode the app renders on the alternate screen (`alternateScreen: true`), but the Kitty keyboard progressive-enhancement flags were pushed only once at startup on the main screen. The Kitty spec tracks these flags per screen buffer, so the alternate screen's stack stays empty and the terminal never reports modifiers: Shift+Enter arrives as a bare Enter (submit) or, when the terminal emits an ESC-prefixed variant, as an orphaned Escape that trips the empty-buffer double-Esc rewind prompt — so Shift+Enter can never insert a newline in VP mode even on Kitty-capable terminals (e.g. cmux). Re-push the flags onto the alternate screen right after Ink enters it (Ink writes the enter-alt-screen sequence synchronously inside render(), so the push is correctly ordered). Ink discards the alternate screen and its flag stack on unmount, leaving the startup main-screen push balanced by the existing disableKittyProtocol() on cleanup. Generated with AI Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> * fix(cli): stabilize streaming thinking block height to stop flicker The pending "Thinking…" block renders the tail of the reasoning stream in a content-sized box. As the model emits paragraph separators, a blank line enters and leaves the tail window (and `trimEnd` drops trailing blanks), so the visible line count oscillates and the block flickers 2→3→5 rows during streaming. Track the tallest height the block has reached for the current thought and never render fewer rows than that (capped at the streaming window size), padding at the top so the newest line stays pinned to the bottom. The tracker resets when streaming ends or when the buffer shrinks (a new thought replaced it), so height is monotonic within a thought without leaking across thoughts. Generated with AI Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> * fix(cli): decode xterm modifyOtherKeys Shift/Ctrl/Alt+Enter so it inserts a newline Terminals such as Ghostty report Shift+Enter as the xterm modifyOtherKeys sequence `ESC [ 27 ; <mods> ; <key> ~` (e.g. `ESC [ 27 ; 2 ; 13 ~`) when the Kitty keyboard protocol is not negotiated — which is the default, since Kitty detection does not always succeed. Two bugs kept this from inserting a newline: 1. The CSI-u parser read the leading `27` marker as the key code (matching the Escape key code 27) instead of the real key code in the third parameter, so with Kitty enabled Shift+Enter was mistaken for Escape and tripped the double-Esc rewind prompt. 2. The reassembly path that stitches readline's shredded CSI fragments back together was gated behind `kittyProtocolEnabled`, so with Kitty disabled the `ESC [ 27 ; 2 ;` head plus the stray `13~` tail leaked into the composer as literal text and no newline was inserted. Decode the third parameter as the real key code for the `27;…~` form, and route those sequences through the reassembly buffer even when Kitty is off (only the `ESC [ 27` marker opts in, so keys readline already parses cleanly are untouched). Shift/Ctrl/Alt+Enter now insert a newline in both VP and non-VP mode regardless of Kitty negotiation. Generated with AI Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> * fix(cli): anchor VP viewport to the top until a conversation turn exists On a fresh VP-mode session the virtualized list holds the banner plus startup notices (tips / MOTD / info), so it is longer than one item. Keying the initial scroll anchor off list length alone selected scroll-to-end, which pinned the banner to the bottom of the full-height viewport and left the top half of the screen blank. Anchor to the top until there is an actual conversation turn (a user/user_shell history item or a pending response), then resume scroll-to-end so the latest output stays in view. Startup notices no longer count as content that forces bottom alignment. Generated with AI Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> * fix(cli): stabilize streaming thinking window against availableTerminalHeight drift The grow-only streaming thinking window still flickered because its line cap was derived from availableTerminalHeight. While a thought streams the terminal keeps constrainHeight on, so availableTerminalHeight (and the derived maxLines) drifts up and down as sibling pending content grows, and the grow-only clamp `min(maxLines, …)` shrank the block whenever it dipped. Use a constant window height (MAX_STREAMING_THINKING_VISUAL_LINES) for the pending window instead. The window is only a few lines, so a fixed cap cannot meaningfully overflow (VP scrolls anyway), and the height stays stable while still growing monotonically within a thought. Generated with AI Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> * Revert "fix(cli): anchor VP viewport to the top until a conversation turn exists" This reverts commit fbe86a9e159b75ea1f5b689cc327599c9dc91090. * fix(cli): guard modifyOtherKeys detection against keypresses without a sequence The modifyOtherKeys prefix check ran on every keypress, but some synthetic keypresses (and the useKeypress test harness) emit a key with no `sequence`, so `key.sequence.startsWith(...)` threw an unhandled rejection. Use optional chaining so a missing sequence is simply not a modifyOtherKeys start. Generated with AI Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> * test(cli): mock pushKittyProtocolFlags in gemini.test.tsx kitty mock The kittyProtocolDetector mock omitted the newly added pushKittyProtocolFlags export. Add it so the mock stays in sync with the real module and a VP-mode startup path exercised through this suite cannot hit an undefined call. Generated with AI Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> --------- Co-authored-by: 秦奇 <gary.gq@alibaba-inc.com> Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> * fix(web-shell): open singleton subagent details (#7495) Co-authored-by: ytahdn <ytahdn@gmail.com> * fix(web-shell): avoid redundant git status requests (#7496) Co-authored-by: ytahdn <ytahdn@gmail.com> * fix(agent): ignore empty working_dir placeholders (#7343) * fix(agent): ignore empty working_dir placeholders * test(agent): align empty working_dir expectations * feat(prompts): allow overriding core identity via QWEN_SYSTEM_IDENTITY_MD (#7478) * feat(prompts): update prompts.ts for QWEN_SYSTEM_IDENTITY_MD * feat(prompts): update prompts.test.ts for QWEN_SYSTEM_IDENTITY_MD * fix(prompts): address CR on QWEN_SYSTEM_IDENTITY_MD Keep getDefaultCoreIdentitySentence private, fail loud on path resolution errors, use trimEnd, and resolve identity only on the default-prompt branch. * test(prompts): align identity override tests with CR feedback Sample default identity from live prompt, cover trimEnd trailing whitespace, and assert homedir resolution failures throw. --------- Co-authored-by: 易良 <1204183885@qq.com> * fix(cli): yield to single-slot background agents (#7258) Co-authored-by: hogeheer <267467744+hogeheer499-commits@users.noreply.github.com> * docs(autofix): require evidenced pre-commit verification, not a bare "verified" (#7486) * docs(autofix): require evidenced pre-commit verification, not a bare "verified" The skill already said to run build/typecheck/lint/Vitest before committing, but softly — and #7408 committed a fix with a TS error the gate then rejected while its summary claimed "verified all 3 commits". A self-assessment the gate contradicts wastes a whole round. Strengthens the address-review contract from "run the checks" to: - actually run them, do not assert them from reading the diff; - if typecheck or a touched-package test fails, do NOT commit — treat the feedback as unresolved (failure.md); - end address-summary.md with a `## Verification` section listing each command run and its result; a bare "verified" is not acceptable. The framing is structural, not etiquette: the deterministic gate re-runs the same commands and discards the round on any failure, so skipping them only moves the rejection later. Pinned by a test so it cannot soften back. This is the checkable half of "audit before committing" — the undirected/reverse-audit-until-clean practice does not transfer to an unsupervised agent (no verifiable stopping condition, and it worsens the timeouts seen on large PRs), but "run the gate's own checks first and show the evidence" does. * fix(autofix): clarify Verification section precedes collapsed Chinese translation (#7486) --------- Co-authored-by: wenshao <wenshao@example.com> Co-authored-by: qwen-code-ci-bot <qwen-code-ci-bot@users.noreply.github.com> Co-authored-by: Qwen Code Bot <qwen-code-bot@users.noreply.github.com> * feat(autofix): stop a PR that fails to push for N rounds in a row (#7482) * feat(autofix): stop a PR that fails to push for N rounds in a row Under takeover the round cap is 100, which is right for a PR that needs many PRODUCTIVE rounds. It is wrong for one that fails every round: #6723 ran 7 consecutive failed rounds (3 agent timeouts at 50 min, 4 gate rejections whose fix broke tests) over 8 hours, heading for round 100, because it is a 5700-line, 47-file, 5-day-old PR racing a fast-moving main — every round re-resolves a conflict it cannot finish or that fails the gate. Retrying at the same per-round budget will not converge; a human has to rebase or split it. Adds CONSECUTIVE_FAILURE_CAP (5), distinct from the total round cap. The handoff step already runs only when a round did NOT push, so it counts the unbroken run of prior failure markers — stopping at the first push ("Addressed the latest review feedback") or legitimate no-op ("no changes needed"), either of which proves progress and resets the streak. At the cap it forces the terminal round even under takeover, with a handoff that names the real fix (rebase/split, then /retry). Cause- agnostic: a timeout and a gate rejection both count. * fix(autofix): address review feedback on consecutive-failure circuit breaker (#7482) - Fix misleading comment: the walk is oldest-first (API order) with reset-on-success, not newest-first with early stop - Prefer the already-fetched ic.json over a redundant gh api call, falling back to the API only when the file is missing - Filter eval markers by re-arm window (win=) so pre-re-arm failures do not immediately re-terminate a re-armed PR - Add test coverage for the MARK_ROUND == MAX_ROUNDS guard and for window-scoped streak counting * fix(autofix): exempt transient model errors from consecutive-failure breaker (#7482) --------- Co-authored-by: wenshao <wenshao@example.com> Co-authored-by: qwen-code-dev-bot <qwen-code-dev-bot@users.noreply.github.com> * feat(core): restore background agent roster (#7459) * feat(core): restore background agent roster * fix(web-shell): add list_agents to TOOL_DISPLAY_NAMES The new list_agents core wire tool was added to core's ToolNames but not to the web-shell TOOL_DISPLAY_NAMES map, causing toolFormatting.drift.test.ts to fail (expected ['list_agents'] to deeply equal []). Add the missing 'ListAgents' display-name entry so the browser panel shows a friendly name instead of the raw wire name and the drift guard passes. * fix(cli): reload old-session background agents on failed resume rollback When /resume fails after core has swapped but before the UI swap, the catch block rolls core back to the old session via startNewSession(oldSessionId). However the forward path already called resetBackgroundStateForSessionSwitch, which cleared the old session's in-memory background agents. The rollback did not reload them, so list_agents returned empty for the old session (whose sidecars are still on disk) until the next process start or successful resume. Reload the old session's paused background agents after rolling core back, so the restored roster matches on-disk state. Placed after startNewSession so the loadPausedBackgroundAgents current-session guard is satisfied; best-effort via .catch so it never blocks the rollback path. * fix(web-shell): add zh translation for list_agents tool name The toolFormatting test 'has a zh translation for every tool in the display-name map' failed with expected ['list_agents'] to deeply equal [] because list_agents was added to TOOL_DISPLAY_NAMES without a matching toolName.list_agents zh-CN entry. Add the translation to restore parity. * fix(cli): resolve CI failures for background-agent roster restore - Add toolDisplayName.ListAgents translations (en, zh, zh-TW, ca) so the new list_agents tool has a zh entry; fixes i18n/index.test.ts. - Add loadPausedBackgroundAgents and consumePendingRecoveredAgentsNotice to the acpAgent worktree test config mock, which loadSession now calls via #restoreBackgroundAgentsOnResume; fixes acpAgent.worktree.test.ts. * refactor(core): extract incompatible-isolation blocked reason to a const Move the incompatible-isolation blocked-reason string out of an inline literal into a module-level INCOMPATIBLE_ISOLATION_BLOCKED_REASON const, matching its four sibling reasons so the text is discoverable by constant-name grep and edited alongside the others. * fix(core): preserve retained activity state on failed agent revive Address review feedback on the background-agent roster restore: - On a failed completed-agent revive, restore UI state with a non-empty guard instead of `??`. Because `restorePausedEntry` resets the paused entry's `recentActivities` to `[]`, the previous `failedEntry?.field ?? completedEntry.field` kept that empty array and dropped the pre-revive snapshot (the UI Progress section rendered empty). Applied consistently to pendingMessages, recentActivities, and pendingApprovals. Add regression coverage for previously untested paths: - failed revive preserves pre-revive recentActivities - terminal-agent cap admits only the newest MAX_RETAINED_TERMINAL_AGENTS completed sidecars on restore - /resume rollback reloads the old session's background agents - headless resume prepends the recovered-agents notice to the prompt * test(cli): cover interrupted-turn continuation not consuming recovered-agents notice Add ACP and headless regression tests asserting an interrupted-turn continuation does not consume the one-shot recovered-agents notice (the !isContinue / !continueInterrupted guards), so it is delivered on the user's next ordinary prompt. Mirrors the existing slash-command coverage. --------- Co-authored-by: Claude <noreply@anthropic.com> * feat(cli): support custom skill directories via settings (#7395) * feat(cli): support custom skill directories via settings (#7394) Add skills.directories setting that accepts an array of additional directory paths to scan for skills (SKILL.md files). Paths support ~ expansion. Directories are scanned recursively at user level, after the default ~/.qwen/skills/ directory. Example settings.json: { "skills": { "directories": ["~/.agent/skills", "~/.claude/skills"] } } Changes: - settingsSchema.ts: add skills.directories array setting - core Config: add customSkillDirs param and getCustomSkillDirs() - SkillManager: append custom dirs to user-level skill base dirs - CLI config: read skills.directories and pass to core Config * fix(cli): regenerate settings schema for skills.directories (#7394) * fix(core): address review feedback for custom skill directories (#7395) - Use optional chaining for getCustomSkillDirs() to prevent TypeError on partial Config mocks (workspace-skill-management, workspace-skills-status) - Reuse expandHomeDir utility instead of inline tilde expansion - Fix inaccurate 'scanned recursively' wording to 'one level deep' - Correct JSDoc: paths are raw, expansion happens in SkillManager - Trim whitespace from custom dir entries in CLI layer - Add tests for custom dir expansion, dedup, and partial config safety * fix(core): address review feedback for custom skill directories (#7395) * fix(core): address review feedback for custom skill directories (#7395) * test(core): add relative path resolution test for custom skill dirs (#7395) * fix(cli): add Array.isArray guard for skills.directories and safe mode test (#7395) * fix(skills): address review feedback on custom skill directories (#7395) - Add bare mode test for skills.directories guard - Include resolved absolute path in relative directory warning - Clarify that dedup applies to default user dirs, not bundled skills - Regenerate settings schema --------- Co-authored-by: qwen-code-dev-bot <qwen-code-dev-bot@users.noreply.github.com> Co-authored-by: qwen-code-ci-bot <qwen-code-ci-bot@users.noreply.github.com> Co-authored-by: Qwen Code Bot <qwen-code-bot@users.noreply.github.com> Co-authored-by: Qwen Code Autofix <qwen-code-autofix@users.noreply.github.com> * fix(core): add image modality support for qwen3.8-max and kimi-k3 models (#7491) * fix(core): add image modality support for qwen3.8-max models qwen3.8-max-preview supports image input but was falling through to the catch-all text-only rule because no pattern matched it. This caused the vision bridge to unnecessarily transcribe images via a secondary model instead of sending them directly to the primary model. * fix(core): also add image modality for kimi-k3 Kimi K3 officially supports image + video input but was falling through to the catch-all text-only rule, same issue as qwen3.8-max. * fix(dingtalk): preserve non-bot mention context (#7473) * fix(dingtalk): preserve non-bot mention context * test(dingtalk): cover plural mentions, staffId fallback, and edge cases (#7473) --------- Co-authored-by: qwen-code-dev-bot <qwen-code-dev-bot@users.noreply.github.com> * fix(core): harden the usage salvage around session deletion (#7425) Post-merge review follow-ups on #7391 (three findings): - Salvage the archived transcript in the active-branch deletion too: when both copies co-exist (an interrupted archive) and the fresh active transcript carries no telemetry, the archived copy holds the session's usage history and was deleted unsalvaged. The dedup guard makes the extra call a no-op whenever the active copy already wrote. - Enforce the "never blocks deletion" contract at the call site: a salvageUsageBestEffort wrapper catches and warns, so the guarantee is structural rather than an implementation detail of persistUsageBeforeTranscriptDeletion. The new failure-tolerance test (salvage rejects -> deletion still succeeds) fails without the wrapper — the bare await let the rejection escape through removeSessionFiles' rethrowing catch. - Clear the salvage module mock in beforeEach so the wiring test's invocationCallOrder assertions can never read stale calls. Co-authored-by: Claude Fable 5 <noreply@anthropic.com> * fix(core): make fork subagents discoverable (#7460) * test(core): cover Shell truncation without an artifact (#7470) Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> * fix(ci): autofix route checks existing labels on non-trigger label events (#7481) * fix(ci): autofix route checks existing labels on non-trigger label events When triage adds multiple labels in sequence, per-issue concurrency cancels earlier runs. If the last label is not a trigger label (e.g. scope/build-system), the surviving run skips the issue phase even though the issue already has autofix/approved + status/ready-for-agent. Before ignoring a non-trigger label event, check ISSUE_LABELS_JSON for both required labels. If present and the issue is open, proceed with the issue phase. Trust was already established when the trigger labels were applied (both require triage+ permission). * fix(ci): require trusted sender for label fallback * feat(cli): preserve semantic text when copying VP selections (#7286) * docs(cli): define semantic copy fidelity scope * docs(cli): address semantic frame review gaps * docs(cli): preserve soft-wrap source separators * feat(cli): preserve semantic selection copy * fix(cli): address semantic copy review findings * fix(cli): preserve clipped semantic boundaries * fix(cli): limit separator carrier joiner to visible width in wrap metadata The greedy /\s+/ match in wrapTextWithMetadata could capture more source whitespace than the separator carrier row actually consumed (e.g. a tab following a space), causing duplicated whitespace in semantic copy. Limit the match to visibleLine.length characters and add a mixed space/tab regression test. --------- Co-authored-by: 秦奇 <gary.gq@alibaba-inc.com> * test(core): stub the registry methods agent.ts actually calls (#7538) The shared stubRegistry in agent.test.ts was missing six methods that agent.ts reaches: bridgeApprovalEvents, getQueuedCount, registerResidentAgent, restartCompletedAgent, unregisterResidentAgent and waitForMessages. That is not a benign omission. The background body wraps its work in a try/catch that routes any throw into registry.fail(), so a missing method never surfaces as 'not a function' — it silently converts a successful run into a failed one. On the GOAL completion path unregisterResidentAgent is called immediately before complete(), so the TypeError replaced the completion entirely: registry.fail('fork-...', 'registry2.unregisterResidentAgent is not a function', ...) That is what broke 'runs a non-interactive fork through the background registry' on main. #7460 added the registry.complete assertion, which exposed the incomplete stub — before it, nothing checked whether the background body finished successfully and the TypeError was swallowed. Stub all six with their real return shapes (unregisterResidentAgent returns boolean, bridgeApprovalEvents returns the unsubscribe callback agent.ts later invokes, waitForMessages resolves to a list) and assert registry.fail was not called before asserting completion, so a future gap reports the actual error instead of 'complete: 0 calls'. * perf(startup): lazy-load Google GenAI SDK on first use (#7512) * perf(startup): lazy-load Google GenAI SDK on first use Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> * codex: address PR review feedback (#7512) Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> * codex: address PR review feedback (#7512) Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> --------- Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> * fix(vscode): use file picker image paths for vision input (#7493) * fix(vscode): use image paths from file picker * fix(vscode): keep image picker paths raw * fix(vscode): resolve image picker paths on submit * fix(vscode): send picked images as vision context * fix(vscode): encode prompt image file URIs * fix(vscode): address image path review comments * test(vscode): cover image file reference edge cases * fix(cli): open the actual serve fallback port (#7501) * fix(cli): open actual serve fallback port * test(cli): match serve URL to fallback listener * docs(cli): clarify serve listen error handling --------- Co-authored-by: Shaojin Wen <shaojin.wensj@alibaba-inc.com> * fix(ci): don't let one failing scenario sink the whole visual preview (#7511) The web-shell visuals render runs every screenshot and flow in a single `test:e2e:visuals`, and that step had no `continue-on-error`, while the compose and upload steps had no `if: always()`. So one failing or timing-out scenario failed the job, the artifact was never uploaded, and the publish workflow had nothing to post — the entire preview vanished even when every other scenario passed and its PNG was already on disk. A flow (a long multi-click sequence) is the most fragile scenario kind, so the fragile one silently takes down the deterministic screenshots. PR #7498 hit exactly this: 29 scenarios passed, one new channel-management flow timed out, and the PR got no preview and no comment at all. Make the after-capture step `continue-on-error` so the passing captures survive and the later steps still compose and upload them. The publish job only runs on a `success` conclusion, so the job must stay green — but a masked failure must not read as a clean preview. Ship the step's real `.outcome` (which continue-on-error does NOT mask, unlike `.conclusion`) to the publisher as `render-status.txt`, and have the comment builder use it: an empty preview whose render failed says "one or more scenarios failed to render" and is explicitly NOT the reassuring green check or the coverage-gap prompt (both imply the render ran); a partial preview is labelled partial above the shots that did render. A missing status file (older run) defaults to complete, so this only ever adds a warning, never suppresses a real preview. The failing scenario still needs fixing — it's now surfaced in the comment rather than by silently deleting everyone else's preview. Co-authored-by: wenshao <wenshao@example.com> * feat(web-shell): add selective shadow DOM isolation (#7551) Co-authored-by: 钉萁 <dingqi.jww@alibaba-inc.com> * feat(web-shell): add renderChatHeader slot for custom session header (#7553) * fix(cli): say review coverage gaps in the author's units, not chunk ids (#7550) The posted review body rendered coverage disclosures with the run's own bookkeeping as subjects: bare chunk ids, unsorted, one per subject. On a run that certified nothing (PR #7268) the body enumerated all 49 chunk ids across two sentences while opening with "Reviewed. Suggestions are inline." — the opener certified the exact thing every following sentence took back, and nothing on the PR page maps a chunk id to code. Three changes, all render-time — the structural entries, the caps, the caller-echo dedup and the stderr remediation still key on chunk ids, which is where the id is the selector a reader can act on: - Coverage now returns the plan's chunk→files table (DiffChunk.files was already in the plan JSON; the coverage type slice dropped it). - compose-review renders chunk gaps through describeChunkGap: every planned chunk collapses to "the entire diff", a narrow gap with known files names the files, and anything wider is counted against the plan's total. Applied to the receipt sentence, the uncoverable sentence (bare CLI entries only — caller-authored entries render verbatim) and the grouped per-cause sentences. - The COMMENT opener may no longer say "Reviewed." over a disclosure set that denies it: when no chunk is both covered and undisclosed — or no chunk universe could be read at all — it opens with a zero-certified warning instead. A rewritten launch demonstrably read its chunk, so coverage alone is not the test; certified is covered with no disclosure against it. Co-authored-by: verify <verify@local> * fix(autofix): retry a skipped-Prepare instead of stranding the PR terminal (#7490) * fix(autofix): retry a skipped-Prepare instead of stranding the PR terminal A base/infra failure BEFORE the agent runs was misread as an agent crash and terminated the PR forever. When an early step fails — installing or building the trusted base, checkout, node setup — the `Prepare branch and feedback` step is skipped, so NEWEST is empty, and the report step's "crashed before reading feedback" branch fired: MARK_ROUND=MAX_ROUNDS, terminal, scan skips it on every future tick. Observed: a web-shell TypeScript break on `main` failed `Install dependencies and build` (which builds the trusted base) across a whole scan batch, and SIX healthy PRs were stranded terminal at round=100 in one run — including ones at round 9 and 11 that had nothing to do with the break. `round=100` there is a terminal sentinel, not 100 attempts. NEWEST-empty now splits on steps.prepare.outcome: - 'skipped' (an earlier step failed, the agent never ran) is infra/base and transient: retry with a sentinel ts so the feedback stays live, incrementing the round so a PERSISTENTLY broken base is still bounded and stops at the cap (recoverable with /retry). - 'success'/'failure' (Prepare ran, no feedback produced) is a genuine pre-read agent crash: unchanged terminal behaviour. This is the reverse of the asymmetry #7482 addresses: that bounds a crash AFTER reading that retried forever; this stops a transient failure BEFORE reading from going terminal after one. * docs(autofix): note a pre-Prepare cancel also retries intentionally (#7490) * fix(autofix): also retry a cancelled/empty prepare outcome, not just skipped A previous review comment on this PR noted that a job cancelled before Prepare should retry too. It was right about the intent but the code did not do it: `steps.prepare.outcome` is 'cancelled' for a cancel and '' for a job that stopped before Prepare entered the step context — both DISTINCT from 'skipped', so `== 'skipped'` sent them to the terminal branch, the same over-termination this PR exists to fix. Match on "not a real Prepare run" (`!= 'success' && != 'failure'`) instead, so skipped, cancelled, and empty all retry; only a Prepare that actually ran to a verdict (success/failure) with no feedback stays terminal — the genuine pre-read agent crash. Test extended to drive the cancelled and empty cases (retry) and both real-run outcomes (terminal); mutation-verified that reverting to `== 'skipped'` reddens the cancelled case. * test(autofix): update the pre-read-crash case for the broadened retry The prior commit broadened NEWEST-empty retry to skipped/cancelled/empty but left the older 'replays the handoff decision' test asserting the old terminal behaviour for an unset PREPARE_OUTCOME (which now retries). That test's terminal cases now set PREPARE_OUTCOME=success/failure explicitly — the only outcomes that still terminate — so it exercises the genuine pre-read agent crash rather than the infra/cancel path. * test(autofix): anchor the skipped-Prepare extraction past the CONSEC block CI reddened `retries a skipped-Prepare` after main's consecutive-failure cap (#7482) merged into this branch: that block was inserted between this decision block and the report `{`, and it calls `gh api`. The test's `{`-anchored regex over-captured through it, so the extracted script ran the unstubbed `gh api` and failed. Anchor the end on the same `# Consecutive-failure` comment the sibling gate-crash test already uses, so the extraction stops at this decision block's own closing `fi`. * fix(autofix): exempt skipped-Prepare from the consecutive-failure breaker A broken base build skips Prepare, producing no API error file — so the consecutive-failure breaker ran on the new retry path and, after 5 scans, re-introduced the exact mass-stranding this PR exists to prevent. Exempt pre-agent infra failures (skipped/cancelled/empty outcome) from the breaker, mirroring the transient 429/5xx exemption: same failure class (not the PR's fault, self-heals, hits the whole batch). The round cap + sentinel-ts /retry recovery already bounds a persistently broken base. Also trim "checkout" from the retry headlines (checkout failures do not land in this branch) and hoist the duplicated MARK_TS assignment. * fix(autofix): reset the consecutive-failure streak on prior infra-failure markers The streak walker counted prior infra-failure headlines ("AutoFix could not start —…") as failures, inflating the consecutive-failure count on subsequent rounds. A PR with 3 real agent failures, then 3 rounds of base-build infra failures, then 1 more real failure would trip the cap-5 breaker even though only 4 rounds were the PR's fault. Add the two infra-failure headline patterns as reset strings in the streak walker, alongside the existing push and no-op resets. The genuine agent-crash headline ("AutoFix could not start evaluation —…") is deliberately excluded — it is a real failure and must still count. * fix(autofix): clarify infra-failure headlines and else-branch comment (#7490) Address review nits: the retry headline now mentions cancelled runs, the cap headline says 'reached the round cap' instead of overstating 'could not start for N rounds', the else-branch comment says 'prepare itself crashed' instead of 'agent crash', and the streak-reset pattern is simplified now that both infra headlines share the same prefix. --------- Co-authored-by: wenshao <wenshao@example.com> Co-authored-by: qwen-code-dev-bot <qwen-code-dev-bot@users.noreply.github.com> Co-authored-by: qwen-code-ci-bot <qwen-code-ci-bot@users.noreply.github.com> * fix(cli): keep role codenames and brief paths out of the posted review body (#7560) The posted body still carried two operator registers #7550 left in place: roster role subjects rendered their internal codenames ("Agent 1c: Cross-file tracer", "Test coverage matrix (whole-diff)"), and an unread brief's disclosure interpolated its filesystem path. And when verify and the reverse audit failed the same way, the body said it twice, in two near-identical sentences. - Every Brief now carries a publicLabel — the dimension said as what it checks ("the cross-file consistency pass") — and coverage's structural disclosures carry it as publicSubject beside the internal subject, plus a path-free publicReason for unread briefs. The internal label and the path stay on stderr, where they are the selector an operator acts on; every dedup and certification check still keys on the internal subject. - compose-review renders the public fields and groups by the reason the body PRINTS, so two unread briefs share one path-free sentence instead of repeating it per role. - verificationGaps merges verify and reverse-audit failures of the same delivery shape into one sentence with both subjects and both consequences; mixed shapes keep their precise per-role texts, and the per-role rebuild commands stay on stderr either way. Co-authored-by: verify <verify@local> * fix(autofix): retry an agent timeout instead of advancing past its feedback (#7563) A timeout evaluated NOTHING — the agent ran out of budget before finishing, so nothing was committed and the feedback is unaddressed. It was treated as an evaluated verdict (real ts, watermark advances), which strands that feedback: the next scan sees "nothing new" and never retries. Observed on #7471 (round 13/100), a heavily-reviewed 1871-line PR: rounds 11 and 13 timed out, but round 12 pushed — so a timeout is transient far more often than not, and advancing past it left the round-13 feedback unhandled. run-agent.mjs now drops an `agent-timeout` signal on result.timedOut, and the handoff routes it like a pre-verdict crash: sentinel ts (feedback stays live) and a retry, with a headline that names the real fix at the cap (split the PR or raise the budget). A PR that PERSISTENTLY times out is bounded by the round cap and the consecutive-failure cap, so this cannot loop forever — it just stops treating a one-off budget blip as a verdict. The loop guard stays terminal (a tool-call loop is a real defect, not a budget blip). An API error still routes to its own model-key handoff; the timeout signal is written only when NOT an API error. Co-authored-by: wenshao <wenshao@example.com> * feat(serve): add workspace-level generation (#7552) * feat(serve): add workspace-level generation * docs(serve): document workspace generation capability * fix(serve): align workspace generation contracts --------- Co-authored-by: ytahdn <ytahdn@gmail.com> * ci: matrix ECS runner update + sudo install + repository_dispatch trigger (#7513) * ci: matrix ECS runner update with sudo install - Use matrix strategy (ecs-update-sg, ecs-update-64c) to update both physical ECS hosts in parallel (fail-fast: false). - Always use sudo npm install -g so the package lands in /usr/local (system-wide PATH) instead of the runner user's home directory. - Move concurrency to job level (matrix context not available at workflow level per actionlint). - Add repository_dispatch trigger for release-driven updates. - Register new runner labels in actionlint.yaml. * fix(ci): use dispatch version for runner update --------- Co-authored-by: Shaojin Wen <shaojin.wensj@alibaba-inc.com> * fix(web-shell): include managed id in artifact open requests (#7570) Co-authored-by: 钉萁 <dingqi.jww@alibaba-inc.com> * feat(serve): persist workspace channel configuration (#7514) * feat(serve): persist workspace channel configuration * fix(serve): harden channel settings snapshots * fix(serve): validate startup channel names * fix(serve): reserve all channel name --------- Co-authored-by: Shaojin Wen <shaojin.wensj@alibaba-inc.com> * fix(sdk-python): require canonical form in validate_session_id (#7532) uuid.UUID() accepts several non-canonical spellings — braced {...}, urn:uuid:..., and dash-less hex — so validate_session_id let them through after the RFC 4122 variant check. The value is then forwarded to the CLI verbatim as --session-id/--resume, producing a malformed session id downstream rather than a clear error at the SDK boundary. Reject anything whose canonical form differs from the input. Case is deliberately not part of the comparison: UUID() lowercases, and an all-uppercase spelling is still valid canonical input. Co-authored-by: Shaojin Wen <shaojin.wensj@alibaba-inc.com> * fix(web-shell): sync background agent status (#7561) * fix(web-shell): sync background agent status * fix(web-shell): harden background agent reconciliation --------- Co-authored-by: ytahdn <ytahdn@gmail.com> * feat(core): propagate trusted daemon invocation context (#7279) * feat(core): propagate trusted daemon invocation context * test(cli): update ACP startup expectation * refactor(core): centralize ACP capability env key * test(cli): update worktree ACP core mock * test(integration): run daemon context smoke on PRs * test(ci): update no-AK smoke expectation * test(core): cover invocation context isolation * fix(cli): compare ACP capability safely * fix(docs): restore GitHub action input names * fix(core): sanitize private ACP capability from child env * fix(core): reuse private ACP capability env constant * test(cli): cover malformed trusted invocation context * test(acp-bridge): assert exact child environment --------- Co-authored-by: qwen-code-dev-bot <qwen-code-dev-bot@users.noreply.github.com> Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> Co-authored-by: 易良 <1204183885@qq.com> * fix(feishu): await stream cancels in media download teardown (#7465) * fix(feishu): await stream cancels in media download teardown downloadMedia left two reject paths' stream teardown unawaited: - the oversize-stream path called reader.cancel() without awaiting, so a cancel error during teardown became an unhandled rejection (fatal under Node's default --unhandled-rejections=throw); - the Content-Length reject path returned without cancelling resp.body, leaving the connection pinned until GC. Both were already fixed for the sibling DingTalk downloader in #7361 (which was itself modelled on this Feishu code), so this brings Feishu to parity. Adds a regression test that pins the reader.cancel() await via a rejecting cancel, plus an assertion that the Content-Length path releases the body. * test(feishu): cover a rejecting body.cancel() on the Content-Length path Mirrors the existing reader.cancel() teardown test for the other reject path, per review feedback. Removing the await on resp.body?.cancel() flips execution onto the 'rejected: size ... exceeds' branch and the test fails. * fix(autofix): make the review-address report wrapper lines bilingual (#7569) The agent's address-summary.md / no-action.md already ends with a collapsed Chinese translation, but the workflow-appended wrapper lines around it — the "Addressed/Reviewed the latest feedback" lead-in, the "Base-conflict check" line, and the "Re-review when you have a moment" footer — were English-only and sat outside that block. So the posted comment was only half translated, unlike the takeover-ack comments (full collapsed Chinese block) and the "model/模型" sign-off in this same report (already inline-bilingual). Give each wrapper line an inline Chinese translation, matching the model/模型 idiom. The English halves are preserved verbatim — the streak-reset detector globs on "Addressed the latest review feedback" and "no changes needed", and a test extracts these lines — so behaviour is unchanged and old English-only comments still match. A new test pins each English-Chinese pair so a future reword that drops the Chinese fails. The terminal handoff/failure comment is left English-only for now (SKILL.md keeps it so by design); that is a separate change. Co-authored-by: wenshao <wenshao@example.com> * feat(cli): post the review body bilingually when the PR description is Chinese (#7564) When the PR author writes Chinese, the posted /review body was English-only. fetch-pr now records whether the PR description contains Han characters (prDescriptionHasHan, detected from the same gh pr view call and stamped into the plan report), and compose-review renders the body bilingually off that flag: the English body leads, the complete Chinese version rides collapsed in a <details><summary>中文说明</summary> block, and the model footer stays outside the fold. The signal is the CLI's own — the caller cannot toggle the register of a certified body — and a local plan has no field, so nothing changes for terminal-only reviews. Every deterministic body fragment carries an en/zh pair end to end: compose-review's clause templates and describeChunkGap phrases, the coverage disclosures (reasons, publicLabel role subjects via a new publicLabelZh, the path-free unread-brief reason) and the Step 4/5 gap texts including the combined same-shape sentence. Fragments with no deterministic translation — model-written findings, caller echoes, interpolated errors — ride verbatim in both halves. verificationGaps now returns structural {subject, reason, subjectZh, reasonZh} entries, which also removes compose-review's last recover-the-boundary-from-prose parse. SKILL.md instructs the same format for the model-authored inline comments: English finding first (marker and suggestion block stay in the English half — tooling filters on them), full Chinese translation collapsed beneath, footer last. Co-authored-by: verify <verify@local> * feat(autofix): auto-rerun a check that died on infrastructure, once (#7562) * feat(autofix): auto-rerun a check that died on infrastructure, once A failed check can be red because the machine died, not the code — a self-hosted runner losing the server, the disk filling. #7490's E2E failed with "runner lost communication with the server" and went green on a rerun. The scan now reruns such a check's failed jobs automatically. Detection is a conservative annotation whitelist (INFRA_FAILURE_SIGNATURES) — only unambiguous machine failures, never a test-level timeout, which could be a real regression. The one-shot guard is run_attempt, not a marker: a run already retried to attempt 2 and still infra-failing is persistent, so it is left for a human; after a rerun the attempt increments, so the next scan will not rerun it. Every step is fail-safe (any API error → no rerun), it runs only when the PR actually has a failed check, and the gate carries the same review-address carve-out as the other check selectors so the loop never reruns its own runs. This is the transient-infra sibling of #7554 (stale-base): that merges current main when a check is base-inherited; this reruns when a check died on the runner. Neither touches a check that is a genuine failure. Note: rerun-failed-jobs needs the PAT to hold `actions: write`. * fix(autofix): use POSIX ERE groups in infra-failure regex, cover all signatures in tests (#7562) * fix(autofix): also treat a git fetch/clone transport death as infra #6506's checkout died mid-transfer — "fetch-pack: invalid index-pack output" and "RPC failed; curl 92 ... CANCEL" — which then hung the job into the 20m limit. That is infra, not the PR (it only touches a doc), and a re-run made it green. But the infra-signature whitelist did not cover it, so the auto-rerun did not fire and it waited on a human. Add `invalid index-pack output` and `RPC failed` — the two canonical git-transport-death phrases — to INFRA_FAILURE_SIGNATURES. A co-present job-timeout line does not block the match (one matching line classifies the run), and a BARE timeout with no transport signature is still left alone, since it can be a real regression. Both new signatures are pinned in the test's per-signature loop, plus a case on #6506's real composite annotation and a bare-timeout-is-not-rerun guard. * fix(autofix): paginate annotations and filter Autofix runs in infra-rerun loop (#7562) --------- Co-authored-by: wenshao <wenshao@example.com> Co-authored-by: Qwen Code Bot <qwen-code-bot@users.noreply.github.com> Co-authored-by: qwen-code-ci-bot <qwen-code-ci-bot@users.noreply.github.com> * fix(serve): detect stale SSE cursors across daemon restarts via epoch token; preserve turn attribution and surface compaction failures in replay (#7458) * fix(daemon): epoch-token restart detection, compaction attribution, and degraded-snapshot signaling (DAEMON-001/007/008) * fix(acp-bridge): field-level turn attribution merge and replayDegraded bridge test (#7458) * fix(serve): skip bus epoch lookup for virtual subagent SSE streams (#7458) The REST SSE route looked up the bus epoch for every session id, but virtual subagent sessions ride their own bus and their compound ids are not in the bridge's byId map, so the lookup threw and aborted the subscription — breaking subagent event streams. Skip the lookup for the virtual path and degrade a torn-down real session to a headerless stream (mirrors the /acp route). Also bumps the daemon browser SDK bundle budget (167KB -> 168KB) for the epoch fields and declares eventEpoch on DaemonSession so the create/attach path drops its inline type cast. * fix(serve): stamp eventEpoch on accepted continuations and surface replayDegraded in the SDK (#7458) Address three review suggestions: - POST /session/:id/continue now returns eventEpoch alongside lastEventId, mirroring the prompt 202 envelope so continuation-seeded SSE cursors detect daemon restarts (DAEMON-001) - DaemonSessionClient exposes replayDegraded from the load response so SDK consumers can prefer the full transcript over a degraded snapshot - add /acp dispatch-level regression test for the degraded-snapshot stderr breadcrumb (fires only when snapshot.degraded is set) * test(cli): fix load-reply race in the degraded-breadcrumb transport test Await each session/load reply frame before opening the session stream so the GET cannot race conn.ownSession() into a 403; addresses the review Critical on the deg-0 arm. * fix(serve): allow and expose X-Qwen-Event-Epoch in CORS headers Cross-origin SSE clients must send the epoch header through preflight and read it from the response, or stale-cursor detection (DAEMON-001) is silently disabled for every CORS client. --------- Co-authored-by: qwen-code-bot <qwen-code-bot@users.noreply.github.com> Co-authored-by: qwen-code-dev-bot <qwen-code-dev-bot@users.noreply.github.com> Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> Co-authored-by: Qwen Autofix <qwen-autofix[bot]@users.noreply.github.com> * feat(core): Align GenAI telemetry with ARMS (#7536) * feat(core): align GenAI telemetry with ARMS Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> * fix(core): remove estimated token usage splits Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> * fix(core): address GenAI telemetry review feedback Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> --------- Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> Co-authored-by: Shaojin Wen <shaojin.wensj@alibaba-inc.com> * fix(serve): avoid TOCTOU race dropping live sessions from list response (#7556) * Initial plan * fix(serve): avoid TOCTOU race dropping live sessions from list response --------- Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com> Co-authored-by: 易良 <1204183885@qq.com> * fix(cli): prevent monitor turns after task_stop (#7573) --------- Co-authored-by: 秦奇 <gary.gq@alibaba-inc.com> Co-authored-by: Qwen Code Autofix <qwen-code-autofix@users.noreply.github.com> Co-authored-by: Qwen Code Bot <qwen-code-bot@users.noreply.github.com> Co-authored-by: qwen-code-ci-bot <qwen-code-ci-bot@users.noreply.github.com> Co-authored-by: Qwen Code Autofix <qwen-code-autofix[bot]@users.noreply.github.com> Co-authored-by: qwen-code-dev-bot <qwen-code-dev-bot@users.noreply.github.com> Co-authored-by: destire-mio <qppque@gmail.com> Co-authored-by: destire-mio <248462155+destire-mio@users.noreply.github.com> Co-authored-by: Dragon <52599892+DragonnZhang@users.noreply.github.com> Co-authored-by: Claude <noreply@anthropic.com> Co-authored-by: 易良 <1204183885@qq.com> Co-authored-by: jinye <djy1989418@126.com> Co-authored-by: chinesepowered <nlai@rediffmail.com> Co-authored-by: ovochouovo <18212194+ovochouovo@users.noreply.github.com> Co-authored-by: Edenman <67549719+BZ-D@users.noreply.github.com> Co-authored-by: 克竟 <dingbingzhi.dbz@alibaba-inc.com> Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com> Co-authored-by: ytahdn <1294726970@qq.com> Co-authored-by: ytahdn <ytahdn@gmail.com> Co-authored-by: Truraly <94105924+Truraly@users.noreply.github.com> Co-authored-by: zjgzx1988 <zjgzx1988@hotmail.com> Co-authored-by: hogeheer499-commits <hogeheer499@gmail.com> Co-authored-by: hogeheer <267467744+hogeheer499-commits@users.noreply.github.com> Co-authored-by: Shaojin Wen <shaojin.wensj@alibaba-inc.com> Co-authored-by: wenshao <wenshao@example.com> Co-authored-by: qwen-code-dev-bot <qwen-code-dev@service.alibaba.com> Co-authored-by: Nothing Chan <chenliu.cl@alibaba-inc.com> Co-authored-by: 钉萁 <dingqi.jww@alibaba-inc.com> Co-authored-by: yuanyuanAli <135116774+yuanyuanAli@users.noreply.github.com> Co-authored-by: verify <verify@local> Co-authored-by: qqqys <qys177@gmail.com> Co-authored-by: callmeYe <512217680@qq.com> Co-authored-by: Qwen Autofix <qwen-autofix[bot]@users.noreply.github.com> Co-authored-by: Copilot <198982749+Copilot@users.noreply.github.com>
|
Thanks for the deep post-merge review — verified all six findings against the code and confirmed your analysis. Status, point by point: Fixed on a follow-up branch (
Planned for the same follow-up PR: #4 (budget ledger comment — will reword to 'no bump needed'), #5 ( The forcePush note is recorded as pre-existing amplification, not addressed here. |
PR QwenLM#7458 on main adopted a more structured approach to the same problem PR QwenLM#7499 solves (preserving event attribution through turn compaction). Resolved in favour of main's structured lastTurn/lastSessionId fields plus captureTurnFields/captureSessionId helpers, removing the duplicate spread operators both sides independently added to makeMergedSessionUpdateEvent and mergeToolCallEvent.








What this PR does
This PR hardens the daemon's event replay/reconnect chain in three ways. First, every session event bus now mints a random epoch token when it is constructed, and that token travels with every surface a client learns a resume cursor from: session load/resume/create responses, the non-blocking 202 prompt envelope, and an
X-Qwen-Event-Epochresponse header on both SSE surfaces (REST events stream and the/acpsession stream). Clients echo the token back alongsideLast-Event-IDon reconnect; when it doesn't match the bus's current epoch, the daemon deterministically forces the existing resync path instead of trusting event-id arithmetic. The resync frame carries adetail: 'epoch_mismatch'discriminator so operators can tell the token-based trigger apart from the numeric heuristic. Second, the turn-boundary compaction engine now preserves turn attribution: merged text/thought events re-stamp the latestpromptId/originatorClientId/data.sessionIdcaptured from their source chunks, and folded tool-call events switch to latest-wins attribution, mirroring how the event id is already merged. Third, compaction failures are no longer invisible: the bus latches a degraded flag on the first ingest/seed failure, replay snapshots built after that point are marked degraded, session load responses surface it asreplayDegraded, and the daemon logs one operator breadcrumb per session plus a warning when an/acpinitial replay serves a degraded snapshot. The TypeScript SDK learns the epoch from restore responses and response headers, persists it next to the cursor, and echoes it on every reconnect — including the prompt-envelope-driven subscribe insideprompt().Why it's needed
Event ids restart from 1 on every daemon restart, and the only stale-cursor detection was a numeric heuristic (
lastEventId >= nextId). Once the new epoch's event count catches up with a stale cursor, that heuristic is defeated: a client reconnecting withLast-Event-ID: 50against a bus that has already emitted 60 fresh events looks like a perfectly valid suffix resume, so it silently skips the new epoch's first 50 events and applies deltas on top of reducer state from the dead epoch. Separately, clients rebuilding state from a compacted replay snapshot lost the ability to correlate merged events to their prompt or filter by originator, because compaction dropped those stamps. And when the compaction engine threw, the bus correctly swallowed the error to keeppublish()never-throwing — but nothing recorded that the snapshot was now incomplete, so every later consumer served silently-truncated replay data with no signal to the operator or the client. These are items DAEMON-001, DAEMON-007, and DAEMON-008 from the daemon/SDK reliability audit ( https://github.com/doudouOUC/code_agent/blob/main/qwen-code/feature/daemon-serve-mode/12-daemon-sdk-reliability-audit.md ), following the earlier batches #7386 and #7400 .All changes are backward compatible in both directions: an old client that never sends the epoch header gets today's numeric-heuristic behavior, and a new client talking to an old daemon simply never learns an epoch and falls back the same way. The new response fields are optional and additive.
Reviewer Test Plan
How to verify
qwen serve, create a session, prompt it, and note the SSE cursor. Restart the daemon, load the session again, then subscribe toGET /session/:id/eventswith the OLDLast-Event-IDand the OLDX-Qwen-Event-Epochvalue. Expected: the stream opens with astate_resync_requiredframe withreason: 'epoch_reset'anddetail: 'epoch_mismatch', followed by a full replay — even when the new bus has already emitted more events than the stale cursor (the case the numeric heuristic misses). Without the header, behavior is unchanged.eventEpochnext tolastEventId, that a non-blocking prompt 202 envelope carrieseventEpoch, and that both SSE surfaces respond with anX-Qwen-Event-Epochheader (also on the first, cursor-less subscribe).agent_message_chunk/agent_thought_chunkevents incompactedReplaystill carry top-levelpromptId/originatorClientIdanddata.sessionId; folded tool_call events carry the latest stamp.degraded: true, load responses gainreplayDegraded: true, and the daemon writes a singlecompaction degraded for session=…stderr line.packages/acp-bridge(eventBus epoch + degradation, compactionEngine attribution, bridge restore payloads),packages/cli(REST SSE header/query plumbing,/acptransport epoch pairing, 202 envelope),packages/sdk-typescript(transport header send/learn, session client seeding/refresh).Evidence (Before & After)
N/A (protocol/replay-layer change; no TUI surface). Full test runs after rebasing onto latest main:
packages/acp-bridge541 passed,packages/cliserve suites 771 + 409 passed,packages/sdk-typescript423 passed;npm run buildandnpm run typecheckclean.Tested on
Environment (optional)
Unit tests via vitest per package; build + typecheck from the repo root.
Risk & Scope
/acptransport intentionally ignores the epoch option (it has no resume mechanism, so there is no stale-cursor problem — documented inline); DAEMON-009/010/011 (resource hardening) are a separate batch.Linked Issues
Fixes #7457
中文说明
本 PR 做了什么
本 PR 从三方面加固 daemon 的事件重放/重连链路。其一,每个会话事件总线在构造时生成一个随机 epoch token,并随所有客户端可获取续传游标的通道下发:load/resume/create 响应、非阻塞 202 prompt envelope,以及两个 SSE 通道(REST 事件流与
/acp会话流)的X-Qwen-Event-Epoch响应头。客户端重连时随Last-Event-ID一起回传该 token;与总线当前 epoch 不一致时,daemon 确定性地走既有的 resync 路径而不再依赖事件 id 的数字推断。resync 帧携带detail: 'epoch_mismatch'判别字段,方便运维区分 token 触发与数字启发式触发。其二,turn 边界压缩引擎现在保留 turn 归属:合并后的 text/thought 事件重新盖上源 chunk 中捕获的最新promptId/originatorClientId/data.sessionId,折叠后的 tool_call 事件改为 latest-wins,与事件 id 的合并方式一致。其三,压缩失败不再不可见:总线在首次 ingest/seed 失败时锁存 degraded 标志,此后构建的重放快照标记为 degraded,load 响应以replayDegraded透出,daemon 每会话写一条运维日志,/acp初始重放使用降级快照时也会告警。TypeScript SDK 从 restore 响应和响应头学习 epoch,与游标一起保存,并在每次重连(包括prompt()内部由 202 envelope 驱动的订阅)时回传。为什么需要
daemon 每次重启后事件 id 从 1 重新开始,此前唯一的过期游标检测是数字启发式(
lastEventId >= nextId)。一旦新纪元的事件数追上旧游标,该启发式即失效:客户端带着Last-Event-ID: 50重连、而新总线已发出 60 个事件时,看起来是完全合法的后缀续传,于是客户端静默跳过新纪元前 50 个事件,并把增量应用在死纪元的 reducer 状态之上。另外,从压缩重放快照重建状态的客户端无法再做 prompt 关联和 originator 过滤,因为压缩丢掉了这些标记。压缩引擎抛错时总线为维持publish()永不抛错而正确地吞掉了异常——但没有任何地方记录快照已不完整,之后所有消费方都把被静默截断的重放数据当完整的下发,运维和客户端都得不到信号。这些对应 daemon/SDK 可靠性审计文档( https://github.com/doudouOUC/code_agent/blob/main/qwen-code/feature/daemon-serve-mode/12-daemon-sdk-reliability-audit.md )中的 DAEMON-001、DAEMON-007、DAEMON-008,是 #7386 与 #7400 之后的第三批。所有改动双向向后兼容:旧客户端不发 epoch 头则保持现有数字启发式行为;新客户端连旧 daemon 学不到 epoch,同样回落。新增响应字段均为可选、纯增量。
审阅者验证计划
如何验证
qwen serve,创建会话并 prompt,记下 SSE 游标。重启 daemon 并重新 load 会话,然后带旧的Last-Event-ID和旧的X-Qwen-Event-Epoch订阅GET /session/:id/events。预期:流以state_resync_required(reason: 'epoch_reset'、detail: 'epoch_mismatch')开头并全量重放——即使新总线的事件数已超过旧游标(数字启发式漏掉的场景)。不带该头则行为不变。lastEventId旁携带eventEpoch、非阻塞 prompt 202 envelope 携带eventEpoch、两个 SSE 通道均返回X-Qwen-Event-Epoch响应头(首个无游标订阅也返回)。compactedReplay中合并的agent_message_chunk/agent_thought_chunk仍携带顶层promptId/originatorClientId与data.sessionId;折叠的 tool_call 携带最新标记。degraded: true,load 响应获得replayDegraded: true,daemon 写一条compaction degraded for session=…stderr 日志。packages/acp-bridge(eventBus epoch + 降级、compactionEngine 归属、bridge restore 载荷)、packages/cli(REST SSE 头/查询串联、/acptransport epoch 配对、202 envelope)、packages/sdk-typescript(transport 头收发、session client 播种/刷新)。证据(Before & After)
N/A(协议/重放层改动,无 TUI 表面)。rebase 到最新 main 后全量测试:
packages/acp-bridge541 通过、packages/cliserve 套件 771 + 409 通过、packages/sdk-typescript423 通过;npm run build与npm run typecheck干净。风险与范围
/acptransport 有意忽略 epoch 选项(无续传机制即无过期游标问题——已内联注释说明);DAEMON-009/010/011(资源硬化)属于另一批。关联 Issue
Fixes #7457