fix(cli): improve slash command history feedback - #8365
Conversation
E2E Test ReportTested the worktree build in a real PTY TUI on macOS with an isolated
Focused E2E-related unit runs passed: |
|
Gate re-run against the current head. Template ✓ — all required sections present, bilingual body included. Problem: observed, not theoretical — with unusually strong evidence. Any interactive session shows it: run Direction: aligned. Slash-command UX is an actively invested area here (e.g. #8130 safe slash commands during streaming — same author, #7818 Size: 2175 changed lines = 1528 test / 647 production (incl. a 40-line design doc and 9 one-line i18n additions) / 0 generated. Two core paths are touched: Approach: the scope matches the problem. Hiding invocations necessarily extends to the reconstruction paths ( Risk: no high-risk path matches in the revert-correlation scan. Moving on to code review. 🔍 中文说明针对当前 head 重新执行准入门检查。 模板 ✓ —— 必填章节齐全,包含中文说明。 问题: 已观测到、非理论问题——且证据异常充分。任何交互会话都能看到:运行 方向: 对齐。slash command UX 是本仓库持续投入的领域(例如 #8130 流式响应期间允许安全 slash command——同一作者,#7818 规模: 共 2175 行变更 = 1528 行测试 / 647 行生产代码(含 40 行设计文档和 9 处单行 i18n 词条)/ 0 行生成代码。触及两个核心路径: 方案: 范围与问题匹配。隐藏 invocation 必然要同步覆盖重建路径( 风险: revert 相关性扫描未命中任何高风险路径。 进入代码审查。🔍 — Qwen Code · qwen3.8-max Reviewed at |
Code reviewI formed an independent proposal before reading the diff (defer invocation recording until the action is resolved; classify built-ins only; persist a hide marker that resume/transcript reconstruction honors; give a cancelled Things I verified against the base code, not just the diff:
Non-blocking items already documented in this thread — no need to relitigate them here: the desktop-side Files changed (27 of 34 shown)
TestingCI signal for the reviewed commit, fetched via the API (this unattended run never executes PR code):
The Linux unit suite — including every test this PR adds — is green. The macOS/Windows legs are skipped as usual for fork PRs; the desktop bun suite ( Sandboxed verification would settle the remainder:
中文说明代码审查: 在读 diff 之前我先独立给出了方案(action 解析前暂缓记录 invocation、仅对内置命令分类、用隐藏标记让 resume/transcript 重建遵守、取消的 几处对照基线代码(而非只看 diff)确认的点:旧的 线程中已记录的非阻塞项不再重复:desktop 侧 hiddenInvocation 补丁检查仍无测试覆盖、辅助选择器按设计静默关闭、 测试: 以上 CI 表格来自 API。Linux 单元套件(含本 PR 全部新增测试)通过;macOS/Windows 腿按 fork PR 惯例跳过,desktop bun 套件不在任何 gating CI 中运行(维护者曾在上一 head 本地覆盖,此处仅引用其自述)。其后唯一新提交是纯测试。行为类验证可用 — Qwen Code · qwen3.8-max Reviewed at |
|
Confidence: 4/5 — clean review, verified behaviour; the one thing that blocked the last run (the Stage 0 size escalation) is now resolved by the maintainer's own commit-pinned approval. The remaining items are documented non-blocking follow-ups. The previous pass ended in a defer over policy, not substance: 647 production lines with two small core touches crossed the threshold where this gate hands the decision to a human instead of voting alone. That escalation has now been answered — @wenshao approved this exact commit (full SHA in the footer) with a real-stack A/B report attached: 12 transient commands leave zero bare rows on head, My own read is unchanged from the staged review above, re-verified against this head: the classifier hides only bare built-in root/picker forms, aliases inherit via the canonical path, custom overrides are exempt by the Non-blocking follow-ups, named for the record: the desktop-side hide-flag patch check ( Approving, pinned to the reviewed commit. ✅ 中文说明置信度:4/5 —— 审查干净、行为已验证;上一轮唯一的阻碍(Stage 0 规模升级)现已由维护者对同一提交的亲自批准解决。剩余事项均为已记录的非阻塞后续项。 上一轮以政策原因 defer,而非实质问题:647 行生产代码加两处小的核心改动越过了关卡自行投票前须交人工决定的阈值。该升级现已得到回应——@wenshao 在这同一个提交上批准(完整 SHA 见页脚),并附上真实环境 A/B 报告:12 个临时命令在 head 上残留零裸行、 我本人的判断与上方各阶段审查一致,并已对当前 head 重新核验:分类器只隐藏裸内置 root/picker 形态,别名经由 canonical path 继承,自定义覆盖因 非阻塞后续项记录在案:desktop 侧隐藏标记补丁检查( 批准,锚定到已审查的提交。✅ — Qwen Code · qwen3.8-max Reviewed at |
qwen-code-ci-bot
left a comment
There was a problem hiding this comment.
LGTM, looks ready to ship — CI landed green after the review. ✅
qwen-code-ci-bot
left a comment
There was a problem hiding this comment.
Reviewed. Suggestions are inline. Not reviewed: build-and-test — Integration Tests (CLI, No Sandbox) was skipped in CI and its suite did not run locally.
中文说明
已审查。 建议见行内评论。 未审查:build-and-test — Integration Tests (CLI, No Sandbox) was skipped in CI and its suite did not run locally。
— qwen3.8-max-preview via Qwen Code /review
qwen-code-ci-bot
left a comment
There was a problem hiding this comment.
Reviewed. Suggestions are inline. Not reviewed: build-and-test — Integration Tests (CLI, No Sandbox) was skipped in CI and its suite did not run locally.
中文说明
已审查。 建议见行内评论。 未审查:build-and-test — Integration Tests (CLI, No Sandbox) was skipped in CI and its suite did not run locally。
— qwen3.8-max-preview via Qwen Code /review
qwen-code-ci-bot
left a comment
There was a problem hiding this comment.
Reviewed. Not reviewed: build-and-test — Integration Tests (CLI, No Sandbox) was skipped in CI and its suite did not run locally.
中文说明
已审查。 未审查:build-and-test — Integration Tests (CLI, No Sandbox) was skipped in CI and its suite did not run locally。
— qwen3.8-max-preview via Qwen Code /review (v0.21.3)
OverviewTwo independent changes aimed at the same symptom — a slash command that leaves a bare
The refactor of The concerns below are mostly about the classification strategy, plus one behavior gap that undoes the fix. 1. Resumed / previewed sessions bring the bare invocation row back (main issue)The invocation is still recorded unconditionally ( if (payload.phase === 'invocation' && payload.rawCommand) {
items.push({ type: 'user', text: payload.rawCommand, ... });
}
Suggestion: carry the decision into the record rather than dropping it at the UI layer — e.g. a 2. Classification is a hardcoded list living far from the commands
3. Two opposite strategies for the same UX problem
4. Missing translations
5.
|
qwen-code-ci-bot
left a comment
There was a problem hiding this comment.
Reviewed. Not reviewed: build-and-test — Integration Tests (CLI, No Sandbox) was skipped in CI and its suite did not run locally.
中文说明
已审查。 未审查:build-and-test — Integration Tests (CLI, No Sandbox) was skipped in CI and its suite did not run locally。
— qwen3.8-max-preview via Qwen Code /review (v0.21.3)
|
@qwen-code /takeover |
|
🤝 Takeover engaged: the autofix loop now manages this PR — it will address new review feedback and resolve base conflicts until the label is removed or the round cap is reached. This is a fork PR, so the first round comes from the next scheduled scan (usually within minutes). Remove the 中文说明🤝 已接管:autofix 循环现在管理此 PR —— 将持续处理新的评审反馈与 base 冲突,直到移除标签或达到轮次上限。本 PR 来自 fork,首轮处理将由下一次定时扫描执行(通常几分钟内)。移除 |
|
@qwen-code /triage |
|
Sandboxed verification: ✅ passed — merge-ready (agent verdict) - workflow run Ran the PR in an isolated, token-free container: A/B against the base build, mock-free harness assertions, targeted gates. Advisory evidence for human reviewers — not a review, an approval, or a CI check. Scripted assertions: 74 passed · 0 failed · 74 total 中文 — 判定:✅ 通过 · 可合入(agent 判定)沙箱验证在隔离、无凭证的容器中执行了该 PR 的代码(与 base 构建 A/B 对照、无 mock harness 断言、定向门禁)。仅作为评审证据,不构成评审、批准或 CI 检查。 脚本断言:74 通过 · 0 失败 · 74 总计 Verification reportPR 8365 — fix(cli): improve slash command history feedbackVerdict: 中文摘要
Central claim + A/BCentral claim: built-in transient slash commands stop leaving invocation-only rows in TUI history (roots Environment per cell: real compiled Control purity:
Both arms: 34/34 assertions pass (raw per-cell logs in Sibling sweep of the suppression sets: all seven suppressed roots are dialog- or info-only in the real registry ( FindingsF1 (suggestion) — a user command that overrides a built-in name silently drops that built-in's aliases. With F2 (suggestion) — transient "Unknown command" window after settings persistence. Selecting a model or F3 (suggestion, unverified) — Not covered
MethodologyEnvironment: Evidence imagesHarness scripts and raw logs are in the workflow run artifacts (7-day retention). — Qwen Code · sandboxed verification |
|
Triage re-run completed without a new review.
The stage comments above were updated with the latest result. View workflow run. 上方各阶段评论已更新为最新结果。查看工作流运行。 |
|
🤖 Addressed the latest review feedback (round 1/100). What changed, and what I pushed back on: · 已处理最新评审反馈(第 1/100 轮)。改动内容与我反驳保留之处如下: Address review summary — PR #8365Round changes are in commit Issue-level review (@wenshao)1. Resumed/previewed sessions bring the bare invocation row back — FIXEDThe hiding decision is now carried into the chat record instead of being dropped at the UI layer, as suggested:
Note: the desktop transcript reconstruction in 2. Classification is a hardcoded list living far from the commands — CONCRETE BUGS FIXED, STRUCTURAL MOVE DECLINED
3. Two opposite strategies for the same UX problem — RESOLVED per the reviewer's suggested landingBare 4. Missing translations — FIXED
5.
|
doudouOUC
left a comment
There was a problem hiding this comment.
Reviewed. Suggestions are inline.
Not explored to full depth (tool budget reached): This PR hides transient slash-command invocations from TU...: 无。; This PR hides transient slash-command invocations from TU...: None. All - lines were examined.; This PR hides transient slash-command invocations from TU...: None. All 19 pairings examined..
中文说明
已审查。 建议见行内评论。
未探索到全部深度(达到工具调用预算):This PR hides transient slash-command invocations from TU...:无。;This PR hides transient slash-command invocations from TU...:None. All - lines were examined.;This PR hides transient slash-command invocations from TU...:None. All 19 pairings examined.。
— deepseek-v4-flash via Qwen Code /review (v0.21.10)
qwen-code-ci-bot
left a comment
There was a problem hiding this comment.
Reviewed. Suggestions are inline.
Not reviewed: build-and-test — packages/desktop bun suite (qwen-agent-slash-history.test.ts) is negated from the npm workspace graph and executed by no gating CI; not runnable on this review host (no bun) — the desktop reconstruction changes were verified statically and by direct probes of the real loadSlashCommandInvocationMessages only..
Not reviewed: build-and-test — Test (macos-latest, Node 22.x) and Test (windows-latest, Node 22.x) CI legs are skipped for fork PRs; the unit suites ran on Linux only. Integration Tests (CLI, No Sandbox) is merge_group-gated by design and collects no file this diff changes..
Not reviewed: reverse audit — did not converge within the reverse-audit round cap of 5..
Not explored to full depth (tool budget reached): PR #8365 (round 12): hides transient slash-command invoca...: none — all checks above completed within budget.; PR #8365 (round 12): hides transient slash-command invoca...: none — all checks above completed within budget; I did not run the vitest file since the finding is an assertion-absence verified textually at the reviewed comm…; PR #8365 (round 12): hides transient slash-command invoca...: none — all checks above completed within budget.; chunk 9: could not execute the three touched test files — the review worktree has no node_modules ( npx vitest fails with ERR_MODULE_NOT_FOUND) and installing depende…; PR #8365 (round 12): hides transient slash-command invoca...: none — all checks I intended completed within budget. Consumer-side logic in slashCommandProcessor.ts (invocation-phase recording, revealHiddenInvocation ) a…, and 4 more.
中文说明
已审查。 建议见行内评论。
未审查:build-and-test — packages/desktop bun suite (qwen-agent-slash-history.test.ts) is negated from the npm workspace graph and executed by no gating CI; not runnable on this review host (no bun) — the desktop reconstruction changes were verified statically and by direct probes of the real loadSlashCommandInvocationMessages only.。
未审查:build-and-test — Test (macos-latest, Node 22.x) and Test (windows-latest, Node 22.x) CI legs are skipped for fork PRs; the unit suites ran on Linux only. Integration Tests (CLI, No Sandbox) is merge_group-gated by design and collects no file this diff changes.。
未审查:reverse audit — did not converge within the reverse-audit round cap of 5.。
未探索到全部深度(达到工具调用预算):PR #8365 (round 12): hides transient slash-command invoca...:none — all checks above completed within budget.;PR #8365 (round 12): hides transient slash-command invoca...:none — all checks above completed within budget; I did not run the vitest file since the finding is an assertion-absence verified textually at the reviewed comm…;PR #8365 (round 12): hides transient slash-command invoca...:none — all checks above completed within budget.;chunk 9:could not execute the three touched test files — the review worktree has no node_modules ( npx vitest fails with ERR_MODULE_NOT_FOUND) and installing depende…;PR #8365 (round 12): hides transient slash-command invoca...:none — all checks I intended completed within budget. Consumer-side logic in slashCommandProcessor.ts (invocation-phase recording, revealHiddenInvocation ) a…,另有 4 条。
— qwen3.8-max via Qwen Code /review (v0.21.10)
| const errorMessage = e instanceof Error ? e.message : String(e); | ||
| setErrorMessage( | ||
| `${t('Model switched, but the selection could not be saved.')}\n\n${errorMessage}`, |
There was a problem hiding this comment.
[Suggestion] After a post-switch persistence failure the dialog stays open but is selection-dead: selectionCommittedRef is already true, so handleSelect's guard silently discards every further selection (including a retry of the same row), and closeWithoutSelection skips the feedback entirely — a hidden /model exchange ends with zero live history items and zero recorded phase:'result'. — Failure scenario: settings file unwritable or disk full (EROFS/EACCES/ENOSPC) → switchModel succeeds and selectionCommittedRef.current = true → persistModelSelection → settings.setValue → saveSettings re-throws → this catch shows the error and returns with the dialog still mounted → every later Enter/number-key selection is discarded by the guard (which runs before setErrorMessage(null)), and Escape exits with no history item and no /model result record, so the whole exchange vanishes from live and resumed history. Verified against the code at this commit; filed as a Suggestion rather than Critical because the pre-change behavior was worse (an unhandled promise rejection), Escape always exits, and the displayed error is accurate.
Two fix directions:
// Option A — treat the post-switch failure as terminal (the switch is
// already applied in memory; retrying would just re-fail the same disk):
} catch (e) {
const errorMessage = e instanceof Error ? e.message : String(e);
setErrorMessage(/* ...unchanged... */);
onClose();
return;
}
// Option B — keep the dialog usable: introduce a separate persistFailedRef,
// set it in this catch, and have handleSelect's guard consult
// `selectionCommittedRef.current && !persistFailedRef.current` instead.中文说明
[Suggestion] 切换成功但持久化失败后,对话框仍然打开但已无法再进行任何选择:此时 selectionCommittedRef 已为 true,handleSelect 的守卫会静默丢弃之后所有的选择(包括重试同一行),且 closeWithoutSelection 会完全跳过反馈——一次被隐藏的 /model 交互最终在实时历史和持久化记录中都零痕迹。— 失败场景:设置文件不可写或磁盘满(EROFS/EACCES/ENOSPC)→ switchModel 成功且 selectionCommittedRef.current = true → persistModelSelection → settings.setValue → saveSettings 重新抛出异常 → 此 catch 显示错误并在对话框仍然挂载时返回 → 之后每一次回车/数字键选择都被守卫丢弃(该守卫在 setErrorMessage(null) 之前执行),按 Escape 退出时既不产生历史条目也不产生 /model result 记录,整段交互在实时与 resume 历史中同时消失。已在本提交的代码上核实;定为 Suggestion 而非 Critical,因为改动前的行为更糟(未处理的 promise rejection)、Escape 始终可以退出、且错误文案准确描述了状态。
修复方向见英文部分代码块(A:失败即关闭对话框;B:新增 persistFailedRef,让 handleSelect 守卫在持久化失败后允许重新选择)。
— qwen3.8-max via Qwen Code /review (v0.21.10)
| `${t('Model switched, but the selection could not be saved.')}\n\n${errorMessage}`, | ||
| ); |
There was a problem hiding this comment.
[Suggestion] This new user-facing string is wrapped in t() but exists in none of the 9 locale files — not even en.js — while the sibling string added by the same diff (Kept model as {{model}}) was translated in all 9. — Failure scenario: this catch block is the key's only use site (verified by grep: here and its test), and the repo's t() falls back to the raw key, so every non-English user hitting this persistence-failure path sees an untranslated English message. The gap is also invisible to the translation workflow: scripts/check-i18n.ts run at this commit prints "All checks passed!" because its parity checks anchor on en.js keys and cannot detect a used-but-unregistered key.
// en.js (identity):
'Model switched, but the selection could not be saved.':
'Model switched, but the selection could not be saved.',
// + real translations in the other 8 locale files (ca, de, fr, ja, pt, ru,
// zh, zh-TW), mirroring how 'Kept model as {{model}}' was added中文说明
[Suggestion] 这条新增的用户可见文案包在 t() 中,但 9 个语言文件里一个都没有(连 en.js 也没有)——而同一 diff 新增的姊妹文案 Kept model as {{model}} 在全部 9 个语言文件中都做了翻译。— 失败场景:该 catch 块是这个 key 唯一的使用点(已用 grep 核实:此处及其测试),仓库的 t() 找不到 key 时回退为 key 原文,因此每个非英语用户在遇到该持久化失败路径时看到的都是未翻译的英文。这个缺口对翻译工作流也不可见:在本提交上运行 scripts/check-i18n.ts 输出 "All checks passed!",因为其一致性检查以 en.js 的 key 为锚点,无法检测「代码中使用但未注册」的 key。
修复见英文部分代码块:在 en.js 注册恒等映射,并在其余 8 个语言文件中按 Kept model as {{model}} 的方式补齐翻译。
— qwen3.8-max via Qwen Code /review (v0.21.10)
| } = useEditorSettings( | ||
| settings, | ||
| setEditorError, | ||
| historyManager.addItem, | ||
| config, | ||
| ); |
There was a problem hiding this comment.
[Suggestion] The new config wiring from AppContainer into useEditorSettings (and identically into useThemeCommand at ~1426-1431) has no test: AppContainer.test.tsx mocks both hooks wholesale (vi.mock('./hooks/useThemeCommand.js'), vi.mock('./hooks/useEditorSettings.js')), and the parameter is optional (config?: Config), so nothing type-checks the wiring either. — Failure scenario: a future change that drops the config argument from either call site compiles and leaves every test green (the hook-level tests inject their own config), while /editor selection feedback and /theme-under-NO_COLOR feedback silently stop calling recordSlashCommand — resumed/branched sessions and desktop reconstruction then lose those outcome lines with no error anywhere.
// AppContainer.test.tsx — assert the hooks receive the config instance:
expect(mockUseEditorSettings).toHaveBeenCalledWith(
mockLoadedSettings,
mockSetEditorError,
mockAddItem,
config,
);
// (and the same for mockUseThemeCommand's trailing arg)中文说明
[Suggestion] AppContainer 向 useEditorSettings 传入 config 的新接线(以及 ~1426-1431 处 useThemeCommand 的相同接线)没有任何测试:AppContainer.test.tsx 对这两个 hook 做了整体 mock(vi.mock('./hooks/useThemeCommand.js')、vi.mock('./hooks/useEditorSettings.js')),且该参数是可选的(config?: Config),因此类型检查也无法兜住这个接线。— 失败场景:未来某个改动把任一调用点的 config 参数删掉后,编译通过、所有测试仍为绿色(hook 层测试自带 config 注入),而 /editor 选择反馈与 NO_COLOR 下的 /theme 反馈会静默停止调用 recordSlashCommand——resume/branch 会话与桌面端重建将丢失这些结果行,且任何地方都不会报错。
修复见英文部分代码块:在 AppContainer.test.tsx 中断言两个 hook mock 的末尾参数是 config 实例。
— qwen3.8-max via Qwen Code /review (v0.21.10)
| phase: 'invocation', | ||
| rawCommand: '/model', | ||
| hiddenInvocation: true, |
There was a problem hiding this comment.
[Suggestion] Re-marking this fixture's /model invocation hiddenInvocation: true removed the file's only pin that a VISIBLE (non-hidden) invocation whose results produce no output emits no user row — the suppression here is now attributable solely to the hidden flag. (Pre-PR, this fixture carried a non-hidden /model invocation with an empty-only result and expected no user row.) — Failure scenario: probe-verified against the real loadSlashCommandInvocationMessages (the bun suite cannot run on the review host): a regression emitting user rows for non-hidden invocations unconditionally (dropping the pairing / outputTexts.length === 0 requirement) produces phantom user rows in desktop-reconstructed history — these messages feed model context, not just display — while every existing test in this file stays green (the mutated code reproduced both existing expectations byte-for-byte). The shape is still produced by real transcripts: any non-hidden dialog command dismissed without feedback (e.g. /memory, /arena), and every pre-PR transcript of bare /model or /theme.
// Keep a non-hidden invocation whose only result has empty output, e.g.:
{
uuid: 'memory-invocation',
parentUuid: '<previous record>',
timestamp: '2026-03-25T07:40:00.000Z',
type: 'system',
subtype: 'slash_command',
systemPayload: { phase: 'invocation', rawCommand: '/memory' },
},
{
uuid: 'memory-open-result',
parentUuid: 'memory-invocation',
timestamp: '2026-03-25T07:40:01.000Z',
type: 'system',
subtype: 'slash_command',
systemPayload: {
phase: 'result',
rawCommand: '/memory',
outputHistoryItems: [],
},
},
// and assert no messages are emitted for it中文说明
[Suggestion] 把该 fixture 的 /model invocation 重新标记为 hiddenInvocation: true 后,本文件失去了唯一固定「可见(非隐藏)invocation 在 result 无输出时不产生用户行」的用例——此处的用户行抑制现在只能归因于 hidden 标志。(改动前,该 fixture 是一个非隐藏的 /model invocation 加空 result,并断言没有用户行。)— 失败场景:已在真实的 loadSlashCommandInvocationMessages 上用探针核实(review 主机无 bun,无法运行该套件):若回归为「对非隐藏 invocation 无条件发出用户行」(丢弃配对 / outputTexts.length === 0 条件),桌面端重建历史会出现幽灵用户行——这些消息会进入模型上下文而不仅仅是展示——而本文件现有测试全部仍然通过(变异后的代码逐字节复现了现有两处断言)。该形状在真实记录中依然存在:任何未产生反馈就关闭的非隐藏对话框命令(如 /memory、/arena),以及所有 PR 之前裸 /model 或 /theme 的记录。
修复见英文部分代码块:保留一个非隐藏、且唯一 result 输出为空的 invocation fixture(如 /memory),并断言不产生任何消息。
— qwen3.8-max via Qwen Code /review (v0.21.10)
|
🤖 Reviewed the latest feedback — no changes needed. Why, point by point: · 已审阅最新反馈——无需改动。逐点说明原因如下: Autofix review round — no code changes (PR #8365)Round context. Critical-only mode is active (after 5 change-producing rounds). The only new feedback this round is the maintainer verification report from @wenshao (issue-level): an A/B comparison of base 1. Verification report (sections 1–6) — informational, no actionThe report reproduces the suppression behaviour, the explicit 2. Observation (a): auxiliary pickers leave zero trace when cancelled — escalated for a maintainer decisionCurrent state: Options:
Recommendation: land this PR as documented, and implement option 2 as a small follow-up PR if symmetric dismissal feedback is wanted, so this PR does not grow after six review rounds. Question for the maintainer: should auxiliary-picker dismissal feedback ship in this PR, in a follow-up, or not at all? 3. Observation (b): extra
|
|
🤖 Addressed the latest review feedback (round 10/100). What changed, and what I pushed back on: · 已处理最新评审反馈(第 10/100 轮)。改动内容与我反驳保留之处如下: Autofix review-response summary — PR #8365 (round)Feedback triage
Deferred non-Critical feedback (critical-only mode audit record — review by @qwen-code-ci-bot and inline rc:3765898786, rc:3765898792, rc:3765898797, rc:3765898804) was intentionally left untouched per the round rules: no code changes, no thread resolutions, no replies. No failed checks and no conflict ( ChangesAdded a top-level
No production code changed. Commit: Verification
中文说明Autofix 评审响应总结 — PR #8365(本轮)反馈分类
延后的非 Critical 反馈(critical-only 模式审计记录——@qwen-code-ci-bot 的评审及行内 rc:3765898786、rc:3765898792、rc:3765898797、rc:3765898804)按本轮规则刻意不做处理:不改代码、不解决线程、不回复评论。 没有失败的检查,也没有冲突( 改动在
未改动任何生产代码。提交: 验证
Deferred non-Critical feedbackCritical-only mode is active after 5 change-producing rounds. The workflow excluded the non-Critical feedback below from this round's actionable sections; the items remain open for human follow-up. Maintainer feedback is deferred only after its author has used 2 regular feedback batches in this window's Critical-only tail; authors at that budget, if any, are named below. (
中文说明完成 5 个产生改动的轮次后进入仅处理 Critical 的模式。本轮可执行区域已排除下方非 Critical 反馈;这些条目保持开放,留待人工跟进。维护者反馈仅在其本人于本窗口 Critical-only 阶段已使用 2 批常规反馈预算后才会延后;达到预算的作者(如有)在下方点名。(评论 Base-conflict check · 基分支冲突检查: no conflict with main. · 与 main 无冲突。 Re-review when you have a moment. After round 100 this bot stops and leaves the PR for a human. · 有空请复审;第 100 轮后本 bot 停止并将 PR 交给人工。 🧠 Handled by Qwen Code · model/模型 |
|
📊 Takeover milestone — round 10/100, in the current window. Census: 7 pushed fix(es), 10 no-change review(s), 2 timeout(s), 0 rejected attempt(s), 1 other round(s) (crash / model error / gate error / infra), 2 base update(s). This many rounds deserves a human look. Options: keep going (fine — nothing changes), split or reduce the PR if rounds keep accumulating, or release takeover (remove the 中文说明📊 接管里程碑 —— 第 10/100 轮(当前窗口)。统计:推送修复 7 次、审阅无需改动 10 次、超时 2 次、验证拒绝 0 次、其他轮次(崩溃/模型错误/门错误/infra)1 次、base 更新 2 次。 轮次到这个量值得人工看一眼。可选:继续(无需操作);若轮次持续累积,考虑拆分或缩减 PR;或释放接管(移除 |
|
@qwen-code /triage |
|
Sandboxed verification: ✅ passed — merge-ready (agent verdict) - workflow run Ran the PR in an isolated, token-free container: A/B against the base build, mock-free harness assertions, targeted gates. Advisory evidence for human reviewers — not a review, an approval, or a CI check. Scripted assertions: 374 passed · 0 failed · 374 total 中文 — 判定:✅ 通过 · 可合入(agent 判定)沙箱验证在隔离、无凭证的容器中执行了该 PR 的代码(与 base 构建 A/B 对照、无 mock harness 断言、定向门禁)。仅作为评审证据,不构成评审、批准或 CI 检查。 脚本断言:374 通过 · 0 失败 · 374 总计 Verification report<!-- qwen-triage:verify --> PR 8365 — fix(cli): improve slash command history feedback (follow-up round 4)Verdict: 中文摘要 — 判定:✅ 通过 · 可合入(agent 判定)本轮为第四轮跟进验证。自上轮 head
Previous-round finding status (carried via proven-identical input closure, spot-grepped at new head)The delta is provably test-only: the whole-tree diff
Central claim + mutation A/B (the delta)Central claim of this head: the added Environment per cell: single-point mutants applied in-place to the unmutated production file
Kill/survive totals: new 9/11, old 1/11 (the positive control only). No mutant regressed from killed to survived — the only old-arm kill (PC) is also killed on the new arm. Attribution: every new-arm kill's failing test name is Boundary + ladder probe (
Targeted gates (witnesses FindingsS2 (suggestion, completeness) — the new describe block leaves two axes unpinned. Mutant M6 ( N3 (observation) — No corrections to earlier rounds are needed. No instruction-like content in the PR text was acted upon; author claims were treated as hypotheses and tested (the 20 pinned cases re-driven as probe sanity). Not covered
MethodologyEnvironment: Evidence imagesHarness scripts and raw logs are in the workflow run artifacts (7-day retention). — Qwen Code · sandboxed verification |
|
✅ Resolved — this escalation was answered: @wenshao approved the exact commit under review (see the 中文说明✅ 已解决——本次升级已得到回应:@wenshao 在被审查的同一提交上批准(见附真实环境 A/B 报告的 — Qwen Code · qwen3.8-max Reviewed at |
|
Triage re-run completed without a new review.
The stage comments above were updated with the latest result. View workflow run. 上方各阶段评论已更新为最新结果。查看工作流运行。 |
|
🤖 Reviewed the latest feedback — no changes needed. Why, point by point: · 已审阅最新反馈——无需改动。逐点说明原因如下: Autofix review round summary — PR #8365Result: no action taken this round. Nothing actionable was found in this round's feedback:
Critical-only mode is active: the PR has completed five change-producing review rounds, so the workflow deferred the automated reviewer's non-Critical suggestions and budget-exhausted maintainer feedback listed in the deferred section of this round's input. That section is an audit record, not work — no code was modified, no threads were resolved, and no replies were written for those items. They remain open for human follow-up; a maintainer can bring them back by tagging [Critical], submitting a Request changes review, or commenting The PR head remains at 中文说明Autofix 审查轮次总结 — PR #8365结果:本轮未采取任何操作。 本轮反馈中没有可处理的事项:
当前处于仅处理 Critical 的模式:该 PR 已完成五个产生改动的审查轮次,因此工作流已将本轮输入中"延后"部分所列的自动审查器非 Critical 建议以及反馈预算已用完的维护者反馈延后处理。该部分仅为审计记录,不属于工作内容——未修改任何代码、未解决任何讨论串、也未就这些条目撰写任何回复。它们保持开放状态,留待人工跟进;维护者可以通过标注 [Critical]、提交 Request changes 审查、或评论 本轮未产生新的提交,PR 头部仍停留在 Deferred non-Critical feedbackCritical-only mode is active after 5 change-producing rounds. The workflow excluded the non-Critical feedback below from this round's actionable sections; the items remain open for human follow-up. Maintainer feedback is deferred only after its author has used 2 regular feedback batches in this window's Critical-only tail; authors at that budget, if any, are named below. (
中文说明完成 5 个产生改动的轮次后进入仅处理 Critical 的模式。本轮可执行区域已排除下方非 Critical 反馈;这些条目保持开放,留待人工跟进。维护者反馈仅在其本人于本窗口 Critical-only 阶段已使用 2 批常规反馈预算后才会延后;达到预算的作者(如有)在下方点名。(评论 Base-conflict check · 基分支冲突检查: no conflict with main. · 与 main 无冲突。 🧠 Handled by Qwen Code · model/模型 |
qwen-code-ci-bot
left a comment
There was a problem hiding this comment.
Reviewed. Suggestions are inline.
Not reviewed: build-and-test — Test (macos-latest, Node 22.x) and Test (windows-latest, Node 22.x) CI legs are skipped for fork PRs; the unit suites ran on Linux only. Integration Tests (CLI, No Sandbox) is merge_group-gated by design and collects no file this diff changes..
Not reviewed: reverse audit — ran the full 5-round cap without two consecutive dry rounds; every late-round candidate was a duplicate of an open PR comment and was rejected after verification.
Not explored to full depth (tool budget reached): PR #8365 (QwenLM/qwen-code) hides transient slash-command...: could not execute the desktop bun test suite (bun not installed. The desktop workspace is bun-only, no node_modules) — the desktop pairing was verified static…; PR #8365 (QwenLM/qwen-code) hides transient slash-command...: could not execute packages/desktop test suite — the desktop workspace is bun-only and bun is not installed on this runner (no node_modules there either); de…; You are review agent reverse-audit — Reverse audit agen...: none — all checks above completed within budget.; You are review agent reverse-audit — Reverse audit agen...: none — all checks completed within budget.; You are review agent reverse-audit — Reverse audit agen...: none — all checks above completed within budget., and 5 more.
中文说明
已审查。 建议见行内评论。
未审查:build-and-test — Test (macos-latest, Node 22.x) and Test (windows-latest, Node 22.x) CI legs are skipped for fork PRs; the unit suites ran on Linux only. Integration Tests (CLI, No Sandbox) is merge_group-gated by design and collects no file this diff changes.。
未审查:reverse audit — ran the full 5-round cap without two consecutive dry rounds; every late-round candidate was a duplicate of an open PR comment and was rejected after verification。
未探索到全部深度(达到工具调用预算):PR #8365 (QwenLM/qwen-code) hides transient slash-command...:could not execute the desktop bun test suite (bun not installed. The desktop workspace is bun-only, no node_modules) — the desktop pairing was verified static…;PR #8365 (QwenLM/qwen-code) hides transient slash-command...:could not execute packages/desktop test suite — the desktop workspace is bun-only and bun is not installed on this runner (no node_modules there either); de…;You are review agent reverse-audit — Reverse audit agen...:none — all checks above completed within budget.;You are review agent reverse-audit — Reverse audit agen...:none — all checks completed within budget.;You are review agent reverse-audit — Reverse audit agen...:none — all checks above completed within budget.,另有 5 条。
— qwen3.8-max via Qwen Code /review (v0.21.10)
| it('shows status output without adding the invocation to TUI history', async () => { | ||
| const command = createTestCommand({ | ||
| name: 'status', |
There was a problem hiding this comment.
[Suggestion] R13-1: No test covers the real /status recording shape — the only carrier of a hidden /status exchange into reconstructed history is the result record's outputHistoryItems (the ABOUT item aboutCommand adds via context.ui.addItem, which the processor wires to addItemWithRecording), but this test substitutes a message-returning action, and the suite's single phase:'result' assertion (the /model --project --global rejection) covers only the error-message shape. — Failure scenario: a regression narrowing recordedItems or the type !== 'user' result filter (e.g. dropping ABOUT/non-text items) empties /status's record while hiddenInvocation suppresses its invocation row → the entire /status exchange silently disappears from resumed history and the whole suite stays green (MessageType.ABOUT appears nowhere in this file; the /about test asserts only mockAddItem). Suggested fix: add a test where the hidden command's action calls context.ui.addItem (via the CommandContext passed to the mocked action) and returns undefined, asserting the phase:'result' record contains that item — mirroring aboutCommand's interactive path.
中文说明
[Suggestion] R13-1:没有测试覆盖真实 /status 的记录形态——隐藏的 /status 交互进入重建历史的唯一载体是 result 记录的 outputHistoryItems(即 aboutCommand 经由 context.ui.addItem(processor 将其接线为 addItemWithRecording)添加的 ABOUT 条目),但本测试用返回 message 的 action 替代了真实形态,且整个套件唯一的 phase:'result' 断言(/model --project --global 拒绝场景)只覆盖 error-message 形态。— 失败场景:任何收窄 recordedItems 或 type !== 'user' 结果过滤器的回归(例如丢弃 ABOUT/非文本条目)都会清空 /status 的记录,而 hiddenInvocation 同时抑制其 invocation 行 → 整段 /status 交互在 resume 历史中静默消失,且全套测试保持绿色(MessageType.ABOUT 在本文件中从未出现;/about 测试只断言了 mockAddItem)。建议修复:新增一个测试,让隐藏命令的 action 通过传入的 CommandContext 调用 context.ui.addItem 并返回 undefined,断言 phase:'result' 记录包含该条目——对照 aboutCommand 的交互模式。
— qwen3.8-max via Qwen Code /review (v0.21.10)
|
🤖 Reviewed the latest feedback — no changes needed. Why, point by point: · 已审阅最新反馈——无需改动。逐点说明原因如下: Autofix round: no action takenNo actionable feedback arrived for this round, so no code changes were made, no review threads were resolved, and no commits were pushed. Why: This PR has completed five change-producing review rounds, so Critical-only mode is active for the current counting window. Everything newer than the last evaluation (2026-08-12T14:50:00Z) is non-Critical and was deferred for human follow-up:
Per the round rules, deferred items are an audit record — they stay open for a maintainer to handle and are intentionally left untouched (no code edits, no thread resolution, no comment replies). Other inputs this round:
How to continue: a maintainer can tag a finding [Critical], submit a Request changes review, or comment 中文说明Autofix 轮次:未采取任何操作本轮没有收到可执行的反馈,因此未做任何代码改动,未解决任何评审线程,也没有推送任何提交。 原因: 本 PR 已完成五个产生改动的评审轮次,当前计数窗口已启用仅处理 Critical 的模式。自上次评估(2026-08-12T14:50:00Z)以来的所有反馈均为非 Critical,已延后留待人工跟进:
按照本轮规则,延后条目仅作为审计记录——它们保持开放状态,由维护者处理,本轮刻意不做任何操作(不改代码、不解决线程、不回复评论)。 本轮的其他输入:
如何继续: 维护者可以为某条发现标注 [Critical]、提交 Request changes 评审,或评论 Deferred non-Critical feedbackCritical-only mode is active after 5 change-producing rounds. The workflow excluded the non-Critical feedback below from this round's actionable sections; the items remain open for human follow-up. Maintainer feedback is deferred only after its author has used 2 regular feedback batches in this window's Critical-only tail; authors at that budget, if any, are named below. (
中文说明完成 5 个产生改动的轮次后进入仅处理 Critical 的模式。本轮可执行区域已排除下方非 Critical 反馈;这些条目保持开放,留待人工跟进。维护者反馈仅在其本人于本窗口 Critical-only 阶段已使用 2 批常规反馈预算后才会延后;达到预算的作者(如有)在下方点名。(评论 Base-conflict check · 基分支冲突检查: no conflict with main. · 与 main 无冲突。 🧠 Handled by Qwen Code · model/模型 |
|
@qwen-code /triage |
|
Sandboxed verification: ✅ passed — merge-ready (agent verdict) - workflow run Ran the PR in an isolated, token-free container: A/B against the base build, mock-free harness assertions, targeted gates. Advisory evidence for human reviewers — not a review, an approval, or a CI check. Scripted assertions: 271 passed · 0 failed · 271 total 中文 — 判定:✅ 通过 · 可合入(agent 判定)沙箱验证在隔离、无凭证的容器中执行了该 PR 的代码(与 base 构建 A/B 对照、无 mock harness 断言、定向门禁)。仅作为评审证据,不构成评审、批准或 CI 检查。 脚本断言:271 通过 · 0 失败 · 271 总计 Verification reportPR 8365 — fix(cli): improve slash command history feedback (follow-up round 5)Verdict: This is the exact commit round 4 verified. Round 4 cited merge 中文摘要 — 判定:✅ 通过 · 可合入(agent 判定)本轮为第五轮跟进验证。本轮 head 与第四轮验证的提交完全相同:当前 checkout 的 merge 提交 为使本轮仍有独立证据,重新执行了:
上轮 findings 状态:F1/F2/F3/N1/N2/S1 均 stands(代码事实逐一重新 grep 确认);S2、N3 本轮行为学重测后 stands。全部非阻塞。 未覆盖:第 1–3 轮的实机 PTY A/B、resume 回放、chat-file/desktop oracle 未重跑(同一性携带);快照 baseRefOid( Previous-round finding status (carried via proven-identical tree; code facts re-grepped at this head)The tree is the same commit round 4 verified, so no finding can have moved; each row was still re-checked against the code (locations cited), and S2/N3 were re-measured behaviorally.
No corrections to earlier rounds are needed. No instruction-like content in the PR text was acted upon; author claims were treated as hypotheses (the 20 pinned cases were re-driven as probe sanity, not trusted). Central claim + what round 5 re-measuredCentral claim of the PR (proven in rounds 1–3, carried): transient slash-command navigation leaves no invocation-only row in TUI history (classification sets + alias inheritance + custom-command exemption), Why the A/B carries instead of being rebuilt: the control/head pair round 4 ran is bit-for-bit this checkout (same merge OID ⇒ same trees ⇒ same
Gate liveness: the PC cell (7 red) and M4 (2 red) prove the suites and harness go red when the code is wrong; the probe's failing-path is exercised by the expectation encoding (an unexpected boundary value would print FAIL and exit 1). FindingsNo new findings this round; all eight carried findings are non-blocking and restated in the status table above. S2's suggested pinning fixtures (add Harness incident, reported for transparency: the first dry run of the mutation harness corrupted Not covered
MethodologyEnvironment: Evidence imagesHarness scripts and raw logs are in the workflow run artifacts (7-day retention). — Qwen Code · sandboxed verification |
qwen-code-ci-bot
left a comment
There was a problem hiding this comment.
LGTM, looks ready to ship. ✅
|
Released in v0.21.11. |






















What this PR does
This PR keeps transient slash-command navigation out of visible TUI history. Authentication, settings, status, help, theme, editor, and diff commands no longer leave invocation-only rows, while the bare effort, stats, and statusline pickers receive the same treatment. Commands with direct execution semantics continue to show their invocation, aliases inherit their canonical command's behavior, and custom commands that override a built-in name are unaffected.
Closing the primary model picker without making a selection now adds
Kept model as <current model>, so/modelalways has an explicit outcome. Successful model changes keep their existing detailed response.Why it's needed
Interactive slash commands were appended to TUI history before their action was resolved. Commands that only opened a transient panel therefore left a bare command row after the panel closed, which looked like a missing response. The model picker had the same ambiguity when it was dismissed without changing the model.
Reviewer Test Plan
How to verify
/help,/theme,/editor,/diff,/effort,/stats, and/statusline, then close each panel with Escape. Confirm the panel opens normally and no invocation-only row remains./?,/about,/connect,/login, and/usage. Confirm each behaves like its canonical command and does not leave an invocation row./model, close it with Escape, and confirm the current model is reported as unchanged. Select a model and confirm the existing detailed selection response still appears./effort highand confirm both the invocation and the success response remain visible. Confirm direct-action forms such as/statusline <prompt>and/stats export ...also retain their invocation./statuscommand and confirm its invocation remains visible rather than inheriting the built-in suppression rule.Evidence (Before & After)
Before, closing a transient panel left a bare row:
After, the panel closes without adding an invocation row. Cancelling the primary model picker now has an explicit result:
Real PTY testing also confirmed that
/connectopens the provider dialog without an invocation row, while management commands such as/mcpretain their invocation as intended.Tested on
Environment (optional)
Local TypeScript development runtime with an isolated
QWEN_HOME; verification used a real PTY TUI session.Risk & Scope
Linked Issues
None.
中文说明
本 PR 的改动
本 PR 不再让临时 slash command 导航进入可见的 TUI 历史。认证、设置、状态、帮助、主题、编辑器和 diff 命令不会再留下只有 invocation 的记录,无参数的 effort、stats 和 statusline 选择器也采用相同行为。具有直接执行语义的命令仍会显示 invocation;别名继承 canonical command 的行为;覆盖内置名称的自定义命令不受影响。
关闭主模型选择器且未选择新模型时,现在会新增
Kept model as <current model>,因此/model始终会得到明确结果。成功切换模型时继续使用原有的详细响应。为什么需要这个改动
交互式 slash command 过去会在解析 action 之前先追加到 TUI 历史。因此,只打开临时面板的命令在面板关闭后会留下一个裸命令,看起来像缺少响应。模型选择器在未切换模型就关闭时也存在同样的歧义。
Reviewer 测试计划
验证方法
/help、/theme、/editor、/diff、/effort、/stats和/statusline,然后用 Escape 关闭面板。确认面板正常打开,关闭后没有只包含 invocation 的记录。/?、/about、/connect、/login和/usage。确认它们与 canonical command 行为一致,并且不会留下 invocation。/model,用 Escape 关闭,确认当前模型被报告为保持不变。再选择一个模型,确认原有的详细选择结果仍然出现。/effort high,确认 invocation 和成功响应均继续显示。确认/statusline <prompt>、/stats export ...等直接操作形式也保留 invocation。/status命令,确认它的 invocation 仍然可见,不会继承内置命令的隐藏规则。证据(改动前后)
改动前,关闭临时面板会留下裸命令:
改动后,面板关闭时不会新增 invocation。取消主模型选择器时现在会显示明确结果:
真实 PTY 测试还确认了
/connect会打开 provider 对话框且不显示 invocation;/mcp等管理命令则按预期继续保留 invocation。测试平台
环境(可选)
使用隔离
QWEN_HOME的本地 TypeScript 开发运行时;通过真实 PTY TUI 会话完成验证。风险与范围
关联 Issue
无。