fix(review): name each certification bar and defer degrade notes past admission (#9259) - #9272
Conversation
|
Thanks for the continued work! Template — complete, with the bilingual translation ✓ Problem — observed, not theoretical. Same footing as the last pass: this lands the deferred Suggestions recorded in #9259 (each names the #9213 review thread it came from) on top of the original never-retire bug #9206 (observed in a real 12-chunk run, since closed). The production changes map 1:1 to recorded findings. Direction — aligned. This is the diagnostic channel of the repo's own reverse-audit retirement machinery. No auth/sandbox/model-selection/public-contract surface — the changed diagnostic strings are stderr NOTE text with no consumer anywhere in the tree (re-verified at the new head: the retired Size — not core-module territory ( Approach — the scope still feels right. Three small production changes plus test-pinning traced to the mutation each pin names. The three substantive rounds since the last triage pass (rounds 5, 7, 8) kept evolving the same two spots instead of widening scope: the polarity-guard family — each lookahead shape falsified by execution the round it landed — was replaced by a single restatement bar that closes the whole passive/reduced/spliced entrance family by form, and the degrade-note claim fence moved from file mtimes to the claim's own content and shape. Risk — no elevated risk signals: none of the changed files matches the revert-correlated high-risk paths. Re-run note: the head moved from Moving on to code review. 🔍 中文说明感谢持续的跟进! 模板 — 完整,含中文翻译 ✓ 问题 — 已观测而非理论。与上一轮相同的依据:落地 #9213 评审线程中延期、记录在 #9259 的 Suggestions(每条都注明来源线程),其基础是原始的静默不退役 bug #9206(在真实 12-chunk 运行中观测到,已关闭)。生产改动与记录一一对应。 方向 — 对齐。这是仓库自身反向审计退役机制的诊断通道。不触及 auth/沙箱/模型选择/公共契约——改动的诊断字符串只是 stderr NOTE 文本,整个代码树中无任何消费者(已在新 head 复查:退役的 规模 — 非核心模块路径( 方案 — 范围依然合理。三项小的生产改动,其余均为溯源到各自命名变异的测试钉住。上次 triage 之后的三个实质轮次(第 5、7、8 轮)持续演化同一两处而非扩张范围:极性守卫家族——每个 lookahead 形态都在落地的当轮被执行证伪——被单一的复述门槛取代,以形式封闭整个被动/缩减/拼接入口家族;降级 NOTE 认领的围栏从文件 mtime 改为认领自身的内容与形状。 风险 — 无升级风险信号:改动文件均未命中与 revert 相关的高风险路径。 Re-run 说明:head 从 进入代码审查 🔍 — Qwen Code · qwen3.8-max Reviewed at |
Code reviewIndependent proposal first (written before reading the diff): split the collapsed failure name into one bar per cause, move both schedule-catch NOTEs past the admission gate, re-key the per-chunk suppression off the admission stamp onto an atomic per-round claim file, and — given that rounds 4–6's executed probes kept falsifying polarity lookaheads — stop enumerating entrance shapes and refuse any clause that restates the all-clear core, with the brief mandating the same. That is exactly what shipped; no simpler path missed. No Critical findings at the new head. Verified in the code at the reviewed commit, not just in the diff — the three substantive rounds since the last pass are the focus:
Three non-blocking observations:
Carried from the review job's round 6 (recorded, explicitly non-blocking): eleven probes on test-label overclaims and one corner of Test evidence — the PR's own CI, fetched via API (PR code is never executed in triage)All checks on this head are completed; everything that ran is green — including
One row per check name (latest run); skipped checks omitted; failures sort first. / 每个检查名一行(取最新一次运行),省略 skipped,失败项排在最前。 Not verified: the claimed 533/533 count and the per-mutation pin checks, because triage does not run PR-derived code — CI substantiates the suite passing; whether the pins bite is the Sandboxed verification would settle the one open claim: 中文说明代码审查先写独立方案(读 diff 之前):把归并的失败名拆成每因一门槛、把两处调度 catch NOTE 移到准入门之后、把逐 chunk 抑制键从准入戳记换成按轮次的原子认领文件,并且——既然第 4–6 轮的执行探针持续证伪极性 lookahead——停止枚举入口形状,直接拒绝任何复述全清核心的从句,brief 同步强制。落地的正是这些;没有发现更简路径。 新 head 无 Critical 发现。 经被审提交的代码本身验证,重点在上次之后的三个实质轮次:
三点非阻塞观察:
承接 review 作业第 6 轮(已记录、明确非阻断):11 条关于测试标签过度声明的探针,以及 测试证据 —— PR 自身 CI,经 API 获取(triage 从不执行 PR 代码)本 head 所有检查均已完成,执行的全部绿——包括运行本 PR 所指六个测试文件的 未验证:533/533 计数与逐变异钉住核查——triage 不执行 PR 代码;CI 证实套件通过,钉住是否咬合由 沙箱验证可以了结唯一未决声明: — Qwen Code · qwen3.8-max Reviewed at |
|
Confidence: 4/5 — clean, minimal, exhaustively pinned follow-up; the reservations are the non-blocking nits named above (two stale spots in the PR body, the marginally-strong equivalence comment, and mutation numbers that rest on the author's word until Stepping back: the three rounds since the last pass made this PR better in exactly the way the deferral discipline hopes for. Each shipped guard shape that execution falsified was replaced rather than patched — the polarity lookaheads gave way to a single restatement bar that closes the whole entrance family by form, with the brief updated to mandate the same, and the mtime fence (the review job's R5-7) gave way to a content-and-shape fence because mtimes are not reliable across runners. The two round-5 Criticals are resolved: the review job's round 6 came back with an empty findings ledger, and my own pass at the new head found no Critical either — gate placement verified in both build paths, the bar order matching the pins end to end, every reachable claim failure still failing toward printing. CI is green and settled on this head, including the ubuntu suite that landed red two passes ago. The body drift is cosmetic; in six months the pin-per-mutation discipline here is something to thank, not curse. Approving, pinned to the reviewed commit. If the pin-bearing claim matters before merge, 中文说明置信度:4/5 —— 干净、最小化、钉住充分的后续 PR;保留意见是上方点名的非阻塞小项(PR 正文两处滞后、略强的等价性注释、以及在 退一步看:上次之后的三个轮次让这个 PR 恰好按延期纪律期望的方向变好。每个被执行证伪的守卫形态都被替换而非打补丁——极性 lookahead 让位于以形式封闭整个入口家族的单一复述门槛,brief 同步强制同样的规则;mtime 围栏(review 作业 R5-7)让位于内容与形状围栏,因为 mtime 在不同 runner 间不可靠。第 5 轮的两个 Critical 已解决:review 作业第 6 轮返回空的发现台账,我在新 head 的独立审查也未发现 Critical——两条构建路径的门位置均经代码验证,门槛顺序与钉住端到端一致,所有可达的认领失败仍倒向打印。本 head 的 CI 全绿且已稳定——包括两个 pass 前曾红的 ubuntu 套件。正文滞后只是表面;六个月后回看,这种"每钉对应一个变异"的纪律值得感谢而非抱怨。 批准,钉住被审提交。若合并前想独立核实钉住有效性, — Qwen Code · qwen3.8-max Reviewed at |
…#9091 findings-read bar)
|
@qwen-code /takeover |
|
🤝 Takeover engaged: the autofix loop now manages this PR — it will address new review feedback and resolve base conflicts until the label is removed or the round cap is reached. This is a fork PR, so the first round comes from the next scheduled scan (usually within minutes). Remove the 中文说明🤝 已接管:autofix 循环现在管理此 PR —— 将持续处理新的评审反馈与 base 冲突,直到移除标签或达到轮次上限。本 PR 来自 fork,首轮处理将由下一次定时扫描执行(通常几分钟内)。移除 |
qwen-code-ci-bot
left a comment
There was a problem hiding this comment.
Partially reviewed — gaps disclosed.
Not reviewed: reverse audit — did not converge within the reverse-audit round cap of 5.
中文说明
仅完成部分审查,审查缺口已披露。
未审查:反向审计——在 5 轮的反审轮数上限内未收敛。
— qwen3.8-max via Qwen Code /review (v0.21.12)
…ss-process (QwenLM#9272 round 2)
|
🤖 Reviewed the latest feedback — no changes needed. Why, point by point: · 已审阅最新反馈——无需改动。逐点说明原因如下: Autofix round — no code change needed: all 9 findings re-verified as already fixed at HEADThis round found no new actionable work. All 9 findings from review Critical findings
Suggestions
The nine reply comments ( Verification
中文说明Autofix 轮次 — 无需代码变更:全部 9 条发现已在 HEAD 修复并重新验证本轮没有新的可行动工作。审查 Critical 发现
Suggestion 发现
九条回复评论( 验证
Base-conflict check · 基分支冲突检查: no conflict with main. · 与 main 无冲突。 🧠 Handled by Qwen Code · model/模型 |
qwen-code-ci-bot
left a comment
There was a problem hiding this comment.
Partially reviewed — gaps disclosed.
Not explored to full depth (tool budget reached): "agent 1c": none — but note I did not execute the six test suites (another lane owns test runs); my verification was static cross-file tracing..
Not reviewed: reverse audit — did not converge within the reverse-audit round cap of 5.
中文说明
仅完成部分审查,审查缺口已披露。
未探索到全部深度(达到工具调用预算):"agent 1c":none — but note I did not execute the six test suites (another lane owns test runs); my verification was static cross-file tracing.。
未审查:反向审计——在 5 轮的反审轮数上限内未收敛。
— qwen3.8-max via Qwen Code /review (v0.21.12)
…e-note claim (QwenLM#9272 round 3)
qwen-code-ci-bot
left a comment
There was a problem hiding this comment.
Reviewed — no blockers. Suggestions are inline.
中文说明
已审查——无阻断问题。 建议见行内评论。
— qwen3.8-max via Qwen Code /review (v0.21.12)
|
🤖 Addressed the latest review feedback (round 1/100). What changed, and what I pushed back on: · 已处理最新评审反馈(第 1/100 轮)。改动内容与我反驳保留之处如下: 🤖 Addressed the round-2 review findings — one reproduced defect fixed in Autofix round — R2-1's fix claim was partially falsified by probe; mechanism 3 now closedThis round re-verified all five findings from review R2-1
|
…ew-retirement-followups-9259
|
@qwen-code /triage |
|
Sandboxed verification: ❌ not passed — findings reported (agent verdict) - workflow run Ran the PR in an isolated, token-free container: A/B against the base build, mock-free harness assertions, targeted gates. Advisory evidence for human reviewers — not a review, an approval, or a CI check. Scripted assertions: 1062 passed · 0 failed · 1062 total 中文 — 判定:❌ 不通过 · 报告了发现(agent 判定)沙箱验证在隔离、无凭证的容器中执行了该 PR 的代码(与 base 构建 A/B 对照、无 mock harness 断言、定向门禁)。仅作为评审证据,不构成评审、批准或 CI 检查。 脚本断言:1062 通过 · 0 失败 · 1062 总计 Verification reportPR #9272 deep verification —
|
| Cell | BASE (control) | HEAD (PR) |
|---|---|---|
bar-lead-hedge (…found but only skimmed.) |
unknown — receipt clause not substantive |
unknown — receipt lead contradicts the phrase |
bar-clause-admission (…but skipped the generated files) |
unknown — collapsed | unknown — receipt clause contradicts the phrase |
bar-no-walk (everything looks fine this round) |
unknown — collapsed | unknown — receipt clause names no walk |
bar-too-thin (walked lexing) |
unknown — collapsed | unknown — receipt clause too thin |
leak-passive-issues (…; no issues were verified.) |
dry — admission certified | unknown — polarity |
leak-passive-findings (…; no findings were verified.) |
dry — admission certified | unknown — polarity |
leak-passive-gaps (…; no gaps are verified outstanding.) |
dry — admission certified | unknown — polarity |
leak-adverb (…; no issues at all were verified.) |
dry — admission certified | unknown — polarity |
leak-filler-seat (there were no issues verified this round…) |
unknown — collapsed (walk gate, by accident¹) | unknown — polarity, named |
control-object-seat (verified no issues in it or its callers) |
dry | dry |
residue-honest (verified no regressions…) |
unknown — collapsed | unknown — polarity (see F1) |
control-passive-regressions (no regressions were verified) |
unknown — collapsed | unknown — polarity |
| control-clean-dry (canonical receipt) | dry | dry |
4 of 5 leak shapes flip from dry (base certified the admission) to unknown (head refuses it); all four bars flip from a collapsed name to self-naming; all 4 controls are byte-stable. Head 13/13, base 13/13 against the measured expectation tables.
¹ Base refused the filler-seat shape only because its filler-greedy saturation strip swallowed verified at the walk gate — the collapsed failure name made this coincidence indistinguishable from a polarity refusal, which is precisely the diagnosability gap change 1 closes. My initial base prediction for this cell was dry; measurement corrected it (disclosed in Methodology).
Sibling sweep of the strip guard (head): 7/7 adjacent shapes — aux-have passive, participial, traced, no new issues, singular no issue all refused; echo shapes (no gaps: none., doubled no issues found) still strip and retire. Scaling ladder on the new lookahead regex: linear — 2k/10k/40k-char hostile clauses classified in 0.3 / 0.5–1.1 / 1.1–2.7 ms (2 s cap), both the refusal and the admit scan.
A/B proof — harness 2: degrade-NOTE truthfulness
Four scenarios driven through the real agentPromptCommand.handler (real fs; only the stdio writers mocked — the observation seam). Evidence: 03-ab-notes-head.png, 04-ab-notes-base.png.
| Scenario | BASE | HEAD | flip |
|---|---|---|---|
S1 round-cap-refused round, unreadable history (--all-chunks) |
NOTE printed, THEN refusal | refusal only, no NOTE | ✅ |
| S2 c13 admission build, schedule reads cleanly | silent | silent | — |
| S2 c14 repair build after clean admission (history died) | silent — stamp-keyed suppression | NOTE once, reason carried (unavailable this round — <why> — auditing the chunk.) |
✅ |
| S2 c15 second repair, same round | silent | silent — claim spent | — |
| S2 r4 later round, history still dead | NOTE | NOTE — per-round claim | — |
S3 round-cap-refused round (--chunk) |
NOTE printed, THEN refusal | refusal only | ✅ |
| S4 ADMITTED round, unreadable history | NOTE, 3 chunks built | NOTE, 3 chunks built | — (deferral does not swallow the note) |
Both arms 4/4 against arm-specific expectations. S2-c14 is the exact never-retire-with-no-word shape named in the PR: base silences it, head names it once.
Concurrency probe (the claim exists for inter-process exclusion): 24 parallel node processes claiming the same round — exactly 1 CLAIMED, 23 silent; sequential second claim false; a different round claims independently; re-capturing the plan (retry run) re-arms. Evidence: 05-claim-concurrency.png. Claim-file lifecycle checked against cleanup.ts: written mid-run (mtime > plan) it cannot fake the previous-run retention predicate; it rides the record dir's retention, no orphaned artifact, and nothing reads its content for a decision.
Vacuity / mutation matrix
Scratch worktree at HEAD, one mutation at a time, suites re-run (unmutated control is the green gate above):
| Mutant | Named pin(s) | Result |
|---|---|---|
M0 remove unexamined from marker list (positive control) |
the un-examined admission case | 1 red — exactly the named pin; suite proven live |
| M1 revert bar split to collapsed base names | the four bar-naming tests + admission diagnostics | 19 red, all name mismatches — due/skipped still pass under the mutant, proving the per-domain polarity split is behavior-preserving as claimed |
| M2 remove strip lookahead (restore base phrase-core) | strip-dead-noun family | 5 red — 4 of them because the mutant retires the admission (due=[]), the 5th flips bar attribution |
| M3 move round-builder NOTE back before the gate | refused-round NOTE test | 1 red — mutant prints the NOTE on a refused round |
M4 claim create flag wx→w |
once-per-round claim test + chunk-15 silence | 2 red |
No survivors. Every guard the PR introduces is pinned by a test that fails the intended behavioral assertion (failure messages quoted in logs/m0–m4.log show expected-vs-actual, not import/compile breakage).
Targeted gates
| Gate | Result |
|---|---|
| Six suites at HEAD (the PR's own list) | 518/518 pass (body says 509, 中文 says 500 — see F2) |
| Same six suites at BASE with base's test files | 485/485 pass → +33 tests, 0 regressions, all deltas are this PR's pins |
tsc --noEmit in packages/cli |
exit 0 |
Findings
F1 — Reviewer Test Plan claims the opposite of the merged behavior (non-blocking; PR description)
The plan's second behavior says: "a clause like No issues found — verified no regressions in the reconnect path and re-walked its call sites. now certifies dry instead of reading unknown."
Measured: that exact return reads unknown — receipt clause contradicts the phrase — on BOTH arms (A/B cell residue-honest), and the PR's own dedicated test ("an honest absence-of-problems receipt stays under audit — the accepted residue (#9272)") pins exactly that. The sentence describes the absence-of-problems exception class that rounds 2–3 tried and round 4 removed fail-closed — PR body §3 states the residue correctly (verified no regressions reads unknown, chunk stays under audit). So the code and the body agree; the Test Plan step contradicts both. A reviewer walking the plan step by step meets the opposite of what it promises.
Suggested fix (prose only, nothing to measure in code): replace the second behavior with e.g. "…reads unknown and the chunk stays under audit — the accepted residue, pinned by its own test; the certification that changed is the bar naming and the NOTE timing above."
F2 — test-count drift in the PR body (nit)
English body: "expect 509/509"; 中文: "500/500"; measured 518 at head / 485 at base. Counts drifted across the PR's four review rounds; the verified numbers are the ones above.
Not covered
- Per-commit attribution: depth-2 checkout —
git rev-list HEAD^1..HEAD^2returns 1 (the head merge) vs 7 commits in the metadata snapshot; the shallow boundary is proven, not assumed. Only the aggregateHEAD^1..HEADdiff was verified. - Live-model E2E: declared out of scope by the PR itself; not attempted. Harnesses reproduce the wire shapes (records + transcripts as the harness writes them), not an orchestrator run.
- Repo-wide gates: only the affected workspace's six suites + package typecheck ran; no repo-wide
npm run lint/npm run test. - Windows:
wxatomicity relies on Node's O_EXCL mapping; measured on Linux only. The PR self-declares Windows untested. - Run 1 of the base arm of harness 1 was 12/13: my prediction for
leak-filler-seatwasdry; measurement showed base refuses it at the walk gate (footnote ¹). The expectation table was corrected to measured ground truth and re-run clean — a harness-prediction error, not a PR defect, disclosed per the counts-are-sacred rule. - The 24-process probe exercises the claim function directly; a real multi-process orchestrator run was not staged.
Methodology
Environment: node v22.23.2, CI container, merge-ref checkout (HEAD = merge, HEAD^1 = base tip, HEAD^2 = verified head). Control hygiene: base worktree at HEAD^1; the PR touches no package.json/package-lock.json (verified by diff), so per-package node_modules (ajv etc. are not hoisted to root) were symlinked head→base — third-party deps byte-identical, only source differs; vitest's alias resolves @qwen-code/qwen-code-core into the worktree's own source (observed base-tree paths in transform output), so the base arm ran base code. Harnesses: zz-verify-ab-classifier.test.ts (13-cell matrix), zz-verify-ab-notes.test.ts (4 handler-driven scenarios), zz-verify-siblings.test.ts (sweep + ladder), claim-concurrency.mjs (24-process probe) — all preserved in this dir for rerun; expectation tables in expect-*.json. Mutation edits applied via scripted exact-block replacement in a scratch worktree (logs m1–m4.log + M0 inline). Raw gate logs: logs/head-six-suites.log. Counts: 518 (head gate) + 485 (base gate) + 26 (A/B1) + 8 (A/B2) + 13 (siblings/ladder) + 6 (concurrency comparisons) + 5 (mutation verdicts; the 28 expected-red tests roll up into these five) + 1 (typecheck) = 1062 pass / 0 fail.
Evidence images
Harness scripts and raw logs are in the workflow run artifacts (7-day retention).
— Qwen Code · sandboxed verification
qwen-code-ci-bot
left a comment
There was a problem hiding this comment.
Partially reviewed — gaps disclosed.
Not reviewed: reverse audit — did not converge within the reverse-audit round cap of 5.
中文说明
仅完成部分审查,审查缺口已披露。
未审查:反向审计——在 5 轮的反审轮数上限内未收敛。
— qwen3.8-max via Qwen Code /review (v0.21.12)
qwen-code-ci-bot
left a comment
There was a problem hiding this comment.
LGTM, looks ready to ship. ✅
…wenLM#9272 round 5) The polarity guard named walk verbs one by one; round 4's probe retired chunks on five entrance families of passive admissions (non-walk participles, hyphenated and prefixed ones, NBSP-split runs, parentheticals). Close the surface on a closed class instead: refuse the strip when the word-run after `no <noun>` reaches a passive head (the be/get auxiliaries) or a walk verb, across every whitespace width and parenthetical punctuation, exempting runs that carry the phrase's own filler `found` so honest echoes keep stripping. Group the EN alternative before the guard binds. Also split the degrade-note claim's mkdir into its own fail-open catch (a recursive mkdir's EEXIST on a regular file is not a claim), and re-pin the fail-open tests on stimuli that fault for any uid — permission bits do not fault for root.
|
🤖 Addressed the latest review feedback (round 2/100). What changed, and what I pushed back on: · 已处理最新评审反馈(第 2/100 轮)。改动内容与我反驳保留之处如下: Round 5 summary — PR #9272Commit: Findings and dispositionsR4-1 — [Critical] the walk-verb guard is a shape-by-shape patch (rc:3793116734) — FIXEDReproduced first, on the pre-round tree: all five entrance families retired the chunk on receipts that admit the check was not done — 11 new tests failed at HEAD ( The fix closes the surface on a closed grammatical class instead of enumerating participles: the marker strip now refuses a R4-7 — the guard fired on honest receipts (rc:3793116751) — FIXEDThe fix above carries the exemption the finding suggested: a word-run carrying the phrase's OWN filler R4-6 — guard binds only because EN is one alternation (rc:3793116750) — FIXEDEN is grouped at the composition site — R4-2 — mkdir's EEXIST conflated with the claim's (rc:3793116736) — FIXED
R4-3 — the fail-open test's stimulus does not fault under uid 0 (rc:3793116738) — FIXEDThe permission-based test (chmod 0o555 → EACCES) is replaced by two uid-independent tests: the record-dir-as-file shape fault above (mkdir branch), and an injected EACCES throw on the R4-4 — hand-rolled findings-pointer extraction (rc:3793116740) — FIXEDBoth hand-rolled copies ( R4-5 — budget-env isolation pasted as string literals (rc:3793116744) — FIXEDBoth isolation blocks ( R3-1 / R3-2 / R3-3 (rc:3792790796 / rc:3792790815 / rc:3792790832) — re-verified RESOLVEDFixed in R3-4 — PR body still documents the removed exception class (rc:3792790836, rc:3793116754) — ESCALATED, threads left openVerified real: the live PR body's Reviewer Test Plan still says the honest-absence clause "now certifies dry instead of reading unknown", the opposite of the shipped fail-closed behavior and of the residue test. The fix is a PR-description edit — a GitHub write this headless run is not permitted to perform. The exact replacement text is posted as a reply on both threads; they stay open for a maintainer. Verification
中文说明第 5 轮总结 — PR #9272提交: 发现与处置R4-1 — [Critical] 行走动词守卫是逐形状补丁(rc:3793116734)— 已修复先复现后修复:在轮前代码树上,五个入口类全部在"自认未检查"的回执上退役了 chunk — 11 个新测试在 HEAD 处失败( 修复改用封闭语法类收口,而非逐一分词枚举:标记剥离现在拒绝 R4-7 — 守卫在诚实回执上误触发(rc:3793116751)— 已修复上述修复携带了发现所建议的豁免:词串中只要带有短语自身的填充词 R4-6 — 守卫的绑定依赖 EN 恰为单一分支(rc:3793116750)— 已修复组合处对 EN 分组—— R4-2 — mkdir 的 EEXIST 与申领的 EEXIST 混同(rc:3793116736)— 已修复
R4-3 — 失败开放测试的刺激在 uid 0 下不触发(rc:3793116738)— 已修复基于权限的测试(chmod 0o555 → EACCES)被两个 uid 无关测试替换:上述记录路径为文件的形状故障(mkdir 分支),以及对 R4-4 — 手工实现的 findings 指针提取(rc:3793116740)— 已修复两处手工副本( R4-5 — 预算环境隔离块以字符串字面量粘贴(rc:3793116744)— 已修复两个隔离块( R3-1 / R3-2 / R3-3(rc:3792790796 / rc:3792790815 / rc:3792790832)— 复核确认已解决已在 R3-4 — PR 正文仍在描述已移除的例外类(rc:3792790836、rc:3793116754)— 已升级,线程保持打开核实属实:线上 PR 正文的评审验证计划仍称诚实缺省从句"现在认证为 dry 而不再读作 unknown",与上线的失败关闭行为及残留测试相反。修复是 PR 描述编辑——本次无头运行无权执行的 GitHub 写操作。确切的替换文本已作为回复发布在两个线程上,保持打开等待维护者处理。 验证
Base-conflict check · 基分支冲突检查: no conflict with main. · 与 main 无冲突。 Re-review when you have a moment. After round 100 this bot stops and leaves the PR for a human. · 有空请复审;第 100 轮后本 bot 停止并将 PR 交给人工。 🧠 Handled by Qwen Code · model/模型 |
qwen-code-ci-bot
left a comment
There was a problem hiding this comment.
Partially reviewed — gaps disclosed.
Not reviewed: reverse audit — did not converge within the reverse-audit round cap of 5.
中文说明
仅完成部分审查,审查缺口已披露。
未审查:反向审计——在 5 轮的反审轮数上限内未收敛。
— qwen3.8-max via Qwen Code /review (v0.21.12)
…lose over the guard family (QwenLM#9272 round 7)
|
🤖 AutoFix updated a stale base — the fix did not pass verification, but this PR was behind Round 8 — verification round over the round-5 review findings, plus two small fixesEvery finding of the round-5 review ( Dispositions
Why it was not pushed: Note: the base has since been auto-updated; the verdict below predates that update, and the next round's re-measurement may charge the round. bite check: changed tests pass on the pre-round tree (claimed defect does not reproduce) Run log: https://github.com/QwenLM/qwen-code/actions/runs/31992805230 🧠 Handled by Qwen Code · model/模型 |
…ile mtimes (QwenLM#9272 round 8)
…ew-retirement-followups-9259
qwen-code-ci-bot
left a comment
There was a problem hiding this comment.
Partially reviewed — gaps disclosed.
Not reviewed: build-and-test — Integration Tests (CLI, No Sandbox) was skipped in CI and its suite did not run locally.
Not explored to full depth (tool budget reached): chunk 2: did not run the two vitest files in the worktree (no dependency install attempted); the Critical's evidence is the executed predicate-vs-production-output corne….
Not reviewed: reverse audit — did not converge within the reverse-audit round cap of 5.
Deferred under the convergence posture (round 6, not a blocker) — recorded, not requested in this round:
packages/cli/src/commands/review/cleanup.test.ts:531 — [probe] the Nothing to clean assertion cannot fail for the regression it pins — the mock shape always sets removedAny firstpackages/cli/src/commands/review/lib/retirement.test.ts:1100 — [probe] fixture comment claims a marker-strip pin, but the restatement bar refuses the clause first — the claimed pin does not existpackages/cli/src/commands/review/lib/retirement.test.ts:1369 — [probe] 'a nested phrase occurrence' test name/comment outlives its fixtures — the named scenario is now structurally unadmittablepackages/cli/src/commands/review/lib/deadline.ts:229 — [probe] claimRetirementDegradeNote fails toward silence when an unremovable stale claim survives to the wx createpackages/cli/src/commands/review/lib/retirement.test.ts:787 — [probe] the zh 无-core restatement shape is pinned by nothing — the only core with no bare-marker backstoppackages/cli/src/commands/review/lib/retirement.test.ts:1733 — [probe] doubled-stock-sentence test name credits the parrot bar while the comment credits the restatement bar — neither is pinnedpackages/cli/src/commands/review/lib/retirement.test.ts:1248 — [probe] comment overclaims: the fenced shape is inert under the plain-line-filter mutation — only the blockquoted shape pins itpackages/cli/src/commands/review/lib/retirement.ts:381 — [probe] declared-residue comment cites examples the walk gate refuses — the residue list overstates the holepackages/cli/src/commands/review/lib/deadline.ts:236 — [probe] the EEXIST fail-open branch (!isFile → print) is reached by none of the five new claim testspackages/cli/src/commands/review/cleanup.test.ts:480 — [probe] comment claims a slack/sign-slip pin the fixture is arithmetically insensitive to — that pin lives in the pre-existing repro testpackages/cli/src/commands/review/issue-9206-repro.test.ts:444 — [probe] mentionsChunk oracle is blind to 5 of 14 certification reasons — keyword-less diagnostics read as silence
中文说明
仅完成部分审查,审查缺口已披露。
未审查:build-and-test — Integration Tests (CLI, No Sandbox) was skipped in CI and its suite did not run locally。
未探索到全部深度(达到工具调用预算):chunk 2:did not run the two vitest files in the worktree (no dependency install attempted); the Critical's evidence is the executed predicate-vs-production-output corne…。
未审查:反向审计——在 5 轮的反审轮数上限内未收敛。
收敛姿态下延后(第 6 轮,非阻断)——已记录,本轮不要求修改:共 11 条(原文未翻译,列表见上方英文部分)。
— qwen3.8-max via Qwen Code /review (v0.21.13)
|
🤖 Reviewed the latest feedback — no changes needed. Why, point by point: · 已审阅最新反馈——无需改动。逐点说明原因如下: Round 9 is a no-code re-verification round. Every actionable item in this The recorded rejection is superseded — measured, not assertedThe rejection's all-green measurement (test run starting 13:26 / bite run
Findings re-verified on the pushed head (resolved via resolved-comments.txt)
Left open, with reasons already on their threads
The round-6 review (the newest) found zero findings; its eleven deferred Verification
中文说明第 9 轮是无代码改动的复验轮。本轮反馈中所有可处理的事项,要么已在已推送的分支头上修复,要么已带着理由在对应线程中拒绝,要么是本无头运行无法执行的 GitHub 写操作。本轮不做任何代码改动——这正是 bite 检查拒绝意见对"缺陷已被本分支更早提交修复"的情形所开出的路径。 记录在案的拒绝已被取代——以实测为凭,而非口断该拒绝的全绿测量(测试运行始于 13:26、bite 运行 13:28,均为运行器本地时间)早于已推送的分支头(deadline 围栏提交创建于 14:13 +0800),也早于基分支自动更新(13:29 +0800)——与"该判定早于那次更新"的备注一致。它测量的是一个被丢弃的中间尝试,而非上线的代码。在当前分支头上重新测量:
已在推送头上复验的反馈(经 resolved-comments.txt 标记解决)
保持打开、理由已在线程上的事项
第 6 轮审查(最新一轮)未发现任何缺陷;其 11 条延后的探针条目明确标注"已记录,本轮不要求修改"——维持记录状态。其中两条 deadline 相关的(不可移除的过期申领;未触达的失败开放分支)已被推送的围栏及其测试关闭。 验证
Base-conflict check · 基分支冲突检查: no conflict with main. · 与 main 无冲突。 🧠 Handled by Qwen Code · model/模型 |
|
@qwen-code /triage |
|
Sandboxed verification: ❌ not passed — findings reported (agent verdict) - workflow run Ran the PR in an isolated, token-free container: A/B against the base build, mock-free harness assertions, targeted gates. Advisory evidence for human reviewers — not a review, an approval, or a CI check. Scripted assertions: 355 passed · 0 failed · 355 total 中文 — 判定:❌ 不通过 · 报告了发现(agent 判定)沙箱验证在隔离、无凭证的容器中执行了该 PR 的代码(与 base 构建 A/B 对照、无 mock harness 断言、定向门禁)。仅作为评审证据,不构成评审、批准或 CI 检查。 脚本断言:355 通过 · 0 失败 · 355 总计 Verification reportPR #9272 deep verification (round 2) —
|
| # | Finding | Severity | Status at head 526462f4 |
|---|---|---|---|
| F1 | Reviewer Test Plan claimed the residue clause "now certifies dry" — the opposite of merged behavior | non-blocking, docs | fixed — the Test Plan now says the residue clause keeps reading unknown; re-measured at the new head (A/B cell residue-honest: unknown on both arms, named receipt clause contradicts the phrase on head). All three plan steps were walked and matched measurement — see "Reviewer Test Plan walk" below. |
| F2 | Test-count drift in the PR body (then: 509/500 claimed vs 518 measured) | nit | stands — the body was updated to 533/533 (both languages), measured is now 542 at head / 493 at base. Same failure mode: the count stales across review rounds. |
Carried-over measurements, re-run at the new head (input closure changed — four new commits touch the measured code — so no carry-forward shortcut applied):
| Measurement | Previous round | This round |
|---|---|---|
| Classifier A/B (cells × 2 arms) | 13 cells, head 13/13, base 13/13 | 17 cells, head 18/18 tests, base 18/18 tests |
| Leak shapes base certifies dry | 4 (round-4 guard) | 6 (restatement bar closes more doors, incl. the old control-object-seat long form) |
| Degrade-NOTE A/B | 4 scenarios both arms | same 4 scenarios, re-measured, all flips hold |
| Claim concurrency (24 processes) | exactly 1 CLAIMED | exactly 1 CLAIMED, 23 silent; re-capture re-arms |
| Mutation matrix | M0–M4, no survivors | M0–M5 (M5 new for the round-8 fence), no survivors |
| Six-suite gate | head 518/518, base 485/485 | head 542/542, base 493/493 |
| Lookahead-regex ladder (2k/10k/40k) | linear, 0.3→2.7 ms | superseded — the lookahead guards no longer exist (round 7 replaced the family with the restatement bar); new ladder over the current classifier: linear, head refuse 1–4 ms / admit 0–4 ms |
Windows wx atomicity |
not covered (Linux only) | not covered (Linux only) |
Scope
Central claim (three production changes, all in the reverse-audit retirement path):
- Four self-naming certification bars replace the collapsed
receipt clause not substantive— now five names: round 7 addedreceipt clause restates the all-clear. - Degrade NOTEs deferred past the budget/round-cap gate; the per-chunk note claimed cross-process via
claimRetirementDegradeNote(wxcreate + content/shape fence — round 8 re-keyed the stale check from file mtimes to the claim's ownatMsvs the strict plan mtime). - Marker list stays bare (
unexaminedadded); the polarity guard is now closed by FORM: a clause containing the receipt's phrase core at all is refused, with no lookahead and no exception list.
Secondary: the new test pins are non-vacuous (mutation-checked, incl. the round-8 fence pins); six suites green on both arms.
Delta since the previous round (commits b7e92c9e round 5, 8c7f2b00 round 7, 9ab02ccf round 8, per the metadata snapshot — individual diffs unreachable at depth 2): round 5's passive-head lookahead and uid-proof claim tests were superseded by round 7, which deleted the whole lookahead guard family in favor of the restatement bar and updated the auditor brief to match; round 8 re-fenced the claim on content/shape. New probes this round target exactly that delta: the restatement bar's entrance family, the brief self-consistency, and the claim fence's shape/fault matrix.
A/B proof — harness 1: the certification classifier
17 identical fixture cells driven through the real agentPromptCommand.handler over real-fs plan/records/transcripts (only the stdio writers mocked), on the PR build and on a base worktree at HEAD^1 (control hygiene in Methodology). Oracles per chunk 13/14: retired vs the diagnostic line's exact named bar. Evidence: 01-ab-classifier-base-vs-head.png.
| Cell | BASE (control) | HEAD (PR) |
|---|---|---|
bar-lead-hedge (…found but only skimmed.) |
unknown — collapsed | unknown — receipt lead contradicts the phrase |
bar-clause-admission (…but skipped the generated files…) |
unknown — collapsed | unknown — receipt clause contradicts the phrase |
bar-no-walk (everything looks fine this round…) |
unknown — collapsed | unknown — receipt clause names no walk |
bar-too-thin (walked lexing) |
unknown — collapsed | unknown — receipt clause too thin |
leak-passive-issues (…; no issues were verified.) |
dry — admission certified | unknown — receipt clause restates the all-clear |
leak-passive-findings (…; no findings were verified.) |
dry — admission certified | unknown — restates |
leak-passive-gaps (…; no gaps are verified outstanding.) |
dry — admission certified | unknown — restates |
leak-adverb (…; no issues at all were verified.) |
dry — admission certified | unknown — restates |
leak-object-seat (verified no issues in the reconnect state machine and both call sites) |
dry — admission certified | unknown — restates |
leak-reduced-passive (…; no issues verified outstanding.) |
dry — admission certified | unknown — restates |
leak-object-seat-short (verified no issues in it or its callers) |
unknown — collapsed (substance floor¹) | unknown — restates |
leak-nbsp (no\u00A0issues were verified — NBSP splits the core) |
unknown — collapsed | unknown — clause polarity (core match broken, marker catches it) |
residue-honest (verified no regressions in the reconnect path and re-walked its call sites) |
unknown — collapsed | unknown — clause polarity (F1 clause, accepted residue) |
residue-passive-regressions (…; no regressions were verified.) |
unknown — collapsed | unknown — clause polarity |
| control-clean-dry (canonical receipt) | dry | dry |
| control-example-parrot (the brief's own example receipt, verbatim) | unknown — collapsed | unknown — too thin (parrot refusal, pre-existing) |
| control-yield (a finding return) | hot, no diagnostic | hot, no diagnostic |
6 of 8 leak shapes flip from dry (base certified the admission) to unknown (head refuses it); all four bars plus the NBSP/short-object boundary cells flip from the collapsed name to self-naming; all 3 controls are behavior-stable. Head 18/18, base 18/18 against arm-specific expectation tables.
¹ Harness-prediction correction, disclosed per the counts-are-sacred rule: I predicted base would certify the short object-seat dry (the previous round's analogous cell was dry at the old base). Measurement showed base's substance floor trips first — the base's greedy filler strip caps at 32 chars and leaves verified . (<20 chars), so base refuses too, collapsed. The expectation table was corrected to measured ground truth; the long object-seat cell above preserves the intended dry→unknown flip. A second prediction correction in the entrance harness is noted there.
Reviewer Test Plan walk (per step, at the new head):
- "a round refused at the round cap with an unreadable transcript history now emits only the ROUND CAP refusal" — measured, A/B scenario S1/S3: head emits refusal only, no NOTE; base emits the NOTE then the refusal. ✅
- "
No issues found — verified no regressions in the reconnect path and re-walked its call sites.keeps readingunknown" — measured, cellresidue-honest: unknown on head, barreceipt clause contradicts the phrase. ✅ (this is the F1 fix) - "
No issues found — re-walked the path; no regressions were verified.now readsreceipt clause contradicts the phrase" — measured, cellresidue-passive-regressions. ✅
A/B proof — harness 2: degrade-NOTE truthfulness
Four scenarios driven through the real handler (both --all-chunks and --chunk paths; real fs, stdio writers mocked). Evidence: 02-ab-degrade-notes-base-vs-head.png.
| Scenario | BASE | HEAD | flip |
|---|---|---|---|
S1 round-cap-refused round, unreadable history (--all-chunks) |
NOTE printed, THEN refusal | refusal only, no NOTE | ✅ |
| S2 c13 admission build, schedule reads cleanly | silent | silent | — |
| S2 c14 repair build after clean admission (history died) | silent — stamp-keyed suppression | NOTE once, reason carried (unavailable this round — <why> — auditing the chunk.) |
✅ |
| S2 c15 second repair, same round | silent | silent — claim spent | — |
| S2 r4 later round, history still dead | NOTE | NOTE — per-round claim | — |
S3 round-cap-refused round (--chunk) |
NOTE printed, THEN refusal | refusal only | ✅ |
| S4 ADMITTED round, unreadable history | NOTE, 3 chunks built | NOTE, 3 chunks built | — (deferral does not swallow the note) |
Both arms 4/4 against arm-specific expectations; all three repair builds still recorded their chunk builds on both arms (the safe direction holds).
Delta probe — the restatement bar's entrance family (round 7)
Driven through the real scheduleReverseAuditRound over seeded histories; expectations arm-specific. The PR's own pin table covers the NBSP/U+3000 splits after issues; my harness adds the siblings the other side of the core and the non-English doors:
| Entrance | BASE | HEAD |
|---|---|---|
comma-spliced (…, no issues, everything checks out.) |
dry | unknown — restates |
zh restatement (未发现新问题,重新走查了…,未发现问题。) |
dry (core stripped, walk + CJK substance pass) | unknown — restates |
parenthetical (no issues (none) verified) |
dry | unknown — restates |
restatement inside a path span (re-walked docs/no issues.md and the callers) |
dry | unknown — restates (no quoted-span exemption; fail-closed, consistent with #9213 precedent) |
no issues were found because nothing was verified (the pair no regex separates) |
unknown² — collapsed | unknown — restates |
DECLARED RESIDUE: nothing was verified (no core, no listed marker) |
dry | dry — the accepted never-retire cost, both arms; declared in code and body, pinned by the PR's own test |
false-positive guard: re-walked the notifier path and the findings index |
dry | dry (marker-adjacent words are not markers) |
overlooked the files (unlisted hedge, no walk verb) |
unknown — collapsed | unknown — receipt clause names no walk |
² Second harness-prediction correction: predicted base-dry; base's substance floor trips first (same mechanism as footnote ¹). Corrected to measured ground truth.
Scaling ladder over the current classifier (the previous round's lookahead no longer exists): hostile clause (restatement + marker scan) and benign clause (admit scan) at 2k / 10k / 40k chars — head refuse 1 / 1 / 4 ms, admit 0 / 1 / 4 ms; base refuse 1 / 2 / 3 ms, admit 1 / 1 / 7 ms (an earlier cold base run measured refuse 1 / 30 / 22 ms — shared-runner variance, same linear shape). Linear, no rung near the 2 s cap. 01-… image shows the cell matrix; the numbers above are logs/ladder-head.log / logs/ladder-base.log.
Brief self-consistency (the brief half of change 3): the reverse-audit brief now carries "The clause narrates the walk in the walk's own words and NEVER restates the all-clear … not even as the walk's object (verified no issues in X)" — asserted present; and the brief's example receipt clause itself carries no phrase core (the brief does not teach a form the tooling always refuses). Both pass.
Delta probe — the claim fence (rounds 5 + 8)
Direct probes of the exported claimRetirementDegradeNote (real fs; the base arm asserts the symbol's ABSENCE — the runtime control proving the base tree executed base code; it passed):
| Probe | Result |
|---|---|
| once per round per run; round isolation; plan re-capture (+1 h mtime) re-arms | ✅ |
fence is CONTENT: claim atMs vs STRICT plan mtime (<); equality edge keeps the claim (recorded boundary³); +1 ms re-arms |
✅ |
corrupt occupant ({) reclaimed, not EEXIST-silenced |
✅ |
| directory occupant removed, claim lands | ✅ |
| record dir blocked by a REGULAR FILE → fails OPEN (any uid) | ✅ |
| record dir permission-faulted (0o555; this sandbox runs uid 1000, so mode bits fault for real) → fails OPEN, no claim file | ✅ |
| dangling symlink at the claim path → fails OPEN | ✅ |
round-less claim keys on round-x |
✅ |
³ Boundary, not a finding: at atMs == plan mtime exactly, the strict < keeps the claim (suppression side). Reaching it requires a plan re-captured in the same millisecond as the claim it should supersede — a retry run re-captures minutes later, and within one run the claim always postdates the capture. The PR's own deadline test pins the strict fence deliberately.
Concurrency (the claim exists for inter-process exclusion; probe against the compiled dist/): 24 parallel node processes claiming the same round — exactly 1 CLAIMED, 23 silent, 0 errors; sequential re-claim silent; plan re-capture re-arms. Evidence: 04-gates-head-base-claim-concurrency.png. (Disclosure: my first run printed 23 CLAIMED — a harness bug, my child script mis-parsed node -e argv and every child raced on a nonexistent plan path, which fail-open-reclaimed in a loop; fixing the argv produced the clean result. The mis-run did incidentally demonstrate the plan-missing edge fails toward printing.)
Cleanup interaction (static, re-verified at this head): cleanup.ts is unchanged by the PR; its retention predicate hasPreviousRunRecords reads only file mtimes against runEpochMs(plan) (plan mtime − slack). The claim file is written mid-run (mtime ≫ plan mtime), can never fake the previous-run evidence, is never read for its content, and rides the record dir's retention. Same conclusion as the previous round, re-derived against round-8 content.
Vacuity / mutation matrix
Scratch worktree at HEAD, one mutation at a time, pinned suites re-run (unmutated control green: 468/468 over the three pinned suites). Evidence: 03-mutation-matrix-all-killed.png; raw logs logs/matrix.log, m0–m5.log.
| Mutant | Named pin(s) | Result |
|---|---|---|
M0 remove unexamined from marker list (positive control) |
the un-examined admission case | 1 red — exactly the named pin; suite proven live |
| M1 revert bar split to one collapsed name | the four bar-naming groups + restatement pins | 33 red; the #9206 repro suite stays green under the mutant — detection intact, only attribution broken, which is the change's point |
| M2 delete the restatement bar (restore strip-then-marker) | restatement family | 20 red — the mutant retires the admissions again |
| M3 move round-builder NOTE back before the gate | refused-round NOTE test | 1 red — mutant prints the NOTE on a refused round |
M4 claim create flag wx→w |
once-per-round claim + reclaim tests | 2 red |
M5 disable the content fence (stale = false) |
re-arm on re-capture, strict-mtime fence, stale reclaim | 3 red |
No survivors. Every guard the PR ships — including both round-8 additions — is pinned by a test that fails the intended behavioral assertion under the mutant.
Targeted gates
| Gate | Result |
|---|---|
| Six suites at HEAD (the PR's own list) | 542/542 pass (retirement 146, agent-prompt 266, deadline 56, cleanup 39, audit-layers 25, issue-9206-repro 10) — body claims 533 (see F2) |
| Same six suites at BASE (base's own test files) | 493/493 pass → +49 tests, 0 regressions |
tsc --noEmit in packages/cli |
exit 0, zero diagnostics |
Findings
F2 (carried over) — test-count drift in the PR body (nit; non-blocking, documentation)
English body and 中文 both say "expect 533/533" for the six named suites; measured at the verified head is 542 (base: 493). Same failure mode as the previous round (then 509/500 claimed vs 518 measured): the count is not re-synced after later rounds add pins (rounds 5–8 added 24 tests to the head suites since the last verification). The plan itself works — all three behavior steps verified above; only the number is stale. Suggested fix (prose only): update to 542/542, or drop the exact count ("all tests pass").
No other findings. Specifically probed and not findings: the claim fence's equality edge (boundary ³, effectively unreachable, deliberately strict); the nothing was verified residue (declared, pinned, fails toward retirement which the module states as its direction); the two harness-prediction corrections (footnotes ¹ ² — harness errors, not PR defects, disclosed per the counts rule).
Not covered
- Per-commit attribution: depth-2 checkout — only the merge commit,
HEAD^1, andHEAD^2exist locally;git rev-list HEAD^1..HEAD^2returns 1 vs 12 commits in the metadata snapshot, and the repository is shallow. The delta since the previous headfe0ec470was reconstructed from the metadata commit messages + the final code, not from per-commit diffs; only the aggregateHEAD^1..HEADdiff was verified. - Live-model E2E: declared out of scope by the PR itself; harnesses reproduce the wire shapes (records + transcripts as the orchestrator writes them), not a live orchestrator run.
- Repo-wide gates: only the affected workspace's six suites + package typecheck. No repo-wide
npm run lint/npm run test. - Windows:
wx/O_EXCLatomicity and the claim's fs semantics measured on Linux only; the PR self-declares Windows untested (CI). - Multi-process orchestrator run: the 24-process probe exercises the claim function directly against
dist/; a real Step-3B fan-out was not staged. - The entrance probe's two prediction corrections (footnotes ¹ ²) mean the base-arm expectations for those two cells were measured twice; final table is measured ground truth.
Methodology
Environment: node v22.23.2, CI container (node:22-bookworm), merge-ref checkout (HEAD = merge 3d3c48f1, HEAD^1 = base tip d9d210eb, HEAD^2 = verified head 526462f4), uid 1000 (non-root — permission-bit fault stimuli work for real). Control hygiene: base worktree at HEAD^1 wired to the already-installed head node_modules via symlinks (the PR touches no package.json/package-lock.json — verified by diff — so third-party deps are byte-identical); vitest's alias resolves @qwen-code/qwen-code-core to each worktree's own source, and the base arm additionally passed a runtime control (the claim symbol's absence on base) proving base code executed. Harnesses (preserved in this dir for rerun): zz-verify-ab-classifier.test.ts (17-cell matrix + arm control), zz-verify-ab-notes.test.ts (4 scenarios), zz-verify-claim.test.ts (8 fence probes + base-absence control), zz-verify-entrances.test.ts (8 entrances + ladder + brief check), claim-concurrency.mjs (24-process dist probe), run-matrix.sh (mutation runner). Mutation edits applied via scripted exact-block replacement in a scratch worktree, restored with git restore after each (tree verified clean at the end). Counts: classifier 104 head + 105 base, notes 17 + 17, claim 18 + 1, entrances 42 + 36, concurrency 5, mutation verdicts 7 (control + six mutants), gates 3 (head suites, base suites, typecheck) = 355 pass / 0 fail. Raw logs in logs/; evidence images referenced inline.
Evidence images
Harness scripts and raw logs are in the workflow run artifacts (7-day retention).
— Qwen Code · sandboxed verification
qwen-code-ci-bot
left a comment
There was a problem hiding this comment.
LGTM, looks ready to ship. ✅
|
🤖 Reviewed the latest feedback — no changes needed. Why, point by point: · 已审阅最新反馈——无需改动。逐点说明原因如下: No action needed this roundReviewed the feedback batch for this evaluation window (items newer than 2026-08-17T09:35:43Z). No code changes were made, and no commit was created. What was in the batch:
Cross-check against raw data: verified that no review bodies or inline comments newer than the evaluation cutoff exist outside the prepared feedback, and that the check-run data shows no failures or pending runs. Diff growth: source 74 / test 221 net lines vs budgets of 400/400 — within budget, 0 prior rounds over budget. No trajectory concern. The branch stays as-is at 中文说明本轮无需处理已审阅本次评估窗口内的反馈批次(2026-08-17T09:35:43Z 之后的新增条目)。未做任何代码改动,也未创建任何提交。 本批次内容:
与原始数据交叉核对: 已确认评估截止时间之后,不存在准备反馈之外的新评审正文或行内评论;同时检查数据中也没有失败或未完成的运行。 Diff 增长情况: 源码净增 74 行 / 测试净增 221 行,预算为 400/400——在预算之内,此前没有任何一轮超出预算。增长趋势无需关注。 分支保持在 Base-conflict check · 基分支冲突检查: no conflict with main. · 与 main 无冲突。 🧠 Handled by Qwen Code · model/模型 |
|
Released in v0.21.14. |









What this PR does
Follow-up to #9213, landing the deferred Suggestions tracked in #9259 (that PR was frozen to Critical-only after 7 review rounds; every deferred item was recorded so nothing would be silently dropped).
Production changes (three, all in the reverse-audit retirement path):
receipt clause not substantivefailure name collapsed four semantically distinct bars (lead-side polarity, clause-side polarity, the walk gate, the substance floor), so the /review: chunk retirement silently does not fire in the reverse-audit loop, and cleanup destroys the evidence #9206 diagnostic could not tell a silent never-retire's cause apart. The bars now reportreceipt lead contradicts the phrase/receipt clause contradicts the phrase/receipt clause names no walk/receipt clause too thin. Splitting the polarity test into per-domain calls classifies identically to the combined domain it replaces — no token can span the lead/clause boundary (the separator's trailing line-bound whitespace sits between them).auditing every chunk./auditing the chunk.) printed before the budget/round-cap gate, so a refused round's stderr promised an audit that never happened; they now print only after admission succeeds. The per-chunk note's suppression also re-keys from the admission stamp to a per-process record of the note having been printed: the stamp lands on the admission build whether or not that build's schedule read failed, so a round whose first build read cleanly and whose later repair builds began to throw was silenced on every remaining build — the exact never-retire shape /review: chunk retirement silently does not fire in the reverse-audit loop, and cleanup destroys the evidence #9206 exists to name. The re-key prints once per round per process (a retried headless run re-prints — the safe side).no regressions have been verified, 未来得及, 没有回归测试 — the surface has no last corner, the fix(review): fix silent reverse-audit retirement failures and keep non-converged evidence #9213 R2-2 lesson), so the list keeps only witnessed admission markers (addingunexamined) and the honest mirror is the stated, accepted residue:verified no regressions/ 确认没有回归 readunknownand the chunk stays under audit — the failure direction the module declares, preferable to certifying one admission. The residue is pinned by a dedicated test.Everything else is test-pinning, each pin verified against the mutation it names:
no read of the diff/territory read missingbar order, the yield-suppression gate, the CJK substance-floor branch in both directions, the substance floor's phrase-stripped measurement, quoted layer lines (fence/blockquote) never being stripped,no successful tool calls, and the malformed-marker shape gate inreadBudgetStopUnfenced.writeStderrLineSafemock is cleared per test (assertions previously read accumulated output — an order-dependent oracle).statSynctests for the killed-run mtime signal, the plan-missing second-cleanup signal, the previous-run unfenced marker read, and theNothing to clean-vs-preserved guard.yieldedrather than emitting the diagnostic noise the suite exists to catch.Declined with justification (recorded in #9259): the stamped-build schedule re-read "perf" item (the pairing walk is tens of small file reads per chunk build — negligible against an LLM round); the mid-run plan-rewrite mtime edge and the previous-marker-content overwrite (both fail toward over-keeping — the safe direction for evidence); an automatic sweeper for plan-missing retention (invents a retention policy; the Kept line's manual-removal instruction is the deliberate exit).
Why it's needed
#9213 made silent retirement failures diagnosable, but the deferred findings are what make the diagnostic channel trustworthy in practice: one collapsed failure name hid which bar fell, the degrade notes could claim an audit that a gate then refused, and common honest phrasing (
no regressions) never retired — the same never-retire cost #9206 reported, wearing a different coat. The test pins keep every one of these from regressing green.Reviewer Test Plan
How to verify
Run the six suites:
cd packages/cli && npx vitest run src/commands/review/lib/retirement.test.ts src/commands/review/lib/audit-layers.test.ts src/commands/review/lib/deadline.test.ts src/commands/review/agent-prompt.test.ts src/commands/review/cleanup.test.ts src/commands/review/issue-9206-repro.test.ts— expect 533/533. Behavior a reviewer can confirm by reading: a round refused at the round cap with an unreadable transcript history now emits only the ROUND CAP refusal (noauditing every chunk.NOTE); a clause likeNo issues found — verified no regressions in the reconnect path and re-walked its call sites.keeps readingunknown(the accepted residue — the exception class that would spare it licensed admissions no enumeration closes, removed fail-closed in round 3), whileNo issues found — re-walked the path; no regressions were verified.now readsreceipt clause contradicts the phrase(the strip no longer eats the marker's domain overissues/findings/gaps).Evidence (Before & After)
N/A (non-UI; unit-level evidence is the 500-test suite plus mutation checks — each new pin was executed against the mutant it names: greedy strip, unanchored matcher, stamp-keyed suppression, bare markers, inverted mtime comparison, removed yield gate, deleted CJK branch, fixed-YIELD harness — all fail their pins under the mutant and pass restored).
Tested on
Environment (optional)
N/A — unit tests only (vitest, real-fs tmpdir harnesses).
Risk & Scope
unexaminedand nothing else — the absence-of-problems exception class was probed, falsified twice, and removed fail-closed; honestno regressions-style phrasing keeps readingunknown(the accepted never-retire cost, pinned by a test).Linked Issues
Closes #9259
中文说明
本 PR 内容
#9213 的后续,落地 #9259 中跟踪的延期 Suggestion(该 PR 在 7 轮评审后冻结为仅修 Critical,所有延期项均已记录,确保无静默丢弃)。
生产改动(三处,均在反审退役路径):
receipt clause not substantive失败名归并了四个语义不同的门槛(引导侧极性、从句侧极性、行走门、实质性下限),/review: chunk retirement silently does not fire in the reverse-audit loop, and cleanup destroys the evidence #9206 的诊断因此无法区分静默不退役的原因。现在分别报告receipt lead contradicts the phrase/receipt clause contradicts the phrase/receipt clause names no walk/receipt clause too thin。按域拆分极性测试与原有的合并域判定完全等价——引导/从句边界不可能有标记词跨越(分隔符的行绑定尾随空白位于其间)。auditing every chunk./auditing the chunk.)原先在预算/轮数上限门之前输出,被拒绝的轮次 stderr 会承诺一场并不发生的审计;现改为准入成功后才输出。逐 chunk NOTE 的抑制键从准入戳记改为"本进程已打印"记录:戳记在准入构建落下时无论其调度读取是否失败,因此"首轮读取干净、后续修复构建开始抛异常"的轮次此前在剩余每次构建上都被静默——正是 /review: chunk retirement silently does not fire in the reverse-audit loop, and cleanup destroys the evidence #9206 要命名的静默不退役形态。改键后每轮每进程打印一次(无头重试会重新打印——安全方向)。no regressions have been verified、未来得及、没有回归测试——该表面没有最后一个角,与 fix(review): fix silent reverse-audit retirement failures and keep non-converged evidence #9213 R2-2 同教训),因此词表只保留已见证的自认标记(新增unexamined),诚实镜像作为明示的、被接受的残留:verified no regressions/ 确认没有回归 读作unknown、chunk 继续受审——模块声明的失败方向,优于认证一个自认。残留由专门测试钉住。其余均为测试钉住,每个钉住都对其命名的变异执行过验证(详见英文正文与 #9259):贪婪剥离 fixture 现在残留中命名行走、锚定钉住、极性先于对象逃逸、
no read of the diff门序、yield 抑制门、CJK 下限双向、剥离后测量、引用 layer 行不剥离、no successful tool calls、畸形标记形状门、逐 chunk 诊断窄化的缺席断言、writeStderrLineSafemock 逐测试清空、cleanup retention 的 statSync 模拟测试、repro 骨架的环境变量隔离/Safe 流捕获/逐 chunk 断言/对照组每轮新发现,以及过时测试标签修正。论证后放弃(已记录于 #9259): 戳记构建的调度重读性能项(配对遍历仅为每构建数十次小文件读取,相对 LLM 轮次可忽略);运行中计划重写的 mtime 边角与上次标记内容覆写(均倒向多保留——证据的安全方向);plan 缺失保留的自动清扫器(会发明保留策略;Kept 行的手动移除指引是刻意设计的出口)。
为什么需要
#9213 让静默退役失败可诊断,而延期的这些发现决定诊断通道在实践中是否可信:归并的失败名掩盖了倒在哪个门槛、降级 NOTE 可能在门拒绝后谎称继续、常见诚实措辞(
no regressions)永不退役——正是 #9206 报告的代价换了个外衣。测试钉住确保这些不会以全绿回退。评审验证
六个套件 533/533;变异验证:8 个新钉住各自在其命名的变异下失败、还原后通过。