feat(core): Add privacy-safe tool-result boundary diagnostics - #9039
Conversation
E2E Test ReportTested commit: Fixture: deterministic MCP tool result containing 499,999 ASCII bytes with retained head/tail sentinels. Source SHA-256:
The producer artifact remained 499,999 bytes with the fixture SHA-256 on every route. Enabling diagnostics did not change target wire bytes or artifact contents; with Privacy verification checked the new Batch tests verified that multi-tool Headless JSON and ACP bulk writer events retain tool-call HMAC and artifact-summary slot order. Subagent tests verified that only the closed Latest-main verification after syncing to
Known limitation: a fully model-driven nested-subagent E2E was not run. The nested path is covered by cross-package unit tests and exact static consumer tracing. The ACP harness also emits a known private teardown-method error after the target response is fully captured; it does not affect the measured tool-result frame or artifact. |
|
Re-run on an unchanged head —
Moving on to code review. 🔍 中文说明在未变更的 head 上重跑 gate——
进入代码审查 🔍 — Qwen Code · qwen3.8-max Reviewed at |
Code reviewNo code delta — the head is still
These match the Aug 16 Two behavioral side-notes from prior passes, re-verified in code at this head:
Privacy posture re-verified at this head: events carry sizes, process-local HMACs, and closed-enum artifact summaries only; identifiers and tool names are HMAC'd, never raw (confirmed at the event-construction site, lines 290-320). Gate conclusion: no merge blockers. TestingCI evidence for the reviewed head, recomputed today via the check-runs API (per the gate rules I do not build or run PR code):
1,152 check runs at the head: 35 success, 0 failure, 1,089 skipped by path routing, 28 cancelled or stale-queued routing jobs (bot-routing artifacts). All three Behavioral attestation at this exact head: four completed/success sandboxed verify runs — 31884380047, 31890726987, 31938443110, and 31952997718. This pass re-read the latest report rather than quoting it from memory: verdict 中文说明代码审查没有代码增量——head 仍是
以上与 8 月 16 日 上轮的两项行为说明本轮再次在代码中复核:
隐私姿态在本 head 上复核:事件只携带尺寸、进程内 HMAC 与闭集 artifact 摘要;标识符与工具名经 HMAC、从不落原文(已在事件构造处 290-320 行确认)。 Gate 结论:无合入阻塞项。 测试reviewed head 的 CI 证据今天经 check-runs API 重新计算(表格见英文部分;按 gate 规则不构建、不运行 PR 代码):head 上共 1,152 个 check run——35 success、0 failure、1,089 个被路径路由跳过、28 个 cancelled 或滞留的路由任务(机器人路由产物)。三个 本 head 的行为背书:四次已完成/成功的沙箱验证运行(31884380047、31890726987、31938443110、31952997718)。本轮重新读取了最新一份报告而非凭记忆引用:判定 — Qwen Code · qwen3.8-max Reviewed at |
Code Coverage Summary
CLI Package - Full Text ReportCore Package - Full Text ReportFor detailed HTML reports, please see the 'coverage-reports-22.x-ubuntu-latest' artifact from the main CI run. |
|
Confidence: 3/5 — clean review at an unchanged head, findings dispositioned by a human; the cap, not doubt, is still what keeps the bot short of approval. Reflection: my independent baseline for this problem — gate on the existing debug-file flag, measure sizes plus per-process keyed hashes at each boundary, capture exact writer bytes at the transport, rate-limit, and isolate every diagnostic failure — still matches this PR's shape, and I still don't see a materially simpler path that delivers attribution. This pass checked out the exact head in a worktree and re-traced all five contested findings in the source again: each behaves exactly as the reviewers describe, and each is confined to the diagnostic pipeline (the duplicate-callId prompt map feeds only Thread state re-counted through the GraphQL API this pass rather than carried forward: all 143 review threads are resolved, zero unresolved (up from 137 — the six new threads are the Aug 16 review round's, all resolved). @yiliang114's full-pass APPROVED review at this exact head stands. Why the bot still does not approve:
So the honest end-state is unchanged from the last pass, and re-running 中文说明置信度:3/5 —— 未变更 head 上的干净审查,发现已有人的处置;让机器人停在批准之前的仍是上限而非疑虑。 反思:我对这个问题的独立基线——挂在现有 debug 文件日志开关上、在每个边界测量尺寸加进程内 keyed 哈希、在传输层捕获精确 writer 字节、限流、隔离所有诊断失败——与本 PR 的形态依然一致,我也仍未找到仍能完成归因的更简路径。本轮在 worktree 中精确检出该 head,再次在源码里逐一追溯五个争议发现:每一个的行为都与 reviewer 的描述完全一致,且都局限于诊断管线内部(重复 callId 的 prompt Map 只供给 线程状态本轮经 GraphQL API 重新计数而非沿用旧值:143 条 review thread 全部 resolved,零未决(从 137 增至 143——新增六条来自 8 月 16 日的 review 轮,均已 resolved)。@yiliang114 在本 head 上的完整通过批准仍然有效。 机器人仍不批准的原因:
因此诚实的终态与上轮相同,在未变更的 head 上重跑 — Qwen Code · qwen3.8-max Reviewed at |
|
⏸️ Deferring to @wenshao — re-run on an unchanged head (
— Qwen Code · qwen3.8-max Reviewed at |
Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>
|
Addressed the Stage 2 ACP wire observation in Validation: focused emitter tests 63/63, independent ACP emitter/projector tests 95/95, Headless diagnostics/progress tests 59/59, recorder tests 3/3, changed-file Prettier/ESLint, and repository No review items were rejected or deferred. The unconditional disabled observer call was left unchanged because it immediately returns and has only the intended trivial cost. There were no inline review threads to resolve (resolved: 0). Qwen Review was not triggered. |
qwen-code-ci-bot
left a comment
There was a problem hiding this comment.
Partially reviewed — gaps disclosed. Suggestions are inline.
Not reviewed: build-and-test — Integration Tests (CLI, No Sandbox) was skipped in CI and its suite did not run locally.
Not explored to full depth (tool budget reached): "You are review agent reverse-audit — Reverse audit agent…": none — all planned checks completed within budget.; "You are review agent reverse-audit — Reverse audit agent…": none — all checks above completed within budget.; "You are review agent reverse-audit — Reverse audit agent…": none — all checks above completed within budget.; "You are review agent reverse-audit — Reverse audit agent…": none — all planned checks completed within budget.; "You are review agent reverse-audit — Reverse audit agent…": none — all checks above completed within budget., and 6 more.
Not reviewed: reverse audit — did not converge within the reverse-audit round cap of 5.
中文说明
仅完成部分审查,审查缺口已披露。 建议见行内评论。
未审查:build-and-test — Integration Tests (CLI, No Sandbox) was skipped in CI and its suite did not run locally。
未探索到全部深度(达到工具调用预算):"You are review agent reverse-audit — Reverse audit agent…":none — all planned checks completed within budget.;"You are review agent reverse-audit — Reverse audit agent…":none — all checks above completed within budget.;"You are review agent reverse-audit — Reverse audit agent…":none — all checks above completed within budget.;"You are review agent reverse-audit — Reverse audit agent…":none — all planned checks completed within budget.;"You are review agent reverse-audit — Reverse audit agent…":none — all checks above completed within budget.,另有 6 条。
未审查:反向审计——在 5 轮的反审轮数上限内未收敛。
— qwen3.8-max via Qwen Code /review (v0.21.10)
Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>
Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>
|
Review feedback addressed in commits 3bdd498 and cf00808. Fixed:
Deferred:
Verification:
Resolved review threads: 41. |
qwen-code-ci-bot
left a comment
There was a problem hiding this comment.
Partially reviewed — gaps disclosed. Suggestions are inline.
Not reviewed: reverse audit — reached the 5-round cap without two consecutive dry rounds (rounds 3–5 each reported findings; every reported finding was verified).
Not reviewed: build-and-test — Integration Tests (CLI, No Sandbox) was skipped in CI and its suite did not run locally.
Not explored to full depth (tool budget reached): "You are review agent reverse-audit — Reverse audit agent…": none — all checks above completed within budget.; "You are review agent reverse-audit — Reverse audit agent…": none (~11 of ~35 tool calls used).; chunk 3: direct inspection of @agentclientprotocol/sdk 's SessionUpdate type declaration (node_modules absent; substituted with repo precedent check).; "You are review agent reverse-audit — Reverse audit agent…": none — completed within budget (~17 of ~54 calls).; "You are review agent reverse-audit — Reverse audit agent…": none — all checks I started were completed within budget (6 tool calls used)., and 18 more.
中文说明
仅完成部分审查,审查缺口已披露。 建议见行内评论。
未审查:reverse audit — reached the 5-round cap without two consecutive dry rounds (rounds 3–5 each reported findings; every reported finding was verified)。
未审查:build-and-test — Integration Tests (CLI, No Sandbox) was skipped in CI and its suite did not run locally。
未探索到全部深度(达到工具调用预算):"You are review agent reverse-audit — Reverse audit agent…":none — all checks above completed within budget.;"You are review agent reverse-audit — Reverse audit agent…":none (~11 of ~35 tool calls used).;chunk 3:direct inspection of @agentclientprotocol/sdk 's SessionUpdate type declaration (node_modules absent; substituted with repo precedent check).;"You are review agent reverse-audit — Reverse audit agent…":none — completed within budget (~17 of ~54 calls).;"You are review agent reverse-audit — Reverse audit agent…":none — all checks I started were completed within budget (6 tool calls used).,另有 18 条。
— qwen3.8-max via Qwen Code /review (v0.21.11)
Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>
|
Addressed the latest review batch in
Validation: Core targeted 401/401, CLI diagnostics/replay 26/26, full ACP Session 587/587, repository build, typecheck, lint, formatting, and |
Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>
Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>
Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>
Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>
ed9a1c0 to
2a941e3
Compare
|
Rebased the branch onto latest Post-rebase focused validation: Core 401/401 and CLI diagnostics/replay 28/28 passed. The full build is currently blocked by an unrelated |
qwen-code-ci-bot
left a comment
There was a problem hiding this comment.
Partially reviewed — gaps disclosed.
Not reviewed: build-and-test — Integration Tests (CLI, No Sandbox) was skipped in CI and its suite did not run locally.
[Critical] R9-7 (packages/cli/src/ui/hooks/useGeminiStream.ts, re-check of existing blocker, comment 3787318576): prompt correlation is still keyed only by callId in finalizeToolResponses — duplicate call IDs across different prompts collapse attribution to the later prompt. Still stands at HEAD 21bb6fd; mechanism unchanged (author deferred as diagnostic-fidelity-only).
[Critical] R9-8 (packages/core/src/core/coreToolScheduler.ts, re-check of existing blocker, comment 3787318566): the producer fallback is still installed only after throwable scheduler prelude reads (getMessageBus()/getDisableAllHooks()); a throwing prelude publishes a terminal response with no producer observation. Still stands at HEAD 21bb6fd (author deferred as diagnostic-coverage-only).
[Critical] R9-9 (packages/core/src/followup/speculation.ts, re-check of existing blocker, comment 3787318582): the speculation prompt ID is still absent from every boundary event — observeSpeculationProducer sets sessionId/toolCallId/toolName only, and finalizeToolResponses receives no promptIds map. Still stands at HEAD 21bb6fd (author deferred as diagnostic-correlation-only).
[Critical] R9-10 (packages/core/src/utils/tool-result-boundary-diagnostics.ts, re-check of existing blocker, comment 3787318590): the broad artifact catch still returns { state: 'undecided', kinds: [] } on artifact iteration failure even when persistedOutputFiles already proves the result reusable/file-backed. Still stands at HEAD 21bb6fd (author deferred as diagnostic-summary-fidelity-only).
[Critical] R9-11 (packages/core/src/utils/tool-result-boundary-diagnostics.ts, re-check of existing blocker, comment 3787318571): rate-limit state (suppressedCount reset + emittedInWindow increment) is still consumed before logger.debug succeeds — a throwing logger consumes quota without writing the event. Still stands at HEAD 21bb6fd (author deferred as diagnostics-only recovery behavior).
中文说明
仅完成部分审查,审查缺口已披露。
未审查:build-and-test — Integration Tests (CLI, No Sandbox) was skipped in CI and its suite did not run locally。
[Critical] R9-7 (packages/cli/src/ui/hooks/useGeminiStream.ts, re-check of existing blocker, comment 3787318576): prompt correlation is still keyed only by callId in finalizeToolResponses — duplicate call IDs across different prompts collapse attribution to the later prompt. Still stands at HEAD 21bb6fd; mechanism unchanged (author deferred as diagnostic-fidelity-only).
[Critical] R9-8 (packages/core/src/core/coreToolScheduler.ts, re-check of existing blocker, comment 3787318566): the producer fallback is still installed only after throwable scheduler prelude reads (getMessageBus()/getDisableAllHooks()); a throwing prelude publishes a terminal response with no producer observation. Still stands at HEAD 21bb6fd (author deferred as diagnostic-coverage-only).
[Critical] R9-9 (packages/core/src/followup/speculation.ts, re-check of existing blocker, comment 3787318582): the speculation prompt ID is still absent from every boundary event — observeSpeculationProducer sets sessionId/toolCallId/toolName only, and finalizeToolResponses receives no promptIds map. Still stands at HEAD 21bb6fd (author deferred as diagnostic-correlation-only).
[Critical] R9-10 (packages/core/src/utils/tool-result-boundary-diagnostics.ts, re-check of existing blocker, comment 3787318590): the broad artifact catch still returns { state: 'undecided', kinds: [] } on artifact iteration failure even when persistedOutputFiles already proves the result reusable/file-backed. Still stands at HEAD 21bb6fd (author deferred as diagnostic-summary-fidelity-only).
[Critical] R9-11 (packages/core/src/utils/tool-result-boundary-diagnostics.ts, re-check of existing blocker, comment 3787318571): rate-limit state (suppressedCount reset + emittedInWindow increment) is still consumed before logger.debug succeeds — a throwing logger consumes quota without writing the event. Still stands at HEAD 21bb6fd (author deferred as diagnostics-only recovery behavior).
— qwen3.8-max via Qwen Code /review (v0.21.12)
|
Thanks — rechecked at 21bb6fd. No code change is warranted in this round. These are repetitions of the previously reviewed R9-7 through R9-11 findings, and the mechanisms remain as previously documented:
The skipped integration suite is acknowledged as a review gap, not a failing branch check; the current check snapshot is terminal with 20 successful checks and no failures. Under the repository rule for a PR beyond roughly five review rounds, these diagnostic fidelity, coverage, and recovery items do not meet the correctness, security, data-loss, or regression threshold for further expansion. They remain deferred to a focused follow-up. |
|
@qwen-code /triage |
|
Sandboxed verification: ✅ passed — merge-ready (agent verdict) - workflow run Ran the PR in an isolated, token-free container: A/B against the base build, mock-free harness assertions, targeted gates. Advisory evidence for human reviewers — not a review, an approval, or a CI check. Scripted assertions: 92 passed · 0 failed · 92 total 中文 — 判定:✅ 通过 · 可合入(agent 判定)沙箱验证在隔离、无凭证的容器中执行了该 PR 的代码(与 base 构建 A/B 对照、无 mock harness 断言、定向门禁)。仅作为评审证据,不构成评审、批准或 CI 检查。 脚本断言:92 通过 · 0 失败 · 92 总计 Verification reportPR #9039 — feat(core): Add privacy-safe tool-result boundary diagnostics (follow-up round, base moved again)Verdict: 中文 — 判定:✅ 通过 · 可合入(agent 判定)· 跟进轮(base 再次前移)跟进轮验证:PR head 与上两轮完全相同(
Previous-finding status (follow-up round)The previous round verified this identical PR head (
Carried measurements re-run at the new merge tree (all hold; tables below): A/B 7-vs-0 flip per format (13 in multi), exact wire byte accounting, the 65,536-JSON-byte content cap (present on base too), HMAC correlation chain with first mismatch at Round-over-round deltas, all benign:
Central claim + A/BCentral claim: when file debug logging is enabled, oversized/mutated tool-result representations emit Secondary claims: (1) equal values share one HMAC per process and the first changed boundary produces the first HMAC mismatch; (2) multi-tool aggregate writer events keep per-call HMACs and artifact summaries slot-aligned (Reviewer Test Plan step 5). Merge-integrity precondition (scripted, 2 assertions): for the 16 files both the PR and main touched, (a) 14 diagnostic marker names occur with identical counts in the PR-head blob and the merged tree (0 mismatches), and (b) every one of the 1,367 lines the PR added to those files is present verbatim in the merged tree (0 missing). Harness: Scenario (identical in every cell): esbuild-bundled repo
Witnesses: Key measured facts (70 E2E assertions, all scripted in
One pipeline observation, not a defect (unchanged from the previous round): in stream-json the wire frame is emitted at delivery time (before the recorder events), whereas json writes the aggregate frame after recording. The analyzer therefore asserts stage-set + counts for the stream cell and exact line order for the json cell — both pass. Finalizer de-gating (the previous rounds' finding-2 fix) re-proven through the compiled dist ( CorrectionsNone — the previous rounds' descriptions of the code checked out on re-inspection. FindingsNo blocking findings. Two non-blocking coverage gaps, both carried over and re-measured (completeness reporting, not merge conditions):
Mutation matrix (12 single-point mutants incl. the positive control; unmutated controls green at 18/18, 368/368, 89/89, 25/25, 33/33, cli 11/11; witness
No mutant regressed a previously-killed guard to survived; the two survivors are byte-for-byte the same unpinned axes the previous rounds reported, re-measured on identical test counts for those suites (18 and 33). Not covered
MethodologyEnvironment: CI verify container ( Evidence imagesHarness scripts and raw logs are in the workflow run artifacts (7-day retention). — Qwen Code · sandboxed verification |
|
Triage re-run completed — no new commits since the previous pass, so this run re-verified rather than re-reviewed. What this pass did: re-ran the full gate on the unchanged head What did not change: the verdict. Confidence remains 3/5 — the Stage 0 size escalation (1,821 production logic lines, cross-package core) caps the bot below approval by policy, and @wenshao's Aug 15 CHANGES_REQUESTED review with its five findings is still open on this exact head. The stage comments above were updated in place with today's recomputed evidence; the defer comment still names the one decision this PR is waiting on. No approval was issued, and no new review was submitted — the bot's existing position stands. 本次重跑完成——head 未变,重新验证而非重新评审:gate 全量重跑,五个争议发现已在代码中再次逐一追溯(均与评审描述一致,且均局限于诊断管线);CI 保持全绿(899 个 check run、零失败)。今天早些时候重触发的确认验证已 ✅(31938443110),本次重触发又启动一次确认(31952997718,进行中)。结论不变:置信度 3/5——Stage 0 规模升级使机器人按政策不自动批准,且 @wenshao 8 月 15 日携五项发现的 CHANGES_REQUESTED 仍悬在本 head 上。上方各阶段评论已就地更新为今日重算的证据;defer 评论仍然写明本 PR 唯一等待的那个决定。未发出批准,也未提交新评审——机器人既有立场不变。 |
|
@qwen-code /triage |
|
Sandboxed verification: ✅ passed — merge-ready (agent verdict) - workflow run Ran the PR in an isolated, token-free container: A/B against the base build, mock-free harness assertions, targeted gates. Advisory evidence for human reviewers — not a review, an approval, or a CI check. Scripted assertions: 121 passed · 0 failed · 121 total 中文 — 判定:✅ 通过 · 可合入(agent 判定)沙箱验证在隔离、无凭证的容器中执行了该 PR 的代码(与 base 构建 A/B 对照、无 mock harness 断言、定向门禁)。仅作为评审证据,不构成评审、批准或 CI 检查。 脚本断言:121 通过 · 0 失败 · 121 总计 Verification reportPR #9039 — feat(core): Add privacy-safe tool-result boundary diagnostics (follow-up round, identical closure)Verdict: 中文 — 判定:✅ 通过 · 可合入(agent 判定)跟进轮验证:本轮 merge ref 与上轮验证的完全相同(merge
Previous-finding status (follow-up round)The previous round verified this identical merge commit (
Central claim + A/BCentral claim: when file debug logging is enabled, oversized/mutated tool-result representations emit Secondary claims: (1) equal values share one HMAC per process and the first changed boundary produces the first HMAC mismatch; (2) multi-tool aggregate writer events keep per-call HMACs and artifact summaries slot-aligned (Reviewer Test Plan step 5). Scenario (identical in every cell): esbuild-bundled repo
Witnesses: Key measured facts (79 E2E assertions, all scripted in
Finalizer de-gating re-proven through the compiled dist ( CorrectionsNone — the previous rounds' descriptions of the code checked out on re-inspection. FindingsNo blocking findings. Two non-blocking coverage gaps, both carried over and re-measured (completeness reporting, not merge conditions):
Mutation matrix (14 single-point mutants incl. the positive control; unmutated controls green at 18/18, 25/25, 89/89, 33/33, 368/368 core and 11/11 cli; witness
No mutant regressed a previously-killed guard to survived; the two survivors are byte-for-byte the same unpinned axes the previous rounds reported, re-measured on identical test counts for those suites (18 and 33). Targeted gatesAll green, run this round at the verified merge tree (vitest JSON reporters, logs
The first seven core rows sum to 809 and the seven cli rows excluding useGeminiStream sum to 1,245 — exactly the previous round's "core 809/809, cli 1,245/1,245" gate set, reproduced at the identical closure; Not covered
MethodologyEnvironment: CI verify container ( Evidence imagesHarness scripts and raw logs are in the workflow run artifacts (7-day retention). — Qwen Code · sandboxed verification |
|
Triage re-run completed — no new commits since the previous pass, so this run re-verified rather than re-reviewed. No approval was issued, and no new review was submitted. What changed since the last pass is the thread state, not the code: @yiliang114 (core code owner, admin) submitted a full-pass APPROVED review at What did not change: the verdict. Confidence remains 3/5 — the Stage 0 size escalation (1,821 production logic lines, cross-package core) caps the bot below approval by policy, @wenshao's Aug 15 CHANGES_REQUESTED and the review agent's change request still stand on this head, and 本次重跑完成——head 未变,重新验证而非重新评审;未发出批准,也未提交新评审。变化的是线程状态而非代码:@yiliang114(核心 code owner、admin)在 |
yiliang114
left a comment
There was a problem hiding this comment.
Approved after a full pass over the 39-file diff at 21bb6fd.
Verified:
- Privacy holds: boundary events serialize only sizes, mutation flags, closed-set artifact state/kind enums, and HMAC-SHA-256 values keyed by a lazily generated process-local random 32-byte key. Raw tool-result values are destructured out of summaries before JSON.stringify; the only new logger.debug call emits the sanitized event. Session/prompt/tool-call IDs and tool names are HMAC'd, never raw.
- Gating and containment: observation early-returns before value extraction unless file debug logging is enabled (expensive call sites pass lazy thunks); 50 events/minute cap with clock-rewind-safe window reset; every integration point wraps diagnostics in try/catch, and headless/ACP wire observers exit via WeakMap lookups before any byte scanning when nothing was eligible.
- No wire or schema leakage: boundaryArtifact is stripped from task_execution rawOutput before ACP emission, deleted before chat recording, and never added to Headless frames; wire byte counts reuse strings already built for writing.
- All 137 review threads are resolved at this head (7 Criticals fixed, incl. fixes landing in 21bb6fd; the deferred Criticals are diagnostics-only completeness cases acknowledged under the late-round rule).
- CI on this exact head: no failures — only bot-routing skips/cancels (fork-gated build/test workflows don't run without maintainer approval).
|
@qwen-code /triage |
|
Sandboxed verification: ✅ passed — merge-ready (agent verdict) - workflow run Ran the PR in an isolated, token-free container: A/B against the base build, mock-free harness assertions, targeted gates. Advisory evidence for human reviewers — not a review, an approval, or a CI check. Scripted assertions: 105 passed · 0 failed · 105 total 中文 — 判定:✅ 通过 · 可合入(agent 判定)沙箱验证在隔离、无凭证的容器中执行了该 PR 的代码(与 base 构建 A/B 对照、无 mock harness 断言、定向门禁)。仅作为评审证据,不构成评审、批准或 CI 检查。 脚本断言:105 通过 · 0 失败 · 105 总计 Verification report<!-- qwen-triage:verify --> Sandboxed verification: ✅ passed — merge-ready (agent verdict) - [workflow run](this run) Ran the PR in an isolated, token-free container: A/B against a freshly rebuilt base control at the NEW base tip, mock-free harness assertions, mutation matrix, targeted gates. Advisory evidence for human reviewers — not a review, an approval, or a CI check. Scripted assertions: 105 passed · 0 failed · 105 total 中文 — 判定:✅ 通过 · 可合入(agent 判定)跟进轮验证:PR head 与上轮逐字节相同(
Verification reportPR #9039 — feat(core): Add privacy-safe tool-result boundary diagnostics (follow-up round at a NEW base)Verdict: Unlike the previous round (which verified the same PR head against base Previous-finding status (follow-up round)
Both re-measurements ran on the new merged tree; the suites' test counts are identical to the previous round (18 and 33), i.e. the base delta did not touch those fixtures. Central claim + A/BCentral claim: when file debug logging is enabled, oversized/mutated tool-result representations emit Secondary claims: (1) equal values share one HMAC per process and the first changed boundary produces the first HMAC mismatch; (2) multi-tool aggregate writer events keep per-call HMACs and artifact summaries slot-aligned (Reviewer Test Plan step 5). Scenario (identical in every cell): esbuild-bundled repo
Witnesses: Key measured facts (65 E2E assertions, all scripted in
Reviewer Test Plan walk-through
Finalizer de-gating re-proven through the compiled dist ( Corrections
FindingsNo blocking findings. Two non-blocking coverage gaps, both carried over and re-measured (completeness reporting, not merge conditions):
Completeness notes from the matrix (not findings about the PR):
Mutation matrix (9 single-point mutants + 6 unmutated controls + 1 E2E adapter control; sources byte-restored and sha-verified after each mutant; witness
No mutant regressed a previously-killed guard to survived; the two survivors are the same unpinned axes the previous rounds reported, re-measured at identical suite counts. Targeted gatesAll green, run this round at the verified merge tree (vitest JSON reporters, logs
The first six rows are the mutation-matrix unmutated controls; the rest are the remaining changed suites (the ACP session suites moved under Suite liveness is proven by the mutation matrix itself (8 mutants/controls turned suites red this round). Not covered
MethodologyEnvironment: CI verify container ( Evidence imagesHarness scripts and raw logs are in the workflow run artifacts (7-day retention). — Qwen Code · sandboxed verification Evidence imagesHarness scripts and raw logs are in the workflow run artifacts (7-day retention). — Qwen Code · sandboxed verification |
|
Triage re-run completed without a new review. The bot already has a review of its own on 机器人在 The stage comments above were updated with the latest result. View workflow run. 上方各阶段评论已更新为最新结果。查看工作流运行。 |
qwen-code-ci-bot
left a comment
There was a problem hiding this comment.
Partially reviewed — gaps disclosed.
Not reviewed: reverse audit — reached the 5-round cap without two consecutive dry rounds (rounds 3-5 each reported candidates; every confirmed finding verified).
Not reviewed: build-and-test — Integration Tests (CLI, No Sandbox) was skipped in CI and its suite did not run locally.
Not explored to full depth (tool budget reached): "agent reverse-audit (round 5)": none — no check was cut short.; "agent reverse-audit (round 4)": none — all checks above ran to completion within budget.; "agent reverse-audit (round 1)": none — no check was cut short.; chunk 8: npm install in worktree absent; could not run vitest run coreToolScheduler.test.ts to confirm the new tests pass — verified assertions statically against sche…; chunk 8: worktree has no node_modules; could not run vitest run coreToolScheduler.test.ts to confirm the new tests pass — assertions were verified statically against t….
[Critical] R9-7 (packages/cli/src/ui/hooks/useGeminiStream.ts, re-check of existing blocker, comment 3787318576): prompt correlation is still keyed only by callId — duplicate call IDs across different prompts collapse attribution to the later prompt. Still stands at HEAD 21bb6fd (re-verified at useGeminiStream.ts:4412 this round); mechanism unchanged (author deferred as diagnostic-fidelity-only).
[Critical] R9-8 (packages/core/src/core/coreToolScheduler.ts, re-check of existing blocker, comment 3787318566): the producer fallback is still installed only after throwable scheduler prelude reads (getMessageBus()/getDisableAllHooks() at 4455-4456 precede the producerObserved/observeSyntheticProducer declarations at 4457-4458); a throwing prelude publishes a terminal response with no producer observation. Still stands at HEAD 21bb6fd (re-verified this round; author deferred as diagnostic-coverage-only).
[Critical] R9-9 (packages/core/src/followup/speculation.ts, re-check of existing blocker, comment 3787318582): the speculation prompt ID is still absent from every boundary event — observeSpeculationProducer (speculation.ts:340) sets sessionId/toolCallId/toolName only, and finalizeToolResponses (speculation.ts:511) receives no promptIds map. Still stands at HEAD 21bb6fd (re-verified this round; author deferred as diagnostic-correlation-only).
[Critical] R9-10 (packages/core/src/utils/tool-result-boundary-diagnostics.ts, re-check of existing blocker, comment 3787318590): the broad artifact catch (line 165) still returns { state: 'undecided', kinds: [] } on artifact iteration failure, even when persistedOutputFiles already proves the result reusable/file-backed. Still stands at HEAD 21bb6fd (re-verified this round; author deferred as diagnostic-summary-fidelity-only).
[Critical] R9-11 (packages/core/src/utils/tool-result-boundary-diagnostics.ts, re-check of existing blocker, comment 3787318571): rate-limit state (suppressedCount reset + emittedInWindow increment at 321-322) is still consumed before logger.debug (323) succeeds — a throwing logger consumes quota without writing the event. Still stands at HEAD 21bb6fd (re-verified this round; author deferred as diagnostics-only recovery behavior).
中文说明
仅完成部分审查,审查缺口已披露。
未审查:reverse audit — reached the 5-round cap without two consecutive dry rounds (rounds 3-5 each reported candidates; every confirmed finding verified)。
未审查:build-and-test — Integration Tests (CLI, No Sandbox) was skipped in CI and its suite did not run locally。
未探索到全部深度(达到工具调用预算):"agent reverse-audit (round 5)":none — no check was cut short.;"agent reverse-audit (round 4)":none — all checks above ran to completion within budget.;"agent reverse-audit (round 1)":none — no check was cut short.;chunk 8:npm install in worktree absent; could not run vitest run coreToolScheduler.test.ts to confirm the new tests pass — verified assertions statically against sche…;chunk 8:worktree has no node_modules; could not run vitest run coreToolScheduler.test.ts to confirm the new tests pass — assertions were verified statically against t…。
[Critical] R9-7 (packages/cli/src/ui/hooks/useGeminiStream.ts, re-check of existing blocker, comment 3787318576): prompt correlation is still keyed only by callId — duplicate call IDs across different prompts collapse attribution to the later prompt. Still stands at HEAD 21bb6fd (re-verified at useGeminiStream.ts:4412 this round); mechanism unchanged (author deferred as diagnostic-fidelity-only).
[Critical] R9-8 (packages/core/src/core/coreToolScheduler.ts, re-check of existing blocker, comment 3787318566): the producer fallback is still installed only after throwable scheduler prelude reads (getMessageBus()/getDisableAllHooks() at 4455-4456 precede the producerObserved/observeSyntheticProducer declarations at 4457-4458); a throwing prelude publishes a terminal response with no producer observation. Still stands at HEAD 21bb6fd (re-verified this round; author deferred as diagnostic-coverage-only).
[Critical] R9-9 (packages/core/src/followup/speculation.ts, re-check of existing blocker, comment 3787318582): the speculation prompt ID is still absent from every boundary event — observeSpeculationProducer (speculation.ts:340) sets sessionId/toolCallId/toolName only, and finalizeToolResponses (speculation.ts:511) receives no promptIds map. Still stands at HEAD 21bb6fd (re-verified this round; author deferred as diagnostic-correlation-only).
[Critical] R9-10 (packages/core/src/utils/tool-result-boundary-diagnostics.ts, re-check of existing blocker, comment 3787318590): the broad artifact catch (line 165) still returns { state: 'undecided', kinds: [] } on artifact iteration failure, even when persistedOutputFiles already proves the result reusable/file-backed. Still stands at HEAD 21bb6fd (re-verified this round; author deferred as diagnostic-summary-fidelity-only).
[Critical] R9-11 (packages/core/src/utils/tool-result-boundary-diagnostics.ts, re-check of existing blocker, comment 3787318571): rate-limit state (suppressedCount reset + emittedInWindow increment at 321-322) is still consumed before logger.debug (323) succeeds — a throwing logger consumes quota without writing the event. Still stands at HEAD 21bb6fd (re-verified this round; author deferred as diagnostics-only recovery behavior).
— qwen3.8-max via Qwen Code /review (v0.21.12)
|
@qwen-code /triage |
|
Sandboxed verification: ✅ passed — merge-ready (agent verdict) - workflow run Ran the PR in an isolated, token-free container: A/B against the base build, mock-free harness assertions, targeted gates. Advisory evidence for human reviewers — not a review, an approval, or a CI check. Scripted assertions: 77 passed · 0 failed · 77 total 中文 — 判定:✅ 通过 · 可合入(agent 判定)沙箱验证在隔离、无凭证的容器中执行了该 PR 的代码(与 base 构建 A/B 对照、无 mock harness 断言、定向门禁)。仅作为评审证据,不构成评审、批准或 CI 检查。 脚本断言:77 通过 · 0 失败 · 77 总计 Verification report<!-- qwen-triage:verify --> Sandboxed verification: ✅ passed — merge-ready (agent verdict) - [workflow run](this run) Ran the PR in an isolated, token-free container: A/B against a freshly rebuilt base control, mock-free harness assertions, mutation matrix, targeted gates. Advisory evidence for human reviewers — not a review, an approval, or a CI check. Scripted assertions: 77 passed · 0 failed · 77 total 中文 — 判定:✅ 通过 · 可合入(agent 判定)跟进轮验证:PR head、base、merge 三个 OID 与上轮逐位相同(
Verification reportPR #9039 — feat(core): Add privacy-safe tool-result boundary diagnostics (follow-up round at an identical closure)Verdict: Follow-up status: previous findings (re-measured at the same head)The input closure is provably identical to the previous round —
No declined or deferred rows existed in the previous round; nothing worsened (both survivors re-measured at identical suite counts). Central claim + A/BCentral claim: when file debug logging is enabled, oversized/mutated tool-result representations emit Secondary claims: (1) equal values share one HMAC per process and the first changed boundary produces the first HMAC mismatch; (2) multi-tool aggregate writer events keep per-call HMACs and artifact summaries slot-aligned (Reviewer Test Plan step 5). Scenario (identical in every cell): repo
Witness: Key measured facts (43 E2E + 4 budget assertions, all scripted in
Reviewer Test Plan walk-through
CorrectionsNone this round. (The previous round's FindingsNo blocking findings. Two non-blocking coverage gaps, both carried over and re-measured (completeness reporting, not merge conditions):
Completeness notes (not findings about the PR):
Mutation matrix (6 mutants + 2 wiring probes; sources byte-restored and git-status-verified clean after each; witness
No mutant regressed a previously-killed guard to survived; the two survivors are the same unpinned axes the previous rounds reported. Targeted gatesAll green, run this round at the verified merge tree (vitest JSON reporters,
Every count is byte-identical to the previous round's, consistent with the bit-identical closure. Suite liveness is proven by the matrix itself (M1/M7/M10/MPC-AGENT/M8-ADAPTER-E2E turned suites or cells red this round). Not covered
MethodologyEnvironment: CI verify container ( Evidence imagesHarness scripts and raw logs are in the workflow run artifacts (7-day retention). — Qwen Code · sandboxed verification Evidence imagesHarness scripts and raw logs are in the workflow run artifacts (7-day retention). — Qwen Code · sandboxed verification |
|
Triage re-run completed without a new review. The bot already has a review of its own on 机器人在 The stage comments above were updated with the latest result. View workflow run. 上方各阶段评论已更新为最新结果。查看工作流运行。 |
Already have 2 approves,3ks.
|
Released in v0.21.14. |


























What this PR does
This PR adds opt-in, privacy-safe debug events that correlate oversized or mutated tool-result representations across producer, model finalization, session recording, ACP and Headless projection, and the actual transport writer. Each event records exact size measurements, process-local HMAC-SHA-256 values, mutation state, and a closed artifact state/kind summary without recording result text, paths, prompts, or raw identifiers.
The correlation covers built-in interactive, Headless JSON/stream-json, persistent SDK transport, ACP live/replay, subagent, and speculative execution paths. Aggregate writer events preserve one artifact summary per tool call in call order, while weak associations keep diagnostic metadata out of ACP and Headless wire schemas. Diagnostics remain disabled unless file debug logging is active, are capped at 50 eligible events per minute, and swallow all diagnostic failures.
Why it's needed
Large tool results can be persisted, finalized, recorded, projected, or serialized at different boundaries. Existing logs could show that a large frame occurred but could not safely identify the first boundary that changed a representation. These events make that transition observable without exposing sensitive tool output or stable cross-process fingerprints, completing the diagnostics contract tracked by #8448.
Reviewer Test Plan
How to verify
QWEN_DEBUG_LOG_FILEand run a deterministic tool that returns an oversized text result through Headless JSON, stream-json, and ACP. Confirm boundary events appear only for oversized or mutated representations, equal values retain the same HMAC within the process, and the first changed representation produces the first HMAC mismatch.Evidence (Before & After)
Before: the deterministic 499,999-byte fixture completed with the expected bounded transport previews and intact producer artifact, but the debug log contained no tool-result boundary events, so the first mutating boundary could not be identified.
After: Headless JSON recorded a 71,252-byte writer buffer exactly matching stdout, with
tool_result.contentat exactly 65,536 JSON UTF-8 bytes. Stream-json recorded a 65,800-byte target frame, and ACP recorded a 131,349-byte target frame; both matched the captured NDJSON bytes. The producer artifact remained 499,999 bytes with SHA-2560004a69109d580ed184c0772fd37b888f02e7f5bcd05a18a7a8288bee395adeb. Debug-on and debug-off runs produced identical target wire bytes, and all new boundary events were free of fixture text, prompts, raw identifiers, tool names, and artifact paths. This is a non-UI diagnostic change; screenshots are N/A.Tested on
Environment (optional)
macOS 26.4.1, Node.js 24.12.0, npm 10.9.8, local unsandboxed CLI with deterministic OpenAI and MCP fixtures.
Risk & Scope
Linked Issues
Closes #8448
中文说明
本 PR 做了什么
本 PR 新增了按需启用、隐私安全的调试事件,用于关联 tool result 在 producer、模型终结、会话录制、ACP 与 Headless 投影以及实际传输 writer 各边界上的超大或变异表示。每条事件记录精确的大小指标、进程内 HMAC-SHA-256、变异状态以及闭合集合的 artifact 状态/类型摘要,但不会记录结果正文、路径、提示词或原始标识符。
关联范围覆盖内置交互模式、Headless JSON/stream-json、持久 SDK transport、ACP live/replay、subagent 和 speculative execution 路径。聚合 writer 事件按 tool call 顺序保留逐项 artifact 摘要,弱关联机制保证诊断元数据不会进入 ACP 或 Headless wire schema。只有文件调试日志处于启用状态时诊断才会运行,每分钟最多记录 50 条 eligible 事件,并吞掉所有诊断异常。
为什么需要它
大型 tool result 可能在持久化、终结、录制、投影或序列化等不同边界发生变化。现有日志只能显示发生了大帧,无法在不泄露内容的前提下定位第一个改变表示的边界。这些事件可以在不暴露敏感工具输出、也不创建跨进程稳定指纹的情况下观察该变化,完成 #8448 跟踪的诊断契约。
Reviewer 测试计划
如何验证
QWEN_DEBUG_LOG_FILE,通过 Headless JSON、stream-json 和 ACP 运行一个返回超大文本的确定性工具。确认只为超大或已变异表示生成边界事件;相同值在同一进程内保持相同 HMAC;第一个发生变化的表示产生第一个 HMAC 不匹配。证据(Before & After)
Before:确定性的 499,999 字节 fixture 会生成预期的有界传输预览,producer artifact 也保持完整,但调试日志中没有 tool-result 边界事件,因此无法定位第一个发生变异的边界。
After:Headless JSON 记录的 writer buffer 为 71,252 字节,与 stdout 精确一致,且
tool_result.content恰好为 65,536 JSON UTF-8 字节。stream-json 记录的目标 frame 为 65,800 字节,ACP 记录的目标 frame 为 131,349 字节,均与捕获的 NDJSON 字节完全一致。producer artifact 仍为 499,999 字节,SHA-256 为0004a69109d580ed184c0772fd37b888f02e7f5bcd05a18a7a8288bee395adeb。启用和关闭调试诊断时目标 wire 字节完全相同,所有新增边界事件均不包含 fixture 正文、提示词、原始标识符、工具名或 artifact 路径。这是非 UI 诊断变更,截图不适用。测试平台
环境(可选)
macOS 26.4.1、Node.js 24.12.0、npm 10.9.8、本地无 sandbox CLI,以及确定性的 OpenAI 和 MCP fixture。
风险与范围
关联 Issue
Closes #8448