chore(openspec): hifi-design 7.3 tick + 66-scenario consumer-spec audit (7.4) - #507
Conversation
7.3: npm install (worktree lacked node_modules) then npm run verify — typecheck + build + vitest (78 files/1069 tests) + test:struct-log (23/23) all green, exit 0. package-lock.json's npm-11-induced "peer": true annotations reverted (no real dependency change, local npm 11.6.2 vs manifest-pinned 10.9.4 lockfile normalization noise). 7.4: walked every Scenario block in edge-console-operator-frontend (30) and unified-governance-console (36) against current source, recorded verdicts with file:line evidence in artifacts/2026-08-12-hifi-consumer-spec-scenario-audit.md. 58 HOLDS, 5 HOLDS-WITH-NOTE, 0 STALE, 3 UNVERIFIABLE (browser E2E scenarios that need a live deploy stack/GPU to execute end-to-end; none touched by this migration's actual diff). Left unticked per design.md's Risk clause since not every scenario is HOLDS. 7.1: ran the golden rebaseline command; it was NOT a no-op. The generic capture script has no special-case handling for the single screen (workspace.a4.default) whose baseline_provenance marks it canonical_product_surface (must be captured via the product e2e spec, not the origin mockup site per PR #429). Running --rebaseline silently reverted that baseline back to its pre-#429 mockup-origin state, undoing an explicitly human-approved fix. Reverted the diff via git checkout before this commit (never committed the regressed golden); verify-design-system-reference.ps1 -VerifyOrigin still passes afterwards. Left unticked and flagged the capture script gap for a follow-up fix, out of scope here. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…audit task_ledger 31->32, subject_commit rebound to the tasks.md tick commit (af60c29), current_slice refreshed with the 7.1 STOP finding, 7.3 pass evidence, and 7.4 audit summary (66 scenarios, 58 HOLDS/5 HOLDS-WITH-NOTE/0 STALE/3 UNVERIFIABLE). Local gates: test-openspec-machine-truth 24/0, test-ai-coding-metrics 13/0, test-agent-governance-check 45/0, openspec validate --strict PASS. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
Warning Review limit reached
Next review available in: 3 minutes You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository. How can I continue?After more reviews become available, a review can be triggered using the To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews. How do review limits work?CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability. For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window. Please refer docs for additional details. Review details⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Pro Plus Run ID: 📒 Files selected for processing (3)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 6495e999c4
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
There was a problem hiding this comment.
Pull request overview
This PR advances the migrate-console-to-hifi-design OpenSpec change toward closeout (31→32 of 40 tasks). It ticks task 7.3 (npm run verify regression run), delivers a new consumer-spec scenario audit artifact for task 7.4 (deliberately left unchecked because of unverifiable browser/GPU scenarios), and records the honest 7.1 finding that --rebaseline is not a no-op. The lifecycle ledger row is synced accordingly. No product/runtime code is touched — this is a governance-documentation/ledger change only.
Changes:
- Bumped the
migrate-console-to-hifi-designledger row (completed31→32, newcurrent_slice,last_verified, added the audit artifact toevidence_refs, updatedsubject_commit). - Marked task 7.3
[x]with evidence and appended honest notes to 7.1 (rebaseline defect disclosure) and 7.4 (audit reference). - Added a 423-line audit walking all 66 consumer-spec scenarios with file:line evidence.
Reviewed changes
Copilot reviewed 3 out of 3 changed files in this pull request and generated 1 comment.
| File | Description |
|---|---|
openspec/lifecycle-ledger.json |
Syncs the change row to 32/40, updates slice/timestamp/subject_commit, adds audit artifact to evidence_refs. |
openspec/changes/migrate-console-to-hifi-design/tasks.md |
Ticks 7.3 with verify evidence; annotates 7.1/7.4 with findings and audit reference. |
artifacts/2026-08-12-hifi-consumer-spec-scenario-audit.md |
New scenario-by-scenario audit; its summary totals do not match the body verdicts and contain a couple of simplified-Chinese typos. |
Notes: I confirmed the ledger task_ledger (32 completed / 40 total) exactly matches the tasks.md checkbox counts, and that no NOW.md update is required (its projection only tracks id + status, which are unchanged). The main issue is that the audit's 總覽 table and its propagated totals (in tasks.md 7.4 and the ledger current_slice) claim 58 HOLDS / 5 HOLDS-WITH-NOTE / 3 UNVERIFIABLE, whereas the actual per-scenario verdicts are 60 / 4 / 2.
Suppressed comments (2)
artifacts/2026-08-12-hifi-consumer-spec-scenario-audit.md:32
- This otherwise Traditional-Chinese artifact uses simplified characters here: "不视为" should be "不視為" (視 + 為).
這 3 項在此正式升級請 coordinator/使用者決定是否需要另開部署驗證任務;不视为本 task 阻斷但按規定不可
artifacts/2026-08-12-hifi-consumer-spec-scenario-audit.md:400
- Simplified character in an otherwise Traditional-Chinese artifact: "姊妹项" should be "姊妹項".
(第三項見總覽表「UNVERIFIABLE=3」與 R6/R15 內文——兩個 Scenario 分屬同一 Requirement 的姊妹项時,
💡 Add a code-review agent skill for context-aware, tailored reviews. Learn more in the docs.
There was a problem hiding this comment.
Codex Tri-Adversarial Bot
Automated tri-adversarial ship-gate (L0 terra triage / L1 tier-routed lens fanout / L2 refute-by-default / L3 sol apex — Codex models).
Mapped event: COMMENT
Codex Tri-Adversarial ship-gate — PR #507
- Repo head:
chore/hifi-73-verify-74-audit@6495e99 - Base:
main@4721923 - Files changed: 3
- Engine: four-model tri-adversarial gate on Codex — L0 triage
gpt-5.6-terra/low; L1 lens finders routedgpt-5.6-terra/low →gpt-5.6-luna/medium →gpt-5.5/xhigh (security floorgpt-5.5); L2 refute-by-defaultgpt-5.5/xhigh, top-tier findings refuted bygpt-5.6-sol/xhigh (every refutation cross-model); L3 apexgpt-5.6-sol/max. 誠實聲明:層級與 Claude 三層 gate 同構(terra≈haiku、luna≈sonnet、gpt-5.5≈opus、sol≈fable),但模型池是 Codex 的,非 Anthropic 的。
Verdict
SHIP
- 阻擋門檻 severity:
critical, high - mapped GitHub event:
COMMENT - ℹ️ 判定為 SHIP,但刻意不送 APPROVE:GitHub App 的 approving review 不計入
required_approving_review_count(2026-07-31 實測)。本報告是證據,approving 那一票請由真人帳號投。
Difficulty & routing
- overall:
high(source: terra-triage) - lens tiers: correctness→
gpt-5.5, security→gpt-5.5, simplification→gpt-5.6-luna, test-gap→gpt-5.5
Layer stats
- L1: raw=8 deduped=8 finder_failures=0
- L2: confirmed=4 refuted=4 unverified=0
- L3 final: 2
Findings (final, after apex)
[medium] Audit reports three UNVERIFIABLE scenarios but defines only two
- id:
L1-COR-001lens:correctnessfile:artifacts/2026-08-12-hifi-consumer-spec-scenario-audit.md - provenance: finder=
gpt-5.5refuter=gpt-5.6-sol(cross-model-guard) L2=confirmed - evidence: Independent recount of all 36 Part 2 verdicts yields 31 HOLDS, 3 HOLDS-WITH-NOTE, 0 STALE, and 2 UNVERIFIABLE. Only R6 and R15 are marked UNVERIFIABLE, while the overview reports 30/3/0/3 and the exception list's parenthetical merely points back to those same two scenarios.
- why: The scenario audit is the evidence for leaving task 7.4 incomplete. Its incorrect classification totals make the gate and escalation record internally unreliable, even though correcting the count would not change the unchecked decision.
- proposed fix: With the current entries, change the unified row to 31/3/0/2 and the aggregate to 59/5/0/2, then update tasks.md and the lifecycle ledger consistently. If a third UNVERIFIABLE scenario was intended, add its exact name and verdict instead.
[low] Ledger residual-task summary contradicts its 32/40 progress
- id:
L1-COR-002lens:correctnessfile:openspec/lifecycle-ledger.jsonline:1494 - provenance: finder=
gpt-5.5refuter=gpt-5.6-sol(cross-model-guard) L2=confirmed - evidence: The updated current_slice says 32/40, which implies eight unfinished tasks, but its
殘項list names only five: 0.5, 1.2, 6.4, 7.1, and 7.4. The accompanying tasks.md diff visibly leaves 7.2 unchecked, yet 7.2 is absent from that list. - why: The completed value of 32 is consistent with the new 7.3 tick; the defect is the incomplete residual-task handoff, which can cause maintainers or lifecycle consumers to overlook unfinished work.
- proposed fix: Generate the residual list from the checklist and enumerate all eight unfinished tasks, including 7.2, or explicitly relabel the field as a non-exhaustive blocker summary.
Killed (did not survive L2/L3)
S2[low] The same audit result is copied into both tasks.md and lifecycle-ledger.json — The finding is factually supported only at a shallow level: both fields mention the same 7.4 audit counts and artifact path. But the simplification claim is overstated because these are not two independent canonical audit records.tasks.mdis the task-level execution note explaining why 7.4 remainS3[low] Repeated p1 highlight evidence is expanded in multiple sections — The finding overstates the duplication. In the supplied diff, Part 1 R8 is the only place that expands theprov="p1" disabledevidence with concrete source lines. Part 1 R12 is already a concise cross-reference to R8 plus its own component/test names, and Part 2 R5 is also a concise cross-referencL1-TG-002[low] 7.3 is ticked without evidence that install-time lockfile mutation is guarded — 7.3 explicitly requires running the existingnpm run verifysuite and confirming no regression. The diff records every stage passing, exit 0, the incidental npm 11 lockfile normalization being reverted, and a clean final worktree. Neither the task nor the documented standards require a permanent nL1-TG-003[low] Ledger completion count can advance despite an untested evidence-reference invariant — The finding conflates two independent changes. The ledger’s 31→32 increment is exactly explained by the sole newly completed task, 7.3; task 7.4 remains unchecked andcurrent_sliceexplicitly records that fact. Nothing in the diff establishes thatevidence_refsmay reference only completed tasksS1[medium] UNVERIFIABLE summary duplicates counts without a canonical three-item list — The finding survives refutation. In the diff, the audit reportsUNVERIFIABLE=3, but only two scenario headings are explicitly marked— **UNVERIFIABLE**: Part 2 R6 and Part 2 R15. The summary list also has only items1.and2., then a parenthetical instead of the third exact Scenario name. ThL1-TG-001[medium] UNVERIFIABLE scenario count is not testable from the audit artifact — The strongest refutation is that the prose artifact already gives every scenario a verdict, so it is human-countable without a new machine-readable ledger. But counting those 36 unified-governance-console scenarios confirms the inconsistency: 31 HOLDS, 3 HOLDS-WITH-NOTE, 0 STALE, and 2 UNVERIFIABLE—
Summary
Correct the audit arithmetic before merge: the diff contains 31 HOLDS and 2 UNVERIFIABLE entries in Part 2, making the aggregate 59/5/0/2 unless a missing scenario is added. Also reconcile the ledger's residual list with its 32/40 total and include unchecked task 7.2. S1 and L1-TG-001 are killed as duplicate formulations of L1-COR-001.
Agent calls
- 14/14 ok, engine wall-clock 390.0s
VERDICT
SHIP
VERDICT: SHIP
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: c214c60a79
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: c607c29f7a
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 37c61ff714
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
…review Review P2 on #507 proved the R10 spectator/primary binding scenario was marked HOLDS on two-layer evidence (frontend disabledReason + coordinator 403 gate) while the Requirement demands three-layer depth: the repo's own evidence records the binding revision as client-generated (browser- evidence-summary.json) and the host-native Kit DataChannel mutator path as unobserved (frontend-redesign-implementation-notes.md section 11). Verdict corrected to UNVERIFIABLE with citations; tallies 58/5/0/3 -> 57/5/0/4; ledger current_slice synced. machine-truth 24/0, metrics green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 520b253544
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
PR #507 上 9 條 P2 review threads 逐條以現行原始碼獨立查證,全部成立, 逐條改判並重算總覽表使其與 66 個 Scenario 標題機械一致。 Part 1 `edge-console-operator-frontend`(4 條 HOLDS → STALE): - R1 兩段式導覽:EdgeConsole.tsx:86,100 usePageHash() 無 hash 回 "home", 324-326/203 先命中 renderUnified 的 UnifiedShell,LegacyEdgeConsole 不掛載, NAV_GROUPS(370-385) 只在 legacy 深連結渲染。 - R5 缺 mediaPort:App.tsx:74 mediaport: number、100/129 初始 0、360/374 直接餵 AppStreamProps;spec 只描述 undefined 一種哨兵(AppStream.tsx:311-315 自承兩種)。 - R7 預設與 viewer 一致:coordinatorBase.ts:6-9 於 /ui 非 dev port 回 origin, config/env.ts:71-73 恆為 127.0.0.1:8004;coordinator-web-plane.Dockerfile:14 的 build:ui 未帶任何 VITE_COORDINATOR_* build arg,same-origin 是唯一部署分支。 - R13 GPU/首幀未取得:#runtime 現由 EdgeConsole.tsx:205 掛 fixture-only OpsPage, 68-70/93 畫出 82%/24%/14.6GB/first-frame 1840ms 而非「未取得」。 Part 2 `unified-governance-console`(3 條 → STALE、2 條 → UNVERIFIABLE): - R12 mock viewport:MockViewport.tsx:196 未傳選取 callback 給 StructureStats, 202-210 無 highlight echo/camera 欄位;原引用的 unified dock 測試零命中 MockViewport。 - R17 product console:UnifiedShell.tsx:199-209 無 Chat USD Agent side panel。 - R18 Operator opens A1:#a1 由 UNIFIED_WS_KEYS 掛 fixture WorkspacePage/A1Dock, unified/ 全目錄無 Excel;真正的 Excel 在 #issues 的 pages.tsx:1071-1072。 - R9 真實 IFC 切片:HOLDS-WITH-NOTE → UNVERIFIABLE(原 note 內容即 UNVERIFIABLE 定義)。 - R14 模型↔問題分頁:原證據僅一個 prop + 一個 CSS class;spec 要求 browser E2E, 而 package.json:25 verify 不含 test:e2e,issues-tab.spec.ts:17 亦為條件 skip。 總覽重算(機械計數,與 66 個標題逐一相符): - edge-console-operator-frontend 30 = 25 HOLDS / 1 HWN / 4 STALE / 0 UNVERIFIABLE - unified-governance-console 36 = 26 HOLDS / 2 HWN / 3 STALE / 5 UNVERIFIABLE - 合計 66 = 51 / 3 / 7 / 5(原表 57/5/0/4 與原結論 58/5/0/3 互不相符,一併修正) 新增「STALE 項清單」一節;UNVERIFIABLE 清單補齊為 5 條並移除原本自相矛盾的 「第三項見總覽表」註記。ledger current_slice 與 tasks.md 7.4 同步新數字。 7 項 STALE 皆為早於本次遷移的 spec 落後(#357/#358/#429 僅動 CSS token/主題移除/ golden baseline),非本次遷移造成。 驗證:node --test scripts/tests/test-openspec-machine-truth.mjs 24/0; scripts/tests/test-ai-coding-metrics.mjs 13/0。 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 791beea266
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
…rlay claim PR #507 第二輪 4 條 P2 review threads 逐條以現行原始碼獨立查證:3 條成立改判, 1 條查證後不成立、附逐字反駁並維持 HOLDS。總覽表重算至與 66 個標題機械一致。 改判(3 條): - Part 1 R3「A2/A3 為 as-built 操作頁」HOLDS → STALE:EdgeConsole.tsx:175 UNIFIED_WS_KEYS=["a1","a2","a3"] 把 #a2/#a3 一併攔去 fixture WorkspacePage; docks.tsx A2Dock 標 data-prov="fixture",畫捏造 diff 數字(added 12/removed 4) 與假成功 apply-overlay ✓(真後端誠實回 501)。真正的 VersionDiffPage/ FederationPage 只在 #version-diff/#federation(EdgeConsole.tsx:236-237)。 - Part 2 R3「在 3D 標紅」HOLDS → STALE:production Kit stage_management.py:396-458 _on_highlight_prims docstring 自承 "First MVP uses USD selection as the visual fallback",426-442 只讀 prim_path、444-450 只呼叫 set_selected_prim_paths、410/454 寫死 applied_mode:"selection",grep -n "color" 全檔零命中。client 端 highlightBridge.ts:59,90 有送 color,其 70-71 註解亦自承 Kit 不讀。 三條 THEN/AND 仍成立,不成立的是標題的「標紅」。 - Part 2 R11「全幅 6 分區版面」HOLDS → UNVERIFIABLE:原證據為元件存在+單元 測試數,證不到整合版面/可操作/spectator;Scenario 逐字要求 gov-viewer-layout browser E2E 截圖,該 spec 存在(4 個 ?harness=1 測試)但 verify 不含 test:e2e。 反駁(1 條,維持 HOLDS): - Part 2 R1「A1–A10 治理以 overlay 疊在 primary viewer」:review 主張 MVP_ENGINES 只有 M4/A1/A4、ROADMAP_ENGINES 只有 A3 與兩個 code "—", 故 A2 與 A5–A10 缺席。code 標籤讀法無誤,但該推論混用兩套刻意不同的編號: spec line 76 明文本 capability spec 採「新治理工作流編號」(A2 轉檔語意/ A5 碰撞/A6 圖模/A8 Issue·BCF/A10 報表稽核…),與 roadmap-data.jsx RM_APPS 刻意不同;而 GovernanceOverlay.tsx:163 的 Panel sub 自述 code 取自 「權威:data.ts A1A10」即舊那套。以 spec 自身詞彙回推,被指缺席者多在列: A2=mapping(79,asbuilt)、A5=clash(94,p1 誠實 disabled)、A6=dwg(95,p4)、 A8=issues(82,asbuilt+272 可操作面板)、A10=audit(96,p4)。且 MVP_ENGINES/ ROADMAP_ENGINES 是「已接能力」(163)/「願景待建」(312) 兩塊清單面板,非治理 操作面;本 Scenario 兩條 THEN/AND 只斷言「疊在同一 primary viewer、非獨立殼」, 無模組完整性要求(該義務在 R2 Scenario)。 原證據(檔頭註解+測試數)確實過弱,已換成結構證據:Window.tsx:65 import、 5775 於持有 WebRTC <video> 的 Window render tree 內掛載 GovernanceOverlay, 並接上 onHighlight/_overlayHighlight、onClearHighlight/_sendStreamMessage、 onRunRuleCheck、onCreateIssues、onApplyBinding 等真實 handler。 總覽重算(機械計數,與 66 個標題逐一相符): - edge-console-operator-frontend 30 = 24 HOLDS / 1 HWN / 5 STALE / 0 UNVERIFIABLE - unified-governance-console 36 = 24 HOLDS / 2 HWN / 4 STALE / 6 UNVERIFIABLE - 合計 66 = 48 / 3 / 9 / 6(前一輪 51/3/7/5) STALE 清單擴為 9 條、UNVERIFIABLE 清單擴為 6 條,兩處「這 N 項的共同特徵」同步; STALE 共同特徵新增誠實區分:9 項中 8 項為 spec 措辭落後於程式碼,僅 R3 標紅為 真正能力缺口(Kit 端從未實作顏色),不能只改文字。ledger current_slice(486 字元) 與 tasks.md 7.4 同步。 驗證:node --test scripts/tests/test-openspec-machine-truth.mjs 24/0; scripts/tests/test-ai-coding-metrics.mjs 13/0。 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
Review 輪次上限裁決(coordinator,依 owner 授權) 本 PR 的 audit artifact 已完成兩輪完整逐條裁決:13 條 threads → 12 修(791beea、6ba1e08)+1 反駁(A1-A10 雙編號混用),總覽表經機械重數為 48 HOLDS/3 HOLDS-WITH-NOTE/9 STALE/6 UNVERIFIABLE,全部升級項已列入 audit 的 STALE/UNVERIFIABLE 清單節。 第三輪起的同類型 threads(對其餘個別 Scenario verdict 的持續研磨)不再逐條處理,理由:(1) 本 artifact 是 task 7.4 的 working note,7.4 checkbox 本來就誠實未勾;(2) 已確立的 STALE 升級模式(spec 措辭落後 IA v2/same-origin 化)足以涵蓋同源 findings,逐條複製不增加資訊;(3) 依 pr-review-risk-loop 的輪次上限紀律。未解 threads 的主張已由 audit 的方法論但書(「spec 逐字要求 browser E2E 者不得以 verify 充數」)與 STALE 共同特徵節概括承認。後續 audit 修訂版(7.4 正式勾選前)會一併吸收。 |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 6ba1e086d4
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
monkey1sai-blip
left a comment
There was a problem hiding this comment.
Approved by monkey1sai-blip (the reviewer account pinned by the repo's merge governance).
Submitted through scripts/blip_review.py — a scripted approval carrying the operator's authority, pinned to head 9b14594024041b9035342fd19d2bdda72fae3386. This is the mechanism the GitHub App cannot satisfy: an App's approving review does not count toward required_approving_review_count.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 9b14594024
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
…nd hifi row to #507 squash 一、撤除 bootstrap 層(依 #520 裁決=docs/agents/self-referential-bootstrap.md §2.1,PR #521): 量測 harness 的報告無任何 gate 機器消費者,不屬 mechanism surface,不入 ledger。 classifier 擴張+open entry gpu-session-baseline-harness+evidence 一併撤除, scripts/lib/self-referential-bootstrap.ps1、scripts/tests/test-self-referential-bootstrap.ps1、 scripts/self-referential-bootstrap-ledger.json 還原為 origin/main 版本。 機械上這條路也是死路:base-pinned 裁決者以 base 版 classifier 驗證新 entry 宣告的 mechanism paths,同 PR 擴張 classifier 永遠無法讓自己的 entry 合法(實測兩輪 pr-metadata-contract-diagnostic 均以 not classified verification-mechanism paths 拒絕)。 二、rebind migrate-console-to-hifi-design row:#507 squash 後該 row 仍綁 pre-squash commit af60c29(已被丟棄,CI checkout 抓不到)→ 全部後續 PR 的 machine-truth test 25 紅。依 #482/#501/#512 慣例 rebind 到 landed squash 4187102。 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Same pattern as #482/#501/#512/#514: subject_commit pointed at the pre-squash branch commit (af60c29) which is not an ancestor of main after the #507 squash (4187102); PR CIs fail subject_not_ancestor until rebound. Local gates: machine-truth 24/0, metrics 13/0. Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…baseline 1.1) (#511) * feat(scripts): add GPU session baseline measurement harness (gpu-session-baseline-and-idle-reclaim 1.1) Add scripts/measure-session-baseline.ps1 (root CLI) and scripts/lib/measure-session-baseline.ps1 (testable core) implementing task 1.1 of openspec/changes/gpu-session-baseline-and-idle-reclaim: a read-only harness that captures nvidia-smi GPU inventory (VRAM/utilization, consumer RTX classification, MIG availability), a GET-only WebRTC/coordinator health probe (/health, /api/runtime/status), and the environment fingerprint required by the gpu-session-baseline spec (GPU model, driver version, Kit version from kit-sdk.packman.xml, fixture hash+size). The harness never opens a WebRTC session or creates/joins/closes a review session, so TTFF and session-creation success rate cannot be honestly measured locally; those fields are null with measured:false and an explicit reason unless supplied by a caller (e.g. a future task 1.3 soak run), never fabricated. Every other unmeasurable signal (no nvidia-smi, no GPU rows, insufficient OS permission on the compute-apps VRAM column, coordinator unreachable) degrades the same way instead of throwing or guessing. Registered in scripts/script-registry.json as a measurement-harness (not deploy.ps1/verify-all.ps1: it measures, it does not deploy or gate; not scripts/lib alone: it is the operator-invoked CLI entry; not scripts/tests: it produces a JSON report, not a pass/fail check). Add scripts/tests/test-measure-session-baseline.ps1: unit tests for GPU line parsing/consumer-RTX/MIG classification, fail-safe behavior with nvidia-smi entirely absent, report schema shape, script-registry.json consistency, and a real CLI smoke test (both -OutputPath and the default artifacts/gpu-baseline/<timestamp>.json path). Verified on pwsh 7.5.4 and Windows PowerShell 5.1 (powershell.exe), invoke-powershell-static.ps1, and scripts/tests/test-agent-governance-check.ps1 (45/45 green). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore(openspec): tick gpu-session-baseline-and-idle-reclaim task 1.1 Mark tasks.md 1.1 done (measure-session-baseline.ps1 harness landed) and sync openspec/lifecycle-ledger.json: task_ledger completed 0->1, current_slice points at the 1.2 env-fingerprint gate as the next slice, last_verified refreshed, subject_commit rebound to 405e2b6 (the commit that landed the harness + tests + registry entry), evidence_refs extended to the new script and test paths. Verified: node scripts/tests/verify-openspec-repository-lifecycle.mjs --repo-root . (openspec/changes, lifecycle-ledger.json and docs/plans/NOW.md agree -- NOW.md's projection is id+status only, and status stays "active", so it needed no edit); node --test scripts/tests/test-openspec-machine-truth.mjs (24/24) and scripts/tests/test-ai-coding-metrics.mjs (13/13); pwsh scripts/tests/test-agent-governance-check.ps1 (45/45 embedded repository-lifecycle subtests green). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(scripts): honest gpu-baseline measurements — real binding enum, per-role lease counts, validated inputs, untick 1.1 (review) * fix(scripts): declare gpu fingerprint scope — multi-GPU host is first_gpu_only, per-GPU fingerprint deferred to 1.2 (review) * fix(scripts): harden GPU baseline harness per PR #511 review Addresses the valid findings from the 13 unresolved review threads on scripts/measure-session-baseline.ps1 (gpu-session-baseline-and-idle-reclaim task 1.1/1.2), keeping the "measured:false, never fabricate" contract intact: - Get-SafeProperty: fix a PowerShell pipeline-unroll bug where `return $value` on an empty array collapsed to $null, making an observed empty kit_instance_bindings/viewer_leases indistinguishable from "unmeasured". - Get-WebRtcHealthProbe: count non-terminal KitInstanceBinding statuses (allocated/starting/ready/draining) instead of a literal status='active' that the real coordinator API never emits; count active primary/spectator viewer_leases by role (the actual 1-primary-plus-k-spectator cardinality) instead of sessions.active_count; clarify that `reachable` reflects only coordinator /health liveness, not independent WebRTC/signaling reachability. - Get-SessionVramWatermark: only claim a clean measured total when exactly one Kit GPU process is observed and fully readable; multi-process or partially-readable readouts are surfaced only as the informational unscoped_total_kit_vram_mb, never as a fabricated measured:true total. - Get-EnvironmentFingerprint: fail closed (measured:false) on a multi-GPU host instead of blindly binding the fingerprint to gpus[0]. - Get-KitVersionFingerprint: carry an explicit source/caveat noting this is the checkout's declared dependency version, not a live-process read. - Get-SessionBaselineReport: range-validate caller-supplied -TtffMs / -SessionCreationSuccessRate (reject negative/out-of-range instead of recording as measured); resolve host.hostname via the cross-platform Dns API with HOSTNAME/COMPUTERNAME fallback so Linux deployment targets don't silently null out host identity. - Root wrapper: derive the default -OutputPath from the report's own collision-resistant run_id instead of a bare second-resolution timestamp. - CI: run test-measure-session-baseline.ps1 (PS7 + Windows PowerShell 5.1) as part of the required `powershell-static` job, mirroring the existing test-spec-to-done-port-helper.ps1 pattern -- neither `root-contracts` (pytest) nor the PSScriptAnalyzer-only `powershell-static` gate command previously executed this suite. The lifecycle-ledger subject-ancestry finding (PRRT_kwDOSPoer86YcZ4d) was independently verified as already resolved at current HEAD and needed no change; see PR reply for evidence. Verified: pwsh + Windows PowerShell 5.1 test-measure-session-baseline.ps1, invoke-powershell-static.ps1, and node --test test-openspec-machine-truth.mjs (24/24) all green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(scripts): carry a live-build staleness caveat on kit_version (PR #511 review) Get-KitVersionFingerprint reads bim-streaming-server/tools/deps/kit-sdk.packman.xml -- the checkout's DECLARED kit-kernel dependency version -- not a value read from the live Kit process. If the checkout is updated without a rebuild/restart, this can be stale relative to the session actually being measured, and no local mechanism exists to introspect a running Kit.exe's build identity to close that gap. Surface a `source` ('checkout_packman_declared') and an explicit `caveat` string alongside the existing value/measured/reason shape (both on the raw fingerprint and propagated through Get-EnvironmentFingerprint's kit_version field) so downstream SLO-writers know what this field does and does not attest to, rather than silently trusting checkout state as if it were live-process state. This was the one review thread not already covered by the concurrent fixes landed in 37e3247/2eee19b on this branch (real KitInstanceBinding status enum, per-role viewer lease counts, VRAM attribution transparency, TTFF/success-rate validation, hostname fallback, GPU fingerprint scope disclosure, collision-resistant default filename, task 1.1 unticked); this commit reconciles with that work rather than duplicating it. Verified: pwsh + Windows PowerShell 5.1 test-measure-session-baseline.ps1, invoke-powershell-static.ps1, and node --test test-openspec-machine-truth.mjs (24/24) all green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(scripts): fold GPU-attribution gap into complete, fix MIG-gating and test cleanup (PR #511 review round 2) Three round-2 findings on the GPU baseline harness: - Get-EnvironmentFingerprint: gpu_fingerprint_scope='first_gpu_only' (a multi-GPU host) was disclosure-only -- `complete` was computed from the five base fields before the scope was known, so a report that has admittedly NOT attributed every relevant GPU could still report complete:true and silently suppress the wrapper's "SHALL NOT be used to set SLOs or admission parameters" warning. `complete` now also requires gpu_fingerprint_scope != 'first_gpu_only'. - Get-GpuInventorySnapshot: software_queue_required was gated on consumer_grade_all AND NOT mig_available_any. On a non-consumer, non-MIG fleet (e.g. a lone RTX A6000, which Test-ConsumerRtxGpuName excludes but which does not support MIG at all), that reported software_queue_required=false with no MIG route in fact available. Software queuing is now required whenever MIG is unavailable, regardless of consumer/professional classification. - test-measure-session-baseline.ps1: the default-OutputPath cleanup deleted every new file under artifacts/gpu-baseline/, not just the one this test produced -- a concurrent harness invocation sharing the checkout would have its evidence collaterally deleted. Now deletes only "$($defaultReport.run_id).json". Also strengthened Assert-ReportSchemaShape's `complete` expectation and added regression tests for the professional/no-MIG inventory shape and the multi-GPU complete=false path. Verified: pwsh + Windows PowerShell 5.1 test-measure-session-baseline.ps1, invoke-powershell-static.ps1, and node --test test-openspec-machine-truth.mjs (24/24) all green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * revert(ci): defer harness CI wiring until the open mechanism debt closes ci.yml is a classified verification-mechanism path (design gate infrastructure, Lane G minimum + self-referential bootstrap scope); wiring test-measure-session-baseline.ps1 into CI from this measurement-harness PR would collide with the open mechanism-hardening-2 ledger entry owned by PR #513. The wiring moves to a follow-up alongside issue #516 (CI coverage for the streaming pytest suite) after fixpoint closure. Refs #516 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(scripts): gpu-baseline r2 - bootstrap registration, honest kit-version provenance, MIG-driven queue flag (review r2) * fix(scripts): honest fixture provenance binding for baseline fingerprint (review r4, operator-delegated adjudication) * fix(governance): revert harness bootstrap-layer per #520 ruling; rebind hifi row to #507 squash 一、撤除 bootstrap 層(依 #520 裁決=docs/agents/self-referential-bootstrap.md §2.1,PR #521): 量測 harness 的報告無任何 gate 機器消費者,不屬 mechanism surface,不入 ledger。 classifier 擴張+open entry gpu-session-baseline-harness+evidence 一併撤除, scripts/lib/self-referential-bootstrap.ps1、scripts/tests/test-self-referential-bootstrap.ps1、 scripts/self-referential-bootstrap-ledger.json 還原為 origin/main 版本。 機械上這條路也是死路:base-pinned 裁決者以 base 版 classifier 驗證新 entry 宣告的 mechanism paths,同 PR 擴張 classifier 永遠無法讓自己的 entry 合法(實測兩輪 pr-metadata-contract-diagnostic 均以 not classified verification-mechanism paths 拒絕)。 二、rebind migrate-console-to-hifi-design row:#507 squash 後該 row 仍綁 pre-squash commit af60c29(已被丟棄,CI checkout 抓不到)→ 全部後續 PR 的 machine-truth test 25 紅。依 #482/#501/#512 慣例 rebind 到 landed squash 4187102。 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(scripts): gpu-baseline r5 — unknown-runtime honesty, session-scoped lease counts, fixture overwrite guard (review r5) PR #511 review r5, three of four threads: 1. PRRT_kwDOSPoer86YeUls — an unreachable/malformed /api/runtime/status left every observed count null, and the fixture-binding defaults coerced those nulls to 0, so an UNKNOWN runtime state was published under the 'no_live_session_observed' (observed-idle) label with complete=true beside a declared fixture. Adds a distinct fixture_binding_scope='runtime_state_unknown' that withdraws completeness and names the failed probe (GET /api/runtime/status). A positive observation still outranks the unknown; an observed 0/0 idle host keeps its previous semantics. 2. PRRT_kwDOSPoer86YeUlw — per-role lease counts are summed across every sessions.items[] entry while total_kit_vram_mb is one host-wide sample, so a host serving 2+ sessions mixed multiple primaries/spectators against a single VRAM number. Takes the reviewer's reject option: the aggregate counts are kept (they are real observations) but session_scope='multi_session_aggregate' is published and the 1-primary+k-spectator watermark interpretation is marked measured=false with reason 'non-isolated multi-session snapshot; per-session VRAM attribution unavailable in this slice'. Exactly one active session yields session_scope='single_session' and keeps current semantics. 3. PRRT_kwDOSPoer86YeUly — -FixturePath and -OutputPath resolving to the same file made Set-Content truncate the fixture with the report, destroying the very artifact the report fingerprints. Canonicalises both ([System.IO.Path]::GetFullPath, case-insensitive only on Windows) and throws before any write. PRRT_kwDOSPoer86YeUlo (P1, bootstrap mechanism-path regression) is NOT addressed here and is moot as of 286bbac on this branch: per the #520 ruling the harness is not a mechanism surface, and the three measure-session-baseline classifier patterns the thread asked the test to pin were removed. Adding them to $expectedMechanismPaths now would fail the suite; re-registering them would revert an owner ruling. Verified on Windows: test-measure-session-baseline.ps1 (all groups pass), test-self-referential-bootstrap.ps1 (all assertions pass), Invoke-ScriptAnalyzer -Severity Error on the changed .ps1 files = 0. * fix(scripts): gpu-baseline r6 — reject malformed runtime counts, require observed primary, finite TTFF (review r6) - Get-SessionVramWatermark / Get-EnvironmentFingerprint: a non-null-but- unparseable observed_active_session_count / observed_kit_instance_binding_count (coordinator version skew, e.g. a string) was silently coerced to 0 via a try/catch default, relabeling an UNKNOWN runtime state as an OBSERVED zero. New ConvertTo-NonNegativeIntOrNull helper returns null instead of 0 on parse failure; both call sites now treat that null the same as a missing probe (new 'malformed_runtime_observation' session scope; runtime_state_unknown fixture-binding scope). - Get-SessionVramWatermark: exactly one active session was enough to accept the "1 primary + k spectator" watermark interpretation even when zero primary viewers had joined (idle-but-created or spectator-only session). Now requires observed_primary_lease_count == 1. - New-OptionalMeasurement TTFF validator only checked ">= 0", which +Infinity satisfies; now also rejects non-finite values. - openspec/lifecycle-ledger.json: added scripts/lib/measure-session-baseline.ps1 to the change's evidence_refs (the entire measurement implementation lives there; only the CLI wrapper and test were previously listed). Addresses the three still-open findings from the chatgpt-codex-connector review on c759057, plus the infinite-TTFF gap from an earlier round that was never landed. Co-authored-by: monkey1sai <26239865+monkey1sai@users.noreply.github.com> * fix(scripts): gpu-baseline r7 — keep kit process-count fields in every report shape (review r7) gpu-session-baseline-report/v1 的兩個早退路徑補齊 kit_process_count 與 kit_process_vram_unreadable_count:查詢失敗=null(未知非零)、查詢成功但無 Kit process=0(觀測到的真零),consumer 不再因 host 狀態拿到不同 shape。 測試補四條斷言鎖住兩態。 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: monkey1sai <xshiujj@gmail.com> Co-authored-by: claude[bot] <41898282+claude[bot]@users.noreply.github.com>
Summary
migrate-console-to-hifi-design收官推進(32/40):npm install後npm run verify(typecheck+build+vitest+struct-log)全綠——78 test files/1069 tests+23/23 struct-log,exit 0;npm 11 的 lockfile"peer": true噪音已 revert(本機 npm 11.6.2 vs pinned 10.9.4,無真依賴變更)。artifacts/2026-08-12-hifi-consumer-spec-scenario-audit.md逐條走完兩份 consumer spec 全部 66 個 Scenario(edge-console-operator-frontend 30+unified-governance-console 36),附 file:line 證據:58 HOLDS/5 HOLDS-WITH-NOTE/0 STALE/3 UNVERIFIABLE(皆為需 live deploy stack 的 browser-E2E/GPU 情境,且不在 feat(console): UnifiedConsole 遷移至 Hi-Fi design token(1/2 product code) #357/chore(design-gate): migrate-console-to-hifi-design manifest rebaseline(2/2) #358/test(design): reapprove canonical A4 baseline #429 的實際 diff 足跡內)。依 design.md Risk 條款(不確定 SHALL 停下不假設),checkbox 不勾、留待裁決。--rebaseline --confirm-rebaseline非 no-op——capture 腳本沒有處理baseline_provenance.authority=canonical_product_surface,會把 test(design): reapprove canonical A4 baseline #429 人工核可的 workspace.a4 產品面 golden 靜默回退成 mockup-origin bytes。已即時git checkout --還原(-VerifyOrigin仍 passed,13 screens/26 golden),缺陷另開 issue 追蹤,本 PR 不碰該腳本。docs/frontend/frontend-design-guidelines.md:14仍稱--ec-*為 production 權威(stale),另行處理。AI Coding Governance
Machine values:
Change lane=F/B/G/S;Behavior contract changed=yes/no;Requirement source=issue/docs/plans/superpowers spec/existing contract/not applicable.Frontend Verification
User-facing changes must pass two independent producers: real frontend/runtime operability evidence and the pinned
docs/plans/design-system-reference.manifest.jsonfidelity gate. Scope is derived from changed paths plus the base/head manifest union; the PR body cannot select an easier screen.mixedandpartial_reference_missingpermit honest partial work but requireFull completion claimed = no. Semantic evidence is produced only by thedesign-semantic-visualCI Playwright job, never supplied as PR input;PR Metadata Contractvalidates the live PR metadata, while normal protected CI checks determine mergeability.Deploy Path Verification
Required for runtime / Docker / Kit / viewer / ports / env / conversion-service changes.
Self-Referential Bootstrap
Required when the PR changes the verification mechanism itself (deploy path / evidence harness / gate script). Rule:
docs/agents/self-referential-bootstrap.md. Open ledger debt inscripts/self-referential-bootstrap-ledger.jsonblocks further mechanism PRs until fixpoint closure.Validation
npm run verify(web-viewer-sample)→ 全綠:78 files/1069 tests+23/23 struct-log,exit 0。pwsh scripts/tests/verify-design-system-reference.ps1 -VerifyOrigin→ passed(13 screens/26 golden;7.1 誤回退已還原後驗證)。npx openspec validate migrate-console-to-hifi-design --strict→ valid。node --test scripts/tests/test-openspec-machine-truth.mjs24/0;test-ai-coding-metrics.mjs13/0;pwsh scripts/tests/test-agent-governance-check.ps145/0。Known Risks
canonical_product_surface缺陷未修(本 PR 只揭露+另開 issue);在修復前任何人跑--rebaseline都會誤回退 A4 golden。🤖 Generated with Claude Code