docs(governance): self-referential-bootstrap 契約增補 §2.1 範圍界定與 scope 反例(issue #520) - #521
Conversation
- 機制面判準單一謂詞化:行為改變會改變其他 PR 裁決或 canonical deployment 驗證 - evidence harness 僅限輸出被 gate/deploy verification 機器消費者 - 明確排除產品量測/遙測腳本;新增 promotion rule(接線 PR 同步補登清單) - §5 記錄 PR #511 measure-session-baseline scope 反例裁決 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
You have reached your Codex usage limits for security reviews. Please try again later. |
📝 WalkthroughWalkthroughThe changes define the mechanism-surface scope for self-referential bootstrap governance and add evidence records for PR 521. The records document gate-suite results, stale-checkout remediation, verification references, and an open post-merge fixpoint. ChangesBootstrap scope and evidence
Estimated code review effort: 2 (Simple) | ~10 minutes Possibly related issues
Possibly related PRs
Suggested reviewers: 🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches 💡 1🛠️ Fix failing CI checks 💡
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
…ification (PR #521) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
There was a problem hiding this comment.
Pull request overview
This PR is a documentation-only change to the governance contract docs/agents/self-referential-bootstrap.md, implementing the adjudication recorded in issue #520. It narrows (does not expand) the scope of what counts as a "mechanism surface" for the self-referential bootstrap gate, resolving an ambiguity surfaced during PR #511's review where scripts/measure-session-baseline.ps1 was argued to belong on the mechanism list. The contract clarification aligns the prose triggers with the actual enumeration in Get-SelfReferentialMechanismPaths (scripts/lib/self-referential-bootstrap.ps1), whose machine pattern list is intentionally left unchanged.
Changes:
- Adds §2.1 defining a single predicate for mechanism-surface inclusion (a path whose behavior change alters another PR's adjudication or canonical deployment verification), narrows "evidence harness" to machine-consumed outputs, explicitly excludes product measurement/telemetry scripts, and adds a wiring-PR upgrade rule.
- Adds a §5 scope counterexample citing PR #511 / issue #520, ruling
measure-session-baseline.ps1out of the mechanism surface (no gate machine-consumer of its report). - The document itself is an adjudicator surface (listed in
SelfReferentialAdjudicatorPaths), so the PR self-declaresbootstrap=yeswith an open ledger entry, deferring the fixpoint attestation to a post-merge closure PR.
💡 Add a code-review agent skill for context-aware, tailored reviews. Learn more in the docs.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 92737674a3
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
…ce(gate 四套件實錄 @8b0efaf) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
monkey1sai-blip
left a comment
There was a problem hiding this comment.
Approved by monkey1sai-blip (the reviewer account pinned by the repo's merge governance).
Submitted through scripts/blip_review.py — a scripted approval carrying the operator's authority, pinned to head 302fde3e0ee68b00aa14e79b2a5ea1f369b4bad4. This is the mechanism the GitHub App cannot satisfy: an App's approving review does not count toward required_approving_review_count.
There was a problem hiding this comment.
Actionable comments posted: 3
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@docs/agents/self-referential-bootstrap.md`:
- Line 45: Make the mechanism-surface rules executable around
Get-SelfReferentialMechanismPaths by adding regression checks that preserve the
current exclusion and promote a producer when a machine consumer is introduced
through an existing generic gate. Mark the reference to
scripts/measure-session-baseline.ps1 as historical evidence, since it is
documentation/OpenSpec-only and not implemented.
In
`@docs/evidence/evidence-harness-scope-clarification/self-referential-bootstrap/gate-suites.txt`:
- Around line 6-27: Update gate-suites.txt entries to record the exact
verification invocation used for each suite, matching the workflows and manifest
with pwsh -NoProfile -NonInteractive -File; if any suite used the current manual
variant, explicitly label that evidence as such.
In
`@docs/evidence/evidence-harness-scope-clarification/self-referential-bootstrap/README.md`:
- Around line 1-3: Add a required document-nature declaration near the title in
the README, using one allowed value such as “working note” alongside the
existing bootstrap evidence heading and stack_kind metadata.
🪄 Autofix
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: CHILL
Plan: Pro Plus
Run ID: 20c5e777-dfaf-42a1-b2ab-95f243c2ae78
📒 Files selected for processing (4)
docs/agents/self-referential-bootstrap.mddocs/evidence/evidence-harness-scope-clarification/self-referential-bootstrap/README.mddocs/evidence/evidence-harness-scope-clarification/self-referential-bootstrap/gate-suites.txtscripts/self-referential-bootstrap-ledger.json
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 302fde3e0e
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
…nd hifi row to #507 squash 一、撤除 bootstrap 層(依 #520 裁決=docs/agents/self-referential-bootstrap.md §2.1,PR #521): 量測 harness 的報告無任何 gate 機器消費者,不屬 mechanism surface,不入 ledger。 classifier 擴張+open entry gpu-session-baseline-harness+evidence 一併撤除, scripts/lib/self-referential-bootstrap.ps1、scripts/tests/test-self-referential-bootstrap.ps1、 scripts/self-referential-bootstrap-ledger.json 還原為 origin/main 版本。 機械上這條路也是死路:base-pinned 裁決者以 base 版 classifier 驗證新 entry 宣告的 mechanism paths,同 PR 擴張 classifier 永遠無法讓自己的 entry 合法(實測兩輪 pr-metadata-contract-diagnostic 均以 not classified verification-mechanism paths 拒絕)。 二、rebind migrate-console-to-hifi-design row:#507 squash 後該 row 仍綁 pre-squash commit af60c29(已被丟棄,CI checkout 抓不到)→ 全部後續 PR 的 machine-truth test 25 紅。依 #482/#501/#512 慣例 rebind 到 landed squash 4187102。 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… threads) - README 補 Document nature 標籤(CodeRabbit:docs/**/*.md 須宣告文件性質) - §5 反例標明 harness 由 PR #511 引入(本樹尚無該檔,屬裁決紀錄)並補述 升級規則維持 review 強制的互鎖理由(機器化=再改 adjudicator=另開 debt) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…baseline 1.1) (#511) * feat(scripts): add GPU session baseline measurement harness (gpu-session-baseline-and-idle-reclaim 1.1) Add scripts/measure-session-baseline.ps1 (root CLI) and scripts/lib/measure-session-baseline.ps1 (testable core) implementing task 1.1 of openspec/changes/gpu-session-baseline-and-idle-reclaim: a read-only harness that captures nvidia-smi GPU inventory (VRAM/utilization, consumer RTX classification, MIG availability), a GET-only WebRTC/coordinator health probe (/health, /api/runtime/status), and the environment fingerprint required by the gpu-session-baseline spec (GPU model, driver version, Kit version from kit-sdk.packman.xml, fixture hash+size). The harness never opens a WebRTC session or creates/joins/closes a review session, so TTFF and session-creation success rate cannot be honestly measured locally; those fields are null with measured:false and an explicit reason unless supplied by a caller (e.g. a future task 1.3 soak run), never fabricated. Every other unmeasurable signal (no nvidia-smi, no GPU rows, insufficient OS permission on the compute-apps VRAM column, coordinator unreachable) degrades the same way instead of throwing or guessing. Registered in scripts/script-registry.json as a measurement-harness (not deploy.ps1/verify-all.ps1: it measures, it does not deploy or gate; not scripts/lib alone: it is the operator-invoked CLI entry; not scripts/tests: it produces a JSON report, not a pass/fail check). Add scripts/tests/test-measure-session-baseline.ps1: unit tests for GPU line parsing/consumer-RTX/MIG classification, fail-safe behavior with nvidia-smi entirely absent, report schema shape, script-registry.json consistency, and a real CLI smoke test (both -OutputPath and the default artifacts/gpu-baseline/<timestamp>.json path). Verified on pwsh 7.5.4 and Windows PowerShell 5.1 (powershell.exe), invoke-powershell-static.ps1, and scripts/tests/test-agent-governance-check.ps1 (45/45 green). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore(openspec): tick gpu-session-baseline-and-idle-reclaim task 1.1 Mark tasks.md 1.1 done (measure-session-baseline.ps1 harness landed) and sync openspec/lifecycle-ledger.json: task_ledger completed 0->1, current_slice points at the 1.2 env-fingerprint gate as the next slice, last_verified refreshed, subject_commit rebound to 405e2b6 (the commit that landed the harness + tests + registry entry), evidence_refs extended to the new script and test paths. Verified: node scripts/tests/verify-openspec-repository-lifecycle.mjs --repo-root . (openspec/changes, lifecycle-ledger.json and docs/plans/NOW.md agree -- NOW.md's projection is id+status only, and status stays "active", so it needed no edit); node --test scripts/tests/test-openspec-machine-truth.mjs (24/24) and scripts/tests/test-ai-coding-metrics.mjs (13/13); pwsh scripts/tests/test-agent-governance-check.ps1 (45/45 embedded repository-lifecycle subtests green). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(scripts): honest gpu-baseline measurements — real binding enum, per-role lease counts, validated inputs, untick 1.1 (review) * fix(scripts): declare gpu fingerprint scope — multi-GPU host is first_gpu_only, per-GPU fingerprint deferred to 1.2 (review) * fix(scripts): harden GPU baseline harness per PR #511 review Addresses the valid findings from the 13 unresolved review threads on scripts/measure-session-baseline.ps1 (gpu-session-baseline-and-idle-reclaim task 1.1/1.2), keeping the "measured:false, never fabricate" contract intact: - Get-SafeProperty: fix a PowerShell pipeline-unroll bug where `return $value` on an empty array collapsed to $null, making an observed empty kit_instance_bindings/viewer_leases indistinguishable from "unmeasured". - Get-WebRtcHealthProbe: count non-terminal KitInstanceBinding statuses (allocated/starting/ready/draining) instead of a literal status='active' that the real coordinator API never emits; count active primary/spectator viewer_leases by role (the actual 1-primary-plus-k-spectator cardinality) instead of sessions.active_count; clarify that `reachable` reflects only coordinator /health liveness, not independent WebRTC/signaling reachability. - Get-SessionVramWatermark: only claim a clean measured total when exactly one Kit GPU process is observed and fully readable; multi-process or partially-readable readouts are surfaced only as the informational unscoped_total_kit_vram_mb, never as a fabricated measured:true total. - Get-EnvironmentFingerprint: fail closed (measured:false) on a multi-GPU host instead of blindly binding the fingerprint to gpus[0]. - Get-KitVersionFingerprint: carry an explicit source/caveat noting this is the checkout's declared dependency version, not a live-process read. - Get-SessionBaselineReport: range-validate caller-supplied -TtffMs / -SessionCreationSuccessRate (reject negative/out-of-range instead of recording as measured); resolve host.hostname via the cross-platform Dns API with HOSTNAME/COMPUTERNAME fallback so Linux deployment targets don't silently null out host identity. - Root wrapper: derive the default -OutputPath from the report's own collision-resistant run_id instead of a bare second-resolution timestamp. - CI: run test-measure-session-baseline.ps1 (PS7 + Windows PowerShell 5.1) as part of the required `powershell-static` job, mirroring the existing test-spec-to-done-port-helper.ps1 pattern -- neither `root-contracts` (pytest) nor the PSScriptAnalyzer-only `powershell-static` gate command previously executed this suite. The lifecycle-ledger subject-ancestry finding (PRRT_kwDOSPoer86YcZ4d) was independently verified as already resolved at current HEAD and needed no change; see PR reply for evidence. Verified: pwsh + Windows PowerShell 5.1 test-measure-session-baseline.ps1, invoke-powershell-static.ps1, and node --test test-openspec-machine-truth.mjs (24/24) all green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(scripts): carry a live-build staleness caveat on kit_version (PR #511 review) Get-KitVersionFingerprint reads bim-streaming-server/tools/deps/kit-sdk.packman.xml -- the checkout's DECLARED kit-kernel dependency version -- not a value read from the live Kit process. If the checkout is updated without a rebuild/restart, this can be stale relative to the session actually being measured, and no local mechanism exists to introspect a running Kit.exe's build identity to close that gap. Surface a `source` ('checkout_packman_declared') and an explicit `caveat` string alongside the existing value/measured/reason shape (both on the raw fingerprint and propagated through Get-EnvironmentFingerprint's kit_version field) so downstream SLO-writers know what this field does and does not attest to, rather than silently trusting checkout state as if it were live-process state. This was the one review thread not already covered by the concurrent fixes landed in 37e3247/2eee19b on this branch (real KitInstanceBinding status enum, per-role viewer lease counts, VRAM attribution transparency, TTFF/success-rate validation, hostname fallback, GPU fingerprint scope disclosure, collision-resistant default filename, task 1.1 unticked); this commit reconciles with that work rather than duplicating it. Verified: pwsh + Windows PowerShell 5.1 test-measure-session-baseline.ps1, invoke-powershell-static.ps1, and node --test test-openspec-machine-truth.mjs (24/24) all green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(scripts): fold GPU-attribution gap into complete, fix MIG-gating and test cleanup (PR #511 review round 2) Three round-2 findings on the GPU baseline harness: - Get-EnvironmentFingerprint: gpu_fingerprint_scope='first_gpu_only' (a multi-GPU host) was disclosure-only -- `complete` was computed from the five base fields before the scope was known, so a report that has admittedly NOT attributed every relevant GPU could still report complete:true and silently suppress the wrapper's "SHALL NOT be used to set SLOs or admission parameters" warning. `complete` now also requires gpu_fingerprint_scope != 'first_gpu_only'. - Get-GpuInventorySnapshot: software_queue_required was gated on consumer_grade_all AND NOT mig_available_any. On a non-consumer, non-MIG fleet (e.g. a lone RTX A6000, which Test-ConsumerRtxGpuName excludes but which does not support MIG at all), that reported software_queue_required=false with no MIG route in fact available. Software queuing is now required whenever MIG is unavailable, regardless of consumer/professional classification. - test-measure-session-baseline.ps1: the default-OutputPath cleanup deleted every new file under artifacts/gpu-baseline/, not just the one this test produced -- a concurrent harness invocation sharing the checkout would have its evidence collaterally deleted. Now deletes only "$($defaultReport.run_id).json". Also strengthened Assert-ReportSchemaShape's `complete` expectation and added regression tests for the professional/no-MIG inventory shape and the multi-GPU complete=false path. Verified: pwsh + Windows PowerShell 5.1 test-measure-session-baseline.ps1, invoke-powershell-static.ps1, and node --test test-openspec-machine-truth.mjs (24/24) all green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * revert(ci): defer harness CI wiring until the open mechanism debt closes ci.yml is a classified verification-mechanism path (design gate infrastructure, Lane G minimum + self-referential bootstrap scope); wiring test-measure-session-baseline.ps1 into CI from this measurement-harness PR would collide with the open mechanism-hardening-2 ledger entry owned by PR #513. The wiring moves to a follow-up alongside issue #516 (CI coverage for the streaming pytest suite) after fixpoint closure. Refs #516 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(scripts): gpu-baseline r2 - bootstrap registration, honest kit-version provenance, MIG-driven queue flag (review r2) * fix(scripts): honest fixture provenance binding for baseline fingerprint (review r4, operator-delegated adjudication) * fix(governance): revert harness bootstrap-layer per #520 ruling; rebind hifi row to #507 squash 一、撤除 bootstrap 層(依 #520 裁決=docs/agents/self-referential-bootstrap.md §2.1,PR #521): 量測 harness 的報告無任何 gate 機器消費者,不屬 mechanism surface,不入 ledger。 classifier 擴張+open entry gpu-session-baseline-harness+evidence 一併撤除, scripts/lib/self-referential-bootstrap.ps1、scripts/tests/test-self-referential-bootstrap.ps1、 scripts/self-referential-bootstrap-ledger.json 還原為 origin/main 版本。 機械上這條路也是死路:base-pinned 裁決者以 base 版 classifier 驗證新 entry 宣告的 mechanism paths,同 PR 擴張 classifier 永遠無法讓自己的 entry 合法(實測兩輪 pr-metadata-contract-diagnostic 均以 not classified verification-mechanism paths 拒絕)。 二、rebind migrate-console-to-hifi-design row:#507 squash 後該 row 仍綁 pre-squash commit af60c29(已被丟棄,CI checkout 抓不到)→ 全部後續 PR 的 machine-truth test 25 紅。依 #482/#501/#512 慣例 rebind 到 landed squash 4187102。 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(scripts): gpu-baseline r5 — unknown-runtime honesty, session-scoped lease counts, fixture overwrite guard (review r5) PR #511 review r5, three of four threads: 1. PRRT_kwDOSPoer86YeUls — an unreachable/malformed /api/runtime/status left every observed count null, and the fixture-binding defaults coerced those nulls to 0, so an UNKNOWN runtime state was published under the 'no_live_session_observed' (observed-idle) label with complete=true beside a declared fixture. Adds a distinct fixture_binding_scope='runtime_state_unknown' that withdraws completeness and names the failed probe (GET /api/runtime/status). A positive observation still outranks the unknown; an observed 0/0 idle host keeps its previous semantics. 2. PRRT_kwDOSPoer86YeUlw — per-role lease counts are summed across every sessions.items[] entry while total_kit_vram_mb is one host-wide sample, so a host serving 2+ sessions mixed multiple primaries/spectators against a single VRAM number. Takes the reviewer's reject option: the aggregate counts are kept (they are real observations) but session_scope='multi_session_aggregate' is published and the 1-primary+k-spectator watermark interpretation is marked measured=false with reason 'non-isolated multi-session snapshot; per-session VRAM attribution unavailable in this slice'. Exactly one active session yields session_scope='single_session' and keeps current semantics. 3. PRRT_kwDOSPoer86YeUly — -FixturePath and -OutputPath resolving to the same file made Set-Content truncate the fixture with the report, destroying the very artifact the report fingerprints. Canonicalises both ([System.IO.Path]::GetFullPath, case-insensitive only on Windows) and throws before any write. PRRT_kwDOSPoer86YeUlo (P1, bootstrap mechanism-path regression) is NOT addressed here and is moot as of 286bbac on this branch: per the #520 ruling the harness is not a mechanism surface, and the three measure-session-baseline classifier patterns the thread asked the test to pin were removed. Adding them to $expectedMechanismPaths now would fail the suite; re-registering them would revert an owner ruling. Verified on Windows: test-measure-session-baseline.ps1 (all groups pass), test-self-referential-bootstrap.ps1 (all assertions pass), Invoke-ScriptAnalyzer -Severity Error on the changed .ps1 files = 0. * fix(scripts): gpu-baseline r6 — reject malformed runtime counts, require observed primary, finite TTFF (review r6) - Get-SessionVramWatermark / Get-EnvironmentFingerprint: a non-null-but- unparseable observed_active_session_count / observed_kit_instance_binding_count (coordinator version skew, e.g. a string) was silently coerced to 0 via a try/catch default, relabeling an UNKNOWN runtime state as an OBSERVED zero. New ConvertTo-NonNegativeIntOrNull helper returns null instead of 0 on parse failure; both call sites now treat that null the same as a missing probe (new 'malformed_runtime_observation' session scope; runtime_state_unknown fixture-binding scope). - Get-SessionVramWatermark: exactly one active session was enough to accept the "1 primary + k spectator" watermark interpretation even when zero primary viewers had joined (idle-but-created or spectator-only session). Now requires observed_primary_lease_count == 1. - New-OptionalMeasurement TTFF validator only checked ">= 0", which +Infinity satisfies; now also rejects non-finite values. - openspec/lifecycle-ledger.json: added scripts/lib/measure-session-baseline.ps1 to the change's evidence_refs (the entire measurement implementation lives there; only the CLI wrapper and test were previously listed). Addresses the three still-open findings from the chatgpt-codex-connector review on c759057, plus the infinite-TTFF gap from an earlier round that was never landed. Co-authored-by: monkey1sai <26239865+monkey1sai@users.noreply.github.com> * fix(scripts): gpu-baseline r7 — keep kit process-count fields in every report shape (review r7) gpu-session-baseline-report/v1 的兩個早退路徑補齊 kit_process_count 與 kit_process_vram_unreadable_count:查詢失敗=null(未知非零)、查詢成功但無 Kit process=0(觀測到的真零),consumer 不再因 host 狀態拿到不同 shape。 測試補四條斷言鎖住兩態。 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: monkey1sai <xshiujj@gmail.com> Co-authored-by: claude[bot] <41898282+claude[bot]@users.noreply.github.com>
…evidence-harness-scope # Conflicts: # scripts/self-referential-bootstrap-ledger.json
…式;review threads) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…evidence-harness-scope
…l debt gate The first fixpoint run of a mechanism is also the first time that mechanism executes as the canonical path, so it can surface a regression in itself. The contract left no legal way to fix one: naming the existing open entry hit the impersonation guard, self-registering a second entry hit the other-open-debt gate, and declaring bootstrap=no hit that same gate — three-way interlock (issue #494). Ledger entries gain an OPTIONAL append-only `repair_prs` array (strictly increasing positive integers; absent reads as empty, so the four immutable closed entries stay byte-identical). `Assert-SelfReferentialJsonObjectShape` grows `-OptionalProperties` so the exact-property-set rule can admit a new field without requiring it of entries written before it existed. A repair PR is admitted only when all five hold: 1. the named entry is pre-existing debt, open at base AND still open at head; 2. the transition's only change to it is a repair_prs tail append whose appended value is exactly this PR number (no live PR number => refused); 3. every mechanism path this PR changes is already inside that entry's declared verification_mechanism_paths (case-sensitive); 4. the PR touches none of this gate's own adjudicators; 5. the transition neither opens nor closes any entry. Invariants preserved: the ledger stays append-only (repair_prs elements are never rewritten or dropped); every other entry field stays immutable and the only status transition is still one open -> closed; closed entries remain fully immutable — repair_prs is not a back door into them; the closure still demands a complete, all-green fixpoint attestation against the frozen verification contract, so a repaired mechanism must prove itself; and a single transition still may not both settle and incur debt. Refs #494 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 52488e8189
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
…-lane head Folding issue #494's repair lane into this PR grew the entry's declared surface from two paths to four (the gate library and its suite joined the contract prose and the ledger). The bootstrap evidence still recorded the pre-fold head 82c6aa4 and described the fixpoint as having 'one thing to prove', which understated what the entry now covers. Reran the four contract suites at 52488e8 (all exit 0, timings recorded) and corrected the README's scope paragraph to state both obligations. Refs #494 #520 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
併入通知(coordinator,依 owner 指示):issue #494 的 regression-repair lane 已折進本 PR(commit 已同步處理的連動:
repair lane 的五個放行條件(全滿足才放行,否則逐條具名拒絕):
不變式全保:ledger append-only、entry 唯一 未更動:§2.1 條文本身、evidence 的既有結論、machine pattern 清單。若您(#521 原作者 session)對範圍擴張有異議,請回覆,我可把 repair lane 拆回獨立 PR(代價=多一輪 debt 與 fixpoint)。 |
There was a problem hiding this comment.
Codex Tri-Adversarial Bot
Automated tri-adversarial ship-gate (L0 terra triage / L1 tier-routed lens fanout / L2 refute-by-default / L3 sol apex — Codex models).
Mapped event: COMMENT
Codex Tri-Adversarial ship-gate — PR #521
- Repo head:
governance/bootstrap-evidence-harness-scope@2af2697 - Base:
main@c5d423c - Files changed: 6
- Engine: four-model tri-adversarial gate on Codex — L0 triage
gpt-5.6-terra/low; L1 lens finders routedgpt-5.6-terra/low →gpt-5.6-luna/medium →gpt-5.5/xhigh (security floorgpt-5.5); L2 refute-by-defaultgpt-5.5/xhigh, top-tier findings refuted bygpt-5.6-sol/xhigh (every refutation cross-model); L3 apexgpt-5.6-sol/max. 誠實聲明:層級與 Claude 三層 gate 同構(terra≈haiku、luna≈sonnet、gpt-5.5≈opus、sol≈fable),但模型池是 Codex 的,非 Anthropic 的。
Verdict
HELD — 三層驗證未能完成,本次不投同意票(fail-closed)。
- held reason:
apex_unavailable_or_failed - mapped GitHub event:
COMMENT
Difficulty & routing
- overall:
critical(source: terra-triage) - lens tiers: correctness→
gpt-5.5, security→gpt-5.5, simplification→gpt-5.5, test-gap→gpt-5.5
Layer stats
- L1: raw=5 deduped=5 finder_failures=0
- L2: confirmed=0 refuted=5 unverified=0
- L3 final: 0
Agent calls
- 10/12 ok, engine wall-clock 233.9s
VERDICT
HELD
VERDICT: HELD
…uire real repair work Two connector threads, both verified before acting. (A) Condition 4 banned adjudicator edits outright, on the theory that a PR repairing the rule that judges it could wave itself through. Measured against the workflow, that theory does not hold: pr-review-agent.yml checks out pull_request.base.sha (:25), materializes the gate from BASE via git archive (:82-86), and exits non-zero with base_gate_incomplete_external_approval_required (:90-94) rather than ever falling back to head; :113 always resolves the checker under GATE_ROOT, which is only set on the base-pinned branch. The only other invocation is scripts/dev/check-pr-local-preflight.ps1, a developer preflight with no merge authority. test-base-gate-capability.ps1 is the executable form of the invariant and passes. The ban also deadlocked the debt it protected: this repository's own open entry declares two adjudicator paths, so a failing fixpoint on it would have had no lane at all — issue #494 reproduced one level up. Adjudicator paths are now repairable, bounded by the declared surface. Safety rests on base-pinned adjudication (no self-clearance), declared-subset (no scope expansion), and the unchanged fixpoint obligation (the repaired mechanism must still prove itself against the entry's frozen verification contract). (B) A repair could append a repair_prs record while fixing nothing. The proposed test — "changed at least one mechanism path" — is vacuous: measured, the ledger is itself a classified mechanism path and a repair PR necessarily edits it to append repair_prs, so the count is >= 1 by construction, and a PR with zero mechanism paths returns from the body gate before the lane is reached. The check that actually bites requires a NON-LEDGER mechanism path — the exact mirror of the closure rule, which permits the ledger and nothing else. Demonstrated: with the check disabled, a ledger-only repair was admitted. Conditions stay five and stay aligned with §2.2: (3) is now real repair work, (4) is the declared-surface bound that carries the adjudicator carve-out. The condition-5 closure check is retained as defence in depth and labelled as such — with (3) in place, a repair that also closes debt is refused earlier on both reachable paths, and both are now pinned by tests. Refs #494 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The four contract suites were re-run on 7f590ea (declared adjudicator paths are repairable; a repair must change a non-ledger mechanism path) with a clean tree; tested_head, run_at and every duration record that run rather than the superseded aa32597 one. This commit changes only this evidence file, so no untested code tree is being claimed as passing. Refs #494 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
There was a problem hiding this comment.
Codex Tri-Adversarial Bot
Automated tri-adversarial ship-gate (L0 terra triage / L1 tier-routed lens fanout / L2 refute-by-default / L3 sol apex — Codex models).
Mapped event: COMMENT
Codex Tri-Adversarial ship-gate — PR #521
- Repo head:
governance/bootstrap-evidence-harness-scope@1f298a4 - Base:
main@c5d423c - Files changed: 6
- Engine: four-model tri-adversarial gate on Codex — L0 triage
gpt-5.6-terra/low; L1 lens finders routedgpt-5.6-terra/low →gpt-5.6-luna/medium →gpt-5.5/xhigh (security floorgpt-5.5); L2 refute-by-defaultgpt-5.5/xhigh, top-tier findings refuted bygpt-5.6-sol/xhigh (every refutation cross-model); L3 apexgpt-5.6-sol/max. 誠實聲明:層級與 Claude 三層 gate 同構(terra≈haiku、luna≈sonnet、gpt-5.5≈opus、sol≈fable),但模型池是 Codex 的,非 Anthropic 的。
Verdict
HELD — 三層驗證未能完成,本次不投同意票(fail-closed)。
- held reason:
all_finders_failed - mapped GitHub event:
COMMENT
Difficulty & routing
- overall:
high(source: fail-safe-default) - lens tiers: correctness→
gpt-5.6-luna, security→gpt-5.5, simplification→gpt-5.6-luna, test-gap→gpt-5.6-luna
Agent calls
- 0/10 ok, engine wall-clock 66.9s
VERDICT
HELD
VERDICT: HELD
|
Claude finished @monkey1sai's task in 2m 57s —— View job Independent final-head review — head
|
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 1f298a45ea
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
Claude finished @monkey1sai's task in 3m 21s —— View job Independent final-head review — head
|
|
Codex Review: Didn't find any major issues. Nice work! Reviewed commit: ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |
|
@codex review @claude review Exact head: The PR body now matches the linked-successor implementation. Four exact-code-head governance gates, local PR preflight, hosted CI, hosted Agent Governance, hosted PR Metadata Contract, and Windows deploy dry-run passed. Sol/Terra adversarial review found no remaining blocker. Please review only this exact head. |
|
Claude finished @monkey1sai's task in 2m 1s —— View job Independent final-head review — head
|
| Check | Result |
|---|---|
| CI | ✅ success |
| Agent Governance | ✅ success |
| PR Metadata Contract | ✅ success (one earlier transient failure at 05:31:29Z, succeeded on the 05:35:27Z re-run) |
Governance Base Audit (governance-trust-root.yml) |
❌ failing, twice, at this head |
The Governance Base Audit failure is not a defect in this PR's diff — scripts/dev/check_governance_trust_root.py (unchanged by this PR) requires a fresh exact-head approval from monkey1sai-blip (_has_exact_head_approval, line 515), and the PR body itself already discloses this as outstanding ("monkey1sai-blip exact-head approval pending"). The last recorded monkey1sai-blip approval was pinned to an earlier head (302fde3), not bc5a06b, so this check is expected to stay red until that approval is re-submitted at the current head — it's a gating state, not a code correctness issue. Flagging it because the trigger comment's claim list didn't mention this workflow and I wanted to verify status independently rather than take the summary at face value.
Verdict
No new correctness issues at bc5a06b — this head is unchanged from the one I already reviewed in depth. The remaining blocker to merge is procedural: monkey1sai-blip's exact-head approval, which the PR body already tracks as pending and which Governance Base Audit is correctly fail-closing on.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: bc5a06bc74
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
Claude finished @monkey1sai's task in 3m 33s —— View job Independent final-head review — head
|
|
Codex Review: Didn't find any major issues. More of your lovely PRs please. Reviewed commit: ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |
Summary
repair_prs, current-PR binding, and non-ledger repair-work requirements.repair_prs, new debt attaches to the unique open leaf, and new path ownership is computed from the whole active-chain surface union.A -> B -> Cto close only one oldest open root at a time, leaving the exact contiguous successor suffix open.Verification
Exact code commit
e5e2c9829b49dae5b9964f553fea73788977c345:pwsh -NoProfile -NonInteractive -File scripts/tests/test-self-referential-bootstrap.ps1— exit 0, 37spwsh -NoProfile -NonInteractive -File scripts/tests/test-pr-body-evidence.ps1— exit 0, 9spwsh -NoProfile -NonInteractive -File scripts/tests/test-agent-governance-check.ps1— exit 0, 116spwsh -NoProfile -NonInteractive -File scripts/tests/invoke-powershell-static.ps1— exit 0, 2sEvidence-only commit
e670d116c6c868094d6cbbea63147afc79448406records those exact-code results and bindstested_headto the code commit. Sol/Terra/Luna cross-review found and closed the recursive-chain deadlock, discontinuous-status/multi-close integrity bypass, and emptyrepair_prsparser gap. Final frozen-diff Sol and Terra reviews report no P0/P1/P2 blocker.Change Classification
docs/agents/self-referential-bootstrap.md; issues #520 and #494; exact-head review discussions#discussion_r3772464308and#discussion_r3772731806AI Coding Governance
docs/agents/self-referential-bootstrap.mdmonkey1sai-blipexact-head approval pendingWindows On-Demand Verification
e670d116c6c868094d6cbbea63147afc79448406;pwsh -NoProfile -NonInteractive -File scripts/deploy.ps1 -DryRunexited 0 in 8s, performed no Phase 2 actions, and reported only local would-create/would-build plus example-env fallback diagnostics. Exact-head GitHub Actions run: https://github.com/monkey1sai/AI-BIM-governance/actions/runs/31698089451Self-referential bootstrap