Skip to content

docs(governance): self-referential-bootstrap 契約增補 §2.1 範圍界定與 scope 反例(issue #520) - #521

Merged
monkey1sai merged 22 commits into
mainfrom
governance/bootstrap-evidence-harness-scope
Aug 13, 2026
Merged

monkey1sai merged 22 commits into
mainfrom
governance/bootstrap-evidence-harness-scope

Conversation

@monkey1sai

@monkey1sai monkey1sai commented Aug 12, 2026 •

Copy link
Copy Markdown
Owner

Summary

Verification

Exact code commit e5e2c9829b49dae5b9964f553fea73788977c345:

  • pwsh -NoProfile -NonInteractive -File scripts/tests/test-self-referential-bootstrap.ps1 — exit 0, 37s
  • pwsh -NoProfile -NonInteractive -File scripts/tests/test-pr-body-evidence.ps1 — exit 0, 9s
  • pwsh -NoProfile -NonInteractive -File scripts/tests/test-agent-governance-check.ps1 — exit 0, 116s
  • pwsh -NoProfile -NonInteractive -File scripts/tests/invoke-powershell-static.ps1 — exit 0, 2s

Evidence-only commit e670d116c6c868094d6cbbea63147afc79448406 records those exact-code results and binds tested_head to the code commit. Sol/Terra/Luna cross-review found and closed the recursive-chain deadlock, discontinuous-status/multi-close integrity bypass, and empty repair_prs parser gap. Final frozen-diff Sol and Terra reviews report no P0/P1/P2 blocker.

Change Classification

Item Result
Change lane G
Behavior contract changed yes
Requirement source existing contract docs/agents/self-referential-bootstrap.md; issues #520 and #494; exact-head review discussions #discussion_r3772464308 and #discussion_r3772731806

AI Coding Governance

Item Result
Linked issue #520; #494
Requirement source existing contract docs/agents/self-referential-bootstrap.md
CODEOWNERS / owner review required — monkey1sai-blip exact-head approval pending
GitNexus evidence UNKNOWN — reviewed CLI 1.6.9 ran impact/detect-changes, but its main-checkout index is 83 commits stale and incorrectly reported no changes; exact diff, adversarial tests, four exact-code gates, and Sol/Terra/Luna review are the compensating evidence
Browser E2E evidence not applicable — no frontend/runtime path changed
Agent workflow changed? yes — bootstrap repair, successor-chain, path ownership, parser integrity, and ordered closure contracts changed
Required checks expected CI, Agent Governance, PR Metadata Contract, and the base-pinned Governance Base Audit

Windows On-Demand Verification

Item Result
Windows verification tier deploy_dryrun
Windows verification evidence Exact PR head e670d116c6c868094d6cbbea63147afc79448406; pwsh -NoProfile -NonInteractive -File scripts/deploy.ps1 -DryRun exited 0 in 8s, performed no Phase 2 actions, and reported only local would-create/would-build plus example-env fallback diagnostics. Exact-head GitHub Actions run: https://github.com/monkey1sai/AI-BIM-governance/actions/runs/31698089451

Self-referential bootstrap

Item Result
Self-referential bootstrap yes
Bootstrap ledger entry evidence-harness-scope-clarification
Bootstrap reason This PR changes the base-pinned adjudicator contract and its executable gate. The new repair and linked-successor-chain behavior cannot adjudicate its own opening transition from the old main gate, so the existing open debt remains mandatory and must be closed only by a separate post-merge exact-main fixpoint with fresh attestations.

- 機制面判準單一謂詞化:行為改變會改變其他 PR 裁決或 canonical deployment 驗證
- evidence harness 僅限輸出被 gate/deploy verification 機器消費者
- 明確排除產品量測/遙測腳本;新增 promotion rule(接線 PR 同步補登清單)
- §5 記錄 PR #511 measure-session-baseline scope 反例裁決

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Copilot AI balanced review requested due to automatic review settings August 12, 2026 05:20
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for security reviews. Please try again later.

@coderabbitai

coderabbitai Bot commented Aug 12, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

📝 Walkthrough

Walkthrough

The changes define the mechanism-surface scope for self-referential bootstrap governance and add evidence records for PR 521. The records document gate-suite results, stale-checkout remediation, verification references, and an open post-merge fixpoint.

Changes

Bootstrap scope and evidence

Layer / File(s) Summary
Mechanism-surface scope definition
docs/agents/self-referential-bootstrap.md
Defines mechanism-surface inclusion criteria, exclusions for manually consumed scripts, escalation rules for newly machine-consumed scripts, and the GPU session baseline counterexample.
Bootstrap evidence and ledger records
docs/evidence/evidence-harness-scope-clarification/self-referential-bootstrap/README.md, docs/evidence/evidence-harness-scope-clarification/self-referential-bootstrap/gate-suites.txt, scripts/self-referential-bootstrap-ledger.json
Records the bootstrap evidence, four successful gate suites, stale-checkout remediation, verification references, and an open fixpoint for PR 521.

Estimated code review effort: 2 (Simple) | ~10 minutes

Possibly related issues

Possibly related PRs

Suggested reviewers: monkey1sai-blip

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title clearly identifies the governance contract update, scope clarification, and scope counterexample addressed by the pull request.
Description check ✅ Passed The description directly explains the scope changes, bootstrap ledger entry, repair-chain behavior, and verification evidence.
✨ Finishing Touches 💡 1
🛠️ Fix failing CI checks 💡
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch governance/bootstrap-evidence-harness-scope

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

…ification (PR #521)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR is a documentation-only change to the governance contract docs/agents/self-referential-bootstrap.md, implementing the adjudication recorded in issue #520. It narrows (does not expand) the scope of what counts as a "mechanism surface" for the self-referential bootstrap gate, resolving an ambiguity surfaced during PR #511's review where scripts/measure-session-baseline.ps1 was argued to belong on the mechanism list. The contract clarification aligns the prose triggers with the actual enumeration in Get-SelfReferentialMechanismPaths (scripts/lib/self-referential-bootstrap.ps1), whose machine pattern list is intentionally left unchanged.

Changes:

  • Adds §2.1 defining a single predicate for mechanism-surface inclusion (a path whose behavior change alters another PR's adjudication or canonical deployment verification), narrows "evidence harness" to machine-consumed outputs, explicitly excludes product measurement/telemetry scripts, and adds a wiring-PR upgrade rule.
  • Adds a §5 scope counterexample citing PR #511 / issue #520, ruling measure-session-baseline.ps1 out of the mechanism surface (no gate machine-consumer of its report).
  • The document itself is an adjudicator surface (listed in SelfReferentialAdjudicatorPaths), so the PR self-declares bootstrap=yes with an open ledger entry, deferring the fixpoint attestation to a post-merge closure PR.

💡 Add a code-review agent skill for context-aware, tailored reviews. Learn more in the docs.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 92737674a3

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread docs/agents/self-referential-bootstrap.md
Comment thread docs/agents/self-referential-bootstrap.md Outdated
Comment thread docs/agents/self-referential-bootstrap.md
…ce(gate 四套件實錄 @8b0efaf)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@monkey1sai
monkey1sai enabled auto-merge (squash) August 12, 2026 05:34

@monkey1sai-blip monkey1sai-blip left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved by monkey1sai-blip (the reviewer account pinned by the repo's merge governance).

Submitted through scripts/blip_review.py — a scripted approval carrying the operator's authority, pinned to head 302fde3e0ee68b00aa14e79b2a5ea1f369b4bad4. This is the mechanism the GitHub App cannot satisfy: an App's approving review does not count toward required_approving_review_count.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@docs/agents/self-referential-bootstrap.md`:
- Line 45: Make the mechanism-surface rules executable around
Get-SelfReferentialMechanismPaths by adding regression checks that preserve the
current exclusion and promote a producer when a machine consumer is introduced
through an existing generic gate. Mark the reference to
scripts/measure-session-baseline.ps1 as historical evidence, since it is
documentation/OpenSpec-only and not implemented.

In
`@docs/evidence/evidence-harness-scope-clarification/self-referential-bootstrap/gate-suites.txt`:
- Around line 6-27: Update gate-suites.txt entries to record the exact
verification invocation used for each suite, matching the workflows and manifest
with pwsh -NoProfile -NonInteractive -File; if any suite used the current manual
variant, explicitly label that evidence as such.

In
`@docs/evidence/evidence-harness-scope-clarification/self-referential-bootstrap/README.md`:
- Around line 1-3: Add a required document-nature declaration near the title in
the README, using one allowed value such as “working note” alongside the
existing bootstrap evidence heading and stack_kind metadata.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: 20c5e777-dfaf-42a1-b2ab-95f243c2ae78

📥 Commits

Reviewing files that changed from the base of the PR and between 20ec48e and de78087.

📒 Files selected for processing (4)
  • docs/agents/self-referential-bootstrap.md
  • docs/evidence/evidence-harness-scope-clarification/self-referential-bootstrap/README.md
  • docs/evidence/evidence-harness-scope-clarification/self-referential-bootstrap/gate-suites.txt
  • scripts/self-referential-bootstrap-ledger.json

Comment thread docs/agents/self-referential-bootstrap.md

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 302fde3e0e

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread scripts/self-referential-bootstrap-ledger.json
monkey1sai added a commit that referenced this pull request Aug 12, 2026
…nd hifi row to #507 squash

一、撤除 bootstrap 層(依 #520 裁決=docs/agents/self-referential-bootstrap.md §2.1,PR #521):
量測 harness 的報告無任何 gate 機器消費者,不屬 mechanism surface,不入 ledger。
classifier 擴張+open entry gpu-session-baseline-harness+evidence 一併撤除,
scripts/lib/self-referential-bootstrap.ps1、scripts/tests/test-self-referential-bootstrap.ps1、
scripts/self-referential-bootstrap-ledger.json 還原為 origin/main 版本。
機械上這條路也是死路:base-pinned 裁決者以 base 版 classifier 驗證新 entry 宣告的
mechanism paths,同 PR 擴張 classifier 永遠無法讓自己的 entry 合法(實測兩輪
pr-metadata-contract-diagnostic 均以 not classified verification-mechanism paths 拒絕)。

二、rebind migrate-console-to-hifi-design row:#507 squash 後該 row 仍綁 pre-squash
commit af60c29(已被丟棄,CI checkout 抓不到)→ 全部後續 PR 的 machine-truth
test 25 紅。依 #482/#501/#512 慣例 rebind 到 landed squash 4187102。

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… threads)

- README 補 Document nature 標籤(CodeRabbit:docs/**/*.md 須宣告文件性質)
- §5 反例標明 harness 由 PR #511 引入(本樹尚無該檔,屬裁決紀錄)並補述
  升級規則維持 review 強制的互鎖理由(機器化=再改 adjudicator=另開 debt)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
monkey1sai added a commit that referenced this pull request Aug 12, 2026
…baseline 1.1) (#511)

* feat(scripts): add GPU session baseline measurement harness (gpu-session-baseline-and-idle-reclaim 1.1)

Add scripts/measure-session-baseline.ps1 (root CLI) and
scripts/lib/measure-session-baseline.ps1 (testable core) implementing task
1.1 of openspec/changes/gpu-session-baseline-and-idle-reclaim: a read-only
harness that captures nvidia-smi GPU inventory (VRAM/utilization, consumer
RTX classification, MIG availability), a GET-only WebRTC/coordinator health
probe (/health, /api/runtime/status), and the environment fingerprint
required by the gpu-session-baseline spec (GPU model, driver version, Kit
version from kit-sdk.packman.xml, fixture hash+size).

The harness never opens a WebRTC session or creates/joins/closes a review
session, so TTFF and session-creation success rate cannot be honestly
measured locally; those fields are null with measured:false and an explicit
reason unless supplied by a caller (e.g. a future task 1.3 soak run), never
fabricated. Every other unmeasurable signal (no nvidia-smi, no GPU rows,
insufficient OS permission on the compute-apps VRAM column, coordinator
unreachable) degrades the same way instead of throwing or guessing.

Registered in scripts/script-registry.json as a measurement-harness (not
deploy.ps1/verify-all.ps1: it measures, it does not deploy or gate; not
scripts/lib alone: it is the operator-invoked CLI entry; not scripts/tests:
it produces a JSON report, not a pass/fail check).

Add scripts/tests/test-measure-session-baseline.ps1: unit tests for GPU line
parsing/consumer-RTX/MIG classification, fail-safe behavior with nvidia-smi
entirely absent, report schema shape, script-registry.json consistency, and
a real CLI smoke test (both -OutputPath and the default
artifacts/gpu-baseline/<timestamp>.json path). Verified on pwsh 7.5.4 and
Windows PowerShell 5.1 (powershell.exe), invoke-powershell-static.ps1, and
scripts/tests/test-agent-governance-check.ps1 (45/45 green).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(openspec): tick gpu-session-baseline-and-idle-reclaim task 1.1

Mark tasks.md 1.1 done (measure-session-baseline.ps1 harness landed) and
sync openspec/lifecycle-ledger.json: task_ledger completed 0->1,
current_slice points at the 1.2 env-fingerprint gate as the next slice,
last_verified refreshed, subject_commit rebound to 405e2b6 (the commit that
landed the harness + tests + registry entry), evidence_refs extended to the
new script and test paths.

Verified: node scripts/tests/verify-openspec-repository-lifecycle.mjs
--repo-root . (openspec/changes, lifecycle-ledger.json and
docs/plans/NOW.md agree -- NOW.md's projection is id+status only, and
status stays "active", so it needed no edit); node --test
scripts/tests/test-openspec-machine-truth.mjs (24/24) and
scripts/tests/test-ai-coding-metrics.mjs (13/13); pwsh
scripts/tests/test-agent-governance-check.ps1 (45/45 embedded
repository-lifecycle subtests green).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(scripts): honest gpu-baseline measurements — real binding enum, per-role lease counts, validated inputs, untick 1.1 (review)

* fix(scripts): declare gpu fingerprint scope — multi-GPU host is first_gpu_only, per-GPU fingerprint deferred to 1.2 (review)

* fix(scripts): harden GPU baseline harness per PR #511 review

Addresses the valid findings from the 13 unresolved review threads on
scripts/measure-session-baseline.ps1 (gpu-session-baseline-and-idle-reclaim
task 1.1/1.2), keeping the "measured:false, never fabricate" contract intact:

- Get-SafeProperty: fix a PowerShell pipeline-unroll bug where `return
  $value` on an empty array collapsed to $null, making an observed empty
  kit_instance_bindings/viewer_leases indistinguishable from "unmeasured".
- Get-WebRtcHealthProbe: count non-terminal KitInstanceBinding statuses
  (allocated/starting/ready/draining) instead of a literal status='active'
  that the real coordinator API never emits; count active primary/spectator
  viewer_leases by role (the actual 1-primary-plus-k-spectator cardinality)
  instead of sessions.active_count; clarify that `reachable` reflects only
  coordinator /health liveness, not independent WebRTC/signaling reachability.
- Get-SessionVramWatermark: only claim a clean measured total when exactly
  one Kit GPU process is observed and fully readable; multi-process or
  partially-readable readouts are surfaced only as the informational
  unscoped_total_kit_vram_mb, never as a fabricated measured:true total.
- Get-EnvironmentFingerprint: fail closed (measured:false) on a multi-GPU
  host instead of blindly binding the fingerprint to gpus[0].
- Get-KitVersionFingerprint: carry an explicit source/caveat noting this is
  the checkout's declared dependency version, not a live-process read.
- Get-SessionBaselineReport: range-validate caller-supplied -TtffMs /
  -SessionCreationSuccessRate (reject negative/out-of-range instead of
  recording as measured); resolve host.hostname via the cross-platform Dns
  API with HOSTNAME/COMPUTERNAME fallback so Linux deployment targets don't
  silently null out host identity.
- Root wrapper: derive the default -OutputPath from the report's own
  collision-resistant run_id instead of a bare second-resolution timestamp.
- CI: run test-measure-session-baseline.ps1 (PS7 + Windows PowerShell 5.1)
  as part of the required `powershell-static` job, mirroring the existing
  test-spec-to-done-port-helper.ps1 pattern -- neither `root-contracts`
  (pytest) nor the PSScriptAnalyzer-only `powershell-static` gate command
  previously executed this suite.

The lifecycle-ledger subject-ancestry finding (PRRT_kwDOSPoer86YcZ4d) was
independently verified as already resolved at current HEAD and needed no
change; see PR reply for evidence.

Verified: pwsh + Windows PowerShell 5.1 test-measure-session-baseline.ps1,
invoke-powershell-static.ps1, and node --test test-openspec-machine-truth.mjs
(24/24) all green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(scripts): carry a live-build staleness caveat on kit_version (PR #511 review)

Get-KitVersionFingerprint reads bim-streaming-server/tools/deps/kit-sdk.packman.xml
-- the checkout's DECLARED kit-kernel dependency version -- not a value read
from the live Kit process. If the checkout is updated without a
rebuild/restart, this can be stale relative to the session actually being
measured, and no local mechanism exists to introspect a running Kit.exe's
build identity to close that gap. Surface a `source` ('checkout_packman_declared')
and an explicit `caveat` string alongside the existing value/measured/reason
shape (both on the raw fingerprint and propagated through
Get-EnvironmentFingerprint's kit_version field) so downstream SLO-writers
know what this field does and does not attest to, rather than silently
trusting checkout state as if it were live-process state.

This was the one review thread not already covered by the concurrent fixes
landed in 37e3247/2eee19b on this branch (real KitInstanceBinding status
enum, per-role viewer lease counts, VRAM attribution transparency,
TTFF/success-rate validation, hostname fallback, GPU fingerprint scope
disclosure, collision-resistant default filename, task 1.1 unticked); this
commit reconciles with that work rather than duplicating it.

Verified: pwsh + Windows PowerShell 5.1 test-measure-session-baseline.ps1,
invoke-powershell-static.ps1, and node --test test-openspec-machine-truth.mjs
(24/24) all green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(scripts): fold GPU-attribution gap into complete, fix MIG-gating and test cleanup (PR #511 review round 2)

Three round-2 findings on the GPU baseline harness:

- Get-EnvironmentFingerprint: gpu_fingerprint_scope='first_gpu_only' (a
  multi-GPU host) was disclosure-only -- `complete` was computed from the
  five base fields before the scope was known, so a report that has
  admittedly NOT attributed every relevant GPU could still report
  complete:true and silently suppress the wrapper's "SHALL NOT be used to
  set SLOs or admission parameters" warning. `complete` now also requires
  gpu_fingerprint_scope != 'first_gpu_only'.
- Get-GpuInventorySnapshot: software_queue_required was gated on
  consumer_grade_all AND NOT mig_available_any. On a non-consumer, non-MIG
  fleet (e.g. a lone RTX A6000, which Test-ConsumerRtxGpuName excludes but
  which does not support MIG at all), that reported
  software_queue_required=false with no MIG route in fact available.
  Software queuing is now required whenever MIG is unavailable, regardless
  of consumer/professional classification.
- test-measure-session-baseline.ps1: the default-OutputPath cleanup deleted
  every new file under artifacts/gpu-baseline/, not just the one this test
  produced -- a concurrent harness invocation sharing the checkout would
  have its evidence collaterally deleted. Now deletes only
  "$($defaultReport.run_id).json".

Also strengthened Assert-ReportSchemaShape's `complete` expectation and
added regression tests for the professional/no-MIG inventory shape and the
multi-GPU complete=false path.

Verified: pwsh + Windows PowerShell 5.1 test-measure-session-baseline.ps1,
invoke-powershell-static.ps1, and node --test test-openspec-machine-truth.mjs
(24/24) all green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* revert(ci): defer harness CI wiring until the open mechanism debt closes

ci.yml is a classified verification-mechanism path (design gate
infrastructure, Lane G minimum + self-referential bootstrap scope); wiring
test-measure-session-baseline.ps1 into CI from this measurement-harness PR
would collide with the open mechanism-hardening-2 ledger entry owned by
PR #513. The wiring moves to a follow-up alongside issue #516 (CI coverage
for the streaming pytest suite) after fixpoint closure.

Refs #516

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(scripts): gpu-baseline r2 - bootstrap registration, honest kit-version provenance, MIG-driven queue flag (review r2)

* fix(scripts): honest fixture provenance binding for baseline fingerprint (review r4, operator-delegated adjudication)

* fix(governance): revert harness bootstrap-layer per #520 ruling; rebind hifi row to #507 squash

一、撤除 bootstrap 層(依 #520 裁決=docs/agents/self-referential-bootstrap.md §2.1,PR #521):
量測 harness 的報告無任何 gate 機器消費者,不屬 mechanism surface,不入 ledger。
classifier 擴張+open entry gpu-session-baseline-harness+evidence 一併撤除,
scripts/lib/self-referential-bootstrap.ps1、scripts/tests/test-self-referential-bootstrap.ps1、
scripts/self-referential-bootstrap-ledger.json 還原為 origin/main 版本。
機械上這條路也是死路:base-pinned 裁決者以 base 版 classifier 驗證新 entry 宣告的
mechanism paths,同 PR 擴張 classifier 永遠無法讓自己的 entry 合法(實測兩輪
pr-metadata-contract-diagnostic 均以 not classified verification-mechanism paths 拒絕)。

二、rebind migrate-console-to-hifi-design row:#507 squash 後該 row 仍綁 pre-squash
commit af60c29(已被丟棄,CI checkout 抓不到)→ 全部後續 PR 的 machine-truth
test 25 紅。依 #482/#501/#512 慣例 rebind 到 landed squash 4187102。

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(scripts): gpu-baseline r5 — unknown-runtime honesty, session-scoped lease counts, fixture overwrite guard (review r5)

PR #511 review r5, three of four threads:

1. PRRT_kwDOSPoer86YeUls — an unreachable/malformed /api/runtime/status left
   every observed count null, and the fixture-binding defaults coerced those
   nulls to 0, so an UNKNOWN runtime state was published under the
   'no_live_session_observed' (observed-idle) label with complete=true beside
   a declared fixture. Adds a distinct fixture_binding_scope='runtime_state_unknown'
   that withdraws completeness and names the failed probe (GET /api/runtime/status).
   A positive observation still outranks the unknown; an observed 0/0 idle host
   keeps its previous semantics.

2. PRRT_kwDOSPoer86YeUlw — per-role lease counts are summed across every
   sessions.items[] entry while total_kit_vram_mb is one host-wide sample, so a
   host serving 2+ sessions mixed multiple primaries/spectators against a single
   VRAM number. Takes the reviewer's reject option: the aggregate counts are kept
   (they are real observations) but session_scope='multi_session_aggregate' is
   published and the 1-primary+k-spectator watermark interpretation is marked
   measured=false with reason 'non-isolated multi-session snapshot; per-session
   VRAM attribution unavailable in this slice'. Exactly one active session yields
   session_scope='single_session' and keeps current semantics.

3. PRRT_kwDOSPoer86YeUly — -FixturePath and -OutputPath resolving to the same
   file made Set-Content truncate the fixture with the report, destroying the very
   artifact the report fingerprints. Canonicalises both ([System.IO.Path]::GetFullPath,
   case-insensitive only on Windows) and throws before any write.

PRRT_kwDOSPoer86YeUlo (P1, bootstrap mechanism-path regression) is NOT addressed
here and is moot as of 286bbac on this branch: per the #520 ruling the harness is
not a mechanism surface, and the three measure-session-baseline classifier patterns
the thread asked the test to pin were removed. Adding them to $expectedMechanismPaths
now would fail the suite; re-registering them would revert an owner ruling.

Verified on Windows: test-measure-session-baseline.ps1 (all groups pass),
test-self-referential-bootstrap.ps1 (all assertions pass),
Invoke-ScriptAnalyzer -Severity Error on the changed .ps1 files = 0.

* fix(scripts): gpu-baseline r6 — reject malformed runtime counts, require observed primary, finite TTFF (review r6)

- Get-SessionVramWatermark / Get-EnvironmentFingerprint: a non-null-but-
  unparseable observed_active_session_count / observed_kit_instance_binding_count
  (coordinator version skew, e.g. a string) was silently coerced to 0 via a
  try/catch default, relabeling an UNKNOWN runtime state as an OBSERVED zero.
  New ConvertTo-NonNegativeIntOrNull helper returns null instead of 0 on parse
  failure; both call sites now treat that null the same as a missing probe
  (new 'malformed_runtime_observation' session scope; runtime_state_unknown
  fixture-binding scope).
- Get-SessionVramWatermark: exactly one active session was enough to accept
  the "1 primary + k spectator" watermark interpretation even when zero
  primary viewers had joined (idle-but-created or spectator-only session).
  Now requires observed_primary_lease_count == 1.
- New-OptionalMeasurement TTFF validator only checked ">= 0", which
  +Infinity satisfies; now also rejects non-finite values.
- openspec/lifecycle-ledger.json: added scripts/lib/measure-session-baseline.ps1
  to the change's evidence_refs (the entire measurement implementation lives
  there; only the CLI wrapper and test were previously listed).

Addresses the three still-open findings from the chatgpt-codex-connector
review on c759057, plus the infinite-TTFF gap from an earlier round that
was never landed.

Co-authored-by: monkey1sai <26239865+monkey1sai@users.noreply.github.com>

* fix(scripts): gpu-baseline r7 — keep kit process-count fields in every report shape (review r7)

gpu-session-baseline-report/v1 的兩個早退路徑補齊 kit_process_count 與
kit_process_vram_unreadable_count:查詢失敗=null(未知非零)、查詢成功但無
Kit process=0(觀測到的真零),consumer 不再因 host 狀態拿到不同 shape。
測試補四條斷言鎖住兩態。

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: monkey1sai <xshiujj@gmail.com>
Co-authored-by: claude[bot] <41898282+claude[bot]@users.noreply.github.com>
monkey1sai and others added 2 commits August 12, 2026 20:16
…evidence-harness-scope

# Conflicts:
#	scripts/self-referential-bootstrap-ledger.json
…式;review threads)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
monkey1sai and others added 2 commits August 12, 2026 20:38
…l debt gate

The first fixpoint run of a mechanism is also the first time that mechanism
executes as the canonical path, so it can surface a regression in itself. The
contract left no legal way to fix one: naming the existing open entry hit the
impersonation guard, self-registering a second entry hit the other-open-debt
gate, and declaring bootstrap=no hit that same gate — three-way interlock
(issue #494).

Ledger entries gain an OPTIONAL append-only `repair_prs` array (strictly
increasing positive integers; absent reads as empty, so the four immutable
closed entries stay byte-identical). `Assert-SelfReferentialJsonObjectShape`
grows `-OptionalProperties` so the exact-property-set rule can admit a new
field without requiring it of entries written before it existed.

A repair PR is admitted only when all five hold:
1. the named entry is pre-existing debt, open at base AND still open at head;
2. the transition's only change to it is a repair_prs tail append whose
   appended value is exactly this PR number (no live PR number => refused);
3. every mechanism path this PR changes is already inside that entry's
   declared verification_mechanism_paths (case-sensitive);
4. the PR touches none of this gate's own adjudicators;
5. the transition neither opens nor closes any entry.

Invariants preserved: the ledger stays append-only (repair_prs elements are
never rewritten or dropped); every other entry field stays immutable and the
only status transition is still one open -> closed; closed entries remain
fully immutable — repair_prs is not a back door into them; the closure still
demands a complete, all-green fixpoint attestation against the frozen
verification contract, so a repaired mechanism must prove itself; and a single
transition still may not both settle and incur debt.

Refs #494

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 52488e8189

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread docs/agents/self-referential-bootstrap.md Outdated
Comment thread scripts/lib/self-referential-bootstrap.ps1 Outdated
…-lane head

Folding issue #494's repair lane into this PR grew the entry's declared
surface from two paths to four (the gate library and its suite joined the
contract prose and the ledger). The bootstrap evidence still recorded the
pre-fold head 82c6aa4 and described the fixpoint as having 'one thing to
prove', which understated what the entry now covers.

Reran the four contract suites at 52488e8 (all exit 0, timings recorded)
and corrected the README's scope paragraph to state both obligations.

Refs #494 #520

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@monkey1sai

monkey1sai commented Aug 12, 2026 •

Copy link
Copy Markdown
Owner Author

併入通知(coordinator,依 owner 指示):issue #494 的 regression-repair lane 已折進本 PR(commit 52488e8 實作+cb457f8 evidence 重綁),而非另開第二筆 bootstrap entry——兩者共用同一筆 debt 與同一輪 fixpoint。

已同步處理的連動:

  1. entry 的 verification_mechanism_paths 由 2 條擴為 4 條(以 Get-SelfReferentialMechanismPaths 對本 PR diff 實跑結果為準);command_ids 未動故 contract_sha256 不變。
  2. bootstrap evidence 於新 head cb457f8 重跑四條契約套件(全 exit 0)並更正 README 的範圍段落——原文「fixpoint 只有一件事要證」在擴張後已不準確。
  3. PR body 補上先前未涵蓋的 Windows On-Demand Verification(tier deploy_dryrun)。觸發原因是 ^scripts/lib/(?!platform/)[^/]+\.ps1$ 這條寬 glob 命中 gate library,而非真實 deploy-path 行為變更(deploy.ps1 並未 dot-source 本 library,已 grep 確認);仍依規補上 -DryRun exit 0 的實跑證據與 CI run URL。

repair lane 的五個放行條件(全滿足才放行,否則逐條具名拒絕):

  1. 命名的 entry 於 base 存在且 status=open,head 仍 open
  2. 只對它做 repair_prs 尾端追加,且追加值恰為本 PR 號
  3. 本 PR 的 mechanism paths ⊆ 該 entry 已宣告範圍(case-sensitive)
  4. 完全不觸及任何 adjudicator path
  5. 同一 transition 不新增 entry、不同時關帳

不變式全保:ledger append-only、entry 唯一 open→closed 轉移、closure attestation 全綠、closed entry 不可變。repair_prs 為選填欄位(shape 檢查新增 -OptionalProperties),四筆既有 closed entry 維持 byte-identical。

未更動:§2.1 條文本身、evidence 的既有結論、machine pattern 清單。若您(#521 原作者 session)對範圍擴張有異議,請回覆,我可把 repair lane 拆回獨立 PR(代價=多一輪 debt 與 fixpoint)。

@codex-tri-adversarial-bot codex-tri-adversarial-bot Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Codex Tri-Adversarial Bot

Automated tri-adversarial ship-gate (L0 terra triage / L1 tier-routed lens fanout / L2 refute-by-default / L3 sol apex — Codex models).
Mapped event: COMMENT


Codex Tri-Adversarial ship-gate — PR #521

  • Repo head: governance/bootstrap-evidence-harness-scope @ 2af2697
  • Base: main @ c5d423c
  • Files changed: 6
  • Engine: four-model tri-adversarial gate on Codex — L0 triage gpt-5.6-terra/low; L1 lens finders routed gpt-5.6-terra/low → gpt-5.6-luna/medium → gpt-5.5/xhigh (security floor gpt-5.5); L2 refute-by-default gpt-5.5/xhigh, top-tier findings refuted by gpt-5.6-sol/xhigh (every refutation cross-model); L3 apex gpt-5.6-sol/max. 誠實聲明:層級與 Claude 三層 gate 同構(terra≈haiku、luna≈sonnet、gpt-5.5≈opus、sol≈fable),但模型池是 Codex 的,非 Anthropic 的。

Verdict

HELD — 三層驗證未能完成,本次不投同意票(fail-closed)。

  • held reason: apex_unavailable_or_failed
  • mapped GitHub event: COMMENT

Difficulty & routing

  • overall: critical (source: terra-triage)
  • lens tiers: correctness→gpt-5.5, security→gpt-5.5, simplification→gpt-5.5, test-gap→gpt-5.5

Layer stats

  • L1: raw=5 deduped=5 finder_failures=0
  • L2: confirmed=0 refuted=5 unverified=0
  • L3 final: 0

Agent calls

  • 10/12 ok, engine wall-clock 233.9s

VERDICT

HELD

VERDICT: HELD

monkey1sai and others added 2 commits August 12, 2026 22:47
…uire real repair work

Two connector threads, both verified before acting.

(A) Condition 4 banned adjudicator edits outright, on the theory that a PR
repairing the rule that judges it could wave itself through. Measured against
the workflow, that theory does not hold: pr-review-agent.yml checks out
pull_request.base.sha (:25), materializes the gate from BASE via git archive
(:82-86), and exits non-zero with base_gate_incomplete_external_approval_required
(:90-94) rather than ever falling back to head; :113 always resolves the
checker under GATE_ROOT, which is only set on the base-pinned branch. The only
other invocation is scripts/dev/check-pr-local-preflight.ps1, a developer
preflight with no merge authority. test-base-gate-capability.ps1 is the
executable form of the invariant and passes.

The ban also deadlocked the debt it protected: this repository's own open entry
declares two adjudicator paths, so a failing fixpoint on it would have had no
lane at all — issue #494 reproduced one level up. Adjudicator paths are now
repairable, bounded by the declared surface. Safety rests on base-pinned
adjudication (no self-clearance), declared-subset (no scope expansion), and the
unchanged fixpoint obligation (the repaired mechanism must still prove itself
against the entry's frozen verification contract).

(B) A repair could append a repair_prs record while fixing nothing. The
proposed test — "changed at least one mechanism path" — is vacuous: measured,
the ledger is itself a classified mechanism path and a repair PR necessarily
edits it to append repair_prs, so the count is >= 1 by construction, and a PR
with zero mechanism paths returns from the body gate before the lane is
reached. The check that actually bites requires a NON-LEDGER mechanism path —
the exact mirror of the closure rule, which permits the ledger and nothing
else. Demonstrated: with the check disabled, a ledger-only repair was admitted.

Conditions stay five and stay aligned with §2.2: (3) is now real repair work,
(4) is the declared-surface bound that carries the adjudicator carve-out. The
condition-5 closure check is retained as defence in depth and labelled as such
— with (3) in place, a repair that also closes debt is refused earlier on both
reachable paths, and both are now pinned by tests.

Refs #494

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The four contract suites were re-run on 7f590ea (declared adjudicator paths are
repairable; a repair must change a non-ledger mechanism path) with a clean
tree; tested_head, run_at and every duration record that run rather than the
superseded aa32597 one. This commit changes only this evidence file, so no
untested code tree is being claimed as passing.

Refs #494

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.
To continue using code reviews, add credits to your account and enable them for code reviews in your settings.

@codex-tri-adversarial-bot codex-tri-adversarial-bot Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Codex Tri-Adversarial Bot

Automated tri-adversarial ship-gate (L0 terra triage / L1 tier-routed lens fanout / L2 refute-by-default / L3 sol apex — Codex models).
Mapped event: COMMENT


Codex Tri-Adversarial ship-gate — PR #521

  • Repo head: governance/bootstrap-evidence-harness-scope @ 1f298a4
  • Base: main @ c5d423c
  • Files changed: 6
  • Engine: four-model tri-adversarial gate on Codex — L0 triage gpt-5.6-terra/low; L1 lens finders routed gpt-5.6-terra/low → gpt-5.6-luna/medium → gpt-5.5/xhigh (security floor gpt-5.5); L2 refute-by-default gpt-5.5/xhigh, top-tier findings refuted by gpt-5.6-sol/xhigh (every refutation cross-model); L3 apex gpt-5.6-sol/max. 誠實聲明:層級與 Claude 三層 gate 同構(terra≈haiku、luna≈sonnet、gpt-5.5≈opus、sol≈fable),但模型池是 Codex 的,非 Anthropic 的。

Verdict

HELD — 三層驗證未能完成,本次不投同意票(fail-closed)。

  • held reason: all_finders_failed
  • mapped GitHub event: COMMENT

Difficulty & routing

  • overall: high (source: fail-safe-default)
  • lens tiers: correctness→gpt-5.6-luna, security→gpt-5.5, simplification→gpt-5.6-luna, test-gap→gpt-5.6-luna

Agent calls

  • 0/10 ok, engine wall-clock 66.9s

VERDICT

HELD

VERDICT: HELD

@monkey1sai

Copy link
Copy Markdown
Owner Author

@codex review

@claude review

Exact head: 1f298a4. Current-head local PR preflight passed; all review threads are resolved; please perform an independent final-head review only.

@claude

claude Bot commented Aug 13, 2026 •

Copy link
Copy Markdown

Claude finished @monkey1sai's task in 2m 57s —— View job


Independent final-head review — head 1f298a4

  • Gather context (PR description, prior review threads, diff)
  • Verify exact head 1f298a4 and diff against origin/main (c5d423c)
  • Independent review of scripts/lib/self-referential-bootstrap.ps1 repair-lane logic
  • Independent review of docs/agents/self-referential-bootstrap.md contract prose
  • Independent review of ledger/evidence files
  • Check regression test coverage for repair-lane conditions
  • Post final review

Scope of this review: static/manual trace of the code and doc diff (git diff origin/main...HEAD), plus verification that docs/evidence/.../gate-suites.txt's tested_head=7f590ea is an ancestor of 1f298a4 and that 7f590ea..1f298a4 only touches the evidence file itself (confirmed via git diff --stat). I was not able to actually execute pwsh in this sandboxed environment (command approval was denied), so I did not re-run test-self-referential-bootstrap.ps1 myself — this review relies on reading the implementation against the tests rather than re-executing them. Flagging this so you can weigh it accordingly; PR-provided evidence already shows a fresh exit-0 run at the pinned head.

Findings

All four defects raised across the three prior codex-tri-adversarial-bot rounds and the Codex inline comments appear correctly fixed at this head, each with a matching regression test:

  1. New entries can't fabricate repair_prs history (self-referential-bootstrap.ps1:891-900) — Compare-SelfReferentialLedgerTransition throws for any head-only entry that declares repair_prs, including []. Covered by the new entry born with repair_prs: [...] loop over @(), @(123), @(123, 456).
  2. Case-sensitive entry immutability (:944, :955, :966) — all three canonical-entry comparisons (base open→open, base open→closed, base closed) now use -cne instead of -ne. I confirmed the PowerShell 7.5.4 case-insensitivity claim independently is consistent with documented .NET/PowerShell string-comparison semantics for -ne. Covered by four case-only-mutation tests (reason, mechanism path, closed-entry reason, closure-comparison field).
  3. Single-door repair lane (L1-COR-001) (Assert-SelfReferentialRepairLane, the $allRepairs.Count -ne 1 check) — the fix correctly validates against the full Transition.RepairEntries collection, not just the body-named entry, so a second entry's repair_prs can no longer ride along unbound. Covered by the two-open-entry test pair.
  4. Real repair work required / adjudicator paths repairable within declared scope — condition 3 (non-ledger mechanism path required) and condition 4 (declared-subset, adjudicator paths now included) both match the doc's §2.2 five-condition list exactly, in the same order, and the base-pinned-adjudication argument for relaxing the adjudicator ban (pr-review-agent.yml checks out base.sha via git archive, never falls back to head) is architecturally sound — self-clearance isn't possible because the gate that judges this PR is materialized from main, not from this branch.

I traced condition ordering, the Get-SelfReferentialRepairAppend tail-append detection, and the interaction between the repair lane and the pre-existing closure/new-entry checks (e.g. a repair can't combine with closing a different entry in the same PR — it fails closed via the closure's "may only change the ledger" check rather than an explicit repair-lane rule, which is safe but slightly implicit). No correctness gaps found in this pass.

scripts/self-referential-bootstrap-ledger.json's new evidence-harness-scope-clarification entry correctly omits repair_prs, and its four verification_mechanism_paths match the actual diff surface (doc + gate library + ledger + test suite).

Non-blocking, pre-existing observation (not introduced by this PR)

The earlier Codex review comment about web-viewer-sample/e2e/conv-history.spec.ts and design-system-visual.spec.ts not being classified by Get-SelfReferentialMechanismPaths despite feeding required-CI gate consumers appears to still be unaddressed. This is a gap in the existing machine pattern list (which this PR intentionally leaves unchanged — "機器 pattern 清單不變"), not something this PR's diff introduces or worsens, so I don't consider it a blocker for this PR, but it may be worth a follow-up issue if §2.1's new "output consumed by a gate machine-consumer" predicate is meant to be exhaustively enforced by the classifier rather than just documented.

Verdict

No new correctness issues found at 1f298a4. All previously-flagged findings (Copilot P1, Codex P1/P2, and all three codex-tri-adversarial-bot rounds' confirmed findings) have matching code fixes and regression tests at this head. The two most recent tri-adversarial bot runs (HELD) failed due to engine/apex unavailability, not a substantive finding.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 1f298a45ea

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread scripts/lib/self-referential-bootstrap.ps1 Outdated
@monkey1sai

Copy link
Copy Markdown
Owner Author

@codex review @claude review Exact head: bc5a06b. PR body now matches the linked-successor implementation. Four exact-code-head governance gates and Windows deploy dry-run exited 0; Sol/Terra adversarial review found no remaining blocker. Please review only this exact head.

@claude

claude Bot commented Aug 13, 2026 •

Copy link
Copy Markdown

Claude finished @monkey1sai's task in 3m 21s —— View job


Independent final-head review — head bc5a06b

  • Gather context (PR description, prior review threads, diff)
  • Verify exact head bc5a06b and diff against origin/main (c5d423c)
  • Review the linked-successor transition logic in scripts/lib/self-referential-bootstrap.ps1
  • Review docs/agents/self-referential-bootstrap.md contract prose (§2.1/§2.2/§2.2.1)
  • Check regression coverage for orphan/fork/cycle/second-successor/three-open-debt/downgrade/child-first/close-both
  • Review ledger/evidence files for this head
  • Post final review

Scope of this review: static/manual trace of git diff origin/main...HEAD plus the incremental diff from the previously-reviewed head 1f298a4. I could not get pwsh execution approved in this sandbox (same limitation the prior review at 1f298a4 hit — command approval denied), so this is not an independent re-execution of the four gate suites; it relies on reading the implementation against the tests and against git history directly.

What changed since the last independent review (1f298a4 → bc5a06b)

  • 1f298a4..1ef0ea8 ("fix(governance): add linked bootstrap successor debt") adds §2.2.1 to the doc and ~200/~555 lines to the gate library/test suite, implementing exactly the governed linked-successor transition the PR description promises. This directly resolves the one P1 Codex finding still open against 1f298a4 ("Permit repairs that require a new mechanism dependency" — a repair needing an outside-declared-surface classified path had no legal path forward).
  • 1ef0ea8..bc5a06b is evidence-only: gate-suites.txt rebinds tested_head to 1ef0ea8. Confirmed directly: 1ef0ea8 is the sole parent of bc5a06b, and git diff 1ef0ea8..bc5a06b --stat touches only that one evidence file (8 lines changed). So the evidence is bound to the actual reviewed code tree, not a stale one.

Linked-successor implementation trace

Read Assert-SelfReferentialRepairLane (scripts/lib/self-referential-bootstrap.ps1:1155-1333) and the ledger-parse-time integrity checks (:469-563) against the doc's §2.2.1 rules and cross-checked each against its regression test:

  • Adjudicator-repair deadlock removed: condition 4 now permits repairing declared-surface adjudicator paths (base-pinned adjudication via pr-review-agent.yml checking out base.sha is the safety argument, same as the prior review verified) — closing the exact gap Codex flagged at 1f298a4.
  • Outside-surface repair requires exactly one successor (:1266-1323): successor_of must name the repaired entry, no pre-existing successor, predecessor must be the only open debt at base, successor's verification_mechanism_paths must equal ledger+outside-paths exactly (no missing/extra/duplicate), and successor's command_ids must preserve the predecessor's as an ordered prefix (no downgrade/reorder). Each has a matching Assert-Throws test (test-self-referential-bootstrap.ps1:1351-1424).
  • Orphan/cycle/fork integrity is enforced at ledger-parse time regardless of transition path (:536-562): missing predecessor, second successor, and multi-entry cycles all fail closed, each with a direct test (:395-439).
  • Three-open-chain prevented: repairing a successor to spawn a second-generation successor is blocked while its own predecessor is still open, via the "only open debt at base" check reused from the first-generation case (test-self-referential-bootstrap.ps1:1430-1456).
  • Closure ordering: Compare-SelfReferentialLedgerTransition:1127-1139 rejects predecessor+successor closing in the same transition; the body-level "no other open debt" exception (:1463-1496) only fires for a single predecessor-only closure whose sole remaining open debt is its direct successor, and only allows ledger+that-entry's-fixpoint-evidence paths. Tests cover the success case (predecessor-first, :1706-1711), successor-first rejection (:1751-1758), and simultaneous close-both rejection (:1810-1820) — matching the three adversarial cases the PR description says Sol/Terra added.
  • Ledger at head (scripts/self-referential-bootstrap-ledger.json) has exactly one open entry (evidence-harness-scope-clarification, PR docs(governance): self-referential-bootstrap 契約增補 §2.1 範圍界定與 scope 反例(issue #520) #521) with no repair_prs/successor_of, declaring the 4 mechanism paths that match the diff (doc, gate library, ledger, test suite) — consistent with the PR body.

Minor, non-blocking observation

Assert-SelfReferentialBootstrapBody:1523 filters $transition.OpenDebt with [string]$_.id -ne $entryId (case-insensitive -ne), while nearly every other comparison in this file was deliberately hardened to -cne/-ccontains after the earlier case-sensitivity findings. In isolation this would be a gap, but entry id is constrained by schema to ^[a-z0-9][a-z0-9-]{2,63}$ (:435) — lowercase only — so no two ids can differ only by case, and this comparison is inert. Not a functional defect at this head; flagging only because it's the one remaining -ne on an id-like comparison in a file that otherwise treats that as a bug class.

Verdict

No new correctness issues found at bc5a06b. The linked-successor transition matches its documentation precisely, is well-covered by adversarial regression tests for the specific failure modes called out in the PR description (orphan/fork/cycle/second-successor/three-open-debt/downgrade/child-first/close-both), and closes the one P1 finding that was still open against the previous reviewed head. Evidence (gate-suites.txt, tested_head=1ef0ea8) is correctly bound to the code that's actually at this head.

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. Nice work!

Reviewed commit: bc5a06bc74

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

@monkey1sai

Copy link
Copy Markdown
Owner Author

@codex review

@claude review

Exact head: bc5a06bc74ccfd217106537d7b71207e0c851953.

The PR body now matches the linked-successor implementation. Four exact-code-head governance gates, local PR preflight, hosted CI, hosted Agent Governance, hosted PR Metadata Contract, and Windows deploy dry-run passed. Sol/Terra adversarial review found no remaining blocker. Please review only this exact head.

@claude

claude Bot commented Aug 13, 2026 •

Copy link
Copy Markdown

Claude finished @monkey1sai's task in 2m 1s —— View job


Independent final-head review — head bc5a06b

  • Gather context (PR description, prior review threads, diff)
  • Verify exact head bc5a06b matches current PR HEAD and diff against origin/main
  • Check hosted CI / Agent Governance / PR Metadata Contract status
  • Confirm no code changed since my prior review of this same head
  • Post final review

This is the same head I already reviewed. git log -1 on this checkout confirms HEAD = bc5a06bc74ccfd217106537d7b71207e0c851953, merge-base with origin/main is c5d423c — identical to the state I independently reviewed in my prior comment on this PR (job 31670859552). No new commits exist between that review and this request, so the code-level findings from that pass still stand:

  • Linked-successor transition (Assert-SelfReferentialRepairLane, scripts/lib/self-referential-bootstrap.ps1:1155-1333) matches §2.2.1 of the doc precisely — outside-surface repairs require exactly one successor_of entry, exact-set mechanism paths, ordered-prefix command_ids (no downgrade), and predecessor-first closure with no simultaneous close-both. Each rule has a matching regression test.
  • Orphan/cycle/fork integrity is enforced at ledger-parse time regardless of transition path, with direct tests.
  • Evidence (gate-suites.txt, tested_head=1ef0ea8) is correctly bound to the actual code tree (1ef0ea8..bc5a06b touches only that one evidence file).
  • One pre-existing, non-blocking observation carried over: Assert-SelfReferentialBootstrapBody:1523 still uses case-insensitive -ne on OpenDebt.id, which is inert because entry id is schema-constrained to lowercase (:435), so no functional defect at this head.

New in this request — hosted check status (verified via mcp__github_ci__get_ci_status, not just the PR-body claim):

Check Result
CI ✅ success
Agent Governance ✅ success
PR Metadata Contract ✅ success (one earlier transient failure at 05:31:29Z, succeeded on the 05:35:27Z re-run)
Governance Base Audit (governance-trust-root.yml) ❌ failing, twice, at this head

The Governance Base Audit failure is not a defect in this PR's diff — scripts/dev/check_governance_trust_root.py (unchanged by this PR) requires a fresh exact-head approval from monkey1sai-blip (_has_exact_head_approval, line 515), and the PR body itself already discloses this as outstanding ("monkey1sai-blip exact-head approval pending"). The last recorded monkey1sai-blip approval was pinned to an earlier head (302fde3), not bc5a06b, so this check is expected to stay red until that approval is re-submitted at the current head — it's a gating state, not a code correctness issue. Flagging it because the trigger comment's claim list didn't mention this workflow and I wanted to verify status independently rather than take the summary at face value.

Verdict

No new correctness issues at bc5a06b — this head is unchanged from the one I already reviewed in depth. The remaining blocker to merge is procedural: monkey1sai-blip's exact-head approval, which the PR body already tracks as pending and which Governance Base Audit is correctly fail-closing on.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: bc5a06bc74

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread scripts/lib/self-referential-bootstrap.ps1 Outdated
@monkey1sai

Copy link
Copy Markdown
Owner Author

@codex review exact head e670d11. Focus on the linear active successor-chain state machine, repair-target versus attachment-leaf separation, active-chain surface union, status monotonicity, exactly-one oldest-root closure, and non-empty repair history.

@monkey1sai

Copy link
Copy Markdown
Owner Author

@claude review exact head e670d11. Focus on governance contract consistency, adversarial reachability, fail-closed parser integrity, and exact-code evidence binding.

@claude

claude Bot commented Aug 13, 2026 •

Copy link
Copy Markdown

Claude finished @monkey1sai's task in 3m 33s —— View job


Independent final-head review — head e670d11

  • Gather context (PR description, prior review threads, diff)
  • Verify exact head e670d11 and diff since last reviewed head bc5a06b
  • Governance contract consistency (doc §2.2.1 vs implementation)
  • Adversarial reachability (successor-chain deadlock, closure ordering)
  • Fail-closed parser integrity (ledger parse-time checks)
  • Exact-code evidence binding (gate-suites.txt tested_head)
  • Post final review

Scope: static/manual trace of git diff bc5a06b..e670d11 (the incremental change since my last independent review of this PR) against docs/agents/self-referential-bootstrap.md §2.2/§2.2.1 and the regression suite. pwsh execution is not available in this sandbox, so this is not a re-run of the four gate suites — it relies on reading the implementation against the tests and the PR-provided exact-head evidence.

What changed since bc5a06b

bc5a06b..e5e2c98 ("fix(governance): permit linear bootstrap successor chains") replaces the earlier "exactly-one predecessor→successor pair" model with a full linear active-chain model. e5e2c98..e670d11 is evidence-only (gate-suites.txt + README.md, confirmed via git diff --stat e5e2c98..e670d11 — no executable path touched). e5e2c98 is the sole parent of e670d11, so tested_head=e5e2c9829b49dae5b9964f553fea73788977c345 in gate-suites.txt is bound to the actual reviewed code tree.

1. Governance contract consistency

docs/agents/self-referential-bootstrap.md §2.2.1's five linked-successor rules match Assert-SelfReferentialRepairLane (scripts/lib/self-referential-bootstrap.ps1:1227-1410) and Get-SelfReferentialOpenSuccessorChain (:588-651) point-for-point:

  • Rule 2 (chain must be a single contiguous open lineage, no fork/orphan/cycle/unrelated debt) ⇔ Get-SelfReferentialOpenSuccessorChain -RequireAllOpen (:1355-1356).
  • Rule 3 (successor surface = ledger + active-chain-union-external paths, no duplicate ownership) ⇔ $expectedSuccessorPaths/$activeSurface construction (:1327-1334, :1372-1384).
  • Rule 4 (command_ids ordered-prefix against the attachment leaf, not necessarily the repair target) ⇔ :1387-1399, using $attachmentLeaf (activeChain[-1]) rather than $headEntries[0].
  • Rule 5 ("repair target may be any open chain entry; new successor always attaches to leaf") ⇔ $attachmentLeafId is computed independently of $EntryId and used for both the successor_of check (:1361) and the existing-successor check (:1364-1369), while repair_prs binding stays keyed to $EntryId (:1270). Test 'nested successor forks from non-leaf repair target' and the open-grandparent-debt repair case (test file, ~line 1500) exercise exactly this target≠leaf split.
  • The closure exception in Assert-SelfReferentialBootstrapBody (:1540-1583) correctly requires closing the oldest open root and matches the remaining debt against Get-SelfReferentialOpenSuccessorChain -RequireAllOpen | Select-Object -Skip 1, i.e. the exact contiguous suffix — matching the doc's "只能依 close A → close B → close C 關帳" rule.

2. Adversarial reachability — the one open P1 is now closed

The Codex finding posted against bc5a06b ("Permit successor chains after later fixpoint failures" — repairing B while A is still open was previously rejected, recreating the deadlock) is resolved: the otherBaseOpen/"only open debt at PR base" check that caused that deadlock is removed, replaced by the chain-aware RequireAllOpen check, so a later fixpoint on any open chain member can legally extend the chain as long as the whole base-open set is one lineage. This is exercised end-to-end by the new A -> B -> C lifecycle test block (test file ~1455-2010): open-chain repair creating a nested successor, non-leaf repair target, fork-from-non-leaf rejection, ancestor-surface-overlap rejection, downgrade rejection, unrelated-open-debt rejection, and the full three-step close A → close B → close C closure sequence with real fixture commits.

3. Fail-closed parser integrity

  • New: ledger-parse-time rejection of a closed successor whose predecessor is still open (:551-561), independent of transition-time checks — this is a genuine strengthening, since it's enforced regardless of which code path produced the malformed ledger, not just the closure-transition body check.
  • New: repair_prs must be non-empty when present, enforced at raw JSON parse time (Assert-SelfReferentialJsonRepairPrs, :268-271) — closes the gap the two codex-tri-adversarial-bot rounds flagged (new entries fabricating empty-array repair history), and is layered under the existing new-entry-must-not-declare-repair_prs check so it also catches a base entry mutated to repair_prs: [].
  • Get-SelfReferentialOpenSuccessorChain's own cycle/fork detection (:622, :630) is reachable only in principle — the ledger-parse-time checks (:549, :571) already reject multi-successor forks and cycles across the entire ledger (open or closed) before this function ever runs on a validated BaseLedger/HeadLedger. That makes the backward root-walk in this function (which would otherwise infinite-loop on an undetected open↔open cycle) safe in practice. This is consistent with the file's existing "defence in depth" style elsewhere (e.g. the comment at :1401-1405) — not a defect, just worth noting these particular branches have no direct unit test isolating them from the parser's earlier check.

4. Exact-code evidence binding

Confirmed directly: e670d11^ == e5e2c98, git diff --stat e5e2c98..e670d11 touches only the two evidence files, and gate-suites.txt's tested_head matches e5e2c98 exactly. The four recorded suites use the canonical pwsh -NoProfile -NonInteractive -File <suite> invocation and all report exit_code=0.

Non-blocking, pre-existing observation (unchanged by this diff)

Assert-SelfReferentialBootstrapBody:1610 ($otherOpen = ... Where-Object { [string]$_.id -ne $entryId }) still uses case-insensitive -ne rather than -cne, same as the instance flagged non-blocking in my bc5a06b review — inert because entry id is schema-constrained to lowercase. Not touched by this increment.

Verdict

No new correctness issues found at e670d11. The linear active-chain model is a real fix for the deadlock the previous head's P1 finding identified, doc/implementation stay in lockstep across all five §2.2.1 rules, the new parser-level checks (closure-order, non-empty repair_prs) fail closed independent of the transition path that produces them, and the evidence binding is exact.

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. More of your lovely PRs please.

Reviewed commit: e670d116c6

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants