Skip to content

Fail closed on unavailable Windows startup evidence - #10238

Closed
lawrencecchen wants to merge 23 commits into
feat-tui-startup-benchmark-runner-recutfrom
fix-pr10131-windows-evidence
Closed

lawrencecchen wants to merge 23 commits into
feat-tui-startup-benchmark-runner-recutfrom
fix-pr10131-windows-evidence

Conversation

@lawrencecchen

@lawrencecchen lawrencecchen commented Aug 17, 2026 •

Copy link
Copy Markdown
Contributor

Stacked on #10131.

The trusted preflight currently reports Windows job, process, privilege, and handle observations that it does not collect. The red test requires unavailable observations to stay null. The follow-up commit will relay only the optional child job observation and leave unsupported signals unavailable, so existing validation rejects an unproven Windows claim.

Regression proof uses two commits: test first, implementation second. Hosted CI only; no local Rust/Cargo/Zig/Xcode tests.


View with [code]smith Autofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.


Summary by cubic

Fails closed when Windows startup preflight evidence is unavailable and publishes an explicit skipped Windows claim. Previously, missing Windows signals implied success; now Windows-only unavailability exits 78 with a Windows-specific reason, sets claim_status=unverified and claim_reason, attests the account-process AppContainer probe and staging failures, and gates packaging, paired runs, and profile capture.

Review notes

  • Enforces a shared evidence contract in startup_benchmark_contract.py; verify-startup-benchmark.py validates schema v8, rejects malformed/foreign fields, and classifies Windows verified vs unverified. Links claims to attested inputs by checking bootstrap SHA-256 and approved imports plus Windows API set DLLs, and requires linked AppContainer feasibility evidence and the attested account-process probe.
  • Windows preflight relays only windows_grandchild_in_job; true is required for verification, None exits 78 with a Windows reason (unverified), and false hard-fails. AppContainer feasibility unavailability exits 78; access denied or core failures hard-fail. The account-process probe and staging capability failures are attested before any skip.
  • CI runs contract and claim tests, handles Windows exit 78, writes the skipped claim via startup_benchmark_claim.py, and gates downstream on claim_status/claim_reason.

Required actions

  • Do not synthesize unsupported Windows observations.
  • Gate all downstream steps on claim_status; treat Windows exit 78 as unverified, not error.
  • When unverified on Windows, use startup_benchmark_claim.py to write the skipped claim with expected SHA-256s; include and link Windows AppContainer feasibility evidence and the attested account-process probe.

Written for commit e488456. Summary will update on new commits.

Review in cubic

@coderabbitai

coderabbitai Bot commented Aug 17, 2026 •

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: a4a60854-e59d-4a86-b098-6b947963d61a

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@cursor

cursor Bot commented Aug 17, 2026

Copy link
Copy Markdown

Bugbot is paused — on-demand spend limit reached

Bugbot uses usage-based billing for this team and has hit its on-demand spend limit.

A team admin can raise the spend limit in the Cursor dashboard, or wait for the next billing cycle to continue.

@cursor

cursor Bot commented Aug 17, 2026

Copy link
Copy Markdown

Bugbot is paused — on-demand spend limit reached

Bugbot uses usage-based billing for this team and has hit its on-demand spend limit.

A team admin can raise the spend limit in the Cursor dashboard, or wait for the next billing cycle to continue.

@blacksmith-sh

This comment has been minimized.

@blacksmith-sh

This comment has been minimized.

@lawrencecchen

Copy link
Copy Markdown
Contributor Author

Closing the stale stacked startup-benchmark child. The parent #10131 is retired and this head conflicts with current main. Carry any needed Windows evidence into the fresh benchmark design in #11697; do not merge this aggregate.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant