Skip to content

test(reborn): complete WS9 generated state machines - #6886

Merged
serrrfirat merged 3 commits into
mainfrom
codex/ws9-equivalence-state-machines
Jul 30, 2026
Merged

serrrfirat merged 3 commits into
mainfrom
codex/ws9-equivalence-state-machines

Conversation

@serrrfirat

@serrrfirat serrrfirat commented Jul 29, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • Complete WS9’s typed seven-dimension equivalence model with 37 representative rows and 10 selected high-risk pairwise interactions, without a Cartesian explosion.
  • Generate restart and concurrent same-thread double-submit/interleaving sequences alongside the existing trigger/retry/cancel/duplicate lifecycle generator.
  • Assert process ownership and the exact production-composed capability resource governor after every transition; add sabotage cases for each new generator/invariant.
  • Preserve the existing lease_wedge truthful uncertain-outcome proof and map it into the mechanical dimension × sequence × invariant registry. The registry reports 0 remaining WS9 gaps.

Overlap reconciliation: #6794 remains the source of the initial lifecycle/focused-boundary slice and is not duplicated. This branch was rebased and adapted after #6696 made the row-native process journal the sole lifecycle authority. Open #5981 and #6096 do not currently implement this same-thread generated restart/double-submit scope; #5981 may intentionally change the current busy-admission outcome later.

Change Type

  • Bug fix
  • New feature
  • Refactor
  • Documentation
  • CI/Infrastructure
  • Security
  • Dependencies

Linked Issue

Related #6524

Validation

  • cargo fmt --all -- --check
  • Scoped zero-warning clippy for both touched packages: cargo clippy -p ironclaw_reborn_composition -p ironclaw_reborn_integration_tests --tests --all-features -- -D warnings
  • cargo build — not applicable: compilation is covered by both integration targets and scoped all-feature clippy; no shipping build behavior changed.
  • Relevant tests pass: 29 generated gate tests, 15 restart tests, depth-3 generated gate run, lease-wedge/reopen integration targets, architecture suite, and 50 Python coverage/meta-tests listed below.
  • cargo test --features integration if database-backed or integration behavior changed — not applicable: no database implementation/schema behavior changed; the targeted LibSQL restart integration target exercises the affected persistence seam.
  • Manual testing — not applicable: deterministic test infrastructure only; no user interface or manual product workflow changed.
  • pr-shepherd maintainer-quality review completed before requesting review; its architecture/evidence findings were fixed and the focused validation was repeated on the final rebased head.

Test Strategy

User behavior:
Generated Reborn runs retain one authoritative process tree and no leaked capability resource holds across trigger, retry, cancel, duplicate, restart, and concurrent same-thread double-submit interleavings. Durable approval outcomes remain recoverable after a real runtime restart.

Risk areas:

  • Model behavior
  • Browser
  • Side effect
  • Persistence
  • Security or permissions
  • External provider
  • Cross-component behavior

Tests added or updated:

  • Unit or contract: tests/e2e/scenarios/test_state_machine_coverage.py mechanically checks all 7 dimensions, 37 representative rows, 10 required high-risk pairs, all 6 sequences, all 5 invariants, duplicate case-id rejection, and 0 retained gaps. Journey-derived claims require matching declared observable evidence, preventing admission-only journeys from being credited with terminal stability.
  • Reborn integration: reborn_generated_gate_sequences checks per-transition ownership/resource invariants and all same-thread double-submit schedules; reborn_generated_restart_sequences rebuilds the coordinator/executor/scheduler/scope gateway/process journal on a fresh LibSQL connection for approve/deny/cancel after restart.
  • Recorded fixture: Not applicable: no HTTP/provider fixture or recorded boundary changed.
  • Browser E2E: Not applicable: no browser surface or frontend behavior changed.
  • Backend or runtime: Targeted LibSQL restart integration plus production-composed capability-governor read-back through a feature-gated composition test-support builder.
  • Live canary: Not applicable: no live provider, credential, network, or deployment behavior changed.

What the tests prove:

  • Dimension coverage: 7/7 typed axes; representative equivalence rows cover each value and all selected high-risk pairs.
  • Sequence coverage: 6/6 (trigger, retry, cancel, duplicate, restart, concurrent_double_submit).
  • Invariant coverage: 5/5, checked mechanically and after generated transitions where applicable.
  • Restart: gate identity/state survives a fresh LibSQL connection and rebuilt runtime; approve emits exactly one effect, deny/cancel emit none.
  • Concurrent double-submit: first/second/simultaneous schedules yield exactly one accepted run and one busy result referencing that run.
  • No orphans/reservations: every expected process ID is specifically an AgentTurn, every parent exists, no unexpected agent-turn process appears, and the exact production-composed governor has no leaked holds at quiescent blocked/terminal boundaries. In-flight transitions check process ownership without misclassifying a legitimate executing-capability reservation as an orphan.
  • Truthful uncertainty: the existing reborn_integration_lease_wedge proof is the canonical evidence and is referenced by the registry rather than weakened or duplicated.
  • Sabotage: missing restart-deny generation, two accepted submissions, omitted agent-turn ownership, an expected ID under the wrong process kind, a missing process parent, a leaked governor hold, unsupported journey evidence, and an incomplete composite registry all fail loudly.
  • Effect evidence: restart approve reads back exact persisted file contents while deny/cancel prove file absence; every double-submit interleaving proves exactly one capability result plus durable file read-back.
  • Depth policy: deterministic depth 2 remains the PR default (12 gate/lifecycle sequences plus 3 double-submit schedules); the nightly workflow uses reproducible depth 3. A depth-3 local run exercised 39 gate/lifecycle sequences.
  • Remaining proven WS9 gap count: 0.

Commands run:

cargo test -p ironclaw_reborn_integration_tests --test reborn_generated_gate_sequences --test reborn_generated_restart_sequences
IRONCLAW_GENERATED_SEQUENCE_DEPTH=3 cargo test -p ironclaw_reborn_integration_tests --test reborn_generated_gate_sequences generated_gate_sequences_preserve_lifecycle_invariants -- --nocapture
cargo test -p ironclaw_reborn_integration_tests --test reborn_integration_lease_wedge --test reborn_integration_reopen_resume_through_gate
uv run --project tests/e2e pytest -q tests/e2e/scenarios/test_state_machine_coverage.py tests/e2e/scenarios/test_journey_coverage.py
uvx ruff check tests/e2e/state_machine_coverage.py tests/e2e/scenarios/test_state_machine_coverage.py
uvx ruff format --check tests/e2e/state_machine_coverage.py tests/e2e/scenarios/test_state_machine_coverage.py
cargo test -p ironclaw_architecture
cargo test -p ironclaw_architecture --test reborn_struct_test_support_ratchet reborn_production_struct_test_support_and_dead_code_members_do_not_grow -- --nocapture
cargo clippy -p ironclaw_reborn_composition -p ironclaw_reborn_integration_tests --tests --all-features -- -D warnings
cargo fmt --all -- --check
git diff --check
scripts/ci/reborn-coverage-ratchet.sh <PR-6886-merged.lcov> tests/integration/coverage-exemptions.toml tests/integration/coverage-floor.toml

Additional package-suite evidence:

  • cargo test -p ironclaw_reborn_composition: 541 passed; four timeout-shaped failures passed when rerun individually. observability::trace_capture::tests::capture_skips_when_policy_missing_or_disabled fails identically on an untouched origin/main worktree, so it is a pre-existing, non-attributable base failure.
  • The coverage floor was recaptured from PR test(reborn): complete WS9 generated state machines #6886’s successful df5081ad4 merged artifact: global 85.54% (315436/368757), runner 86.87% (14669/16887), processes 88.07% (5839/6630), and turns 86.21% (9453/10965). The ratchet passes against that exact artifact with the new floors.

Security Impact

None. The only composition change is feature-gated to test-support and returns the exact already-composed resource governor alongside the runtime for read-only invariant assertions. It does not alter authorization, approvals, network mediation, secrets, sandboxing, or production capability dispatch.

Reborn Trust-Boundary Checklist

  • Public policy/evidence/trust-bearing types: no new production trust-bearing types or construction paths.
  • Untrusted content enters prompts only through an envelope/escaping primitive: not applicable; no prompt/content ingress changed.
  • Hashes declare purpose; trust/binding/authenticity uses SHA-256/BLAKE3 or separate authenticity check: not applicable; no hashes changed.
  • New/changed status, exit, policy, runtime, or error variants: downstream match sites audited. Command/output: no variants added; targeted rg and compilation/clippy cover the touched runtime/test-support accessors.
  • Security/durability serde(default) fields fail closed or have migration tests: not applicable; no serialized fields changed.
  • Queues/maps/buffers/counters have bounds and overflow-safe arithmetic: generator depth is explicitly bounded; the representative registry is finite and duplicate-checked.
  • Driver/operator-visible errors have stable class semantics: no production error classes changed; unsupported restart backends fail loudly in the harness.
  • Sandbox/native/host names accurately describe trust boundary: no production trust-boundary names changed; the seam returns the exact production-composed governor only under test support.

Database Impact

None. No migration, schema, query contract, or backend implementation changes. The restart generator uses LibSQL because it is the hermetic backend with an independent reopen recipe; unsupported InMemory/Postgres restart modes fail loudly rather than pretending to restart.

Blast Radius

Test infrastructure under tests/e2e and tests/integration, the Reborn test workflow’s documented generated counts, and a feature-gated composition test-support builder. The main review risks are keeping the representative registry synchronized with supported enums and preserving current same-thread busy-admission semantics if #5981 later changes them.

Rollback Plan

Revert commits 50348a75c, 7f78b108a, and 513478eb1. There is no migration, persisted format change, configuration change, or production default to roll back. CI would return to #6794’s prior generated coverage.

Review Follow-Through

Maintainer self-review found and fixed two issues before readiness: the governor read seam initially grew the production runtime service shape, and generic journey rows initially inferred stronger lifecycle evidence than their declared assertions proved. The final seam is owned by composition test support, and the registry now derives lifecycle/sequence/invariant claims only from explicit observable evidence. If #5981 lands first and changes busy admission into queued steering, rebase and update the double-submit expected outcome deliberately rather than weakening the invariant.

The complete eight-reviewer code-review-multi pass then found eight actionable issues, all fixed: process-kind validation, in-flight reservation timing, restart and double-submit durable effect read-back, restart denominator simplification, coverage-floor recapture, owning-guide documentation, and private-helper naming. One additional suggestion—treating equal axes/evidence with different case IDs as duplicate rows—was reproduced and rejected because case_id is the executable parameter identity for distinct journey cases sharing one parametrized test.

Three pre-existing bot findings on the reviewed old head were also fixed and resolved: provider read rows no longer claim an effect-count invariant, approval denial now uses the actually observed Completed lifecycle with an explicit generated assertion, and the typed registry selector helpers have concrete return annotations.

Checkbox-to-test/PR map

  • 7 typed dimensions + pairwise representatives → Python registry/meta-tests in this PR.
  • Restart sequences → reborn_generated_restart_sequences in this PR.
  • Concurrent double-submit/interleavings → generated gate target in this PR.
  • No orphan runs/reservations + exact resource governor → per-transition assertion helper and sabotage tests in this PR.
  • Truthful uncertain outcome → existing reborn_integration_lease_wedge, referenced mechanically by the registry.
  • Fast PR / deeper nightly reproducibility → default depth 2 and nightly depth 3 workflow; depth-3 evidence recorded above.
  • Remaining proven WS9 gaps → 0.

Review track: A (tests/CI infrastructure; production behavior unchanged)

@coderabbitai

coderabbitai Bot commented Jul 29, 2026 •

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Summary by CodeRabbit

  • Reliability

    • Improved lifecycle handling for concurrent submissions, cancellations, approvals, and restarts.
    • Added validation that completed workflows leave no orphaned processes or resource reservations.
    • Added durable restart coverage to verify expected outcomes after runtime recovery.
  • Tests

    • Expanded state-machine coverage across authentication, policy, provider outcomes, sequences, and invariants.
    • Added completeness checks to detect missing, duplicate, or unsupported lifecycle scenarios.
    • Updated performance and sequence-count descriptions for depth-3 test runs.

Walkthrough

Changes

The PR adds typed state-machine coverage inventories and fail-loud gap tests, exposes runtime resource governors to integration harnesses, adds LibSQL planned-runtime restart testing for approval outcomes, and strengthens orphan ownership/reservation and double-submit invariant checks.

Reborn lifecycle validation

Layer / File(s) Summary
Runtime governor observability
crates/ironclaw_reborn_composition/..., tests/integration/support/harness/*
Runtime construction and host harnesses now expose the composed ResourceGovernor; alternate harness profiles explicitly use None.
Orphan and interleaving invariants
tests/integration/generated_gate_sequences.rs, tests/integration/support/assertions.rs
Generated gate tests validate process ownership, parent relationships, orphan runs, reservations, and same-thread double-submit interleavings.
Planned runtime restart recovery
tests/integration/generated_restart_sequences.rs, tests/integration/support/group.rs, Cargo.toml, .github/workflows/reborn-tests.yml
LibSQL groups retain their builder wiring for restart, and generated tests cover approve, deny, and cancel recovery branches.
State-machine coverage inventory
tests/e2e/state_machine_coverage.py, tests/e2e/scenarios/test_state_machine_coverage.py
Typed coverage cases, evidence checks, duplicate detection, dimension gaps, pairwise gaps, sequence gaps, and invariant gaps are added.

Estimated code review effort: 4 (Complex) | ~45 minutes

Possibly related issues

Possibly related PRs

  • nearai/ironclaw#6609 — Related Cargo test-target discovery changes affect registration of the new generated restart sequence test.

Sequence Diagram(s)

sequenceDiagram
  participant PreRestartHarness
  participant RebornIntegrationGroup
  participant LibSQL
  participant PostRestartHarness
  PreRestartHarness->>LibSQL: persist parked approval-gated run
  PreRestartHarness->>RebornIntegrationGroup: request planned runtime restart
  RebornIntegrationGroup->>LibSQL: reopen durable composite
  RebornIntegrationGroup->>PostRestartHarness: rebuild runtime and thread
  PostRestartHarness->>LibSQL: apply approve, deny, or cancel
  PostRestartHarness-->>PreRestartHarness: verify terminal status and effect count
Loading

Suggested reviewers: ilblackdragon

🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title uses Conventional Commits style and accurately summarizes the WS9 generated state-machine coverage work.
Description check ✅ Passed The description matches the repository template and fills the required summary, issue, validation, test strategy, and review sections.

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@railway-app

railway-app Bot commented Jul 29, 2026 •

Copy link
Copy Markdown

🚅 Deployed to the ironclaw-pr-6886 environment in ironclaw-ci-preview

Service Status Web Updated (UTC)
ironclaw ✅ Success (View Logs) Web Jul 30, 2026 at 10:50 am

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6886 July 29, 2026 22:56 Destroyed
@github-actions github-actions Bot added scope: ci CI/CD workflows scope: dependencies Dependency updates size: S 10-49 changed lines labels Jul 29, 2026
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Caution

The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased.

@github-actions github-actions Bot added risk: medium Business logic, config, or moderate-risk modules contributor: core 20+ merged PRs labels Jul 29, 2026
@ironloopai

ironloopai Bot commented Jul 29, 2026 •

Copy link
Copy Markdown
Contributor

🔎 Review · PR #6886

🟢 Completed · Review submitted

1 actionable findings →

Reviewed the complete trusted base-to-head comparison. The restart, concurrent-admission, harness, composition seam, workflow, and coverage-registry changes are generally coherent, but the registry overstates one journey’s executable invariant evidence.

Automatic · PR opened + CI failed · attempt 1 of 3 · completed in 1m 56s

Run details
  • Repository: nearai/ironclaw
  • Base: main at bed3f68
  • Head: codex/ws9-equivalence-state-machines at c91a54b
  • Created: Jul 29, 2026, 11:01 PM UTC
  • Updated: Jul 29, 2026, 11:03 PM UTC
  • Run: cb549eef-89e6-4ddf-a7c9-aa5a357eb542
  • Latest attempt: 1 · Completed · e177a39b-84db-4128-9c2f-c31da59b1f09

@ironloopai ironloopai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔍 Review complete · PR #6886

💬 1 finding

Reviewed the complete trusted base-to-head comparison. The restart, concurrent-admission, harness, composition seam, workflow, and coverage-registry changes are generally coherent, but the registry overstates one journey’s executable invariant evidence.

Findings

  1. 🟠 Medium · Admission-only webhook evidence is credited with terminal stability — tests/e2e/state_machine_coverage.py:161-164
    Details are attached to the relevant diff.
Validation and technical details
  • Inspected all 17 changed files and surrounding composition, resource-governor, process-journal, integration-harness, journey-registry, and evidence-checker code.
  • git diff --check refs/ironloop/base...refs/ironloop/head passed.
  • Repository knowledge graph was unavailable; followed the required guidance-and-targeted-search fallback.
  • Rust and Python tests could not be executed because cargo and uv are not installed in the review environment.
  • Base: main
  • Head: codex/ws9-equivalence-state-machines at c91a54b
  • Run: cb549eef-89e6-4ddf-a7c9-aa5a357eb542

Comment thread tests/e2e/state_machine_coverage.py Outdated
@serrrfirat
serrrfirat force-pushed the codex/ws9-equivalence-state-machines branch from c91a54b to df5081a Compare July 30, 2026 09:28
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6886 July 30, 2026 09:28 Destroyed
@serrrfirat
serrrfirat marked this pull request as ready for review July 30, 2026 09:31
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Caution

The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased.

@ironloopai

ironloopai Bot commented Jul 30, 2026 •

Copy link
Copy Markdown
Contributor

🔎 Review · PR #6886

🟢 Completed · Review submitted

1 actionable findings →

Reviewed the complete trusted base-to-head comparison. The runtime test-support seam and restart/double-submit harness changes are coherent, but the new coverage registry overstates its mechanically proven lifecycle coverage: a required denial row is backed by a test that never verifies the claimed denied state.

Automatic · PR opened · attempt 1 of 3 · completed in 2m 4s

Run details
  • Repository: nearai/ironclaw
  • Base: main at a378b39
  • Head: codex/ws9-equivalence-state-machines at df5081a
  • Created: Jul 30, 2026, 9:36 AM UTC
  • Updated: Jul 30, 2026, 9:38 AM UTC
  • Run: 8a1dd829-716b-430f-b118-783e8490e047
  • Latest attempt: 1 · Completed · 7382c57b-9f12-4991-abfa-e829748ac670

@ironloopai ironloopai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔍 Review complete · PR #6886

⚠️ 1 finding · 1 blocking

Reviewed the complete trusted base-to-head comparison. The runtime test-support seam and restart/double-submit harness changes are coherent, but the new coverage registry overstates its mechanically proven lifecycle coverage: a required denial row is backed by a test that never verifies the claimed denied state.

Findings

  1. 🟠 Medium · Denial coverage is credited without observing a denied lifecycle state — tests/e2e/state_machine_coverage.py:277-290
    Details are attached to the relevant diff.
Validation and technical details
  • Compared refs/ironloop/base a378b39 through refs/ironloop/head df5081a across all 18 changed files.
  • Inspected composition governor wiring, all test-support consumers, group restart reconstruction, generated gate/restart sequences, process/resource assertions, Python coverage registry/meta-tests, workflow, and Cargo target registration.
  • python3 -m py_compile tests/e2e/state_machine_coverage.py tests/e2e/scenarios/test_state_machine_coverage.py passed.
  • git diff --check refs/ironloop/base..refs/ironloop/head passed.
  • Rust and pytest suites could not be executed because this review environment does not provide cargo or uv.
  • Base: main
  • Head: codex/ws9-equivalence-state-machines at df5081a
  • Run: 8a1dd829-716b-430f-b118-783e8490e047

Comment thread tests/e2e/state_machine_coverage.py Outdated

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@tests/e2e/state_machine_coverage.py`:
- Around line 105-121: Annotate _cargo, _operation, and _fault with their
appropriate return types. Import the provider operation and fault case types
from their respective modules, using those types for the selector helpers and
the existing CargoEvidence type for _cargo.
- Around line 178-195: Update _provider_case to derive invariants from
operation_class instead of unconditionally assigning
StateMachineInvariant.AT_MOST_ONCE_EFFECT. Include the effect invariant only for
write-capable operations, while READ cases such as github_get_issue and
google_sheets_read_values_empty use no effect invariant.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: bf12ae04-8488-4ee6-ad11-731eda5107fd

📥 Commits

Reviewing files that changed from the base of the PR and between a378b39 and df5081a.

📒 Files selected for processing (18)
  • .github/workflows/reborn-tests.yml
  • Cargo.toml
  • crates/ironclaw_reborn_composition/src/runtime.rs
  • crates/ironclaw_reborn_composition/src/test_support/mod.rs
  • tests/e2e/scenarios/test_state_machine_coverage.py
  • tests/e2e/state_machine_coverage.py
  • tests/integration/generated_gate_sequences.rs
  • tests/integration/generated_restart_sequences.rs
  • tests/integration/support/assertions.rs
  • tests/integration/support/builder.rs
  • tests/integration/support/group.rs
  • tests/integration/support/harness/mod.rs
  • tests/integration/support/harness/profiles/core_builtin.rs
  • tests/integration/support/harness/profiles/github.rs
  • tests/integration/support/harness/profiles/mock_mcp.rs
  • tests/integration/support/harness/profiles/qa_smoke.rs
  • tests/integration/support/harness/profiles/web_access.rs
  • tests/integration/support/harness/recorder.rs

Comment thread tests/e2e/state_machine_coverage.py Outdated
Comment thread tests/e2e/state_machine_coverage.py
@github-actions

github-actions Bot commented Jul 30, 2026 •

Copy link
Copy Markdown
Contributor

Coverage ratchet

Ratchet mode: ENFORCING

RATCHET PASS: global
  observed: 85.55% (315466 / 368757 lines)
  floor:    85.54% (tolerance 0.5pp -> effective floor 85.04%)
  denominator: 368757 lines now vs 368757 at floor capture (+0 lines, +0%) — not a material change

RATCHET PASS: ironclaw_runner
  observed: 86.87% (14669 / 16887 lines)
  floor:    86.87% (tolerance 0.5pp -> effective floor 86.37%)
  floor_covered_lines: 14669 (tolerance 20 lines -> effective floor 14649)
  denominator: 16887 lines now vs 16887 at floor capture (+0 lines, +0%) — not a material change

RATCHET PASS: ironclaw_processes
  observed: 88.05% (5838 / 6630 lines)
  floor:    88.07% (tolerance 0.5pp -> effective floor 87.57%)
  floor_covered_lines: 5839 (tolerance 20 lines -> effective floor 5819)
  denominator: 6630 lines now vs 6630 at floor capture (+0 lines, +0%) — not a material change

RATCHET PASS: ironclaw_turns
  observed: 86.21% (9453 / 10965 lines)
  floor:    86.21% (tolerance 0.5pp -> effective floor 85.71%)
  floor_covered_lines: 9453 (tolerance 20 lines -> effective floor 9433)
  denominator: 10965 lines now vs 10965 at floor capture (+0 lines, +0%) — not a material change

⚠️ 2 Reborn crate(s) have 0 int-tier coverage (target: 0) — ironclaw_prompt_envelope, ironclaw_scripts

Reborn integration-tier coverage

Line coverage (Reborn crates): 85.55% — 315466 / 368757 lines

Per-crate breakdown (60 crates, lowest-covered first)
Crate Line % Covered / Total
ironclaw_prompt_envelope 0% 0 / 88
ironclaw_scripts 0% 0 / 345
ironclaw_process_sandbox 34.38% 120 / 349
ironclaw_host_ingress 42.5% 17 / 40
ironclaw_event_projections 43.51% 684 / 1572
ironclaw_libsql_runtime 59.91% 127 / 212
ironclaw_observability 61.54% 16 / 26
ironclaw_authorization 63.02% 610 / 968
ironclaw_memory 64.41% 959 / 1489
ironclaw_telegram_v2_adapter 69.69% 731 / 1049
ironclaw_trust 73.21% 664 / 907
ironclaw_wasm_limiter 74.6% 47 / 63
ironclaw_extractors 74.72% 538 / 720
ironclaw_capabilities 74.95% 2882 / 3845
ironclaw_projects 76.48% 400 / 523
ironclaw_mcp 76.6% 779 / 1017
ironclaw_filesystem 77.54% 5989 / 7724
ironclaw_reborn_cli 78.37% 10795 / 13774
ironclaw_wasm 79.72% 735 / 922
ironclaw_llm 80.82% 23319 / 28854
ironclaw_memory_native 81.02% 3299 / 4072
ironclaw_auth 81.88% 6679 / 8157
ironclaw_first_party_extensions 82.38% 6682 / 8111
ironclaw_host_api 83.68% 9728 / 11625
ironclaw_events 83.8% 1536 / 1833
ironclaw_reborn_identity 83.8% 450 / 537
ironclaw_operator 84.41% 5561 / 6588
ironclaw_secrets 84.53% 2797 / 3309
ironclaw_network 85.09% 959 / 1127
ironclaw_reborn_config 85.23% 2101 / 2465
ironclaw_skills 85.27% 4493 / 5269
ironclaw_telegram_extension 85.4% 1299 / 1521
ironclaw_extension_host 85.52% 19926 / 23299
ironclaw_approvals 86.08% 1818 / 2112
ironclaw_triggers 86.11% 2803 / 3255
ironclaw_turns 86.21% 9453 / 10965
ironclaw_reborn_composition 86.24% 22054 / 25573
ironclaw_webui 86.37% 11171 / 12934
ironclaw_hooks 86.73% 9950 / 11473
ironclaw_runner 86.87% 14669 / 16887
ironclaw_common 86.99% 1772 / 2037
ironclaw_reborn_event_store 87.19% 1327 / 1522
ironclaw_extensions 87.45% 4927 / 5634
ironclaw_product 87.49% 21354 / 24406
ironclaw_reborn_openai_compat 88% 3724 / 4232
ironclaw_processes 88.05% 5838 / 6630
ironclaw_reborn_traces 88.13% 11987 / 13601
ironclaw_threads 88.57% 4943 / 5581
ironclaw_host_runtime 89.09% 20539 / 23054
ironclaw_slack_extension 89.26% 2027 / 2271
ironclaw_conversations 90.03% 3171 / 3522
ironclaw_resources 90.84% 4474 / 4925
ironclaw_loop_host 90.93% 17735 / 19505
ironclaw_event_streams 91.24% 1063 / 1165
ironclaw_attachments 93.06% 630 / 677
ironclaw_outbound 93.91% 4101 / 4367
ironclaw_agent_loop 94.42% 10422 / 11038
ironclaw_safety 95.22% 3941 / 4139
ironclaw_first_party_extension_ports 95.71% 3837 / 4009
ironclaw_runtime_policy 96.56% 814 / 843

This table itself is informational and never gates the PR on its own — not the percentage, not the per-crate holes, not the 0-coverage callout. A separate coverage ratchet (dry-run until enforce=true; see tests/integration/coverage-floor.toml) can fail the build on specific configured floors.

Exemptions (3 entry/entries excluded from the accounting above)
Module / Crate Reason Issue
crate: ironclaw_embeddings v1-only: consumed only by root ironclaw (src/app.rs, src/tools/builtin/memory.rs, src/workspace/mod.rs, src/config/{mod,embeddings}.rs); no crates/* dependents. Covered by "Tests (Legacy)". #5657
crate: ironclaw_gateway v1-only: consumed only by root ironclaw (src/channels/web/platform/static_files.rs, src/channels/web/handlers/frontend.rs); no crates/* dependents. Covered by "Tests (Legacy)". #5657
crate: ironclaw_tui v1-only: consumed only by root ironclaw (src/main.rs, src/channels/tui.rs); no crates/* dependents. Crate's own doc comment confirms it bridges INTO v1, not Reborn. Covered by "Tests (Legacy)". #5657

@serrrfirat
serrrfirat force-pushed the codex/ws9-equivalence-state-machines branch from df5081a to 7f78b10 Compare July 30, 2026 10:39
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6886 July 30, 2026 10:39 Destroyed
@github-actions github-actions Bot added the scope: docs Documentation label Jul 30, 2026
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6886 July 30, 2026 10:43 Destroyed
@serrrfirat
serrrfirat merged commit 6535edc into main Jul 30, 2026
65 checks passed
@serrrfirat
serrrfirat deleted the codex/ws9-equivalence-state-machines branch July 30, 2026 11:19
l3ocifer pushed a commit to l3ocifer/frick-ironclaw that referenced this pull request Sep 3, 2026
* test(reborn): complete WS9 generated state machines

* test(reborn): strengthen WS9 invariant evidence

* test(reborn): align registry claims with evidence

This branch was successfully deployed

No deployments
ironclaw-ci-preview / ironclaw-pr-6886 — 50348a75 Deployed Jul 30, 2026 by railway-app[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: medium Business logic, config, or moderate-risk modules scope: ci CI/CD workflows scope: dependencies Dependency updates scope: docs Documentation size: S 10-49 changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant