Skip to content

test(e2e): isolate mutable Emulate provider worlds - #6525

Merged
serrrfirat merged 3 commits into
mainfrom
codex/isolate-emulate-provider-worlds
Jul 23, 2026
Merged

serrrfirat merged 3 commits into
mainfrom
codex/isolate-emulate-provider-worlds

Conversation

@serrrfirat

@serrrfirat serrrfirat commented Jul 22, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • Isolate mutable Emulate provider state between full-path QA journeys without rebuilding Ironclaw or restarting Reborn per case.
  • Restart seeded Google state on stable ports and remove Slack deliveries using provider-issued timestamps so existing OAuth state remains valid.
  • Add clean-baseline assertions and immediately repeated mutation journeys to catch order-dependent tests.
  • Keep read-only providers warm and document the fixture lifecycle.

Change Type

  • Bug fix
  • New feature
  • Refactor
  • Documentation
  • CI/Infrastructure
  • Security
  • Dependencies

Linked Issue

Part of #6524

Validation

  • cargo fmt --all -- --check — Not applicable: no Rust files changed.
  • cargo clippy --all --benches --tests --examples --all-features -- -D warnings — Not applicable: no Rust files changed.
  • cargo build — cargo build -p ironclaw --bin ironclaw
  • Relevant tests pass: 96 Emulate provider-contract, recorded-replay, and full-path tests.
  • cargo test --features integration if database-backed or integration behavior changed — Not applicable: test harness only; no database or product behavior changed.
  • Manual testing: exercised the exact combined CI pytest selection against the pinned local Emulate fork.
  • If a coding agent was used and supports it, review-pr or pr-shepherd --fix was run before requesting review — iterative autoreview completed with no actionable findings.

Test Strategy

User behavior: Full-path recorded QA journeys run hermetically without inheriting mutable provider state from a previous case, while CI retains one built binary and one module-scoped Reborn process.

Risk areas:

  • Model behavior
  • Browser
  • Side effect
  • Persistence
  • Security or permissions
  • External provider
  • Cross-component behavior

Tests added or updated:

  • Unit or contract: Clean-baseline assertions verify expected Gmail, Drive, and Slack mutations are absent before each applicable journey.
  • Reborn integration: The provider full-path matrix runs every Emulate-supported QA journey through standalone Reborn, extension auth, capability dispatch, and provider readback.
  • Recorded fixture: Existing harvested QA traces remain the deterministic model/tool-call inputs; two mutation journeys are repeated immediately to prove isolation.
  • Browser E2E: Not applicable: this changes the backend/provider E2E harness and no browser behavior.
  • Backend or runtime: The combined provider-contract, replay, and full-path selection passes with one binary build and one Reborn runtime.
  • Live canary: Not applicable: no production behavior or live-provider contract changed; this PR improves the hermetic tier.

What the tests prove: Provider mutation cases start clean, produce the expected provider-observable side effect, and cannot pass only because an earlier case left matching data behind. Read-only provider processes remain reused.

Commands run:

cargo build -p ironclaw --bin ironclaw
IRONCLAW_EMULATE_CLI=/Volumes/NVME/codex-bac3-emulate/packages/emulate/dist/index.js \
  ./tests/e2e/.venv/bin/pytest -q \
  tests/e2e/scenarios/test_emulate_reborn_provider_contracts.py \
  tests/e2e/scenarios/test_reborn_qa_trace_replay.py \
  tests/e2e/scenarios/test_reborn_qa_trace_full_path.py \
  --timeout=240
./tests/e2e/.venv/bin/python -m py_compile \
  tests/e2e/conftest.py \
  tests/e2e/scenarios/test_reborn_qa_trace_full_path.py
bash scripts/ci/check-reborn-qa-fixtures.sh
git diff --check

Result: 96 passed in 107.35s; fixture validation passed for 61 files.

Security Impact

None. The change only controls local Emulate processes and test data cleanup; it does not alter product permissions, secrets, network policy, file access, tool execution, or sandbox behavior.

Reborn Trust-Boundary Checklist

N/A: test-harness-only change; no Reborn policy, runtime, persistence, status, error, serialization, or trust-bearing production types changed.

Database Impact

None.

Blast Radius

Limited to the Emulate-backed E2E fixtures and full-path QA trace matrix. A fixture lifecycle bug could cause provider tests to fail or leak local test state; product binaries and runtime behavior are unchanged.

Rollback Plan

Revert this commit to restore the previous module-scoped provider fixtures. No data migration or compatibility step is required.

Review Follow-Through

CodeRabbit thread 3634027367 was addressed in de7e32be6: threaded Slack sends now use conversations.replies for outcome, baseline, and exact-timestamp cleanup, with a real Emulate regression test. No known review follow-up remains.


Review track: C (CI/test infrastructure)

@ironloopai

ironloopai Bot commented Jul 22, 2026 •

Copy link
Copy Markdown
Contributor

🔎 IronLoop Review Status

Head: c160bbdedd513b738b3fffbafe5681df484cbc02
Result: One or more review results were superseded by a newer PR head.
Next: Run @ironloopai review on the latest PR head.
Updated: 2026-07-23T07:22:39.329Z

Current reviewers:

Reviewer State Verdict Findings Last update
ironloop/common-reviewer (reviewer) Superseded N/A N/A 2026-07-22T22:27:18.555Z
Reviewer summaries
Reviewer Detail
ironloop/common-reviewer (reviewer) Superseded by a newer PR head. New head: de7e32b. Previous verdict: Needs validation.
Recent activity
Time Reviewer State Detail
2026-07-22T21:37:18.622Z ironloop/common-reviewer (reviewer) Queued Accepted review request for head e0a598d.
2026-07-22T21:37:18.622Z ironloop/common-reviewer (reviewer) Queued Waiting for this reviewer lane to become available.
2026-07-22T21:37:19.130Z ironloop/common-reviewer (reviewer) Started Reviewer worker started.
2026-07-22T21:37:21.791Z ironloop/common-reviewer (reviewer) Workspace ready Prepared isolated checkout (merge_ref) at 048c620.
2026-07-22T21:40:55.751Z ironloop/common-reviewer (reviewer) Result captured Needs validation; 0 blocking findings.
2026-07-22T21:40:55.751Z ironloop/common-reviewer (reviewer) Completed Review completed and terminal status was persisted.
2026-07-22T22:27:18.555Z ironloop/common-reviewer (reviewer) Superseded A newer PR head replaced this review (de7e32b).
Available commands
  • @ironloopai help
  • @ironloopai agents
  • @ironloopai review
  • @ironloopai review --agent <agent>
Run metadata

Admission: webhook accepted the request and IronLoop persisted reviewer state before this projection.

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6525 July 22, 2026 21:37 Destroyed
@github-actions github-actions Bot added size: XS < 10 changed lines (excluding docs) risk: low Changes to docs, tests, or low-risk modules contributor: core 20+ merged PRs labels Jul 22, 2026
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Caution

The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased.

@coderabbitai

coderabbitai Bot commented Jul 22, 2026 •

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Walkthrough

Emulate-backed QA journeys reuse a module-scoped Reborn runtime, run providers on stable ports, reset mutated provider state between cases, and assert clean provider baselines before replay. Repeated isolation cases and supporting E2E documentation were added.

Changes

Provider journey isolation

Layer / File(s) Summary
Stable Emulate provider world
tests/e2e/conftest.py
Emulate servers accept pinned ports, while ResettableEmulateProviderWorld manages provider startup, reset, and teardown.
Shared runtime and per-case reset
tests/e2e/scenarios/test_reborn_qa_trace_full_path.py
Journey traces identify mutated providers, reuse one Reborn runtime, and reset or clean provider state between cases.
Baseline assertions and cleanup
tests/e2e/scenarios/test_reborn_qa_trace_full_path.py, tests/e2e/CLAUDE.md
Google and Slack baselines are checked before replay, Slack thread mutations are cleaned, isolation cases are tested, and the workflow is documented.

Estimated code review effort: 4 (Complex) | ~45 minutes

Sequence Diagram(s)

sequenceDiagram
  participant QAJourney
  participant ProviderFixture
  participant ProviderWorld
  participant RebornRuntime
  participant EmulateProviders
  QAJourney->>ProviderFixture: load trace and identify mutations
  ProviderFixture->>ProviderWorld: reset mutated services
  ProviderWorld->>EmulateProviders: restart or clean provider state
  ProviderFixture->>RebornRuntime: reuse shared runtime
  RebornRuntime->>EmulateProviders: replay journey
  ProviderFixture->>EmulateProviders: assert baseline and clean Slack mutations
Loading

Possibly related PRs

Suggested reviewers: think-in-universe

🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title uses Conventional Commits style and accurately summarizes the E2E fixture isolation change.
Description check ✅ Passed The description matches the template closely and includes all major required sections, including validation and rollout details.

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions github-actions Bot added the scope: docs Documentation label Jul 22, 2026

@ironloopai ironloopai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ IronLoop Review: reviewer

Review at a glance

Verdict Blocking Notes Inline Head
⚠️ Needs validation 0 0 0 e0a598d59f4d

Head: e0a598d59f4dc53cff9cb8f4b64227761654dcdb
Next: Human review or validation is required before merging.

Run details

Status: Current
Needs human: no
Needs validation: yes

Summary

Scoped review of the 3-file, 332-addition E2E fixture change found no concrete correctness or security defect. Runtime validation is still required because this environment lacks Python, the E2E virtualenv, and the built Reborn binary.

Findings

None.

Developer follow-up

After fixing this feedback:

  1. Push the fix to this PR branch.
  2. Re-run this reviewer with @ironloopai review --agent reviewer if you only changed this reviewer's findings.
  3. Re-run all reviewers with @ironloopai review when the fix may affect multiple areas.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@tests/e2e/scenarios/test_reborn_qa_trace_full_path.py`:
- Around line 982-1016: The Slack replay isolation logic in
_cleanup_slack_provider_mutations must include threaded messages: at
tests/e2e/scenarios/test_reborn_qa_trace_full_path.py:982-1016, enumerate thread
sends via the same conversations.replies pagination path used at
tests/e2e/scenarios/test_reborn_qa_trace_full_path.py:952-979, then match and
remove them with chat.delete; update the baseline-absence checks in the sibling
range as needed, while preserving existing non-thread history handling.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: ff3c4f0c-d1f4-4e92-93c0-dbf27f8ce436

📥 Commits

Reviewing files that changed from the base of the PR and between 3c51d17 and e0a598d.

📒 Files selected for processing (3)
  • tests/e2e/CLAUDE.md
  • tests/e2e/conftest.py
  • tests/e2e/scenarios/test_reborn_qa_trace_full_path.py

Comment thread tests/e2e/scenarios/test_reborn_qa_trace_full_path.py
@github-actions

github-actions Bot commented Jul 22, 2026 •

Copy link
Copy Markdown
Contributor

Coverage ratchet

Ratchet mode: ENFORCING

RATCHET PASS: global
  observed: 86.32% (309582 / 358644 lines)
  floor:    86.27% (tolerance 0.5pp -> effective floor 85.77%)
  denominator: 358644 lines now vs 354049 at floor capture (+4595 lines, +1.3%) — not a material change

⚠️ 2 Reborn crate(s) have 0 int-tier coverage (target: 0) — ironclaw_prompt_envelope, ironclaw_scripts

Reborn integration-tier coverage

Line coverage (Reborn crates): 86.32% — 309582 / 358644 lines

Per-crate breakdown (62 crates, lowest-covered first)
Crate Line % Covered / Total
ironclaw_prompt_envelope 0% 0 / 88
ironclaw_scripts 0% 0 / 345
ironclaw_telegram_extension 34.88% 60 / 172
ironclaw_event_projections 43.31% 673 / 1554
ironclaw_product_context 57.78% 26 / 45
ironclaw_dispatcher 60% 72 / 120
ironclaw_observability 61.54% 16 / 26
ironclaw_authorization 62.98% 609 / 967
ironclaw_memory 69.2% 773 / 1117
ironclaw_trust 72.88% 661 / 907
ironclaw_filesystem 73.5% 4543 / 6181
ironclaw_capabilities 74.07% 2717 / 3668
ironclaw_wasm_limiter 74.6% 47 / 63
ironclaw_extractors 74.72% 538 / 720
ironclaw_projects 76.48% 400 / 523
ironclaw_triggers 77.33% 2531 / 3273
ironclaw_mcp 77.56% 736 / 949
ironclaw_reborn_cli 78.25% 10379 / 13264
ironclaw_llm 78.63% 20824 / 26485
ironclaw_wasm 79.72% 735 / 922
ironclaw_process_sandbox 80.46% 671 / 834
ironclaw_memory_native 81.17% 3195 / 3936
ironclaw_events 81.95% 1594 / 1945
ironclaw_first_party_extensions 82.34% 6620 / 8040
ironclaw_reborn_event_store 83.03% 1169 / 1408
ironclaw_telegram_v2_adapter 83.07% 2017 / 2428
ironclaw_product_adapter_registry 83.43% 574 / 688
ironclaw_processes 83.78% 940 / 1122
ironclaw_reborn_identity 83.8% 450 / 537
ironclaw_secrets 83.8% 2550 / 3043
ironclaw_reborn_config 84.17% 1962 / 2331
ironclaw_common 84.48% 1769 / 2094
ironclaw_product_adapters 84.65% 3308 / 3908
ironclaw_auth 85.01% 4011 / 4718
ironclaw_run_state 85.61% 458 / 535
ironclaw_network 85.97% 913 / 1062
ironclaw_product_workflow 86.51% 16271 / 18809
ironclaw_hooks 86.6% 9931 / 11468
ironclaw_extensions 87.02% 3795 / 4361
ironclaw_threads 87.2% 4851 / 5563
ironclaw_host_api 87.55% 5407 / 6176
ironclaw_skills 87.58% 4470 / 5104
ironclaw_reborn_traces 88.11% 11972 / 13587
ironclaw_reborn_composition 88.17% 57298 / 64987
ironclaw_turns 88.63% 14581 / 16452
ironclaw_host_runtime 88.66% 18459 / 20819
ironclaw_reborn_openai_compat 88.92% 3580 / 4026
ironclaw_slack_extension 89.18% 2439 / 2735
ironclaw_extension_host 89.59% 2856 / 3188
ironclaw_approvals 90.18% 1598 / 1772
ironclaw_conversations 90.39% 3123 / 3455
ironclaw_webui 90.42% 9102 / 10066
ironclaw_resources 90.85% 4477 / 4928
ironclaw_event_streams 91.24% 1063 / 1165
ironclaw_runner 91.27% 17073 / 18707
ironclaw_loop_host 92.26% 16260 / 17624
ironclaw_attachments 93.06% 630 / 677
ironclaw_outbound 93.98% 3733 / 3972
ironclaw_agent_loop 94.93% 9840 / 10365
ironclaw_safety 95.15% 3749 / 3940
ironclaw_first_party_extension_ports 95.62% 3672 / 3840
ironclaw_runtime_policy 96.55% 811 / 840

This table itself is informational and never gates the PR on its own — not the percentage, not the per-crate holes, not the 0-coverage callout. A separate coverage ratchet (dry-run until enforce=true; see tests/integration/coverage-floor.toml) can fail the build on specific configured floors.

Exemptions (3 entry/entries excluded from the accounting above)
Module / Crate Reason Issue
crate: ironclaw_embeddings v1-only: consumed only by root ironclaw (src/app.rs, src/tools/builtin/memory.rs, src/workspace/mod.rs, src/config/{mod,embeddings}.rs); no crates/* dependents. Covered by "Tests (Legacy)". #5657
crate: ironclaw_gateway v1-only: consumed only by root ironclaw (src/channels/web/platform/static_files.rs, src/channels/web/handlers/frontend.rs); no crates/* dependents. Covered by "Tests (Legacy)". #5657
crate: ironclaw_tui v1-only: consumed only by root ironclaw (src/main.rs, src/channels/tui.rs); no crates/* dependents. Crate's own doc comment confirms it bridges INTO v1, not Reborn. Covered by "Tests (Legacy)". #5657

@railway-app

railway-app Bot commented Jul 22, 2026 •

Copy link
Copy Markdown

🚅 Deployed to the ironclaw-pr-6525 environment in ironclaw-ci-preview

Service Status Web Updated (UTC)
ironclaw ✅ Success (View Logs) Web Jul 22, 2026 at 10:36 pm

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6525 July 22, 2026 22:27 Destroyed
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6525 July 23, 2026 07:23 Destroyed
@serrrfirat
serrrfirat merged commit 68b4ded into main Jul 23, 2026
61 checks passed
@serrrfirat
serrrfirat deleted the codex/isolate-emulate-provider-worlds branch July 23, 2026 07:36
serrrfirat added a commit that referenced this pull request Jul 28, 2026
* test(e2e): replay the provider journeys in reverse order nightly

Workstream 3 of #6524 owes four isolation proofs: a journey must pass
alone, twice consecutively, after another mutating journey, and in
reversed order. #6525 landed the doubled-repeat arm for two journeys.
This adds the reversed arm.

Reversing is the cheapest arrangement that puts every case somewhere it
has never run, so it is the arm that actually catches state leaking
between journeys through the shared Google/Slack/GitHub worlds — a case
that only passes because an earlier one left the world in a convenient
shape fails here and nowhere else.

Ordering is a parameter of `provider_journey_runs(reverse=...)` rather
than a hidden environment read, because the failure mode to design out
is a "reversed" lane that silently runs forward: it would pass exactly
like the ordinary lane and quietly retire the proof it was added to
provide. So the reversal is asserted directly — order flipped, multiset
unchanged, and more than one run so the comparison cannot pass
vacuously — plus the env parsing that CI actually sets. Verified by
sabotage: making `reverse=True` a no-op fails the new test.

Wired through the existing `reborn-e2e` workflow_call rather than a new
job, so the Emulate setup is reused; off by default and enabled only by
nightly-deep-ci, since it doubles the slowest lane.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(e2e): make the reverse replay share one provider world

Review caught that the original lane could not fail. The ordinary
parametrized replay runs behind `reborn_qa_emulate_provider_server`, a
function-scoped fixture whose `finally` cleans Slack and resets every
other mutated provider after each case. Every journey therefore starts
from seed state regardless of position, so reversing that list could not
expose cross-journey leakage — there was no surviving state to leak. The
lane was a no-op wearing the shape of a proof, which is exactly the
failure the guards in this PR were supposed to rule out; they proved the
reversal was configured, not that reversing could detect anything.

The reverse proof now lives in its own scenario that drives
`reborn_qa_emulate_runtime` directly, bypassing the per-case reset
fixture: it replays only the mutating journeys, back to front, keeping
every prior provider effect in place, and cleans up once after the whole
sequence. A journey that only passes because an earlier one left the
world in a convenient shape now fails there and nowhere else.

The scenario also fails closed when `IRONCLAW_JOURNEY_ORDER` is not set
to reverse, so a workflow edit cannot quietly downgrade the expensive
lane back into a forward replay.

Ordinary replay ordering is unchanged and stays per-case isolated —
reversing it was never meaningful. `pytest.mark.parametrize` values are
a list of tuples per the repo lint invariant (PT007).

Sabotage-verified: making the shared-world reversal a no-op fails
`test_shared_world_replay_reverses_each_mutating_journey_once`.
Reverted; 33 pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(e2e): keep the shared-world replay out of the ordinary lane

Two CI failures from the previous commit, both real.

The shared-world scenario lives in test_reborn_qa_trace_full_path.py,
which the ordinary PR step runs whole. That step does not set
IRONCLAW_JOURNEY_ORDER, so the scenario's fail-closed assert fired on
every PR. Deselect it there by node id rather than softening the assert
to a skip: the assert is what stops a workflow edit from quietly
downgrading the nightly proof into a forward replay, and a skip would
look identical to a pass in exactly that case.

Extracting the replay body into `_replay_qa_journey_provider_leg` also
broke the journey-evidence checker, which requires the declared test to
call its readback helper — the calls now sit one hop away. The checker
now follows a single level of module-level delegation, but only through
an *awaited* call: an un-awaited coroutine never runs, and arbitrary
depth would let a readback buried behind several hops vouch for itself
where no reviewer can see it at the call site. Following nothing is the
default, so an omitted argument makes the check stricter, not looser.

Three self-tests pin that: an awaited delegate is accepted, an
un-awaited one is rejected, and two levels are rejected.

54 tests pass across the two affected files.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(e2e): filter the shared-world replay by marker, not --deselect

The `--deselect` added in the previous commit silently did nothing. It
matched when the step passed a single file, which is how I checked it,
and stopped matching once the real step passed six — so CI ran the
shared-world scenario in the ordinary lane anyway and it failed its
fail-closed assert exactly as before.

That is the same failure shape this PR keeps circling: a guard that
looks applied, reports nothing, and changes no behaviour. The local
check was too narrow to see it.

A marker does not depend on nodeid path resolution. The scenario is
marked `shared_world`, the marker is registered in pyproject so an
unknown-marker typo cannot pass silently, and the ordinary lane filters
with `-m "not shared_world"`. The nightly lane still selects it by node
id, and its fail-closed assert is untouched.

Verified both directions with the full six-file list the step actually
uses: the ordinary lane collects zero shared-world tests, and the
nightly node id still collects exactly one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
l3ocifer pushed a commit to l3ocifer/frick-ironclaw that referenced this pull request Sep 3, 2026
…rai#6728)

* test(e2e): replay the provider journeys in reverse order nightly

Workstream 3 of nearai#6524 owes four isolation proofs: a journey must pass
alone, twice consecutively, after another mutating journey, and in
reversed order. nearai#6525 landed the doubled-repeat arm for two journeys.
This adds the reversed arm.

Reversing is the cheapest arrangement that puts every case somewhere it
has never run, so it is the arm that actually catches state leaking
between journeys through the shared Google/Slack/GitHub worlds — a case
that only passes because an earlier one left the world in a convenient
shape fails here and nowhere else.

Ordering is a parameter of `provider_journey_runs(reverse=...)` rather
than a hidden environment read, because the failure mode to design out
is a "reversed" lane that silently runs forward: it would pass exactly
like the ordinary lane and quietly retire the proof it was added to
provide. So the reversal is asserted directly — order flipped, multiset
unchanged, and more than one run so the comparison cannot pass
vacuously — plus the env parsing that CI actually sets. Verified by
sabotage: making `reverse=True` a no-op fails the new test.

Wired through the existing `reborn-e2e` workflow_call rather than a new
job, so the Emulate setup is reused; off by default and enabled only by
nightly-deep-ci, since it doubles the slowest lane.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(e2e): make the reverse replay share one provider world

Review caught that the original lane could not fail. The ordinary
parametrized replay runs behind `reborn_qa_emulate_provider_server`, a
function-scoped fixture whose `finally` cleans Slack and resets every
other mutated provider after each case. Every journey therefore starts
from seed state regardless of position, so reversing that list could not
expose cross-journey leakage — there was no surviving state to leak. The
lane was a no-op wearing the shape of a proof, which is exactly the
failure the guards in this PR were supposed to rule out; they proved the
reversal was configured, not that reversing could detect anything.

The reverse proof now lives in its own scenario that drives
`reborn_qa_emulate_runtime` directly, bypassing the per-case reset
fixture: it replays only the mutating journeys, back to front, keeping
every prior provider effect in place, and cleans up once after the whole
sequence. A journey that only passes because an earlier one left the
world in a convenient shape now fails there and nowhere else.

The scenario also fails closed when `IRONCLAW_JOURNEY_ORDER` is not set
to reverse, so a workflow edit cannot quietly downgrade the expensive
lane back into a forward replay.

Ordinary replay ordering is unchanged and stays per-case isolated —
reversing it was never meaningful. `pytest.mark.parametrize` values are
a list of tuples per the repo lint invariant (PT007).

Sabotage-verified: making the shared-world reversal a no-op fails
`test_shared_world_replay_reverses_each_mutating_journey_once`.
Reverted; 33 pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(e2e): keep the shared-world replay out of the ordinary lane

Two CI failures from the previous commit, both real.

The shared-world scenario lives in test_reborn_qa_trace_full_path.py,
which the ordinary PR step runs whole. That step does not set
IRONCLAW_JOURNEY_ORDER, so the scenario's fail-closed assert fired on
every PR. Deselect it there by node id rather than softening the assert
to a skip: the assert is what stops a workflow edit from quietly
downgrading the nightly proof into a forward replay, and a skip would
look identical to a pass in exactly that case.

Extracting the replay body into `_replay_qa_journey_provider_leg` also
broke the journey-evidence checker, which requires the declared test to
call its readback helper — the calls now sit one hop away. The checker
now follows a single level of module-level delegation, but only through
an *awaited* call: an un-awaited coroutine never runs, and arbitrary
depth would let a readback buried behind several hops vouch for itself
where no reviewer can see it at the call site. Following nothing is the
default, so an omitted argument makes the check stricter, not looser.

Three self-tests pin that: an awaited delegate is accepted, an
un-awaited one is rejected, and two levels are rejected.

54 tests pass across the two affected files.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(e2e): filter the shared-world replay by marker, not --deselect

The `--deselect` added in the previous commit silently did nothing. It
matched when the step passed a single file, which is how I checked it,
and stopped matching once the real step passed six — so CI ran the
shared-world scenario in the ordinary lane anyway and it failed its
fail-closed assert exactly as before.

That is the same failure shape this PR keeps circling: a guard that
looks applied, reports nothing, and changes no behaviour. The local
check was too narrow to see it.

A marker does not depend on nodeid path resolution. The scenario is
marked `shared_world`, the marker is registered in pyproject so an
unknown-marker typo cannot pass silently, and the ordinary lane filters
with `-m "not shared_world"`. The nightly lane still selects it by node
id, and its fail-closed assert is untouched.

Verified both directions with the full six-file list the step actually
uses: the ordinary lane collects zero shared-world tests, and the
nightly node id still collects exactly one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

This branch was successfully deployed

No deployments
ironclaw-ci-preview / ironclaw-pr-6525 — c160bbde Deployed Jul 23, 2026 by railway-app[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: low Changes to docs, tests, or low-risk modules scope: docs Documentation size: XS < 10 changed lines (excluding docs)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant