Skip to content

test(e2e): inventory provider capability coverage - #6526

Merged
serrrfirat merged 4 commits into
mainfrom
codex/capability-coverage-inventory
Jul 23, 2026
Merged

serrrfirat merged 4 commits into
mainfrom
codex/capability-coverage-inventory

Conversation

@serrrfirat

@serrrfirat serrrfirat commented Jul 22, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • Derive the static provider-capability denominator from shipped first-party extension manifests.
  • Classify all 123 current capabilities as 31 hermetically tested operations or 92 owned waivers tracked by Epic: Hermetic capability and journey testing platform #6524; make unsupported and live-only claims explicit rather than inferred.
  • Fail CI when a shipped capability is missing/stale, classified twice, lacks waiver ownership, or is marked tested without harvested or typed full-path evidence.
  • Drive full-path Emulate selection from the same tested registry and cover missing operations through reusable typed provider-operation cases, while retaining the existing single-build E2E lane.

Change Type

  • Bug fix
  • New feature
  • Refactor
  • Documentation
  • CI/Infrastructure
  • Security
  • Dependencies

Linked Issue

Part of #6524

Validation

  • cargo fmt --all -- --check — Not applicable: no Rust files changed by this PR.
  • cargo clippy --all --benches --tests --examples --all-features -- -D warnings — Not applicable: no Rust files changed by this PR.
  • cargo build — Reborn ironclaw binary built successfully after merging current main.
  • Relevant tests pass: 102 inventory, provider-contract, recorded-replay, isolation, and full-path tests.
  • cargo test --features integration if database-backed or integration behavior changed — Not applicable: no database or product integration behavior changed.
  • Manual testing: ran the exact Emulate selection against the pinned fork and verified the inventory reports 123/123 classified capabilities.
  • If a coding agent was used and supports it, review-pr or pr-shepherd --fix was run before requesting review — iterative autoreview completed with no actionable findings before the merge-conflict update.

Test Strategy

User behavior: Developers get an actionable CI failure when the shipped provider surface expands without an explicit test classification, instead of relying on harvested traces to discover coverage gaps later.

Risk areas:

  • Model behavior
  • Browser
  • Side effect
  • Persistence
  • Security or permissions
  • External provider
  • Cross-component behavior

Tests added or updated:

  • Unit or contract: test_provider_capability_inventory.py proves exact manifest/inventory equality, non-overlapping classifications, waiver metadata, and executable evidence for tested operations.
  • Reborn integration: Existing full-path QA journeys and typed provider-operation cases execute through the standalone Reborn runtime.
  • Recorded fixture: The fixture closed-set check detects unknown or misspelled operations while allowing explicitly classified operations.
  • Browser E2E: Not applicable: no browser behavior changed.
  • Backend or runtime: The complete Emulate provider-contract and full-path selection passes using the same tested inventory.
  • Live canary: Not applicable: this PR makes no live-only coverage claims; operations without hermetic evidence remain owned waivers.

What the tests prove: Every statically shipped first-party provider capability is classified exactly once; adding or renaming one fails CI until its coverage status is owned. A capability cannot be marked tested merely by editing metadata: it must have either a harvested full-path call or a typed operation case, and the same registry drives full-path execution.

Commands run:

./tests/e2e/.venv/bin/python -m py_compile \
  tests/e2e/provider_capability_inventory.py \
  tests/e2e/provider_operation_cases.py \
  tests/e2e/scenarios/test_provider_capability_inventory.py \
  tests/e2e/scenarios/test_reborn_qa_trace_replay.py \
  tests/e2e/scenarios/test_reborn_qa_trace_full_path.py

./tests/e2e/.venv/bin/pytest -q \
  tests/e2e/scenarios/test_provider_capability_inventory.py \
  tests/e2e/scenarios/test_reborn_qa_trace_replay.py \
  --timeout=240

bash scripts/ci/check-reborn-qa-fixtures.sh

CARGO_TARGET_DIR=/Volumes/NVME/ironclaw/target \
  cargo build -p ironclaw --bin ironclaw

IRONCLAW_EMULATE_CLI=/Volumes/NVME/codex-bac3-emulate/packages/emulate/dist/index.js \
CARGO_TARGET_DIR=/Volumes/NVME/ironclaw/target \
  ./tests/e2e/.venv/bin/pytest -q \
  tests/e2e/scenarios/test_provider_capability_inventory.py \
  tests/e2e/scenarios/test_emulate_reborn_provider_contracts.py \
  tests/e2e/scenarios/test_reborn_qa_trace_replay.py \
  tests/e2e/scenarios/test_reborn_qa_trace_full_path.py \
  --timeout=240

git diff --check

Results:

  • Fast inventory/replay checks: 51 passed in 1.37s
  • Fixture scrub validator: 61 files
  • Full Emulate lane: 102 passed in 79.02s
  • Reborn build: passed (one warning already present on current main)

Security Impact

None. This is test metadata and validation only; it does not alter production permissions, credentials, provider traffic, file access, tool execution, or sandbox policy.

Reborn Trust-Boundary Checklist

N/A: test-harness-only change; no trust-bearing production types, runtime behavior, serialization, queues, errors, or sandbox boundaries changed.

Database Impact

None.

Blast Radius

Limited to the harvested-QA/Emulate E2E lane. Incorrect manifest parsing or classification could fail that lane; product binaries and runtime behavior are unchanged. Runtime impact is lightweight inventory validation plus four typed cases inside the existing single-build provider process.

Rollback Plan

Revert the inventory and provider-operation commits to restore trace-local Emulate classification. No data migration or compatibility step is required.

Review Follow-Through

The 92 waivers are intentionally visible debt owned by #6524, not claims of coverage. Follow-up provider-operation PRs should move capabilities from the waiver list to tested only when their executable Emulate/full-path evidence lands.


Review track: C (CI/test infrastructure)

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Caution

The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased.

@ironloopai

ironloopai Bot commented Jul 22, 2026 •

Copy link
Copy Markdown
Contributor

🔎 IronLoop Review Status

Head: 24d76541b345c08bec1003a07ef4affef1278fec
Result: One or more review results were superseded by a newer PR head.
Next: Run @ironloopai review on the latest PR head.
Updated: 2026-07-23T07:49:28.846Z

Current reviewers:

Reviewer State Verdict Findings Last update
ironloop/common-reviewer (reviewer) Superseded N/A N/A 2026-07-23T07:49:28.834Z
Reviewer summaries
Reviewer Detail
ironloop/common-reviewer (reviewer) Superseded by a newer PR head. New head: 24d7654. Previous verdict: Approved.
Recent activity
Time Reviewer State Detail
2026-07-22T22:02:30.941Z ironloop/common-reviewer (reviewer) Queued Accepted review request for head f98251c.
2026-07-22T22:02:30.941Z ironloop/common-reviewer (reviewer) Queued Waiting for this reviewer lane to become available.
2026-07-22T22:02:31.782Z ironloop/common-reviewer (reviewer) Started Reviewer worker started.
2026-07-22T22:02:34.269Z ironloop/common-reviewer (reviewer) Workspace ready Prepared isolated checkout (merge_ref) at 2980000.
2026-07-22T22:05:42.817Z ironloop/common-reviewer (reviewer) Result captured Approved; 0 blocking findings.
2026-07-22T22:05:42.817Z ironloop/common-reviewer (reviewer) Completed Review completed and terminal status was persisted.
2026-07-23T07:49:28.834Z ironloop/common-reviewer (reviewer) Superseded A newer PR head replaced this review (24d7654).
Available commands
  • @ironloopai help
  • @ironloopai agents
  • @ironloopai review
  • @ironloopai review --agent <agent>
Run metadata

Admission: webhook accepted the request and IronLoop persisted reviewer state before this projection.

@coderabbitai

coderabbitai Bot commented Jul 22, 2026 •

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Walkthrough

Adds a manifest-backed provider capability inventory, typed Google operation cases with provider readback, exact-name Emulate trace filtering, and two completeness tests. The inventory scenario is wired into the WebUI v2 smoke workflow and documented in the E2E coverage guidance.

Changes

Provider capability coverage

Layer / File(s) Summary
Capability inventory contract
tests/e2e/fixtures/provider_capability_coverage.toml, tests/e2e/provider_capability_inventory.py, tests/e2e/scenarios/test_provider_capability_inventory.py
Classifies capabilities, derives wire names from shipped manifests, and validates exclusive ownership, manifest parity, waiver metadata, and recorded or typed-operation evidence.
Typed provider operation cases
tests/e2e/provider_operation_cases.py, tests/e2e/CLAUDE.md
Defines four Google Drive/Gmail cases with typed arguments and baseline/outcome assertions for provider readback.
Emulate trace execution
tests/e2e/scenarios/test_reborn_qa_trace_full_path.py, tests/e2e/scenarios/test_reborn_qa_trace_replay.py
Uses exact classified provider tool names, adds resettable operation-case execution, and validates completed Reborn previews plus provider outcomes.
Coverage gate wiring
.github/workflows/reborn-e2e.yml, tests/e2e/CLAUDE.md
Runs the capability inventory scenario in the WebUI v2 smoke job and lists it in the Reborn E2E coverage gate.

Estimated code review effort: 4 (Complex) | ~45 minutes

Sequence Diagram(s)

sequenceDiagram
  participant E2ESmoke
  participant Reborn
  participant MockLLM
  participant Emulate
  participant ProviderOracle
  E2ESmoke->>Reborn: execute provider operation case
  Reborn->>MockLLM: request inline trace response
  MockLLM-->>Reborn: capability_info and provider tool call
  Reborn->>Emulate: execute provider tool
  Emulate-->>Reborn: operation result
  Reborn-->>ProviderOracle: completed preview and replay metadata
  ProviderOracle->>Emulate: verify provider state
Loading

Possibly related PRs

  • nearai/ironclaw#6439: Introduced the harvested trace replay scenario and its provider-tool validation.
  • nearai/ironclaw#6466: Changed the shared full-path harness’s provider trace filtering and loading flow.
  • nearai/ironclaw#6525: Added the resettable Emulate provider fixture used by the replay harness.

Suggested reviewers: think-in-universe

🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title follows Conventional Commits style and accurately describes the new E2E provider capability inventory work.
Description check ✅ Passed The description matches the repository template and fills the required summary, linkage, validation, test strategy, and risk sections.

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6526 July 22, 2026 22:02 Destroyed
@github-actions github-actions Bot added scope: ci CI/CD workflows scope: docs Documentation size: XS < 10 changed lines (excluding docs) risk: medium Business logic, config, or moderate-risk modules contributor: core 20+ merged PRs labels Jul 22, 2026

@ironloopai ironloopai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ IronLoop Review: reviewer

Review at a glance

Verdict Blocking Notes Inline Head
✅ Approved 0 0 0 f98251cb6dba

Head: f98251cb6dbab2d3741af7d2484f60c53b216d54
Next: No reviewer action needed.

Run details

Status: Current
Needs human: no
Needs validation: no

Summary

Approved. Reviewed the focused stacked-layer change (1 commit; 7 files; 309 additions, 63 deletions) across the CI wiring, inventory source, helper, and replay/full-path consumers. The manifest inventory, trace evidence gate, and executable full-path selection are consistent.

Findings

None.

Developer follow-up

After fixing this feedback:

  1. Push the fix to this PR branch.
  2. Re-run this reviewer with @ironloopai review --agent reviewer if you only changed this reviewer's findings.
  3. Re-run all reviewers with @ironloopai review when the fix may affect multiple areas.

@railway-app

railway-app Bot commented Jul 22, 2026 •

Copy link
Copy Markdown

🚅 Deployed to the ironclaw-pr-6526 environment in ironclaw-ci-preview

Service Status Web Updated (UTC)
ironclaw ✅ Success (View Logs) Web Jul 23, 2026 at 7:58 am

* test(e2e): add typed provider operation cases

* test(e2e): address provider operation review
Base automatically changed from codex/isolate-emulate-provider-worlds to main July 23, 2026 07:36
…rage-inventory

# Conflicts:
#	tests/e2e/scenarios/test_reborn_qa_trace_full_path.py
@railway-app
railway-app Bot temporarily deployed to ironclaw-ci-preview / ironclaw-pr-6526 July 23, 2026 07:49 Destroyed
@github-actions

Copy link
Copy Markdown
Contributor

Coverage ratchet

Ratchet mode: ENFORCING

RATCHET PASS: global
  observed: 86.32% (309596 / 358644 lines)
  floor:    86.27% (tolerance 0.5pp -> effective floor 85.77%)
  denominator: 358644 lines now vs 354049 at floor capture (+4595 lines, +1.3%) — not a material change

⚠️ 2 Reborn crate(s) have 0 int-tier coverage (target: 0) — ironclaw_prompt_envelope, ironclaw_scripts

Reborn integration-tier coverage

Line coverage (Reborn crates): 86.32% — 309596 / 358644 lines

Per-crate breakdown (62 crates, lowest-covered first)
Crate Line % Covered / Total
ironclaw_prompt_envelope 0% 0 / 88
ironclaw_scripts 0% 0 / 345
ironclaw_telegram_extension 34.88% 60 / 172
ironclaw_event_projections 43.31% 673 / 1554
ironclaw_product_context 57.78% 26 / 45
ironclaw_dispatcher 60% 72 / 120
ironclaw_observability 61.54% 16 / 26
ironclaw_authorization 62.98% 609 / 967
ironclaw_memory 69.2% 773 / 1117
ironclaw_trust 72.88% 661 / 907
ironclaw_filesystem 73.5% 4543 / 6181
ironclaw_capabilities 74.07% 2717 / 3668
ironclaw_wasm_limiter 74.6% 47 / 63
ironclaw_extractors 74.72% 538 / 720
ironclaw_projects 76.48% 400 / 523
ironclaw_triggers 77.33% 2531 / 3273
ironclaw_mcp 77.56% 736 / 949
ironclaw_reborn_cli 78.25% 10379 / 13264
ironclaw_llm 78.63% 20824 / 26485
ironclaw_wasm 79.72% 735 / 922
ironclaw_process_sandbox 80.46% 671 / 834
ironclaw_memory_native 81.17% 3195 / 3936
ironclaw_events 81.95% 1594 / 1945
ironclaw_first_party_extensions 82.34% 6620 / 8040
ironclaw_reborn_event_store 83.03% 1169 / 1408
ironclaw_telegram_v2_adapter 83.07% 2017 / 2428
ironclaw_product_adapter_registry 83.43% 574 / 688
ironclaw_processes 83.78% 940 / 1122
ironclaw_reborn_identity 83.8% 450 / 537
ironclaw_secrets 83.8% 2550 / 3043
ironclaw_reborn_config 84.17% 1962 / 2331
ironclaw_common 84.48% 1769 / 2094
ironclaw_product_adapters 84.65% 3308 / 3908
ironclaw_auth 85.01% 4011 / 4718
ironclaw_run_state 85.61% 458 / 535
ironclaw_network 85.97% 913 / 1062
ironclaw_product_workflow 86.51% 16271 / 18809
ironclaw_hooks 86.6% 9931 / 11468
ironclaw_extensions 87.02% 3795 / 4361
ironclaw_threads 87.2% 4851 / 5563
ironclaw_host_api 87.55% 5407 / 6176
ironclaw_skills 87.58% 4470 / 5104
ironclaw_reborn_traces 88.11% 11972 / 13587
ironclaw_reborn_composition 88.17% 57298 / 64987
ironclaw_host_runtime 88.66% 18459 / 20819
ironclaw_turns 88.71% 14595 / 16452
ironclaw_reborn_openai_compat 88.92% 3580 / 4026
ironclaw_slack_extension 89.18% 2439 / 2735
ironclaw_extension_host 89.59% 2856 / 3188
ironclaw_approvals 90.18% 1598 / 1772
ironclaw_conversations 90.39% 3123 / 3455
ironclaw_webui 90.42% 9102 / 10066
ironclaw_resources 90.85% 4477 / 4928
ironclaw_event_streams 91.24% 1063 / 1165
ironclaw_runner 91.27% 17073 / 18707
ironclaw_loop_host 92.26% 16260 / 17624
ironclaw_attachments 93.06% 630 / 677
ironclaw_outbound 93.98% 3733 / 3972
ironclaw_agent_loop 94.93% 9840 / 10365
ironclaw_safety 95.15% 3749 / 3940
ironclaw_first_party_extension_ports 95.62% 3672 / 3840
ironclaw_runtime_policy 96.55% 811 / 840

This table itself is informational and never gates the PR on its own — not the percentage, not the per-crate holes, not the 0-coverage callout. A separate coverage ratchet (dry-run until enforce=true; see tests/integration/coverage-floor.toml) can fail the build on specific configured floors.

Exemptions (3 entry/entries excluded from the accounting above)
Module / Crate Reason Issue
crate: ironclaw_embeddings v1-only: consumed only by root ironclaw (src/app.rs, src/tools/builtin/memory.rs, src/workspace/mod.rs, src/config/{mod,embeddings}.rs); no crates/* dependents. Covered by "Tests (Legacy)". #5657
crate: ironclaw_gateway v1-only: consumed only by root ironclaw (src/channels/web/platform/static_files.rs, src/channels/web/handlers/frontend.rs); no crates/* dependents. Covered by "Tests (Legacy)". #5657
crate: ironclaw_tui v1-only: consumed only by root ironclaw (src/main.rs, src/channels/tui.rs); no crates/* dependents. Crate's own doc comment confirms it bridges INTO v1, not Reborn. Covered by "Tests (Legacy)". #5657

@serrrfirat
serrrfirat merged commit 495d814 into main Jul 23, 2026
61 checks passed
@serrrfirat
serrrfirat deleted the codex/capability-coverage-inventory branch July 23, 2026 08:09

This branch was successfully deployed

No deployments
ironclaw-ci-preview / ironclaw-pr-6526 — 24d76541 Deployed Jul 23, 2026 by railway-app[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: medium Business logic, config, or moderate-risk modules scope: ci CI/CD workflows scope: docs Documentation size: XS < 10 changed lines (excluding docs)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant