ci(spike): experimental full-suite gate - nextest archive + mold + sccache + sharding - #5086
serrrfirat wants to merge 5 commits into
Conversation
|
Note Gemini is unable to generate a review for this pull request due to the file types involved not being currently supported. |
|
🚅 Deployed to the ironclaw-pr-5086 environment in ironclaw-ci-preview
|
…robe The spike failed at the mold-verification step before any build ran. The rustc probe grepped the linker output for "mold", but rustc swallows the linker stdout on a successful link, so that check could never pass even with mold working. Switch to explicit clang --ld-path=/usr/bin/mold (no ld.mold name resolution, no silent GNU ld fallback) and assert the rustc->clang->mold link produces a runnable binary instead. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
📝 WalkthroughWalkthroughAdds two new GitHub Actions workflows: ChangesExperimental Full Suite Measurement Workflow
Experimental Compile Benchmark Workflow
Sequence DiagramssequenceDiagram
participant Runner as Build Runner
participant Mold as mold Linker
participant sccache as sccache
participant nextest as cargo nextest
participant Postgres as Postgres Service
participant Artifacts as GitHub Artifacts
rect rgba(0, 100, 200, 0.5)
Note over Runner: build-archive job
Runner->>Mold: verify clang --ld-path=mold linkage
Mold-->>Runner: confirmed
Runner->>sccache: start and restore cache
Runner->>nextest: cargo nextest archive --all-targets
nextest-->>Runner: success or failure
alt archive failed
Runner->>nextest: retry without --all-targets
end
nextest-->>Artifacts: upload nextest-LABEL.tar.zst
end
rect rgba(100, 150, 50, 0.5)
Note over Runner,Postgres: test-shard job (×4 shards)
Artifacts->>Runner: download nextest archive
Runner->>Postgres: start service
Runner->>nextest: cargo nextest run --partition count:N/4
nextest-->>Runner: JSONL output with exec_time per test
Runner->>Artifacts: upload per-shard measurement JSON
end
rect rgba(200, 100, 100, 0.5)
Note over Runner,Artifacts: report job
Artifacts->>Runner: download all build + shard artifacts
Runner->>Runner: aggregate durations, compute max/p50, derive wall-clock
Runner->>Runner: compare against BUDGET_SECONDS → PASS/FAIL
Runner->>Runner: write Markdown table to GITHUB_STEP_SUMMARY
end
sequenceDiagram
participant Runner as Compile Bench Runner
participant Rust as Rust Toolchain
participant Backend as Backend Selector
participant nextest as cargo nextest
participant BuildLog as Build Logs
participant Artifacts as GitHub Artifacts
rect rgba(100, 150, 200, 0.5)
Note over Runner,Rust: Setup (both backends)
Runner->>Rust: install wasm32-wasip2, cargo-nextest
Runner->>Backend: verify mold linkage
Backend-->>Runner: confirmed
end
rect rgba(200, 150, 100, 0.5)
Note over Runner,Backend: Matrix arm: LLVM vs Cranelift
alt llvm-baseline on stable
Runner->>Backend: probe LLVM version
else cranelift on nightly
Runner->>Backend: install cranelift component, write .cargo/config.toml
Runner->>Rust: compile probe crate
Rust-->>Backend: verify backend in rustc output
end
Backend-->>Runner: backend status to GITHUB_ENV
end
rect rgba(100, 200, 100, 0.5)
Note over Runner,nextest: Core measurement
Runner->>nextest: cargo nextest archive (under timeout)
nextest-->>BuildLog: compiler output + logs
Runner->>Runner: capture exit status, wall-clock duration
Runner->>Runner: parse build logs → JSON (cache-hit metrics, failing crates)
Runner->>Artifacts: upload measurement JSON
end
rect rgba(200, 100, 200, 0.5)
Note over Runner,BuildLog: LLVM-only timings
Runner->>Rust: cargo build --timings --json=short
Rust-->>BuildLog: timing JSON + HTML
Runner->>Runner: parse timing data → top 15 slow units
Runner->>Artifacts: upload timing HTML/JSON
end
rect rgba(150, 150, 50, 0.5)
Note over Runner,Artifacts: report job
Artifacts->>Runner: download all variant + baseline artifacts
Runner->>Runner: load build JSONs, compute trusted deltas
Runner->>Runner: aggregate top slow units
Runner->>Runner: generate comparison table + Verdict
Runner->>Runner: write to GITHUB_STEP_SUMMARY
end
Estimated code review effort🎯 4 (Complex) | ⏱️ ~60 minutes Possibly related issues
Poem
🚥 Pre-merge checks | ✅ 4✅ Passed checks (4 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. Warning Review ran into problems🔥 ProblemsStopped waiting for pipeline failures after 30000ms. One of your pipelines takes longer than our 30000ms fetch window to run, so review may not consider pipeline-failure results for inline comments if any failures occurred after the fetch window. Increase the timeout if you want to wait longer or run a Comment |
…, measurement-mode shards Two failures, both downstream of a working harness (mold + build-once + shard all proven): 1. libsql-only could not archive: cargo nextest archive --workspace --features libsql errors because Reborn crates expose libsql only via dep:libsql behind their own feature names; no uniform libsql feature exists across the 36 crates. Dropped it (needs cargo-hakari to express). 2. Shards ran the entire heterogeneous workspace suite against one un-migrated Postgres in parallel, so env-dependent tests fail (missing memory_documents relation, pg_type schema races, CLI smoke env). These are environmental, not harness bugs. The shard is now a measurement: green when nextest executes; per-shard pass/fail counts are recorded and surfaced in the report. Making run-everything fully green needs migrations + per-test DB isolation + env — a separate workstream this spike has now quantified. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In @.github/workflows/experimental-full-suite.yml:
- Around line 454-468: The current logic in this shard treats both exit code 0
(all tests passed) and exit code 100 (tests ran with failures) as successful
completions by exiting with status 0. Instead, only exit 0 when the status is
exactly 0 (indicating all shards passed). For the case when status equals 100
(nextest ran but had test failures), exit with a different status code that
indicates an INCOMPLETE or non-blocking outcome rather than a PASS. Keep other
exit codes unchanged to propagate infrastructure errors. Apply this same fix
pattern to all occurrences mentioned in the comment (lines 599-605 and 612-620)
to ensure consistent reporting across all measurement shards.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro Plus
Run ID: 2c60f9dc-2918-4439-8f51-730051db0bdc
📒 Files selected for processing (1)
.github/workflows/experimental-full-suite.yml
| # Measurement spike: this shard runs the ENTIRE heterogeneous workspace | ||
| # suite against one Postgres in parallel, so env-dependent failures | ||
| # (un-migrated schema -> "relation memory_documents does not exist", | ||
| # parallel schema races -> pg_type duplicate key, CLI smoke env) are | ||
| # EXPECTED and are recorded as data (exit_code + pass/fail counts in the | ||
| # JSON, surfaced in the report) — they are not a gate. The job is green | ||
| # when nextest actually executed; it fails only if nextest could not run. | ||
| # nextest exit codes: 0 = all passed, 100 = ran with test failures, | ||
| # others = could-not-run / infra error. | ||
| if [ "${status}" -eq 0 ] || [ "${status}" -eq 100 ]; then | ||
| echo "nextest executed for ${LABEL} shard ${SHARD_INDEX} (exit ${status}); measurement captured." | ||
| exit 0 | ||
| fi | ||
| echo "::error::nextest could not run for ${LABEL} shard ${SHARD_INDEX} (exit ${status})" | ||
| exit "${status}" |
There was a problem hiding this comment.
Don’t emit a 30m PASS from failing shards.
exit_code == 100 means nextest ran with failures; migration/schema/env failures can fail fast, so build + slowest shard is only a lower bound. Keep the job non-blocking, but mark PASS only when all expected shards exit 0; otherwise report INCOMPLETE unless the lower bound already exceeds budget.
Proposed fix
- # Measurement spike: this shard runs the ENTIRE heterogeneous workspace
- # suite against one Postgres in parallel, so env-dependent failures
+ # Measurement spike: this shard runs the ENTIRE heterogeneous workspace
+ # suite against one Postgres in parallel, so known env-dependent failures
# (un-migrated schema -> "relation memory_documents does not exist",
# parallel schema races -> pg_type duplicate key, CLI smoke env) are
- # EXPECTED and are recorded as data (exit_code + pass/fail counts in the
+ # possible and are recorded as data (exit_code + pass/fail counts in the
# JSON, surfaced in the report) — they are not a gate. The job is green
# when nextest actually executed; it fails only if nextest could not run. build_seconds = build.get("build_seconds") if build else None
max_shard = max(shard_seconds) if shard_seconds else None
p50_shard = statistics.median(shard_seconds) if shard_seconds else None
total = build_seconds + max_shard if build_seconds is not None and max_shard is not None else None
budget_verdict = verdict(total)
+ tests_passed = sum(row.get("tests_passed") or 0 for row in shard_rows)
+ tests_failed = sum(row.get("tests_failed") or 0 for row in shard_rows)
+ expected_shards = int(os.environ["SHARDS"])
+ infra_failed_shards = [
+ row for row in shard_rows
+ if row.get("exit_code") not in (0, 100, None)
+ ]
+ test_failed_shards = [
+ row for row in shard_rows
+ if row.get("exit_code") == 100 or (row.get("tests_failed") or 0) > 0
+ ]
+ timing_complete = (
+ len(shard_seconds) == expected_shards
+ and not infra_failed_shards
+ and not test_failed_shards
+ )
+ display_verdict = (
+ "INCOMPLETE"
+ if budget_verdict == "PASS" and not timing_complete
+ else budget_verdict
+ )
+ timing_note = (
+ "Complete timing."
+ if timing_complete
+ else "Timing is a lower bound until every shard exits 0."
+ )
if build:
mold = "yes" if build.get("mold_active") else "no"
parallel = "yes" if build.get("parallel_linking_succeeded_without_oom") else "no/unknown"
scope = build.get("archive_scope") or "n/a"
@@
failed_shards = [row for row in shard_rows if row.get("exit_code") not in (0, None)]
shard_cell = f"{len(shard_seconds)}/{os.environ['SHARDS']}"
if failed_shards:
- shard_cell += f" ({len(failed_shards)} failed)"
+ shard_cell += f" ({len(failed_shards)} with failures)"
out.write(
f"| `{label}` | {scope} | {seconds(build_seconds)} | {hit_rate} | {shard_cell} | "
- f"{seconds(p50_shard)} | {seconds(max_shard)} | {seconds(total)} | **{budget_verdict}** | "
+ f"{seconds(p50_shard)} | {seconds(max_shard)} | {seconds(total)} | **{display_verdict}** | "
f"mold: {mold}; parallel without OOM: {parallel} |\n"
)
- tests_passed = sum(row.get("tests_passed") or 0 for row in shard_rows)
- tests_failed = sum(row.get("tests_failed") or 0 for row in shard_rows)
verdict_lines.append(
- f"- `{label}`: {budget_verdict} for build + slowest shard "
+ f"- `{label}`: {display_verdict} for build + slowest shard "
f"({seconds(total)} against {budget // 60}m budget). "
- f"Tests: {tests_passed} passed, {tests_failed} failed "
- f"(failures are env-dependent in this monolithic run — see notes)."
+ f"Tests: {tests_passed} passed, {tests_failed} failed. "
+ f"{timing_note}"
)- "- Test failures in this run are **environmental**, not harness bugs: the whole "
- "heterogeneous workspace suite runs against one Postgres in parallel, so tests that "
+ "- Test failures in this run are **not attributed by this workflow**. Known "
+ "environment gaps include tests that "
"expect migrations (`relation memory_documents does not exist`), per-test DB isolation "
"(`pg_type` duplicate-key races), or CLI-smoke env fail here. Making 'run everything' "
- "go fully green is a separate env workstream (migrations + per-test DB + env), which "
- "this spike has now quantified.\n"
+ "go fully green is a separate env workstream (migrations + per-test DB + env).\n"Also applies to: 599-605, 612-620
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In @.github/workflows/experimental-full-suite.yml around lines 454 - 468, The
current logic in this shard treats both exit code 0 (all tests passed) and exit
code 100 (tests ran with failures) as successful completions by exiting with
status 0. Instead, only exit 0 when the status is exactly 0 (indicating all
shards passed). For the case when status equals 100 (nextest ran but had test
failures), exit with a different status code that indicates an INCOMPLETE or
non-blocking outcome rather than a PASS. Keep other exit codes unchanged to
propagate infrastructure errors. Apply this same fix pattern to all occurrences
mentioned in the comment (lines 599-605 and 612-620) to ensure consistent
reporting across all measurement shards.
…lvm, mold, non-blocking)
Builds only the ironclaw_reborn_cli crate tree (-p, excludes the root ironclaw v1 monolith that profiling showed is ~55% of compile). Same toolchain/flags as llvm-baseline so the report delta is the win. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In @.github/workflows/experimental-compile-bench.yml:
- Around line 47-57: The reborn-only benchmark configuration asserts that
ironclaw_reborn_cli does NOT depend on the root ironclaw crate, but this
dependency invariant is never verified before the benchmark runs. To fix this,
add a pre-benchmark verification step that runs the cargo tree command
(referenced in the comment: cargo tree -p ironclaw_reborn_cli -i ironclaw) to
confirm the dependency relationship before the reborn-only label's benchmark
executes. This prevents the benchmark from silently measuring a different scope
if the dependency graph drifts in the future.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro Plus
Run ID: d2be7d35-591a-47de-afae-ca2f301b6c13
📒 Files selected for processing (1)
.github/workflows/experimental-compile-bench.yml
| # Measures the compile win of dropping the v1 monolith: build only the | ||
| # Reborn binary's crate tree. `cargo tree -p ironclaw_reborn_cli -i ironclaw` | ||
| # confirms reborn_cli does NOT depend on the root `ironclaw` crate, so this | ||
| # scope skips the root monolith (~55% of workspace compile) entirely. Same | ||
| # toolchain/backend/flags as the baseline, scoped with -p, so the report's | ||
| # delta vs llvm-baseline is exactly the dropping-v1 win. | ||
| - label: reborn-only | ||
| toolchain: stable | ||
| backend: llvm | ||
| runs_on: ubuntu-latest | ||
| cargo_scope: "-p ironclaw_reborn_cli --all-targets --all-features" |
There was a problem hiding this comment.
🧩 Analysis chain
🏁 Script executed:
#!/bin/bash
set -euo pipefail
cargo metadata --format-version=1 --locked > /tmp/ironclaw-cargo-metadata.json
python3 - /tmp/ironclaw-cargo-metadata.json <<'PY'
import json
import sys
with open(sys.argv[1], encoding="utf-8") as f:
metadata = json.load(f)
packages = metadata.get("packages", [])
resolve = metadata.get("resolve") or {}
nodes = {node["id"]: node for node in resolve.get("nodes", [])}
ids_by_name = {}
for package in packages:
ids_by_name.setdefault(package.get("name"), []).append(package.get("id"))
reborn_ids = [pkg_id for pkg_id in ids_by_name.get("ironclaw_reborn_cli", []) if pkg_id]
ironclaw_ids = {pkg_id for pkg_id in ids_by_name.get("ironclaw", []) if pkg_id}
if len(reborn_ids) != 1:
raise SystemExit(f"Expected exactly one ironclaw_reborn_cli package, found {len(reborn_ids)}")
if not ironclaw_ids:
print("OK: no package named ironclaw is present.")
raise SystemExit(0)
seen = set()
stack = list(reborn_ids)
while stack:
current = stack.pop()
if current in seen:
continue
seen.add(current)
stack.extend(dep["pkg"] for dep in nodes.get(current, {}).get("deps", []))
hits = ironclaw_ids & seen
if hits:
raise SystemExit("reborn-only currently reaches root ironclaw: " + ", ".join(sorted(hits)))
print("OK: ironclaw_reborn_cli does not depend on root ironclaw.")
PYRepository: nearai/ironclaw
Length of output: 1895
🏁 Script executed:
# Find and read the workflow file
find . -name "experimental-compile-bench.yml" -type f 2>/dev/null | head -5Repository: nearai/ironclaw
Length of output: 1913
🏁 Script executed:
# Look for Cargo.toml files to understand workspace structure
find . -name "Cargo.toml" -type f 2>/dev/null | head -20Repository: nearai/ironclaw
Length of output: 1913
🏁 Script executed:
# Search for references to ironclaw and ironclaw_reborn_cli in the codebase
rg "ironclaw_reborn_cli|ironclaw" --type yaml --type toml -l 2>/dev/null | head -20Repository: nearai/ironclaw
Length of output: 1913
Add pre-measurement verification that ironclaw_reborn_cli excludes root ironclaw before the reborn-only benchmark runs.
The comment asserts this dependency invariant, but the workflow produces the "dropping-v1" delta without verifying it. If the dependency graph drifts, the benchmark silently measures a different scope.
Suggested fix: add a pre-benchmark step using the provided script or equivalent, or document why drift is not a concern for this workflow.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In @.github/workflows/experimental-compile-bench.yml around lines 47 - 57, The
reborn-only benchmark configuration asserts that ironclaw_reborn_cli does NOT
depend on the root ironclaw crate, but this dependency invariant is never
verified before the benchmark runs. To fix this, add a pre-benchmark
verification step that runs the cargo tree command (referenced in the comment:
cargo tree -p ironclaw_reborn_cli -i ironclaw) to confirm the dependency
relationship before the reborn-only label's benchmark executes. This prevents
the benchmark from silently measuring a different scope if the dependency graph
drifts in the future.
PR CI's reborn-tests matrix was a name-prefix allowlist of 21 reborn/ product-family crates. But only 10 of those overlap the actual `ironclaw_reborn_cli` dependency closure (53 workspace crates) — so 43 crates the shipped Reborn binary links (auth, host_runtime, skills, first_party_extensions, extensions, dispatcher, llm, safety, memory, network, turns, host_api, loop_support, threads, ...) ran their own test suites ONLY via the nightly/manual closure path, never on a PR. They were exercised on PR only indirectly by the 4 root integration partitions, which don't run those crates' own unit/contract tests. That gap is exactly why the bugs fixed in #5105/#5108 (host_runtime github surface, skills TOCTOU, gsuite wrong-account egress, loop_support/ threads/auth) slipped through normal PR CI and only surfaced when the closure was run by hand. This makes the closure the default PR matrix — "run everything on every PR": - package-matrix: discover the union of the existing allowlist and the `cargo tree -p ironclaw_reborn_cli -e normal,build` closure ∩ workspace members (64 crates). Union (not replace) so non-closure reborn-family crates — channel adapters, webui_v2 — are never dropped from coverage. - package-feature-flags.sh: derive fallback features (default/libsql when declared) for closure crates without an explicit recipe; keep the previously-allowlisted no-flag crates flag-free so their behavior is unchanged. The 3 closure reds this would have caught are already fixed on main (#5105, #5108), so the closure should be green — this PR's own CI run is the 64/64 verification. Tradeoff: 21 -> 64 parallel crate jobs raises compute and, if runner concurrency is capped, may raise wall-clock as jobs queue. Follow-ups: (1) build-once `nextest archive` + shard to cut redundant compiles (spike #5086); (2) bake a few runs, then promote reborn-tests to a required check; (3) cut v1 (`test.yml` / Tests (all-features), ~29m) so the gate drops to ~10-12m. Automated agent-authored. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
PR CI's reborn-tests matrix was a name-prefix allowlist of 21 reborn/ product-family crates. But only 10 of those overlap the actual `ironclaw_reborn_cli` dependency closure (53 workspace crates) — so 43 crates the shipped Reborn binary links (auth, host_runtime, skills, first_party_extensions, extensions, dispatcher, llm, safety, memory, network, turns, host_api, loop_support, threads, ...) ran their own test suites ONLY via the nightly/manual closure path, never on a PR. They were exercised on PR only indirectly by the 4 root integration partitions, which don't run those crates' own unit/contract tests. That gap is exactly why the bugs fixed in #5105/#5108 (host_runtime github surface, skills TOCTOU, gsuite wrong-account egress, loop_support/ threads/auth) slipped through normal PR CI and only surfaced when the closure was run by hand. This makes the closure the default PR matrix — "run everything on every PR": - package-matrix: discover the union of the existing allowlist and the `cargo tree -p ironclaw_reborn_cli -e normal,build` closure ∩ workspace members (64 crates). Union (not replace) so non-closure reborn-family crates — channel adapters, webui_v2 — are never dropped from coverage. - package-feature-flags.sh: derive fallback features (default/libsql when declared) for closure crates without an explicit recipe; keep the previously-allowlisted no-flag crates flag-free so their behavior is unchanged. The 3 closure reds this would have caught are already fixed on main (#5105, #5108), so the closure should be green — this PR's own CI run is the 64/64 verification. Tradeoff: 21 -> 64 parallel crate jobs raises compute and, if runner concurrency is capped, may raise wall-clock as jobs queue. Follow-ups: (1) build-once `nextest archive` + shard to cut redundant compiles (spike #5086); (2) bake a few runs, then promote reborn-tests to a required check; (3) cut v1 (`test.yml` / Tests (all-features), ~29m) so the gate drops to ~10-12m. Automated agent-authored. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…5110) * ci(reborn): run the full reborn_cli dependency closure on every PR PR CI's reborn-tests matrix was a name-prefix allowlist of 21 reborn/ product-family crates. But only 10 of those overlap the actual `ironclaw_reborn_cli` dependency closure (53 workspace crates) — so 43 crates the shipped Reborn binary links (auth, host_runtime, skills, first_party_extensions, extensions, dispatcher, llm, safety, memory, network, turns, host_api, loop_support, threads, ...) ran their own test suites ONLY via the nightly/manual closure path, never on a PR. They were exercised on PR only indirectly by the 4 root integration partitions, which don't run those crates' own unit/contract tests. That gap is exactly why the bugs fixed in #5105/#5108 (host_runtime github surface, skills TOCTOU, gsuite wrong-account egress, loop_support/ threads/auth) slipped through normal PR CI and only surfaced when the closure was run by hand. This makes the closure the default PR matrix — "run everything on every PR": - package-matrix: discover the union of the existing allowlist and the `cargo tree -p ironclaw_reborn_cli -e normal,build` closure ∩ workspace members (64 crates). Union (not replace) so non-closure reborn-family crates — channel adapters, webui_v2 — are never dropped from coverage. - package-feature-flags.sh: derive fallback features (default/libsql when declared) for closure crates without an explicit recipe; keep the previously-allowlisted no-flag crates flag-free so their behavior is unchanged. The 3 closure reds this would have caught are already fixed on main (#5105, #5108), so the closure should be green — this PR's own CI run is the 64/64 verification. Tradeoff: 21 -> 64 parallel crate jobs raises compute and, if runner concurrency is capped, may raise wall-clock as jobs queue. Follow-ups: (1) build-once `nextest archive` + shard to cut redundant compiles (spike #5086); (2) bake a few runs, then promote reborn-tests to a required check; (3) cut v1 (`test.yml` / Tests (all-features), ~29m) so the gate drops to ~10-12m. Automated agent-authored. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * ci(reborn): test ironclaw_host_runtime with test-support feature in the closure host_runtime's integration tests (tests/) link the lib as a normal dependency, so cfg(test) is false there and the deterministic test-mode behavior they assert is gated behind `feature = "test-support"`. The generic default/libsql fallback runs the lib in production mode, so give host_runtime an explicit `--features test-support,libsql` recipe — libsql exercises the embedded-DB paths without needing a Postgres server (which the crate-tests job does not provision). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
📊 Measured Results (non-blocking measurement spikes)
This PR contains two self-contained, non-blocking experimental workflows used to settle the "can we run everything on the merge gate, and how do we make compile fast" questions with real numbers:
.github/workflows/experimental-full-suite.yml— build-oncenextest archive+ mold + sccache + 4-way sharding..github/workflows/experimental-compile-bench.yml— compile-time A/B (llvm vs cranelift),--timingsprofiling, and areborn-onlyscope to measure the cost of the v1 monolith.All numbers below are from the workflows' own CI runs. Caches were cold-to-partial, so absolute times are near-worst-case; ratios are the robust signal.
A. Can we run the full deterministic suite on the merge gate? — Yes, ~21–24 min
all-featuresdefaultCARGO_BUILD_JOBS=1can be lifted. This was the core unknown; answered.relation "memory_documents" does not exist,pg_typeschema races, CLI-smoke env). A true all-pass "run everything" needs migrations + per-test DB isolation + env — a finite (~150-test), systematic workstream this spike has now quantified.B. Compile-time A/B + profiling
wasmtime(backend incompatible). Dead end for this workspace.cargo tree -i aws-lc-syson default features → empty). The 67saws-lc-sysC build and theaws-smithy-*tree appear only under--all-features/--features bedrock. No default-build AWS cost to remove — no-op.Where compile time actually goes (top units,
cargo --timings):ironclawlib (test)ironclawlibironclaw_reborn_composition(test+lib)ironclawbuild-scriptaws-lc-sysbuild-script--all-featurestests/*.rsbinaries→ The root
ironclawv1 crate (lib + test) is ~55% of total compile time in one crate. A single 426s unit is serial — more cores don't help until it's split.C. The big one — cost of the v1 monolith: building only Reborn is ~4× faster
ironclaw_reborn_clidoes not depend on the rootironclawcrate (cargo tree -p ironclaw_reborn_cli -i ironclaw→ no match), so a Reborn-only build skips the monolith entirely:--workspace --all-targets --all-features)-p ironclaw_reborn_cli --all-targets --all-features)Dropping the v1 codebase roughly quarters the build — by far the single biggest compile lever, bigger than mold + sccache + cranelift combined.
D. Conclusions → roadmap (fully measured)
CARGO_BUILD_JOBS=1acrossreborn-tests.yml/reborn-integration.yml— proven, immediate, no-OOM.Original spike description
This PR adds self-contained, non-blocking experimental CI workflows. They are measurement spikes, not production merge-gate wiring — they run on
pull_request+workflow_dispatch, are not wired into any existing workflow, and are not required checks.Technique (full-suite): build one
cargo nextest archiveper feature config; compile with workflow-scopedmold+sccache; deliberately leaveCARGO_BUILD_JOBSparallel to test whether mold makes parallel linking viable; fan each archive across 4 nextest shards with no recompile (--workspace-remap .); collect duration JSON for a later duration-balanced (LPT) sharding pass.Deliverable: each workflow's CI run emits its measurement report in the job summary (build seconds, sccache hit rate, p50/max shard, build+slowest-shard vs a 30-min budget, mold/parallel-link status; and for compile-bench, the llvm-vs-cranelift delta, top-15 slowest units, and the reborn-only delta).
Follow-ups if promoted: duration-balanced LPT sharding from the uploaded artifacts; a real env layer (migrations + per-test DB) for a true all-pass gate; promote to a required
merge_groupgate only after the report shows it fits budget reliably.Automated agent-authored spike. Note: the
regression-test-checkred on this PR is a false-positive of that gate on a CI-only change (no product code → no regression test applies).