ci: revive the nightly deep tier + make Platform & Compat and Reborn E2E requirable - #5841
Conversation
…E2E requirable Nightly Deep CI has startup-failed every night since 2026-06-20 with zero jobs: reborn-tests.yml references secrets.SCCACHE_* (sccache-dist) and nightly's call site did not pass secrets, which fails workflow validation at trigger time. The in-run nightly-alert job dies with the run, so nothing reported it. Separately, Platform & Compat's deep jobs (Windows build, bench compile, docker build) were gated on github.event_name == 'workflow_call', which never matches — in a reusable workflow event_name reflects the caller's event — so nightly's "deep reuse" of those jobs silently skipped (all three are 'skipped' in the last nightly run that executed jobs at all). - nightly-deep-ci.yml: pass `secrets: inherit` to the reborn-tests call - platform-and-compat.yml: add a `deep` workflow_call marker input (default true, materializes only under workflow_call) and gate windows-build / wasm-wit-compat / bench-compile / docker-build on it instead of the never-true event_name comparison - nightly-watchdog.yml (new): dead-man's switch that inspects the latest scheduled Nightly Deep CI run from outside it — startup failures and never-fired crons now raise/update the same "Nightly Deep CI failed" issue via .github/scripts/nightly-alert-issue.sh - platform-and-compat.yml: add a stable "Platform & Compat" roll-up job (skip-tolerant, requirable as a status check), delete the vestigial matrix-config job (test_matrix had no consumer; windows_matrix's SLIM branch was unreachable because windows-build never runs on PR or merge_group), and build Docker images in the merge queue when the merge group touches Dockerfile inputs - reborn-e2e.yml: run in the merge queue — merge_group trigger plus a changes job mirroring the pull_request/push paths filters (merge_group does not support paths), with the "Reborn E2E" roll-up reporting on every queue entry so it can become a required check - .github/workflows/README.md (new): the CI tier contract, required-check inventory, deep-tier gotchas, and deliberately accepted gaps Verified with actionlint (no findings beyond pre-existing SC2129 style nits). Reborn E2E queue cost is ~5-9 min based on recent main runs. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
🔎 IronLoop Review StatusHead: Current reviewers:
Reviewer summaries
Recent activity
Available commands
Run metadataAdmission: webhook accepted the request and IronLoop persisted review state before this projection. |
|
Note Gemini is unable to generate a review for this pull request due to the file types involved not being currently supported. |
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: ASSERTIVE Plan: Pro Plus Run ID: 📒 Files selected for processing (1)
📝 WalkthroughSummary by CodeRabbit
WalkthroughThis PR updates CI contracts and gating across reusable workflows, switches E2E scope detection to job logic, moves nightly alerting out of workflow runs, and raises one nightly stress threshold. ChangesCI workflow contract and gating changes
Estimated code review effort: 4 (Complex) | ~45 minutes Possibly related PRs
🚥 Pre-merge checks | ✅ 3 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (3 passed)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
❌ IronLoop Review: reviewer
Review at a glance
| Verdict | Blocking | Notes | Inline | Head |
|---|---|---|---|---|
| ❌ Changes requested | 1 | 0 | 1 | 4779368199f7 |
Head: 4779368199f73e24c025263c34b5dbba8a1d8f21
Next: Fix the blocking findings, push the PR branch, then re-run this reviewer.
Run details
Status: Current
Needs human: no
Needs validation: no
Summary
Found one blocking CI gating issue in the new Platform & Compat merge-queue Docker path.
Findings
Blocking: 1 / Notes: 0
Blocking findings
1. ❌ [MEDIUM] Docker-risk merge groups still skip the Docker build
Location: .github/workflows/platform-and-compat.yml:347-351
The new merge_group Docker path is still guarded by needs.changes.outputs.has_legacy_tests == 'true'. scripts/ci/classify-test-scope.sh returns has_legacy_tests=false for Docker inputs such as Dockerfile.reborn and .dockerignore, while the new has_docker_risk regex returns true for them. As a result, a merge group that only changes Dockerfile.reborn or .dockerignore skips docker-build, and the roll-up accepts the skipped job, leaving the deterministic Docker failure to show up only after merge. Either classify all Docker-risk paths as test-relevant or exempt the Docker-risk merge_group arm from the legacy-tests guard.
Developer follow-up
After fixing this feedback:
- Push the fix to this PR branch.
- Re-run this reviewer with
@ironloopai review --agent reviewerif you only changed this reviewer's findings. - Re-run all reviewers with
@ironloopai reviewwhen the fix may affect multiple areas. - Use
@ironloopai statusto check queued/running/completed/stale/stalled state while reviewers run.
| (github.event_name == 'push' || (github.event_name == 'workflow_call' && inputs.include_docker)) | ||
| (github.event_name == 'push' || | ||
| (inputs.deep == true && inputs.include_docker) || | ||
| (github.event_name == 'merge_group' && needs.changes.outputs.has_docker_risk == 'true')) |
There was a problem hiding this comment.
This merge_group Docker arm is still behind has_legacy_tests == 'true'. The classifier reports has_legacy_tests=false for Dockerfile.reborn and .dockerignore, even though the new has_docker_risk regex reports true, so those changes skip docker-build and the roll-up treats the skip as passing.
There was a problem hiding this comment.
Fixed in 176fb97 — the merge_group arm no longer requires has_legacy_tests: it now gates on docs_only != 'true' && has_docker_risk == 'true' alone, so a Dockerfile.reborn/.dockerignore-only merge group builds the images. The push/deep arms keep the has_legacy_tests gate unchanged.
Coverage ratchetReborn integration-tier coverageLine coverage (Reborn crates): 85.17% — 282279 / 331449 lines Per-crate breakdown (65 crates, lowest-covered first)
This table itself is informational and never gates the PR on its own — not the percentage, not the per-crate holes, not the 0-coverage callout. A separate coverage ratchet (dry-run until enforce=true; see tests/integration/coverage-floor.toml) can fail the build on specific configured floors. Exemptions (4 entry/entries excluded from the accounting above)
|
A dispatched validation run of the previous commit still startup-failed: secrets: inherit fixed one violation, but called-workflow jobs that declare job-level permissions beyond the caller's grant also fail validation at trigger time (even when those jobs are event-gated off schedule runs). platform-and-compat's version-check declares pull-requests/issues read; reborn-tests' coverage-report declares pull-requests: write. Grant those supersets at the call sites. Also corrects the incident window in the comments and README: retained history shows zero successful Nightly Deep CI runs since its creation on 2026-05-06 (65 of 74 runs are startup_failures), not merely since 2026-06-20 — the permissions violations date to day one, the secrets one to the 2026-07-03 sccache rollout. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
There was a problem hiding this comment.
Actionable comments posted: 2
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In @.github/workflows/nightly-watchdog.yml:
- Around line 73-83: The “Drive the nightly alert issue” step can be skipped
when the evaluate step fails, which breaks the watchdog’s dead-man’s switch
behavior. Add an always-run condition to this step in the nightly-watchdog
workflow, matching the existing pattern used by the nightly-alert job in
nightly-deep-ci.yml, so the alert issue is still created or refreshed even if
steps.evaluate exits non-zero.
In @.github/workflows/platform-and-compat.yml:
- Around line 14-26: The `has_engine_replay_risk` job condition is still using
`github.event_name == 'workflow_call'`, which fails inside this reusable
workflow because it sees the caller’s event instead of `workflow_call`. Update
the job’s condition to use `inputs.deep` consistently with the other jobs, or
make it unconditional if that is the intended behavior, and reference the
existing reusable-workflow gating pattern in
`.github/workflows/platform-and-compat.yml`.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro Plus
Run ID: 5995fdac-d1f9-4021-b4ae-80e20af0bf0b
📒 Files selected for processing (5)
.github/workflows/README.md.github/workflows/nightly-deep-ci.yml.github/workflows/nightly-watchdog.yml.github/workflows/platform-and-compat.yml.github/workflows/reborn-e2e.yml
…eshold Reborn Playwright and IronClaw Stress were failing nightly with no alerting at all — same silent-failure class the deep-CI revival fixed, in two more places. Diagnosis of the standing failures: - IronClaw Stress (red every retained scheduled run): the nightly bottleneck suite caps p95 at 1500ms, but its model-tail case injects a synthetic 2.0s model wait by design — the ceiling was structurally unsatisfiable (measured p95 2.05s = the 2.0s wait + ~50ms real work; every case at 0.00% failures). Raised to 2500ms with a comment; the PR-mode variant already ran at 3000ms. Per-case ceilings in ironclaw_stress are the proper follow-up. - Reborn Playwright (flaky-red): last night's failure was test_reborn_legacy_looping_tool_calls_stop_at_low_iteration_boundary asserting failure_category == "driver_protocol_violation" while the runtime now emits "iteration_limit" — already realigned on main by the error-classification refactor (#5651); no change needed here. Alerting changes: - ironclaw-stress.yml + reborn-playwright.yml: schedule-only alert jobs driving .github/scripts/nightly-alert-issue.sh, same contract as the Nightly E2E / Nightly Deep CI alerts. - nightly-watchdog.yml: generalized to a matrix over all four nightlies (Deep CI, E2E, Playwright, Stress) — startup failures and cron-never-fired now alarm for every nightly, not just Deep CI. - nightly-alert-issue.sh: optional Slack mirror. When the SLACK_CI_ALERTS_WEBHOOK_URL repo secret is set (Slack incoming webhook), every failure posts a one-liner with run + issue links and every recovery posts a close-out; absent secret = silent no-op. Delivery is best-effort and never fails the alert job. All five call sites pass the secret through. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Per review direction: failures post to Slack, nothing else, one code path, no GitHub issues. - nightly-watchdog.yml is now the only alerting mechanism: at 08:00 UTC it checks each nightly's latest scheduled run (Nightly Deep CI, Nightly E2E, Reborn Playwright, IronClaw Stress) and posts failures — workflow, conclusion, failed job names, run link — to the Slack channel behind the existing secrets.SLACK_WEBHOOK_URL (the same webhook live-canary reports through; no new secret needed). Missing runs, stale runs (>26h, cron never fired), and startup_failures alarm too — the cases an in-run alert job structurally cannot see. A detected failure turns the watchdog matrix job red so its run history doubles as the failure record. Successes post nothing. - Removed the GitHub-issue alerting entirely: nightly-alert-issue.sh and its test harness are deleted, and the in-run nightly-alert jobs are removed from nightly-deep-ci.yml and nightly-e2e.yml along with the round-2 stress/playwright alert jobs. Housekeeping after merge: close the open "Nightly E2E failed" issue (#4108) manually — nothing auto-closes it now. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
🚅 Deployed to the ironclaw-pr-5841 environment in ironclaw-ci-preview
|
There was a problem hiding this comment.
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (1)
.github/workflows/README.md (1)
107-110: 📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick winDrop the stale
paths:reference.
reborn-e2e.ymlnow derives scope from itschangesjob and merge-group trigger; there are no PRpaths:filters to keep in sync anymore. Leaving this bullet as-is will send future edits to the wrong contract.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In @.github/workflows/README.md around lines 107 - 110, Remove the outdated note about keeping `reborn-e2e.yml`’s `paths:` filters in sync, since that workflow now derives scope from its `changes` job and merge-group trigger only. Update the guidance in the README section about scope classifiers to reference the current contract in `reborn-e2e.yml` and the `changes` job, so future edits point to the right workflow mechanism.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Outside diff comments:
In @.github/workflows/README.md:
- Around line 107-110: Remove the outdated note about keeping `reborn-e2e.yml`’s
`paths:` filters in sync, since that workflow now derives scope from its
`changes` job and merge-group trigger only. Update the guidance in the README
section about scope classifiers to reference the current contract in
`reborn-e2e.yml` and the `changes` job, so future edits point to the right
workflow mechanism.
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro Plus
Run ID: 517cdec9-e740-4e59-ab29-ddf9a33d9ddd
📒 Files selected for processing (7)
.github/scripts/nightly-alert-issue.sh.github/scripts/nightly-alert-issue.test.sh.github/workflows/README.md.github/workflows/ironclaw-stress.yml.github/workflows/nightly-deep-ci.yml.github/workflows/nightly-e2e.yml.github/workflows/nightly-watchdog.yml
💤 Files with no reviewable changes (2)
- .github/scripts/nightly-alert-issue.sh
- .github/scripts/nightly-alert-issue.test.sh
- docker-build: the merge_group arm no longer requires has_legacy_tests. The scope classifier matches only the literal `Dockerfile`, so a Dockerfile.reborn/.dockerignore-only merge group reported has_legacy_tests=false and skipped the Docker build exactly when it should run (IronLoop blocking finding). push/deep arms keep the gate. - changes: drop has_engine_replay_risk — no consumer in this workflow, and its workflow_call arm could never fire (github.event_name is the caller's event in reusable workflows). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… nightly Deliberate freeze pending v1 (src/) removal, per team decision: nightly no longer calls the Legacy Tests workflow, leaving test.yml invoked nowhere. Documented loudly in the workflow header and the CI README — including the consequence that test.yml is the only place the root `ironclaw` package's tests run, and that a v1 fix landing before src/ is deleted should temporarily restore the call job. This is an explicit freeze with a paper trail, not the silent-death mode this workflow's history is infamous for; delete test.yml together with src/. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
There was a problem hiding this comment.
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (1)
.github/workflows/README.md (1)
111-114: 📐 Maintainability & Code Quality | 🟠 Major | ⚡ Quick winStale guidance: references
reborn-e2e.ymlpaths:filters that this PR removes.Line 114 instructs keeping
reborn-e2e.yml'schangesregex in sync with itspaths:filters. However, this PR's scope-based E2E gating change removes PR path filtering fromreborn-e2e.ymlin favor of ahas_e2e_scopechangesjob. The guidance should be updated to reflect the new scope-detection model rather than the removedpaths:filters.📝 Suggested update
- **Scope classifiers** (`scripts/ci/classify-test-scope.sh` and per-workflow `changes` jobs) are curated allowlists. Adding a new crate or test directory - requires updating them, or the queue's scoped checks silently narrow. Keep - `reborn-e2e.yml`'s `changes` regex in sync with its `paths:` filters. + requires updating them, or the queue's scoped checks silently narrow. Keep + `reborn-e2e.yml`'s `changes` regex in sync with the scope it is intended to + detect (formerly `paths:` filters, now the `has_e2e_scope` job).🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In @.github/workflows/README.md around lines 111 - 114, Update the scope-classifier guidance in the README to match the new E2E gating flow: remove the outdated reference to keeping reborn-e2e.yml’s changes regex in sync with its paths filters, and instead describe that the has_e2e_scope changes job is now the source of truth for PR-based E2E scope detection. Keep the note about scripts/ci/classify-test-scope.sh and other per-workflow changes jobs being curated allowlists, but align the reborn-e2e.yml guidance with the new scope-based model.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Outside diff comments:
In @.github/workflows/README.md:
- Around line 111-114: Update the scope-classifier guidance in the README to
match the new E2E gating flow: remove the outdated reference to keeping
reborn-e2e.yml’s changes regex in sync with its paths filters, and instead
describe that the has_e2e_scope changes job is now the source of truth for
PR-based E2E scope detection. Keep the note about
scripts/ci/classify-test-scope.sh and other per-workflow changes jobs being
curated allowlists, but align the reborn-e2e.yml guidance with the new
scope-based model.
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro Plus
Run ID: 16d6fbee-b2ba-4a44-a9da-a02de7155c50
📒 Files selected for processing (2)
.github/workflows/README.md.github/workflows/nightly-deep-ci.yml
Completes the legacy freeze: nightly-e2e.yml (the scheduler for e2e.yml's full v1 browser suite) is deleted and Nightly E2E is dropped from the watchdog matrix. Like Nightly Deep CI before its revival, it had zero successful runs in retained history — its own alert issue notes there is no prior green run on main to attribute against — and its standing failures (v1 /api/chat auth-gate SSE dedup and approval tests) are v1 work the team has frozen pending src/ removal. e2e.yml itself stays (workflow_call/workflow_dispatch), frozen alongside test.yml; both are documented in the CI README to be deleted together with src/. The nightly fleet is now Nightly Deep CI, Reborn Playwright, and IronClaw Stress — all three validated green today — with the watchdog covering exactly those three. After merge: manually close the open "Nightly E2E failed" issue #4108 (the freeze resolves it; nothing auto-closes it). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…eeze The Emulate-backed full-path tests are the only scheduled coverage for install -> OAuth -> model-routed tool call -> provider mutation, and per tests/e2e/CLAUDE.md they boot the legacy gateway binary, so they froze with v1. Note in the CI README that a Reborn-native port through `ironclaw-reborn serve` is the follow-up that restores this tier — deliberately NOT re-homed into the Reborn nightlies as-is, which would have smuggled the legacy binary back in. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
There was a problem hiding this comment.
Actionable comments posted: 1
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (1)
.github/workflows/nightly-watchdog.yml (1)
39-91: 🩺 Stability & Availability | 🟡 Minor | ⚡ Quick winAdd a timeout to the Slack webhook POST. The empty
gh run listcase already falls through tomissing; the remaining failure mode iscurlhanging forever if Slack stalls.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In @.github/workflows/nightly-watchdog.yml around lines 39 - 91, Add a timeout to the Slack webhook POST in the “Check latest scheduled run and post failures to Slack” step so the workflow cannot hang indefinitely if Slack is slow or unresponsive. Update the curl invocation that sends the JSON payload to use an explicit timeout/retry-safe setting, and keep the existing fallback warning behavior intact in the same run block where text and payload are built.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In @.github/workflows/README.md:
- Around line 106-113: Tighten the v1 freeze wording in the README section that
discusses the legacy suites and nightly scheduling: the current “invoked
nowhere” phrasing is too absolute given the documented manual/temporary restore
paths. Update the text around the legacy v1 suites, nightly-e2e, test.yml,
e2e.yml, and nightly-deep-ci.yml references so it says the suites are not
scheduled by nightly rather than unusable, while preserving the guidance about
temporary restoration and deletion alongside src/.
---
Outside diff comments:
In @.github/workflows/nightly-watchdog.yml:
- Around line 39-91: Add a timeout to the Slack webhook POST in the “Check
latest scheduled run and post failures to Slack” step so the workflow cannot
hang indefinitely if Slack is slow or unresponsive. Update the curl invocation
that sends the JSON payload to use an explicit timeout/retry-safe setting, and
keep the existing fallback warning behavior intact in the same run block where
text and payload are built.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro Plus
Run ID: 94ec6578-cce5-4426-bd39-f3c4b8844744
📒 Files selected for processing (3)
.github/workflows/README.md.github/workflows/nightly-e2e.yml.github/workflows/nightly-watchdog.yml
💤 Files with no reviewable changes (1)
- .github/workflows/nightly-e2e.yml
| - **The legacy v1 suites are deliberately invoked nowhere** — v1 (`src/`) is | ||
| frozen pending removal. `test.yml` (the only place the root `ironclaw` | ||
| package's tests run) is no longer called by nightly, and the former | ||
| `nightly-e2e.yml` scheduler for the v1 browser suite (`e2e.yml` full mode) | ||
| is deleted — it had zero successful runs in retained history. Until `src/` | ||
| is deleted, a v1 bug fix that must land should temporarily restore the | ||
| `deterministic-deep-tests` call in `nightly-deep-ci.yml` (and/or dispatch | ||
| `e2e.yml` manually). Delete `test.yml` and `e2e.yml` together with `src/`. |
There was a problem hiding this comment.
📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win
🧩 Analysis chain
🏁 Script executed:
#!/bin/bash
set -euo pipefail
printf '\n== repo rule files ==\n'
git ls-files | rg '(^|/)(CLAUDE\.md|AGENTS\.md|\.claude/.*|\.github/workflows/README\.md)$' || true
printf '\n== relevant README excerpt ==\n'
nl -ba .github/workflows/README.md | sed -n '96,120p'
printf '\n== repo rules excerpt ==\n'
for f in CLAUDE.md AGENTS.md .claude/rules.md .claude/rules/*; do
[ -f "$f" ] && { echo "--- $f"; nl -ba "$f" | sed -n '1,220p'; }
doneRepository: nearai/ironclaw
Length of output: 6378
🏁 Script executed:
#!/bin/bash
set -euo pipefail
printf '\n== README lines 100-118 ==\n'
sed -n '100,118p' .github/workflows/README.md | cat -n
printf '\n== doc/review rules ==\n'
for f in .claude/rules/doc-hygiene.md .claude/rules/review-discipline.md .claude/rules/tool-evidence.md; do
echo "--- $f"
sed -n '1,220p' "$f" | cat -n
doneRepository: nearai/ironclaw
Length of output: 12969
Tighten the v1 freeze wording. .github/workflows/README.md:106-113
“Invoked nowhere” overstates the contract; the next sentence still leaves manual/temporary restore paths. Say “not scheduled by nightly” so the doc doesn’t read as if test.yml/e2e.yml are unusable. .claude/rules/review-discipline.md treats stale contract wording as a bug.
♻️ Suggested wording fix
-- **The legacy v1 suites are deliberately invoked nowhere** — v1 (`src/`)
- is frozen pending removal.
+- **The legacy v1 suites are not scheduled by nightly** — v1 (`src/`)
+ is frozen pending removal.📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| - **The legacy v1 suites are deliberately invoked nowhere** — v1 (`src/`) is | |
| frozen pending removal. `test.yml` (the only place the root `ironclaw` | |
| package's tests run) is no longer called by nightly, and the former | |
| `nightly-e2e.yml` scheduler for the v1 browser suite (`e2e.yml` full mode) | |
| is deleted — it had zero successful runs in retained history. Until `src/` | |
| is deleted, a v1 bug fix that must land should temporarily restore the | |
| `deterministic-deep-tests` call in `nightly-deep-ci.yml` (and/or dispatch | |
| `e2e.yml` manually). Delete `test.yml` and `e2e.yml` together with `src/`. | |
| - **The legacy v1 suites are not scheduled by nightly** — v1 (`src/`) is | |
| frozen pending removal. `test.yml` (the only place the root `ironclaw` | |
| package's tests run) is no longer called by nightly, and the former | |
| `nightly-e2e.yml` scheduler for the v1 browser suite (`e2e.yml` full mode) | |
| is deleted — it had zero successful runs in retained history. Until `src/` | |
| is deleted, a v1 bug fix that must land should temporarily restore the | |
| `deterministic-deep-tests` call in `nightly-deep-ci.yml` (and/or dispatch | |
| `e2e.yml` manually). Delete `test.yml` and `e2e.yml` together with `src/`. |
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In @.github/workflows/README.md around lines 106 - 113, Tighten the v1 freeze
wording in the README section that discusses the legacy suites and nightly
scheduling: the current “invoked nowhere” phrasing is too absolute given the
documented manual/temporary restore paths. Update the text around the legacy v1
suites, nightly-e2e, test.yml, e2e.yml, and nightly-deep-ci.yml references so it
says the suites are not scheduled by nightly rather than unusable, while
preserving the guidance about temporary restoration and deletion alongside src/.
Follow-up to #5840 (merge-queue full clippy matrix). That PR closes the gate for feature-matrix lints; this one revives the silently-dead deep tier and makes the remaining deterministic workflows requirable as status checks.
Why
Nightly Deep CI has never had a successful run — not once since its creation on 2026-05-06. 65 of its 74 retained runs are
startup_failurewith zero jobs executed. Root cause: reusable-workflow call-contract violations, validated by GitHub at trigger time, which kill the entire run before any job — including the in-runnightly-alertjob, so nothing ever reported it (contrast: "Nightly E2E failed" #4108 exists because that workflow fails after starting). Three stacked violations:platform-and-compat.yml'sversion-checkjob declares job-levelpull-requests: read+issues: read; nightly granted onlycontents: read. A called workflow may not request more than the caller grants — validated even though the job is event-gated off schedule runs. (Present since roughly day one.)reborn-tests.yml'scoverage-reportjob declarespull-requests: write— same class (landed with the coverage ratchet).reborn-tests.ymlreferencessecrets.SCCACHE_*(2026-07-03 sccache rollout) and the call site didn't passsecrets: inherit.Separately: Platform & Compat's deep jobs never ran under nightly anyway.
github.event_name == 'workflow_call'never matches — in a reusable workflow,event_namereflects the caller's event (schedule). Windows Build, Benchmark Compilation, and Docker Build are allskippedin the last nightly run that executed jobs at all (run 28702916492).Also, per the CI-contract direction: Platform & Compat had no stable roll-up (unrequirable), and Reborn E2E didn't run in the merge queue.
What
secrets: inherit+ caller-sidepermissionssupersets on the reborn-tests and platform-and-compat calls — unbreaks nightly startup.deepworkflow_call marker input (defaulttrue, only materializes under workflow_call);windows-build/wasm-wit-compat/bench-compile/docker-buildgate on it instead of the never-trueevent_name == 'workflow_call'..github/scripts/nightly-alert-issue.shissue contract. Alarms on: missing run, run older than 26h (cron never fired), or any non-success conclusion — includingstartup_failure, which the in-run alert structurally cannot report.Platform & Compatroll-up job (skip-tolerant of event/scope-gated jobs, requirable as a status check). Deleted the vestigialmatrix-configjob:test_matrixhad no consumer, andwindows_matrix's SLIM branch was unreachable (windows-build never runs on PR/merge_group). Docker images now also build in the merge queue when the merge group touchesDockerfile*/.dockerignore.merge_grouptrigger + achangesscope job (merge_group doesn't supportpaths:); thepull_requesttrigger drops itspaths:filter in favor of the same scope job so theReborn E2Eroll-up reports on every PR — a prerequisite for making it a required check without wedging the queue. Out-of-scope PRs/merge-groups skip the heavy jobs and the roll-up passes fast.Validation
Dispatched Nightly Deep CI from this branch via
workflow_dispatch:startup_failure, 0 jobs — which is how the permissions class was found.actionlint clean on all touched workflows (pre-existing style nits only). The Windows/bench/docker skip diagnosis is empirical (job conclusions from run 28702916492).
Queue cost
Reborn E2E adds ~5–9 min wall-clock per merge group (recent main-push run times), running in parallel with Tests (Reborn) (~10–15 min), so queue latency is roughly unchanged. Docker builds enter the queue only for Dockerfile-touching merge groups.
After this lands
Platform & CompatandReborn E2Eto the required checks in ruleset "Main" (Settings → Rules → Rulesets → Main), alongside the existingCode Style (fmt + clippy)andTests (Reborn).🤖 Generated with Claude Code
Round 2 (9f9f223): alert every nightly, Slack mirror, stress threshold
Reborn Playwright and IronClaw Stress turned out to be failing nightly with no alerting at all — the same silent-failure class, two more instances. Diagnoses:
model-tailcase injects a synthetic 2.0s model wait by design — the ceiling was structurally unsatisfiable (measured p95 2.05s = the synthetic wait + ~50ms of real work; every case at 0.00% failures — the system itself is healthy). Raised the nightly cap to 2500ms (PR mode already ran at 3000ms). Proper follow-up: per-case ceilings inironclaw_stress.test_reborn_legacy_looping_tool_calls_stop_at_low_iteration_boundaryassertingfailure_category == "driver_protocol_violation"while the runtime now emits"iteration_limit"— the expectation was already realigned on main by the error-classification refactor refactor(errors): static enforcement that failures surface, not swallow #5651 (merged 2026-07-08); no change needed here, tonight should pass.test_auth_required_sse_without_duplicate_response,test_chat_reply_always_auto_approves_next_same_tool) — real burn-down items, tracked by the existing Nightly E2E failed #4108 alert issue; not addressed in this CI-plumbing PR.Changes:
ironclaw-stress.yml+reborn-playwright.yml: schedule-only alert jobs driving the sharednightly-alert-issue.sh(issue titlesIronClaw Stress nightly failed/Reborn Playwright nightly failed).nightly-watchdog.yml: generalized to a matrix over all four nightlies — startup failures and cron-never-fired alarm everywhere, not just Deep CI.nightly-alert-issue.sh: optional Slack mirror — when theSLACK_CI_ALERTS_WEBHOOK_URLrepo secret is set (a Slack incoming-webhook URL), every failure posts a one-liner (run + issue links) and every recovery posts a close-out to the CI alerts channel. Absent secret = silent no-op, delivery is best-effort and never fails the alert job. All five call sites pass it through.Operator setup for Slack (one-time): create a Slack app → enable Incoming Webhooks → add a webhook for the alerts channel →
gh secret set SLACK_CI_ALERTS_WEBHOOK_URL --repo nearai/ironclaw.Round 3 (ec493db): alerting simplified per review direction — Slack-only, one path
nightly-alert-issue.sh+ its test harness deleted; in-runnightly-alertjobs removed fromnightly-deep-ci.ymlandnightly-e2e.yml(and the round-2 stress/playwright ones). Net −745 lines.nightly-watchdog.ymlis the single alerting path: 08:00 UTC, checks all four nightlies' latest scheduled runs, posts failures (workflow, conclusion, failed job names, run link) to Slack via the existingsecrets.SLACK_WEBHOOK_URL— the same webhook live-canary already reports through, so no new secret or channel setup is needed. Successes post nothing. A detected failure also turns the watchdog's matrix job red, so its run history doubles as the failure record.Round 4 (09e0a8f): review fixes + legacy suite frozen
has_legacy_tests(the classifier only matches the literalDockerfile, soDockerfile.reborn/.dockerignore-only merge groups were skipping the build). Dropped the deadhas_engine_replay_riskscope output (CodeRabbit). The watchdogif: always()comment is superseded by the round-3 single-step design. All threads replied.test.yml, leaving it invoked nowhere untilsrc/is removed. Documented loudly in the workflow header + README, including the consequence (test.yml is the only place the rootironclawpackage's tests run) and the revert path (restore the call job for a v1 fix). Explicit freeze with a paper trail — not the silent-death mode.Round 5: legacy fully out of the nightly fleet
nightly-e2e.ymldeleted and Nightly E2E removed from the watchdog matrix. It was the v1 browser suite's scheduler and — like pre-revival Deep CI — had zero successful runs in retained history (its own alert issue Nightly E2E failed #4108 says there is no prior green run to attribute against). Its standing failures are v1/api/chattests, frozen with the rest of v1.test.yml, no legacy v1 test suite is invoked by any scheduled workflow.test.ymlande2e.ymlremain in-tree (callable manually), documented in the README to be deleted together withsrc/.