Repository navigation
ci: report scheduled-workflow failures as a tracking issue - #702
Conversation
|
Important Review skippedAuto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: ASSERTIVE Plan: Pro Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Note Reviews pausedUse the following commands to manage reviews:
Use the checkboxes below for quick actions:
WalkthroughThe pull request adds a composite action that creates or updates tracking issues for scheduled workflow failures. The Miri and stress workflows invoke it for failed or cancelled scheduled runs with scoped issue-write permission. ChangesScheduled failure reporting
Estimated code review effort: 3 (Moderate) | ~20 minutes Sequence Diagram(s)sequenceDiagram
participant ScheduledWorkflow
participant ReportCronFailure
participant GitHubIssues
ScheduledWorkflow->>ReportCronFailure: invoke after scheduled failure
ReportCronFailure->>GitHubIssues: ensure label and query open issues
GitHubIssues-->>ReportCronFailure: return exact-title match
alt matching issue exists
ReportCronFailure->>GitHubIssues: add failure comment
else no matching issue
ReportCronFailure->>GitHubIssues: create tracking issue
end
Possibly related PRs
Suggested labels: 🚥 Pre-merge checks | ✅ 3✅ Passed checks (3 passed)
Comment |
|
@coderabbitai pause |
✅ Action performedReviews paused. |
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #702 +/- ##
==========================================
- Coverage 94.04% 94.03% -0.01%
==========================================
Files 178 178
Lines 108835 108835
==========================================
- Hits 102353 102347 -6
- Misses 6482 6488 +6 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
4b45c2e to
3efb8ec
Compare
|
@coderabbitai review |
✅ Action performedReview finished.
|
There was a problem hiding this comment.
Actionable comments posted: 2
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In @.github/actions/report-cron-failure/action.yml:
- Around line 50-54: Update the gh issue list invocation used by the existing
issue lookup to retrieve all relevant open issues, either by setting an explicit
limit covering the expected maximum or by paginating through results. Preserve
the exact-title jq matching and existing variable flow so tracking issues are
still identified without creating duplicates.
In @.github/workflows/miri.yml:
- Around line 87-95: Replace the inline “Report failure” steps in
.github/workflows/miri.yml lines 87-95 and .github/workflows/stress.yml lines
60-68 with separate reporting jobs. Each job must independently check out the
repository, depend on its corresponding miri or stress-tests job, and run only
when always() && needs.<job>.result == 'failure' && github.event_name ==
'schedule', while preserving the report-cron-failure action and github.token
input.
🪄 Autofix
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro
Run ID: acc3181e-f430-469d-84ea-2826ac33f505
📒 Files selected for processing (3)
.github/actions/report-cron-failure/action.yml.github/workflows/miri.yml.github/workflows/stress.yml
3efb8ec to
02cbcbb
Compare
|
@coderabbitai review |
✅ Action performedReview finished.
|
|
Note GitHub couldn't provide a complete incremental comparison for this pull request, so CodeRabbit is performing a full review instead. This review may take a little longer. |
A failing cron notifies nobody. GitHub emails only the person who last touched the workflow file, which is private, easy to filter, and invisible to everyone else. `miri.yml` failed on every one of its runs for two weeks without that surfacing anywhere. Add a composite action that opens a tracking issue when a scheduled workflow fails, and wire it into both crons as a dedicated reporting job. A job rather than a final step in the job that failed, because the two failures most worth reporting are the ones a step could not report on. A step `uses: ./.github/actions/...`, which only resolves once the repository is on disk, so a failed checkout takes the reporting down with it; and `timeout-minutes` *cancels* the job, which skips `if: failure()` steps entirely — so the hang that bound exists to catch would go unreported. The reporting job checks out for itself and reads the outcome from `needs`, which survives both. It matches `cancelled` as well as `failure`. GitHub documents a `timeout-minutes` overrun as cancelling the job and does not state which of the two it then reports, so matching both is what makes a hang reportable either way. The cost is that a deliberately cancelled scheduled run also files — but a human cancelling a nightly cron was watching it, and the dedupe folds that into one issue they can close. A hang has nobody watching, which is the whole problem. The issue text says "did not succeed" rather than "failed" for the same reason. It is gated on `github.event_name == 'schedule'`, not on failure alone. A `workflow_dispatch` failure is already in front of whoever dispatched it, and if either workflow later also runs on `pull_request` a failure there shows up in that PR's checks. Only the unwatched trigger needs an issue; filing one for the other two would reintroduce the noise the dedupe exists to avoid. The dedupe is the substance: a persistent failure comments on the existing open issue for that workflow rather than filing a new one each night. One issue per fortnight-long outage is a signal; fourteen issues is the same silence reached by a different route. Matching is on the exact issue title, so the two workflows never adopt each other's issue. The lookup passes `--limit` because `gh issue list` returns only the first 30 by default — a match paged out of that window would file the duplicate the dedupe exists to prevent. `stress.yml` gets the same job. It has been green every day so far, which means its blind spot is untested rather than absent. `issues: write` is scoped to the reporting job alone rather than the whole workflow; neither workflow writes to repository contents, and `persist-credentials: false` on checkout is unchanged. The `ci-failure` label is created on demand so a fresh checkout needs no out-of-band setup. Verified by extracting the script and running it against a stubbed `gh`: files an issue when none is open, comments when a matching one exists, does not confuse an issue belonging to the other workflow for its own, and finds a match sitting past the default page.
02cbcbb to
e3284d9
Compare
Summary
A failing cron notifies nobody. GitHub emails only the person who last touched the workflow file — private, easy to filter, invisible to the rest of the team.
miri.ymlfailed on every one of its 13 runs over two weeks and that surfaced nowhere. It was found by chance, by manually dispatching the workflow for an unrelated reason.This adds a composite action that opens a tracking issue when a scheduled workflow fails, wired into both crons as an
if: failure()step.The dedupe is the point
A naive "file an issue on failure" would have produced fourteen issues for the outage above, which is just as easy to tune out as producing none — the same silence reached by a different route.
So: one open issue per workflow, reused. The first failure files it; every subsequent failure comments on it. The comment count then reads as how long this has been broken. Closing the issue once the workflow is green means the next failure opens a fresh one.
Matching is on the exact issue title (
Scheduled workflow failing: <name>) rather thangh's full-text--search, which would also match issues that merely mention the workflow — and would let the two crons adopt each other's issue.Scope
Both cron workflows, not just the broken one.
stress.ymlhas the identical blind spot; it has simply been green every day so far, so its blind spot is untested rather than absent.Permissions
Both workflows gain
issues: write, used only by the reporting step. Neither writes to repository contents, andpersist-credentials: falseon checkout is unchanged — the step authenticatesghthroughgithub.tokenin the environment, not through git credentials.The
ci-failurelabel is created on demand inside the action (idempotent, failure tolerated), so this needs no out-of-band setup and no manual step before the first failure can be reported.No third-party action is introduced —
ghandjqare both preinstalled on GitHub-hosted runners, so there is nothing new to pin by commit hash.Verification
The failure path is the one that never runs in normal operation, which is exactly how the original bug survived, so I tested it directly rather than by inspection. I extracted the action's script and ran it against a stubbed
ghacross three cases:Also confirmed the issue body renders flush-left — indented heredoc content would have turned the whole body into a Markdown code block.
What this does not solve
miri.ymlfailed on its very first run, so there was never a green baseline to regress from. Notification would have caught it on day one, which is enough here. But nothing automatable enforces "confirm a new gate actually passes before trusting it" — that stays a review-time habit.Related
miri.ymlfailure. The two PRs touch different regions of that file and should not conflict.Risk: command output changes: yes, issue reporting is pinned by exact workflow-title matching;
unsafe: none, CLAUDE.md allowlist: unchanged; memory, queue, and thread/backpressure policy: none.Fix: The composite action reports scheduled
miri.ymlandstress.ymlfailures through one open tracking issue per workflow. It creates theci-failurelabel when needed, comments on an existing exact-title issue, or creates a new issue after closure. Both workflows now grantissues: write.