Skip to content

fix(ci): collapse Runtime CI onto one hosted runner - #347

Closed
seonghobae wants to merge 2 commits into
mainfrom
fix/ci-single-runtime-runner-20260824
Closed

seonghobae wants to merge 2 commits into
mainfrom
fix/ci-single-runtime-runner-20260824

Conversation

@seonghobae

@seonghobae seonghobae commented Aug 24, 2026 •

Copy link
Copy Markdown
Contributor

RCA

Fresh exact-head runs across multiple open PRs are currently stalling before any repository step executes: Runtime CI, Security, SAST, SBOM, and provenance runs remain queued, and Runtime CI alone requests three independent ubuntu-latest allocations per pull request. The first failing boundary is therefore hosted-runner allocation, not checkout, PostgreSQL startup, compilation, tests, coverage, credentials, or repository code.

PR #285 already removed an avoidable second SBOM runner allocation while preserving its artifact handoff. Runtime CI still duplicates exact-head checkout, the pinned PostgreSQL service, and ephemeral database setup across three jobs solely to parallelize quality, line coverage, and branch coverage. Under the current allocation pressure that parallelism increases scheduling boundaries and blocks evidence arrival.

Falsifiable hypothesis: reducing Runtime CI from three hosted-runner allocations to one will reduce allocation pressure while preserving every existing quality and exact-coverage gate. Acceptance is observable on this exact head: Runtime CI must expose one job, execute all existing formatting/compile/lockfile/Clippy/test/rustdoc/line-coverage/branch-coverage checks against the same exact PR head, and complete without weakening permissions or evidence semantics.

TDD

  • RED 4e2a8fb14d57b63de3b7010f7c75c7d5f27fc476 changes the executable CI contract to require a single ubuntu-latest allocation, a single exact-head checkout/PostgreSQL service, and one pinned cargo-llvm-cov install while retaining both coverage diagnostics.
  • GREEN 8a53ed1c72243246dbe2dea394003cc238db67af collapses the three Runtime CI jobs into one sequential evidence job. The branch was kept private from PR-triggered CI until GREEN, so the RED/GREEN history is preserved without amplifying the already-congested runner queue.

Preserved gates

  • exact PR-head checkout remains immutable-SHA-pinned and uses persist-credentials: false;
  • workflow/job authority remains contents: read only;
  • PostgreSQL 18 image remains digest-pinned with per-run ephemeral credentials and health startup allowance;
  • formatting, locked compilation, Cargo.lock drift evidence, Clippy, all-target tests, Python coverage-contract regressions, and rustdoc remain mandatory;
  • exact 100% owned production line and branch coverage remain separate mandatory steps with the existing diagnostics;
  • Rust stable/nightly versions and cargo-llvm-cov remain pinned;
  • no branch protection, review, security, provenance, SBOM, or release gate is removed or weakened.

The only intentional tradeoff is wall-clock serialization inside one runner. Timeout is raised from the former per-job 15 minutes to 45 minutes so the same checks can complete sequentially; this does not relax any assertion or coverage threshold.

Superseded

Closed without merge after fresh live-state reconciliation found that PR #288 already carries the safer allocation remedy and preserves the long-lived Runtime check identities (Format, lint, test, and rustdoc, Production line coverage, and Production branch coverage). The available connector cannot prove that no organization ruleset still binds those identities, so replacing three check identities with this PR's single renamed job would introduce an avoidable governance risk. PR #288 has now been non-destructively reconciled with current protected main and remains the sole landing lane.

A fresh Devin finding on this branch also correctly notes that the collapsed job's bare if: failure() coverage diagnostics can fire after unrelated earlier failures and obscure the primary failure. That defect is not patched here because this branch is superseded; the thread is intentionally left unresolved rather than falsely marked addressed.

Acceptance

Do not revive this PR as a competing writer lane. Continue the root-cause work in #288, preserving live check identities and exact-head evidence.

@coderabbitai

coderabbitai Bot commented Aug 24, 2026

Copy link
Copy Markdown

Warning

Review limit reached

Next included review available in 49 minutes.

View limit details

Limit details: You’ve used the included review currently available.

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 4fcc6b6e-af84-4702-a2c0-8d5fd7e3c9a3

📥 Commits

Reviewing files that changed from the base of the PR and between 8c6b433 and 8a53ed1.

📒 Files selected for processing (2)
  • .github/workflows/ci.yml
  • tests/ci_contract.rs

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Devin Review found 1 potential issue.

Open in Devin Review

Comment thread .github/workflows/ci.yml

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

📝 Info: Coverage diagnostics fire on unrelated step failures

Collapsing the three jobs into one changes the two bare if: failure() diagnostics (ci.yml and ci.yml) from job-scoped to job-wide. Any earlier failure (cargo fmt, cargo clippy, cargo test) now triggers them, and they crash reading a coverage.json/coverage-branches.json that was never generated. The job still fails correctly, but the real cause gets buried under spurious tracebacks. Scoping each diagnostic to a specific step id (as the lockfile gate does at ci.yml) would restore the original behavior.

(Refers to this code)

Open in Devin Review

Was this helpful? React with 👍 or 👎 to provide feedback.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant