Skip to content

ci: read CLA policy snapshots through git to spare the GITHUB_TOKEN budget - #16225

Open
lawrencecchen wants to merge 1 commit into
mainfrom
fix-cla-guard-git-snapshots
Open

lawrencecchen wants to merge 1 commit into
mainfrom
fix-cla-guard-git-snapshots

Conversation

@lawrencecchen

@lawrencecchen lawrencecchen commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Summary

CLA policy guard runs on every pull_request_target event, including edited. Between 12:00 and 13:00 PDT on 2026-09-30 it ran 494 times for 264 head SHAs. 175 of those runs came from review bots editing PR descriptions (cubic-dev-ai 134, coderabbitai 29, blacksmith 12). Each run fetched twelve policy files through the REST contents API, for 14 GITHUB_TOKEN requests a run. That is about 6,900 requests an hour, out of a 15,000-an-hour budget that every workflow in the repository shares. After ci-fail-fast.yml was removed (#16160), this guard was the largest remaining consumer. When the budget ran out, the guard failed on every PR with a bare 403.

fetch_snapshot now reads each revision through git. It fetches the commit's trees into a blob-less bare repository, with --depth=1 --filter=blob:none and one promisor remote per repository. It then fetches only the policy blobs, by object id. Git transfers are not charged to the REST budget. Objects are addressed by SHA, so the bytes are the ones the contents API returned. I checked this for a main base, a fork head (#16198 from wanjinhao1/cmux) and a test-merge commit. The public repositories are fetched anonymously, and no credential is written. A missing path is still nil. A directory or symlink is still "not a regular file". The size limit is checked before the blob is read. One difference: a revision that cannot be fetched now fails the check. Before, every file silently read as missing.

If the budget is exhausted anyway, api_json asks GET /rate_limit for the reset time (that endpoint is not charged). If the reset is at most five minutes away, it waits and retries once. Otherwise it fails with the reset time rather than a bare 403.

This changes the guard script, so the guard running from main requires a trusted reviewer's approval of this exact head (@austinywang or @azooz2003-bit). The check stays red until one of them approves. EXPECTED_GUARD_SCRIPT_DIGEST is updated to the new self-digest.

Companion: #16224 cuts the merge-group watcher's polling.

Testing

  • REST requests per run, counted with a gh shim against live PRs (fix: recover interrupted Cloud vm run creates #16221, same-repo; Fix #16193: refuse case-only duplicate label names in the labels manifest #16198, fork): 14 before, 2 after (repos/{repo} and pulls/{n}). Wall time is about 4 s. The full validator, with all its regression matrices, passes on Ruby 2.6 and 4.0.
  • Snapshot equivalence: the old contents-API fetch_snapshot and the new git one gave identical results for the 7 policy paths, including missing paths, at 4 revisions.
  • Rate-limit path: with a fake gh whose first call returns "API rate limit exceeded", the validator waits for the reported reset and then succeeds. With the wait cap set to 0, it fails with "exhausted until HH:MM UTC".
  • Not verified: the guard-changed and test-merge branches running on a real runner. This PR's own guard run exercises them once it is approved.

Changelog

none

🤖 Generated with Claude Code


View with [code]smith Autofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.


Summary by cubic

Reduces the CLA policy guard's GITHUB_TOKEN usage by reading policy snapshots through git instead of the REST contents API, dropping a run from 14 to 2 REST requests.

  • Transfers policy blobs by object id from blob-less bare repositories, which is not charged to the shared rate-limit budget.
  • On rate-limit exhaustion, waits for the reported reset when it is at most five minutes away, else fails with the reset time.
  • A revision that cannot be fetched now fails the check instead of silently reading every file as missing.

Written for commit fdcf940. Summary will update on new commits.

Review in cubic

The CLA policy guard runs on every pull_request_target event, including
review-bot description edits (494 runs from 12:00 to 13:00 PDT on
2026-09-30), and fetched twelve policy files per run through the REST
contents API: 14 GITHUB_TOKEN requests a run, the largest share of the
repository's shared budget once ci-fail-fast.yml was gone. Fetch each
revision's tree into a blob-less bare repository and its policy blobs
by object id instead. Git transfers are not charged to that budget, and
the bytes are identical (checked for base, fork head and test-merge
revisions). A run now makes 2 REST requests.

When the budget is spent anyway, wait for the reset GET /rate_limit
reports (it is not charged) if it is at most five minutes away, and
otherwise fail with the reset time instead of a bare 403.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@cursor

cursor Bot commented Sep 30, 2026

Copy link
Copy Markdown

Bugbot is paused — on-demand spend limit reached

Bugbot uses usage-based billing for this team and has hit its on-demand spend limit.

A team admin can raise the spend limit in the Cursor dashboard, or wait for the next billing cycle to continue.

@coderabbitai

coderabbitai Bot commented Sep 30, 2026

Copy link
Copy Markdown

Warning

Review limit reached

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

Next included review available in 28 minutes.

Check out review usage here.

View limit details

Limit details: You’ve used all 10 included reviews currently available.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: Repository: manaflow-ai/cmux/.coderabbit.yaml

Review profile: ASSERTIVE

Plan: Advanced

Run ID: a588d20f-5043-4104-b389-1a48600bf773

📥 Commits

Reviewing files that changed from the base of the PR and between 3016cf3 and fdcf940.

📒 Files selected for processing (1)
  • scripts/ci/validate-cla-policy.rb
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions

Copy link
Copy Markdown
Contributor

All contributors have signed the CLA ✍️ ✅
Posted by the CLA Assistant Lite bot.

@teamleaderleo

Copy link
Copy Markdown
Collaborator

Independent confirmation of the problem, plus one finding on the fix.

Confirming the impact from the other side. I hit this while attributing
twelve reds on an unrelated PR. Counting the CLA policy guard workflow
(347343695) by creation window rather than by gh run list, which is dominated
by whatever is newest:

gh api --paginate "repos/manaflow-ai/cmux/actions/workflows/347343695/runs?created=2026-09-30T19:20..2026-09-30T20:00&per_page=100"

151 of 288 runs failed in those forty minutes, across 51 distinct head branches.
That matches your 12:00 to 13:00 PDT window and the budget exhaustion you
describe. Reading policy revisions through git rather than the contents API
looks like the correct fix, and it removes the largest consumer rather than
rationing it.

Finding: the new rate-limit diagnostic cannot reach the annotation.
fail! raises PolicyError (line 417), and the exhaustion branch calls it:

fail!("the repository's GITHUB_TOKEN rate limit is exhausted until #{reset_at}; re-run this check after that (#{endpoint})")

but the script ends with

rescue PolicyError
  warn "::error::CLA policy validation rejected the proposed policy"

So after the single wait-and-retry, a still-exhausted budget is annotated as a
rejected policy. That is the specific misdiagnosis that cost me time today: the
check asserts the branch proposed a bad CLA policy when nothing about the branch
is at fault, and the comment above that rescue explains why no detail can be
added, so there is nothing for a maintainer to go on. The same applies to the
generic fail!("GitHub API request failed for ...") on line 483, which turns
any 403, 502 or network error into a policy rejection.

The classification you want already exists one line below:

rescue StandardError
  warn "::error::CLA policy validation could not complete"

Suggested fix, small and contained: add an error class for infrastructure
failures and raise it from the two API paths instead of PolicyError.

class InfrastructureError < StandardError; end

def infra_fail!(message)
  raise InfrastructureError, message
end

then rescue InfrastructureError alongside the existing StandardError branch,
or simply let it fall through to it. This keeps the annotation generic, so it
does not weaken the rule about not copying candidate-controlled diagnostics into
a public check, while telling a maintainer whether to fix their branch or re-run
the check. Worth doing here since this PR already touches both paths.

Not a blocker on the approach, which I think is right.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants