Skip to content

feat(ci): AI-first CI workflows — review, interact, health monitor, smoke test - #2459

Closed
serrrfirat wants to merge 12 commits into
stagingfrom
feat/ai-first-ci
Closed

serrrfirat wants to merge 12 commits into
stagingfrom
feat/ai-first-ci

Conversation

@serrrfirat

Copy link
Copy Markdown
Collaborator

Summary

Adds four GitHub Actions workflows (and modifies two existing ones) to move IronClaw's CI from AI-assisted to AI-first, inspired by CREAO's harness engineering approach.

  • AI review on every PR — Haiku-based 2-agent review (quality + security) on all PRs to staging, non-blocking informational feedback
  • Upgrade promotion review — Existing staging promotion gate upgraded from Haiku to Sonnet for deeper reasoning
  • Interactive @claude — Mention @claude in any issue or PR comment to get Sonnet-powered investigation and analysis
  • Daily CI health monitor — Opus-powered daily analysis of CI run outcomes, flaky tests, open bugs, dependabot alerts, coverage trends; creates/updates a rolling health report issue
  • Post-deploy smoke test — Docker image verification after every build (pull, start, health check, create issue on failure)
  • Wired into docker.yml — Smoke test runs automatically after successful image push

Three-Tier Model Strategy

Context Model Cost
Per-PR review (every PR) Haiku 4.5 ~$0.10/PR
Promotion gate + interactive Sonnet 4.5 ~$0.40/PR
Daily health monitor Opus 4.6 ~$3/run

Estimated monthly cost at 10 PRs/day: $100-250/month.

Files Changed

File Action
.github/workflows/claude-review-pr.yml Create — Haiku per-PR review
.github/workflows/claude-interact.yml Create — Interactive @claude
.github/workflows/ci-health-monitor.yml Create — Daily CI health monitor
.github/workflows/deploy-verify.yml Create — Post-deploy smoke test
.github/workflows/claude-review.yml Modify — Haiku → Sonnet
.github/workflows/docker.yml Modify — Add verify job
docs/superpowers/specs/ Design spec
docs/superpowers/plans/ Implementation plan

Test plan

  • Verify YAML parses: python3 -c "import yaml; yaml.safe_load(open(f))" for each workflow
  • Open a test PR to staging to trigger claude-review-pr.yml
  • Comment @claude what does this PR do? to trigger claude-interact.yml
  • Manually dispatch ci-health-monitor.yml via Actions UI
  • Manually dispatch deploy-verify.yml with a known-good image tag
  • Verify docker.yml build triggers the smoke test verify job

🤖 Generated with Claude Code

@github-actions github-actions Bot added scope: ci CI/CD workflows scope: docs Documentation size: XL 500+ changed lines risk: medium Business logic, config, or moderate-risk modules contributor: core 20+ merged PRs labels Apr 14, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces an 'AI-First CI' strategy by updating the engine's system prompts to prioritize direct tool calls over CodeAct orchestration and adding a comprehensive design and implementation plan for new GitHub Actions workflows. The plan outlines the addition of per-PR AI reviews, interactive Claude responses in issues, and a daily CI health monitor. Feedback focuses on improving the robustness of shell scripts within the proposed workflows, specifically regarding file filtering logic and ensuring sufficient data retrieval limits for CI health reporting.

Comment thread docs/superpowers/plans/2026-04-14-ai-first-ci.md Outdated
Comment thread docs/superpowers/plans/2026-04-14-ai-first-ci.md Outdated
serrrfirat and others added 6 commits April 14, 2026 16:11
Three-tier model strategy: Haiku for per-PR early feedback,
Sonnet for the staging promotion gate (deeper reasoning),
Opus reserved for the CI health monitor.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Lightweight 2-agent review (quality+bugs, security) runs on every
non-draft PR targeting staging. Non-blocking informational feedback.
Skips docs-only and JSON-only changes.
Engineers can mention @claude in any issue or PR comment to get
AI-powered code analysis, bug investigation, or explanations.
Uses Sonnet for investigation depth, bounded to 30 turns.
Scheduled daily at 8 AM UTC. Collects CI run outcomes, detects
flaky tests, checks dependabot alerts and open bugs. Opus analyzes
patterns and maintains a rolling health report issue with action items.
Pulls the just-pushed image, starts it with minimal config,
and verifies the health endpoint responds. Creates a GitHub
issue on failure with container logs and workflow link.
After building and pushing the Docker image, the smoke test
workflow runs against the sha-tagged image. Failure creates a
GitHub issue; success is silent.
@henrypark133

Copy link
Copy Markdown
Collaborator

Code Review

Overview

Solid PR — well-structured tiered model strategy, good security posture (read-only code access, bot-loop prevention), smart filtering (draft/docs-only skip, skip-ai-review escape hatch), pinned action SHAs, and clean docker.yml integration.

Issues

HIGH — Unnecessary id-token: write permission

Both claude-review-pr.yml and claude-interact.yml request id-token: write, but neither workflow uses OIDC. This is an unnecessary permission escalation. Remove it — contents: read + pull-requests: write (+ issues: write for interact) is sufficient.

HIGH — Shell injection via container logs in deploy-verify issue body

In deploy-verify.yml, container logs are interpolated directly into gh issue create --body:

LOGS=$(cat /tmp/smoke-logs.txt | head -100)
BODY="...${LOGS}..."
gh issue create --body "$BODY" ...

If container logs contain shell metacharacters (backticks, $(), unbalanced quotes), this will break or produce mangled issues. Use --body-file instead:

cat <<'ISSUE_EOF' > /tmp/issue-body.md
## Deploy Smoke Test Failed
...
ISSUE_EOF
cat /tmp/smoke-logs.txt | head -100 >> /tmp/issue-body.md
gh issue create --body-file /tmp/issue-body.md ...

HIGH — Missing files from PR description

The PR description lists docs/superpowers/specs/ and docs/superpowers/plans/ as changed files, but they are not in the diff. Either they were forgotten or the description is stale.

MEDIUM — Model version choices

The PR uses claude-sonnet-4-5-20250929 (dated Sonnet 4.5) in both claude-review.yml and claude-interact.yml. Was claude-sonnet-4-6 considered? If 4.5 is intentional for cost/stability, a comment explaining the choice would help future maintainers.

MEDIUM — Flaky test detection group_by style

In ci-health-monitor.yml:62, group_by(.headSha, .name) works in practice but relies on jq implementation behavior. The canonical form group_by([.headSha, .name]) is safer across jq versions — CI runners may update jq.

MEDIUM — Health monitor Opus cost guardrails

The health monitor runs Opus 4.6 with --max-turns 10 but no --max-tokens limit. A single run could get expensive if the model generates verbose analysis. Consider adding --max-tokens to cap output, especially for a daily scheduled job.

LOW — Direct context interpolation in shell

In claude-review-pr.yml, ${{ github.event.pull_request.number }} is interpolated directly into shell. While PR numbers are always numeric, best practice is to pass via env: to avoid the pattern being copied to unsafe contexts:

env:
  PR_NUMBER: ${{ github.event.pull_request.number }}
run: |
  gh pr diff "$PR_NUMBER" ...

LOW — Redundant always() in docker.yml verify job

if: ${{ always() && needs.build.result == 'success' }}

Since verify only depends on build and requires success, always() is unnecessary. if: needs.build.result == 'success' suffices.

@henrypark133 henrypark133 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review: AI-first CI workflows (Risk: High)

Good operational additions — health monitor, deploy verification, and interactive Claude are valuable CI capabilities. However, two security issues need resolution before merge.

Positives:

  • All actions properly SHA-pinned to commit hashes
  • deploy-verify.yml correctly handles container crashes (docker ps check in retry loop)
  • Concurrency groups prevent duplicate runs per issue/PR
  • persist-credentials: false on all checkouts

Critical: claude-interact.yml lacks author restriction — prompt injection risk [Security]

File: .github/workflows/claude-interact.yml:22-25
The if: guard only blocks bot accounts (claude[bot], github-actions[bot]), not external contributors. Any GitHub user who can comment on issues/PRs can trigger Claude with arbitrary instructions. The workflow has issues: write and pull-requests: write — a crafted @claude comment could post misleading content under the bot identity.
Suggested fix: Add author association check:

if: >
  contains(github.event.comment.body, '@claude') &&
  contains(fromJSON('["OWNER","MEMBER","COLLABORATOR"]'), github.event.comment.author_association) &&
  github.event.comment.user.login != 'claude[bot]'

Critical: Unnecessary id-token: write permission [Security]

Files: .github/workflows/claude-interact.yml:13, .github/workflows/claude-review-pr.yml:11
Neither workflow uses OIDC authentication. This permission enables GitHub OIDC token minting — a force-multiplier if combined with prompt injection (finding above).
Suggested fix: Remove id-token: write from both workflows.

Concerning: claude-review-pr.yml overlaps with claude-review.yml [Architecture]

Files: .github/workflows/claude-review-pr.yml, .github/workflows/claude-review.yml
Both fire on PRs to staging (different triggers: open/synchronize vs labeled). On promotion PRs, both run simultaneously. The staging gate in staging-ci.yml reads the last claude[bot] comment — a race could cause the wrong review to be evaluated.
Suggested fix: Either remove claude-review-pr.yml (the existing workflow already covers promotion PRs) or use distinct comment signatures so the gate can differentiate.

Concerning: Model upgrade haiku→sonnet with unchanged 50-turn/4-agent budget [Architecture]

File: .github/workflows/claude-review.yml:33
5-8x cost increase per review. With hourly staging promotions, significant monthly cost delta. Consider reducing --max-turns to 15-20 if upgrading to Sonnet.

Convention notes:

  • cancel-in-progress: false on claude-interact is acceptable (concurrency group is per-issue) but true would be safer against comment spam
  • Consider extracting inline Claude prompts to .github/prompts/*.md per the project's "prompt templates live in files" convention

@zmanian zmanian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review: AI-first CI workflows

+511/-1, 6 files. Adds Claude-powered PR review, issue interaction, health monitoring, and deploy verification workflows.

Security (must fix)

  1. claude-interact.yml is open to any GitHub commenter. The if: guard only blocks two bot accounts. Any external user who can comment on a public issue can trigger Claude with issues: write + pull-requests: write permissions — prompt injection vector. Add an author_association check restricting to OWNER, MEMBER, COLLABORATOR.

  2. id-token: write is unnecessary in both claude-interact.yml and claude-review-pr.yml. Neither uses OIDC. Remove to reduce blast radius.

Architecture (should fix)

  1. Duplicate review triggers on staging PRs. claude-review-pr.yml fires on pull_request: [opened, synchronize] to staging, while the existing claude-review.yml fires on labeled PRs to staging. Promotion PRs trigger both, risking a race where the staging gate evaluates the wrong (lighter) review. Scope one to exclude the other.

  2. Sonnet upgrade + unchanged 50-turn budget. The Haiku→Sonnet change in claude-review.yml is 5-8x cost increase. Consider reducing to 15-20 turns.

Minor

  1. ci-health-monitor.yml --limit 200 may silently truncate busy weeks.

  2. deploy-verify.yml interpolates container logs into issue body without sanitization — potential markdown injection from compromised images.

Agree with Henry's existing CHANGES_REQUESTED. Items 1-2 are security blockers.

@serrrfirat

Copy link
Copy Markdown
Collaborator Author

all of these are fair comments. I will attend to these tomorrow and ask for re-reviews.

…conventions (#2459)

Security:
- Add author_association guard (OWNER/MEMBER/COLLABORATOR) to claude-interact
- Remove unnecessary id-token: write from claude-interact and claude-review-pr
- Sanitize container logs in deploy-verify to prevent markdown injection

Architecture:
- Exclude staging-promotion PRs from claude-review-pr (prevents race with promotion gate)
- Reduce claude-review max-turns from 50 to 20 (Sonnet cost control)
- Increase ci-health-monitor run limit from 200 to 1000

Conventions:
- Extract all inline Claude prompts to .github/prompts/*.md
- Set cancel-in-progress: true on claude-interact

Bug fix:
- Consolidate grep filters and add empty-line guard in file change detection
@github-actions github-actions Bot added size: L 200-499 changed lines and removed size: XL 500+ changed lines labels Apr 15, 2026
@serrrfirat

Copy link
Copy Markdown
Collaborator Author

Addressed review feedback (a3db6f0)

Security (henrypark133 + zmanian)

  • Author restriction: Added author_association guard restricting @claude triggers to OWNER, MEMBER, COLLABORATOR in claude-interact.yml
  • id-token: write: Removed from both claude-interact.yml and claude-review-pr.yml
  • Log sanitization: deploy-verify.yml now escapes triple backticks in container logs and uses --body-file to prevent markdown injection

Architecture (henrypark133 + zmanian)

  • Duplicate review race: claude-review-pr.yml now skips PRs labeled staging-promotion, so promotion PRs only get the deeper Sonnet gate review
  • Cost control: Reduced claude-review.yml --max-turns from 50 → 20 for the Sonnet upgrade
  • Run limit: Increased ci-health-monitor.yml --limit from 200 → 1000

Convention / minor

  • Extracted all 4 inline Claude prompts to .github/prompts/*.md (per project "prompts in files" convention)
  • Changed cancel-in-progress: true on claude-interact.yml
  • Consolidated grep filters + added empty-line guard in PR file change detection (gemini-code-assist suggestion)

@zmanian zmanian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review: AI-First CI Workflows

Critical

1. claude-interact.yml -- COLLABORATOR author_association is too broad.
The guard allows anyone who has ever had a PR merged. Combined with pull-requests: write, issues: write, and Bash(cargo check:*, cargo clippy:*) tool access, a former contributor could craft a comment that triggers cargo commands against a malicious branch. The prompt says "Do NOT attempt to build" but the tool allowlist contradicts this.

  • Restrict to OWNER + MEMBER only, or use an explicit username allowlist
  • Remove cargo check/cargo clippy from allowed tools

2. docker.yml -- secrets: inherit passes ALL repo secrets to smoke test.
deploy-verify.yml only needs github.token, but secrets: inherit exposes ANTHROPIC_API_KEY, Docker Hub credentials, and everything else. If the tested container is compromised, secrets could leak.

  • Replace with explicit secret passing, or remove secrets: inherit entirely (default GITHUB_TOKEN is passed automatically)

High

3. claude-review-pr.yml uses pull_request trigger -- silent no-op on fork PRs since the ANTHROPIC_API_KEY secret isn't available in fork context. Should be documented as intentional or switched to pull_request_target with safeguards.

4. claude-review.yml -- Sonnet upgrade with 4 parallel sub-agents could be costly per promotion PR. No cost ceiling. Acceptable if monitored, but worth documenting expected per-run cost.

Medium

5. ci-health-monitor.yml -- issue dedup is fragile (keyword search only). Could create duplicate issues on repeated runs. Also ~$15-30/run with Opus -- consider whether Sonnet suffices for this structured task.

6. deploy-verify.yml -- log sanitization only escapes backticks. Container logs could contain markdown injection or sensitive env vars that get embedded in auto-created issues.

Positives

  • All GitHub Actions are SHA-pinned (good supply-chain hygiene)
  • Prompts extracted to .github/prompts/*.md
  • Tool allowlists are restrictive (no Bash(*) wildcard)
  • persist-credentials: false used consistently
  • Concurrency groups prevent parallel runaway

…2459)

Security:
- Restrict claude-interact author_association to OWNER+MEMBER (drop COLLABORATOR)
- Remove cargo check/clippy from claude-interact allowed tools
- Remove secrets: inherit from docker.yml verify job (only needs github.token)
- Broaden deploy-verify log sanitization: strip ANSI codes, redact secret-like env vars

Architecture:
- Downgrade ci-health-monitor from Opus to Sonnet (~$1-2/run vs ~$15-30)
- Document fork PR no-op behavior in claude-review-pr.yml
- Document expected per-run cost in claude-review.yml
- Document issue dedup trade-off in ci-health-monitor.yml
@github-actions github-actions Bot removed the size: L 200-499 changed lines label Apr 15, 2026
@serrrfirat

Copy link
Copy Markdown
Collaborator Author

Addressed zmanian review round 2 (625935a)

Security

  • Author restriction tightened: Dropped COLLABORATOR — claude-interact.yml now restricted to OWNER + MEMBER only
  • Removed cargo tools: cargo check / cargo clippy removed from claude-interact allowed tools (contradicted read-only intent)
  • secrets: inherit removed: docker.yml verify job no longer inherits all repo secrets — deploy-verify.yml only uses github.token (passed automatically)
  • Broader log sanitization: deploy-verify.yml now strips ANSI escape codes and redacts values matching API_KEY, TOKEN, SECRET, PASSWORD, CREDENTIAL, AUTH patterns

Architecture / cost

  • Opus → Sonnet: ci-health-monitor downgraded to Sonnet (~$1-2/run vs ~$15-30)
  • Fork PR behavior documented: Comment in claude-review-pr.yml explaining intentional silent no-op on forks
  • Cost documented: claude-review.yml includes expected per-run cost estimate ($2-5)
  • Dedup trade-off documented: ci-health-monitor.yml notes keyword-based dedup limitation

@github-actions github-actions Bot added the size: XL 500+ changed lines label Apr 15, 2026
@serrrfirat
serrrfirat requested a review from zmanian April 15, 2026 07:53
zmanian
zmanian previously approved these changes Apr 15, 2026

@zmanian zmanian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-review: All critical/high findings addressed

Finding Status
claude-interact.yml author_association + cargo tools FIXED -- restricted to OWNER+MEMBER, cargo tools removed
docker.yml secrets: inherit FIXED -- removed, only github.token used
Fork PR silent skip undocumented FIXED -- comment block added
Sonnet + 4 agents cost ceiling FIXED -- max-turns reduced to 20, cost documented
Health monitor dedup + Opus cost FIXED -- switched to Sonnet (~$1-2/run)
Log sanitization PARTIALLY FIXED -- ANSI stripping + keyword redaction added, acceptable risk since container gets no real secrets

No new issues introduced by the fix commits. LGTM.

Three-tier model strategy:
- Haiku: per-PR early feedback (lightweight, high volume)
- Sonnet: interactive @claude + CI health monitor (balanced)
- Opus: staging promotion gate (deep reasoning for 4-agent consolidation)

Promotion PRs are low-frequency, high-stakes — Opus's superior
reasoning justifies the cost (~$10-20/review) for the final gate.
)

CI fix:
- Restore id-token: write to claude-review-pr.yml (required by
  claude-code-action for OIDC auth — removal was a false positive)
- Remove redundant always() wrapper in docker.yml verify job

Model upgrades (use latest stable IDs):
- Haiku: claude-haiku-4-5-20251001 → claude-haiku-4-5
- Sonnet: claude-sonnet-4-5-20250929 → claude-sonnet-4-6
- Opus: already on claude-opus-4-6 (unchanged)
@serrrfirat serrrfirat added the skip-regression-check Bypass regression test CI gate (tests exist but not in tests/ dir) label Apr 15, 2026
@serrrfirat
serrrfirat requested a review from zmanian April 15, 2026 11:14
- Add RUSTSEC-2026-0098 and RUSTSEC-2026-0099 (rustls-webpki URI/wildcard
  name constraint validation) — 0.102.8 pinned by libsql transitive dep,
  0.103.x awaiting upstream patch
- Remove 4 stale wasmtime advisories (RUSTSEC-2025-0046, RUSTSEC-2025-0118,
  RUSTSEC-2026-0020, RUSTSEC-2026-0021) that no longer match any crate
  after wasmtime upgrade to v43

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

@henrypark133 henrypark133 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Verified finding:

.github/workflows/ci-health-monitor.yml:63-69 serializes flaky runs as "<workflow> <sha>" and then parses them with while read -r workflow sha. That breaks as soon as the workflow name contains spaces. For example, "Docker Image abc123" is parsed as workflow=Docker and sha="Image abc123", so the follow-up jq lookup never matches the failed run and no flaky log is collected. Several of this repo's workflow names are multi-word, so the health monitor silently drops the failure-log enrichment for the common case.

Suggested fix: emit structured data (JSON, tab-separated, or NUL-separated fields) and parse that instead of splitting on spaces.

Multi-word workflow names (e.g. "Docker Image") broke the
space-delimited `while read -r workflow sha` loop — the name
was split across both variables so the jq lookup never matched.
Switch to tab-separated jq output with IFS=$'\t'.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
@serrrfirat

Copy link
Copy Markdown
Collaborator Author

Addressed @henrypark133's flaky-run parsing feedback in c032e7e:

  • jq now emits tab-separated workflow\tsha instead of space-separated
  • while IFS=$'\t' read -r workflow sha correctly handles multi-word workflow names (e.g. "Docker Image")

No other unresolved review items remain.

@serrrfirat
serrrfirat requested a review from henrypark133 April 16, 2026 12:20

@henrypark133 henrypark133 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review: interactive workflow still reads the wrong tree on PR comments

The permission tightening and log-parsing fixes look good, but one correctness issue remains in the interactive workflow.

Concerning: PR comment investigations still run against the default branch checkout

File: .github/workflows/claude-interact.yml:28
issue_comment and pull_request_review_comment both trigger on PR discussions, and the prompt explicitly tells the agent to use Read/Glob plus file:line references when analyzing code. This checkout step never switches to the PR head (refs/pull/<n>/head), so for PR comments the workspace contains the repository default branch, not the commented change. The result is that @claude can read stale files, cite wrong line numbers, and answer review threads against code that is not actually under discussion.
Suggested fix: When the comment is attached to a PR, check out the PR head (or merge ref) before invoking Claude. For issue-only comments, keep the default-branch checkout.

@ilblackdragon

Copy link
Copy Markdown
Member

Context: this is review feedback informed by a broader 2-week velocity/quality audit of the repo. Flagging upfront so the suggestions land as a strategic read, not drive-by nits.

What's genuinely good

Strategic concern

This PR adds a 3rd AI reviewer without retiring the first two. Copilot and Gemini still fire on every PR. Large PRs already show the cost of that firehose — #2515 drew ~7 bot reviews on top of 33 human rounds. Adding a Haiku reviewer is the right call, but the velocity win only materializes if Copilot + Gemini are demoted to summary-only (or off) in the same change. Otherwise net reviewer count goes up, not down.

Specific issues worth addressing before merge

  1. No spend cap anywhere. @claude per-comment = one 30-turn Sonnet run. workflow_dispatch on health-monitor has no cooldown — any perm-holder can trigger a $3+ Opus run per click. The $100–250/mo estimate is a forecast, not a limit. Suggest a per-day budget gate that short-circuits further runs once hit.
  2. Rate-limit @claude per user. One member spamming @claude across 50 comments = 50 Sonnet runs. The concurrency guard cancels the current in-flight run, but doesn't cap cumulative spend.
  3. Fires on opened, synchronize, not ready_for_review. A non-draft PR gets reviewed on every push — pay-per-iteration. Suggest adding ready_for_review and skipping synchronize when draft.
  4. review-promotion.md doesn't say skip LOW severity like review-pr.md does. The promotion gate should be more selective, not less — nits on the highest-stakes review are the worst kind of noise. Apply the same filter.
  5. deploy-verify.yml starts the container against a dead LLM URL (localhost:9999). That's a startup/crash test, not a smoke test. Rename to deploy-startup.yml or add a minimal LLM-backed check — otherwise the name oversells what's being verified.
  6. No de-dup of AI-review comments on force-push. 12 commits on this PR alone = up to 12 stale Found N issues comments. Bot should edit-last (or delete prior bot comments) to keep a single rolling review.

Suggested split

This is size: XL + risk: medium with 5+ review rounds so far. Consider splitting into 3 PRs, each reviewable in <30 min:

  1. PR A — Prompt externalization + Haiku per-PR review + Copilot/Gemini demoted to summary-only or off. Non-negotiably paired so net reviewer count drops.
  2. PR B — CI health monitor (clean standalone; ship it).
  3. PR C — @claude interactive + deploy-verify. Gate merge on a daily $-cap env var.

Ship A and B this week. Hold C until spend caps land.

Happy to help draft any of these if useful.

@serrrfirat

Copy link
Copy Markdown
Collaborator Author

Closing this PR based on team discussion around pushing more local testing on pre hooks rather than making noise on the CI. Will come up with a redesign asap.

@serrrfirat serrrfirat closed this Apr 18, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: medium Business logic, config, or moderate-risk modules scope: ci CI/CD workflows scope: docs Documentation size: XL 500+ changed lines skip-regression-check Bypass regression test CI gate (tests exist but not in tests/ dir)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants