Repository navigation
ci: make deep-review label reviews investigate like a prompted session - #347
Conversation
|
Warning Review limit reached
Next review available in: 54 minutes Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available. How can I continue?After more reviews become available, a review can be triggered using the To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews. How do review limits work?CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability. For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window. Please refer docs for additional details. Review details⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Pro Plus Run ID: 📒 Files selected for processing (3)
WalkthroughThe deep-review workflow now accepts verified ChangesDeep-review workflow
Estimated code review effort: 3 (Moderate) | ~25 minutes Sequence Diagram(s)sequenceDiagram
participant Maintainer
participant PRComments
participant CollaboratorsAPI
participant ClaudeReview
participant Claude
Maintainer->>PRComments: Post deep-review focus comment
ClaudeReview->>PRComments: Fetch and deduplicate comments
ClaudeReview->>CollaboratorsAPI: Verify comment authors
CollaboratorsAPI-->>ClaudeReview: Return author permissions
ClaudeReview->>Claude: Inject focus and review instructions
Claude-->>ClaudeReview: Return blocking/advisory summary
Possibly related PRs
Suggested labels: 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Preview DeploymentPreview URL: https://84b65f19.bestax.pages.dev |
|
@claude perform a deep review |
|
Claude finished @allxsmith's task in 5m 10s —— View job Deep review
Evidence check on the PR's central claim: pulled Verified locally:
🟠 Major · Correctness — docs describe a feature that doesn't exist in this repo
None of this exists in the repository as of this diff:
A maintainer who follows this doc's own instruction — 🔵 Advisory · Robustness — the 5-distinct-author cap can bury a genuine steer
Overall: The prompt-engineering changes are the well-evidenced core of this PR and hold up — they're a direct, cited fix for the exact gap #341 describes, and the new focus-steer step's trust logic (live role re-verification, newest-per-author dedup, random-delimiter heredoc against injection) is sound on manual trace. The one real defect is scope creep that shipped broken: the "Screenshots at handoff" doc section describes tooling that isn't in this repo. Residual risk: Beyond the two items above, I didn't find a way the original #340 failure mode (deep-review missing evidence-backed residual bugs) could still occur through this diff — the new steps 2-4 explicitly require fetching cited evidence and enumerating residual-risk variants before judging code, which is exactly the gap that let #340's Finding 1 through. |
0507832 to
10a8b3b
Compare
Preview DeploymentPreview URL: https://128f7906.bestax.pages.dev |
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@docs/docs/guides/getting-started/ai-development.md`:
- Around line 90-100: Update the “Screenshots at handoff” section to remove
unsupported claims about story-screenshots.yml, its pr input, artifact
retention, story-screenshots branch, and “nothing to screenshot” behavior unless
corresponding workflow source is checked into the repository; otherwise replace
the details with behavior verifiable from existing checked-in sources.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: CHILL
Plan: Pro Plus
Run ID: ea650ed6-1263-42b6-bdbb-1dd2464ce89c
📒 Files selected for processing (3)
.github/workflows/claude-review.ymlCLAUDE.mddocs/docs/guides/getting-started/ai-development.md
🚧 Files skipped from review as they are similar to previous changes (2)
- CLAUDE.md
- .github/workflows/claude-review.yml
The deep-review label prompt validated diffs in isolation (0 findings on PR #340) while a prompted @claude session on the same diff found 5, including a real residual bug. Close the gap in what the reviewer is told to do, not the model: - Evidence phase: chase PR claims into linked issues and run/job logs, then verify the fix empirically — reading the diff is not verification. - Residual-risk hunt: enumerate how the addressed failure class could still occur; refute with evidence or post as a finding. A Residual risk section is now required in the summary review. - Advisory tier: 🔵 Advisory for real limitations/trade-offs worth putting on the record. Summary-table only — never inline, so advisory notes cannot become fixer work items and spin the AI loop. - PR-type-aware lens: keep the component checklist for bulma-ui/docs diffs; add an infra lens (guard bypasses, shell+jq edge cases, untrusted input, silent failures, idempotency) for .github/.claude/ scripts diffs. - Optional focus steer: a triage+ user may pre-post a PR comment starting with deep-review: — a new gate step verifies each candidate author's live role (newest steer per distinct author, newest-first, max 5 role checks, so an outsider can neither inject text nor displace a maintainer's steer), and injects it as a FOCUS block via a random-delimiter heredoc output. Fail-soft by design: API failures degrade to an unfocused review, never a red run. Unchanged: opus model, 120 turns, dedupe marker, one-review invariant, allowed_bots, show_full_output debug-only, the is_error gate. Docs: mention the steer in the AI development guide's label table and CLAUDE.md's deep-review sentence. Fixes #341
10a8b3b to
0497843
Compare
Preview DeploymentPreview URL: https://d4e97c5f.bestax.pages.dev |
|
@claude run another deep review |
|
Claude finished @allxsmith's task in 6m 49s —— View job Deep review (round 2)
Evidence check. Current HEAD is Re-traced the
No new blocking defect found in this step. One pre-existing, already-disclosed advisory still applies and is unchanged by this round: the 5-distinct-author cap is applied before the permission check, so a burst of ≥5 other One additional minor/non-blocking observation (not worth a blocking tag): Residual risk: I did not find a way the original #340/#341 failure mode (deep-review missing evidence-backed residual bugs because it was never told to hunt for them) could still occur through this diff. The new steps 2–4 explicitly require fetching cited evidence and enumerating residual-risk variants before judging code — the same gap the first deep-review round on this PR (using the old, unstructured prompting style) actually caught: it found the real "Screenshots at handoff" documentation-of-nonexistent-feature bug and it's now fixed. Overall: No blocking findings this round. The one prior blocking issue is resolved by removal; the one prior advisory (author-cap-before-permission-check ordering) remains a known, accepted trade-off, not a regression introduced since. This PR is in a mergeable state as far as this review is concerned — the only irreducible caveat (already called out in the PR body) is that the label-triggered workflow path itself can't be exercised against its own PR before merge, so the definitive test remains the planned post-merge re-run against PR #340. |
|
🎉 This PR is included in version 3.6.0 🎉 The release is available on: Your semantic-release bot 📦🚀 |
|
🎉 This PR is included in version 2.0.0 🎉 The release is available on: Your semantic-release bot 📦🚀 |
|
🎉 This PR is included in version 5.8.1 🎉 The release is available on: Your semantic-release bot 📦🚀 |
|
🎉 This PR is included in version 1.0.0 🎉 The release is available on: Your semantic-release bot 📦🚀 |
Fixes #341
On PR #340 the
deep-reviewlabel path (opus, 120 turns) found 0 findings while a prompted@claudesession (sonnet, 60 turns) found 5, including a real residual bug fixed ind2ffb51. The gap was what each session was told to do, not the model. This PR closes it inclaude-review.yml's prompt plus one new gate step — no new tools, no new write surfaces.Prompt changes
gh apirun/job logs — already allowlisted) and verify the fix against it empirically. Reading the diff is not verification.bulma-ui//docs/diffs; a new infra lens (guard/permission bypasses, shell+jq edge cases, untrusted-input handling, silent-failure shapes, idempotency/concurrency) applies to.github//.claude//scripts/diffs.reviewThreads, which only inline comments create). Zero-blocking summaries list advisories instead of a bare "No blocking defects found."; the heading becomes<N> blocking · <M> advisory.Focus steer (new
Extract focus steerstep)A triage+ user may pre-post a PR comment starting with
deep-review:before applying the label; it's injected into the prompt as a FOCUS block. Security posture, since anyone can comment on a public-repo PR and the text enters the LLM prompt:EOF,key=value, …) can terminate the value early or smuggle step outputs.pipefail,jq -sstill prints[]whenghdies, so an inner|| echowould concatenate both and corrupt the count).Unchanged: opus model, 120 turns, dedupe marker, one-review invariant,
allowed_bots,show_full_outputdebug-only, theis_errorgate.Docs: the steer is mentioned in the AI-development guide's label table and CLAUDE.md's deep-review sentence.
Verification
check:conformancegreen on changed files.EOF/allowed=false/block=/fake-delimiter lines stay inside the value; onlyblocksurfaces as an output key), newest-own-steer-wins, 5-author cap behavior, truncation, CRLF, whitespace-only, comments-API failure (exit 0, empty block), per-author permission-API failure fall-through, and a live no-steer run against PR ci: fail AI triage loudly when the session bails before posting (#338) #340.seq 0 -1counts down (empty candidate list would loop on a Mac dev box; replaced with awhilecounter) and the fail-soft||placement above. Side finding, documented in a comment: on a public repo the collaborators-permission endpoint returnsreadfor any user rather than 404, which the case statement already handles.Post-merge (this workflow can't run on its own PR — the action's OIDC exchange requires the file to match main): re-apply
deep-reviewto a PR fixing a documented failure (#340 is the benchmark), with adeep-review:steer comment posted first, and confirm the review (a) cites the linked issue's evidence, (b) contains a Residual risk section, (c) reports advisories in the summary without creating inline threads.Summary by CodeRabbit
New Features
deep-review:-prefixed pull request comments.Documentation