Repository navigation
feat: implement issue #1419 — [docs] metrics-baseline denominator note misstates the ~12% figure's basis — claims consistency where populations differ - #1429
Conversation
…e misstates the ~12% figure's basis — claims consistency where populations differ
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
🤖 CodeAnt AI — Review Status
|
|
Warning Review limit reached
Next review available in: 18 minutes Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available. How can I continue?After more reviews become available, a review can be triggered using the To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews. How do review limits work?CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability. For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window. Please refer docs for additional details. Review details⚙️ Run configurationConfiguration used: Organization UI Review profile: ASSERTIVE Plan: Pro Plus Run ID: 📒 Files selected for processing (1)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
PR Summary by QodoDocs: correct metrics-baseline note on ~12% noise estimate denominator scope
AI Description
High-Level Assessment
Files changed (1)
|
There was a problem hiding this comment.
Code Review
This pull request updates docs/metrics-baseline.md to add a dated correction clarifying the denominator scope of the ~12% noise estimate, explaining why the original claim of denominator consistency was incorrect. The review feedback suggests clarifying the distinction between different comment and review totals (180 vs. 254) to prevent confusion, which is a helpful improvement.
Dev-Lead — review-changes (no-changes)No changes were needed for this PR. |
|
Advisory bots were rate-limited; auto-approval is withheld until they recover. pr-review-sweep will re-review this PR after 2026-08-02T09:46:20Z. |
Dev-Lead — fix-reviews (applied)Changes committed and pushed. |
Code Review by Qodo
Context used✅ Compliance rules (platform):
50 rules 1.
|
don-petry
left a comment
There was a problem hiding this comment.
Review — PR #1429 (#1419 baseline denominator correction)
No findings. This is exactly the correction the issue asked for, and handled better than I specified.
Every AC satisfied, verified against the branch:
- AC #1 — the ~12% figure's real basis is now stated: all PR comments in a sampled window, third-party bots included, classified by phrase scan. Provenance recorded precisely (10-PR sample;
total comments = 180;est. no-action = 21→ 21/180 ≈ 12%), and the254/254 machine-authored, 0 humanfigure is carried with the explanation that it is only meaningful because third-party bots were in scope. I derived these numbers originally; they are reproduced correctly, including naming the six bots. - AC #2 — the false sentence is struck, and the doc now says plainly that the denominators are not consistent and that the first measured value is expected to differ.
- AC #3 — the interpretation rule is there, and sharper than I asked for: a population difference is "not a regression and not an improvement", and it defines what would count as a genuine correction (the classifier disagreeing with a hand-count at a fixed first-party denominator), so a future reader cannot mistake the gap itself for a change in noise.
- AC #4 — answered explicitly: no all-comments series is maintained, the first-party share is the only noise time-series here, and the reviewer scorecard is the only third-party view.
- AC #5 — the dated-append convention is honoured, and beyond the letter of the ask: the wrong sentence is struck in place with a visible
[Corrected …]marker rather than deleted, so the error and its correction are both legible. That is the right call for a document whose entire purpose is to be a trustworthy "before".
The shape here is now a model for the other corrections in this epic — including my own. Ready from my side.
|
CI checks on this PR are still running. Once they complete, re-mention Posted by the donpetry-bot PR-review cascade. |
Dev-Lead — fix-reviews (applied)Changes committed and pushed. |
|
Dev-Lead — waiting on PR blockers (intent: review-changes)PR: #1429 |
|
Note @don-petry I reviewed this PR and no code changes were needed, but it still has blocking checks or reviews (failing or cancelled checks, or changes-requested reviews), so I cannot mark it done yet. I'll re-check automatically. |
|
Advisory bots were rate-limited; auto-approval is withheld until they recover. pr-review-sweep will re-review this PR after 2026-08-02T09:57:27Z. |
donpetry-bot
left a comment
There was a problem hiding this comment.
Automated review — APPROVED ✓
Risk: LOW
Reviewed commit: d74fc0e50bb1f26c12d46aec53f6c1a8c92cc8aa
Review mode: triage-approved (single reviewer)
Summary
Docs-only PR (single file: docs/metrics-baseline.md, +67/-19) correcting the documented basis of the ~12% reviewer-noise estimate. The triage tier's low-risk assessment is confirmed: no code, workflow, script, or security surface is touched. The correction accurately states the ~12% figure was an all-comments phrase-scan estimate (third-party bots included), not a first-party marker count, and follows the doc's dated-append convention.
Linked issue analysis
Closes #1419. All five acceptance criteria are substantively addressed in the diff: (1) the ~12% figure's actual basis is stated with full provenance (10-PR sample, 21/180 phrase-scan, 254 machine-authored review activities); (2) the incorrect 'denominators are consistent' sentence is struck in place and corrected — the doc now states the classifier measures a narrower population and a first-run divergence is a population change, not a regression/improvement; (3) a clear interpretation rule distinguishes genuine measurement corrections from the known population difference; (4) the doc explicitly states no all-comments series is maintained and names the reviewer scorecard as the only third-party view; (5) the correction is a dated append (Correction 2026-08-02) with the original text retained struck-through, per the doc's convention.
Findings
No blocking findings. Earlier bot-flagged issues (Qodo/Graphite: 180-vs-254 inconsistency and contradictory ~12% guidance) were fixed by dev-lead follow-up commits — the doc now distinguishes 180 comment bodies (phrase-scan denominator) from 254 distinct review activities — and all 4 review threads are resolved. Secret-scanning MCP tool unavailable in this environment; gitleaks CI check is green and the diff contains no secret-like content (markdown prose only).
CI status
All substantive checks green: Lint, ShellCheck, CodeQL, Agent Security Scan, agent-shield, Secret scan (gitleaks), SonarCloud (quality gate passed), unit-tests, holdout-guard, Compile agentic workflows, CodeRabbit, Graphite. The five CANCELLED entries are superseded agent-orchestration jobs (dev-lead dispatch/ci-relay, PR-review trigger runs), not failing CI. Dependency-audit jobs skipped (no matching ecosystems).
Reviewed automatically by the PR-review agent (single-reviewer mode: fable 5). Reply if you need a human review.



User description
Closes #1419
Implemented by dev-lead agent. Please review.
CodeAnt-AI Description
Correct the documented basis of the reviewer-noise estimate
What Changed
Impact
✅ Accurate reviewer-noise comparisons✅ Clearer interpretation of first-party metrics✅ Fewer false regression or improvement claims💡 Usage Guide
Checking Your Pull Request
Every time you make a pull request, our system automatically looks through it. We check for security issues, mistakes in how you're setting up your infrastructure, and common code problems. We do this to make sure your changes are solid and won't cause any trouble later.
Talking to CodeAnt AI
Got a question or need a hand with something in your pull request? You can easily get in touch with CodeAnt AI right here. Just type the following in a comment on your pull request, and replace "Your question here" with whatever you want to ask:
This lets you have a chat with CodeAnt AI about your pull request, making it easier to understand and improve your code.
Example
Preserve Org Learnings with CodeAnt
You can record team preferences so CodeAnt AI applies them in future reviews. Reply directly to the specific CodeAnt AI suggestion (in the same thread) and replace "Your feedback here" with your input:
This helps CodeAnt AI learn and adapt to your team's coding style and standards.
Example
Retrigger review
Ask CodeAnt AI to review the PR again, by typing:
Check Your Repository Health
To analyze the health of your code repository, visit our dashboard at https://app.codeant.ai. This tool helps you identify potential issues and areas for improvement in your codebase, ensuring your repository maintains high standards of code health.