Skip to content

feat: implement issue #1419 — [docs] metrics-baseline denominator note misstates the ~12% figure's basis — claims consistency where populations differ - #1429

Merged
don-petry merged 3 commits into
mainfrom
dev-lead/issue-1419-20260802-0840
Aug 2, 2026
Merged

don-petry merged 3 commits into
mainfrom
dev-lead/issue-1419-20260802-0840

Conversation

@don-petry

@don-petry don-petry commented Aug 2, 2026 •

Copy link
Copy Markdown
Collaborator

User description

Closes #1419

Implemented by dev-lead agent. Please review.


CodeAnt-AI Description

Correct the documented basis of the reviewer-noise estimate

What Changed

  • Corrects the documentation to state that the approximate 12% no-action estimate included all PR comments, including third-party reviewer bots
  • Explains that the current first-party-only measurement uses a different population and should not be compared directly with the 12% estimate
  • Clarifies how to interpret the first scheduled measurement and directs readers to the reviewer scorecard for all-comment coverage

Impact

✅ Accurate reviewer-noise comparisons
✅ Clearer interpretation of first-party metrics
✅ Fewer false regression or improvement claims

💡 Usage Guide

Checking Your Pull Request

Every time you make a pull request, our system automatically looks through it. We check for security issues, mistakes in how you're setting up your infrastructure, and common code problems. We do this to make sure your changes are solid and won't cause any trouble later.

Talking to CodeAnt AI

Got a question or need a hand with something in your pull request? You can easily get in touch with CodeAnt AI right here. Just type the following in a comment on your pull request, and replace "Your question here" with whatever you want to ask:

@codeant-ai ask: Your question here

This lets you have a chat with CodeAnt AI about your pull request, making it easier to understand and improve your code.

Example

@codeant-ai ask: Can you suggest a safer alternative to storing this secret?

Preserve Org Learnings with CodeAnt

You can record team preferences so CodeAnt AI applies them in future reviews. Reply directly to the specific CodeAnt AI suggestion (in the same thread) and replace "Your feedback here" with your input:

@codeant-ai: Your feedback here

This helps CodeAnt AI learn and adapt to your team's coding style and standards.

Example

@codeant-ai: Do not flag unused imports.

Retrigger review

Ask CodeAnt AI to review the PR again, by typing:

@codeant-ai: review

Check Your Repository Health

To analyze the health of your code repository, visit our dashboard at https://app.codeant.ai. This tool helps you identify potential issues and areas for improvement in your codebase, ensuring your repository maintains high standards of code health.

…e misstates the ~12% figure's basis — claims consistency where populations differ
@don-petry
don-petry requested a review from a team as a code owner August 2, 2026 08:45
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

@codeant-ai

codeant-ai Bot commented Aug 2, 2026 •

Copy link
Copy Markdown

🤖 CodeAnt AI — Review Status

Status Commit Started (UTC) Finished (UTC)
✅ Reviewed your PR bfa02f7 Aug 02, 2026 · 08:45 08:45

@coderabbitai

coderabbitai Bot commented Aug 2, 2026 •

Copy link
Copy Markdown

Warning

Review limit reached

@don-petry, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 18 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 3089e95f-cc17-43ed-92e7-df6cefd4ef6b

📥 Commits

Reviewing files that changed from the base of the PR and between 8c0aa7e and d74fc0e.

📒 Files selected for processing (1)
  • docs/metrics-baseline.md

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@codeant-ai codeant-ai Bot added the size:M This PR changes 30-99 lines, ignoring generated files label Aug 2, 2026
@qodo-code-review

Copy link
Copy Markdown

PR Summary by Qodo

Docs: correct metrics-baseline note on ~12% noise estimate denominator scope

📝 Documentation 🕐 Less than 10 minutes

Grey Divider

AI Description

• Mark prior denominator-consistency claim as incorrect per dated-correction convention.
• Add a correction section explaining the ~12% figure’s all-comments sampling basis.
• Clarify how to interpret first-party-only measurements vs all-comments estimates.
High-Level Assessment

The chosen approach (strike-through the incorrect sentence and add a dated correction section) matches the document’s stated dated-correction convention and preserves auditability. Alternatives like silently rewriting or moving the correction to an external issue would reduce clarity and historical traceability.

Files changed (1) +44 / -4

Documentation (1) +44 / -4
metrics-baseline.mdAdd dated correction clarifying ~12% noise estimate denominator and interpretation +44/-4

Add dated correction clarifying ~12% noise estimate denominator and interpretation

• Replaces the prior claim that the ~12% figure shared the same first-party-marker denominator by striking it in place and adding an explicit correction notice. Adds a new dated correction section detailing the original all-comments sampling basis (including third-party bots), explains the population mismatch vs the first-party-only classifier, and provides guidance for interpreting initial scheduled measurements.

docs/metrics-baseline.md

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request updates docs/metrics-baseline.md to add a dated correction clarifying the denominator scope of the ~12% noise estimate, explaining why the original claim of denominator consistency was incorrect. The review feedback suggests clarifying the distinction between different comment and review totals (180 vs. 254) to prevent confusion, which is a helpful improvement.

Comment thread docs/metrics-baseline.md Outdated
@don-petry

Copy link
Copy Markdown
Collaborator Author

Dev-Lead — review-changes (no-changes)

No changes were needed for this PR.

@don-petry
don-petry enabled auto-merge (squash) August 2, 2026 08:45
@donpetry-bot

Copy link
Copy Markdown
Contributor

Advisory bots were rate-limited; auto-approval is withheld until they recover. pr-review-sweep will re-review this PR after 2026-08-02T09:46:20Z.

@don-petry
don-petry disabled auto-merge August 2, 2026 08:46
coderabbitai[bot]
coderabbitai Bot previously approved these changes Aug 2, 2026
@don-petry

Copy link
Copy Markdown
Collaborator Author

Dev-Lead — fix-reviews (applied)

Changes committed and pushed.

@qodo-code-review

qodo-code-review Bot commented Aug 2, 2026 •

Copy link
Copy Markdown

Code Review by Qodo

🐞 Bugs (0) 📘 Rule violations (0) 📜 Skill insights (0)

Context used
✅ Compliance rules (platform): 50 rules

Grey Divider


Action required

1. Contradictory ~12% guidance ✓ Resolved 🐞 Bug ≡ Correctness
Description
The new correction states the ~12% estimate was computed over all PR comments and that the
first-party measurement is expected to differ, but earlier baseline text still labels ~12% as a
first-party share and says any divergence on the first run is a “measurement correction”. This makes
the baseline internally inconsistent and can lead readers to misinterpret the first scheduled run as
a regression/improvement rather than a population change.
Code

docs/metrics-baseline.md[R139-142]

+- **The two populations genuinely differ (AC #2).** `cn_render_noise_section` counts only first-party,
+  marker-bearing comments — a **narrower** population than the all-comments basis of the ~12% estimate.
+  The denominators are therefore **not** consistent. The first measured first-party value is
+  **expected to differ** from ~12%, and a difference of that kind is a **population change — not a
Relevance

●●● Strong

Similar “denominator/baseline mismatch” doc issue was explicitly raised and accepted in PR #1414.

PR-#1414

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The correction section explicitly states the all-comments estimate and the expectation of a
different first-party value, while earlier baseline text still treats ~12% as a first-party metric
and interprets divergence as correction; the underlying classifier code confirms the first-party
marker-only denominator.

docs/metrics-baseline.md[78-105]
docs/metrics-baseline.md[129-150]
scripts/lib/comment-noise.sh[13-18]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

### Issue description
The new correction section explains that the historical ~12% estimate used an **all-comments** denominator and is **not comparable** to the first-party, marker-only `cn_render_noise_section` series. However, earlier unstruck baseline text still presents ~12% as a first-party share and treats first-run divergence as a “measurement correction”, which conflicts with the correction and can mislead readers.

### Issue Context
- `cn_render_noise_section` is explicitly first-party marker–based (denominator = bodies carrying our automation markers).
- The correction section says the earlier ~12% was an all-comments phrase-scan estimate and that the first-party value is expected to differ.

### Fix Focus Areas
- docs/metrics-baseline.md[78-105]
- docs/metrics-baseline.md[129-150]
- scripts/lib/comment-noise.sh[13-18]

### What to change
- Update the baseline table row and the “Noise” narrative paragraphs so ~12% is described as an **all-comments estimate** (historical) and explicitly **not** “of first-party agent comments”.
- Rewrite the “Any divergence… is a measurement correction” guidance to match the correction section (i.e., first-party-vs-all-comments gap is expected; only within-denominator changes are comparable over time).
- Ensure the “confirmed on first run” language doesn’t imply validating the all-comments estimate with a first-party-only classifier.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools



Remediation recommended

2. Inconsistent sample totals ✓ Resolved 🐞 Bug ⚙ Maintainability
Description
The correction section states the 10-PR provenance sample has “total comments … = 180” but then says
“254/254 comments+reviews were machine-authored” for the same sample without defining a different
denominator. As written, the provenance note is internally inconsistent and not reproducible.
Code

docs/metrics-baseline.md[R132-135]

+  marker-bearing subset. Provenance: a 10-PR sample (#1355, #1366, #1372, #1359, #1351, #1357, #1356,
+  #1347, #1346, #1345) with **total comments (issue + review) = 180** and
+  **est. no-action/advisory = 21** → 21/180 ≈ 12%. In that same sample **254/254 comments+reviews were
+  machine-authored and 0 were human** — a figure only meaningful because third-party bots
Relevance

●●● Strong

Team frequently accepts doc numeric/consistency fixes (e.g., rounding/metric mismatches accepted in
PRs #891, #1218).

PR-#891
PR-#1218
PR-#1414

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The inconsistency is present in the newly added correction bullet itself, where both totals are
asserted for the same sample.

docs/metrics-baseline.md[132-137]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

### Issue description
The correction section includes two different totals for the same sample window (180 vs 254) without clarifying what each total counts. This creates a direct internal contradiction and undermines the provenance of the ~12% estimate.

### Issue Context
The text currently says:
- “total comments (issue + review) = 180”
- “In that same sample 254/254 comments+reviews were machine-authored …”

### Fix Focus Areas
- docs/metrics-baseline.md[129-137]

### What to change
- Either correct one of the totals, or explicitly define two denominators (e.g., “comment bodies only” vs “comment bodies + review events”) and explain how 180 relates to 254.
- Keep the 21/180 ≈ 12% derivation consistent with the defined denominator(s).

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


Grey Divider

To customize comments, go to the Qodo configuration screen, or learn more in the docs.

Qodo Logo

Comment thread docs/metrics-baseline.md
Comment thread docs/metrics-baseline.md
Comment thread docs/metrics-baseline.md Outdated
coderabbitai[bot]
coderabbitai Bot previously approved these changes Aug 2, 2026
@don-petry
don-petry disabled auto-merge August 2, 2026 08:48

@don-petry don-petry left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review — PR #1429 (#1419 baseline denominator correction)

No findings. This is exactly the correction the issue asked for, and handled better than I specified.

Every AC satisfied, verified against the branch:

  • AC #1 — the ~12% figure's real basis is now stated: all PR comments in a sampled window, third-party bots included, classified by phrase scan. Provenance recorded precisely (10-PR sample; total comments = 180; est. no-action = 21 → 21/180 ≈ 12%), and the 254/254 machine-authored, 0 human figure is carried with the explanation that it is only meaningful because third-party bots were in scope. I derived these numbers originally; they are reproduced correctly, including naming the six bots.
  • AC #2 — the false sentence is struck, and the doc now says plainly that the denominators are not consistent and that the first measured value is expected to differ.
  • AC #3 — the interpretation rule is there, and sharper than I asked for: a population difference is "not a regression and not an improvement", and it defines what would count as a genuine correction (the classifier disagreeing with a hand-count at a fixed first-party denominator), so a future reader cannot mistake the gap itself for a change in noise.
  • AC #4 — answered explicitly: no all-comments series is maintained, the first-party share is the only noise time-series here, and the reviewer scorecard is the only third-party view.
  • AC #5 — the dated-append convention is honoured, and beyond the letter of the ask: the wrong sentence is struck in place with a visible [Corrected …] marker rather than deleted, so the error and its correction are both legible. That is the right call for a document whose entire purpose is to be a trustworthy "before".

The shape here is now a model for the other corrections in this epic — including my own. Ready from my side.

@donpetry-bot

Copy link
Copy Markdown
Contributor

CI checks on this PR are still running. Once they complete, re-mention @donpetry-bot to trigger a fresh review.

Posted by the donpetry-bot PR-review cascade.

@don-petry

Copy link
Copy Markdown
Collaborator Author

Dev-Lead — fix-reviews (applied)

Changes committed and pushed.

@don-petry
don-petry disabled auto-merge August 2, 2026 08:56
@sonarqubecloud

sonarqubecloud Bot commented Aug 2, 2026

Copy link
Copy Markdown

@don-petry

Copy link
Copy Markdown
Collaborator Author

Dev-Lead — waiting on PR blockers (intent: review-changes)

PR: #1429
No changes were committed, but the PR still has blocking checks or reviews (failing or cancelled checks, or changes-requested reviews). The retry cron will re-attempt automatically. Next attempt after: 2026-08-02T09:27:09Z

@don-petry

Copy link
Copy Markdown
Collaborator Author

Note

@don-petry I reviewed this PR and no code changes were needed, but it still has blocking checks or reviews (failing or cancelled checks, or changes-requested reviews), so I cannot mark it done yet. I'll re-check automatically.
Next attempt after: 2026-08-02T09:27:09Z

@don-petry
don-petry enabled auto-merge (squash) August 2, 2026 08:57
@donpetry-bot

Copy link
Copy Markdown
Contributor

Advisory bots were rate-limited; auto-approval is withheld until they recover. pr-review-sweep will re-review this PR after 2026-08-02T09:57:27Z.

@donpetry-bot donpetry-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Automated review — APPROVED ✓

Risk: LOW
Reviewed commit: d74fc0e50bb1f26c12d46aec53f6c1a8c92cc8aa
Review mode: triage-approved (single reviewer)

Summary

Docs-only PR (single file: docs/metrics-baseline.md, +67/-19) correcting the documented basis of the ~12% reviewer-noise estimate. The triage tier's low-risk assessment is confirmed: no code, workflow, script, or security surface is touched. The correction accurately states the ~12% figure was an all-comments phrase-scan estimate (third-party bots included), not a first-party marker count, and follows the doc's dated-append convention.

Linked issue analysis

Closes #1419. All five acceptance criteria are substantively addressed in the diff: (1) the ~12% figure's actual basis is stated with full provenance (10-PR sample, 21/180 phrase-scan, 254 machine-authored review activities); (2) the incorrect 'denominators are consistent' sentence is struck in place and corrected — the doc now states the classifier measures a narrower population and a first-run divergence is a population change, not a regression/improvement; (3) a clear interpretation rule distinguishes genuine measurement corrections from the known population difference; (4) the doc explicitly states no all-comments series is maintained and names the reviewer scorecard as the only third-party view; (5) the correction is a dated append (Correction 2026-08-02) with the original text retained struck-through, per the doc's convention.

Findings

No blocking findings. Earlier bot-flagged issues (Qodo/Graphite: 180-vs-254 inconsistency and contradictory ~12% guidance) were fixed by dev-lead follow-up commits — the doc now distinguishes 180 comment bodies (phrase-scan denominator) from 254 distinct review activities — and all 4 review threads are resolved. Secret-scanning MCP tool unavailable in this environment; gitleaks CI check is green and the diff contains no secret-like content (markdown prose only).

CI status

All substantive checks green: Lint, ShellCheck, CodeQL, Agent Security Scan, agent-shield, Secret scan (gitleaks), SonarCloud (quality gate passed), unit-tests, holdout-guard, Compile agentic workflows, CodeRabbit, Graphite. The five CANCELLED entries are superseded agent-orchestration jobs (dev-lead dispatch/ci-relay, PR-review trigger runs), not failing CI. Dependency-audit jobs skipped (no matching ecosystems).


Reviewed automatically by the PR-review agent (single-reviewer mode: fable 5). Reply if you need a human review.

@don-petry
don-petry merged commit 01c6221 into main Aug 2, 2026
30 of 35 checks passed
@don-petry
don-petry deleted the dev-lead/issue-1419-20260802-0840 branch August 2, 2026 08:59
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:M This PR changes 30-99 lines, ignoring generated files

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[docs] metrics-baseline denominator note misstates the ~12% figure's basis — claims consistency where populations differ

2 participants