docs(bench): RFC 0031 §9.13 comparative entry draft (runs #8–#18) — maintainer-gated fold-in - #494
Conversation
§9.13 compiles the RFC 0031 comparative program's honest-metric era (runs #8–#17 on corpus/otel-demo-v8 vs digest-pinned Loki 3.5.3): L1 and L3 provisional must-win passes on both channels, L2 parity-plus storage-side with named levers, the time-window losses published, the Loki flag deviations and nondeterminism recorded, and the §7 freeze inputs listed as open maintainer decisions. Fold-in is maintainer-gated; this is the draft. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
📝 WalkthroughWalkthroughAdds Section 9.13 to document provisional RFC 0031 comparisons between Ourios and Grafana Loki on the otel-demo-v8 corpus, including methodology, run identifiers, L1–L3 and L6 outcomes, determinism observations, configuration deviations, and deferred calibration decisions. ChangesRFC 0031 benchmark documentation
Estimated code review effort: 1 (Trivial) | ~5 minutes Possibly related PRs
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Pull request overview
Adds a new draft benchmark results entry to docs/benchmarks.md documenting RFC 0031’s comparative program runs (#8–#17) of Ourios vs Grafana Loki (otel-demo-v8, Loki 3.5.3), including metric definition notes and per-class outcome tables intended as §7 calibration inputs (not final gate verdicts).
Changes:
- Add §9.13 “Results — 2026-07-12” narrative framing + run provenance table for comparative dispatch runs.
- Document the amended §3.6 “total bytes” metric and retire earlier biased (count-scan-only) run figures as non-citable.
- Record L1/L2/L3/L6 per-class result tables and Loki configuration deviations used for replay integrity.
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
There was a problem hiding this comment.
Actionable comments posted: 3
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@docs/benchmarks.md`:
- Around line 1370-1375: Update the “Determinism note” in docs/benchmarks.md to
scope Ourios’s byte-identical claim to repeated measurements of the same fixed
build and configuration, rather than every optimization run in the ledger.
Preserve the explanation that this enables reliable repetition comparisons,
while acknowledging that totals may differ between optimization runs such as
`#8`–#10 and the L6 `#8/`#10 entries.
- Around line 1220-1224: Update the “Reference system” section to replace the
truncated Loki image digest with the complete sha256 digest, or link directly to
the committed configuration containing the full digest. Preserve the existing
image tag, deployment mode, endpoint, and deviation details.
- Around line 1244-1251: Clarify the benchmark pass-streak accounting in the
surrounding RFC0031.1 results text: either add the omitted run `#16` to the table
with its relevant outcomes, or explicitly state that omitted no-delta runs count
toward consecutive-pass streaks. Ensure the “third consecutive” L1 and “three in
a row” L3 claims can be verified from the documented runs.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
…row, scoped determinism Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WQY9wfrfRggqSpMLH8Xj3Y
… bytes Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WQY9wfrfRggqSpMLH8Xj3Y
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WQY9wfrfRggqSpMLH8Xj3Y
…rom the entry alone Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WQY9wfrfRggqSpMLH8Xj3Y
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WQY9wfrfRggqSpMLH8Xj3Y
…ay so Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WQY9wfrfRggqSpMLH8Xj3Y
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WQY9wfrfRggqSpMLH8Xj3Y
…tency gate; .11 citation Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WQY9wfrfRggqSpMLH8Xj3Y
…; full run table Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WQY9wfrfRggqSpMLH8Xj3Y
…es as written Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WQY9wfrfRggqSpMLH8Xj3Y
… counted runs Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WQY9wfrfRggqSpMLH8Xj3Y
What
The
docs/benchmarks.md§9.13 draft: the RFC 0031 comparative program's honest record, compiling runs #8–#18 (incl. the run #18 latency channel) on otel-demo-v8 against digest-pinned grafana/loki 3.5.3.What the entry carries
Checks run
mdbook build(clean; §9.13 anchor renders), no code touched. Arithmetic cross-checks close exactly (component sums and quoted ratios reproduce from raw bytes).🤖 Generated with Claude Code
https://claude.ai/code/session_01WQY9wfrfRggqSpMLH8Xj3Y
Summary by CodeRabbit
otel-demo-v8capture.