Skip to content

docs: cell-4 q34b rescore — both full-1540 cells dual-judge complete - #68

Merged
jaylfc merged 2 commits into
masterfrom
docs/cell4-q34b-may11
May 12, 2026
Merged

docs: cell-4 q34b rescore — both full-1540 cells dual-judge complete#68
jaylfc merged 2 commits into
masterfrom
docs/cell4-q34b-may11

Conversation

@jaylfc

@jaylfc jaylfc commented May 11, 2026

Copy link
Copy Markdown
Owner

Summary

Follow-up to PR #67. Cell 4 (llama3.1:8b + RRF + temp 0.2, full 1540) qwen3:4b rescore finished — both judges now agree the qwen+mem0 leader wins Overall and Single-hop at full scale.

Cell qwen3:4b Overall / SH gemma4:e2b Overall / SH
qwen3.5:9b + mem0_additive + temp 0.2 0.54 / 0.28 0.68 / 0.55
llama3.1:8b + RRF + temp 0.2 0.53 / 0.27 0.64 / 0.49

The qwen3:4b half confirms the gemma4:e2b finding: llama+RRF trails qwen+mem0 by 0.01 Overall and 0.01 SH under the strict judge too. Subset 200 was over-representing the slice where llama+RRF won; both judges agree this at full 1540.

Test plan

  • Numbers cross-checked against /tmp/llama_rrf_gapfill_summary.tsv
  • No conflict with master (rebased before push)

Summary by CodeRabbit

  • Documentation
    • Updated benchmark documentation with completed validation results across all evaluator configurations.
    • Refined methodology analysis with improved insights on dataset scaling behavior and performance variations.

Review Change Stack

…omplete

Follow-up to PR #67. Cell 4 (llama3.1:8b + RRF + temp 0.2 full 1540) qwen3:4b
rescore landed at 0.53 Overall / 0.27 SH, vs cell 3's 0.54 / 0.28. Both judges
now agree the qwen+mem0 leader wins Overall and Single-hop at full scale —
subset 200 was over-representing the slice where llama+RRF won.

Removes the 'qwen3:4b rescore still running' note and adds a small
both-judges-both-cells summary table so readers don't have to cross-reference
two paragraphs.
@coderabbitai

coderabbitai Bot commented May 11, 2026

Copy link
Copy Markdown

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 2a0a7ed5-6cd1-4d17-a680-7624d810a2c2

📥 Commits

Reviewing files that changed from the base of the PR and between c58a1d9 and 74bc35f.

📒 Files selected for processing (1)
  • docs/benchmarks.md

📝 Walkthrough

Walkthrough

Documentation update to docs/benchmarks.md completing dual-judge validation for the qwen+mem0 and llama+RRF headline recipes. Adds confirmed qwen3:4b results alongside existing gemma4:e2b values and updates methodology takeaways with final conclusions on drift behavior.

Changes

Benchmark Results and Methodology

Layer / File(s) Summary
Dual-Judge Validation Results and Takeaways
docs/benchmarks.md
Adds dual-judge results table with qwen3:4b and gemma4:e2b scores for both headline recipes, removes pending-rescore note, and updates methodology takeaways to finalize conclusions on subset-to-full drift (small overall) and single-hop drift (recipe-dependent, with llama+RRF showing larger regression under lenient judge).

Estimated code review effort

🎯 1 (Trivial) | ⏱️ ~3 minutes

Possibly related PRs

  • jaylfc/taosmd#67: Both PRs modify docs/benchmarks.md to report full-1540 revalidation results for the same temp-sweep headline recipes and update methodology text about subset-to-full and single-hop drift.
  • jaylfc/taosmd#65: This PR completes the dual-judge reporting and methodology framework introduced in PR #65, now with qwen3:4b results finalized.
  • jaylfc/taosmd#32: Both PRs make documentation-only updates to benchmarking results and judge-accuracy reporting in the same file.

Poem

🐰 The judges both agree at last,
Two recipes, two verdicts cast,
From subset small to full so vast,
The drift is mild, the die is passed,
Yet Single-hop tells us contrasts fast!

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately reflects the main change: completing the dual-judge validation by adding qwen3:4b rescore results for cell-4 (llama+RRF) at full scale (1540 examples), making both full-1540 cells complete under both judges.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch docs/cell4-q34b-may11

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@kilo-code-bot

kilo-code-bot Bot commented May 12, 2026

Copy link
Copy Markdown

Code Review Summary

Status: No Issues Found | Recommendation: Merge

Files Reviewed (1 files)
  • docs/benchmarks.md - 0 issues

Reviewed by grok-code-fast-1:optimized:free · 113,607 tokens

@jaylfc
jaylfc merged commit 95dedf4 into master May 12, 2026
2 checks passed
@jaylfc
jaylfc deleted the docs/cell4-q34b-may11 branch May 12, 2026 11:49
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant