docs: cell-4 q34b rescore — both full-1540 cells dual-judge complete - #68
Conversation
…omplete Follow-up to PR #67. Cell 4 (llama3.1:8b + RRF + temp 0.2 full 1540) qwen3:4b rescore landed at 0.53 Overall / 0.27 SH, vs cell 3's 0.54 / 0.28. Both judges now agree the qwen+mem0 leader wins Overall and Single-hop at full scale — subset 200 was over-representing the slice where llama+RRF won. Removes the 'qwen3:4b rescore still running' note and adds a small both-judges-both-cells summary table so readers don't have to cross-reference two paragraphs.
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Pro Plus Run ID: 📒 Files selected for processing (1)
📝 WalkthroughWalkthroughDocumentation update to ChangesBenchmark Results and Methodology
Estimated code review effort🎯 1 (Trivial) | ⏱️ ~3 minutes Possibly related PRs
Poem
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Code Review SummaryStatus: No Issues Found | Recommendation: Merge Files Reviewed (1 files)
Reviewed by grok-code-fast-1:optimized:free · 113,607 tokens |
Summary
Follow-up to PR #67. Cell 4 (llama3.1:8b + RRF + temp 0.2, full 1540) qwen3:4b rescore finished — both judges now agree the qwen+mem0 leader wins Overall and Single-hop at full scale.
The qwen3:4b half confirms the gemma4:e2b finding: llama+RRF trails qwen+mem0 by 0.01 Overall and 0.01 SH under the strict judge too. Subset 200 was over-representing the slice where llama+RRF won; both judges agree this at full 1540.
Test plan
Summary by CodeRabbit