fix(benchmarks): point SWE-bench metadata at README - #11620
Conversation
There was a problem hiding this comment.
Your trial has ended. Reactivate Greptile to resume code reviews.
|
Important Review skippedAuto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Pro Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Your trial has ended. Reactivate Greptile to resume code reviews.
|
Same-account review note, not an approval. I found one small packaging-doc follow-up and pushed it to this PR branch:
Local Windows validation after the follow-up:
No benchmark execution/scoring/model behavior changed, so real-model trajectory and screenshot/video evidence remain N/A for this PR. |
LifeOps Multi-Tier BenchmarkSuite:
|
|
Claude encountered an error —— View job I'll analyze this and get back to you. |
Summary
project.readmepoints at the existingREADME.mdinstead of deletedRESEARCH.mdValidation
python3 - <<'PY' ... PYmetadata check:README.mdexistspython3 -m pytest packages/benchmarks/swe_bench/tests -q(131 passed)python3 -m build packages/benchmarks/swe_bench --sdist --wheel --outdir /tmp/eliza-swebench-buildgit diff --checkEvidence
.github/issue-evidence/11346-swebench-readme-metadata/README.mdIssue
Progress toward #11346. This does not close the full benchmark-matrix issue; it removes a packaging/readiness break in the SWE-bench harness metadata.