-
Notifications
You must be signed in to change notification settings - Fork 193
PyPy3 execution support for LiveCodeBench evaluation #614
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Merged
Merged
Changes from 1 commit
Commits
Show all changes
848 commits
Select commit
Hold shift + click to select a range
9dc70bd
Fix timeout bug for LEAN4 code execution (#647)
fchen97 27f5e44
rename folder, add datasets (#651)
ekmb 62d0845
Improve client (#652)
smahdavi4 3806646
Fixes for code execution (#656)
Kipok 446ed09
Add skip_special_tokens=False for completions (#657)
Kipok 33a8baa
Fix timeout raise in sandbox
Kipok 4cfd725
Fix timeout for code exec
Kipok 5cd44ab
Update annotation
Kipok 8755ab5
Update timeout to 4 hours
Kipok 4e931e6
Online GenSelect (#655)
shtoshni ed68025
Adding long context benchmark MRCR (#634)
fayejf 45ac117
Fix a small bug in generation with chunks (#661)
fchen97 1bd199d
Small fix for mrcr prepare.py (#662)
fayejf 58af382
Fix base checkpoints in the docs
Kipok 6a12754
Fix formatting
Kipok c9d361d
Fix type mismatch for max code executions (#665)
Kipok 642b92a
Allow generation type or custom module in `eval` pipeline (#666)
activatedgeek f200cac
update grpo with megatron backend (#653)
wedu-nvidia 34b90a5
bugfix: missing generation module arg in eval pipeline cmd script (#668)
activatedgeek eab2d58
add support for nsys profile (#667)
wedu-nvidia 8c27f28
Fixing BFCL (#669)
shtoshni eba855e
Minor fixes to dataset defaults (#672)
shtoshni be737a0
Enable system_message for openai prompt format (#670)
Kipok e88c084
Reproducing Llama Nemotron Results with NeMo-Skills (#676)
shtoshni e600c78
Fixes for docs
Kipok 3022a9b
Add SWE-bench inference & evaluation (#671)
ludwig-n 5b00e48
Remove prompt template (#673)
Kipok 86ee4e4
allow overlapping sandbox with run_cmd (#680)
SeanNaren ff19286
Majority + Pass@k fix (#679)
shtoshni f075214
Remove broken benchmark
Kipok f1fb1b3
Pin openai version to v1.99.9 (#684)
ludwig-n 9a47605
Allow customizing SWE-agent/OpenHands repo & commit for SWE-bench (#682)
ludwig-n 5df438b
Fix 'Argument list too long' error in get_remaining_jobs (#674)
tamohannes 7dace78
fixing conflicts
wasiahmad b178cc9
fixing conflicts
wasiahmad e020abd
fix data cache bug (#685)
wedu-nvidia 71e4fa3
Fixes for ruler, simplified recipe and some docs (#686)
Kipok 397943f
Update prepare.py (#687)
fayejf 4cd131d
lcb updates to work with pypy3
wasiahmad a78f0e7
lcb updates to work with pypy3
wasiahmad e2a687a
Patch for LCB score calculation fix (#688)
wasiahmad cc69fd4
Fix cache variable (#692)
smahdavi4 0dab4a8
Add docs for ruler repro in tutorial (#693)
Kipok 47f7fe7
New evaluation docs (#694)
Kipok c672325
Add support for api_key_env_var (#696)
Glorf 722422b
fix: scicode missing functions (#699)
jubick1337 2e21e9b
BFCL Multi turn fixes (#704)
shtoshni b5c289a
Scicode fixed numbers for Llama Nemotron (#705)
shtoshni 1d8b332
Fix links (#706)
Kipok 85a66ce
Update CONTRIBUTING.md (#697)
activatedgeek 096f52a
gpt-oss integration (#711)
Kipok db8f2e5
Error handing for long context stuff (#708)
shtoshni 2fa7da7
Support gemini api (#698)
hsiehjackson 7685b2f
add beyondaime (#707)
wedu-nvidia e6f6a89
Update lcb splits (#714)
shtoshni fcaf706
handle exp name too long (#716)
fayejf 6eecc91
Add IOI (#701)
SeanNaren fc1c13f
Nano v2 tutorial (#718)
shtoshni 4b9b2d0
Fix link in readme
Kipok b26df04
Removing conversion tests (#720)
shtoshni 4bcdcf9
Fix shell execution interface (#719)
Kipok 6c1166e
BFCL v3 data prep - Repo version pinning (#725)
shtoshni 34d961b
add hf_home check (#702)
wedu-nvidia 095f1aa
Ipython session affinity (#622)
gwarmstrong afd373d
Add SFT data translation module (#721)
shuoyangd 1d3fd35
Fix streaming in async code execution (#722)
Kipok 120af1e
Add SWE-bench docs (#726)
ludwig-n 4cf1f48
fix model path in docs/basics aime eval example (#728)
stephencge d921a1a
Adding MCP clients (#713)
gwarmstrong c0b5022
Cleaning up references to TRTLLM model (#731)
shtoshni 53f83c7
Small fix for web search (#732)
smahdavi4 07f1163
Fix judge for non-default benchmarks (#727)
Kipok 75e84d1
Fix SciCode evaluation (#730)
jubick1337 31bcbfd
Fix apptainer installation in docker (#734)
Kipok bd998ff
Defer server wait for generation (#735)
smahdavi4 a5f3bcc
HLE Fix (#733)
shtoshni d922da3
Update evaluation numbers and a default split for scicode (#739)
Kipok 173f811
Add quotes
Kipok 1899ad1
Tool support based on Chat Completion APIs (#717)
activatedgeek 07ad14c
Fix ruff config (#740)
activatedgeek 06cdacb
Remove pin from openai req (#742)
Kipok 0367809
Parallel evaluation on a single node (#743)
Kipok e256ffd
Small fix to parallel mode
Kipok 3ac390f
Scicode updates to plot (#744)
shtoshni 85baa02
Soft failure (#723)
shtoshni 2fecec3
feat: unit test on slurm (#675)
wedu-nvidia c8c6382
AIMO inference tutorial (#745)
darraghdog 02b6c78
Bug Fix w/ Translation Module (#746)
shuoyangd a740b54
Updates to file formatting (#747)
shtoshni 69ff32b
Precommit fixes more complex (#748)
shtoshni 6f1a687
Add a tutorial on running gpt-oss with python tool (#750)
Kipok 9e7774c
Fix missing import
Kipok f902d45
BFCL Docs (#753)
shtoshni d2ee015
Slurm tests enhancements (#754)
Kipok 391652d
Remove wandb_project
Kipok 38de272
Fixes for slurm tests (#755)
Kipok 2c5f31d
More fixes for slurm tests and nemo-rl sft (#756)
Kipok b092a9c
update nemo-rl to latest main (#752)
wedu-nvidia 984e00c
update nemo-rl to latest main (#752)
wedu-nvidia 36f9c91
Server container now can be passed with CLI (#758)
vmendelev 7eed4be
BFCL v3 Testing + Refactoring (#761)
shtoshni b80d807
Add MMLU-Pro-X (#751)
shuoyangd 6a3a0b0
Prompt examples added + Minor changes to prompt construction (#764)
shtoshni 383b428
Generate docs with context length error part added (#765)
shtoshni 61ac015
Natural Language math docs (#767)
avem-nv e2f23a1
Reduce default `max_concurrent_requests` back to 512 (#770)
shuoyangd 0fee3e4
disable validation for nemo-rl if no validation data is provided (#766)
wedu-nvidia 27e4800
Adjust slurm parameters (#771)
Kipok 367ae37
Update to slurm tests setup
Kipok c6ad546
Update instruction for cron
Kipok 51ee492
Update test constraints
Kipok 2c372b6
Small doc fix
Kipok d463ba3
MCP interface updates (#772)
gwarmstrong 5e94850
Fix to not use completions api when soft_fail=True (#774)
Kipok 3be5070
Fix MCP env propagation (#775)
gwarmstrong af217f6
Fixing docs (#776)
shtoshni 0c26b5a
Update slurm tests (#777)
Kipok aaffceb
Update constraints
Kipok 1a27645
Fix async tool registration (#781)
gwarmstrong 988130c
Fix OpenHands patch issue & add more info to SWE-bench docs (#778)
ludwig-n f59b6c0
Potential solution to logging duplication (#782)
shtoshni fccb153
TRTLLM + Soft Fail (#786)
shtoshni 60430e5
Add litellm cache (#789)
smahdavi4 75fd5b8
GenSelect Online (#783)
shtoshni 52d544a
Add standard deviation metrics for benchmark variance analysis (#757)
AdamRajfer 3413dc6
GenSelect -> GenEvolution (#791)
shtoshni a9f31e5
Slurm tests refactoring (#795)
Kipok 0364637
GenEvolution docs (#792)
shtoshni fda85a3
Fix vllm multi-node + add conversion to int for gpus (#796)
Kipok d9dc1b9
Gpt-oss-python slurm test + more small refactoring (#797)
Kipok 2729707
gen select/synth default_factory (#800)
stephencge 1aca76e
Update constraints
Kipok 26f591d
Fix SWE-bench parallel runs and other issues (#794)
ludwig-n db07c74
updated config for sft data prep (#785)
wasiahmad cb4c0be
lean eval last code block and has sorry logic fix (#801)
stephencge a02d527
update max_tokens to max_completion_tokens for openai api (#804)
jiacheng-xu e6c27aa
Add build stage for docker images (#805)
gwarmstrong 5730a9e
FIX Cleanup Sessions on timeout (#803)
gwarmstrong 47740e8
Add search to docs (#807)
darraghdog 6780b53
enable msg format data passing for sft (#806)
wasiahmad e03d046
fix hf_model as None bug (#808)
wedu-nvidia f5614ad
add lr scheduler for nemo-rl sft with fsdp as backend (#759)
wedu-nvidia a9b3ca5
GenSynthesis prompt updates (#809)
shtoshni 4d3f817
Make OpenHands use uploaded dataset instead of redownloading from HF …
ludwig-n ed25dda
Update constraints
Kipok c74bdeb
Unifying context length error handling (#812)
shtoshni 1f45d0c
add MathOlympiadBench (#814)
stephencge 174d271
update miniF2F dataset (#813)
stephencge 58f3d55
fix typo (#816)
wedu-nvidia 678cced
Fix tool calling (#815)
Kipok a2a68ed
Update filters.trim_solutions=false to be default (#817)
Kipok 6decc47
Fix typo in grpo and tests (#818)
Kipok e2bd5dd
Evaluation on BigCodeBench (#547)
wasiahmad f8897ac
A small bug fix (#819)
wasiahmad c6a656f
Disallow empty prepare_data (#820)
Kipok 75c1148
Move sequence parallel to top level policy (#823)
Kipok 00ebe52
Evaluation on LiveBench-Coding (#821)
wasiahmad e0a26f7
Expose timeout parameter for individual calls (#827)
Kipok 6492a70
Implement token std statistics (#826)
AdamRajfer 478765d
Add Long context benchmark AA-LCR (#798)
fayejf 8160c6f
fixing merge conflicts
wasiahmad 906246d
fixing merge conflicts
wasiahmad bb21b6b
code logic reorganized
wasiahmad b15b3a0
code logic reorganized
wasiahmad 3a8b9ad
lcb eval harness main branch need to be used with pypy3
wasiahmad 2d8a042
lcb eval harness main branch need to be used with pypy3
wasiahmad df93e20
Update default sandbox parameters (#830)
Kipok 3163ea8
update to latest commit (#831)
wedu-nvidia c2ba406
Bump nemo-rl version to 0.7.1 (#832)
Kipok 68d6144
Merge branch 'main' into feat/lcb_eval
wasiahmad ebd5bab
Merge branch 'main' into feat/lcb_eval
wasiahmad 2da40f1
Add aa lcr to aai (#836)
fayejf 69b501f
Add new benchmark SimpleQA to nemo_skills (#828)
jiacheng-xu 2afdbfe
hle with detail splits (#837)
jiacheng-xu 9b7ce0c
timeout should be int
wasiahmad 66dbaa6
timeout should be int
wasiahmad 6b7af16
separating lcb code eval into a different file
wasiahmad 93ec8b4
separating lcb code eval into a different file
wasiahmad e1842df
fixing minor issue
wasiahmad a680edc
fixing minor issue
wasiahmad b3d4612
fixing minor issue
wasiahmad 4153e04
fixing minor issue
wasiahmad d64fa43
Merge branch 'main' into feat/lcb_eval
wasiahmad abb4805
Merge branch 'main' into feat/lcb_eval
wasiahmad e6d6b02
minor updates
wasiahmad 6af340e
minor updates
wasiahmad 1abbb9c
Asynchronous eval in Generation Loop (#825)
gwarmstrong 2d294ff
further optimizations
wasiahmad d2388c4
further optimizations
wasiahmad 6e68efa
NeMo-RL SFT sample printing (to verify if template is applied) (#833)
wasiahmad 5196289
Update default MCQ prompts (GPQA, MMLU-Pro) to non-boxed format (#843)
ekmb b25de59
Only run env var check for identity file when key available (#850)
activatedgeek ba480a6
minor issue fix
wasiahmad 20888db
minor issue fix
wasiahmad a9c0bf2
minor issue fix
wasiahmad b887bc7
minor issue fix
wasiahmad 1f2985c
fixing indent issue
wasiahmad 976e6da
fixing indent issue
wasiahmad 0fb1b8e
fixing file issues
wasiahmad c10e26a
fixing file issues
wasiahmad bf6231a
Sandbox history restoration fix (#838)
i-vainn 1d64b6c
changing lcb eval harness url
wasiahmad 27f886b
changing lcb eval harness url
wasiahmad 268d971
Merge remote-tracking branch 'origin/main' into feat/lcb_eval
wasiahmad 31db2de
Merge remote-tracking branch 'origin/main' into feat/lcb_eval
wasiahmad 7798899
pypy3 testing with datasets
wasiahmad 0c73f0e
pypy3 testing with datasets
wasiahmad 694c107
keeping test cases
wasiahmad aa6c209
keeping test cases
wasiahmad 3c3a73d
keeping test cases for pypy3 use
wasiahmad 3df9308
keeping test cases for pypy3 use
wasiahmad c759a74
keeping test cases for pypy3 use
wasiahmad 758333b
keeping test cases for pypy3 use
wasiahmad 2829c20
debugging
wasiahmad 07a13cb
debugging
wasiahmad c3f6a23
fix data prep issues
wasiahmad a6eab74
fix data prep issues
wasiahmad 193c940
dataset preparation updated
wasiahmad 47798dc
dataset preparation updated
wasiahmad 8597857
changing lcb eval harness branch name
wasiahmad b096d9b
changing lcb eval harness branch name
wasiahmad 9e2e0bd
Slurm: fix time format, and allow default timeout (#853)
artbataev 7c61cf1
Prompt sensitivity (multiprompt eval) support (#847)
gnalbandyan 59430b6
Fix HF_TOKEN assignment. Fix env vars priority: config -> environment…
artbataev 8fa3fc2
Merge branch 'main' into feat/lcb_eval
wasiahmad 1ed6bb1
Merge branch 'main' into feat/lcb_eval
wasiahmad 5e9eaf7
Megatron backend changes: minor fix, add random ports (#862)
lizziew b849bfb
fix wandb (#859)
wedu-nvidia 9ee8ea4
Allow setting random seeds for benchmark groups (#860)
Kipok 1a59903
Generation time + Input Sequence Length (#865)
shtoshni d2c5863
Small for for isl calc (#868)
Kipok 680d00b
Proper fix for isl (#869)
Kipok 2650a75
Adding support for arm64 containers (#856)
Kipok a5bede1
Merge branch 'main' into feat/lcb_eval
wasiahmad 0751a89
Merge branch 'main' into feat/lcb_eval
wasiahmad dd27f01
adding LCB docs
wasiahmad 0ccde9f
adding LCB docs
wasiahmad cc875fd
revert nemo-rl patch (#871)
activatedgeek bdf6b37
Remove sharding docs (#872)
smahdavi4 65e99b2
Adding support for training with megatron-lm (#873)
Kipok 1050ecf
Merge branch 'main' into feat/lcb_eval
wasiahmad 60087ac
Merge branch 'main' into feat/lcb_eval
wasiahmad ca09081
Evaluation on OJBench (#848)
wasiahmad 9028635
fixing merge conflicts
wasiahmad 4970d6e
fixing merge conflicts
wasiahmad 4042286
resolving conflicts
wasiahmad 27fde4f
merging
wasiahmad 3729e9a
updating docs
wasiahmad 29dad0f
Merge branch 'main' into feat/lcb_eval
wasiahmad d050d9f
minor doc update
wasiahmad File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
Empty file.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Empty file.
129 changes: 129 additions & 0 deletions
129
nemo_skills/evaluation/evaluator/livecodebench/code_generation.py
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,129 @@ | ||
| import base64 | ||
| import json | ||
| import pickle | ||
| import zlib | ||
| from dataclasses import dataclass | ||
| from datetime import datetime | ||
| from enum import Enum | ||
|
|
||
|
|
||
| class Platform(Enum): | ||
| LEETCODE = "leetcode" | ||
| CODEFORCES = "codeforces" | ||
| ATCODER = "atcoder" | ||
|
|
||
|
|
||
| class Difficulty(Enum): | ||
| EASY = "easy" | ||
| MEDIUM = "medium" | ||
| HARD = "hard" | ||
|
|
||
|
|
||
| class TestType(Enum): | ||
| STDIN = "stdin" | ||
| FUNCTIONAL = "functional" | ||
|
|
||
|
|
||
| @dataclass | ||
| class Test: | ||
| input: str | ||
| output: str | ||
| testtype: TestType | ||
|
|
||
| def __post_init__(self): | ||
| self.testtype = TestType(self.testtype) | ||
| # if self.testtype == TestType.FUNCTIONAL: | ||
| # self.input = json.loads(self.input) | ||
| # self.output = json.loads(self.output) | ||
|
|
||
|
|
||
| @dataclass | ||
| class CodeGenerationProblem: | ||
| question_title: str | ||
| question_content: str | ||
| platform: Platform | ||
| question_id: str | ||
| contest_id: str | ||
| contest_date: datetime | ||
| starter_code: str | ||
| difficulty: Difficulty | ||
| public_test_cases: list[Test] | ||
| private_test_cases: list[Test] | ||
| metadata: dict | ||
| language: str | ||
|
|
||
| def __post_init__(self): | ||
| self.platform = Platform(self.platform) | ||
| self.difficulty = Difficulty(self.difficulty) | ||
| self.contest_date = datetime.fromisoformat(self.contest_date) | ||
|
|
||
| if self.public_test_cases: | ||
| self.public_test_cases = json.loads(self.public_test_cases) # type: ignore | ||
| self.public_test_cases = [Test(**t) for t in self.public_test_cases] | ||
| else: | ||
| self.public_test_cases = [] | ||
|
|
||
| try: | ||
| self.private_test_cases = json.loads(self.private_test_cases) # type: ignore | ||
| except: | ||
| self.private_test_cases = json.loads( | ||
| pickle.loads( | ||
| zlib.decompress(base64.b64decode(self.private_test_cases.encode("utf-8"))) # type: ignore | ||
| ) | ||
| ) # type: ignore | ||
| self.private_test_cases = [Test(**t) for t in self.private_test_cases] | ||
|
|
||
| self.metadata = json.loads(self.metadata) # type: ignore | ||
|
|
||
| def insert_output(self, output_list: list[str], code_list: list[str]) -> dict: | ||
| return { | ||
| "question_title": self.question_title, | ||
| "question_content": self.question_content, | ||
| "platform": self.platform.value, | ||
| "question_id": self.question_id, | ||
| "contest_id": self.contest_id, | ||
| "contest_date": self.contest_date.isoformat(), | ||
| "starter_code": self.starter_code, | ||
| "difficulty": self.difficulty.value, | ||
| "output_list": output_list, | ||
| "code_list": code_list, | ||
| "language": self.language, | ||
| } | ||
|
|
||
| def insert_output_evaluation( | ||
| self, | ||
| output_list: list[str], | ||
| code_list: list[str], | ||
| graded_list: list[bool], | ||
| **kwargs, | ||
| ) -> dict: | ||
| output = self.insert_output(output_list, code_list) | ||
| output["graded_list"] = graded_list | ||
| output["pass@1"] = graded_list.count(True) / len(graded_list) | ||
| for k, v in kwargs.items(): | ||
| output[k] = v | ||
| return output | ||
|
|
||
| def get_evaluation_sample(self): | ||
| return { | ||
| "input_output": json.dumps( | ||
| { | ||
| "inputs": [t.input for t in self.public_test_cases + self.private_test_cases], | ||
| "outputs": [t.output for t in self.public_test_cases + self.private_test_cases], | ||
| "fn_name": self.metadata.get("func_name", None), | ||
| } | ||
| ), | ||
| } | ||
|
|
||
|
|
||
| def load_code_generation_dataset_from_file(filepath, language) -> list[CodeGenerationProblem]: | ||
| dataset = [] | ||
| with open(filepath, "r") as f: | ||
| for line in f: | ||
| p = json.loads(line) | ||
| if "task_id" in p: | ||
| assert "question_id" not in p | ||
| p["question_id"] = p.pop("task_id") | ||
| dataset.append(CodeGenerationProblem(**p, language=language)) | ||
| print(f"Loaded {len(dataset)} problems") | ||
| return dataset |
97 changes: 97 additions & 0 deletions
97
nemo_skills/evaluation/evaluator/livecodebench/evaluate.py
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,97 @@ | ||
| import json | ||
| from datetime import datetime | ||
|
|
||
| from nemo_skills.evaluation.evaluator.livecodebench.code_generation import load_code_generation_dataset_from_file | ||
| from nemo_skills.evaluation.evaluator.livecodebench.metrics import codegen_metrics | ||
| from nemo_skills.evaluation.evaluator.livecodebench.pass_k_utils import extract_instance_results | ||
|
|
||
|
|
||
| def evaluate( | ||
| custom_output_file: str, | ||
| release_version: str = "release_latest", | ||
| k_list=[1], | ||
| language: str = "python", | ||
| test_file: str = None, | ||
| num_process_evaluate: int = 12, | ||
| debug: bool = False, | ||
| timeout: int = 6, | ||
| ): | ||
| assert test_file is not None | ||
| custom_outputs = dict() | ||
| with open(custom_output_file, "r") as f: | ||
| for line in f: | ||
| output = json.loads(line) | ||
| if "question_id" not in output: | ||
| assert "task_id" in output | ||
| output["question_id"] = output.pop("task_id") | ||
| custom_outputs[output["question_id"]] = output | ||
|
|
||
| benchmark = load_code_generation_dataset_from_file(test_file, language) | ||
| benchmark = [problem for problem in benchmark if problem.question_id in custom_outputs] | ||
| assert len(custom_outputs) == len(benchmark), f"{len(custom_outputs)} != {len(benchmark)}" | ||
| assert all(isinstance(custom_output, dict) for custom_output in custom_outputs.values()) | ||
|
|
||
| save_results, combined_results = [], [] | ||
| for instance in benchmark: | ||
| code_list = custom_outputs[instance.question_id]["code_list"] | ||
| output = instance.insert_output(code_list, code_list) | ||
| save_results.append(output) | ||
| combined_results.append((code_list, code_list)) | ||
|
|
||
| eval_samples = [instance.get_evaluation_sample() for instance in benchmark] | ||
| generations = [extracted for _, extracted in combined_results] | ||
|
|
||
| metrics = codegen_metrics( | ||
| eval_samples, | ||
| generations, | ||
| k_list=k_list, | ||
| num_process_evaluate=num_process_evaluate, | ||
| timeout=timeout, | ||
| debug=debug, | ||
| language=language, | ||
| ) | ||
|
|
||
| graded = extract_instance_results(metrics[1]) | ||
|
|
||
| metadatas = metrics[2] | ||
| save_eval_results = [ | ||
| instance.insert_output_evaluation(outputs_list, extracted_list, graded_list, metadata=meta) | ||
| for instance, (outputs_list, extracted_list), graded_list, meta in zip( | ||
| benchmark, combined_results, graded, metadatas | ||
| ) | ||
| ] | ||
|
|
||
| # save_eval_results | ||
| output_results = dict() | ||
| output_results["date"] = datetime.now().strftime("%Y-%m-%d %H:%M") | ||
| for k in metrics[0]: | ||
| if k.startswith("pass@"): | ||
| print(f"{k}: {metrics[0][k]}") | ||
| output_results[k] = metrics[0][k] | ||
|
|
||
| output_results["detail_pass@1"] = dict() | ||
| output_results["eval"] = dict() | ||
| difficulty_wise_pass_at_1 = dict() | ||
| for r in save_eval_results: | ||
| output_results["eval"][r["question_id"]] = r | ||
| if r["difficulty"] not in difficulty_wise_pass_at_1: | ||
| difficulty_wise_pass_at_1[r["difficulty"]] = [] | ||
| difficulty_wise_pass_at_1[r["difficulty"]].append(r["pass@1"]) | ||
|
|
||
| for tag, v in difficulty_wise_pass_at_1.items(): | ||
| pass_at_1 = sum(v) / len(v) | ||
| print(f"{tag} pass@1: {pass_at_1}") | ||
| output_results["detail_pass@1"][tag] = pass_at_1 | ||
|
|
||
| with open(custom_output_file[:-6] + "_eval_results.json", "w") as f: | ||
| json.dump(output_results, f, indent=4) | ||
|
|
||
|
|
||
| def main(): | ||
| from fire import Fire | ||
|
|
||
| Fire(evaluate) | ||
|
|
||
|
|
||
| if __name__ == "__main__": | ||
| main() |
Oops, something went wrong.
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
Uh oh!
There was an error while loading. Please reload this page.