Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
94 commits
Select commit Hold shift + click to select a range
e61e16c
add initial ruler2
hsiehjackson Sep 17, 2025
e970ab2
fix syntax error
hsiehjackson Sep 17, 2025
f62aa6b
fix
hsiehjackson Sep 17, 2025
0d290fd
fix
hsiehjackson Sep 17, 2025
b9eb3d4
fix
hsiehjackson Sep 17, 2025
e7211ab
fix
hsiehjackson Sep 17, 2025
16e7e21
fix
hsiehjackson Sep 17, 2025
f567774
fix
hsiehjackson Sep 17, 2025
c30d3f3
fix
hsiehjackson Sep 17, 2025
810f5f6
add ruler2 metrics
hsiehjackson Sep 18, 2025
4877ea8
fix
hsiehjackson Sep 18, 2025
0de020f
fix
hsiehjackson Sep 18, 2025
aeb325e
add wer
hsiehjackson Sep 18, 2025
39e08c7
Evaluation on LiveBench-Coding (#821)
wasiahmad Sep 18, 2025
fe6d707
add initial ruler2
hsiehjackson Sep 17, 2025
7d2cc35
fix wer
hsiehjackson Oct 8, 2025
e63de3c
fix conflict
hsiehjackson Oct 8, 2025
d4f31ee
fix conflict
hsiehjackson Oct 8, 2025
5189cca
resolve conflict
hsiehjackson Dec 5, 2025
770d781
resolve conflict
hsiehjackson Dec 5, 2025
d73fcd0
resolve conflict
hsiehjackson Dec 5, 2025
f16b8fb
fix
hsiehjackson Dec 5, 2025
d05b538
fix
hsiehjackson Dec 5, 2025
44cf7fd
fix conflict
hsiehjackson Dec 5, 2025
24b6273
rulerv1 reasoning
hsiehjackson Dec 5, 2025
0d78226
save
hsiehjackson Dec 6, 2025
cac7836
remove prefix
hsiehjackson Dec 6, 2025
413ec3b
revert rulerv1
hsiehjackson Dec 12, 2025
2c6b9e3
remove pred postprocess
hsiehjackson Dec 12, 2025
6ae14d9
Add apex-shortlist dataset (#1080)
i-vainn Dec 8, 2025
8a103ab
Introduce regex for small differences of formatting from judge (#1082)
wprazuch Dec 9, 2025
25e53b4
Add LCB Prompts, fix regex bug in robust_eval, remove CR, make summar…
gnalbandyan Dec 9, 2025
47d9312
MAINT pin nemo-evaluator (#1095)
gwarmstrong Dec 10, 2025
8bd2125
Update issue templates
gwarmstrong Dec 11, 2025
3c03013
Delete .github/ISSUE_TEMPLATE directory
gwarmstrong Dec 11, 2025
d1449e0
enable blank issues (#1096)
gwarmstrong Dec 11, 2025
7f5753e
Fix input_file path handling when executor is "none" (#1089)
bzantium Dec 11, 2025
ab87b77
TST for #1089 (#1097)
gwarmstrong Dec 11, 2025
70cc6bf
Stepheng/prover cleanup (#1078)
stephencge Dec 11, 2025
8cfb8d7
add stem dependencies in main python sandbox (#1099)
jiacheng-xu Dec 11, 2025
843b8c6
Audiometrics unification (#1093)
Jorjeous Dec 11, 2025
5738fee
FEAT Add Tavily Search (#1085)
gwarmstrong Dec 11, 2025
a833c4b
updating code extraction logic (#1086)
wasiahmad Dec 11, 2025
3e201c8
Sandbox add stem (#1101)
jiacheng-xu Dec 12, 2025
07986ad
Handle none output in wmtp24++ (#1091)
Froxyy-dev Dec 12, 2025
8a3667a
ENH enable sandbox env overrides in generate (#1107)
gwarmstrong Dec 12, 2025
ed33775
Search Tool Parameter updates (#1112)
gwarmstrong Dec 15, 2025
ebcab2f
autoformalize cleanup (#1098)
stephencge Dec 15, 2025
a02b373
HF ASR Leaderboard Evaluation (#1104)
melllinia Dec 15, 2025
218a0ac
Stepheng/nemotron math proofs docs (#1111)
stephencge Dec 16, 2025
30457b8
Stepheng/prover gpt oss fix (#1114)
stephencge Dec 16, 2025
7c3957d
add Nemotron-Math-V2.pdf (#1113)
wedu-nvidia Dec 16, 2025
81198bc
SWE-bench: don't pass external environment variables into Apptainer c…
ludwig-n Dec 16, 2025
5992dc7
Adding clan PR with AudioBench and Librispeech PC. (#1103)
Jorjeous Dec 16, 2025
7965d51
Schema overrides for tool-calling (#1118)
gwarmstrong Dec 16, 2025
6d5db21
FIX tool call error handling and search tool errors (#1120)
gwarmstrong Dec 17, 2025
07c23ba
Use run.Script for generate pipeline (#1052)
gwarmstrong Dec 17, 2025
5d4cb8e
Port ICPC changes to IOI (#1046)
SeanNaren Dec 17, 2025
8325611
replace raise error with LOG.warning in AA LCR dataset prepare (#1119)
anowaczynski-nvidia Dec 17, 2025
83a0ab0
FIX tavily search results return type (#1123)
gwarmstrong Dec 17, 2025
f81e450
Revert "Use run.Script for generate pipeline (#1052)" (#1125)
gwarmstrong Dec 18, 2025
1c433a7
Fix: add serialized_output on bad request (#1127)
gwarmstrong Dec 18, 2025
0e1e790
update paper link (#1128)
wedu-nvidia Dec 18, 2025
c2c8a56
update paper link, references to dataset, self-correction differences…
stephencge Dec 18, 2025
cd62bf7
FIX ioi ignore (#1131)
gwarmstrong Dec 18, 2025
71d15b6
download AA-LCR_extracted-text.zip via hf_hub_download (#1126)
anowaczynski-nvidia Dec 18, 2025
8ddcdf4
Evaluation on Livecodebench-pro (#1115)
wasiahmad Dec 19, 2025
3a50f7f
Evaluation support for SWE-rebench (#1102)
wasiahmad Dec 24, 2025
26ab834
Trust remote code in tokenizer (#1146)
Kipok Dec 27, 2025
956e8e8
Resolve broken links in docs (#1150)
activatedgeek Jan 5, 2026
7343902
Introduced vLLM_multimodal model to save multimodal outputs (#1136)
vmendelev Jan 6, 2026
3220591
add swe-rebench to excluded datasets (#1154)
gwarmstrong Jan 6, 2026
ba27f63
Fix run.Script refactor (#1133)
gwarmstrong Jan 6, 2026
2dcfe41
BIRD Benchmark (Text-to-SQL) (#1132)
redoctopus Jan 7, 2026
86195df
o3-mini-20250131 -> o3-mini-2025-01-31 (#1149)
bzantium Jan 7, 2026
0c57d24
Unify local sandbox with slurm setup (#1153)
Kipok Jan 7, 2026
257480a
fix: robust judgement handling (#1134)
Froxyy-dev Jan 8, 2026
fcb671c
generation.py to respect separate server type for the client (#1135)
vmendelev Jan 9, 2026
e36619b
Add compute eval (#1158)
blahblahasdf Jan 10, 2026
e256fb6
add musan dataset (#1139)
Jorjeous Jan 12, 2026
5536f54
fix
hsiehjackson Jan 13, 2026
8779b75
fix
hsiehjackson Jan 13, 2026
9490505
fix
hsiehjackson Jan 13, 2026
c41fb80
Merge branch 'main' into chsieh/ruler2
hsiehjackson Jan 13, 2026
53a7d14
fix
hsiehjackson Jan 13, 2026
733844d
fix
hsiehjackson Jan 13, 2026
2581064
resolve comment
hsiehjackson Jan 14, 2026
adf0e08
Merge branch 'main' into chsieh/ruler2
hsiehjackson Jan 14, 2026
979d4f4
Merge branch 'main' into chsieh/ruler2
Kipok Jan 14, 2026
081ca9a
Merge branch 'main' into chsieh/ruler2
hsiehjackson Jan 16, 2026
36d309b
Merge branch 'chsieh/ruler2' of https://github.com/NVIDIA/NeMo-Skills…
Kipok Jan 16, 2026
91cdf7d
Update docs
Kipok Jan 16, 2026
d9eed2a
Add link
Kipok Jan 16, 2026
39c2eda
Merge branch 'main' into chsieh/ruler2
Kipok Jan 16, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
41 changes: 41 additions & 0 deletions docs/evaluation/long-context.md
Original file line number Diff line number Diff line change
Expand Up @@ -26,6 +26,47 @@ Other supported options
* **base**: evaluate base model with answer prefix.
* **chat**: evaluate chat model including non-reasoning and reasoning model without answer prefix.

### ruler2
- Benchmark is defined in [`nemo_skills/dataset/ruler2/__init__.py`](https://github.com/NVIDIA-NeMo/Skills/blob/main/nemo_skills/dataset/ruler2/__init__.py)

It's recommended to use [data_dir parameter](../evaluation/index.md#using-data-on-cluster) when running evaluation.
Ruler2 also requires `setup`, `tokenizer_path` and `max_seq_length` to be specified. Example command to prepare data

```bash
ns prepare_data ruler2 \
--cluster=<cluster config> \
--data_dir=<mounted location to store data into> \
--setup=<typically MODEL_NAME-LENGTH but can be any string> \
--tokenizer_path=<model name, e.g. Qwen/Qwen3-1.7B> \
--max_seq_length=<length you want to evaluate, e.g. 131072>
```

Example evaluation command

```bash
ns eval \
--cluster=<cluster config> \
--data_dir=<must match prepare_data parameter> \
--output_dir=<any mounted output location> \
--benchmarks=ruler2.<what you used for prepare_data setup argument> \
--model=<model name, e.g. Qwen/Qwen3-1.7B> \
--server_nodes=1 \
--server_gpus=8 \
--server_type=vllm
```

Example scores

| Model | Avg | 8192 | 16384 | 32768 | 65536 | 131072 | 262144 | 524288 | 1000000 |
|-----------------------------------------|------|------|-------|-------|-------|--------|--------|--------|---------|
| Gemini 2.5 Flash Think On | 91.4 | 94.3 | 93.7 | 91.4 | 88.4 | 89.0 | - | - | - |
| Gemini 2.5 Flash Think Off | 88.0 | 91.3 | 89.0 | 88.8 | 85.5 | 85.5 | 82.5 | 79.1 | 77.0 |
| GPT 4.1 | 89.2 | 91.2 | 90.8 | 89.8 | 87.7 | 86.5 | 80.6 | 74.5 | 75.2 |
| Qwen3-235B-A22B-Thinking-2507 | 85.2 | 92.9 | 91.3 | 85.3 | 80.6 | 75.7 | - | - | - |
| Qwen3-235B-A22B-Instruct-2507 | 83.7 | 87.3 | 85.8 | 84.5 | 82.5 | 78.2 | 65.3 | 53.0 | 36.1 |

For more details see [https://github.com/NVIDIA/RULER/blob/rulerv2-ns](https://github.com/NVIDIA/RULER/blob/rulerv2-ns/)

### mrcr

- Benchmark is defined in [`nemo_skills/dataset/mrcr/__init__.py`](https://github.com/NVIDIA-NeMo/Skills/blob/main/nemo_skills/dataset/mrcr/__init__.py)
Expand Down
15 changes: 15 additions & 0 deletions nemo_skills/dataset/ruler2/__init__.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,15 @@
# Copyright (c) 2025, NVIDIA CORPORATION. All rights reserved.
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.

DATASET_GROUP = "long-context"
Loading