Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
20 commits
Select commit Hold shift + click to select a range
3fa1141
[TRTLLM-15040][test] Prune legacy model coverage
xinhe-nv Aug 27, 2026
ced7831
[TRTLLM-15040][test] Update retained coverage to supported models
xinhe-nv Aug 27, 2026
1830350
[TRTLLM-15040][test] Preserve AutoDeploy support matrix
xinhe-nv Aug 27, 2026
85fbf8f
[TRTLLM-15040][test] Limit legacy cleanup to tests
xinhe-nv Aug 27, 2026
00c1146
[TRTLLM-15040][test] Complete legacy model test cleanup
xinhe-nv Aug 27, 2026
eba2c02
[TRTLLM-15040][test] Remove deprecated DeciLM coverage
xinhe-nv Aug 27, 2026
6201558
[TRTLLM-15040][test] Fix supported-model test configurations
xinhe-nv Aug 27, 2026
4d949a6
[TRTLLM-15040][test] Restore actively maintained test coverage
xinhe-nv Aug 28, 2026
f9539bf
[TRTLLM-15040][test] Remove legacy Llama V2 scheduler coverage
xinhe-nv Aug 28, 2026
3f15e08
[TRTLLM-15040][test] Migrate LoRA coverage to Qwen3
xinhe-nv Aug 28, 2026
ff99b9b
[TRTLLM-15040][test] Finalize Qwen3 LoRA migration
xinhe-nv Aug 28, 2026
fc2698e
[TRTLLM-15040][test] Restore maintained scheduler coverage
xinhe-nv Aug 28, 2026
f5b3f80
[TRTLLM-15040][test] Restore Eagle3 LoRA coverage
xinhe-nv Aug 28, 2026
bc817c3
[TRTLLM-15040][test] Retire Nemotron NAS LoRA test
xinhe-nv Aug 28, 2026
857d51c
[None][test] Remove redundant Llama 3.3 LoRA test
xinhe-nv Aug 28, 2026
3ef7702
[TRTLLM-15040][test] Remove remaining Nemotron Super metadata
xinhe-nv Aug 28, 2026
446b213
[TRTLLM-15040][test] Strengthen DraftTarget greedy assertion
xinhe-nv Aug 31, 2026
b90448c
[None][test] Validate DraftTarget draft-length scheduling
xinhe-nv Aug 31, 2026
2404eff
[TRTLLM-15040][test] Accept valid draft schedule transitions
xinhe-nv Aug 31, 2026
925be22
[TRTLLM-15040][test] Stabilize DraftTarget parity coverage
xinhe-nv Aug 31, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 0 additions & 11 deletions tests/integration/defs/.test_durations
Original file line number Diff line number Diff line change
Expand Up @@ -74,7 +74,6 @@
"accuracy/test_llm_api_autodeploy.py::TestModelRegistryAccuracy::test_autodeploy_from_registry[deepseek-ai_DeepSeek-R1-0528-True]": 807.0042,
"accuracy/test_llm_api_autodeploy.py::TestModelRegistryAccuracy::test_autodeploy_from_registry[google_gemma-3-1b-it-False]": 54.751599999999996,
"accuracy/test_llm_api_autodeploy.py::TestModelRegistryAccuracy::test_autodeploy_from_registry[meta-llama_Llama-3.1-8B-Instruct-False]": 43.825,
"accuracy/test_llm_api_autodeploy.py::TestModelRegistryAccuracy::test_autodeploy_from_registry[meta-llama_Llama-3.3-70B-Instruct-False]": 141.4636,
"accuracy/test_llm_api_autodeploy.py::TestModelRegistryAccuracy::test_autodeploy_from_registry[mistralai_Codestral-22B-v0.1-False]": 85.2512,
"accuracy/test_llm_api_autodeploy.py::TestModelRegistryAccuracy::test_autodeploy_from_registry[mistralai_Ministral-8B-Instruct-2410-False]": 58.8104,
"accuracy/test_llm_api_autodeploy.py::TestModelRegistryAccuracy::test_autodeploy_from_registry[nvidia_DeepSeek-R1-0528-NVFP4-v2-True]": 1474.047,
Expand Down Expand Up @@ -461,11 +460,6 @@
"accuracy/test_llm_api_pytorch.py::TestLlama3_1_8BInstruct::test_pard[overlap_scheduler=False]": 774.0882071823204,
"accuracy/test_llm_api_pytorch.py::TestLlama3_1_8BInstruct::test_pard[overlap_scheduler=True]": 699.8562759562841,
"accuracy/test_llm_api_pytorch.py::TestLlama3_1_8B_Instruct_RocketKV::test_auto_dtype": 922.8317,
"accuracy/test_llm_api_pytorch.py::TestLlama3_3_70BInstruct::test_fp4_tp2pp2[torch_compile=True-enable_gemm_allreduce_fusion=False]": 535.9970232558139,
"accuracy/test_llm_api_pytorch.py::TestLlama3_3_70BInstruct::test_fp8_tp4[torch_compile=False]": 607.9050751879699,
"accuracy/test_llm_api_pytorch.py::TestLlama3_3_70BInstruct::test_fp8_tp4[torch_compile=True]": 807.2856390977444,
"accuracy/test_llm_api_pytorch.py::TestLlama3_3_70BInstruct::test_nvfp4_tp4[torch_compile=False]": 633.9083157894737,
"accuracy/test_llm_api_pytorch.py::TestLlama3_3_70BInstruct::test_nvfp4_tp4[torch_compile=True]": 815.445175572519,
"accuracy/test_llm_api_pytorch.py::TestMiniMaxM2::test_4gpus[attention_dp=False-cuda_graph=True-overlap_scheduler=True-tp_size=4-ep_size=4]": 292.3591666666667,
"accuracy/test_llm_api_pytorch.py::TestMiniMaxM3::test_nvfp4[use_msa=True]": 863.916775510204,
"accuracy/test_llm_api_pytorch.py::TestMistralLarge3_675B::test_nvfp4_4gpus[latency_moe_trtllm]": 476.5565,
Expand Down Expand Up @@ -917,7 +911,6 @@
"perf/test_perf_sanity.py::test_e2e[aggr_upload-gpt_oss_120b_fp4_grace_blackwell-gpt_oss_fp4_tp1_mtp0_8k1k]": 509.8764444444444,
"perf/test_perf_sanity.py::test_e2e[aggr_upload-gpt_oss_120b_fp4_grace_blackwell-gpt_oss_fp4_tp2_1k8k]": 459.85290000000003,
"perf/test_perf_sanity.py::test_e2e[aggr_upload-host_perf_deepseek_v3_lite-v3lite_fp8_bs8_128_256]": 596.7645321782178,
"perf/test_perf_sanity.py::test_e2e[aggr_upload-host_perf_llama8b-llama8b_fp16_bs8_128_256]": 257.1823105134474,
"perf/test_perf_sanity.py::test_e2e[aggr_upload-qwen3_5_397b_fp4_blackwell-qwen3_5_397b_fp4_dep8_8k1k]": 698.3583333333333,
"perf/test_perf_sanity.py::test_e2e[aggr_upload-qwen3_5_397b_fp4_blackwell-qwen3_5_397b_fp4_dep8_mtp3_8k1k]": 531.4513333333334,
"perf/test_perf_sanity.py::test_e2e[aggr_upload-qwen3_5_397b_fp4_blackwell-qwen3_5_397b_fp4_tep4_mtp3_8k1k]": 404.2888333333333,
Expand Down Expand Up @@ -968,8 +961,6 @@
"perf/test_visual_gen_perf_sanity.py::test_visual_gen_e2e[vg_upload-wan22_i2v_a14b_blackwell-wan22_i2v_a14b_nvfp4_trtllm_cfg2_ulysses4]": 342.01559999999995,
"ray_orchestrator/RL/test_rl_perf_reproduce.py::test_rl_perf_reproduce[tp1_4instances]": 161.67143564356437,
"ray_orchestrator/RL/test_rl_perf_reproduce.py::test_rl_perf_reproduce[tp2_2instances]": 103.54573267326732,
"stress_test/stress_test.py::test_run_stress_test[llama-v3-8b-instruct-hf_tp1-stress_time_300s_timeout_450s-GUARANTEED_NO_EVICT-pytorch-stress-test]": 710.0681999999999,
"stress_test/stress_test.py::test_run_stress_test[llama-v3-8b-instruct-hf_tp1-stress_time_300s_timeout_450s-MAX_UTILIZATION-pytorch-stress-test]": 702.3580999999999,
"test_doc.py::test_url_validity": 49.48444444444444,
"test_e2e.py::test_draft_token_tree_quickstart_advanced_eagle3[Llama-3.1-8b-Instruct-llama-3.1-model/Llama-3.1-8B-Instruct-EAGLE3-LLaMA3.1-Instruct-8B]": 54.35407142857143,
"test_e2e.py::test_draft_token_tree_quickstart_advanced_eagle3_depth_1_tree[Llama-3.1-8b-Instruct-llama-3.1-model/Llama-3.1-8B-Instruct-EAGLE3-LLaMA3.1-Instruct-8B]": 55.181,
Expand Down Expand Up @@ -1080,7 +1071,6 @@
"unittest/_torch/modeling -k \"modeling_llama\"": 116.07080075662043,
"unittest/_torch/modeling -k \"modeling_mixtral\"": 72.60693384223919,
"unittest/_torch/modeling -k \"modeling_nemotron_nano_v2_vl\"": 421.9780921658986,
"unittest/_torch/modeling -k \"modeling_nemotron_nas\"": 42.0892577092511,
"unittest/_torch/modeling -k \"modeling_out_of_tree\"": 144.867929456112,
"unittest/_torch/modeling -k \"modeling_phi3\"": 35.700971046770604,
"unittest/_torch/modeling -k \"modeling_siglip\"": 204.05294545454547,
Expand Down Expand Up @@ -1494,7 +1484,6 @@
"unittest/llmapi/test_llm_pytorch.py -m \"part1\"": 249.2021185086551,
"unittest/llmapi/test_llm_pytorch.py -m \"part2\"": 349.5203164893617,
"unittest/llmapi/test_llm_pytorch.py -m \"part3\"": 279.06796533333335,
"unittest/llmapi/test_llm_pytorch.py::test_nemotron_nas_lora": 194.5492,
"unittest/llmapi/test_llm_quant.py": 21.754477644492912,
"unittest/llmapi/test_llm_telemetry.py": 56.14217090620032,
"unittest/llmapi/test_llm_telemetry.py::TestTelemetryArchitectureExtraction": 61.80126423690205,
Expand Down
3 changes: 0 additions & 3 deletions tests/integration/defs/.test_durations_aws_dfw
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,6 @@
"accuracy/test_llm_api_pytorch.py::TestDeepSeekV3Lite::test_bfloat16_4gpus_python_scheduler[tp4-mtp_nextn=0]": 293.95560641004704,
"accuracy/test_llm_api_pytorch.py::TestDeepSeekV3Lite::test_nvfp4_4gpus[moe_backend=CUTLASS-mtp_nextn=0-tp4-fp8kv=True-attention_dp=True-cuda_graph=True-overlap_scheduler=True-low_precision_combine=False-torch_compile=False]": 275.8909912491217,
"accuracy/test_llm_api_pytorch.py::TestDeepSeekV3Lite::test_nvfp4_4gpus[moe_backend=CUTLASS-mtp_nextn=0-tp4-fp8kv=True-attention_dp=True-cuda_graph=True-overlap_scheduler=True-low_precision_combine=True-torch_compile=False]": 266.709816042101,
"accuracy/test_llm_api_pytorch.py::TestLlama3_3_70BInstruct::test_fp4_tp2pp2[torch_compile=True-enable_gemm_allreduce_fusion=False]": 908.2052957660053,
"accuracy/test_llm_api_pytorch.py::TestNemotronV3Super::test_nvfp4_4gpus_online_eplb[moe_backend=CUTEDSL]": 641.9124943269417,
"accuracy/test_llm_api_pytorch.py::TestDeepSeekV3Lite::test_bfloat16_4gpus_kv_cache_aware_routing[mtp_nextn=0]": 287.03202630905434,
"accuracy/test_llm_api_pytorch.py::TestDeepSeekV3Lite::test_bfloat16_4gpus_kv_cache_aware_routing[mtp_nextn=2]": 350.3160223159939,
Expand All @@ -15,8 +14,6 @@
"accuracy/test_llm_api_pytorch.py::TestDeepSeekV3Lite::test_nvfp4_4gpus[moe_backend=CUTLASS-mtp_nextn=0-tp2pp2-fp8kv=True-attention_dp=True-cuda_graph=True-overlap_scheduler=True-low_precision_combine=False-torch_compile=True]": 430.816315329168,
"accuracy/test_llm_api_pytorch.py::TestGPTOSS::test_w4_4gpus[v1_kv_cache-dp4-cutlass-auto]": 1067.7517247761134,
"accuracy/test_llm_api_pytorch.py::TestGPTOSS::test_w4_4gpus_online_eplb[fp8]": 813.0690455089789,
"accuracy/test_llm_api_pytorch.py::TestLlama3_3_70BInstruct::test_fp8_tp4[torch_compile=False]": 702.5942377618048,
"accuracy/test_llm_api_pytorch.py::TestLlama3_3_70BInstruct::test_nvfp4_tp4[torch_compile=False]": 848.7312017090153,
"accuracy/test_llm_api_pytorch.py::TestDeepSeekV3Lite::test_bfloat16_4gpus_online_eplb[mtp_nextn=2-moe_backend=CUTLASS]": 413.68237058399245,
"accuracy/test_llm_api_pytorch.py::TestDeepSeekV3Lite::test_nvfp4_4gpus[moe_backend=CUTLASS-mtp_nextn=0-ep4-fp8kv=True-attention_dp=True-cuda_graph=True-overlap_scheduler=True-low_precision_combine=False-torch_compile=True]": 319.0723833630327,
"accuracy/test_llm_api_pytorch.py::TestDeepSeekV3Lite::test_nvfp4_4gpus[moe_backend=TRTLLM-mtp_nextn=0-ep4-fp8kv=True-attention_dp=True-cuda_graph=True-overlap_scheduler=True-low_precision_combine=False-torch_compile=False]": 332.56335388694424,
Expand Down

This file was deleted.

57 changes: 0 additions & 57 deletions tests/integration/defs/accuracy/references/cnn_dailymail.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -34,12 +34,6 @@ gpt2-medium:
accuracy: 22.249
gpt-next:
- accuracy: 25.516
microsoft/phi-2:
- accuracy: 31.255
bigcode/starcoder2-7b:
- accuracy: 26.611
- quant_algo: FP8
accuracy: 26.611
state-spaces/mamba-130m-hf:
- accuracy: 19.470
lmsys/vicuna-7b-v1.3:
Expand Down Expand Up @@ -71,13 +65,6 @@ TinyLlama/TinyLlama-1.1B-Chat-v1.0:
accuracy: 27.882
- extra_acc_spec: pp_size=4
accuracy: 15.123
meta-llama/Meta-Llama-3-8B-Instruct:
- accuracy: 34.957
- quant_algo: FP8
kv_cache_quant_algo: FP8
accuracy: 34.737
- quant_algo: W8A16_GPTQ
accuracy: 34.858
meta-llama/Llama-3.1-8B:
- accuracy: 24.360
- quant_algo: W8A8_SQ_PER_CHANNEL_PER_TOKEN_PLUGIN
Expand Down Expand Up @@ -122,48 +109,6 @@ meta-llama/Llama-3.1-8B-Instruct:
- quant_algo: FP8
extra_acc_spec: beam_width=2
accuracy: 31.201
meta-llama/Llama-3.2-1B:
- accuracy: 27.427
- quant_algo: W8A8_SQ_PER_CHANNEL_PER_TOKEN_PLUGIN
accuracy: 27.931
- quant_algo: W8A8_SQ_PER_CHANNEL
accuracy: 25.631
- quant_algo: W4A16_AWQ
accuracy: 25.028
- quant_algo: W4A16_AWQ
kv_cache_quant_algo: INT8
accuracy: 24.354
- quant_algo: FP8
kv_cache_quant_algo: FP8
accuracy: 27.029
- quant_algo: FP8
accuracy: 27.029
- quant_algo: FP8_PER_CHANNEL_PER_TOKEN
accuracy: 27.257
- quant_algo: FP8_PER_CHANNEL_PER_TOKEN
extra_acc_spec: meta_recipe
accuracy: 27.614
- extra_acc_spec: max_attention_window_size=960
accuracy: 27.259
- extra_acc_spec: max_attention_window_size=960;beam_width=4
accuracy: 0
meta-llama/Llama-3.2-3B:
- accuracy: 25.495
- quant_algo: FP8
kv_cache_quant_algo: FP8
accuracy: 33.629
meta-llama/Llama-3.3-70B-Instruct:
- quant_algo: FP8
spec_dec_algo: Eagle
accuracy: 33.244
- quant_algo: FP8
spec_dec_algo: Eagle3
accuracy: 33.244
- quant_algo: NVFP4
kv_cache_quant_algo: FP8
accuracy: 34.383
- quant_algo: FP8
accuracy: 34.927
mistralai/Mistral-Small-3.1-24B-Instruct-2503:
- accuracy: 29.20
- quant_algo: FP8
Expand Down Expand Up @@ -239,5 +184,3 @@ Qwen3/Qwen3-8B:
- quant_algo: FP8_BLOCK_SCALES
accuracy: 30
- accuracy: 30
nvidia/Llama-3_3-Nemotron-Super-49B-v1:
- accuracy: 34.003
37 changes: 0 additions & 37 deletions tests/integration/defs/accuracy/references/gpqa_diamond.yaml
Original file line number Diff line number Diff line change
@@ -1,16 +1,3 @@
meta-llama/Llama-3.3-70B-Instruct:
- accuracy: 45.96
- quant_algo: NVFP4
kv_cache_quant_algo: FP8
accuracy: 45.55
- quant_algo: FP8
kv_cache_quant_algo: FP8
accuracy: 48.03
- quant_algo: FP8
accuracy: 48.03
- quant_algo: FP8
kv_cache_quant_algo: FP8
accuracy: 48.03
deepseek-ai/DeepSeek-R1:
- quant_algo: NVFP4
accuracy: 70.45
Expand All @@ -35,30 +22,6 @@ deepseek-ai/DeepSeek-V3.2-Exp:
- quant_algo: NVFP4
spec_dec_algo: MTP
accuracy: 80.0
nvidia/Llama-3_3-Nemotron-Super-49B-v1:
- accuracy: 44.95
- quant_algo: FP8
kv_cache_quant_algo: FP8
accuracy: 42.42
# GPQA diamond only contains 198 samples, so the score tends to have large variance.
# We repeated evaluation 7 times to choose a lower bound score for FP8, 42.42.
# random_seed=0: 47.98
# random_seed=1: 42.42
# random_seed=2: 52.02
# random_seed=3: 51.52
# random_seed=4: 48.48
# random_seed=5: 47.47
# random_seed=6: 45.96
nvidia/Llama-3.1-Nemotron-Nano-8B-v1:
- accuracy: 40.40
- quant_algo: FP8
kv_cache_quant_algo: FP8
accuracy: 39.39
nvidia/Llama-3_1-Nemotron-Ultra-253B-v1:
- accuracy: 58.08
- quant_algo: FP8
kv_cache_quant_algo: FP8
accuracy: 57.07
GPT-OSS/120B-MXFP4:
- accuracy: 65.0
- spec_dec_algo: Eagle
Expand Down
35 changes: 0 additions & 35 deletions tests/integration/defs/accuracy/references/gsm8k.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -33,24 +33,6 @@ meta-llama/Llama-3.1-8B-Instruct:
- quant_algo: NVFP4
kv_cache_quant_algo: FP8
accuracy: 66.03
meta-llama/Llama-3.3-70B-Instruct:
- accuracy: 83.78
- quant_algo: NVFP4
kv_cache_quant_algo: FP8
accuracy: 87.33
- quant_algo: FP8
kv_cache_quant_algo: FP8
accuracy: 90.30
- quant_algo: FP8
accuracy: 90.30
meta-llama/Llama-4-Scout-17B-16E-Instruct:
- accuracy: 89.70
- quant_algo: NVFP4
kv_cache_quant_algo: FP8
accuracy: 88.61
- quant_algo: FP8
kv_cache_quant_algo: FP8
accuracy: 89.45
deepseek-ai/DeepSeek-V3-Lite:
- accuracy: 64.74
- quant_algo: NVFP4
Expand Down Expand Up @@ -318,11 +300,6 @@ moonshotai/Kimi-K3:
# SA spec dec is lossless; scores match the baseline within noise.
- spec_dec_algo: SA
accuracy: 96.5
nvidia/Llama-3_3-Nemotron-Super-49B-v1:
- accuracy: 92.57
- quant_algo: FP8
kv_cache_quant_algo: FP8
accuracy: 92.42
nvidia/Nemotron-MOE:
- accuracy: 88.249
- quant_algo: FP8
Expand All @@ -331,16 +308,6 @@ nvidia/Nemotron-MOE:
- quant_algo: NVFP4
kv_cache_quant_algo: FP8
accuracy: 63.268
nvidia/Llama-3.1-Nemotron-Nano-8B-v1:
- accuracy: 37.15
- quant_algo: FP8
kv_cache_quant_algo: FP8
accuracy: 28.39
nvidia/Llama-3_1-Nemotron-Ultra-253B-v1:
- accuracy: 94.43
- quant_algo: FP8
kv_cache_quant_algo: FP8
accuracy: 94.16
google/gemma-3-1b-it:
- accuracy: 25.52 # score getting from lm-eval with HF implementation
- quant_algo: FP8
Expand Down Expand Up @@ -433,8 +400,6 @@ GPT-OSS/20B-NVFP4:
accuracy: 85.0
ByteDance-Seed/Seed-OSS-36B-Instruct:
- accuracy: 90.8
bigcode/starcoder2-7b:
- accuracy: 26.5
bigcode/starcoder2-15b:
- accuracy: 54.5
poolside/laguna-XS.2:
Expand Down
77 changes: 0 additions & 77 deletions tests/integration/defs/accuracy/references/mmlu.yaml
Original file line number Diff line number Diff line change
@@ -1,8 +1,3 @@
meta-llama/Meta-Llama-3-8B-Instruct:
- accuracy: 67.74
- quant_algo: NVFP4
kv_cache_quant_algo: FP8
accuracy: 63.47
meta-llama/Llama-3.1-8B:
- accuracy: 66.06
- quant_algo: NVFP4
Expand Down Expand Up @@ -35,59 +30,6 @@ meta-llama/Llama-3.1-8B-Instruct:
- quant_algo: NVFP4
kv_cache_quant_algo: FP8
accuracy: 65.11
meta-llama/Llama-3.2-1B:
- quant_algo: W8A8_SQ_PER_CHANNEL_PER_TOKEN_PLUGIN
accuracy: 32.72
- quant_algo: W8A8_SQ_PER_CHANNEL
accuracy: 32.07
- quant_algo: W4A16_AWQ
accuracy: 30.56
- quant_algo: W4A16_AWQ
kv_cache_quant_algo: INT8
accuracy: 31.29
- quant_algo: FP8
kv_cache_quant_algo: FP8
accuracy: 31.02
- quant_algo: FP8_PER_CHANNEL_PER_TOKEN
accuracy: 33.97
- quant_algo: FP8_PER_CHANNEL_PER_TOKEN
extra_acc_spec: meta_recipe
accuracy: 33.87
- extra_acc_spec: max_attention_window_size=960
accuracy: 32.82
meta-llama/Llama-3.2-3B:
- accuracy: 57.92
- spec_dec_algo: Eagle3
accuracy: 57.92
- quant_algo: FP8
kv_cache_quant_algo: FP8
accuracy: 60.60
meta-llama/Llama-3.3-70B-Instruct:
- accuracy: 81.31
- spec_dec_algo: Eagle3
accuracy: 81.31
- quant_algo: FP8
spec_dec_algo: Eagle
accuracy: 81.31
- quant_algo: FP8
spec_dec_algo: Eagle3
accuracy: 81.31
- quant_algo: NVFP4
kv_cache_quant_algo: FP8
accuracy: 78.78
- quant_algo: FP8
kv_cache_quant_algo: FP8
accuracy: 80.40
- quant_algo: FP8
accuracy: 80.40
meta-llama/Llama-4-Scout-17B-16E-Instruct:
- accuracy: 80.00
- quant_algo: NVFP4
kv_cache_quant_algo: FP8
accuracy: 79.60
- quant_algo: FP8
kv_cache_quant_algo: FP8
accuracy: 78.58
mistralai/Mistral-Small-3.1-24B-Instruct-2503:
- accuracy: 81.7
- quant_algo: FP8
Expand Down Expand Up @@ -256,26 +198,7 @@ moonshotai/Kimi-K2.5:
- quant_algo: NVFP4
kv_cache_quant_algo: FP8
accuracy: 89.67
nvidia/Llama-3_3-Nemotron-Super-49B-v1:
- accuracy: 79.43
- quant_algo: FP8
kv_cache_quant_algo: FP8
accuracy: 79.26
nvidia/Llama-3.1-Nemotron-Nano-8B-v1:
- accuracy: 57.97
- quant_algo: FP8
kv_cache_quant_algo: FP8
accuracy: 57.12
bigcode/starcoder2-7b:
- accuracy: 41.35
- quant_algo: FP8
accuracy: 41.35
# TODO: update once https://nvbugs/5393849 is fixed.
nvidia/Llama-3_1-Nemotron-Ultra-253B-v1:
- accuracy: 83.70
- quant_algo: FP8
kv_cache_quant_algo: FP8
accuracy: 83.36
mistralai/Ministral-8B-Instruct-2410:
- accuracy: 66.35
- quant_algo: FP8
Expand Down

This file was deleted.

Loading
Loading