Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions .github/workflows/nightly-benchmark.yml
Original file line number Diff line number Diff line change
Expand Up @@ -121,6 +121,9 @@ jobs:
- { id: Qwen/Qwen2.5-7B-Instruct, slug: Qwen-Qwen2.5-7B-Instruct, test_class: TestNightlyQwen7bSingle }
- { id: Qwen/Qwen3-30B-A3B, slug: Qwen-Qwen3-30B-A3B, test_class: TestNightlyQwen30bSingle }
- { id: openai/gpt-oss-20b, slug: openai-gpt-oss-20b, test_class: TestNightlyGptOss20bSingle }
- { id: meta-llama/Llama-4-Scout-17B-16E-Instruct, slug: meta-llama-Llama-4-Scout-17B-16E-Instruct, test_class: TestNightlyLlama4ScoutSingle }
- { id: meta-llama/Llama-3.3-70B-Instruct, slug: meta-llama-Llama-3.3-70B-Instruct, test_class: TestNightlyLlama70bSingle }
- { id: RedHatAI/Llama-3.3-70B-Instruct-FP8-dynamic, slug: RedHatAI-Llama-3.3-70B-Instruct-FP8-dynamic, test_class: TestNightlyLlama70bFp8Single }
Comment thread
CatherineSue marked this conversation as resolved.
variant:
- { id: sglang, runtime: sglang, grpc_only: "false", setup_vllm: false, setup_trtllm: false }
- { id: vllm, runtime: vllm, grpc_only: "false", setup_vllm: true, setup_trtllm: false }
Expand Down
21 changes: 21 additions & 0 deletions e2e_test/benchmarks/test_nightly_perf.py
Original file line number Diff line number Diff line change
Expand Up @@ -109,6 +109,27 @@ def _run_nightly(setup_backend, genai_bench_runner, model_id, worker_count=1, **
["http", "grpc"],
{},
),
(
"meta-llama/Llama-4-Scout-17B-16E-Instruct",
"Llama4Scout",
1,
["http", "grpc"],
{},
),
(
"meta-llama/Llama-3.3-70B-Instruct",
"Llama70b",
1,
["http", "grpc"],
{},
),
(
"RedHatAI/Llama-3.3-70B-Instruct-FP8-dynamic",
"Llama70bFp8",
1,
["http", "grpc"],
{},
),
]


Expand Down
54 changes: 46 additions & 8 deletions e2e_test/infra/model_specs.py
Original file line number Diff line number Diff line change
Expand Up @@ -102,14 +102,6 @@ def _resolve_model_path(hf_path: str) -> str:
"tp": 1,
"features": ["chat", "streaming", "multimodal"],
},
# Llama-4-Scout (17B with 16 experts) - Multimodal tests
"meta-llama/Llama-4-Scout-17B-16E-Instruct": {
"model": _resolve_model_path("meta-llama/Llama-4-Scout-17B-16E-Instruct"),
"tp": 4,
"features": ["chat", "streaming", "multimodal", "moe"],
"vllm_args": ["--max-model-len", "196608"],
"startup_timeout": 1200,
},
# Llama-4-Maverick (17B with 128 experts, FP8) - Nightly benchmarks
"meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8": {
"model": _resolve_model_path("meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8"),
Expand All @@ -126,6 +118,52 @@ def _resolve_model_path(hf_path: str) -> str:
"--max-model-len=163840", # 160K context length (vLLM)
"--attention-backend=FLASHINFER", # FLASHINFER attention backend
],
"startup_timeout": 1200, # Large MoE model may need extra download/load time
},
# Llama-4-Scout (17B with 16 experts) - Nightly benchmarks and Multimodal tests
"meta-llama/Llama-4-Scout-17B-16E-Instruct": {
"model": _resolve_model_path("meta-llama/Llama-4-Scout-17B-16E-Instruct"),
"tp": 4,
"features": ["chat", "streaming", "function_calling", "multimodal", "moe"],
"worker_args": [

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We distinguish this by worker_args (sglang) and vllm_args (vllm)? It is quite misleading.
We should rename both better to reflect:

  1. they are for worker
  2. either for sglang or vllm or trtllm

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes I totally agreed the naming is confusing — worker_args actually means SGLang-specific args. This naming convention is pre-existing across all models in model_specs.py. If you like, I can do the rename (worker_args → sglang_args) in a separate refactoring PR to keep this one scoped to model additions. What do you think?

"--context-length=196608",
"--attention-backend=fa3",
"--cuda-graph-max-bs=256",
"--max-running-requests=300",
"--mem-fraction-static=0.85",
],
"vllm_args": [
"--max-model-len=196608",
],
"startup_timeout": 1200, # Large MoE model may need extra download/load time
Comment thread
CatherineSue marked this conversation as resolved.
},
# Llama-3.3-70B - Nightly benchmarks
"meta-llama/Llama-3.3-70B-Instruct": {
"model": _resolve_model_path("meta-llama/Llama-3.3-70B-Instruct"),
"tp": 4,
"features": ["chat", "streaming", "function_calling"],
"worker_args": [
"--mem-fraction-static=0.9",
],
Comment thread
coderabbitai[bot] marked this conversation as resolved.
"vllm_args": [
"--max-model-len=131072",
"--gpu-memory-utilization=0.9",
"--enable-chunked-prefill",
],
},
Comment on lines +140 to +153

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick | 🔵 Trivial

Extract shared 70B args to avoid config drift.

The dense and FP8 70B entries repeat identical worker_args and vllm_args. Pull these into shared constants to keep future tuning changes synchronized.

Also applies to: 146-161

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@e2e_test/infra/model_specs.py` around lines 130 - 145, The repeated 70B model
configuration arrays should be extracted into shared constants to avoid drift:
create two module-level constants (e.g., SHARED_70B_WORKER_ARGS and
SHARED_70B_VLLM_ARGS) containing the identical worker_args and vllm_args values
and replace the inline arrays in the "meta-llama/Llama-3.3-70B-Instruct" entry
and the matching 70B entries (the dense and FP8 entries referenced around lines
146-161) with references to those constants; keep the existing use of
_resolve_model_path and preserve the exact argument values when moving them to
the constants.

# Llama-3.3-70B FP8 - Nightly benchmarks
"RedHatAI/Llama-3.3-70B-Instruct-FP8-dynamic": {
"model": _resolve_model_path("RedHatAI/Llama-3.3-70B-Instruct-FP8-dynamic"),
"tp": 4,
"features": ["chat", "streaming", "function_calling"],
"worker_args": [
"--mem-fraction-static=0.9",
],
"vllm_args": [
"--max-model-len=131072",
"--gpu-memory-utilization=0.9",
"--enable-chunked-prefill",
],
},
}

Expand Down