Repository navigation
test(e2e): add Llama-4 and Llama-3.3 70b to nightly benchmarks #900
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Changes from all commits
e3b68ab
922aa74
812cf0d
8bdbbf2
a091cec
d721496
90543e6
a8ae4a6
43247f5
a4175e2
17da2dc
e73a6c5
91d1775
02a457e
4edab60
File filter
Filter by extension
Conversations
Jump to
Diff view
Diff view
There are no files selected for viewing
| Original file line number | Diff line number | Diff line change |
|---|---|---|
|
|
@@ -102,14 +102,6 @@ def _resolve_model_path(hf_path: str) -> str: | |
| "tp": 1, | ||
| "features": ["chat", "streaming", "multimodal"], | ||
| }, | ||
| # Llama-4-Scout (17B with 16 experts) - Multimodal tests | ||
| "meta-llama/Llama-4-Scout-17B-16E-Instruct": { | ||
| "model": _resolve_model_path("meta-llama/Llama-4-Scout-17B-16E-Instruct"), | ||
| "tp": 4, | ||
| "features": ["chat", "streaming", "multimodal", "moe"], | ||
| "vllm_args": ["--max-model-len", "196608"], | ||
| "startup_timeout": 1200, | ||
| }, | ||
| # Llama-4-Maverick (17B with 128 experts, FP8) - Nightly benchmarks | ||
| "meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8": { | ||
| "model": _resolve_model_path("meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8"), | ||
|
|
@@ -126,6 +118,52 @@ def _resolve_model_path(hf_path: str) -> str: | |
| "--max-model-len=163840", # 160K context length (vLLM) | ||
| "--attention-backend=FLASHINFER", # FLASHINFER attention backend | ||
| ], | ||
| "startup_timeout": 1200, # Large MoE model may need extra download/load time | ||
| }, | ||
| # Llama-4-Scout (17B with 16 experts) - Nightly benchmarks and Multimodal tests | ||
| "meta-llama/Llama-4-Scout-17B-16E-Instruct": { | ||
| "model": _resolve_model_path("meta-llama/Llama-4-Scout-17B-16E-Instruct"), | ||
| "tp": 4, | ||
| "features": ["chat", "streaming", "function_calling", "multimodal", "moe"], | ||
| "worker_args": [ | ||
|
Member
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. We distinguish this by
Contributor
Author
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Yes I totally agreed the naming is confusing — worker_args actually means SGLang-specific args. This naming convention is pre-existing across all models in model_specs.py. If you like, I can do the rename (worker_args → sglang_args) in a separate refactoring PR to keep this one scoped to model additions. What do you think? |
||
| "--context-length=196608", | ||
| "--attention-backend=fa3", | ||
| "--cuda-graph-max-bs=256", | ||
| "--max-running-requests=300", | ||
| "--mem-fraction-static=0.85", | ||
| ], | ||
| "vllm_args": [ | ||
| "--max-model-len=196608", | ||
| ], | ||
| "startup_timeout": 1200, # Large MoE model may need extra download/load time | ||
|
CatherineSue marked this conversation as resolved.
|
||
| }, | ||
| # Llama-3.3-70B - Nightly benchmarks | ||
| "meta-llama/Llama-3.3-70B-Instruct": { | ||
| "model": _resolve_model_path("meta-llama/Llama-3.3-70B-Instruct"), | ||
| "tp": 4, | ||
| "features": ["chat", "streaming", "function_calling"], | ||
| "worker_args": [ | ||
| "--mem-fraction-static=0.9", | ||
| ], | ||
|
coderabbitai[bot] marked this conversation as resolved.
|
||
| "vllm_args": [ | ||
| "--max-model-len=131072", | ||
| "--gpu-memory-utilization=0.9", | ||
| "--enable-chunked-prefill", | ||
| ], | ||
| }, | ||
|
Comment on lines
+140
to
+153
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. 🧹 Nitpick | 🔵 Trivial Extract shared 70B args to avoid config drift. The dense and FP8 70B entries repeat identical Also applies to: 146-161 🤖 Prompt for AI Agents |
||
| # Llama-3.3-70B FP8 - Nightly benchmarks | ||
| "RedHatAI/Llama-3.3-70B-Instruct-FP8-dynamic": { | ||
| "model": _resolve_model_path("RedHatAI/Llama-3.3-70B-Instruct-FP8-dynamic"), | ||
| "tp": 4, | ||
| "features": ["chat", "streaming", "function_calling"], | ||
| "worker_args": [ | ||
| "--mem-fraction-static=0.9", | ||
| ], | ||
| "vllm_args": [ | ||
| "--max-model-len=131072", | ||
| "--gpu-memory-utilization=0.9", | ||
| "--enable-chunked-prefill", | ||
| ], | ||
| }, | ||
| } | ||
|
|
||
|
|
||
Uh oh!
There was an error while loading. Please reload this page.