Skip to content

test(e2e): add Llama-4 and Llama-3.3 70b to nightly benchmarks - #900

Merged
key4ng merged 15 commits into
smg-project:mainfrom
paxiaatucsdedu:pan/llama-nightly-benchmark
Apr 1, 2026
Merged

key4ng merged 15 commits into
smg-project:mainfrom
paxiaatucsdedu:pan/llama-nightly-benchmark

Conversation

@paxiaatucsdedu

@paxiaatucsdedu paxiaatucsdedu commented Mar 24, 2026 •

Copy link
Copy Markdown
Contributor

Description

Problem

The nightly benchmark suite currently lacks coverage for several important Llama model variants: the Llama-4-Scout MoE model, the Llama-3.3-70B dense model, and its FP8-quantized counterpart. Without benchmark data for these models, we cannot track their inference performance over time across SGLang and vLLM engines.

Solution

Add model specifications, workflow entries, and test class definitions for all three models. All models are configured to run on 4-gpu-h100 runners.

Changes

  • e2e_test/infra/model_specs.py:

    • Added meta-llama/Llama-4-Scout-17B-16E-Instruct
    • Added meta-llama/Llama-3.3-70B-Instruct
    • Added RedHatAI/Llama-3.3-70B-Instruct-FP8-dynamic
  • .github/workflows/nightly-benchmark.yml:

    • Added all three models to the single-worker (H100) job matrix.
  • e2e_test/benchmarks/test_nightly_perf.py:

    • Added Llama4Scout, Llama70b, and Llama70bFp8 entries to _NIGHTLY_MODELS list

Test Plan

  • ruff check e2e_test/ passes with no new lint errors.
Checklist
  • cargo +nightly fmt passes
  • cargo clippy --all-targets --all-features -- -D warnings passes
  • (Optional) Documentation updated
  • (Optional) Please join us on Slack #sig-smg to discuss, review, and merge PRs

Summary by CodeRabbit

  • Tests

    • Nightly benchmarks expanded with three additional large-model variants (including a 70B FP8 variant) and configured to run an expanded 128K text traffic scenario set.
    • Benchmark runner now supports passing per-model traffic scenarios so nightly runs cover richer scenario combinations.
  • Chores

    • Nightly CI workflow updated to schedule and run the new model configurations within existing benchmark jobs.

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Repo admins can enable using credits for code reviews in their settings.

@coderabbitai

coderabbitai Bot commented Mar 24, 2026 •

Copy link
Copy Markdown

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review
📝 Walkthrough

Walkthrough

Adds three Llama models to nightly benchmarks and model specs, enables per-model 128K traffic-scenario overrides for nightly tests, and changes the benchmark command builder to derive the --api-backend from the E2E_RUNTIME environment variable. The nightly CI matrix is updated to include the new models.

Changes

Cohort / File(s) Summary
Nightly Workflow
.github/workflows/nightly-benchmark.yml
Expanded single-worker job matrix with three new model entries and matching test_class values; matrix usage unchanged.
Nightly Tests
e2e_test/benchmarks/test_nightly_perf.py
Added _TEXT_SCENARIOS_128K; _run_nightly now accepts traffic_scenario via extra_kwargs; added three nightly model entries that set traffic_scenario to the 128K scenarios.
Model Specs
e2e_test/infra/model_specs.py
Added three MODEL_SPECS entries (meta-llama/Llama-4-Scout-17B-16E-Instruct, meta-llama/Llama-3.3-70B-Instruct, RedHatAI/Llama-3.3-70B-Instruct-FP8-dynamic) with resolved model paths, tp=4, feature tags, and worker/vLLM runtime flags (long context, memory/tuning, Scout-specific flags).
Benchmark CLI Builder
e2e_test/benchmarks/conftest.py
_build_command(...) now derives --api-backend from the E2E_RUNTIME env var (lowercased), mapping known runtimes (sglang→sglang, vllm→vllm) and defaulting to openai otherwise; command construction otherwise unchanged.

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Possibly related PRs

  • #795 — Modifies the same model-specs and nightly benchmark workflow entries; likely overlaps the added model entries.
  • #411 — Overlaps changes to e2e_test/benchmarks/conftest.py (command-building and runtime flags).
  • #381 — Related to nightly-run behavior in test_nightly_perf.py (per-model timeouts / run configuration adjustments).

Suggested labels

ci, tests

Suggested reviewers

  • CatherineSue
  • key4ng
  • slin1237

Poem

🐇 I hopped in code fields, added three models to the night,
Gave them longer context so their tokens stretch just right.
Scenarios extended, flags tucked into place,
I nibble tests and matrix rows with carrot-speedy grace,
Nightly blooms anew — a soft, benchmarked light.

🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The PR title clearly and specifically describes the main change: adding three Llama model variants (Llama-4 Scout and two Llama-3.3 70B variants) to nightly benchmarks, which aligns precisely with the changeset across all four modified files.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request significantly enhances the nightly benchmark coverage by incorporating several new Llama model variants: Llama-4-Scout, Llama-3.3-70B, and its FP8-quantized counterpart. The primary goal is to establish a robust system for continuously monitoring the inference performance of these critical models across both SGLang and vLLM engines, ensuring that any performance regressions or improvements are promptly identified. This expansion provides a more comprehensive performance overview for key large language models.

Highlights

  • Expanded Nightly Benchmarks: Added Llama-4-Scout-17B-16E-Instruct, Llama-3.3-70B-Instruct, and Llama-3.3-70B-Instruct-FP8-dynamic models to the nightly benchmark suite to track their inference performance.
  • Model Specification Definitions: Defined detailed model specifications for the new Llama models, including tensor parallelism, features (chat, streaming, function_calling, moe), and specific worker and vLLM arguments.
  • Benchmark Workflow Integration: Integrated the new models into the single-worker job matrix within the GitHub Actions nightly benchmark workflow and configured necessary Hugging Face environment variables for model downloads.
Ignored Files
  • Ignored by pattern: .github/workflows/** (1)
    • .github/workflows/nightly-benchmark.yml
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request integrates three new Llama models (Llama-4-Scout-17B-16E-Instruct, Llama-3.3-70B-Instruct, and Llama-3.3-70B-Instruct-FP8-dynamic) into the nightly performance benchmarks. This includes adding them to the benchmark test list and defining their comprehensive model specifications, including worker and vLLM arguments. A critical issue was identified in the model specifications for Llama-3.3-70B-Instruct and RedHatAI/Llama-3.3-70B-Instruct-FP8-dynamic, where the SGLang worker argument --mem-frac is invalid and should be corrected to --mem-fraction-static to ensure proper GPU memory management and prevent worker failures.

Comment thread e2e_test/infra/model_specs.py Outdated
Comment thread e2e_test/infra/model_specs.py Outdated

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In @.github/workflows/nightly-benchmark.yml:
- Around line 124-126: The inline mapping entries currently have extra spaces
inside the braces and fail YAMLlint; update each inline mapping for the entries
referencing meta-llama/Llama-4-Scout-17B-16E-Instruct,
meta-llama/Llama-3.3-70B-Instruct, and
RedHatAI/Llama-3.3-70B-Instruct-FP8-dynamic to remove spaces directly after '{'
and before '}' (e.g., change " - { id: ..., slug: ..., test_class: ... }" to " -
{id: ..., slug: ..., test_class: ...}") so the braces have no internal
leading/trailing spaces and the YAML linter passes.

In `@e2e_test/infra/model_specs.py`:
- Around line 130-145: The repeated 70B model configuration arrays should be
extracted into shared constants to avoid drift: create two module-level
constants (e.g., SHARED_70B_WORKER_ARGS and SHARED_70B_VLLM_ARGS) containing the
identical worker_args and vllm_args values and replace the inline arrays in the
"meta-llama/Llama-3.3-70B-Instruct" entry and the matching 70B entries (the
dense and FP8 entries referenced around lines 146-161) with references to those
constants; keep the existing use of _resolve_model_path and preserve the exact
argument values when moving them to the constants.
- Around line 135-138: In the worker_args arrays (the "worker_args" lists in
this diff), replace the unsupported SGLang flag "--mem-frac" with the correct
flag "--mem-fraction-static" so workers start correctly; update both occurrences
(the list containing "--trust-remote-code", "--mem-frac=0.9" and the second
similar list later in the file) to use "--mem-fraction-static=0.9" (preserving
the numeric value).

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: e1a23a6f-97e8-41fa-ab0e-1618131c9fbf

📥 Commits

Reviewing files that changed from the base of the PR and between 84b7593 and 9b32a03.

📒 Files selected for processing (3)
  • .github/workflows/nightly-benchmark.yml
  • e2e_test/benchmarks/test_nightly_perf.py
  • e2e_test/infra/model_specs.py

Comment thread .github/workflows/nightly-benchmark.yml
Comment on lines +130 to +145
# Llama-3.3-70B - Nightly benchmarks
"meta-llama/Llama-3.3-70B-Instruct": {
"model": _resolve_model_path("meta-llama/Llama-3.3-70B-Instruct"),
"tp": 4,
"features": ["chat", "streaming", "function_calling"],
"worker_args": [
"--trust-remote-code",
"--mem-frac=0.9",
],
"vllm_args": [
"--trust-remote-code",
"--max-model-len=131072",
"--gpu-memory-utilization=0.9",
"--enable-chunked-prefill",
],
},

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick | 🔵 Trivial

Extract shared 70B args to avoid config drift.

The dense and FP8 70B entries repeat identical worker_args and vllm_args. Pull these into shared constants to keep future tuning changes synchronized.

Also applies to: 146-161

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@e2e_test/infra/model_specs.py` around lines 130 - 145, The repeated 70B model
configuration arrays should be extracted into shared constants to avoid drift:
create two module-level constants (e.g., SHARED_70B_WORKER_ARGS and
SHARED_70B_VLLM_ARGS) containing the identical worker_args and vllm_args values
and replace the inline arrays in the "meta-llama/Llama-3.3-70B-Instruct" entry
and the matching 70B entries (the dense and FP8 entries referenced around lines
146-161) with references to those constants; keep the existing use of
_resolve_model_path and preserve the exact argument values when moving them to
the constants.

Comment thread e2e_test/infra/model_specs.py

@CatherineSue CatherineSue left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: PR title needs to be updated

Comment thread .github/workflows/nightly-benchmark.yml Outdated
"model": _resolve_model_path("meta-llama/Llama-4-Scout-17B-16E-Instruct"),
"tp": 4,
"features": ["chat", "streaming", "function_calling", "moe"],
"worker_args": [

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We distinguish this by worker_args (sglang) and vllm_args (vllm)? It is quite misleading.
We should rename both better to reflect:

  1. they are for worker
  2. either for sglang or vllm or trtllm

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes I totally agreed the naming is confusing — worker_args actually means SGLang-specific args. This naming convention is pre-existing across all models in model_specs.py. If you like, I can do the rename (worker_args → sglang_args) in a separate refactoring PR to keep this one scoped to model additions. What do you think?

@paxiaatucsdedu paxiaatucsdedu changed the title Add Llama-4 and Llama-3.3 70b to nightly benchmarks test(e2e): Add Llama-4 and Llama-3.3 70b to nightly benchmarks Mar 24, 2026
@paxiaatucsdedu paxiaatucsdedu changed the title test(e2e): Add Llama-4 and Llama-3.3 70b to nightly benchmarks test(e2e): add Llama-4 and Llama-3.3 70b to nightly benchmarks Mar 24, 2026
@paxiaatucsdedu

paxiaatucsdedu commented Mar 24, 2026 •

Copy link
Copy Markdown
Contributor Author

nit: PR title needs to be updated

@CatherineSue Sure I updated the PR title.

@paxiaatucsdedu

paxiaatucsdedu commented Mar 24, 2026 •

Copy link
Copy Markdown
Contributor Author

@CatherineSue
Do we have 8* H100 available? If yes, I would like to use 8* H100 to benchmark Llama 4 maverick model in the nightly benchmark. Currently we are using 8*H200.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
e2e_test/benchmarks/test_nightly_perf.py (1)

70-76: ⚠️ Potential issue | 🟠 Major

Keep BENCH_TEST_MODE pinned to the smoke scenario.

Line 76 now lets model-specific traffic_scenario values override _TEST_TRAFFIC_SCENARIO. All three new nightly entries pass _TEXT_SCENARIOS_128K, so a smoke run will still execute the full six-scenario set, including D(128000,200), which defeats the fast-path and makes BENCH_TEST_MODE much more timeout-prone.

💡 Proposed fix
     if _TEST_MODE:
+        kwargs.pop("traffic_scenario", None)
         genai_bench_runner(
             router_url=gateway.base_url,
             model_path=model_path,
             experiment_folder=experiment_folder,
             num_concurrency=_TEST_NUM_CONCURRENCY,
-            traffic_scenario=kwargs.pop("traffic_scenario", _TEST_TRAFFIC_SCENARIO),
+            traffic_scenario=_TEST_TRAFFIC_SCENARIO,
             max_requests_per_run=_TEST_MAX_REQUESTS,
             timeout_sec=600,
             server_engine=runtime_display,
             gpu_type=gpu_type,
             gpu_count=gpu_count,
             **kwargs,
         )
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@e2e_test/benchmarks/test_nightly_perf.py` around lines 70 - 76, The test-mode
branch is letting callers override the smoke traffic via
kwargs.pop("traffic_scenario", _TEST_TRAFFIC_SCENARIO); change the
genai_bench_runner invocation inside the _TEST_MODE block to explicitly pass
traffic_scenario=_TEST_TRAFFIC_SCENARIO (and stop popping it from kwargs) so
BENCH_TEST_MODE always runs the lightweight smoke scenario—update the call that
uses genai_bench_runner and remove or ignore kwargs.pop("traffic_scenario", ...)
to enforce the pinned smoke scenario.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Outside diff comments:
In `@e2e_test/benchmarks/test_nightly_perf.py`:
- Around line 70-76: The test-mode branch is letting callers override the smoke
traffic via kwargs.pop("traffic_scenario", _TEST_TRAFFIC_SCENARIO); change the
genai_bench_runner invocation inside the _TEST_MODE block to explicitly pass
traffic_scenario=_TEST_TRAFFIC_SCENARIO (and stop popping it from kwargs) so
BENCH_TEST_MODE always runs the lightweight smoke scenario—update the call that
uses genai_bench_runner and remove or ignore kwargs.pop("traffic_scenario", ...)
to enforce the pinned smoke scenario.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 67175f9d-b50e-4909-9e1b-e26f296d239b

📥 Commits

Reviewing files that changed from the base of the PR and between 3e6b563 and 0438feb.

📒 Files selected for processing (2)
  • e2e_test/benchmarks/conftest.py
  • e2e_test/benchmarks/test_nightly_perf.py

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
e2e_test/benchmarks/test_nightly_perf.py (1)

71-84: ⚠️ Potential issue | 🟠 Major

Keep BENCH_TEST_MODE pinned to the smoke scenario.

Line 77 now lets per-model traffic_scenario overrides win in test mode. For the new 128K entries, e2e_test/benchmarks/conftest.py:103-108 expands that list into six --traffic-scenario flags, so BENCH_TEST_MODE stops being the cheap D(100,100) sanity path and starts exercising the full 128K matrix.

🛠️ Proposed fix
     if _TEST_MODE:
+        kwargs.pop("traffic_scenario", None)
         genai_bench_runner(
             router_url=gateway.base_url,
             model_path=model_path,
             experiment_folder=experiment_folder,
             num_concurrency=_TEST_NUM_CONCURRENCY,
-            traffic_scenario=kwargs.pop("traffic_scenario", _TEST_TRAFFIC_SCENARIO),
+            traffic_scenario=_TEST_TRAFFIC_SCENARIO,
             max_requests_per_run=_TEST_MAX_REQUESTS,
             timeout_sec=600,
             server_engine=runtime_display,
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@e2e_test/benchmarks/test_nightly_perf.py` around lines 71 - 84, When
_TEST_MODE is true we must force the smoke scenario instead of allowing
per-model overrides: update the genai_bench_runner call inside the _TEST_MODE
block to pass traffic_scenario=_TEST_TRAFFIC_SCENARIO explicitly (do not use
kwargs.pop("traffic_scenario", ...)) and ensure any incoming traffic_scenario in
kwargs is removed or ignored so it cannot override the test-mode value; key
symbols to change are the _TEST_MODE conditional, the genai_bench_runner call,
the traffic_scenario argument, _TEST_TRAFFIC_SCENARIO, and kwargs to prevent
leakage.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@e2e_test/benchmarks/test_nightly_perf.py`:
- Around line 38-45: The _TEXT_SCENARIOS_128K list is a duplicated copy of the
baseline scenarios; instead, create a single baseline constant (e.g.,
TEXT_SCENARIOS_BASELINE) and derive _TEXT_SCENARIOS_128K from that constant
(e.g., _TEXT_SCENARIOS_128K = derive_128k(TEXT_SCENARIOS_BASELINE) or
slice/transform as needed), then replace all uses of the duplicated list in the
model table and elsewhere to reference the new baseline-derived variable (ensure
any other size-specific variants like 121-141 are similarly derived from the
same baseline constant).

---

Outside diff comments:
In `@e2e_test/benchmarks/test_nightly_perf.py`:
- Around line 71-84: When _TEST_MODE is true we must force the smoke scenario
instead of allowing per-model overrides: update the genai_bench_runner call
inside the _TEST_MODE block to pass traffic_scenario=_TEST_TRAFFIC_SCENARIO
explicitly (do not use kwargs.pop("traffic_scenario", ...)) and ensure any
incoming traffic_scenario in kwargs is removed or ignored so it cannot override
the test-mode value; key symbols to change are the _TEST_MODE conditional, the
genai_bench_runner call, the traffic_scenario argument, _TEST_TRAFFIC_SCENARIO,
and kwargs to prevent leakage.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 520c47e8-0940-41c0-a2fd-db4034d8aa9b

📥 Commits

Reviewing files that changed from the base of the PR and between 0438feb and 217cc38.

📒 Files selected for processing (1)
  • e2e_test/benchmarks/test_nightly_perf.py

Comment thread e2e_test/benchmarks/test_nightly_perf.py Outdated
@paxiaatucsdedu
paxiaatucsdedu force-pushed the pan/llama-nightly-benchmark branch from 3659f8b to 3e6b563 Compare March 25, 2026 01:11

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@e2e_test/benchmarks/conftest.py`:
- Around line 75-84: The code silently falls back to "openai" when an
unrecognized E2E_RUNTIME is provided; modify the block that computes runtime and
api_backend so that after setting runtime = os.environ.get("E2E_RUNTIME",
"").lower() and api_backend = {"sglang": "sglang", "vllm": "vllm"}.get(runtime,
"openai"), you emit a warning if runtime is non-empty and not a key in the
mapping (e.g., using Python's logging.warning or pytest logging) indicating the
unrecognized E2E_RUNTIME value and that you're falling back to "openai"; keep
the existing cmd.extend(...) behavior unchanged.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 5849f5d2-0ab0-42f1-8a36-a5bb328bb0c0

📥 Commits

Reviewing files that changed from the base of the PR and between 217cc38 and 551f09c.

📒 Files selected for processing (1)
  • e2e_test/benchmarks/conftest.py

Comment thread e2e_test/benchmarks/conftest.py Outdated
@github-actions github-actions Bot added ci CI/CD configuration changes tests Test changes labels Mar 25, 2026
@paxiaatucsdedu

Copy link
Copy Markdown
Contributor Author

@CatherineSue
The Nightly Benchmark workflow failed because of:

huggingface_hub.errors.GatedRepoError: 403 Client Error. (Request ID: Root=1-69c35374-0844edc11ab3def65618d379;2885b508-5452-4192-811e-677fee204d88)

Cannot access gated repo for url https://huggingface.co/meta-llama/Llama-3.3-70B-Instruct/resolve/main/config.json.
Access to model meta-llama/Llama-3.3-70B-Instruct is restricted and you are not in the authorized list. Visit https://huggingface.co/meta-llama/Llama-3.3-70B-Instruct to ask for access.

Is it possible to manually download the model on the H100 or make sure the HF_TOKEN have access to the meta models?

@key4ng
key4ng force-pushed the pan/llama-nightly-benchmark branch from c26bcd4 to f8f5227 Compare March 26, 2026 01:29
@key4ng

key4ng commented Mar 26, 2026

Copy link
Copy Markdown
Member

@paxiaatucsdedu

Copy link
Copy Markdown
Contributor Author

Comment thread e2e_test/infra/model_specs.py
Comment thread e2e_test/benchmarks/conftest.py Outdated
@paxiaatucsdedu
paxiaatucsdedu force-pushed the pan/llama-nightly-benchmark branch from 2ec83f8 to 8148c07 Compare March 26, 2026 16:25
@key4ng

key4ng commented Mar 26, 2026

Copy link
Copy Markdown
Member

Hi @paxiaatucsdedu,

  • could you make changes like this commit b0cc8fd did, so we can trigger the test and don't occupy many resources.
  • why added hicache related params, it should not be the scope of this pr?

@paxiaatucsdedu

paxiaatucsdedu commented Mar 26, 2026 •

Copy link
Copy Markdown
Contributor Author

Hi @paxiaatucsdedu,

  • could you make changes like this commit b0cc8fd did, so we can trigger the test and don't occupy many resources.
  • why added hicache related params, it should not be the scope of this pr?

Hi @key4ng

  1. Sure I made changes same as your commit b0cc8fd did.
  2. We use HiCache related params in OME runtime, so I want to include it in this nightly benchmark so that the benchmark results reflects our model performance in prod.

@key4ng

key4ng commented Mar 27, 2026

Copy link
Copy Markdown
Member

cc: @CatherineSue not sure if we should include hicache param for nightly test, could you take a look

Comment thread e2e_test/benchmarks/conftest.py Outdated
@paxiaatucsdedu
paxiaatucsdedu force-pushed the pan/llama-nightly-benchmark branch from 3450499 to e6459e7 Compare March 27, 2026 17:31
Comment thread e2e_test/infra/model_specs.py Outdated
@paxiaatucsdedu
paxiaatucsdedu force-pushed the pan/llama-nightly-benchmark branch from e6459e7 to 0cf03d9 Compare March 27, 2026 18:56
@paxiaatucsdedu

Copy link
Copy Markdown
Contributor Author

@CatherineSue Tests passed: https://github.com/lightseekorg/smg/actions/runs/23662600790
If the branch looks good to you, I can revert the changes for testing purpose so that it is ready to merge.

paxiaatucsdedu and others added 14 commits March 27, 2026 13:22
Include three Llama variants in nightly performance runs: meta-llama/Llama-4-Scout-17B-16E-Instruct, meta-llama/Llama-3.3-70B-Instruct, and RedHatAI/Llama-3.3-70B-Instruct-FP8-dynamic.

Signed-off-by: paxiaatucsdedu <paxia@ucsd.edu>
Signed-off-by: paxiaatucsdedu <paxia@ucsd.edu>
Signed-off-by: paxiaatucsdedu <paxia@ucsd.edu>
Signed-off-by: paxiaatucsdedu <paxia@ucsd.edu>
- Added pull request trigger for specific paths in the nightly benchmark workflow.
- Commented out previous model configurations to streamline the benchmark process.
- Set job conditions to false for multi-worker and single-worker jobs to prevent execution.

Signed-off-by: key4ng <rukeyang@gmail.com>
Signed-off-by: key4ng <rukeyang@gmail.com>
Signed-off-by: key4ng <rukeyang@gmail.com>
…re_eos for SGLang

Signed-off-by: paxiaatucsdedu <paxia@ucsd.edu>
- Add pull_request trigger for benchmark workflow (per b0cc8fd pattern)
- Comment out existing models to reduce resource usage during testing
- Set multi-worker and single-worker-h200 jobs to if:false

Signed-off-by: paxiaatucsdedu <paxia@ucsd.edu>
Signed-off-by: paxiaatucsdedu <paxia@ucsd.edu>
…cate Scout entry

Signed-off-by: paxiaatucsdedu <paxia@ucsd.edu>
SGLang v0.5.9 and vLLM v0.18.0 both have native support for Llama-4
and Llama-3.3 architectures in their model registries, so
--trust-remote-code is not needed for these models.

Removed from:
- meta-llama/Llama-4-Scout-17B-16E-Instruct (worker_args + vllm_args)
- meta-llama/Llama-3.3-70B-Instruct (worker_args + vllm_args)
- RedHatAI/Llama-3.3-70B-Instruct-FP8-dynamic (worker_args + vllm_args)

Signed-off-by: paxiaatucsdedu <paxia@ucsd.edu>
Signed-off-by: paxiaatucsdedu <paxia@ucsd.edu>
@paxiaatucsdedu
paxiaatucsdedu force-pushed the pan/llama-nightly-benchmark branch from 0cf03d9 to 02a457e Compare March 27, 2026 20:22
@paxiaatucsdedu

Copy link
Copy Markdown
Contributor Author

@CatherineSue
The changes for benchmark tests have been reverted. Would you like to merge this PR? Thanks

@paxiaatucsdedu

Copy link
Copy Markdown
Contributor Author

@slin1237 Addressed all your comments and requested changes. All benchmark tests passed: https://github.com/lightseekorg/smg/actions/runs/23662600790

@paxiaatucsdedu

Copy link
Copy Markdown
Contributor Author

@slin1237
The checks failed because of Qwen/Qwen3-VL-8B-Instruct model died during start up. Not related to this PR.

@key4ng
key4ng merged commit f397534 into smg-project:main Apr 1, 2026
66 of 67 checks passed
smfirmin pushed a commit to smfirmin/smg that referenced this pull request Apr 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci CI/CD configuration changes tests Test changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants