Skip to content

Add GLM-4.6 nightly benchmark coverage - #954

Closed
nishanthp wants to merge 10 commits into
smg-project:mainfrom
nishanthp:add-glm46-nightly-benchmark-v2
Closed

nishanthp wants to merge 10 commits into
smg-project:mainfrom
nishanthp:add-glm46-nightly-benchmark-v2

Conversation

@nishanthp

@nishanthp nishanthp commented Mar 27, 2026 •

Copy link
Copy Markdown
Contributor

Description

Problem

The nightly benchmark workflow does not include coverage for zai-org/GLM-4.6, so regressions for that model are not captured in the nightly run.

Solution

Add zai-org/GLM-4.6 to the nightly benchmark configuration and model specs so it runs as part of the nightly benchmark suite.

Changes

  • Added zai-org/GLM-4.6 to the nightly benchmark workflow.
  • Added the GLM-4.6 nightly benchmark test entry.
  • Added the GLM-4.6 model spec with tp: 8 and the expected feature set.

Test Plan

  • Verify the nightly benchmark workflow includes the GLM-4.6 job.
  • Confirm the benchmark test selects TestNightlyGlm46Single.
  • Confirm the model spec resolves zai-org/GLM-4.6 with tp=8.

Summary by CodeRabbit

  • New Features

    • Added support for GLM-4.6 with chat, streaming, function-calling, and reasoning capabilities.
  • Tests

    • Nightly end-to-end performance tests now include GLM-4.6, generating single- and multi-worker test variants and covering HTTP and gRPC backends.
  • Chores

    • Nightly benchmark workflow updated to include GLM-4.6 in the H200 single-worker job matrix.

@github-actions github-actions Bot added ci CI/CD configuration changes tests Test changes labels Mar 27, 2026
@coderabbitai

coderabbitai Bot commented Mar 27, 2026 •

Copy link
Copy Markdown

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review
📝 Walkthrough

Walkthrough

Adds support for the zai-org/GLM-4.6 model to nightly benchmarks: updates the GitHub Actions H200 single-worker matrix, registers the model in test model specs (tp=4, trust-remote-code args, several features), and adds the model to nightly performance test generation (1- and 2-worker test variants, http+grpc).

Changes

Cohort / File(s) Summary
GitHub Actions workflow
\.github/workflows/nightly-benchmark.yml
Added zai-org/GLM-4.6 entry to the H200 single-worker-h200 job matrix ({ id: zai-org/GLM-4.6, slug: zai-org-GLM-4.6, test_class: TestNightlyGlm46Single }).
Nightly performance tests
e2e_test/benchmarks/test_nightly_perf.py
Appended ("zai-org/GLM-4.6", "Glm46", 2, ["http","grpc"], {}) to _NIGHTLY_MODELS, causing generation of nightly Single and Multi test classes for this model and scheduling http+grpc backends.
Model specifications
e2e_test/infra/model_specs.py
Registered MODEL_SPECS["zai-org/GLM-4.6"] with model=_resolve_model_path("zai-org/GLM-4.6"), tp=4, features=["chat","streaming","function_calling","reasoning"], and worker_args/vllm_args set to ["--trust-remote-code"].

Sequence Diagram(s)

sequenceDiagram
    participant GH as GitHub Actions
    participant Runner as CI Runner (H200 single)
    participant Py as pytest (e2e tests)
    participant Spec as MODEL_SPECS
    participant Worker as Model Worker

    GH->>Runner: start nightly-benchmark (matrix selects zai-org/GLM-4.6)
    Runner->>Py: run generated test class TestNightlyGlm46Single
    Py->>Spec: resolve spec for "zai-org/GLM-4.6"
    Spec-->>Py: return model path, tp=4, args, features
    Py->>Runner: request worker launch with args (--trust-remote-code, tp=4)
    Runner->>Worker: launch model worker
    Worker-->>Py: ready (http/grpc endpoints)
    Py->>Runner: execute benchmarks (http + grpc)
    Runner-->>GH: upload test results
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related PRs

Suggested reviewers

  • CatherineSue
  • key4ng
  • slin1237
  • XinyueZhang369

Poem

🐇 I hopped into CI lanes, so spry,
GLM-4.6 now joins the sky.
Trusting remote code, TP set to four,
Tests hop, HTTP and gRPC soar.
Nightly carrots, benchmarks galore.

🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The PR title clearly and specifically summarizes the main change: adding GLM-4.6 model to nightly benchmark coverage, which aligns with all file modifications shown in the changeset.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@mergify

mergify Bot commented Mar 27, 2026

Copy link
Copy Markdown
Contributor

Hi @nishanthp, the DCO sign-off check has failed. All commits must include a Signed-off-by line.

To fix existing commits:

# Sign off the last N commits (replace N with the number of unsigned commits)
git rebase HEAD~N --signoff
git push --force-with-lease

To sign off future commits automatically:

  • Use git commit -s every time, or
  • VSCode: enable Git: Always Sign Off in Settings
  • PyCharm: enable Sign-off commit in the Commit tool window

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request adds the GLM-4.6 model to the nightly performance benchmarks and defines its infrastructure specifications. Feedback was provided regarding the need for the --trust-remote-code argument to ensure the model loads correctly and an adjustment to the worker count in the benchmark configuration to avoid redundant test execution.

Comment thread e2e_test/infra/model_specs.py
Comment thread e2e_test/benchmarks/test_nightly_perf.py

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In @.github/workflows/nightly-benchmark.yml:
- Line 350: The inline matrix item using braces (the entry containing id:
zai-org/GLM-4.6, slug: zai-org-GLM-4.6, test_class: TestNightlyGlm46Single)
violates YAMLlint brace-spacing; replace the inline mapping with a standard
block mapping instead of braces — expand the list item so it uses separate keys
(id, slug, test_class) on their own indented lines under the dash rather than an
inline "{ ... }" form to satisfy YAMLlint.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: cd3b5ba2-dbba-4924-910b-0c4f8f0faae4

📥 Commits

Reviewing files that changed from the base of the PR and between 34edd2b and 6555d27.

📒 Files selected for processing (3)
  • .github/workflows/nightly-benchmark.yml
  • e2e_test/benchmarks/test_nightly_perf.py
  • e2e_test/infra/model_specs.py

Comment thread .github/workflows/nightly-benchmark.yml
@nishanthp
nishanthp force-pushed the add-glm46-nightly-benchmark-v2 branch from 6555d27 to e2561c4 Compare March 27, 2026 20:34

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: e2561c49ee

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment on lines +147 to +148
if _multi_workers > 1:
variants.append(("Multi", _multi_workers))

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Preserve Multi test aliases for 1-worker nightly models

This new guard stops generating TestNightly*Multi classes when _multi_workers == 1, which removes TestNightlyGptOss20bMulti from this module even though the nightly workflow still schedules that class in the H100 multi-worker matrix (.github/workflows/nightly-benchmark.yml model entry for gpt-oss and run step pytest ... -k "$K_FILTER"). Because -k only runs matching tests, that job now selects nothing and pytest exits non-zero (no tests collected), so the multi-worker gpt-oss nightly jobs fail.

Useful? React with 👍 / 👎.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: f93db73bb9

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread e2e_test/benchmarks/test_nightly_perf.py Outdated

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@e2e_test/benchmarks/test_nightly_perf.py`:
- Around line 146-150: The test variant generation conditionally omits the
"Multi" variant when _multi_workers == 1, which removes the generated
TestNightlyGptOss20bMulti class and breaks the workflow selector; update the
variants creation so "Multi" is always added (e.g., variants = [("Single", 1),
("Multi", max(1, _multi_workers))] or always append ("Multi", _multi_workers))
so the TestNightly*Multi class is emitted even if count is 1, leaving any
runtime skip/behavior decisions to the test logic that consumes variants.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 8e35e324-38a0-466d-9a88-97c827cc0305

📥 Commits

Reviewing files that changed from the base of the PR and between 6555d27 and f93db73.

📒 Files selected for processing (3)
  • .github/workflows/nightly-benchmark.yml
  • e2e_test/benchmarks/test_nightly_perf.py
  • e2e_test/infra/model_specs.py

@mergify

mergify Bot commented Mar 30, 2026

Copy link
Copy Markdown
Contributor

Hi @nishanthp, the DCO sign-off check has failed. All commits must include a Signed-off-by line.

To fix existing commits:

# Sign off the last N commits (replace N with the number of unsigned commits)
git rebase HEAD~N --signoff
git push --force-with-lease

To sign off future commits automatically:

  • Use git commit -s every time, or
  • VSCode: enable Git: Always Sign Off in Settings
  • PyCharm: enable Sign-off commit in the Commit tool window

Signed-off-by: Nishanth Prakash <nishanth.prakash@gmail.com>
(cherry picked from commit e2561c4)
Signed-off-by: Nishanth Prakash <nishanth.prakash@gmail.com>
(cherry picked from commit aa3adb1)
Signed-off-by: Nishanth Prakash <nishanth.prakash@gmail.com>
@nishanthp
nishanthp force-pushed the add-glm46-nightly-benchmark-v2 branch from 488b4d6 to d19bc04 Compare March 30, 2026 17:57

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@e2e_test/benchmarks/test_nightly_perf.py`:
- Around line 146-150: Update the module-level docstring to accurately describe
that the test generates "Single" classes for all models and only generates
"Multi" classes when the _multi_workers variable is greater than 1; mention the
conditional behavior tied to the variants list construction and the for loop
iterating over (_suffix, _count) so readers know Multi classes may be skipped
when _multi_workers == 1 (also update any similar wording near the later
reference to the variants/for loop around the second occurrence).
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 85346372-3e1e-4fd5-98e5-1dcf2f9e4d46

📥 Commits

Reviewing files that changed from the base of the PR and between f93db73 and c294bd5.

📒 Files selected for processing (3)
  • .github/workflows/nightly-benchmark.yml
  • e2e_test/benchmarks/test_nightly_perf.py
  • e2e_test/infra/model_specs.py

Comment thread e2e_test/benchmarks/test_nightly_perf.py
Signed-off-by: Nishanth Prakash <nishanth.prakash@gmail.com>
Comment thread .github/workflows/nightly-benchmark.yml Outdated
Comment thread e2e_test/benchmarks/test_nightly_perf.py
Comment thread e2e_test/infra/model_specs.py Outdated
Signed-off-by: Nishanth Prakash <nishanth.prakash@gmail.com>
Signed-off-by: Nishanth Prakash <nishanth.prakash@gmail.com>
@nishanthp

Copy link
Copy Markdown
Contributor Author

@CatherineSue @key4ng Please take a look.

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Repo admins can enable using credits for code reviews in their settings.

Comment thread e2e_test/benchmarks/test_nightly_perf.py Outdated

@CatherineSue CatherineSue left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the PR. It seems the PR title check is failing. And I don't think it is gonna fit in H100.

"tp": 4,
"features": ["chat", "streaming", "function_calling", "reasoning"],
"worker_args": ["--trust-remote-code"],
"vllm_args": ["--trust-remote-code"],

@CatherineSue CatherineSue Apr 1, 2026 •

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We don't need this parameter for this one. No .py in the model files.

# GLM-4.6 - nightly benchmarks
"zai-org/GLM-4.6": {
"model": _resolve_model_path("zai-org/GLM-4.6"),
"tp": 4,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This model is 665GB in BF16. It won't fit in 8xH100.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Confirmed with @slin1237 and it looks like it needs multi node support, which is not supported by the nightly benchmarking pipeline.

Closing this PR

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the review @CatherineSue

@nishanthp nishanthp closed this Apr 1, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci CI/CD configuration changes tests Test changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants