Skip to content

fix(ci): add memory limits to 4-gpu-h100 runner to prevent benchmark OOM - #669

Closed
CatherineSue wants to merge 1 commit into
mainfrom
chang/fix-benchmark-oom
Closed

CatherineSue wants to merge 1 commit into
mainfrom
chang/fix-benchmark-oom

Conversation

@CatherineSue

@CatherineSue CatherineSue commented Mar 7, 2026 •

Copy link
Copy Markdown
Member

Description

Problem

PR Test and Nightly benchmark jobs on 4-gpu-h100 runners have been consistently failing with exit code -9 (SIGKILL/OOM) for the past 48+ hours. All 4 recent PR Test benchmark failures show the same pattern: genai-bench failed with exit code -9.

Root cause: The runner pod spec in runner-values-4-gpu-h100.yaml (introduced in #659) has no memory requests or limits — only nvidia.com/gpu: 4. Without memory reservations, k8s can over-schedule multiple runner pods onto the same BM.GPU.H100.8 node, causing memory contention and OOM kills when model servers + genai-bench run concurrently.

Solution

Add memory: "128Gi" as both requests and limits on the runner container. This ensures k8s reserves sufficient memory per pod and prevents over-packing nodes.

Changes

  • scripts/k8s-runner-resources/runner-values-4-gpu-h100.yaml: Add memory requests/limits to runner container resources

Test Plan

  • Verify the next PR Test benchmark run completes without OOM (exit code -9)
  • Monitor that pods are scheduled with proper memory reservations via kubectl describe pod
Checklist
  • cargo +nightly fmt passes
  • cargo clippy --all-targets --all-features -- -D warnings passes
  • (Optional) Documentation updated

Summary by CodeRabbit

  • Chores
    • Updated GPU H100 runner configurations with memory resource specifications: memory requests and memory limits both set to 128Gi per instance. This ensures optimal resource allocation within the clusters, improves scheduling efficiency, prevents memory-related performance issues, and enhances the overall stability and reliability for GPU-intensive computational workloads and operations.

PR Test and Nightly benchmarks on 4-gpu-h100 runners have been
consistently OOM killed (exit code -9) because the runner pod has no
memory requests/limits. Without memory reservations, k8s over-schedules
pods on the same BM.GPU.H100.8 node, leaving insufficient RAM for
model servers and genai-bench.

Add 128Gi memory requests and limits to ensure proper scheduling.

Signed-off-by: Chang Su <chang.s.su@oracle.com>
@CatherineSue
CatherineSue requested a review from slin1237 as a code owner March 7, 2026 08:52
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request resolves persistent Out Of Memory (OOM) errors occurring in CI benchmark jobs running on 4-gpu-h100 Kubernetes runners. These failures were due to a lack of explicit memory resource definitions in the runner pod specifications, leading to Kubernetes over-scheduling pods and causing memory contention. The solution introduces a fixed memory request and limit of 128Gi for the runner containers, ensuring proper resource allocation and preventing future OOM kills.

Highlights

  • CI Stability Improvement: Addressed consistent Out Of Memory (OOM) failures in PR Test and Nightly benchmark jobs on 4-gpu-h100 runners by implementing explicit memory limits.
  • Kubernetes Resource Management: Configured Kubernetes runner pods with explicit memory requests and limits of "128Gi" to prevent over-scheduling and memory contention on nodes.
Changelog
  • scripts/k8s-runner-resources/runner-values-4-gpu-h100.yaml
    • Added memory requests and limits to the runner container resources.
Activity
  • The author, CatherineSue, has created this pull request to address critical CI failures.
  • The author has marked cargo +nightly fmt and cargo clippy checks as passing in the checklist.
  • The author has provided a test plan to verify the fix by monitoring future benchmark runs and pod scheduling.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@coderabbitai

coderabbitai Bot commented Mar 7, 2026 •

Copy link
Copy Markdown

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 99c6430c-93a2-474a-9de4-6b610e727507

📥 Commits

Reviewing files that changed from the base of the PR and between 9e6045a and 113b871.

📒 Files selected for processing (1)
  • scripts/k8s-runner-resources/runner-values-4-gpu-h100.yaml

📝 Walkthrough

Walkthrough

Adds memory request and limit specifications (128Gi each) to the Runner container's Kubernetes resource configuration for the 4-GPU H100 node pool, alongside the existing GPU allocation.

Changes

Cohort / File(s) Summary
Kubernetes Runner Configuration
scripts/k8s-runner-resources/runner-values-4-gpu-h100.yaml
Introduces memory requests and limits (128Gi) for the Runner container resource allocation.

Estimated code review effort

🎯 1 (Trivial) | ⏱️ ~3 minutes

Possibly related PRs

Suggested labels

ci

Suggested reviewers

  • slin1237

Poem

🐰 A rabbit hops through k8s clouds,
Memory reserves, 128Gi loud!
GPU runners now fed and blessed,
H100 nodes pass the memory test. ✨

🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately describes the main change: adding memory limits to the 4-gpu-h100 runner configuration to prevent out-of-memory errors.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch chang/fix-benchmark-oom

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request correctly adds memory resource requests and limits to the 4-gpu-h100 runner configuration to address persistent OOM failures. The change is sound and should resolve the issue. The suggestion to explicitly declare the GPU request for improved clarity and consistency remains valid. It may be beneficial to audit and apply similar memory constraints to the other GPU runner configurations in a follow-up to proactively prevent similar issues.

value: '120'
resources:
requests:
memory: "128Gi"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

For improved clarity and consistency, it's good practice to explicitly define the GPU request alongside the memory request. While Kubernetes implicitly sets the request equal to the limit when it's not specified, making it explicit enhances the readability of the pod's resource requirements. This also aligns with the pattern used in runner-values-cpu.yaml where all resources are explicitly requested.

            nvidia.com/gpu: 4
            memory: "128Gi"

@lightseek-bot
lightseek-bot deleted the chang/fix-benchmark-oom branch March 8, 2026 04:07
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant