Repository navigation
fix(ci): add memory limits to 4-gpu-h100 runner to prevent benchmark OOM - #669
CatherineSue wants to merge 1 commit into
Conversation
PR Test and Nightly benchmarks on 4-gpu-h100 runners have been consistently OOM killed (exit code -9) because the runner pod has no memory requests/limits. Without memory reservations, k8s over-schedules pods on the same BM.GPU.H100.8 node, leaving insufficient RAM for model servers and genai-bench. Add 128Gi memory requests and limits to ensure proper scheduling. Signed-off-by: Chang Su <chang.s.su@oracle.com>
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request resolves persistent Out Of Memory (OOM) errors occurring in CI benchmark jobs running on Highlights
Changelog
Activity
Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here. You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension. Footnotes
|
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Organization UI Review profile: ASSERTIVE Plan: Pro Run ID: 📒 Files selected for processing (1)
📝 WalkthroughWalkthroughAdds memory request and limit specifications (128Gi each) to the Runner container's Kubernetes resource configuration for the 4-GPU H100 node pool, alongside the existing GPU allocation. Changes
Estimated code review effort🎯 1 (Trivial) | ⏱️ ~3 minutes Possibly related PRs
Suggested labels
Suggested reviewers
Poem
🚥 Pre-merge checks | ✅ 3✅ Passed checks (3 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
There was a problem hiding this comment.
Code Review
This pull request correctly adds memory resource requests and limits to the 4-gpu-h100 runner configuration to address persistent OOM failures. The change is sound and should resolve the issue. The suggestion to explicitly declare the GPU request for improved clarity and consistency remains valid. It may be beneficial to audit and apply similar memory constraints to the other GPU runner configurations in a follow-up to proactively prevent similar issues.
| value: '120' | ||
| resources: | ||
| requests: | ||
| memory: "128Gi" |
There was a problem hiding this comment.
For improved clarity and consistency, it's good practice to explicitly define the GPU request alongside the memory request. While Kubernetes implicitly sets the request equal to the limit when it's not specified, making it explicit enhances the readability of the pod's resource requirements. This also aligns with the pattern used in runner-values-cpu.yaml where all resources are explicitly requested.
nvidia.com/gpu: 4
memory: "128Gi"
Description
Problem
PR Test and Nightly benchmark jobs on
4-gpu-h100runners have been consistently failing with exit code -9 (SIGKILL/OOM) for the past 48+ hours. All 4 recent PR Test benchmark failures show the same pattern:genai-bench failed with exit code -9.Root cause: The runner pod spec in
runner-values-4-gpu-h100.yaml(introduced in #659) has no memoryrequestsorlimits— onlynvidia.com/gpu: 4. Without memory reservations, k8s can over-schedule multiple runner pods onto the sameBM.GPU.H100.8node, causing memory contention and OOM kills when model servers + genai-bench run concurrently.Solution
Add
memory: "128Gi"as bothrequestsandlimitson the runner container. This ensures k8s reserves sufficient memory per pod and prevents over-packing nodes.Changes
scripts/k8s-runner-resources/runner-values-4-gpu-h100.yaml: Add memory requests/limits to runner container resourcesTest Plan
kubectl describe podChecklist
cargo +nightly fmtpassescargo clippy --all-targets --all-features -- -D warningspassesSummary by CodeRabbit