Skip to content

fix(ci): Use H100 runner for Specific Tests - #357

Merged
slin1237 merged 3 commits into
mainfrom
xz/split-runner
Feb 7, 2026
Merged

slin1237 merged 3 commits into
mainfrom
xz/split-runner

Conversation

@XinyueZhang369

@XinyueZhang369 XinyueZhang369 commented Feb 6, 2026 •

Copy link
Copy Markdown
Collaborator

Description

Problem

The benchmarks and chat-completions-trtllm E2E tests require more GPU resources than the standard k8s-runner-gpu runners provide. Additionally, the benchmark TTFT thresholds were calibrated for A10 GPUs and are far too lenient for H100s (see #351).

Solution

Move these two matrix entries into a new gateway-e2e-heavy job that runs on 4-gpu-h100 runners, and tighten benchmark thresholds to match H100 performance.

Changes

  • Extract benchmarks and chat-completions-trtllm from the gateway-e2e matrix into a new gateway-e2e-heavy job targeting 4-gpu-h100 runners
  • Remove now-unused TRT-LLM backend setup conditional from gateway-e2e steps
  • Update finish job to depend on and check gateway-e2e-heavy
  • Update summarize-benchmarks to depend on gateway-e2e-heavy (where benchmark artifacts are now produced)
  • Tighten benchmark TTFT thresholds for H100 GPUs:
    • test_regular_perf: ttft_mean_max 6s → 0.8s (H100 averages ~0.77s)
    • test_pd_perf: ttft_mean_max 13s → 5s (H100 PD averages ~4.2-4.6s)
  • Add k8s runner resource manifests (autoscaler, CPU/GPU runner configs, RBAC) for better tracking the runner configurations

Test Plan

The k8s runner resource manifests are all applied on cluster, then make sure this change pass the test workflow

Checklist
  • cargo +nightly fmt passes
  • cargo clippy --all-targets --all-features -- -D warnings passes
  • (Optional) Documentation updated

Summary by CodeRabbit

  • Chores
    • CI updated to select GPU-capable runners dynamically and to improve job failure checks for more accurate pipeline results
    • Added Kubernetes runner infrastructure with RBAC, CPU and GPU runner deployments, and autoscaling to better handle workload demand
  • Tests
    • Tightened performance benchmarks to enforce faster time-to-first-task targets

@github-actions github-actions Bot added the ci CI/CD configuration changes label Feb 6, 2026
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello @XinyueZhang369, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request addresses the need for more powerful computational resources for specific, demanding CI tests. It reconfigures the CI pipeline to route resource-intensive benchmarks and E2E tests to dedicated H100 GPU runners. This ensures that these critical tests can execute efficiently and reliably, preventing resource starvation on standard runners. The changes also include the necessary Kubernetes configurations to manage the new runner deployments and their automatic scaling.

Highlights

  • Resource Allocation for Tests: Benchmarks and chat-completions-trtllm E2E tests, which require significant GPU resources, have been moved from the standard gateway-e2e matrix to a new gateway-e2e-heavy job.
  • New Runner Type: The gateway-e2e-heavy job is configured to run on 4-gpu-h100 runners, providing the necessary computational power for these demanding tests.
  • CI Workflow Updates: The finish and summarize-benchmarks jobs have been updated to correctly depend on and check the results of the new gateway-e2e-heavy job, ensuring proper workflow execution and artifact collection.
  • Kubernetes Runner Configuration: New Kubernetes resource manifests have been added to define and track the configurations for ARC runners, including autoscalers for H100 GPU, A10 GPU, and CPU runners, as well as their respective deployments and RBAC settings.
Changelog
  • scripts/k8s-runner-resources/arc-runner-autoscaler.yaml
    • Added a new Kubernetes HorizontalRunnerAutoscaler manifest.
    • Configures autoscaling for H100 GPU, A10 GPU, and CPU runners based on queued and in-progress workflow runs.
    • Defines minimum and maximum replica counts for each runner type.
  • scripts/k8s-runner-resources/arc-runner-cpu.yaml
    • Added a new Kubernetes RunnerDeployment manifest for CPU runners.
    • Specifies CPU runner configuration, including image, resource requests/limits (8 CPU, 16Gi memory), and Docker-in-Docker setup.
  • scripts/k8s-runner-resources/arc-runner-gpu.yaml
    • Added a new Kubernetes RunnerDeployment manifest for GPU runners.
    • Defines configurations for both H100 GPU and A10 GPU runners, including GPU limits (4 GPUs), node selectors, tolerations, and pod affinity.
    • Includes volume mounts for model caching, Docker socket, and shared memory (/dev/shm).
  • scripts/k8s-runner-resources/arc-runner-rbac.yaml
    • Added a new Kubernetes RBAC manifest.
    • Defines a ServiceAccount (arc-runner-sa), Role (arc-runner), and RoleBinding (arc-runner-rb) for ARC runners.
    • Grants necessary permissions for accessing secrets and managing pods within the actions-runner-system namespace.
Ignored Files
  • Ignored by pattern: .github/workflows/** (1)
    • .github/workflows/pr-test-rust.yml
Activity
  • No activity (comments, reviews, or progress updates) has been recorded for this pull request yet.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@coderabbitai

coderabbitai Bot commented Feb 6, 2026 •

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

Adds Kubernetes RunnerDeployments (GPU/CPU), autoscalers and RBAC for self-hosted runners, updates the CI workflow to select per-job runners and reference python-unit-tests in failure checks, and tightens two e2e benchmark thresholds. (≤50 words)

Changes

Cohort / File(s) Summary
Workflow changes
.github/workflows/pr-test-rust.yml
Adds per-job matrix.runner and switches jobs to `runs-on: ${{ matrix.runner
GPU Runner manifests
scripts/k8s-runner-resources/arc-runner-gpu.yaml
Adds two RunnerDeployment resources: arc-runner-gpu-h100 (10 replicas, H100 selector, GPU limits, volumes/tolerations/affinity) and arc-runner-gpu-a10 (2 replicas, A10 selector); includes DinD and runner containers and shared volumes/env.
CPU Runner manifest
scripts/k8s-runner-resources/arc-runner-cpu.yaml
Adds arc-runner-cpu RunnerDeployment (4 replicas) with runner and DinD containers and fixed CPU/memory requests/limits.
Autoscaler config
scripts/k8s-runner-resources/arc-runner-autoscaler.yaml
Adds two HorizontalRunnerAutoscaler resources targeting GPU H100 and CPU RunnerDeployments with min/max replicas and PercentageRunnersBusy metrics and thresholds/factors.
RBAC for runners
scripts/k8s-runner-resources/arc-runner-rbac.yaml
Adds ServiceAccount arc-runner-sa, Role arc-runner (permissions for secrets, pods, pods/log, pods/exec) and RoleBinding arc-runner-rb in actions-runner-system.
E2E benchmark tests
e2e_test/benchmarks/test_pd_perf.py, e2e_test/benchmarks/test_regular_perf.py
Tightens performance thresholds: ttft_mean_max changed 13→5 (pd) and 6→0.8 (regular), altering pass/fail criteria.

Sequence Diagram

sequenceDiagram
    participant Dev as Developer (PR)
    participant GH as GitHub Actions
    participant WF as pr-test-rust.yml
    participant Autoscaler as HorizontalRunnerAutoscaler
    participant K8s as Kubernetes (RunnerDeployments)
    participant Runner as Self-hosted Runner (arc-runner-*)
    participant Bench as Benchmarks

    Dev->>GH: Push PR triggers workflow
    GH->>WF: Start pr-test-rust.yml
    WF->>WF: Evaluate job matrix.runner
    WF->>Runner: Schedule job on selected `${{ matrix.runner }}`
    Autoscaler->>K8s: Adjust replicas based on PercentageRunnersBusy
    K8s->>Runner: Provision runner pods (GPU / CPU)
    Runner->>WF: Execute CI jobs (unit, e2e, benchmarks)
    Runner->>Bench: Run benchmarks and upload results
    WF->>WF: Evaluate `python-unit-tests.result` in final/finish
    WF->>GH: Report combined status
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Suggested reviewers

  • CatherineSue
  • slin1237

"I hopped through YAML and nodes tonight,
I nudged heavy jobs to GPUs bright,
Runners wake and scale with glee,
Benchmarks hum, results set free,
A carrot-coded CI delight!" 🐇

🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately reflects the main change: configuring H100 runners for specific E2E tests (benchmarks and chat-completions-trtllm) that require enhanced GPU resources, along with tightened performance thresholds.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch xz/split-runner

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

The pull request introduces Kubernetes manifests for GitHub Actions self-hosted runners, including autoscalers, CPU and GPU runner deployments, and associated RBAC, aiming to provide dedicated H100 GPU runners for resource-intensive tasks. However, the RBAC configuration for the runner service account is overly permissive, granting broad access to secrets and the ability to execute commands in other pods within the namespace, which poses a significant security risk if a runner is compromised. Beyond this critical security concern, there are also areas for improvement regarding consistency and clarity in the configurations.

Comment thread scripts/k8s-runner-resources/arc-runner-gpu.yaml
Comment thread scripts/k8s-runner-resources/arc-runner-rbac.yaml
Comment thread scripts/k8s-runner-resources/arc-runner-rbac.yaml
repository: lightseekorg/smg
labels:
- k8s-runner-cpu
serviceAccountName: argo-runner-arc

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The serviceAccountName here (argo-runner-arc) is inconsistent with the arc-runner-sa used in the GPU runner deployments and defined in arc-runner-rbac.yaml. This inconsistency could lead to permission issues or confusion during deployment and operation. Please align the service account names for consistency and proper RBAC application.

      serviceAccountName: arc-runner-sa

privileged: true # Required for DinD
env:
- name: DOCKER_TLS_CERTDIR
value: "" # Disables TLS for shared socket use

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

Disabling TLS for the Docker socket (DOCKER_TLS_CERTDIR: "") removes a layer of security for communication with the Docker daemon. This is particularly risky when combined with privileged: true as it makes the Docker daemon susceptible to man-in-the-middle attacks if the network is not fully trusted. It's recommended to enable TLS for Docker communication if possible, or ensure the network path is secure.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Need this for docker in docker

Comment on lines +12 to +13
- 4-gpu-h100
- k8s-runner-gpu

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The podAffinityTerm for arc-runner-gpu-h100 uses runner-deployment-name in its labelSelector. For clarity and explicit control, it would be beneficial to explicitly add this label to the labels section of the runner pod template. This ensures that the affinity rule correctly targets pods belonging to this specific runner deployment.

      labels:
        - 4-gpu-h100
        - k8s-runner-gpu
        - runner-deployment-name: arc-runner-gpu-h100

Comment on lines +92 to +93
- 4-gpu-a10
- k8s-runner-gpu

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Similar to the H100 runner, the podAffinityTerm for arc-runner-gpu-a10 uses runner-deployment-name in its labelSelector. Please explicitly add this label to the labels section of the runner pod template for clarity and to ensure the affinity rule correctly targets pods belonging to this specific runner deployment.

      labels:
        - 4-gpu-a10
        - k8s-runner-gpu
        - runner-deployment-name: arc-runner-gpu-a10

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Fix all issues with AI agents
In @.github/workflows/pr-test-rust.yml:
- Line 604: The conditional that aggregates downstream job failures is missing
the python-unit-tests job: update the elif condition that currently checks
needs.python-lint.result, build-wheel, unit-tests, gateway-e2e,
gateway-e2e-heavy, go-unit-tests, and go-bindings-e2e to also include
needs.python-unit-tests.result == "failure" so the finish job fails when
python-unit-tests fails; locate the conditional in the workflow (the elif line
shown) and add the check for needs.python-unit-tests.result in the same format
as the other checks.

In `@scripts/k8s-runner-resources/arc-runner-cpu.yaml`:
- Line 13: The Deployment's serviceAccountName is set to the incorrect value
"argo-runner-arc" causing RBAC mismatch; update the serviceAccountName in this
YAML (the field serviceAccountName: argo-runner-arc) to the correct service
account "arc-runner-sa" so it matches the RBAC resource (arc-runner-rbac /
arc-runner-sa) and the GPU deployments.
🧹 Nitpick comments (6)
scripts/k8s-runner-resources/arc-runner-rbac.yaml (1)

13-13: Misleading comment: "Argo Workflows" should be "GitHub Actions Runners".

The comment suggests this is for Argo Workflows, but this RBAC is for GitHub Actions Runner Controller (ARC). Consider updating for clarity.

-  # Argo Workflows
+  # GitHub Actions Runner secrets access
scripts/k8s-runner-resources/arc-runner-cpu.yaml (1)

25-26: Docker sidecar missing DinD configuration present in GPU deployments.

The GPU deployments include securityContext: privileged: true, DOCKER_TLS_CERTDIR env var, and volume mounts for Docker socket and storage. The CPU deployment's docker container lacks these, which may prevent Docker-in-Docker from functioning correctly.

Consider aligning the DinD configuration with the GPU deployments if this runner needs Docker capabilities:

♻️ Suggested addition
         - name: docker
           image: fra.ocir.io/idqj093njucb/docker:dind
+          securityContext:
+            privileged: true
+          env:
+            - name: DOCKER_TLS_CERTDIR
+              value: ""
+          volumeMounts:
+            - name: docker-sock
+              mountPath: /var/run
+            - name: docker-storage
+              mountPath: /var/lib/docker

You would also need to add corresponding volume definitions in the spec.

scripts/k8s-runner-resources/arc-runner-autoscaler.yaml (1)

40-40: Minor naming inconsistency.

The naming pattern differs: arc-runner-h100-autoscaler, arc-runner-a10-autoscaler vs arc-cpu-runner-autoscaler. Consider using consistent naming like arc-runner-cpu-autoscaler for uniformity.

-  name: arc-cpu-runner-autoscaler
+  name: arc-runner-cpu-autoscaler
scripts/k8s-runner-resources/arc-runner-gpu.yaml (3)

96-98: Same deprecated label used in A10 deployment.

Apply the same fix as the H100 deployment.

       nodeSelector:
         nvidia.com/gpu: "true"
-        beta.kubernetes.io/instance-type: BM.GPU.A10.4
+        node.kubernetes.io/instance-type: BM.GPU.A10.4

52-67: Runner container missing CPU/memory resource requests.

The runner container only specifies GPU limits. Adding CPU and memory requests/limits (similar to the CPU runner's 8 CPU / 16Gi) would improve scheduling predictability and prevent resource contention on GPU nodes.


16-18: Replace deprecated node label in both runner deployments.

beta.kubernetes.io/instance-type has been deprecated since Kubernetes v1.17. Replace with node.kubernetes.io/instance-type for forward compatibility.

♻️ Fixes required

Line 18 (H100 deployment):

       nodeSelector:
         nvidia.com/gpu: "true"
-        beta.kubernetes.io/instance-type: BM.GPU.H100.8
+        node.kubernetes.io/instance-type: BM.GPU.H100.8

Line 98 (A10 deployment):

       nodeSelector:
         nvidia.com/gpu: "true"
-        beta.kubernetes.io/instance-type: BM.GPU.A10.4
+        node.kubernetes.io/instance-type: BM.GPU.A10.4

Comment thread .github/workflows/pr-test-rust.yml Outdated
Comment thread scripts/k8s-runner-resources/arc-runner-cpu.yaml Outdated
@CatherineSue

Copy link
Copy Markdown
Member

We need to reduce the timeout of the benchmarks test with switching from A10 to H100. For instance, TTFT should be reduced to 0.8 or 0.78. See issue #351

@github-actions github-actions Bot added the tests Test changes label Feb 7, 2026
Comment thread .github/workflows/pr-test-rust.yml Outdated
test_filter: ""
setup_trtllm: true
ignore_opts: ""
runs-on: k8s-runner-gpu

@key4ng key4ng Feb 7, 2026 •

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

could we use runs-on: ${{ matrix.runner }} here, and in each matrix we can set runner: 4-gpu-h100 for benchamrk and trt test, and runner k8s-runner-gpu for others. This eliminates the duplications of new gateway-e2e-heavy workflow and looks more clear

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah that would better! Fixed

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nice!
I'm just thinking can it be more neat to set it something like runs-on: ${{ matrix.runner || 'k8s-runner-gpu' }}, in this way, default will be k8s-runner-gpu, only benchmark and trt need set runner to override

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I see, sure, changed!

@CatherineSue CatherineSue left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This looks good to me. But I'd like to wait for @key4ng or @slin1237 's input as they are experts on this.

scaleUpThreshold: "0.95"
scaleDownThreshold: "0.25"
scaleUpFactor: "0.5"
scaleDownFactor: "0.5" No newline at end of file

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: add EOF line?

- name: docker-sock
mountPath: /var/run
- name: docker-storage
mountPath: /var/lib/docker No newline at end of file

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: same for EOF line.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Fix all issues with AI agents
In `@scripts/k8s-runner-resources/arc-runner-gpu.yaml`:
- Around line 16-19: The nodeSelector uses the deprecated label key
"beta.kubernetes.io/instance-type"; replace that key with the stable
"node.kubernetes.io/instance-type" wherever it appears (e.g., in the
nodeSelector block alongside "nvidia.com/gpu: \"true\"") to ensure scheduling
works on newer clusters—update all other occurrences of
"beta.kubernetes.io/instance-type" in the same manifest as well.
- Around line 69-74: The deployment uses privileged DinD (image:
fra.ocir.io/idqj093njucb/docker:dind) with securityContext.privileged: true and
DOCKER_TLS_CERTDIR set to "", which allows job containers full daemon control;
replace this by either running rootless Docker or a non-privileged build runtime
(e.g., switch to containerd/buildkit-based image and remove
securityContext.privileged), or enable TLS by removing DOCKER_TLS_CERTDIR:"" and
configuring DOCKER_TLS_CERTDIR to a secure path plus mounting scoped TLS
credentials/secrets for the runner; apply the same change for both DinD blocks
that set image: fra.ocir.io/.../dind, securityContext.privileged and env
DOCKER_TLS_CERTDIR.
🧹 Nitpick comments (1)
scripts/k8s-runner-resources/arc-runner-gpu.yaml (1)

52-79: Define CPU/memory requests and limits for runner + docker containers.

Without requests/limits, these pods may be best-effort and prone to eviction or noisy-neighbor contention, especially on GPU nodes. Add explicit CPU/memory requests/limits to stabilize scheduling and performance.

Also applies to: 132-159

Comment on lines +16 to +19
nodeSelector:
nvidia.com/gpu: "true"
beta.kubernetes.io/instance-type: BM.GPU.H100.8

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

🧩 Analysis chain

🌐 Web query:

Kubernetes instance-type label beta.kubernetes.io/instance-type deprecated node.kubernetes.io/instance-type

💡 Result:

beta.kubernetes.io/instance-type is a deprecated Node label. Since Kubernetes v1.17, you should use the stable replacement node.kubernetes.io/instance-type instead. [1]

What to change

  • Update any nodeSelector, nodeAffinity, Helm charts, or controllers that match on:
    • beta.kubernetes.io/instance-type=...
  • To instead match on:
    • node.kubernetes.io/instance-type=... [1]

Notes

  • The kubelet typically sets node.kubernetes.io/instance-type based on the cloud provider’s instance type (when a cloud provider integration is in use). [1]
  • Kubernetes generally recommends scheduling on capabilities (CPU/GPU/memory, taints/tolerations, topology) rather than pinning to specific instance types. [1]

Source

  • Kubernetes “Well-Known Labels, Annotations and Taints” reference. [1]

Replace deprecated instance-type label to avoid scheduling failures on newer clusters.

beta.kubernetes.io/instance-type has been deprecated since Kubernetes v1.17. Use the stable replacement node.kubernetes.io/instance-type instead.

🔧 Suggested update
      nodeSelector:
        nvidia.com/gpu: "true"
-       beta.kubernetes.io/instance-type: BM.GPU.H100.8
+       node.kubernetes.io/instance-type: BM.GPU.H100.8
      nodeSelector:
        nvidia.com/gpu: "true"
-       beta.kubernetes.io/instance-type: BM.GPU.A10.4
+       node.kubernetes.io/instance-type: BM.GPU.A10.4

Also applies to: 96-99

🤖 Prompt for AI Agents
In `@scripts/k8s-runner-resources/arc-runner-gpu.yaml` around lines 16 - 19, The
nodeSelector uses the deprecated label key "beta.kubernetes.io/instance-type";
replace that key with the stable "node.kubernetes.io/instance-type" wherever it
appears (e.g., in the nodeSelector block alongside "nvidia.com/gpu: \"true\"")
to ensure scheduling works on newer clusters—update all other occurrences of
"beta.kubernetes.io/instance-type" in the same manifest as well.

Comment on lines +69 to +74
image: fra.ocir.io/idqj093njucb/docker:dind
securityContext:
privileged: true # Required for DinD
env:
- name: DOCKER_TLS_CERTDIR
value: "" # Disables TLS for shared socket use

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

🧩 Analysis chain

🏁 Script executed:

# First, let's look at the file to understand the context
wc -l scripts/k8s-runner-resources/arc-runner-gpu.yaml

Repository: lightseekorg/smg

Length of output: 112


🏁 Script executed:

# Read the relevant sections - lines 69-74 and 149-154
sed -n '65,80p' scripts/k8s-runner-resources/arc-runner-gpu.yaml
echo "---"
sed -n '145,160p' scripts/k8s-runner-resources/arc-runner-gpu.yaml

Repository: lightseekorg/smg

Length of output: 1131


🌐 Web query:

Docker-in-Docker DinD Kubernetes security best practices 2025 privileged rootless TLS

💡 Result:

DinD on Kubernetes (2025→2026) security best practices (privileged, rootless, TLS)

1) Prefer not running DinD at all (best risk reduction)

  • Use daemonless / “no Docker socket” builders where possible (e.g., BuildKit rootless or Buildx Kubernetes driver with rootless=true) to avoid giving a pod a Docker daemon with elevated host-like power. Docker’s Buildx Kubernetes driver explicitly supports rootless mode and creates pods without securityContext.privileged. [1]

2) If you must run DinD, treat it as high risk

  • Assume privileged == near-host-level power: Kubernetes notes privileged containers override/ignore key hardening controls (seccomp/AppArmor/SELinux) and effectively gain broad capabilities. Avoid unless there’s no alternative. [2]
  • Enforce Pod Security Standards:
    • For most namespaces, enforce Baseline or Restricted so privileged: true, host namespaces, and hostPath mounts are disallowed by policy. [3]
    • Put the DinD workload in a tightly controlled exception namespace with explicit approvals.

3) Rootless: helpful, but don’t assume it removes the need for privilege

  • Rootless Docker runs the daemon and containers as a non-root user via user namespaces (reduces impact of daemon/runtime bugs). [4]
  • However, Docker’s own guidance for “Rootless Docker in Docker” still uses --privileged (to disable seccomp/AppArmor and mount restrictions), meaning you may still be bypassing important isolation controls even if the daemon UID isn’t 0. [5]

4) Use Kubernetes User Namespaces where available (additional containment)

  • Consider opting pods into user namespaces with spec.hostUsers: false so “root in the container” maps to an unprivileged host UID/GID, reducing host impact of a breakout. [6]
  • Note: user namespaces come with constraints/limitations (for example, cannot combine with hostNetwork/hostPID/hostIPC). [6]

5) Never expose an unauthenticated Docker API; use TLS or SSH

  • If you expose the Docker daemon beyond a local UNIX socket, Docker recommends securing it with SSH or mutual TLS (tlsverify). [7]

6) If privileged DinD is unavoidable, reduce blast radius aggressively

  • Dedicate a separate node pool (taints/tolerations) for DinD jobs; assume node compromise is possible.
  • Use tight NetworkPolicies (only allow what’s required) and short-lived, ephemeral workers.
  • Minimize mounts: avoid hostPath; don’t mount the host’s /var/run/docker.sock (that is effectively root on the node in many setups).
  • Keep privileged as a last resort; otherwise, grant only specific Linux capabilities (Kubernetes recommends capabilities over privileged when possible). [2]

Key takeaway: In Kubernetes, “DinD” usually pushes you toward privileged (or “privileged-like”) operation; the safest pattern is to avoid DinD and use rootless BuildKit/Buildx-on-Kubernetes where possible. [1][2][5]

Sources: Docker docs on Buildx Kubernetes rootless mode [1], Kubernetes on privileged container impacts [2], Kubernetes Pod Security Standards [3], Docker rootless mode [4], Docker rootless DinD tip showing --privileged [5], Kubernetes user namespaces (hostUsers: false) [6], Docker daemon TLS/SSH guidance [7].


Privileged DinD with TLS disabled is a security risk for CI runners.

A privileged Docker daemon plus DOCKER_TLS_CERTDIR="" gives any job container full daemon control. For untrusted PRs, this can lead to host escape. Consider rootless Docker, a non-privileged runtime (containerd/buildkit), or enabling TLS with scoped credentials.

This applies to both locations: lines 69-74 and 149-154.

🤖 Prompt for AI Agents
In `@scripts/k8s-runner-resources/arc-runner-gpu.yaml` around lines 69 - 74, The
deployment uses privileged DinD (image: fra.ocir.io/idqj093njucb/docker:dind)
with securityContext.privileged: true and DOCKER_TLS_CERTDIR set to "", which
allows job containers full daemon control; replace this by either running
rootless Docker or a non-privileged build runtime (e.g., switch to
containerd/buildkit-based image and remove securityContext.privileged), or
enable TLS by removing DOCKER_TLS_CERTDIR:"" and configuring DOCKER_TLS_CERTDIR
to a secure path plus mounting scoped TLS credentials/secrets for the runner;
apply the same change for both DinD blocks that set image: fra.ocir.io/.../dind,
securityContext.privileged and env DOCKER_TLS_CERTDIR.

@key4ng key4ng left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

overall lgtm

name: arc-runner-gpu-a10
namespace: actions-runner-system
spec:
replicas: 2

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I thought we have 4 A10, do we only have 2? I also didn't see hpa for a10.

@XinyueZhang369 XinyueZhang369 Feb 7, 2026 •

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we probably don't need to provision all 4 A20 all the time, so here is set to 2 for now. For HPA, today the auto scaler somehow kept creating cpu and a10 runner pods that cannot register to the repo regardless the max number, so I deleted all the old auto scalers for cpu, a10 and h100, this one is a new configuration, I want to bake it for some times, since most resources are h100, I only create for h100 for now for baking, once the scaling strategy works stably, I'll create the same for a10 and update this file

@slin1237
slin1237 merged commit 941d22a into main Feb 7, 2026
19 checks passed
@slin1237
slin1237 deleted the xz/split-runner branch February 7, 2026 09:37
pallasathena92 pushed a commit that referenced this pull request Feb 8, 2026
Co-authored-by: xinyzzha <xinyue.zhang@oracle.com>
ppraneth pushed a commit that referenced this pull request Feb 18, 2026
Co-authored-by: xinyzzha <xinyue.zhang@oracle.com>
Signed-off-by: ppraneth <pranethparuchuri@gmail.com>
yetone added a commit that referenced this pull request Apr 24, 2026
``feat/dense-llama-model-registry`` was the temporary pin added when
lightseekorg/tokenspeed#357 was in flight. That PR merged and the
source branch was deleted, so the pinned clone now fails with
``Remote branch feat/dense-llama-model-registry not found in upstream
origin``. ``main`` includes #357, so we can flip back.
yetone added a commit that referenced this pull request Apr 24, 2026
``feat/dense-llama-model-registry`` was the temporary pin added when
lightseekorg/tokenspeed#357 was in flight. That PR merged and the
source branch was deleted, so the pinned clone now fails with
``Remote branch feat/dense-llama-model-registry not found in upstream
origin``. ``main`` includes #357, so we can flip back.

Signed-off-by: yetone <yetoneful@gmail.com>
yetone added a commit that referenced this pull request Apr 27, 2026
``feat/dense-llama-model-registry`` was the temporary pin added when
lightseekorg/tokenspeed#357 was in flight. That PR merged and the
source branch was deleted, so the pinned clone now fails with
``Remote branch feat/dense-llama-model-registry not found in upstream
origin``. ``main`` includes #357, so we can flip back.

Signed-off-by: yetone <yetoneful@gmail.com>
CatherineSue pushed a commit that referenced this pull request Apr 30, 2026
``feat/dense-llama-model-registry`` was the temporary pin added when
lightseekorg/tokenspeed#357 was in flight. That PR merged and the
source branch was deleted, so the pinned clone now fails with
``Remote branch feat/dense-llama-model-registry not found in upstream
origin``. ``main`` includes #357, so we can flip back.

Signed-off-by: yetone <yetoneful@gmail.com>
key4ng pushed a commit that referenced this pull request Apr 30, 2026
``feat/dense-llama-model-registry`` was the temporary pin added when
lightseekorg/tokenspeed#357 was in flight. That PR merged and the
source branch was deleted, so the pinned clone now fails with
``Remote branch feat/dense-llama-model-registry not found in upstream
origin``. ``main`` includes #357, so we can flip back.

Signed-off-by: yetone <yetoneful@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci CI/CD configuration changes tests Test changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants