Skip to content

Add autobuild of efa node exporter - #612

Merged
mhuguesaws merged 1 commit into
mainfrom
feature/efa_node_exporer_autobuild
Mar 27, 2025
Merged

mhuguesaws merged 1 commit into
mainfrom
feature/efa_node_exporer_autobuild

Conversation

@mhuguesaws

Copy link
Copy Markdown
Contributor

Issue #, if available:

Description of changes:

By submitting this pull request, I confirm that you can use, modify, copy, and redistribute this contribution, under the terms of your choice.

@mhuguesaws
mhuguesaws merged commit d72c630 into main Mar 27, 2025
@mhuguesaws
mhuguesaws deleted the feature/efa_node_exporer_autobuild branch March 27, 2025 15:41
mhuguesaws added a commit that referenced this pull request Mar 27, 2025
KeitaW pushed a commit that referenced this pull request Feb 17, 2026
dmvevents added a commit to dmvevents/awsome-distributed-ai that referenced this pull request Aug 25, 2026
…rect-SHA fetches, benchmark provenance

- kubernetes/vllm-deepep-v2-2node.yaml:145,150 — worker probe bracket form `[v]llm serve` so the exec-shell's own cmdline no longer self-matches pgrep (thread awslabs#1/awslabs#9)
- kubernetes/vllm-deepep-v2-2node.yaml:97 — append `|| true` to the LEADER_IP substitution so the FATAL DNS guard can fire under set -e + pipefail (thread awslabs#2)
- kubernetes/vllm-deepep-v2-2node.yaml:159 — hugepages-2Mi comment: needs pre-allocated 2Mi hugepages or the pod sits Pending (thread awslabs#11)
- kubernetes/vllm-deepep-v2-2node.yaml:115 — DEEPEP_ARCH_LIST commented env now names 10.0 (b200) + 10.3 (b300) so the Blackwell knob is reachable from the manifest (thread awslabs#12)
- recipe/verify-image.sh:24,33,34 — convert fi_info efa-direct, ncclGetLsaDevicePointer, ncclGinPlugin checks from `grep -q` to draining `[ "$(... | grep -c X)" -ge 1 ]` (SIGPIPE-141 flake under pipefail), matching Dockerfile Layer 5b (thread awslabs#3)
- recipe/run-kernel-test.sh:16 — worker requires an explicit node-rank (`${3:?...}`) instead of defaulting to 0 and colliding with the leader; add the `case "$ROLE"` guard matching serve.sh (thread awslabs#4)
- recipe/benchmark_probe.py:18,24 — MODEL reads SERVE_MODEL env with the current value as fallback (+ `import os`) so a non-default model doesn't 100%-fail (thread awslabs#5)
- setup_deepep_v2_efa.sh:30,52 — fetch the immutable PR-head SHA directly instead of the moving `refs/pull/N/head` ref for both aws-ofi-nccl #1351 and DeepEP awslabs#612 (thread awslabs#6)
- Dockerfile:116 — CMD banner names /opt/serve.sh + /opt/build_deepep.sh (the in-image paths), not recipe/ (thread awslabs#7)
- recipe/build_deepep.sh:120 — `tee /tmp/deepep-build.log | tail -30` so a failed build keeps the full nvcc diagnostic (thread awslabs#8)
- Dockerfile:60 — gdrcopy pinned by commit SHA (v2.5.2 == c91ad9f) via fetch/checkout instead of the movable `--branch v2.5.2` tag (thread awslabs#13)
- recipe/serve.sh:124 — comment on why --trust-remote-code is unconditional (DeepSeek/Kimi models the preflight supports); chmod 755 serve.sh + run-kernel-test.sh in-tree to match the other four scripts (thread awslabs#14)
- README.md:114 — in-pod benchmark exec sets OUT_ROOT=/work/benchmarks so results land on the /work volume, not the ephemeral container layer (thread awslabs#16)
- setup/env_vars.example:2,12 — header says build-push.sh sources it and recipe/*.sh need `source` first; reword the AWS_OFI_NCCL_PR_SHA note (empty does NOT skip the cherry-pick) (thread awslabs#17)
- benchmarks/README.md:88 — caveat: no alternative-backend baseline measured; every table is --all2all-backend deepep_v2 (thread awslabs#18)
- recipe/serve.sh:48 — EP_EFA_MAX_QPS comment records the pinned plugin (9c44d34) is 76 commits past the seq-window redesign (6e504db) so the 128-slot cap's precondition no longer holds; default left unchanged (pods down) (thread awslabs#21)

Already present at c025946 (prior commits), verified not duplicated:
- README shared-experts caveat naming #47785 + DeepSeek (thread awslabs#15)
- benchmark_probe.py requests-per-level distribution + unique prompt prefix (threads awslabs#19/awslabs#20)
- probes + publishNotReadyAddresses structure (thread awslabs#9) and requests==limits Guaranteed QoS (thread awslabs#10)

Signed-off-by: Anton Alexander <dmvevents@gmail.com>
dmvevents added a commit to dmvevents/awsome-distributed-ai that referenced this pull request Aug 26, 2026
…ocument plugin/fork trade-offs (PR awslabs#1242 review)

Address KeitaW review threads on pin provenance and honesty:

- NGC base: pin @sha256:bbc2b67e… alongside the readable :26.02-py3 tag. It is
  the ABI anchor the whole image builds around (torch 2.11/CUDA 13 + baked
  TE/apex/flash-attn), so a silent tag re-push is the most consequential drift
  possible here. Override NGC_PYTORCH_BASE to move it (a re-measure event).

- NCCL: pin the commit 1933fdd6 that v2.30.4-1 resolves to, not the bare tag —
  held to the same moving-ref standard as the gdrcopy/DeepEP/Megatron SHA pins.
  Correct the stale "same NCCL line the NGC base bakes / no drift" comment: the
  base bakes an OLDER 2.29.x line, so the 00-nccl-gin.conf ld.so.conf entry is
  LOAD-BEARING (it makes this GIN-capable source build win the loader search).

- aws-ofi-nccl #1351: note it is CLOSED-UNMERGED, so unlike the draft PRs in
  patches/ it does NOT self-neutralize — it is a PERMANENT baseline carry on
  every build. Clarify the "baseline has zero dependence on unmerged PRs" line
  refers to the four draft PRs in the opt-in layer, not this plugin pin.

- DeepEP fork trade-off: document honestly that amazon-contributing/DeepEP
  fixes trap-2 structurally (sysfs get_rdma_gbs, restructured QP allocator)
  while baseline stays stock 01dc3aa on purpose (the awslabs#612 --check-clean base +
  the no-fork convention); the knob clamps are the portable equivalent, and
  fork adoption is a deliberate future re-pin + re-measure.

README pins table updated to match (base digest, NCCL commit).

Signed-off-by: Anton Alexander <dmvevents@gmail.com>
dmvevents added a commit to dmvevents/awsome-distributed-ai that referenced this pull request Sep 3, 2026
…ibuting fork; drop superseded awslabs#612

Repoint the DeepEP source from deepseek-ai/DeepEP + a refs/pull/612/head
fetch-merge to the amazon-contributing/DeepEP fork, pinned at an immutable
SHA (97d8f9bcc1). This matches the house V2/NCCL-Gin canonical
micro-benchmarks/expert-parallelism/deepep-v2-benchmark/setup_deepep_gin.sh,
which already clones the same fork.

The fork carries the EFA delta in-code, including both halves of the draft
deepseek-ai/DeepEP#612 that this sample previously fetch-merged:
  - the get_rdma_gbs() sysfs link-rate fast path (deep_ep/utils/envs.py)
  - the auto-QP overflow clamp (deep_ep/buffers/elastic.py)
So awslabs#612 is superseded: pinning the fork HEAD is strictly ahead of the old
base+PR-merge, and it addresses the review note that pinning stock upstream
at the pre-fix fork point forfeits exactly those AWS fixes.

- setup_deepep_v2_gdaki_efa.sh: DEEPEP_REPO -> amazon-contributing/DeepEP,
  DEEPEP_SHA -> fork HEAD; drop the refs/pull/612/head fetch+merge; add
  fail-loud asserts that both awslabs#612 fix-halves are present in the clone.
- Dockerfile: rewrite the Layer-5 provenance header to name the fork; drop
  the DEEPEP_PR / DEEPEP_PR_SHA ARGs and their pass-through.
- README.md: rewrite the source-pin paragraph + Net summary to the fork,
  naming both superseded awslabs#612 halves.

No local source patches; the pin is override-able via DEEPEP_SHA.

Test Results: docs-and-packaging change (source-repoint + provenance).
An end-to-end GDAKI all-to-all rebuild+serve verification on 2x p5en is
tracked separately as the GPU-verify half of this change.

Signed-off-by: Anton Alexander <dmvevents@gmail.com>
dmvevents added a commit to dmvevents/awsome-distributed-ai that referenced this pull request Sep 3, 2026
…-contributing fork; drop superseded awslabs#612

Repoint the DeepEP source from stock deepseek-ai/DeepEP @ 01dc3aa plus a draft
deepseek-ai/DeepEP#612 opt-in patch to the amazon-contributing/DeepEP fork,
pinned at the immutable HEAD 97d8f9bcc1be31e9036db2ab591ef9b9f4e38619. This
mirrors the sibling examples/inference/vllm/deepep-v2-gdaki-efa repoint and
answers the review note that pinning the pre-fix upstream fork-point forfeits
the AWS EFA fixes: the fork is the AWS EPv2/NCCL-GIN tree and carries the EFA
delta IN-CODE, so no awslabs#612 patch is applied on any flavor.

The fork carries both correctness halves of what was draft awslabs#612 structurally:
  - the get_rdma_gbs() sysfs link-rate fast path (deep_ep/utils/envs.py,
    provider-agnostic — reads /sys/class/infiniband/<nic>/ports/*/rate), and
  - the auto-QP overflow clamp (deep_ep/buffers/elastic.py, clamps the
    allocated QP count to _C.{min,max}_unordered_gin_qps).
It does NOT carry awslabs#612's third commit (a kScaleoutUpdateInterval 6->16 latency
micro-opt); the fork keeps =6. No gate in this example depends on that value,
so the two correctness fixes are what "supersedes awslabs#612" means here.

Because the fork carries those fixes on every flavor, this collapses the
baseline/opt-in distinction for DeepEP only:
  - the awslabs#612 entry drops out of the opt-in draft-PR layer entirely (the layer
    now bakes only Megatron-LM#4632 + NeMo-RL#2410);
  - the two dead env vars EP_EFA_MAX_QPS / EP_EFA_RDMA_GBS (the old awslabs#612 patch
    knobs, zero readers on the fork) are removed from env_vars.example and
    kubernetes/raycluster.yaml, replaced with an absence-note; and
  - the verify-image / setup fail-loud gate is rewritten to fork
    discriminators — it now asserts _get_sysfs_rdma_gbs and unordered_gin_qps
    are present in the installed deep_ep tree on EVERY flavor, not gated on the
    draft-PR marker.
EP_NUM_QPS=2 is a different variable (the probe's explicit num_allocated_qps)
and stays: it survives the fork clamp unchanged (max(2, min(2, max)) == 2).

The Dockerfile pins the full 40-char SHA as DEEPEP_SHA; setup and verify-image
fail-loud if the clone lacks either fix-half. This is the same fork the repo's
micro-benchmarks/expert-parallelism/deepep-v2-benchmark/setup_deepep_gin.sh
already clones.

## Test Results

docs-and-packaging change (source-repoint + provenance). No functional change
to the example's runtime behavior on either flavor; the fork already carried
the fixes the removed patch supplied. markdownlint-cli2 clean on the changed
README against the root .markdownlint.jsonc.

Signed-off-by: Anton Alexander <dmvevents@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant