Add autobuild of efa node exporter - #612
Merged
Merged
Conversation
mhuguesaws
added a commit
that referenced
this pull request
Mar 27, 2025
KeitaW
pushed a commit
that referenced
this pull request
Feb 17, 2026
dmvevents
added a commit
to dmvevents/awsome-distributed-ai
that referenced
this pull request
Aug 25, 2026
…rect-SHA fetches, benchmark provenance - kubernetes/vllm-deepep-v2-2node.yaml:145,150 — worker probe bracket form `[v]llm serve` so the exec-shell's own cmdline no longer self-matches pgrep (thread awslabs#1/awslabs#9) - kubernetes/vllm-deepep-v2-2node.yaml:97 — append `|| true` to the LEADER_IP substitution so the FATAL DNS guard can fire under set -e + pipefail (thread awslabs#2) - kubernetes/vllm-deepep-v2-2node.yaml:159 — hugepages-2Mi comment: needs pre-allocated 2Mi hugepages or the pod sits Pending (thread awslabs#11) - kubernetes/vllm-deepep-v2-2node.yaml:115 — DEEPEP_ARCH_LIST commented env now names 10.0 (b200) + 10.3 (b300) so the Blackwell knob is reachable from the manifest (thread awslabs#12) - recipe/verify-image.sh:24,33,34 — convert fi_info efa-direct, ncclGetLsaDevicePointer, ncclGinPlugin checks from `grep -q` to draining `[ "$(... | grep -c X)" -ge 1 ]` (SIGPIPE-141 flake under pipefail), matching Dockerfile Layer 5b (thread awslabs#3) - recipe/run-kernel-test.sh:16 — worker requires an explicit node-rank (`${3:?...}`) instead of defaulting to 0 and colliding with the leader; add the `case "$ROLE"` guard matching serve.sh (thread awslabs#4) - recipe/benchmark_probe.py:18,24 — MODEL reads SERVE_MODEL env with the current value as fallback (+ `import os`) so a non-default model doesn't 100%-fail (thread awslabs#5) - setup_deepep_v2_efa.sh:30,52 — fetch the immutable PR-head SHA directly instead of the moving `refs/pull/N/head` ref for both aws-ofi-nccl #1351 and DeepEP awslabs#612 (thread awslabs#6) - Dockerfile:116 — CMD banner names /opt/serve.sh + /opt/build_deepep.sh (the in-image paths), not recipe/ (thread awslabs#7) - recipe/build_deepep.sh:120 — `tee /tmp/deepep-build.log | tail -30` so a failed build keeps the full nvcc diagnostic (thread awslabs#8) - Dockerfile:60 — gdrcopy pinned by commit SHA (v2.5.2 == c91ad9f) via fetch/checkout instead of the movable `--branch v2.5.2` tag (thread awslabs#13) - recipe/serve.sh:124 — comment on why --trust-remote-code is unconditional (DeepSeek/Kimi models the preflight supports); chmod 755 serve.sh + run-kernel-test.sh in-tree to match the other four scripts (thread awslabs#14) - README.md:114 — in-pod benchmark exec sets OUT_ROOT=/work/benchmarks so results land on the /work volume, not the ephemeral container layer (thread awslabs#16) - setup/env_vars.example:2,12 — header says build-push.sh sources it and recipe/*.sh need `source` first; reword the AWS_OFI_NCCL_PR_SHA note (empty does NOT skip the cherry-pick) (thread awslabs#17) - benchmarks/README.md:88 — caveat: no alternative-backend baseline measured; every table is --all2all-backend deepep_v2 (thread awslabs#18) - recipe/serve.sh:48 — EP_EFA_MAX_QPS comment records the pinned plugin (9c44d34) is 76 commits past the seq-window redesign (6e504db) so the 128-slot cap's precondition no longer holds; default left unchanged (pods down) (thread awslabs#21) Already present at c025946 (prior commits), verified not duplicated: - README shared-experts caveat naming #47785 + DeepSeek (thread awslabs#15) - benchmark_probe.py requests-per-level distribution + unique prompt prefix (threads awslabs#19/awslabs#20) - probes + publishNotReadyAddresses structure (thread awslabs#9) and requests==limits Guaranteed QoS (thread awslabs#10) Signed-off-by: Anton Alexander <dmvevents@gmail.com>
dmvevents
added a commit
to dmvevents/awsome-distributed-ai
that referenced
this pull request
Aug 26, 2026
…ocument plugin/fork trade-offs (PR awslabs#1242 review) Address KeitaW review threads on pin provenance and honesty: - NGC base: pin @sha256:bbc2b67e… alongside the readable :26.02-py3 tag. It is the ABI anchor the whole image builds around (torch 2.11/CUDA 13 + baked TE/apex/flash-attn), so a silent tag re-push is the most consequential drift possible here. Override NGC_PYTORCH_BASE to move it (a re-measure event). - NCCL: pin the commit 1933fdd6 that v2.30.4-1 resolves to, not the bare tag — held to the same moving-ref standard as the gdrcopy/DeepEP/Megatron SHA pins. Correct the stale "same NCCL line the NGC base bakes / no drift" comment: the base bakes an OLDER 2.29.x line, so the 00-nccl-gin.conf ld.so.conf entry is LOAD-BEARING (it makes this GIN-capable source build win the loader search). - aws-ofi-nccl #1351: note it is CLOSED-UNMERGED, so unlike the draft PRs in patches/ it does NOT self-neutralize — it is a PERMANENT baseline carry on every build. Clarify the "baseline has zero dependence on unmerged PRs" line refers to the four draft PRs in the opt-in layer, not this plugin pin. - DeepEP fork trade-off: document honestly that amazon-contributing/DeepEP fixes trap-2 structurally (sysfs get_rdma_gbs, restructured QP allocator) while baseline stays stock 01dc3aa on purpose (the awslabs#612 --check-clean base + the no-fork convention); the knob clamps are the portable equivalent, and fork adoption is a deliberate future re-pin + re-measure. README pins table updated to match (base digest, NCCL commit). Signed-off-by: Anton Alexander <dmvevents@gmail.com>
dmvevents
added a commit
to dmvevents/awsome-distributed-ai
that referenced
this pull request
Sep 3, 2026
…ibuting fork; drop superseded awslabs#612 Repoint the DeepEP source from deepseek-ai/DeepEP + a refs/pull/612/head fetch-merge to the amazon-contributing/DeepEP fork, pinned at an immutable SHA (97d8f9bcc1). This matches the house V2/NCCL-Gin canonical micro-benchmarks/expert-parallelism/deepep-v2-benchmark/setup_deepep_gin.sh, which already clones the same fork. The fork carries the EFA delta in-code, including both halves of the draft deepseek-ai/DeepEP#612 that this sample previously fetch-merged: - the get_rdma_gbs() sysfs link-rate fast path (deep_ep/utils/envs.py) - the auto-QP overflow clamp (deep_ep/buffers/elastic.py) So awslabs#612 is superseded: pinning the fork HEAD is strictly ahead of the old base+PR-merge, and it addresses the review note that pinning stock upstream at the pre-fix fork point forfeits exactly those AWS fixes. - setup_deepep_v2_gdaki_efa.sh: DEEPEP_REPO -> amazon-contributing/DeepEP, DEEPEP_SHA -> fork HEAD; drop the refs/pull/612/head fetch+merge; add fail-loud asserts that both awslabs#612 fix-halves are present in the clone. - Dockerfile: rewrite the Layer-5 provenance header to name the fork; drop the DEEPEP_PR / DEEPEP_PR_SHA ARGs and their pass-through. - README.md: rewrite the source-pin paragraph + Net summary to the fork, naming both superseded awslabs#612 halves. No local source patches; the pin is override-able via DEEPEP_SHA. Test Results: docs-and-packaging change (source-repoint + provenance). An end-to-end GDAKI all-to-all rebuild+serve verification on 2x p5en is tracked separately as the GPU-verify half of this change. Signed-off-by: Anton Alexander <dmvevents@gmail.com>
dmvevents
added a commit
to dmvevents/awsome-distributed-ai
that referenced
this pull request
Sep 3, 2026
…-contributing fork; drop superseded awslabs#612 Repoint the DeepEP source from stock deepseek-ai/DeepEP @ 01dc3aa plus a draft deepseek-ai/DeepEP#612 opt-in patch to the amazon-contributing/DeepEP fork, pinned at the immutable HEAD 97d8f9bcc1be31e9036db2ab591ef9b9f4e38619. This mirrors the sibling examples/inference/vllm/deepep-v2-gdaki-efa repoint and answers the review note that pinning the pre-fix upstream fork-point forfeits the AWS EFA fixes: the fork is the AWS EPv2/NCCL-GIN tree and carries the EFA delta IN-CODE, so no awslabs#612 patch is applied on any flavor. The fork carries both correctness halves of what was draft awslabs#612 structurally: - the get_rdma_gbs() sysfs link-rate fast path (deep_ep/utils/envs.py, provider-agnostic — reads /sys/class/infiniband/<nic>/ports/*/rate), and - the auto-QP overflow clamp (deep_ep/buffers/elastic.py, clamps the allocated QP count to _C.{min,max}_unordered_gin_qps). It does NOT carry awslabs#612's third commit (a kScaleoutUpdateInterval 6->16 latency micro-opt); the fork keeps =6. No gate in this example depends on that value, so the two correctness fixes are what "supersedes awslabs#612" means here. Because the fork carries those fixes on every flavor, this collapses the baseline/opt-in distinction for DeepEP only: - the awslabs#612 entry drops out of the opt-in draft-PR layer entirely (the layer now bakes only Megatron-LM#4632 + NeMo-RL#2410); - the two dead env vars EP_EFA_MAX_QPS / EP_EFA_RDMA_GBS (the old awslabs#612 patch knobs, zero readers on the fork) are removed from env_vars.example and kubernetes/raycluster.yaml, replaced with an absence-note; and - the verify-image / setup fail-loud gate is rewritten to fork discriminators — it now asserts _get_sysfs_rdma_gbs and unordered_gin_qps are present in the installed deep_ep tree on EVERY flavor, not gated on the draft-PR marker. EP_NUM_QPS=2 is a different variable (the probe's explicit num_allocated_qps) and stays: it survives the fork clamp unchanged (max(2, min(2, max)) == 2). The Dockerfile pins the full 40-char SHA as DEEPEP_SHA; setup and verify-image fail-loud if the clone lacks either fix-half. This is the same fork the repo's micro-benchmarks/expert-parallelism/deepep-v2-benchmark/setup_deepep_gin.sh already clones. ## Test Results docs-and-packaging change (source-repoint + provenance). No functional change to the example's runtime behavior on either flavor; the fork already carried the fixes the removed patch supplied. markdownlint-cli2 clean on the changed README against the root .markdownlint.jsonc. Signed-off-by: Anton Alexander <dmvevents@gmail.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Issue #, if available:
Description of changes:
By submitting this pull request, I confirm that you can use, modify, copy, and redistribute this contribution, under the terms of your choice.