Skip to content

fix: small bug for llama-405b fp16 - #733

Merged
Ying1123 merged 1 commit into
mainfrom
ying-config-fix
Jul 26, 2024
Merged

Ying1123 merged 1 commit into
mainfrom
ying-config-fix

Conversation

@Ying1123

Copy link
Copy Markdown
Contributor

No description provided.

@Ying1123
Ying1123 merged commit 252e0f7 into main Jul 26, 2024
@Ying1123
Ying1123 deleted the ying-config-fix branch July 26, 2024 04:14
timethink pushed a commit to timethink/sglang that referenced this pull request Mar 9, 2025
cen121212 pushed a commit to cen121212/sglang that referenced this pull request Nov 10, 2025
<!-- Thank you for your contribution! Please follow these guidelines to
enhance your pull request. If anything is unclear, submit your PR and
reach out to maintainers for assistance. Join our Slack community at
https://slack.sglang.ai to discuss further. -->

## Motivation

<!-- Describe the purpose and goals of this pull request. -->

## Modifications

<!-- Detail the changes made in this pull request. -->

## Accuracy Tests

<!-- If this pull request affects model outputs (e.g., changes to the
kernel or model forward code), provide accuracy test results. -->

## Benchmarking and Profiling

<!-- If this pull request impacts inference speed, provide benchmarking
and profiling results. -->

## Checklist

- [ ] Format your code according to the [Format code with
pre-commit](https://docs.sglang.ai/developer_guide/contribution_guide.html#format-code-with-pre-commit).
- [ ] Add unit tests according to the [Run and add unit
tests](https://docs.sglang.ai/developer_guide/contribution_guide.html#run-and-add-unit-tests).
- [ ] Update documentation according to [Write
documentations](https://docs.sglang.ai/developer_guide/contribution_guide.html#write-documentations).
- [ ] Provide accuracy and speed benchmark results according to [Test
the
accuracy](https://docs.sglang.ai/developer_guide/contribution_guide.html#test-the-accuracy)
and [Benchmark the
speed](https://docs.sglang.ai/developer_guide/contribution_guide.html#benchmark-the-speed).
efschu pushed a commit to efschu/htsglang that referenced this pull request Aug 17, 2026
…, and I misquoted the row that decides it

Picks the wire for the full-plan's 5 MiB prefill crossing. Successor to
NOTE_732_breakable_crossing.md (9f2a61a), which priced the count and left
the transport open. Desk survey; no benchmark run here.

The best-matched row already existed under a misleading name. uneven_perf.py
:1355-1362 is NCCL dist.send/recv at exactly 512*1024 bf16 = 1 MiB -- right
transport class, right topology class, one binary order below the crossing.
Four boot caches agree within 1.5%: 9.06 / 5.83 / 5.10 GB/s, keyed by GPU
UUID, which sidesteps the torch-vs-NVML ordering trap. Live nvidia-smi
resolves the x4 card as id0, and the matrix reads cleanly: both pairs that
include id0 are dragged to 5.1-5.8, the pair that excludes it gets 9.06.

The correction that changes the answer: my predecessor note cited the BAR1
2-rank weak spot as a flat 0.86-0.99x. The source row (FEATURES_VS_UPSTREAM
.md:1349) says "on the fast x8 pair the transport loses ... down to 0.81x.
On the x4 pair and at three cards it wins everywhere." I quoted the first
clause and dropped the second -- the same error I corrected in others this
week. It matters because the crossings straddle exactly that boundary: 16
run over the fast x8 edge where BAR1 loses, 15 over the x4 edge where it
wins. So a single transport verdict is the wrong shape of answer. NCCL for
the x8 half, BAR1 for the x4 half -- and the x4 half is 62% of the crossing
bill, the opposite of the impression my earlier note left. "BAR1 off the
critical path" narrows to "off the x8 half, on the x4 half".

Link-aware placement, asked and answered both ways. Each FA layer moved off
the x4 card saves 0.899 ms/pass, so 8/8 is NOT free -- it is a 3.59 ms/pass
concession against 12/4, and Slot-2's "free" should be corrected. But KV
prices the other direction and wins: 852 MiB per FA layer at the live
max-total-tokens, so rebalancing costs 4.41 GiB of KV headroom per
percentage point of pass time on a 20 GiB card. Keep 8/8. One lever IS free:
layer 63 is terminal (1 crossing, not 2), so placing it on the x4 card saves
0.45 ms/pass at zero KV cost. The odd crossing belongs on the slow link.

sgl-project#733 coupling: id2 is the x8 3080, i.e. the FAST crossing partner. Because
can_access_peer is false on all six directed pairs, NCCL crossings are
host-staged, so spill and crossings contend for host memory bandwidth and
the same DMA path -- not merely for id2's lanes. d2d_bench.json makes the
staging visible directly: "direct" is at best equal to "staged" and mostly
slower (3.91 vs 4.09 GiB/s at 8 MiB). BAR1 writes peer VRAM via dmabuf and
does not share that bottleneck, so the idle-path ranking may not survive
load -- that is gap 5, the only open item that can invert a recommendation.

Also closed a gap I expected to have to chase: barlink_host is not unranked
for want of a MiB row, it is measured and refused -- NCCL wins 4 of 5 sizes
and the "5.1x" was a ping-pong artefact with an 8-byte return leg. The
20 KiB 7.30us figure is the same artefact family. It should not reappear in
a shortlist.

Two upgrades to the predecessor note, both toward more confidence:
- transport share refines 5.7% -> 5.18% of a pass (per-link instead of a
  uniform 5.6 GB/s); delta over the PP3 baseline ~4.85%.
- the 1 MiB -> 5 MiB extrapolation is now bounded: d2d_bench shows the curve
  flat from ~4 MiB, so 1 MiB sits 1-6% below asymptote and the error has a
  known sign (pessimistic, never optimistic).
- fixed a miscitation of my own: the pairwise values are UUID-keyed entries,
  not "__group__" rows.

Also folds in the sgl-project#494 result: the per-break cost does not exist at any
commit reachable from --all. break_cost_clock.py was built (00ac1da) and
has never produced a data point on a card -- one 2026-08-04 boot armed it and
crashed in graph capture first. Upgrades the earlier scoping from "not
reachable from this branch" to a searched absence.

Arithmetic verified: every split sums to 31 crossings; per-crossing times,
the 0.899 ms/layer lever, the 852 MiB KV figure and the 62.5% share all
recomputed from config.json geometry and the profile rows.

Anchors verified at file:line: uneven_perf.py:1355-1362, barlink_host.py
:1100/:1120, barlink_bar1.py:2562, ANALYSE_732_bar1_repricing.md:212-218,
FEATURES_VS_UPSTREAM.md:1349 (repo root, not docs/dev -- corrected),
INTEGRATION_R3_VALIDATION.md:13024.

Open: gap 5 (BAR1-vs-NCCL under host-path load) first, then gap 8 (BAR1
per-size table for the x4 pair, 1-8 MiB) which sizes the recommendation's
BAR1 half. Build item filed: per-edge transport selection -- the seam picks
per communicator today, so the split recommendation is not yet expressible.
efschu pushed a commit to efschu/htsglang that referenced this pull request Aug 18, 2026
…s own arithmetic

Answer to F4-r4's compiled-extension question: does a barlink spin-wait kernel
stay resident while the TP group parks during PP-prefill, reading the
host-mapped abort word and producing the ~1.4 GB/s host->device stream on
rank2? No, on three independent grounds from barlink_bar1_ext.py.

NO PERSISTENT KERNEL EXISTS. Exactly three __global__ kernels are declared
(:546 mesh, :884 ring, :1115 a2a), all per-collective. No daemon, no
persistent-block design, nothing launched outside a collective. With the TP
group parked no collective is issued, so none is on the device.

NO UNBOUNDED SPIN. All three for(;;) loops (:736, :797, :1267) carry the same
three exits: all-arrived, a clock64/capCycles DEADLINE, and the host abort
probe. A collective whose peers never arrive -- the only way a spinner could
outlive its epoch -- terminates at capCycles. "Launched without a completing
epoch" is structurally prevented, not merely unobserved.

THE PROBE IS ~200x TOO SMALL, which is the decisive number. The word is read as
FOUR bytes (volatile unsigned int), not 64, and once per 1024 iterations
(BARLINK_BAR1_HOST_MASK 1023u) -- a rate limit whose rationale is written at
:231-236, including "a collective that completes never reaches the probe".
Granting the hypothesis 64 B per probe anyway: 1.4 GB/s needs ~22 M probes/s =
~22.4 G spin iterations/s = ~45 G BAR1 peer reads/s, which is unattainable and
would anyway dominate a DIFFERENT direction. Inverted: a spin loop at an
optimistic 100 M iter/s yields ~6 MB/s against ~1.4 GB/s observed.

The remaining observations then read the other way. "Stops while work
continues" is awkward for the hypothesis and natural against it: with the TP
group parked there are no barlink collectives during exactly the window where
the traffic IS present. Same for the phase boundary -- the asymmetric stream
lives in the PP phase, where barlink's TP collectives are what is NOT running.

WHERE I WOULD POINT NEXT, with evidence already in hand rather than a guess:
host-staged NCCL PP crossings. can_access_peer is false for all six directed
pairs (ANALYSE_732:212-218), the fork's own feature doc says NCCL "on this
topology falls back to host staging", and d2d_bench shows it directly --
"direct" is at best equal to "staged" (3.91 vs 4.09 GiB/s at 8 MiB) because the
direct peer copy has no peer path to take. Every PP crossing here is
GPU->host->GPU, so the receiving rank sees a host->device stream that exists
only while the PP phase crosses and vanishes at the flip to TP decode. That is
the observation.

Explicitly NOT claimed: that host-staged NCCL IS the source. Only that it fits
the shape and that the no-P2P property is already established. The
work-invariance (chunk rate 2x, byte rate flat) is the one datum it does not
immediately explain; fixed-size recycled staging buffers would make the path
cadence-bound rather than payload-bound, but that is a hypothesis I have not
tested and it is labelled as such.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant