fix: small bug for llama-405b fp16 - #733
Merged
Merged
Conversation
3 tasks
timethink
pushed a commit
to timethink/sglang
that referenced
this pull request
Mar 9, 2025
cen121212
pushed a commit
to cen121212/sglang
that referenced
this pull request
Nov 10, 2025
<!-- Thank you for your contribution! Please follow these guidelines to enhance your pull request. If anything is unclear, submit your PR and reach out to maintainers for assistance. Join our Slack community at https://slack.sglang.ai to discuss further. --> ## Motivation <!-- Describe the purpose and goals of this pull request. --> ## Modifications <!-- Detail the changes made in this pull request. --> ## Accuracy Tests <!-- If this pull request affects model outputs (e.g., changes to the kernel or model forward code), provide accuracy test results. --> ## Benchmarking and Profiling <!-- If this pull request impacts inference speed, provide benchmarking and profiling results. --> ## Checklist - [ ] Format your code according to the [Format code with pre-commit](https://docs.sglang.ai/developer_guide/contribution_guide.html#format-code-with-pre-commit). - [ ] Add unit tests according to the [Run and add unit tests](https://docs.sglang.ai/developer_guide/contribution_guide.html#run-and-add-unit-tests). - [ ] Update documentation according to [Write documentations](https://docs.sglang.ai/developer_guide/contribution_guide.html#write-documentations). - [ ] Provide accuracy and speed benchmark results according to [Test the accuracy](https://docs.sglang.ai/developer_guide/contribution_guide.html#test-the-accuracy) and [Benchmark the speed](https://docs.sglang.ai/developer_guide/contribution_guide.html#benchmark-the-speed).
efschu
pushed a commit
to efschu/htsglang
that referenced
this pull request
Aug 17, 2026
…, and I misquoted the row that decides it Picks the wire for the full-plan's 5 MiB prefill crossing. Successor to NOTE_732_breakable_crossing.md (9f2a61a), which priced the count and left the transport open. Desk survey; no benchmark run here. The best-matched row already existed under a misleading name. uneven_perf.py :1355-1362 is NCCL dist.send/recv at exactly 512*1024 bf16 = 1 MiB -- right transport class, right topology class, one binary order below the crossing. Four boot caches agree within 1.5%: 9.06 / 5.83 / 5.10 GB/s, keyed by GPU UUID, which sidesteps the torch-vs-NVML ordering trap. Live nvidia-smi resolves the x4 card as id0, and the matrix reads cleanly: both pairs that include id0 are dragged to 5.1-5.8, the pair that excludes it gets 9.06. The correction that changes the answer: my predecessor note cited the BAR1 2-rank weak spot as a flat 0.86-0.99x. The source row (FEATURES_VS_UPSTREAM .md:1349) says "on the fast x8 pair the transport loses ... down to 0.81x. On the x4 pair and at three cards it wins everywhere." I quoted the first clause and dropped the second -- the same error I corrected in others this week. It matters because the crossings straddle exactly that boundary: 16 run over the fast x8 edge where BAR1 loses, 15 over the x4 edge where it wins. So a single transport verdict is the wrong shape of answer. NCCL for the x8 half, BAR1 for the x4 half -- and the x4 half is 62% of the crossing bill, the opposite of the impression my earlier note left. "BAR1 off the critical path" narrows to "off the x8 half, on the x4 half". Link-aware placement, asked and answered both ways. Each FA layer moved off the x4 card saves 0.899 ms/pass, so 8/8 is NOT free -- it is a 3.59 ms/pass concession against 12/4, and Slot-2's "free" should be corrected. But KV prices the other direction and wins: 852 MiB per FA layer at the live max-total-tokens, so rebalancing costs 4.41 GiB of KV headroom per percentage point of pass time on a 20 GiB card. Keep 8/8. One lever IS free: layer 63 is terminal (1 crossing, not 2), so placing it on the x4 card saves 0.45 ms/pass at zero KV cost. The odd crossing belongs on the slow link. sgl-project#733 coupling: id2 is the x8 3080, i.e. the FAST crossing partner. Because can_access_peer is false on all six directed pairs, NCCL crossings are host-staged, so spill and crossings contend for host memory bandwidth and the same DMA path -- not merely for id2's lanes. d2d_bench.json makes the staging visible directly: "direct" is at best equal to "staged" and mostly slower (3.91 vs 4.09 GiB/s at 8 MiB). BAR1 writes peer VRAM via dmabuf and does not share that bottleneck, so the idle-path ranking may not survive load -- that is gap 5, the only open item that can invert a recommendation. Also closed a gap I expected to have to chase: barlink_host is not unranked for want of a MiB row, it is measured and refused -- NCCL wins 4 of 5 sizes and the "5.1x" was a ping-pong artefact with an 8-byte return leg. The 20 KiB 7.30us figure is the same artefact family. It should not reappear in a shortlist. Two upgrades to the predecessor note, both toward more confidence: - transport share refines 5.7% -> 5.18% of a pass (per-link instead of a uniform 5.6 GB/s); delta over the PP3 baseline ~4.85%. - the 1 MiB -> 5 MiB extrapolation is now bounded: d2d_bench shows the curve flat from ~4 MiB, so 1 MiB sits 1-6% below asymptote and the error has a known sign (pessimistic, never optimistic). - fixed a miscitation of my own: the pairwise values are UUID-keyed entries, not "__group__" rows. Also folds in the sgl-project#494 result: the per-break cost does not exist at any commit reachable from --all. break_cost_clock.py was built (00ac1da) and has never produced a data point on a card -- one 2026-08-04 boot armed it and crashed in graph capture first. Upgrades the earlier scoping from "not reachable from this branch" to a searched absence. Arithmetic verified: every split sums to 31 crossings; per-crossing times, the 0.899 ms/layer lever, the 852 MiB KV figure and the 62.5% share all recomputed from config.json geometry and the profile rows. Anchors verified at file:line: uneven_perf.py:1355-1362, barlink_host.py :1100/:1120, barlink_bar1.py:2562, ANALYSE_732_bar1_repricing.md:212-218, FEATURES_VS_UPSTREAM.md:1349 (repo root, not docs/dev -- corrected), INTEGRATION_R3_VALIDATION.md:13024. Open: gap 5 (BAR1-vs-NCCL under host-path load) first, then gap 8 (BAR1 per-size table for the x4 pair, 1-8 MiB) which sizes the recommendation's BAR1 half. Build item filed: per-edge transport selection -- the seam picks per communicator today, so the split recommendation is not yet expressible.
efschu
pushed a commit
to efschu/htsglang
that referenced
this pull request
Aug 18, 2026
…s own arithmetic Answer to F4-r4's compiled-extension question: does a barlink spin-wait kernel stay resident while the TP group parks during PP-prefill, reading the host-mapped abort word and producing the ~1.4 GB/s host->device stream on rank2? No, on three independent grounds from barlink_bar1_ext.py. NO PERSISTENT KERNEL EXISTS. Exactly three __global__ kernels are declared (:546 mesh, :884 ring, :1115 a2a), all per-collective. No daemon, no persistent-block design, nothing launched outside a collective. With the TP group parked no collective is issued, so none is on the device. NO UNBOUNDED SPIN. All three for(;;) loops (:736, :797, :1267) carry the same three exits: all-arrived, a clock64/capCycles DEADLINE, and the host abort probe. A collective whose peers never arrive -- the only way a spinner could outlive its epoch -- terminates at capCycles. "Launched without a completing epoch" is structurally prevented, not merely unobserved. THE PROBE IS ~200x TOO SMALL, which is the decisive number. The word is read as FOUR bytes (volatile unsigned int), not 64, and once per 1024 iterations (BARLINK_BAR1_HOST_MASK 1023u) -- a rate limit whose rationale is written at :231-236, including "a collective that completes never reaches the probe". Granting the hypothesis 64 B per probe anyway: 1.4 GB/s needs ~22 M probes/s = ~22.4 G spin iterations/s = ~45 G BAR1 peer reads/s, which is unattainable and would anyway dominate a DIFFERENT direction. Inverted: a spin loop at an optimistic 100 M iter/s yields ~6 MB/s against ~1.4 GB/s observed. The remaining observations then read the other way. "Stops while work continues" is awkward for the hypothesis and natural against it: with the TP group parked there are no barlink collectives during exactly the window where the traffic IS present. Same for the phase boundary -- the asymmetric stream lives in the PP phase, where barlink's TP collectives are what is NOT running. WHERE I WOULD POINT NEXT, with evidence already in hand rather than a guess: host-staged NCCL PP crossings. can_access_peer is false for all six directed pairs (ANALYSE_732:212-218), the fork's own feature doc says NCCL "on this topology falls back to host staging", and d2d_bench shows it directly -- "direct" is at best equal to "staged" (3.91 vs 4.09 GiB/s at 8 MiB) because the direct peer copy has no peer path to take. Every PP crossing here is GPU->host->GPU, so the receiving rank sees a host->device stream that exists only while the PP phase crosses and vanishes at the flip to TP decode. That is the observation. Explicitly NOT claimed: that host-staged NCCL IS the source. Only that it fits the shape and that the no-P2P property is already established. The work-invariance (chunk rate 2x, byte rate flat) is the one datum it does not immediately explain; fixed-size recycled staging buffers would make the path cadence-bound rather than payload-bound, but that is a hypothesis I have not tested and it is labelled as such.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.