Repository navigation
support command-r - #369
Merged
Merged
Conversation
Qubitium
reviewed
Apr 16, 2024
Qubitium
left a comment
Contributor
There was a problem hiding this comment.
We have tested the code with command-r-v01 on inference with batch size ranging from 1 to 10 with no issue.
|
Confirmed working with Command-R-Plus as well. (GPTQ quants) good work! |
Contributor
|
@ZhouXingg thanks for the contribution! |
timethink
pushed a commit
to timethink/sglang
that referenced
this pull request
Mar 9, 2025
UNIDY2002
pushed a commit
to HanHan009527/sglang
that referenced
this pull request
Jan 14, 2026
Co-authored-by: sunqi.7 <sunqi.7@bytedance.com>
heziiop
pushed a commit
to heziiop/sglang-community
that referenced
this pull request
Apr 13, 2026
efschu
added a commit
to efschu/htsglang
that referenced
this pull request
Jul 31, 2026
…n co-location race, dead builder fails loudly; graph gate stays unflipped until the proof can run (sgl-project#369)
efschu
added a commit
to efschu/htsglang
that referenced
this pull request
Aug 1, 2026
… waiters now run under a build deadline (sgl-project#369) ROUTE 3 WORKS: THE PROOF PASSED `benchmark/bar1_graph_check.py` passed 10/10 on three cards (5090 + 2x 3080) inside the htsglang Docker image on the Proxmox host -- all nine gate cases plus the informational `grid` case, which IS the question the release was waiting on: whether the driver accepts `cudaLaunchCooperativeKernel` from inside a stream capture. It does. PASSED [Gate] 1blk-small PASSED [Gate] pipe PASSED [Gate] 1blk-large PASSED [Gate] pipe-direct PASSED [Info] grid PASSED [Gate] pipe-direct-pool-empty PASSED [Gate] reservation PASSED [Gate] broadcast PASSED [Gate] two-graphs PASSED [Gate] broadcast-two-graphs Evidence: gpu-battery-results/2026-08-01_369_bar1_graph_gate/gate_PASS_docker.log The image is the only place on this rig with BOTH halves: it reaches /dev/dmabuf_holder through --device, and it has a self-consistent toolchain (python 3.12.3, torch 2.11.0+cu130, nvcc 13.0, g++ 13.3 -- python and torch identical to the container venv). The working copy is mounted read-only, so the image needs no rebuild to test a tree. Recipe in docs/rig-runbook.md 4.15. A REAL DEFECT IN MY OWN sgl-project#369 FIX, FOUND BY RUNNING IT The first Docker run failed the same way the host runs did -- one rank per case, randomly which -- with the fix already in place. Not the toolchain: the grouped build parked the WAITING ranks in `bounded_barrier` under the 120 s steady-state peer deadline while rank 0 ran nvcc. The image bakes TORCH_CUDA_ARCH_LIST=7.5 8.0 8.6 8.9 9.0 10.0 12.0, so that build runs for many minutes, and every waiter raised CollectiveTimeoutError. A healthy cold build was indistinguishable from a wedged peer -- the exact confusion the bounded wait exists to remove. Fixed by opening `jit_cold_build.cold_build_window` around the build and the barrier, on EVERY rank: that is the mechanism the device-side deadlines already use, and it multiplies the peer timeout by SGLANG_JIT_COLD_BUILD_TIMEOUT_MULT (40, so 120 s -> 80 min) for the duration of the build and nothing else. Still bounded, and a peer that DIES still ends the wait immediately -- only a slow compiler stops looking like a dead rank. No new constant was invented. THE RELEASE parallel_state._GRAPH_ENABLE_ENV default "0" -> "1". bar1 and matrix are now in `capturable_transports()` unless somebody opts out with SGLANG_BARLINK_GRAPH_ENABLE=0. Reaching bar1 at all already requires a patched driver and the dma-buf holder, so this default cannot surprise a rig that has not deliberately set that up -- without them the transport declines and never becomes capturable. The refusal message in `_enforce_cpu_transport_needs_eager` was rewritten: it used to say the capture is UNMEASURED, which was true before today and would be a lie now. That branch is now only reachable through a deliberate opt-out and says so. Verified in the image against the real code, not only in tests: graph_enable_set() = True capturable = ['bar1', 'device', 'host', 'matrix'] bar1 + CUDA graphs: ACCEPTED (no ValueError) opt-out capturable = ['device', 'host'] sgl-project#366: THE BLOCKER IS GONE, THE NUMBERS ARE NOT IN YET With the release set, a sgl-project#366 boot (Qwen3.6-27B-FP8, TP=3, bar1, CUDA graphs ENABLED) got past the startup refusal that stopped sgl-project#366 and reported ACHIEVED=bar1 on all three groups on all three ranks -- world:0, tp:0, dcp:0, no group fell back. It loaded weights, allocated KV, and entered CUDA-graph capture. Capture had not finished after 18 minutes with all three GPUs pinned at 100% and no stack wedged (py-spy: the engine waiting in _wait_for_scheduler_ready, workers busy), so it was stopped because my card window ended -- NOT because anything failed. First-ever graph capture over this transport, cold graph cache, plus the NEXTN draft graphs. The bar1 column of tabelle_366 therefore stays empty for now. What the next window inherits is a proven boot recipe rather than a blocked one; logs in gpu-battery-results/2026-08-01_369_bar1_graph_gate/boot_bar1_graphs.log and boot_ACHIEVED_bar1.txt. TESTS test_barlink_grouped_jit_build.py grew from 7 to 13: - the barrier is reached INSIDE the cold-build window, on every rank (the defect above, pinned where it bit: in the waiters' deadline) - the window closes again afterwards - 40 x 120 s, pinned because both halves live in other modules - bar1/matrix capturable by default, the opt-out still removes them, and the host-staged transports are still refused (the release must not widen anything else) _bar1_marker_source.py: LINE_PARALLEL_STATE_GROUP_OK 722 -> 727 and LINE_PARALLEL_STATE_GROUP_FALLBACK 730 -> 735. Those constants pin the two ACHIEVED= emitters by line number; the release comment shifted them, and the collection error told me exactly that. test_gpu_battery_checks_bar1.py: 122 passed after repointing. CPU suites (CUDA_VISIBLE_DEVICES="", distributed + planner): base 6ff1ffc 23 failed, 3475 passed, 91 skipped this commit 23 failed, 3488 passed, 91 skipped (+13 new) Failure set identical to base. ruff clean on all four touched .py files. codespell (.codespellrc) clean on everything touched. Cards: BELEGT 23:56:11Z, FREI 00:38:31Z, renewed twice with a heartbeat. Every container ran with --rm; after teardown no compute apps and 0 MiB on all three cards, checked from the host and from CT999. No module reload, no driver action, no container-config change.
efschu
added a commit
to efschu/htsglang
that referenced
this pull request
Aug 1, 2026
… (grid = cooperative launch in capture accepted), SGLANG_BARLINK_GRAPH_ENABLE released, cold_build_window fix for the barrier deadline (sgl-project#369)
efschu
added a commit
to efschu/htsglang
that referenced
this pull request
Aug 1, 2026
…sgl-project#366) sgl-project#369 released the bar1 graph gate, so sgl-project#366 could finally measure. It still could not measure with CUDA graphs, and the reason is the result that matters most here. WHY EAGER. bar1 + CUDA graphs + NEXTN does not cold-capture in a usable time on the 27B FP8 vehicle. One graph boot ran >35 min without finishing capture, all three cards pinned at 100% util the whole time -- forward progress, not a wedge. The log says why, dozens of times: "Using default W8A8 Block FP8 kernel config ... Config file not found ... device_name=NVIDIA_GeForce_RTX_5090, dtype=fp8_w8a8, block_shape=[128,128]". The rig ships no tuned W8A8 block-fp8 configs for the 5090, so every GEMM shape falls into a cold triton autotune with ~3 min between shapes. That is the sgl-project#255/sgl-project#368 tuned-config gap, and warming those configs once would make every future bar1 graph boot fast -- a concrete, cheap follow-up. Running EAGER skips capture, costs ~34 s per boot, and puts both transports in the SAME regime, so the transport question is answered honestly. The table therefore does NOT quote sgl-project#354's graph-mode NCCL baseline; it measures NCCL itself, this window, eager. s17_bar1_eager_arm.sh boots one arm from the htsglang image on the PVE host and is parameterised by transport. Its gate is two-sided and mandatory per boot: a bar1 arm must log ACHIEVED=bar1 on every barlink group (world/tp/dcp x 3 ranks) and an NCCL arm must log ZERO barlink lines, so neither a silent fallback nor an accidental barlink run can enter the table. All 8 arms passed. Three environment defects were found and fixed inside it, each documented at the point of the fix: * TORCH_CUDA_ARCH_LIST in the image made the JIT extension build as _cuda_75_80_86_89_90_100_120, missing the warm _86_120 cache and cold-compiling seven arches. Pinned to '8.6;12.0' -> cache hit in 0.2 s. * The INT8 fork sgl-kernel wheel is built against CUDA 12.9 and the image is cu13, so installing it BROKE sgl_kernel outright (libcublas.so.12, then libcudart.so.12 missing -> common_ops unimportable -> every INT8 rank died with "needs sgl_kernel.int8_scaled_mm, which this build has no code for"). Fixed by mounting a directory of symlinks to every *.so.12* the CT999 venv ships and APPENDING it to the image's own LD_LIBRARY_PATH -- the SONAMEs differ from .so.13, so torch keeps its cu13 libs. * A single quote inside the wheel-inject string closed the docker -c '...' wrapper and killed INT8 boots in 2-3 s before the container existed. s17_bar1_eager_table.py renders the deliverable: both levers separately, never only the diagonal, with sgl-project#354's noise floors carried and any delta under its row floor printed as "within noise". Decode reads the client-side token rate for every arm: under eager + docker-logs buffering the server-log tick parser sees zero ticks, so one consistent source beats a mixed column. Test results (2026-08-01, Qwen3.6-27B TP=3 uneven 5090+2x3080, NEXTN 3/1/4, KV fp8_e4m3, ctx 32768, EAGER, 8 boots, all gates passed): point FP8 NCCL FP8 bar1 d(tr) INT8 NCCL INT8 bar1 d(tr) prefill_s1 1335.8 1395.8 +4.5% 1412.1 1687.5 +19.5% prefill_s8 1329.0 1454.6 +9.5% 1448.2 1674.5 +15.6% decode_bs1 25.4 36.0 +41.7% 35.4 33.8 -4.5% decode_bs8 217.8 219.7 within noise 230.2 260.3 +13.1% Transport lever: bar1 wins prefill in BOTH formats (FP8 +4.5/+9.5%, INT8 +19.5/+15.6%) and is mixed on decode. Format lever: INT8 > FP8 on prefill under both transports (NCCL +5.7/+9.0%, bar1 +20.9/+15.1%). Nothing under python/ was touched, so no repo test suite is affected. ruff, codespell and bash -n clean.
5 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.