Skip to content

support command-r - #369

Merged
merrymercy merged 1 commit into
sgl-project:mainfrom
ZX-ModelCloud:command_r_support
Apr 16, 2024
Merged

merrymercy merged 1 commit into
sgl-project:mainfrom
ZX-ModelCloud:command_r_support

Conversation

@ZX-ModelCloud

Copy link
Copy Markdown
Contributor

No description provided.

@Qubitium Qubitium left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We have tested the code with command-r-v01 on inference with batch size ranging from 1 to 10 with no issue.

@deoxykev

Copy link
Copy Markdown

Confirmed working with Command-R-Plus as well. (GPTQ quants)

good work!

@merrymercy
merrymercy merged commit db61106 into sgl-project:main Apr 16, 2024
@merrymercy

Copy link
Copy Markdown
Contributor

@ZhouXingg thanks for the contribution!

timethink pushed a commit to timethink/sglang that referenced this pull request Mar 9, 2025
UNIDY2002 pushed a commit to HanHan009527/sglang that referenced this pull request Jan 14, 2026
Co-authored-by: sunqi.7 <sunqi.7@bytedance.com>
heziiop pushed a commit to heziiop/sglang-community that referenced this pull request Apr 13, 2026
efschu added a commit to efschu/htsglang that referenced this pull request Jul 31, 2026
…n co-location race, dead builder fails loudly; graph gate stays unflipped until the proof can run (sgl-project#369)
efschu added a commit to efschu/htsglang that referenced this pull request Aug 1, 2026
… waiters now run under a build deadline (sgl-project#369)

ROUTE 3 WORKS: THE PROOF PASSED

`benchmark/bar1_graph_check.py` passed 10/10 on three cards (5090 + 2x 3080)
inside the htsglang Docker image on the Proxmox host -- all nine gate cases
plus the informational `grid` case, which IS the question the release was
waiting on: whether the driver accepts `cudaLaunchCooperativeKernel` from
inside a stream capture. It does.

  PASSED [Gate] 1blk-small        PASSED [Gate] pipe
  PASSED [Gate] 1blk-large        PASSED [Gate] pipe-direct
  PASSED [Info] grid              PASSED [Gate] pipe-direct-pool-empty
  PASSED [Gate] reservation       PASSED [Gate] broadcast
  PASSED [Gate] two-graphs        PASSED [Gate] broadcast-two-graphs

Evidence: gpu-battery-results/2026-08-01_369_bar1_graph_gate/gate_PASS_docker.log

The image is the only place on this rig with BOTH halves: it reaches
/dev/dmabuf_holder through --device, and it has a self-consistent toolchain
(python 3.12.3, torch 2.11.0+cu130, nvcc 13.0, g++ 13.3 -- python and torch
identical to the container venv). The working copy is mounted read-only, so
the image needs no rebuild to test a tree. Recipe in docs/rig-runbook.md 4.15.

A REAL DEFECT IN MY OWN sgl-project#369 FIX, FOUND BY RUNNING IT

The first Docker run failed the same way the host runs did -- one rank per
case, randomly which -- with the fix already in place. Not the toolchain: the
grouped build parked the WAITING ranks in `bounded_barrier` under the 120 s
steady-state peer deadline while rank 0 ran nvcc. The image bakes
TORCH_CUDA_ARCH_LIST=7.5 8.0 8.6 8.9 9.0 10.0 12.0, so that build runs for
many minutes, and every waiter raised CollectiveTimeoutError. A healthy cold
build was indistinguishable from a wedged peer -- the exact confusion the
bounded wait exists to remove.

Fixed by opening `jit_cold_build.cold_build_window` around the build and the
barrier, on EVERY rank: that is the mechanism the device-side deadlines
already use, and it multiplies the peer timeout by
SGLANG_JIT_COLD_BUILD_TIMEOUT_MULT (40, so 120 s -> 80 min) for the duration
of the build and nothing else. Still bounded, and a peer that DIES still ends
the wait immediately -- only a slow compiler stops looking like a dead rank.
No new constant was invented.

THE RELEASE

parallel_state._GRAPH_ENABLE_ENV default "0" -> "1". bar1 and matrix are now
in `capturable_transports()` unless somebody opts out with
SGLANG_BARLINK_GRAPH_ENABLE=0. Reaching bar1 at all already requires a patched
driver and the dma-buf holder, so this default cannot surprise a rig that has
not deliberately set that up -- without them the transport declines and never
becomes capturable.

The refusal message in `_enforce_cpu_transport_needs_eager` was rewritten: it
used to say the capture is UNMEASURED, which was true before today and would
be a lie now. That branch is now only reachable through a deliberate opt-out
and says so.

Verified in the image against the real code, not only in tests:
  graph_enable_set() = True
  capturable         = ['bar1', 'device', 'host', 'matrix']
  bar1 + CUDA graphs: ACCEPTED (no ValueError)
  opt-out capturable = ['device', 'host']

sgl-project#366: THE BLOCKER IS GONE, THE NUMBERS ARE NOT IN YET

With the release set, a sgl-project#366 boot (Qwen3.6-27B-FP8, TP=3, bar1, CUDA graphs
ENABLED) got past the startup refusal that stopped sgl-project#366 and reported
ACHIEVED=bar1 on all three groups on all three ranks -- world:0, tp:0, dcp:0,
no group fell back. It loaded weights, allocated KV, and entered CUDA-graph
capture. Capture had not finished after 18 minutes with all three GPUs pinned
at 100% and no stack wedged (py-spy: the engine waiting in
_wait_for_scheduler_ready, workers busy), so it was stopped because my card
window ended -- NOT because anything failed. First-ever graph capture over
this transport, cold graph cache, plus the NEXTN draft graphs.

The bar1 column of tabelle_366 therefore stays empty for now. What the next
window inherits is a proven boot recipe rather than a blocked one; logs in
gpu-battery-results/2026-08-01_369_bar1_graph_gate/boot_bar1_graphs.log and
boot_ACHIEVED_bar1.txt.

TESTS

test_barlink_grouped_jit_build.py grew from 7 to 13:
- the barrier is reached INSIDE the cold-build window, on every rank (the
  defect above, pinned where it bit: in the waiters' deadline)
- the window closes again afterwards
- 40 x 120 s, pinned because both halves live in other modules
- bar1/matrix capturable by default, the opt-out still removes them, and the
  host-staged transports are still refused (the release must not widen
  anything else)

_bar1_marker_source.py: LINE_PARALLEL_STATE_GROUP_OK 722 -> 727 and
LINE_PARALLEL_STATE_GROUP_FALLBACK 730 -> 735. Those constants pin the two
ACHIEVED= emitters by line number; the release comment shifted them, and the
collection error told me exactly that. test_gpu_battery_checks_bar1.py: 122
passed after repointing.

CPU suites (CUDA_VISIBLE_DEVICES="", distributed + planner):
  base 6ff1ffc  23 failed, 3475 passed, 91 skipped
  this commit      23 failed, 3488 passed, 91 skipped   (+13 new)
Failure set identical to base.

ruff clean on all four touched .py files. codespell (.codespellrc) clean on
everything touched.

Cards: BELEGT 23:56:11Z, FREI 00:38:31Z, renewed twice with a heartbeat.
Every container ran with --rm; after teardown no compute apps and 0 MiB on
all three cards, checked from the host and from CT999. No module reload, no
driver action, no container-config change.
efschu added a commit to efschu/htsglang that referenced this pull request Aug 1, 2026
… (grid = cooperative launch in capture accepted), SGLANG_BARLINK_GRAPH_ENABLE released, cold_build_window fix for the barrier deadline (sgl-project#369)
efschu added a commit to efschu/htsglang that referenced this pull request Aug 1, 2026
…sgl-project#366)

sgl-project#369 released the bar1 graph gate, so sgl-project#366 could finally measure. It still
could not measure with CUDA graphs, and the reason is the result that matters
most here.

WHY EAGER. bar1 + CUDA graphs + NEXTN does not cold-capture in a usable time on
the 27B FP8 vehicle. One graph boot ran >35 min without finishing capture, all
three cards pinned at 100% util the whole time -- forward progress, not a
wedge. The log says why, dozens of times: "Using default W8A8 Block FP8 kernel
config ... Config file not found ... device_name=NVIDIA_GeForce_RTX_5090,
dtype=fp8_w8a8, block_shape=[128,128]". The rig ships no tuned W8A8 block-fp8
configs for the 5090, so every GEMM shape falls into a cold triton autotune
with ~3 min between shapes. That is the sgl-project#255/sgl-project#368 tuned-config gap, and warming
those configs once would make every future bar1 graph boot fast -- a concrete,
cheap follow-up. Running EAGER skips capture, costs ~34 s per boot, and puts
both transports in the SAME regime, so the transport question is answered
honestly. The table therefore does NOT quote sgl-project#354's graph-mode NCCL baseline;
it measures NCCL itself, this window, eager.

s17_bar1_eager_arm.sh boots one arm from the htsglang image on the PVE host and
is parameterised by transport. Its gate is two-sided and mandatory per boot:
a bar1 arm must log ACHIEVED=bar1 on every barlink group (world/tp/dcp x 3
ranks) and an NCCL arm must log ZERO barlink lines, so neither a silent
fallback nor an accidental barlink run can enter the table. All 8 arms passed.

Three environment defects were found and fixed inside it, each documented at
the point of the fix:
  * TORCH_CUDA_ARCH_LIST in the image made the JIT extension build as
    _cuda_75_80_86_89_90_100_120, missing the warm _86_120 cache and
    cold-compiling seven arches. Pinned to '8.6;12.0' -> cache hit in 0.2 s.
  * The INT8 fork sgl-kernel wheel is built against CUDA 12.9 and the image is
    cu13, so installing it BROKE sgl_kernel outright (libcublas.so.12, then
    libcudart.so.12 missing -> common_ops unimportable -> every INT8 rank died
    with "needs sgl_kernel.int8_scaled_mm, which this build has no code for").
    Fixed by mounting a directory of symlinks to every *.so.12* the CT999 venv
    ships and APPENDING it to the image's own LD_LIBRARY_PATH -- the SONAMEs
    differ from .so.13, so torch keeps its cu13 libs.
  * A single quote inside the wheel-inject string closed the docker -c '...'
    wrapper and killed INT8 boots in 2-3 s before the container existed.

s17_bar1_eager_table.py renders the deliverable: both levers separately, never
only the diagonal, with sgl-project#354's noise floors carried and any delta under its row
floor printed as "within noise". Decode reads the client-side token rate for
every arm: under eager + docker-logs buffering the server-log tick parser sees
zero ticks, so one consistent source beats a mixed column.

Test results (2026-08-01, Qwen3.6-27B TP=3 uneven 5090+2x3080, NEXTN 3/1/4,
KV fp8_e4m3, ctx 32768, EAGER, 8 boots, all gates passed):

  point         FP8 NCCL  FP8 bar1  d(tr)      INT8 NCCL  INT8 bar1  d(tr)
  prefill_s1      1335.8    1395.8  +4.5%         1412.1     1687.5  +19.5%
  prefill_s8      1329.0    1454.6  +9.5%         1448.2     1674.5  +15.6%
  decode_bs1        25.4      36.0  +41.7%          35.4       33.8  -4.5%
  decode_bs8       217.8     219.7  within noise   230.2      260.3  +13.1%

  Transport lever: bar1 wins prefill in BOTH formats (FP8 +4.5/+9.5%,
  INT8 +19.5/+15.6%) and is mixed on decode.
  Format lever: INT8 > FP8 on prefill under both transports
  (NCCL +5.7/+9.0%, bar1 +20.9/+15.1%).

Nothing under python/ was touched, so no repo test suite is affected.
ruff, codespell and bash -n clean.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants