Skip to content

Add September 27 benchmark grid and K-pool OOM patch - #2

Merged
stewtong merged 1 commit into
mainfrom
glm53-holistic-20260927
Sep 27, 2026
Merged

stewtong merged 1 commit into
mainfrom
glm53-holistic-20260927

Conversation

@stewtong

Copy link
Copy Markdown
Owner

Summary

This adds a same-day benchmark campaign run on September 27, 2026 on one 8x B200 node. Every result in it was measured with two local SGLang source changes on top of the pinned image, both now published under reproduce/patches/.

  • A full context and concurrency grid for the reference configuration: 16K to 1M input tokens at concurrency 1 to 16, three repetitions per cell, 744 of 744 requests successful. Each repetition is one burst of requests, and these records keep a dispatch and completion clock for every request, so a wall-clock burst output rate is reported with per-repetition ranges.
  • Same-day comparisons: one TP8/EP8 engine against the two TP4/EP4 servers (run after the grid), BF16 against FP8 KV cache on the full GSM8K test set and a 48-request cap against 16 (each run side by side on separate GPU halves), and direct engine access against the loopback proxy.
  • reproduce/patches/kpool-chunked-logits.diff: on the pinned image, two 1M-token prompts in one prefill step made the K-pool indexer request a 29.68 GiB FP32 buffer, which raised a CUDA out-of-memory error in the scheduler while the HTTP server kept answering. Upstream fixed the same defect in [DSA] Chunk the kpool indexer MQA logits by query rows under a free-memory budget sgl-project/sglang#40854 (merged September 26, after the pinned image); this diff is an independent equivalent. No image containing #40854 was tested here. An opt-in comparison mode used to check it is split into reproduce/patches/kpool-logits-verify.diff; the servers carried it with the mode off. The README gives tested steps to rebuild and mount the files.
  • reproduce/patches/parser-overlay.diff: the September 6 tool-call parser overlay the servers also carried, a port of [Fix] Resolve tool argument types through top-level anyOf/oneOf/allOf sgl-project/sglang#36626 (merged) and #24147 (unmerged).
  • A replica-loss observation: when one server was stopped mid-stream, its streams stopped, kept HTTP 200, and closed at SIGKILL with no finish reason or usage, so clients need to check finish_reason.
  • Updated claim ledger, benchmark summary, manifests, and checksums. The TP8 comparison and BF16 control move from not run to completed, and open-loop arrivals to partial.

Verification

  • reproduce/derive-results.py recomputes every new cell's burst output rate, request output rate, TTFT, and repetition range, the warm-extension values, and the GSM8K statistics from the sanitized rows in results/holistic-20260927/raw/, and the existing summaries still re-derive exactly. The claim ledger marks the other new values as retained summaries or first-party records.
  • The derivation self-test passes, including new fail-closed fixtures for the grid, warm-extension, and GSM8K rows and whole-campaign coverage checks. It also fixes an existing self-test defect: the failure counter was reset before the final loop, so a failing success-rule fixture could not fail the run.
  • 12 Python tests and the reproduction script tests pass. All result checksums pass.
  • Applying the three diffs to the pinned image's files, with the README's steps, reproduces the deployed SHA-256 values (58e59be0…, 28d5f11e…, 8f04cda3…); its sha256sum -c check fails on a modified file.
  • A pattern scan of the diff found no hostnames, IP addresses, local paths, or credentials.

Measurement limits

  • No result measures the pinned image alone, and patch overhead was not measured against an unpatched run. run-replicas.sh does not apply the patches; the README gives the mounts.
  • The TP8 comparison is sequential (3.9 to 5.4 hours after the grid), used different prompts, kept the TP4-selected prefill settings, and had a KV pool 21.5% smaller than the two servers combined. DP attention was not tested.
  • The request-cap arms received different prompts, split each node load across two servers, and the cap-48 server was slower even where no cap binds, so the rows compare complete server configurations.
  • The burst output rate is one burst per repetition, not sustained throughput. The Poisson ladder reached its stop threshold at 0.8 requests/s with windows of 236-292 s instead of 300 s.
  • The prefix-cache study and the KV admission boundary are open, and three Anthropic Messages cases (one invalid, two failed) are open. BENCHMARKS.md lists each with its cause.

Full 16K-1M by c1-c16 grid (744/744), sequential TP8/EP8 comparison, side-by-side BF16-KV GSM8K control and request-cap arm, direct versus loopback proxy, warm extension, and a bounded Poisson ladder, measured with a local K-pool chunked-logits patch (independent equivalent of sglang#40854) and the September 6 parser overlay, both published under reproduce/patches/. Sanitized request rows re-derive the cell metrics, warm-extension values, and GSM8K statistics; fixes a derive-results self-test that reset its failure count.
@stewtong
stewtong merged commit a8886a7 into main Sep 27, 2026
@stewtong
stewtong deleted the glm53-holistic-20260927 branch September 27, 2026 21:28
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant