Repository navigation
ci(e2e): stop pinning TokenSpeed's sequence window in the Qwen3.5-9B spec - #2467
Conversation
…spec The spec carried --max-num-seqs since the EPD smoke lane was added, first at 4 and then at 32 after a PD burst deadlocked against the smaller window. Neither value was a property of the model or of PD: a decode worker's window bounds how many requests it can admit, and a pinned small window is exactly what turns a burst into bootstrap timeouts. Let the engine default apply; the PD lanes and the EPD lane on this change show whether anything relied on the pin. Signed-off-by: Alex McC <319643551+hello-alexmcc@users.noreply.github.com>
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Advanced Run ID: 📒 Files selected for processing (1)
💤 Files with no reviewable changes (1)
Included review availability: Your plan provides up to 4 included reviews per hour; 2 remain after this review. 📝 SummarySummary by CodeRabbit
WalkthroughThe Qwen/Qwen3.5-9B TokenSpeed configuration removes the explicit ChangesTokenSpeed configuration
Priority: ⬇️ Low — Impact reflects low issue severity. Estimated code review effort: 1 (Trivial) | ~2 minutes Severity of issue fixed: Low Merge Risk: ⚪ Minimal · up to The TokenSpeed specification now uses the engine default sequence window, with no concrete merge-blocking risk identified. 🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Comment |
| "--max-num-seqs", | ||
| "32", | ||
| "--gpu-memory-utilization", | ||
| "0.8", |
There was a problem hiding this comment.
🟡 Nit: Removing the pin leaves a stale cross-reference in the lane that motivated it. .github/workflows/pr-test-rust.yml:924-926 (the engine: tokenspeed entry in e2e-2gpu-pd) still reads:
# 4 (prefill waits for a decode that is never admitted). The window
# now covers the suite's bursts; the over-window deadlock itself is
# pinned by the topology suite.
After this change there is no window in the spec at all, so "the window now covers the suite's bursts" describes state that no longer exists — and it's exactly the comment the next person debugging a deadlock in this lane will read. Worth updating it in the same PR to say the spec now takes the engine default (and, per the description, that the burst/window case is covered elsewhere rather than by this pin).
| @@ -187,13 +187,6 @@ def _resolve_model_path(hf_path: str) -> str: | |||
| "fa3", | |||
| "--max-model-len", | |||
There was a problem hiding this comment.
🟡 Nit: The replacement coverage this removal leans on isn't in the tree yet, so between this PR and the follow-up nothing guards the case the pin was covering. e2e_test/router/test_pd_mmlu.py:88-90 still drives num_threads=32 (and test_loads.py:120 16 concurrent long requests) against this model on the e2e-2gpu-pd (tokenspeed) lane, and there is no topology test asserting the small-window/burst behaviour today — grep -rn "topology" e2e_test/ only hits test_epd_multimodal.py, whose topology cases are sequential.
That's fine if TokenSpeed's default sequence window is comfortably above 32, which the lanes on this PR will show. The part that won't be caught later is a silent change in that default: if it ever computes below the suite's burst (this spec is a hybrid GDN/MoE pool at --gpu-memory-utilization 0.8, where the derived running-request cap is memory-dependent), the failure mode is the 75-minute lane timing out on a bootstrap deadlock rather than a fast, legible failure. Landing #2464's topology test before or with this removal — or noting the observed default in the spec comment — would keep that diagnosable.
CI: the engine default holds — no scoped bound neededRun 34195823976 is green end to end ( The concern that the EPD lane needed a small window just to get the engine up does not hold. Every TokenSpeed lane came up and passed with
EPD selected 8 of 8 and reported Job totals move with image-pull time, so the test steps are the cleaner comparison: EPD's Why nothing needed a boundAt the pinned ref (
If decode-side capture ever does need bounding, |
Description
Problem
e2e_test/infra/model_specs.pypins--max-num-seqsfor Qwen3.5-9B on TokenSpeed. The pin arrived with the EPD multimodal smoke lane (#1924) at 4, without a stated reason, and #2463 raised it to 32 after the first TokenSpeed PD run deadlocked: with a 4-wide decode window and 32 concurrent requests, the prefill's bootstrap deadline expired on requests that were merely queued behind the decode's admission. A small pinned window is not a property of the model, and it is exactly the configuration that turns a burst into bootstrap timeouts in PD. Nobody runs the engine that way outside this spec.Solution
Remove the pin and let the engine's default window apply, as it does for every other model in the suite. The
e2e-2gpu-pd (tokenspeed),e2e-4gpu-epd (tokenspeed)and, once #2464 lands,e2e-4gpu-pd (tokenspeed)lanes on this PR are the check that nothing relied on it. If the EPD lane turns out to need a bound for the encode worker, it should be scoped to that role inworker.py, not applied to every TokenSpeed worker of the model.The burst-versus-window behaviour itself gets its own coverage: a topology-suite test that sets a deliberately small window and asserts the gateway's admission sheds instead of the engine timing out (follow-up on #2464), alongside the gateway admission fix in #2466 and the engine fix in TokenSpeed.
Changes
e2e_test/infra/model_specs.py: drop--max-num-seqsand its comment from the Qwen3.5-9B TokenSpeed args.Test Plan
ruff check/ruff format --checkclean.e2e-2gpu-pd (tokenspeed)(11 cases, MMLU floor marked expected-fail),e2e-4gpu-epd (tokenspeed),e2e-1gpu-chat (tokenspeed).Checklist
cargo +nightly fmtpasses (no Rust changes)cargo clippy --all-targets --all-features -- -D warningspasses (no Rust changes)