Skip to content
Merged
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 0 additions & 7 deletions e2e_test/infra/model_specs.py
Original file line number Diff line number Diff line change
Expand Up @@ -187,13 +187,6 @@ def _resolve_model_path(hf_path: str) -> str:
"fa3",
"--max-model-len",

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Nit: The replacement coverage this removal leans on isn't in the tree yet, so between this PR and the follow-up nothing guards the case the pin was covering. e2e_test/router/test_pd_mmlu.py:88-90 still drives num_threads=32 (and test_loads.py:120 16 concurrent long requests) against this model on the e2e-2gpu-pd (tokenspeed) lane, and there is no topology test asserting the small-window/burst behaviour today — grep -rn "topology" e2e_test/ only hits test_epd_multimodal.py, whose topology cases are sequential.

That's fine if TokenSpeed's default sequence window is comfortably above 32, which the lanes on this PR will show. The part that won't be caught later is a silent change in that default: if it ever computes below the suite's burst (this spec is a hybrid GDN/MoE pool at --gpu-memory-utilization 0.8, where the derived running-request cap is memory-dependent), the failure mode is the 75-minute lane timing out on a bootstrap deadlock rather than a fast, legible failure. Landing #2464's topology test before or with this removal — or noting the observed default in the spec comment — would keep that diagnosable.

"8192",
# PD legs admit requests independently: a window smaller than the
# number of requests in flight lets prefill and decode admit
# disjoint subsets and wait on each other until the transfer
# timeout (run 34173426995 deadlocked at 4 under 32 concurrent).
# 32 covers every burst the PD suites drive.
"--max-num-seqs",
"32",
"--gpu-memory-utilization",
"0.8",

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Nit: Removing the pin leaves a stale cross-reference in the lane that motivated it. .github/workflows/pr-test-rust.yml:924-926 (the engine: tokenspeed entry in e2e-2gpu-pd) still reads:

# 4 (prefill waits for a decode that is never admitted). The window
# now covers the suite's bursts; the over-window deadlock itself is
# pinned by the topology suite.

After this change there is no window in the spec at all, so "the window now covers the suite's bursts" describes state that no longer exists — and it's exactly the comment the next person debugging a deadlock in this lane will read. Worth updating it in the same PR to say the spec now takes the engine default (and, per the description, that the burst/window case is covered elsewhere rather than by this pin).

# This model's hybrid-attention KV pool opts into tokenspeed's
Expand Down
Loading