Repository navigation
[Bugfix][ROCm] Select page-aligned kernel blocks for pooled indexers so the block table addresses storage pages (GLM-5.3-Flash) - #59412
Conversation
…so the block table addresses storage pages (GLM-5.3-Flash) Signed-off-by: vllmellm <vllm.ellm@embeddedllm.com>
b55ebb1 to
b53f00f
Compare
|
Independent confirmation from MI355X (gfx950), GLM-5.3-Flash, TP=2. I hit the same bug and opened #59704 with a different fix: it expands the 1152-token kernel-block table into 128-token storage pages inside Note: these were measured with #59704's fix, not with this branch. They show the impact of the bug, not a validation of this PR's implementation. GPQA-Diamond (198 questions):
Decode-vs-prefill consistency: for each sequence, generate greedily, then re-score the same tokens with one prefill. The table shows the mean |logprob(decode) − logprob(prefill)| over the first 64 generated tokens, and the share of positions where the argmax differs.
GPQA wall time also dropped from ~3000–3700 s to ~2000–2400 s, because generations stopped running to the token limit. I'm happy to rerun the same GPQA-Diamond and decode-vs-prefill checks on MI355X against this branch if that would help. |
|
Independent confirmation + end-to-end validation on gfx1151 (2-box TP=2) We hit the identical failure on the Sep-28 pin ( Measured with score-level instrumentation in the gather/top-k path:
So: an independent gfx1151 datapoint that (a) the bug reproduces on consumer APUs with the 256-token page, (b) page-granular addressing fixes it fully at 256K context. Notes for reviewers:
|
|
Validation on 8x MI350X (gfx950), GLM-5.3-Flash, at TP4 and TP8, to go with the TP2 results above. I applied this PR's head (
The round trip recomputes each pool's compressed key from the writer's inputs and compares it byte for byte with what the prefill gather reads back. Needle retrieval places a passphrase at 10%, 50% and 90% of documents of 9.5K, 14K, 29K, 57K and 105K tokens, greedy, with a unique first line per prompt so nothing is served from the prefix cache. Unpatched, TP4 already misses the needle at 9.5K tokens, and TP8 starts missing it at 29K. With this PR, reasoning, Both TP8 columns ran with MTP off and |
simondanielsson
left a comment
There was a problem hiding this comment.
Thanks for the fix! I'm thinking if we can simplify more.
- Thought: Is the underlying reason for the bug that throughout the code we generally say that whenever
storage_block_sizeis set, we use that as the kernel block size? (e.g. in create_metadata_builders and hisparse etc). However, inprepare_kernel_block_sizeswe ignore this. So could the fix just be to always prio storage_block_size if available?
# inside prepare_kernel_block_sizes
kv_manager_block_size = kv_cache_group.kv_cache_spec.block_size
group_backends = [g.backend for g in attn_groups[kv_cache_gid]]
# new
spec = kv_cache_group.kv_cache_spec
storage_block_size = spec.storage_block_size isinstance(spec, MLAAttentionSpec) else None
if storage_block_size is not None:
# add some validation that hte backends all support this kernel block size
selected_kernel_size = storage_block_size
else:
# old
selected_kernel_size = select_common_block_size(
kv_manager_block_size, group_backends
)
kernel_block_sizes.append(selected_kernel_size)- Suggestion: Can we rework the tests a bit to find one or two that capture the bug we're solving here? For instance it'd be great to have just one checking that prepare_kernel_block_sizes does indeed returns the right kernel block sizes when using kpool (i.e. on rocm, not 640 but rather 128)
Let me know what you think :)
|
Hi @simondanielsson, Thanks for your feedback, yeah we can totally simplify it. I started with the backend-declaration route first because I thought letting the backends declare the page lattice was the most general way to get everyone to agree on the same kernel block, but you're right that And yes, I can simplify the tests along with it too. I've updated accordingly, feel free to take a look. |
…prepare_kernel_block_sizes (GLM-5.3-Flash) Use the spec's storage_block_size as the kernel block when the group's backends accept it, instead of extending the backend declarations, per review. Signed-off-by: vllmellm <vllm.ellm@embeddedllm.com>
simondanielsson
left a comment
There was a problem hiding this comment.
Thanks, LGTM! Would be good to have someone also with more experience in this part of the code to have a look
…docstring Signed-off-by: vllmellm <vllm.ellm@embeddedllm.com>
|
/ci run |
|
✅ Triggered Buildkite CI #93258 for commit |
CI selector (shadow): 136 test steps (178 jobs) instead of 66 (84 jobs)Shadow mode: this changes nothing about what CI runs. It shows what the evidence-based selector would pick for this PR, next to today's rules. How it works. Feedback welcome: reply here if it would skip a step this change needs, or runs something unrelated.
Selector would run (136)
Would skip (today's rules run them) (14)
Would add (today's rules do not run them) (84)
AMD mirrors: would skip (9)
AMD mirrors: would add (72)
2 changed files · base |
|
/ci retry |
|
✅ Queued 2 failed job(s) for retry in Buildkite CI #93258. |
|
Follow-up to my earlier comment: I've now rerun GPQA-Diamond on MI355X with this PR's head ( Setup: GLM-5.3-Flash, TP=4, MI355X, upstream main
So this PR's fix restores the same accuracy and removes the runaway generations, matching the alternative fix. Caveat: the unfixed and #59704 rows are on an older main ( |
decision-balance 0.3.1 moves the boundary between medium and high effort from 0.8 to 0.675. Repeated stratified cross-validation on the per-effort benchmark samples chose it together with the hard-STEM difficulty of 2 and the code lane in most folds; held-out accuracy rose by about 1.3 points over 0.8 for about 4% more GPU time per request. Probes near the new boundary are replaced by ones with clear margins, and the plan-without-tools negative moves from standard to hard. The Model Card lists the serving notes behind the measurements on AMD GPUs with vLLM 0.31: TRITON_ATTN for the 27B, whose default backend's decode cost grows with input length, and the GLM-5.3-Flash indexer fix of vllm-project/vllm#59412. Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
|
/ci run |
|
✅ Triggered Buildkite CI #93469 for commit |
|
/ci run |
|
✅ Triggered Buildkite CI #93595 for commit |
…ools (#4734) * [Feature] Router: let algorithm.decision ask the decision model An algorithm.decision selector had to name an explicit model_runtime deployment, so the Vela 2.0 decision model that already answers a request's signals could not also choose its model without a second copy of the same weights. A selector that names no deployment now asks the decision model's shared implicit deployment, as a routing.signals.decision question does: the Router resolves it with RouterConfig.DecisionSelectorDeployment, the model runtime serves it even when no signal uses it, and the selection reasoning names the deployment that chose. An explicit deployment keeps working; with decision_model: Vela-1.0 an omitted deployment is a load error that asks for one, in the Router and in vllm-sr config validate. A blank or padded deployment is rejected with the same hint. Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * [Feature] Router: let projections read one option of a decision question The Router publishes every option of a choice, set or span decision question as its own signal value (decision:<question>:<key>), but a projection score input could name only the question. A set question's bare value is its most probable label, so a score such as a reasoning effort could not weigh one label, like needs:deliberation, on its own. A projection input of type decision now takes <question>:<key> for one option of a choice, set or span question; validation rejects a key the question does not declare and an option of a noul or score question. Signal usage asks the question whenever a used projection reads one of its options, and the DSL validator accepts the same form. Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * [Bug] DSL: validate decision-option routes and catalog-backed models cleanly Three DSL findings from authoring a recipe that routes on Vela 2.0 decision questions over catalog-backed models; a maintained recipe's DSL must validate without diagnostics and round-trip byte for byte. - The route guard compared references without their labels, so routes on different options of one decision question or classifier (task:agentic and task:facts) were reported as overlapping. A labelled reference now names its label; the same option in two routes still warns. - A route that answers with fast_response, inline or through a template, calls no model, so it no longer warns that it has no MODEL. - The decompiler copied a catalog-backed model's parameter size from its built-in card into every route, which compiled back as operator metadata. Routes now repeat only an operator's own param_size. Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * [Bug] Router: prepare set and span questions that ask the decision model A routing.signals.decision set or span question that names no deployment asks the Router's decision model, but preparation checked the model card of the empty deployment name. The Router failed to start with "provider \"\" is not served by the model runtime" for any such question, so set and span questions only worked with an explicit deployment. Preparation now checks the card of the deployment the question asks: its own, or the decision model's shared deployment. Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * [Feature] Router: preview the decision selector's choice Routing Preview reported execution_required for every algorithm.decision route, so a routing-only evaluation could not show which model the decision model would choose. The decision selector keeps no state between requests: it asks the decision model one Choice question about the request. Preview now dry-runs it like static, multi_factor and latency_aware and reports the chosen model as selected; a failed or late answer falls back to the first modelRef as at request time. Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * [Bug] Recipes: plan a GPU-only decision model as a hardware requirement Live CPU conformance derived hardware requirements only from explicit model_runtime deployment devices. A recipe whose global.model_catalog.system.decision_model is Vela-2.0-4B or Vela-2.0-9B has no explicit deployment, so the planner scheduled it on CPU, where vllm-sr serve refuses a GPU-only decision model. Such a recipe now lists the gpu requirement and stays out of the CPU matrix like any other hardware-bound recipe. Probe checks also accept the decision algorithm's Preview statuses, selected or execution_required, like the other base selectors. Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * [Feature] Recipes: add the decision-balance Mixture-of-Models recipe Decision Balance serves vllm-sr/auto over GLM-5.3-Flash, Qwen3.8-Flash-Next and Qwen3.8-27B. Vela-2.0-4B answers five System One questions (task choice, difficulty score, precise-facts noul, needs set, correction noul) in the call that also answers the prompt guard and safety signals; heuristics cover images, tools, tool loops, earlier answers, input length and brief-answer requests. An effort projection over the difficulty score and the deliberation and verification labels sets the reasoning effort. Ten decisions route guard (fast_response), long_context, vision, recovery, agentic and facts to fixed models, frontier through algorithm: decision with the decision model itself, hard and standard through multi_factor, and the rest to Qwen3.8-27B with thinking off. Costs are relative GPU-seconds per token measured on MI325X; quality evidence reuses the catalog's third-party records and labels the operator ratings it adds for an effort without them. The CPU conformance plan excludes it as GPU-bound, like vela-amd. Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * [Feature] Router: keep personal data out of replay with a recipe data policy Router Replay could be limited per decision or forbidden per recipe, but not by what a request contains. A recipe that keeps replay for its routing evidence therefore also stored the prompts and answers of requests with names, emails or phone numbers. routing.data_policy.replay_personal_data: false keeps the replay record of a request in which one of the recipe's PII signals matched, with its route, model, signals and detected PII types, but without the request or response body, prompt, tool definitions or tool trace. The recipe's PII signals are then evaluated for every request, even when no decision references them, and a PII classification that fails counts as personal data. The policy needs a routing.signals.pii rule; the Router and vllm-sr config validate reject it without one. The DSL round-trips it and the reference config sets it. Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * [Feature] Dashboard: show the System One answers and routing latency in Playground A Playground reply showed the decision, algorithm and model but not the decision-model answers that chose them, nor how long routing took: the Router sends x-vsr-matched-decision-model, x-vsr-selected-recipe, x-vsr-selected-confidence and x-vsr-routing-latency-ms, but Playground did not collect or label them. It now shows them as System One Answers, Recipe, Decision Confidence and Routing Latency, the latency beside the other timings. Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * [Feature] Recipes: keep personal data out of decision-balance replay Decision Balance 0.2.0 asks the decision model's PII span head in the same call and sets routing.data_policy.replay_personal_data: false, so Router Replay keeps the routing evidence of a request with personal data but none of its content. The prompt guard threshold rises from 0.9 to 0.95: on public chat traffic, role-play and code requests scored between 0.9 and 0.95 while the attacks scored higher. A probe asserts the PII span on a request that names a person, an email address and a phone number. Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * [Test] Dashboard: re-pin the mom-v1 fixture receipt after its probe rewrites #4725 replaced nine built-in mom-v1 probe examples, which changes the materialized message text by 22 bytes, but left the Dashboard's receipt test pinned to the old text. The test fails on main; the new byte count and digest are those of the probes #4725 ships. Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * [Feature] Recipes: send code to Flash-Next and hard STEM to extra-high effort decision-balance 0.3.0 adds two lane rules, each from per-effort measurements on public benchmark samples: - code: a code task at medium or high effort goes to Qwen3.8-Flash-Next at medium effort. On LiveCodeBench it solved more problems at medium effort than either Qwen model at extra-high, with less than half the tokens. - hard: a STEM task the decision model rates at least multi-step (difficulty >= 2) runs at extra-high effort even when its effort score is lower. Asking for only the final answer lowers the deliberation answer, not the reasoning a GPQA-style problem needs; medium effort lost 8 to 15 points there. Probes move the two code examples from standard to the new code group and add collision variants for a letter-only chemistry question (hard), routine algebra (standard) and a one-line code fix (fast). The Model Card states the rules, their evidence and that long_output and creativity are reported but not routed on. Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * [Feature] Router: restrict a listener to named models A listener's api_keys decide who may call it, but any key holder could name a provider model and have it passed through, bypassing the recipes. listeners[].models lists the only request models a standalone listener accepts, by exact `model` value; empty or absent accepts every model. The gateway hands the allow-list to the routing core after the API key check, and the request-body phase checks it against the model request decoding parsed, the one routing uses, before any signal, cache or decision runs. Another model gets 403 model_not_allowed in the client's protocol, and /v1/models on the listener lists only the allowed names. Request-graph hops the Router makes itself are not restricted, so an auto model still reaches the provider models its decisions name, and a restricted listener ignores the skip-processing opt-out, which would bypass the check. --gateway extproc rejects a listener with models as unsupported (the Router's capability check and the CLI's Envoy generation), since the Envoy listener does not enforce it yet. The schema, the CLI model, the reference config, the Dashboard type and the docs follow. Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * [Bug] CLI: start vllm-sr serve on Docker daemons without a default bridge `vllm-sr serve` always passed `--add-host=host.docker.internal:host-gateway`. Docker derives `host-gateway` from its default bridge network, so a daemon configured with `"bridge": "none"` rejected every service container with `unable to derive the IP value for host-gateway`. - `VLLM_SR_HOST_GATEWAY_IP` maps `host.docker.internal` to an explicit address on Docker and Podman. - Without an override, Docker daemons that report no default bridge network skip the mapping with a warning that names the override; any other probe result keeps the previous `host-gateway` mapping. Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * [Bug] CLI: accept PII signal rules without a threshold `vllm-sr config validate` and `vllm-sr serve` rejected a `routing.signals.pii[]` rule without `threshold` ("Field required"), while the Router accepts one and lets the rule take every span the PII model reports. A Vela 2.0 model reports only spans above its size's calibrated threshold, so a recipe for one size had to hard-code another size's operating point to pass the CLI. The CLI schema now makes the rule threshold optional, like the Router, and the PII signal tutorial says what an omitted threshold means. Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * [Bug] Router: stop RouterDC warnings when no decision uses it The selection factory initializes every algorithm at startup, so RouterDC embedded each model's description even when no decision routes with it. A generation prepares the description embedding model only for decisions that use router_dc or hybrid, so every start without one logged a warning per model: `embedding model "mmbert" was not prepared for this generation`. The embedding set now reports that case as ErrModelNotPrepared (same message, still a capability error), and RouterDC stops at the first such error with a debug line. Other embedding failures keep their warning. Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * [Feature] Recipes: let decision-balance's PII rule use the model's own threshold The CLI now accepts a PII rule without a threshold, so the recipe no longer hard-codes 0.05. Vela-2.0-4B reports only spans above its calibrated, length-aware span threshold, which is what the replay data policy should act on. Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * [Feature] Recipes: raise decision-balance's effort to high from 0.675 decision-balance 0.3.1 moves the boundary between medium and high effort from 0.8 to 0.675. Repeated stratified cross-validation on the per-effort benchmark samples chose it together with the hard-STEM difficulty of 2 and the code lane in most folds; held-out accuracy rose by about 1.3 points over 0.8 for about 4% more GPU time per request. Probes near the new boundary are replaced by ones with clear margins, and the plan-without-tools negative moves from standard to hard. The Model Card lists the serving notes behind the measurements on AMD GPUs with vLLM 0.31: TRITON_ATTN for the 27B, whose default backend's decode cost grows with input length, and the GLM-5.3-Flash indexer fix of vllm-project/vllm#59412. Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * [Bug] Router: keep short-request usage out of context token calibration The context signal's counter learns bytes per token from provider prompt usage and tracks the lower tail. Provider usage also counts the chat template, role and default system tokens, so a one-line prompt reports roughly one token per content byte. After ordinary chat traffic the learned ratio fell to about 1.2-1.4 bytes/token and long English prose was over-estimated 3-5x: 204,643 bytes reported 148,755 tokens against 41,827 counted by the backend, and a 200K-token context band matched requests of about 55K real tokens. Candidate context-window checks read the same count. Ignore samples with less than 4 KiB of text, where template overhead dominates; longer samples still calibrate the tokenizer ratio and an uncalibrated deployment keeps the 4 bytes/token default. Tests: calibrated counter unit tests (short samples ignored, long prose estimate near the real count) and an extproc test that replays chat usage through the response calibration path and asserts a ~55K-token prompt does not match a 200K context rule (it matched at 224K before). Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * [Bug] CLI: never leave the stack down over the sr-bench worker serve stops Router and Dashboard, then starts sr-bench, and reconciled an existing worker only when its image alone changed. Any other difference in the launch identity refused the worker after the stack was already stopped. In production that difference was the host alias: the previous CLI launched the worker with host-gateway (a wrapper rewrote it inside Docker), the new one with an explicit VLLM_SR_HOST_GATEWAY_IP, so serve exited 1 with the stack down. Replace an owned worker (sr-bench identity label, store read from its own arguments) whenever its launch identity differs, after the same paused journal check. When replacement is not safe (active runs or preparations, unreadable journal, stopped worker, or a container without the label) keep the worker as it is, log a warning, and start the rest of the stack. Reuse still requires an identical identity. Tests: replacement for image, credential, host-alias and port changes; reuse when unchanged; active-run, unlabelled and stopped workers are kept while Dashboard still starts and nothing is rolled back; existing journal, pause and rollback tests unchanged. Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * [Bug] Recipes: count only identifying types as decision-balance personal data The personal_data PII rule accepted every span type, so a place, an organization or a date matched: "What is the capital of France?" (GPE), "Plan a trip to Kyoto" (GPE) and "How do Apple and Microsoft compete" (ORGANIZATION) all counted as personal data, and replay_personal_data: false dropped their content from Replay. Allow GPE, ORGANIZATION, DATE_TIME, NRP, TITLE and DOMAIN_NAME on their own, as the privacy recipes do; names, contact details, addresses, identity and account numbers still match. Probes: the common-knowledge group (France, Germany) now forbids personal_data; the Dana Lee probe still expects it. Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * [Feature] Recipes: release decision-balance 0.3.2 with production fixes Bug-fix release: lanes, thresholds and models are unchanged. - personal_data counts only identifying PII types. - With the Router fix in this series, long_context matches near 200K backend tokens instead of about 55K after ordinary chat traffic. - serve keeps the stack up when the sr-bench worker cannot be replaced. Model Card: the decision call's cost with and without the PII question (one line 57/89 ms, 2K tokens 0.11/0.39 s, 16K tokens 0.53/1.75 s on one MI325X), when the rule can be removed (Replay off), the context estimate's remaining overcount, and a Changes section. Live conformance on a GPU stack: 45/45 probes. Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * [Docs] Fix translated safety signal scan-budget links Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * [Fix] Dashboard: surface the configured decision model and shared runtime Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * [Feat] Dashboard: manage decision models and identify serving modes Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * [Fix] Dashboard: keep the decision model deploy action readable Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * [Feat] Dashboard: show decision runtime metrics and simplify model management Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * [Feat] Dashboard: add System One testing and model monitoring Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * [Fix] Dashboard: include model runtime API in container builds Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * [Fix] Dashboard: preserve low-rate chart precision Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * feat(dashboard): complete System One workspace and improve page loading Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * feat: clarify recipe defaults and complete decision model management Use global Router Replay capture defaults with explicit per-decision overrides and remove routing data policy. Resolve default and named recipe strategy and fallback independently from global.router. Add a generated Decision runtime catalog, remove Dashboard KB and WizMap management, move MCP into System, and load embeddings only for active consumers. Update canonical config, DSL, CLI, schemas, documentation, and behavioral tests. Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * fix(dashboard): let nested selectors handle Escape before dialogs Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * feat(systemone): unify decision tasks, explicit entrypoints and instance modes Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * feat(runtime): unify frontend modes and model replica pools Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * fix(runtime): complete pending activation and reject failed replica pools Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * fix(dashboard): display model artifacts independently of deployment identity Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * fix(runtime): require complete input evidence for judgment tasks Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * docs(api): regenerate native input coverage contract Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * test(runtime): verify coverage separately from published model answers Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * test(classification): separate external and repository imports Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * fix: keep decision observations responsive and replica outcomes accurate Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * style: satisfy replica dispatch static checks Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * feat: simplify serving CLI and make replica placement explicit Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * fix: keep release projections independent of runtime imports Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * fix: show authorized native API example for engine startup Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * fix: isolate CI Python tooling and satisfy release checks Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * fix: preserve ownership of CLI startup projections Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * fix: include Hub client in CLI reference dependencies Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * test(cli): isolate output validation from cached router images Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * Fix System One authorization and restore routing regression contracts Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * Preserve Playground sessions and complete dashboard workflow regressions Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * Preserve full-input signal contracts and align recipe evaluation Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * Reconcile retained managed Postgres credentials during startup Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * Sort managed storage imports Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * Keep Engine lifecycle test ports within the stack range Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * Reject oversized full-input embeddings before inference Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * Budget CPU worker threads directly and preserve long-input probe semantics Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * Describe generated probe contexts without unverified token counts Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * Allocate the runner CPU budget to serial full-context recipe checks Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * Clarify runtime development installation and existing stack migration Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * Preserve authored preference execution contracts Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * Honor declared Preview budgets in conformance deployments Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * Clarify default startup and full-input model coverage Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * Isolate no-route gateway contracts from default-provider fallback Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * Link Chinese model guidance to the rendered long-input reference Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * docs: clarify runtime calls and inference deadlines Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * fix(runtime): keep offline decision services from blocking startup Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> * test(runtime): query the reranker worker for direct score comparison Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com> --------- Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
Purpose
Fixes #58858
GLM-5.3-Flash serving on ROCm silently corrupts its DSA index cache. Root cause:
DeepseekV32IndexerBackend.get_supported_kernel_block_sizesreturned[1, MultipleOf(16)]on ROCm for every deployment, soselect_common_block_sizekeeps the hybrid KV manager block (640/1152/2176/4352 for TP8..TP1) as the kernel block size, whileprepare_kernel_block_sizesdrops the spec'sstorage_block_sizewhich every other consumer (cache views, metadata builders, hisparse) already uses as the kernel block.Glm5NextIndexerCachewithindex_kpool = 4, however, stores index-cache entries in pool pages ofpage * index_kpool= 128/256 tokens, and both the writer (get_compressed_slot_mapping) and the gather address the block table ascolumn = pos // storage_block_size. A manager-granular table therefore skips the builder's conversion branch (storage % kernel != 0), indexes far outside the request's row (TP2 @32k: raw table 16 columns wide, 256 needed -> page 0 aliased everywhere), and at TP4 writes past the row buffer entirely.This is also the root cause behind these issues, #55280, #54359 and #56380.
Changes
vllm/v1/worker/utils.py-prepare_kernel_block_sizesis the point where the cache group's spec (kv_cache_group.kv_cache_spec) and the group's backends are both in hand, butstorage_block_sizewas dropped before dispatch. It now takesMLAAttentionSpec.storage_block_sizeas the kernel block size if the group's backends support it, otherwise falls back toselect_common_block_sizeunchanged. On CUDA the indexer declares an exact[64], so the storage block fails validation there and the conversion path keeps working the same way as today. The support check is the existingblock_size_is_supportedlogic moved out ofselect_common_block_sizeto module level, so both call sites use the same check (exact==forint, multiple-of forMultipleOf).Test Plan
Sample command gsm8k + bench:
New tests:
tests/v1/attention/test_kpool_indexer_block_sizes.py:prepare_kernel_block_sizesreturns the storage block (128/256), not the manager block (640/1152/2176/4352), for the kpool group at every TP floor with the real group backends (indexer + tail + aiter sparse MLA) - fails on main as[640] == [128]; plus the fallback, backend rejects the storage block (CUDA-style exact[64]) or spec withoutstorage_block_size->select_common_block_sizeas before.Test Result
713ec07)Benchmark (
random1024/1024/100, c32, 4 runs incl.--no-enable-prefix-caching):713ec07In general no difference between main and this PR.
P.S. TP8 is not tested, because on the tested main
713ec07TP8 will produce gibberish outputs.Essential Elements of an Effective PR Description Checklist
supported_models.mdandexamplesfor a new model.