Skip to content

[Bugfix][ROCm] Select page-aligned kernel blocks for pooled indexers so the block table addresses storage pages (GLM-5.3-Flash) - #59412

Merged
tjtanaa merged 7 commits into
vllm-project:mainfrom
EmbeddedLLM:bugfix/kpool-indexer-kernel-block-size
Oct 8, 2026
Merged

tjtanaa merged 7 commits into
vllm-project:mainfrom
EmbeddedLLM:bugfix/kpool-indexer-kernel-block-size

Conversation

@vllmellm

@vllmellm vllmellm commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Fixes #58858

GLM-5.3-Flash serving on ROCm silently corrupts its DSA index cache. Root cause: DeepseekV32IndexerBackend.get_supported_kernel_block_sizes returned [1, MultipleOf(16)] on ROCm for every deployment, so select_common_block_size keeps the hybrid KV manager block (640/1152/2176/4352 for TP8..TP1) as the kernel block size, while prepare_kernel_block_sizes drops the spec's storage_block_size which every other consumer (cache views, metadata builders, hisparse) already uses as the kernel block. Glm5NextIndexerCache with index_kpool = 4, however, stores index-cache entries in pool pages of page * index_kpool = 128/256 tokens, and both the writer (get_compressed_slot_mapping) and the gather address the block table as column = pos // storage_block_size. A manager-granular table therefore skips the builder's conversion branch (storage % kernel != 0), indexes far outside the request's row (TP2 @32k: raw table 16 columns wide, 256 needed -> page 0 aliased everywhere), and at TP4 writes past the row buffer entirely.

This is also the root cause behind these issues, #55280, #54359 and #56380.

Changes

  • vllm/v1/worker/utils.py - prepare_kernel_block_sizes is the point where the cache group's spec (kv_cache_group.kv_cache_spec) and the group's backends are both in hand, but storage_block_size was dropped before dispatch. It now takes MLAAttentionSpec.storage_block_size as the kernel block size if the group's backends support it, otherwise falls back to select_common_block_size unchanged. On CUDA the indexer declares an exact [64], so the storage block fails validation there and the conversion path keeps working the same way as today. The support check is the existing block_size_is_supported logic moved out of select_common_block_size to module level, so both call sites use the same check (exact == for int, multiple-of for MultipleOf).

Test Plan

Sample command gsm8k + bench:

VLLM_ROCM_USE_AITER=1 vllm serve zai-org/GLM-5.3-Flash \
  --tensor-parallel-size {2,4} --max-model-len 262144 \
  --attention-backend ROCM_AITER_MLA_SPARSE \
  --tool-call-parser glm47 --enable-auto-tool-choice \
  --reasoning-parser glm47

lm_eval --model local-chat-completions \
  --model_args model=zai-org/GLM-5.3-Flash,base_url=http://127.0.0.1:8000/v1/chat/completions,num_concurrent=32,max_retries=10,max_length=32000,timeout=60000 \
  --gen_kwargs '{"max_tokens":16000,"until":[]}' --batch_size auto \
  --tasks gsm8k --num_fewshot 20 --apply_chat_template

vllm bench serve --model zai-org/GLM-5.3-Flash --host localhost --port 8000 \
  --dataset-name random --random-input-len 1024 --random-output-len 1024 \
  --num-prompts 100 --max-concurrency 32

If NIAH config is required, I can add in.

New tests:

  • tests/v1/attention/test_kpool_indexer_block_sizes.py: prepare_kernel_block_sizes returns the storage block (128/256), not the manager block (640/1152/2176/4352), for the kpool group at every TP floor with the real group backends (indexer + tail + aiter sparse MLA) - fails on main as [640] == [128]; plus the fallback, backend rejects the storage block (CUDA-style exact [64]) or spec without storage_block_size -> select_common_block_size as before.

Test Result

suite main (713ec07) this fix
NIAH 4096-32000, TP2 12/28 28/28
NIAH 4096-32000, TP4 12/28 28/28
NIAH 131072/261632, 3 seeds, TP2 3/42 32/42
lm-eval gsm8k 20-shot (1319) 0.8552 ± 0.0097 0.9742 ± 0.0044

Likely something else in play for the not 100% NIAH result on long context 131072/261632 case. Fixing that would be out-of-scope of this PR.

Benchmark (random 1024/1024/100, c32, 4 runs incl. --no-enable-prefix-caching):

branch output tok/s total tok/s mean TPOT
main 713ec07 1017.8 / 1066.7 2035.5 / 2133.4 24.58 / 24.71 ms
this PR 1051.4 / 1033.3 2102.8 / 2066.5 24.90 / 24.97 ms

In general no difference between main and this PR.

pytest tests/v1/attention/test_kpool_indexer_block_sizes.py  -> 5 passed
pytest tests/v1/attention/test_kpool_tail_slot_mapping.py \
       tests/v1/attention/test_indexer_deepseek_v4_slot_mapping.py \
       tests/v1/worker/test_attn_utils.py          -> 120 passed
test_gpu_model_runner block-size cohort            -> 12 passed

P.S. TP8 is not tested, because on the tested main 713ec07 TP8 will produce gibberish outputs.


Essential Elements of an Effective PR Description Checklist
  • The purpose of the PR, such as "Fix some issue (link existing issues this PR will resolve)".
  • The test plan, such as providing test command.
  • The test results, such as pasting the results comparison before and after, or e2e results
  • (Optional) The necessary documentation update, such as updating supported_models.md and examples for a new model.

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

@mergify mergify Bot added glm rocm Related to AMD ROCm bug Something isn't working labels Sep 30, 2026
@github-project-automation github-project-automation Bot moved this to Todo in AMD Sep 30, 2026
…so the block table addresses storage pages (GLM-5.3-Flash)

Signed-off-by: vllmellm <vllm.ellm@embeddedllm.com>
@vllmellm
vllmellm force-pushed the bugfix/kpool-indexer-kernel-block-size branch from b55ebb1 to b53f00f Compare September 30, 2026 13:48
@amd-dlimpus

Copy link
Copy Markdown

Independent confirmation from MI355X (gfx950), GLM-5.3-Flash, TP=2. I hit the same bug and opened #59704 with a different fix: it expands the 1152-token kernel-block table into 128-token storage pages inside DeepseekV32IndexerMetadataBuilder, rather than changing kernel block selection. I'm closing #59704 in favor of this PR, but here are end-to-end accuracy numbers in case they help review.

Note: these were measured with #59704's fix, not with this branch. They show the impact of the bug, not a validation of this PR's implementation.

GPQA-Diamond (198 questions):

Build Accuracy
main, unfixed (8 runs, bf16 + MXFP4 MLA) 80.81–85.35%, with degenerate/runaway generations
main + page-addressing fix, MXFP4, seeds 42/43/44 90.40 / 90.40 / 91.41%
main + page-addressing fix, bf16, seed 42 89.90%
older base without the regression, bf16 / MXFP4 88.89–91.41% / 90.40–91.92%

Decode-vs-prefill consistency: for each sequence, generate greedily, then re-score the same tokens with one prefill. The table shows the mean |logprob(decode) − logprob(prefill)| over the first 64 generated tokens, and the share of positions where the argmax differs.

Build Mean gap Argmax flips
older base 0.074 6.1%
main, unfixed 0.53–0.69 23–25%
main + fix, bf16 0.075 5.3%
main + fix, MXFP4 MLA 0.069 4.3%

GPQA wall time also dropped from ~3000–3700 s to ~2000–2400 s, because generations stopped running to the token limit.

I'm happy to rerun the same GPQA-Diamond and decode-vs-prefill checks on MI355X against this branch if that would help.

@davidcanar

davidcanar commented Oct 2, 2026 •

Copy link
Copy Markdown

Independent confirmation + end-to-end validation on gfx1151 (2-box TP=2)

We hit the identical failure on the Sep-28 pin (0.31.0.dev0+git73859fec) and separately validated this fix direction on hardware. Setup: GLM-5.3-Flash-AWQ-W4A16, TP=2 (ray across 2× Ryzen AI Max+ 395 / Strix Halo, gfx1151), aiter sparse MLA, max_model_len 262144, hybrid manager block 2304 tokens, Glm5NextIndexerCache storage page 64 pools × index_kpool 4 = 256 tokens (the 256 page of the lattice).

Measured with score-level instrumentation in the gather/top-k path:

  • Needle-pool indexer scores were exactly 0.0 beyond pool 7,295 in every sparse layer, identical across chunks and context sizes. 7,296 = 114 raw block-table columns × 64 pools per storage page: the raw table (2304-token manager blocks) is 114 columns wide at max_model_len 262144, while the gather addresses it at 64-pool storage-page granularity. Below the raw width, sequential allocation makes the manager-granular ids coincide with page ids (benign); beyond it the kernel's block_table_id < block_table_width bound masks the store → zeroed workspace → the tail is unscoreable. This matches the column = pos // storage_block_size mechanism in [Question][ROCm] GLM-5.3-Flash kpool indexer: does the 640-token block table reach 32-pool pages on gfx942/gfx950 too? #58858 exactly.
  • Expanding the table to storage-page granularity at the gather call site (expanded[c] = table[c // K] * K + c % K, K = ceil(needed_cols / have_cols) — identity under sequential allocation) eliminated the dead zone (per-layer lastnz = ke−1) and fixed retrieval end-to-end: needle-retrieval probes PASS at ~12K / 54K / 75K / 105K / 146K / 204K / 255K prompt tokens (the full rated 256K context), including an adversarial repeated-decoy case that had never passed before. Serving on this since 2026-10-01 with no regression.

So: an independent gfx1151 datapoint that (a) the bug reproduces on consumer APUs with the 256-token page, (b) page-granular addressing fixes it fully at 256K context.

Notes for reviewers:

@jin-amd

jin-amd commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

Validation on 8x MI350X (gfx950), GLM-5.3-Flash, at TP4 and TP8, to go with the TP2 results above. I applied this PR's head (b53f00f) to vllm/vllm-openai-rocm:nightly-0cbac6cd1305f710e12193596b27488397bcb205. Both source files apply cleanly, and the head also merges cleanly into today's main. The hybrid block is 1152 tokens at TP4 and 640 at TP8, so 9 and 5 storage pages per block.

nightly nightly + #59412
Index-cache round trip on a 9,488-token prompt, TP4 2,372 pools written to 320 slots; 2,056 read back with another pool's keys 2,372 distinct slots, 0 mismatches
Decode top-k overlap with a torch reference over the same pool keys 178/512 512/512
Needle retrieval, TP4 (recipe's AMD command) 2/15 15/15
Needle retrieval, TP8 (command builder's MI350X command) 8/15 15/15
Recipe benchmark, TP4 (8K in / 1K out, concurrency 16): output tok/s, median TPOT 789, 17.98 ms 798, 17.96 ms

The round trip recomputes each pool's compressed key from the writer's inputs and compares it byte for byte with what the prefill gather reads back. Needle retrieval places a passphrase at 10%, 50% and 90% of documents of 9.5K, 14K, 29K, 57K and 105K tokens, greedy, with a unique first line per prompt so nothing is served from the prefix cache. Unpatched, TP4 already misses the needle at 9.5K tokens, and TP8 starts missing it at 29K.

With this PR, reasoning, reasoning_effort, tool calling, image and video checks also pass at TP4, and reasoning and tool calling pass at TP8. The PR's new tests pass on gfx950: 31 passed, and the one skip is the non-ROCm native check.

Both TP8 columns ran with MTP off and VLLM_ROCM_USE_AITER_MOE=0, because on current nightlies a TP8 decode batch of one returns gibberish from the AITER MoE (#59413), which is unrelated to this PR.

@simondanielsson simondanielsson left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the fix! I'm thinking if we can simplify more.

  1. Thought: Is the underlying reason for the bug that throughout the code we generally say that whenever storage_block_size is set, we use that as the kernel block size? (e.g. in create_metadata_builders and hisparse etc). However, in prepare_kernel_block_sizes we ignore this. So could the fix just be to always prio storage_block_size if available?
# inside prepare_kernel_block_sizes
kv_manager_block_size = kv_cache_group.kv_cache_spec.block_size
group_backends = [g.backend for g in attn_groups[kv_cache_gid]]
# new
spec = kv_cache_group.kv_cache_spec
storage_block_size = spec.storage_block_size isinstance(spec, MLAAttentionSpec) else None
if storage_block_size is not None:
    # add some validation that hte backends all support this kernel block size
    selected_kernel_size  = storage_block_size    
else:
    # old
    selected_kernel_size = select_common_block_size(
          kv_manager_block_size, group_backends
      )
kernel_block_sizes.append(selected_kernel_size)
  1. Suggestion: Can we rework the tests a bit to find one or two that capture the bug we're solving here? For instance it'd be great to have just one checking that prepare_kernel_block_sizes does indeed returns the right kernel block sizes when using kpool (i.e. on rocm, not 640 but rather 128)

Let me know what you think :)

@vllmellm

vllmellm commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor Author

Hi @simondanielsson,

Thanks for your feedback, yeah we can totally simplify it.

I started with the backend-declaration route first because I thought letting the backends declare the page lattice was the most general way to get everyone to agree on the same kernel block, but you're right that prepare_kernel_block_sizes is the only place actually dropping storage_block_size, and every other caller already treats it as the kernel block. So the simpler version works: if the spec has storage_block_size and the group's backends accept it, use it, otherwise fall back to select_common_block_size as before (fallback rather than raise — on CUDA the indexer declares an exact [64], so the storage block can legitimately fail validation there, which is the path that works today).

And yes, I can simplify the tests along with it too.

I've updated accordingly, feel free to take a look.

…prepare_kernel_block_sizes (GLM-5.3-Flash)

Use the spec's storage_block_size as the kernel block when the group's
backends accept it, instead of extending the backend declarations, per review.

Signed-off-by: vllmellm <vllm.ellm@embeddedllm.com>

@simondanielsson simondanielsson left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, LGTM! Would be good to have someone also with more experience in this part of the code to have a look

Comment thread tests/v1/attention/test_kpool_indexer_block_sizes.py
Comment thread tests/v1/attention/test_kpool_indexer_block_sizes.py Outdated
@tjtanaa tjtanaa added the ready ONLY add when PR is ready to merge/full CI is needed label Oct 7, 2026
@tjtanaa

tjtanaa commented Oct 7, 2026

Copy link
Copy Markdown
Member

/ci run

@github-actions

github-actions Bot commented Oct 7, 2026

Copy link
Copy Markdown

✅ Triggered Buildkite CI #93258 for commit e1979b0f7c93.

@vllm-agent

vllm-agent commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

CI selector (shadow): 136 test steps (178 jobs) instead of 66 (84 jobs)

Shadow mode: this changes nothing about what CI runs. It shows what the evidence-based selector would pick for this PR, next to today's rules. How it works.

Feedback welcome: reply here if it would skip a step this change needs, or runs something unrelated.

steps (jobs) Today's rules Selector Would skip Would add
NVIDIA, CPU and others 66 (84) 136 (178) 14 (22) 84 (116)
AMD mirrors 59 (74) 122 (156) 9 (12) 72 (94)
Selector would run (136)
  • amd-kernels-mi355
  • amd-lm-eval-small-models-harness
  • amd-qwen3-next-mtp-async-eplb-accuracy
  • async-engine-inputs-utils-worker
  • basic-correctness ×2
  • basic-correctness-cpu-offload
  • basic-correctness-cumem
  • basic-correctness-prefetch-offload
  • basic-correctness-sleep-mode
  • basic-models-tests-extra-initialization ×14
  • basic-models-tests-other
  • batch-invariance-b200
  • batch-invariance-h100
  • benchmarks-cli-test
  • cpu-kernel-tests ×2
  • cpu-multimodal-config
  • cpu-spec-decode-tests
  • crosslayer-kv-layout-distributed-nixlconnector-pd-accuracy-tests-4-gpus
  • cudagraph
  • distributed-compile-comm-4-gpus
  • distributed-compile-rpc-tests-2-gpus
  • distributed-dp-tests-2-gpus
  • distributed-dp-tests-4-gpus
  • distributed-flashinfer-nixlconnector-pd-accuracy-4-gpus
  • distributed-model-tests-2-gpus ×3
  • distributed-mooncakeconnector-pd-accuracy-4-gpus
  • distributed-nixlconnector-pd-accuracy-4-gpus
  • distributed-tests-8xh100
  • distributed-torchrun-examples-4-gpus
  • distributed-torchrun-shutdown-tests-2-gpus
  • dp-ep-distributed-nixlconnector-pd-accuracy-tests-4-gpus
  • e2e-core-1-gpu
  • e2e-core-large-memory
  • e2e-scheduling-1-gpu
  • e2e-scheduling-accuracy-1-gpu
  • elastic-ep-scaling-test
  • engine
  • engine-1-gpu
  • entrypoints-integration-api-server ×4
  • entrypoints-integration-api-server-generate
  • entrypoints-integration-api-server-openai-chat_completion
  • entrypoints-integration-api-server-openai-completion
  • entrypoints-integration-llm
  • entrypoints-integration-multimodal
  • entrypoints-integration-pooling
  • entrypoints-integration-responses-api
  • entrypoints-integration-speech_to_text
  • entrypoints-unit-tests
  • examples
  • extract-hidden-states-integration
  • extract-hidden-states-integration-2-gpus
  • fault-tolerance-e2e-2xh100
  • fusion-e2e-config-sweep-h100
  • fusion-e2e-quick-h100
  • fusion-e2e-tp2-ar-rms-config-sweep-h100
  • fusion-e2e-tp2-ar-rms-dynamo-partition-amd
  • fusion-e2e-tp2-ar-rms-inductor-partition-amd
  • fusion-e2e-tp2-asynctp-config-sweep-h100
  • fusion-e2e-tp2-b200
  • fusion-e2e-tp2-quick-h100
  • hybrid-ssm-nixlconnector-pd-accuracy-tests-4-gpus
  • hybrid-ssm-nixlconnector-pd-prefix-cache-2-gpus
  • kernels-b200 ×3
  • kernels-deepgemm-test-h100
  • kernels-root-misc-test-b200
  • kimi-k3-prefix-cache-4xb200 ×2
  • kv-offload-large
  • kv-offload-medium
  • kv-offload-small
  • language-models-tests-extra-standard ×2
  • language-models-tests-granite-l4-compatibility
  • language-models-tests-hybrid ×2
  • language-models-tests-standard
  • lm-eval-dspark-watermark-2xh100
  • lm-eval-small-models
  • lm-eval-turboquant-k3v4nc
  • lm-eval-turboquant-k8v4
  • lm-eval-turboquant-t3nc
  • lm-eval-turboquant-t4nc
  • lm-eval-watermarking
  • lora ×4
  • lora-tp-distributed ×4
  • metrics-tracing-2-gpus
  • model-executor
  • mooncake-ec-tcp-e2e-2-gpus
  • mrcr-eval-small-models
  • multi-modal-accuracy-eval-small-models
  • multi-modal-models-standard-1-qwen2
  • multi-modal-models-standard-2-qwen3-gemma
  • multi-modal-models-standard-3-llava-qwen2-vl
  • multi-modal-models-standard-4-other-whisper
  • multiconnector-nixl-offloading-pd-accuracy-2-gpus
  • multiconnector-nixl-offloading-pd-edge-cases-2-gpus
  • nixlconnector-pd-edge-cases-2-gpus
  • openai-api-correctness
  • pipeline-context-parallelism-4-gpus
  • plugin-tests-2-gpus
  • push-nixlconnector-pp-prefill-pd-accuracy-4-gpus
  • pytorch-compilation-dynamic-shapes
  • pytorch-compilation-unit-tests
  • pytorch-compilation-unit-tests-h100
  • pytorch-fullgraph-cudagraph-l4-compatibility
  • pytorch-fullgraph-test
  • quantization ×4
  • quantized-models-test
  • quantized-moe-test-b200
  • rayexecutorv2-4-gpus
  • regression
  • replayssm-e2e
  • rust-frontend-core-correctness
  • rust-frontend-distributed
  • rust-frontend-openai-coverage
  • rust-frontend-serve-admin-coverage
  • rust-frontend-tool-use
  • samplers-multimodal-beam-search
  • samplers-test
  • scale-out-ec-e2e-2-gpus
  • sharded-rdt-weight-transfer
  • spec-decode-draft-model ×4
  • spec-decode-eagle-1-deepseek-qwen
  • spec-decode-eagle-2-llama3-qwen-vl-other
  • spec-decode-mtp-deepseek-mimo
  • spec-decode-mtp-gemma4
  • spec-decode-mtp-qwen3-5
  • spec-decode-ngram-suffix
  • spec-decode-speculators
  • v1-attention-b200 ×2
  • v1-attention-h100-mi300 ×2
  • v1-core
  • v1-executor-worker
  • v1-kv-connectors ×4
  • v1-logits-oracle
  • v1-metrics-lmeval
  • v1-others-cpu
  • v1-sample
  • v1-spec-decode
Would skip (today's rules run them) (14)
  • ascend-npu-test
  • basic-models-test-other-cpu
  • basic-models-tests-initialization
  • cpu-language-generation-and-pooling-model-tests ×3
  • cpu-params-env-tokenizers-parser
  • cpu-reasoning-renderers
  • cpu-tool-parsers
  • kernels-fla-ops-test-b200
  • kernels-mhc-test-b200
  • multi-modal-processor ×4
  • multi-modal-processor-cpu ×4
  • pytorch-compilation-passes-unit-tests
  • rl-entrypoints-tests
  • v1-kv-offload
Would add (today's rules do not run them) (84)
  • amd-kernels-mi355 (Python record)
  • amd-lm-eval-small-models-harness (Python record)
  • amd-qwen3-next-mtp-async-eplb-accuracy (Python record)
  • basic-models-tests-extra-initialization ×14 (Python record)
  • batch-invariance-b200 (Python record)
  • batch-invariance-h100 (Python record)
  • cpu-kernel-tests ×2 (code map)
  • cpu-spec-decode-tests (code map)
  • crosslayer-kv-layout-distributed-nixlconnector-pd-accuracy-tests-4-gpus (Python record)
  • cudagraph (Python record)
  • distributed-compile-comm-4-gpus (Python record)
  • distributed-dp-tests-4-gpus (code map)
  • distributed-flashinfer-nixlconnector-pd-accuracy-4-gpus (Python record)
  • distributed-model-tests-2-gpus ×3 (Python record)
  • distributed-mooncakeconnector-pd-accuracy-4-gpus (Python record)
  • distributed-nixlconnector-pd-accuracy-4-gpus (Python record)
  • distributed-tests-8xh100 (Python record)
  • distributed-torchrun-examples-4-gpus (Python record)
  • dp-ep-distributed-nixlconnector-pd-accuracy-tests-4-gpus (Python record)
  • elastic-ep-scaling-test (Python record)
  • engine (Python record)
  • engine-1-gpu (code map)
  • entrypoints-unit-tests (Python record)
  • examples (Python record)
  • extract-hidden-states-integration (Python record)
  • extract-hidden-states-integration-2-gpus (Python record)
  • fusion-e2e-config-sweep-h100 (Python record)
  • fusion-e2e-quick-h100 (Python record)
  • fusion-e2e-tp2-ar-rms-config-sweep-h100 (Python record)
  • fusion-e2e-tp2-ar-rms-dynamo-partition-amd (Python record)
  • fusion-e2e-tp2-ar-rms-inductor-partition-amd (Python record)
  • fusion-e2e-tp2-asynctp-config-sweep-h100 (Python record)
  • fusion-e2e-tp2-b200 (Python record)
  • fusion-e2e-tp2-quick-h100 (Python record)
  • hybrid-ssm-nixlconnector-pd-accuracy-tests-4-gpus (Python record)
  • hybrid-ssm-nixlconnector-pd-prefix-cache-2-gpus (Python record)
  • kernels-b200 ×3 (Python record)
  • kernels-deepgemm-test-h100 (Python record)
  • kimi-k3-prefix-cache-4xb200 ×2 (Python record)
  • kv-offload-large (Python record)
  • kv-offload-medium (Python record)
  • kv-offload-small (Python record)
  • language-models-tests-extra-standard ×2 (Python record)
  • lm-eval-dspark-watermark-2xh100 (Python record)
  • lm-eval-small-models (Python record)
  • lm-eval-turboquant-k3v4nc (Python record)
  • lm-eval-turboquant-k8v4 (Python record)
  • lm-eval-turboquant-t3nc (Python record)
  • lm-eval-turboquant-t4nc (Python record)
  • lm-eval-watermarking (Python record)
  • lora ×4 (code map)
  • lora-tp-distributed ×4 (Python record)
  • model-executor (code map)
  • mooncake-ec-tcp-e2e-2-gpus (Python record)
  • mrcr-eval-small-models (Python record)
  • multi-modal-accuracy-eval-small-models (Python record)
  • multiconnector-nixl-offloading-pd-accuracy-2-gpus (Python record)
  • multiconnector-nixl-offloading-pd-edge-cases-2-gpus (Python record)
  • nixlconnector-pd-edge-cases-2-gpus (Python record)
  • openai-api-correctness (Python record)
  • pipeline-context-parallelism-4-gpus (Python record)
  • plugin-tests-2-gpus (Python record)
  • push-nixlconnector-pp-prefill-pd-accuracy-4-gpus (Python record)
  • quantization ×4 (Python record)
  • quantized-models-test (Python record)
  • quantized-moe-test-b200 (Python record)
  • rayexecutorv2-4-gpus (Python record)
  • replayssm-e2e (Python record)
  • rust-frontend-core-correctness (Python record)
  • rust-frontend-openai-coverage (Python record)
  • rust-frontend-serve-admin-coverage (Python record)
  • rust-frontend-tool-use (Python record)
  • samplers-multimodal-beam-search (Python record)
  • samplers-test (Python record)
  • scale-out-ec-e2e-2-gpus (Python record)
  • sharded-rdt-weight-transfer (Python record)
  • spec-decode-draft-model ×4 (Python record)
  • spec-decode-eagle-1-deepseek-qwen (Python record)
  • spec-decode-eagle-2-llama3-qwen-vl-other (Python record)
  • spec-decode-mtp-deepseek-mimo (Python record)
  • spec-decode-mtp-gemma4 (Python record)
  • spec-decode-mtp-qwen3-5 (Python record)
  • spec-decode-ngram-suffix (Python record)
  • spec-decode-speculators (Python record)
AMD mirrors: would skip (9)
  • basic-models-tests-initialization
  • fault-tolerance-e2e-2xh100
  • kernels-fla-ops-test-b200
  • kernels-mhc-test-b200
  • multi-modal-processor ×4
  • platform-tests
  • pytorch-compilation-passes-unit-tests
  • rl-entrypoints-tests
  • v1-kv-offload
AMD mirrors: would add (72)
  • crosslayer-kv-layout-distributed-nixlconnector-pd-accuracy-tests-4-gpus (Python record)
  • cudagraph (Python record)
  • distributed-dp-tests-4-gpus (code map)
  • distributed-flashinfer-nixlconnector-pd-accuracy-4-gpus (Python record)
  • distributed-mooncakeconnector-pd-accuracy-4-gpus (Python record)
  • distributed-nixlconnector-pd-accuracy-4-gpus (Python record)
  • distributed-tests-8xh100 (Python record)
  • distributed-torchrun-examples-4-gpus (Python record)
  • dp-ep-distributed-nixlconnector-pd-accuracy-tests-4-gpus (Python record)
  • engine (Python record)
  • engine-1-gpu (code map)
  • entrypoints-unit-tests (Python record)
  • examples (Python record)
  • extract-hidden-states-integration (Python record)
  • extract-hidden-states-integration-2-gpus (Python record)
  • fusion-and-compile-unit-tests-2xb200 (Python record)
  • fusion-e2e-config-sweep-h100 (Python record)
  • fusion-e2e-quick-h100 (Python record)
  • fusion-e2e-tp2-ar-rms-config-sweep-h100 (Python record)
  • fusion-e2e-tp2-asynctp-config-sweep-h100 (Python record)
  • fusion-e2e-tp2-b200 (Python record)
  • fusion-e2e-tp2-quick-h100 (Python record)
  • hybrid-ssm-nixlconnector-pd-accuracy-tests-4-gpus (Python record)
  • hybrid-ssm-nixlconnector-pd-prefix-cache-2-gpus (Python record)
  • jit-monitor-no-runtime-jit (Python record)
  • kernels-moe-test ×5 (Python record)
  • kernels-quantization-test ×6 (Python record)
  • kv-offload-large (Python record)
  • kv-offload-medium (Python record)
  • kv-offload-small (Python record)
  • language-models-tests-extra-standard ×2 (Python record)
  • lm-eval-dspark-watermark-2xh100 (Python record)
  • lm-eval-small-models (Python record)
  • lm-eval-turboquant-k3v4nc (Python record)
  • lm-eval-turboquant-k8v4 (Python record)
  • lm-eval-turboquant-t3nc (Python record)
  • lm-eval-turboquant-t4nc (Python record)
  • lm-eval-watermarking (Python record)
  • lora ×4 (code map)
  • lora-tp-distributed ×4 (Python record)
  • model-executor (code map)
  • mooncake-ec-tcp-e2e-2-gpus (Python record)
  • mrcr-eval-small-models (Python record)
  • multi-modal-accuracy-eval-small-models (Python record)
  • multiconnector-nixl-offloading-pd-accuracy-2-gpus (Python record)
  • multiconnector-nixl-offloading-pd-edge-cases-2-gpus (Python record)
  • nixlconnector-pd-edge-cases-2-gpus (Python record)
  • nixlconnector-pd-spec-decode-acceptance-2-gpus (Python record)
  • openai-api-correctness (Python record)
  • pipeline-context-parallelism-4-gpus (Python record)
  • plugin-tests-2-gpus (Python record)
  • push-nixlconnector-pp-prefill-pd-accuracy-4-gpus (Python record)
  • quantization ×4 (Python record)
  • quantized-models-test (Python record)
  • quantized-moe-test-b200 (Python record)
  • rayexecutorv2-4-gpus (Python record)
  • rust-frontend-core-correctness (Python record)
  • rust-frontend-openai-coverage (Python record)
  • rust-frontend-serve-admin-coverage (Python record)
  • rust-frontend-tool-use (Python record)
  • samplers-multimodal-beam-search (Python record)
  • samplers-test (Python record)
  • scale-out-ec-e2e-2-gpus (Python record)
  • sharded-rdt-weight-transfer (Python record)
  • spec-decode-draft-model ×4 (Python record)
  • spec-decode-eagle-1-deepseek-qwen (Python record)
  • spec-decode-eagle-2-llama3-qwen-vl-other (Python record)
  • spec-decode-mtp-deepseek-mimo (Python record)
  • spec-decode-mtp-gemma4 (Python record)
  • spec-decode-mtp-qwen3-5 (Python record)
  • spec-decode-ngram-suffix (Python record)
  • spec-decode-speculators (Python record)

2 changed files · base 1995c0fcd0 · head 21a2c672eb · Python record: build 93039 at 1e5d0ea888 · kernel record: table 43b4aae (build 93244), map 43b4aae · not counted: 10 build steps, 5 A100 steps the generator no longer emits, 124 optional steps the selector would also run

@vllmellm

vllmellm commented Oct 7, 2026

Copy link
Copy Markdown
Contributor Author

/ci retry

@github-actions

github-actions Bot commented Oct 7, 2026

Copy link
Copy Markdown

✅ Queued 2 failed job(s) for retry in Buildkite CI #93258.

@amd-dlimpus

Copy link
Copy Markdown

Follow-up to my earlier comment: I've now rerun GPQA-Diamond on MI355X with this PR's head (753fec5) applied, rather than with #59704's fix.

Setup: GLM-5.3-Flash, TP=4, MI355X, upstream main 5ab4e44 + #60321 (an MXFP4 KV cache PR that doesn't touch the indexer), seed 42, temperature 1.0, top_p 0.95, 163,840-token limit. Both arms on the same image.

Build KV cache GPQA-Diamond Generations that hit the token limit (of 198)
main ed3f6d1, unfixed (8 runs) bf16 / MXFP4 80.81–85.35% 13–18
main ed3f6d1 + #59704's fix (4 runs) bf16 / MXFP4 89.90–91.41% 1–6
main 5ab4e44 + this PR @ 753fec5 bf16 90.91% 5
main 5ab4e44 + this PR @ 753fec5 MXFP4 89.90% 3

So this PR's fix restores the same accuracy and removes the runaway generations, matching the alternative fix. Caveat: the unfixed and #59704 rows are on an older main (ed3f6d1), and a few commits since then touch the indexer path (e.g. #59211, DCP for the kpool indexer), so the comparison isn't on an identical base. I can add an unfixed 5ab4e44 arm if that would help.

Xunzhuo added a commit to vllm-project/semantic-router that referenced this pull request Oct 7, 2026
decision-balance 0.3.1 moves the boundary between medium and high effort
from 0.8 to 0.675. Repeated stratified cross-validation on the per-effort
benchmark samples chose it together with the hard-STEM difficulty of 2
and the code lane in most folds; held-out accuracy rose by about 1.3
points over 0.8 for about 4% more GPU time per request.

Probes near the new boundary are replaced by ones with clear margins, and
the plan-without-tools negative moves from standard to hard. The Model
Card lists the serving notes behind the measurements on AMD GPUs with
vLLM 0.31: TRITON_ATTN for the 27B, whose default backend's decode cost
grows with input length, and the GLM-5.3-Flash indexer fix of
vllm-project/vllm#59412.

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
@vllmellm

vllmellm commented Oct 8, 2026

Copy link
Copy Markdown
Contributor Author

/ci run

@github-actions

github-actions Bot commented Oct 8, 2026

Copy link
Copy Markdown

✅ Triggered Buildkite CI #93469 for commit 21a2c672eb60.

@tjtanaa

tjtanaa commented Oct 8, 2026

Copy link
Copy Markdown
Member

/ci run

@tjtanaa
tjtanaa enabled auto-merge (squash) October 8, 2026 12:48
@github-actions

github-actions Bot commented Oct 8, 2026

Copy link
Copy Markdown

✅ Triggered Buildkite CI #93595 for commit b4a3052fe617.

@tjtanaa
tjtanaa merged commit 273e099 into vllm-project:main Oct 8, 2026
161 of 162 checks passed
Xunzhuo added a commit to vllm-project/semantic-router that referenced this pull request Oct 9, 2026
…ools (#4734)

* [Feature] Router: let algorithm.decision ask the decision model

An algorithm.decision selector had to name an explicit model_runtime
deployment, so the Vela 2.0 decision model that already answers a
request's signals could not also choose its model without a second copy
of the same weights.

A selector that names no deployment now asks the decision model's shared
implicit deployment, as a routing.signals.decision question does: the
Router resolves it with RouterConfig.DecisionSelectorDeployment, the model
runtime serves it even when no signal uses it, and the selection reasoning
names the deployment that chose. An explicit deployment keeps working;
with decision_model: Vela-1.0 an omitted deployment is a load error that
asks for one, in the Router and in vllm-sr config validate. A blank or
padded deployment is rejected with the same hint.

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* [Feature] Router: let projections read one option of a decision question

The Router publishes every option of a choice, set or span decision
question as its own signal value (decision:<question>:<key>), but a
projection score input could name only the question. A set question's
bare value is its most probable label, so a score such as a reasoning
effort could not weigh one label, like needs:deliberation, on its own.

A projection input of type decision now takes <question>:<key> for one
option of a choice, set or span question; validation rejects a key the
question does not declare and an option of a noul or score question.
Signal usage asks the question whenever a used projection reads one of
its options, and the DSL validator accepts the same form.

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* [Bug] DSL: validate decision-option routes and catalog-backed models cleanly

Three DSL findings from authoring a recipe that routes on Vela 2.0
decision questions over catalog-backed models; a maintained recipe's DSL
must validate without diagnostics and round-trip byte for byte.

- The route guard compared references without their labels, so routes on
  different options of one decision question or classifier (task:agentic
  and task:facts) were reported as overlapping. A labelled reference now
  names its label; the same option in two routes still warns.
- A route that answers with fast_response, inline or through a template,
  calls no model, so it no longer warns that it has no MODEL.
- The decompiler copied a catalog-backed model's parameter size from its
  built-in card into every route, which compiled back as operator
  metadata. Routes now repeat only an operator's own param_size.

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* [Bug] Router: prepare set and span questions that ask the decision model

A routing.signals.decision set or span question that names no deployment
asks the Router's decision model, but preparation checked the model card
of the empty deployment name. The Router failed to start with "provider
\"\" is not served by the model runtime" for any such question, so set
and span questions only worked with an explicit deployment.

Preparation now checks the card of the deployment the question asks: its
own, or the decision model's shared deployment.

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* [Feature] Router: preview the decision selector's choice

Routing Preview reported execution_required for every algorithm.decision
route, so a routing-only evaluation could not show which model the
decision model would choose. The decision selector keeps no state between
requests: it asks the decision model one Choice question about the
request. Preview now dry-runs it like static, multi_factor and
latency_aware and reports the chosen model as selected; a failed or late
answer falls back to the first modelRef as at request time.

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* [Bug] Recipes: plan a GPU-only decision model as a hardware requirement

Live CPU conformance derived hardware requirements only from explicit
model_runtime deployment devices. A recipe whose
global.model_catalog.system.decision_model is Vela-2.0-4B or Vela-2.0-9B
has no explicit deployment, so the planner scheduled it on CPU, where
vllm-sr serve refuses a GPU-only decision model. Such a recipe now lists
the gpu requirement and stays out of the CPU matrix like any other
hardware-bound recipe.

Probe checks also accept the decision algorithm's Preview statuses,
selected or execution_required, like the other base selectors.

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* [Feature] Recipes: add the decision-balance Mixture-of-Models recipe

Decision Balance serves vllm-sr/auto over GLM-5.3-Flash,
Qwen3.8-Flash-Next and Qwen3.8-27B. Vela-2.0-4B answers five System One
questions (task choice, difficulty score, precise-facts noul, needs set,
correction noul) in the call that also answers the prompt guard and
safety signals; heuristics cover images, tools, tool loops, earlier
answers, input length and brief-answer requests.

An effort projection over the difficulty score and the deliberation and
verification labels sets the reasoning effort. Ten decisions route
guard (fast_response), long_context, vision, recovery, agentic and
facts to fixed models, frontier through algorithm: decision with the
decision model itself, hard and standard through multi_factor, and the
rest to Qwen3.8-27B with thinking off. Costs are relative GPU-seconds
per token measured on MI325X; quality evidence reuses the catalog's
third-party records and labels the operator ratings it adds for an
effort without them.

The CPU conformance plan excludes it as GPU-bound, like vela-amd.

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* [Feature] Router: keep personal data out of replay with a recipe data policy

Router Replay could be limited per decision or forbidden per recipe, but
not by what a request contains. A recipe that keeps replay for its
routing evidence therefore also stored the prompts and answers of
requests with names, emails or phone numbers.

routing.data_policy.replay_personal_data: false keeps the replay record
of a request in which one of the recipe's PII signals matched, with its
route, model, signals and detected PII types, but without the request or
response body, prompt, tool definitions or tool trace. The recipe's PII
signals are then evaluated for every request, even when no decision
references them, and a PII classification that fails counts as personal
data. The policy needs a routing.signals.pii rule; the Router and
vllm-sr config validate reject it without one. The DSL round-trips it and
the reference config sets it.

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* [Feature] Dashboard: show the System One answers and routing latency in Playground

A Playground reply showed the decision, algorithm and model but not the
decision-model answers that chose them, nor how long routing took: the
Router sends x-vsr-matched-decision-model, x-vsr-selected-recipe,
x-vsr-selected-confidence and x-vsr-routing-latency-ms, but Playground
did not collect or label them. It now shows them as System One Answers,
Recipe, Decision Confidence and Routing Latency, the latency beside the
other timings.

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* [Feature] Recipes: keep personal data out of decision-balance replay

Decision Balance 0.2.0 asks the decision model's PII span head in the
same call and sets routing.data_policy.replay_personal_data: false, so
Router Replay keeps the routing evidence of a request with personal data
but none of its content. The prompt guard threshold rises from 0.9 to
0.95: on public chat traffic, role-play and code requests scored between
0.9 and 0.95 while the attacks scored higher. A probe asserts the PII span
on a request that names a person, an email address and a phone number.

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* [Test] Dashboard: re-pin the mom-v1 fixture receipt after its probe rewrites

#4725 replaced nine built-in mom-v1 probe examples, which changes the
materialized message text by 22 bytes, but left the Dashboard's receipt
test pinned to the old text. The test fails on main; the new byte count
and digest are those of the probes #4725 ships.

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* [Feature] Recipes: send code to Flash-Next and hard STEM to extra-high effort

decision-balance 0.3.0 adds two lane rules, each from per-effort
measurements on public benchmark samples:

- code: a code task at medium or high effort goes to Qwen3.8-Flash-Next at
  medium effort. On LiveCodeBench it solved more problems at medium effort
  than either Qwen model at extra-high, with less than half the tokens.
- hard: a STEM task the decision model rates at least multi-step
  (difficulty >= 2) runs at extra-high effort even when its effort score is
  lower. Asking for only the final answer lowers the deliberation answer,
  not the reasoning a GPQA-style problem needs; medium effort lost 8 to 15
  points there.

Probes move the two code examples from standard to the new code group and
add collision variants for a letter-only chemistry question (hard), routine
algebra (standard) and a one-line code fix (fast). The Model Card states
the rules, their evidence and that long_output and creativity are reported
but not routed on.

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* [Feature] Router: restrict a listener to named models

A listener's api_keys decide who may call it, but any key holder could
name a provider model and have it passed through, bypassing the recipes.
listeners[].models lists the only request models a standalone listener
accepts, by exact `model` value; empty or absent accepts every model.

The gateway hands the allow-list to the routing core after the API key
check, and the request-body phase checks it against the model request
decoding parsed, the one routing uses, before any signal, cache or
decision runs. Another model gets 403 model_not_allowed in the client's
protocol, and /v1/models on the listener lists only the allowed names.
Request-graph hops the Router makes itself are not restricted, so an
auto model still reaches the provider models its decisions name, and a
restricted listener ignores the skip-processing opt-out, which would
bypass the check.

--gateway extproc rejects a listener with models as unsupported (the
Router's capability check and the CLI's Envoy generation), since the
Envoy listener does not enforce it yet. The schema, the CLI model, the
reference config, the Dashboard type and the docs follow.

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* [Bug] CLI: start vllm-sr serve on Docker daemons without a default bridge

`vllm-sr serve` always passed `--add-host=host.docker.internal:host-gateway`.
Docker derives `host-gateway` from its default bridge network, so a daemon
configured with `"bridge": "none"` rejected every service container with
`unable to derive the IP value for host-gateway`.

- `VLLM_SR_HOST_GATEWAY_IP` maps `host.docker.internal` to an explicit
  address on Docker and Podman.
- Without an override, Docker daemons that report no default bridge network
  skip the mapping with a warning that names the override; any other probe
  result keeps the previous `host-gateway` mapping.

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* [Bug] CLI: accept PII signal rules without a threshold

`vllm-sr config validate` and `vllm-sr serve` rejected a
`routing.signals.pii[]` rule without `threshold` ("Field required"), while
the Router accepts one and lets the rule take every span the PII model
reports. A Vela 2.0 model reports only spans above its size's calibrated
threshold, so a recipe for one size had to hard-code another size's
operating point to pass the CLI.

The CLI schema now makes the rule threshold optional, like the Router, and
the PII signal tutorial says what an omitted threshold means.

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* [Bug] Router: stop RouterDC warnings when no decision uses it

The selection factory initializes every algorithm at startup, so RouterDC
embedded each model's description even when no decision routes with it.
A generation prepares the description embedding model only for decisions
that use router_dc or hybrid, so every start without one logged a warning
per model: `embedding model "mmbert" was not prepared for this generation`.

The embedding set now reports that case as ErrModelNotPrepared (same
message, still a capability error), and RouterDC stops at the first such
error with a debug line. Other embedding failures keep their warning.

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* [Feature] Recipes: let decision-balance's PII rule use the model's own threshold

The CLI now accepts a PII rule without a threshold, so the recipe no longer
hard-codes 0.05. Vela-2.0-4B reports only spans above its calibrated,
length-aware span threshold, which is what the replay data policy should
act on.

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* [Feature] Recipes: raise decision-balance's effort to high from 0.675

decision-balance 0.3.1 moves the boundary between medium and high effort
from 0.8 to 0.675. Repeated stratified cross-validation on the per-effort
benchmark samples chose it together with the hard-STEM difficulty of 2
and the code lane in most folds; held-out accuracy rose by about 1.3
points over 0.8 for about 4% more GPU time per request.

Probes near the new boundary are replaced by ones with clear margins, and
the plan-without-tools negative moves from standard to hard. The Model
Card lists the serving notes behind the measurements on AMD GPUs with
vLLM 0.31: TRITON_ATTN for the 27B, whose default backend's decode cost
grows with input length, and the GLM-5.3-Flash indexer fix of
vllm-project/vllm#59412.

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* [Bug] Router: keep short-request usage out of context token calibration

The context signal's counter learns bytes per token from provider
prompt usage and tracks the lower tail. Provider usage also counts the
chat template, role and default system tokens, so a one-line prompt
reports roughly one token per content byte. After ordinary chat traffic
the learned ratio fell to about 1.2-1.4 bytes/token and long English
prose was over-estimated 3-5x: 204,643 bytes reported 148,755 tokens
against 41,827 counted by the backend, and a 200K-token context band
matched requests of about 55K real tokens. Candidate context-window
checks read the same count.

Ignore samples with less than 4 KiB of text, where template overhead
dominates; longer samples still calibrate the tokenizer ratio and an
uncalibrated deployment keeps the 4 bytes/token default.

Tests: calibrated counter unit tests (short samples ignored, long prose
estimate near the real count) and an extproc test that replays chat
usage through the response calibration path and asserts a ~55K-token
prompt does not match a 200K context rule (it matched at 224K before).

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* [Bug] CLI: never leave the stack down over the sr-bench worker

serve stops Router and Dashboard, then starts sr-bench, and reconciled
an existing worker only when its image alone changed. Any other
difference in the launch identity refused the worker after the stack
was already stopped. In production that difference was the host alias:
the previous CLI launched the worker with host-gateway (a wrapper
rewrote it inside Docker), the new one with an explicit
VLLM_SR_HOST_GATEWAY_IP, so serve exited 1 with the stack down.

Replace an owned worker (sr-bench identity label, store read from its
own arguments) whenever its launch identity differs, after the same
paused journal check. When replacement is not safe (active runs or
preparations, unreadable journal, stopped worker, or a container
without the label) keep the worker as it is, log a warning, and start
the rest of the stack. Reuse still requires an identical identity.

Tests: replacement for image, credential, host-alias and port changes;
reuse when unchanged; active-run, unlabelled and stopped workers are
kept while Dashboard still starts and nothing is rolled back; existing
journal, pause and rollback tests unchanged.

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* [Bug] Recipes: count only identifying types as decision-balance personal data

The personal_data PII rule accepted every span type, so a place, an
organization or a date matched: "What is the capital of France?" (GPE),
"Plan a trip to Kyoto" (GPE) and "How do Apple and Microsoft compete"
(ORGANIZATION) all counted as personal data, and
replay_personal_data: false dropped their content from Replay. Allow
GPE, ORGANIZATION, DATE_TIME, NRP, TITLE and DOMAIN_NAME on their own,
as the privacy recipes do; names, contact details, addresses, identity
and account numbers still match.

Probes: the common-knowledge group (France, Germany) now forbids
personal_data; the Dana Lee probe still expects it.

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* [Feature] Recipes: release decision-balance 0.3.2 with production fixes

Bug-fix release: lanes, thresholds and models are unchanged.

- personal_data counts only identifying PII types.
- With the Router fix in this series, long_context matches near 200K
  backend tokens instead of about 55K after ordinary chat traffic.
- serve keeps the stack up when the sr-bench worker cannot be replaced.

Model Card: the decision call's cost with and without the PII question
(one line 57/89 ms, 2K tokens 0.11/0.39 s, 16K tokens 0.53/1.75 s on one
MI325X), when the rule can be removed (Replay off), the context
estimate's remaining overcount, and a Changes section.

Live conformance on a GPU stack: 45/45 probes.

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* [Docs] Fix translated safety signal scan-budget links

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* [Fix] Dashboard: surface the configured decision model and shared runtime

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* [Feat] Dashboard: manage decision models and identify serving modes

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* [Fix] Dashboard: keep the decision model deploy action readable

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* [Feat] Dashboard: show decision runtime metrics and simplify model management

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* [Feat] Dashboard: add System One testing and model monitoring

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* [Fix] Dashboard: include model runtime API in container builds

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* [Fix] Dashboard: preserve low-rate chart precision

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* feat(dashboard): complete System One workspace and improve page loading

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* feat: clarify recipe defaults and complete decision model management

Use global Router Replay capture defaults with explicit per-decision overrides and remove routing data policy. Resolve default and named recipe strategy and fallback independently from global.router.

Add a generated Decision runtime catalog, remove Dashboard KB and WizMap management, move MCP into System, and load embeddings only for active consumers. Update canonical config, DSL, CLI, schemas, documentation, and behavioral tests.

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* fix(dashboard): let nested selectors handle Escape before dialogs

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* feat(systemone): unify decision tasks, explicit entrypoints and instance modes

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* feat(runtime): unify frontend modes and model replica pools

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* fix(runtime): complete pending activation and reject failed replica pools

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* fix(dashboard): display model artifacts independently of deployment identity

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* fix(runtime): require complete input evidence for judgment tasks

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* docs(api): regenerate native input coverage contract

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* test(runtime): verify coverage separately from published model answers

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* test(classification): separate external and repository imports

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* fix: keep decision observations responsive and replica outcomes accurate

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* style: satisfy replica dispatch static checks

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* feat: simplify serving CLI and make replica placement explicit

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* fix: keep release projections independent of runtime imports

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* fix: show authorized native API example for engine startup

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* fix: isolate CI Python tooling and satisfy release checks

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* fix: preserve ownership of CLI startup projections

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* fix: include Hub client in CLI reference dependencies

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* test(cli): isolate output validation from cached router images

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* Fix System One authorization and restore routing regression contracts

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* Preserve Playground sessions and complete dashboard workflow regressions

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* Preserve full-input signal contracts and align recipe evaluation

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* Reconcile retained managed Postgres credentials during startup

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* Sort managed storage imports

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* Keep Engine lifecycle test ports within the stack range

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* Reject oversized full-input embeddings before inference

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* Budget CPU worker threads directly and preserve long-input probe semantics

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* Describe generated probe contexts without unverified token counts

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* Allocate the runner CPU budget to serial full-context recipe checks

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* Clarify runtime development installation and existing stack migration

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* Preserve authored preference execution contracts

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* Honor declared Preview budgets in conformance deployments

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* Clarify default startup and full-input model coverage

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* Isolate no-route gateway contracts from default-provider fallback

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* Link Chinese model guidance to the rendered long-input reference

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* docs: clarify runtime calls and inference deadlines

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* fix(runtime): keep offline decision services from blocking startup

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

* test(runtime): query the reranker worker for direct score comparison

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>

---------

Signed-off-by: Xunzhuo Liu <xunzhuo.liu@amd.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working glm ready ONLY add when PR is ready to merge/full CI is needed rocm Related to AMD ROCm

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

[Question][ROCm] GLM-5.3-Flash kpool indexer: does the 640-token block table reach 32-pool pages on gfx942/gfx950 too?

7 participants