fix(mocker): use logical KV tokens for decode timing [DYN-3851] - #12583
Merged
Conversation
Signed-off-by: PeaBrane <yanrpei@gmail.com>
Contributor
WalkthroughChangesDecode modeling
Estimated code review effort: 2 (Simple) | ~10 minutes 🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
Comment |
dreamtalen
approved these changes
Aug 3, 2026
This comment has been minimized.
This comment has been minimized.
Signed-off-by: PeaBrane <yanrpei@gmail.com>
jthomson04
approved these changes
Aug 3, 2026
…active-kv Signed-off-by: PeaBrane <yanrpei@gmail.com>
PeaBrane
enabled auto-merge (squash)
August 3, 2026 21:39
ishandhanani
approved these changes
Aug 3, 2026
hhzhang16
added a commit
that referenced
this pull request
Aug 4, 2026
dyn-3691-extract-shared-target-pid-cuda-customstorage-operation-layer * 'main' of https://github.com/ai-dynamo/dynamo: (50 commits) docs(cli): correct removed vLLM prefill-worker flag reference (#12581) docs(operator): reserve webhook Ignore for emergencies (#12563) ci(docs): make previews and checks match what actually publishes (#12339) refactor(vllm): organize custom encoder modules (#12416) feat(llm): Select reasoning output field via env var (#11464) feat(runtime): add TLS support to TCP request plane (#10921) fix: convert conditional disagg sglang warning to httperror 400 (#12578) feat(operator): add runtime feature gates (#12421) refactor(runtime): extract PushRouter transport seam behind StreamingDispatch trait (#12447) feat(replay): add deterministic canonical offline reports (#12363) build: bump ModelExpress to 0.5.0(OPS-7978) (#12455) fix(mocker): use logical KV tokens for decode timing (#12583) fix(examples): update Triton example for CUDA 13 + fix libdcgm copy (DYN-3697) (#12577) refactor(operator): implement composition-first DGD reconciliation (#12283) feat(frontend): add basetenkenizer backend (#12376) fix(profiler): configure rapid mocker without planner (#12573) docs(vllm): correct worker-role flags and document --kv-transfer-config (#12568) ci: add Kubernetes deploy test to nightly (#12090) fix(container): reuse pinned protoc in runtime image (#12535) feat(self-host): flip DYN_SELF_HOST_METADATA default to ON (gh-8749) (#11417) ... Signed-off-by: Hannah Zhang <hannahz@nvidia.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
This does not change KV allocation, prefix-cache ownership, routing, or direct AIC decode prediction.
Details
Where should the reviewer start?
Start with
lib/mocker/src/scheduler/vllm/core.rsfor the semantic change andlib/mocker/src/common/perf_model.rsfor the timing-model boundary.Related Issues
Validation
cargo test -p dynamo-mocker test_fpm_prefill_and_decode_mixed_batch -- --nocapture(vLLM and SGLang cases passed locally and on a clean Linux checkout)cargo test -p dynamo-mocker --features kvbm-offload speculative_decode_reservation_waits_without_preempting -- --nocapturecargo fmt --all -- --check, commit hooks, andgit diff --checkComputeLab SC-01: single H100, Qwen3-1.7B, vLLM TP1, 38,863 runtime GPU KV blocks, four prefix groups at 68.68-74.85% measured prefix-cache token hit rate, three trials each at concurrency 8 and 128, zero request errors or output-length mismatches.
At concurrency 128, candidate median ITL was 21.69 ms versus live 22.12 ms; baseline was 8.98 ms. Candidate output-throughput error was -19.6% versus baseline +40.8%.
Paired trace at scheduler batch 106/context 4,103 showed candidate
active_kv_tokens=434,972(106.01 logical contexts) versus baseline143,968(35.09 contexts), confirming the corrected timing input tracks the scheduled batch under prefix dedup.Lower-level attribution on the exact live stack (H100 80GB, pinned vLLM 0.23.1 development image, Qwen3-1.7B BF16 TP1, FlashAttention 3) reproduced the scheduler shape at batch 106 with 434,976 logical KV tokens. At fixed batch and logical total, the actual ragged request vector was within 0.5% of the homogeneous attention control.
The old batch-35 proxy underpriced direct paged FlashAttention by 58.2% and steady-state engine decode by 48.6% versus the actual batch-106 workload. This directly validates removing the dedup-aware scheduler cap.
Physical prefix-block aliasing reduced attention time by 35.7% versus an unaliased replay. That is a separate attention-model fidelity opportunity; it does not make deduplicated physical blocks a valid proxy for the scheduled decode batch.
The live image reports vLLM
0.23.1rc1.dev1502; the available AIC interpolation data is for vLLM0.24.0. Candidate mean and upper-tail simulated ITL remained conservative, so the campaign validates the accounting semantics and center-of-distribution parity rather than exact end-to-end calibration across those versions.