Skip to content

[Bug-fix] Measure KV transfer speed per actual transfer for disaggregation (Mooncake) - #31993

Closed
hunhokim wants to merge 5 commits into
sgl-project:mainfrom
hunhokim:bug-fix/chunked-prefill-kv-transfer
Closed

hunhokim wants to merge 5 commits into
sgl-project:mainfrom
hunhokim:bug-fix/chunked-prefill-kv-transfer

Conversation

@hunhokim

@hunhokim hunhokim commented Jul 22, 2026 •

Copy link
Copy Markdown
Contributor

Motivation

The current implementation of KV-transfer metrics can substantially over-estimate transfer speed whenever a transfer overlaps computation. For example, in a setup with Infiniband NDR NICs, KV transfer speed is observed to be well above 100 GB/s while the line rate is only 50 GB/s. This makes diagnostics on KV transfer unreliable.

Cause

This happens only when KV transfer occurs multiple times on a single request: the recorded byte count spans the whole transfer while the recorded latency covers only the last transfer, resulting in the transfer speed being inflated. Two specific examples:

  • Chunked prefill: latency is captured only for the last chunk, but the byte count covers all chunks — a large byte count divided by a small slice of time.
  • Early-send of cached prefix: when available, the cached prefix KV is sent to the decode side at the very start of the prefill phase. Latency is measured only over the newly generated tokens transferred after prefill — pairing a full byte count with a partial latency window.

The fix measures time and bytes together, on every actual transfer, so the ratio stays correct no matter how a transfer overlaps computation or splits across chunks.

Implementation

Common interfaces are introduced in disaggregation/common/conn.py with implementations provided for Mooncake.

Inside transfer_worker, each KV transfer is recorded into a per-request _TransferRecord. On completion, get_transfer_metric() pops the record and sources both latency and bytes from it.

If no record exists, it falls back to the old index-derived byte count with None latency (unchanged behavior for backends that don't record, e.g. NIXL).

Some points to consider:

  • Teardown: prefill.py snapshots the record onto the new Req.transfer_metric field before the scheduler's clear() drains it.
  • Replica factor: the measured path sums bytes across destination ranks and does not re-apply the factor; the fallback path scales by get_kv_replica_factor(). Conflating them would reintroduce the MLA byte over-count.
  • Concurrency: a single status_record_lock guards all sites that touch status and the record together (_record_transfer, update_status, clear()), preventing a late chunk from leaking a stale record into the next request reusing the same bootstrap_room.

Unit tests

python -m pytest test/registered/unit/disaggregation/test_kv_transfer_metrics.py

E2E tests

Environments are as follows:

  • DeepSeek-V2-Lite used Prefill CP2 and Decode TP2
  • Qwen3-8B used TP2 for both Prefill and Decode
  • NVIDIA H100 and Infiniband NDR NICs were used

All the values are from prometheus metrics.

Before the patch, transfer speeds show unrealistic values.
After the patch, transfer speeds show reasonable values.

Chunked prefill

prompt_tokens = 8192+1655 (chunked)

Value = KV transfer speed (GiB/s)

Model Before patch After patch
DeepSeek-V2-Lite 106 18.5
Qwen3-8B 80 35

Early-send of cached prefix

prompt_tokens = 6554

Value = KV transfer speed (GiB/s)

Model Before patch After patch
DeepSeek-V2-Lite 231 19.2
Qwen3-8B 141 35

CI States

Latest PR Test (Base): ⏳ Run #29968160092
Latest PR Test (Extra): ⏳ Run #29968159916

Hun-ho Kim and others added 4 commits July 21, 2026 16:55
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Report the measured KV-cache transfer latency and byte count instead of
a static estimate, falling back to the estimate when no measurement is
available. Apply the replica factor only on the fallback path to avoid
double-counting measured bytes for MLA.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Dummy CP ranks transfer no KV and never bootstrap, so return the
zero-default metric early to avoid a spurious warning.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Caution

The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased.

Rephrase the regression-test comments to describe present-tense
invariants and the mistake to avoid, instead of narrating past
"buggy code" / "pre-guard code" states that are no longer visible in
the tree. Also reword the lock-serialization assertion comment so the
counterintuitive assertFalse reads clearly on its own.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@zcfh

zcfh commented Oct 9, 2026

Copy link
Copy Markdown

May I ask why this PR was closed? Was it replaced by a new PR?

@hunhokim

hunhokim commented Oct 9, 2026

Copy link
Copy Markdown
Contributor Author

May I ask why this PR was closed? Was it replaced by a new PR?

I had a private conversation with one of the maintainers on Slack and was told to work on the transfer engine side (e.g. Mooncake) for accurate metrics collection.

@hunhokim
hunhokim restored the bug-fix/chunked-prefill-kv-transfer branch October 9, 2026 22:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants