Skip to content

[PP] Fix start_layer_id with pp in get kv_buffer_shape - #29887

Merged
ShangmingCai merged 6 commits into
sgl-project:mainfrom
staugust:dcp_bug_fix_for_pp
Jul 7, 2026
Merged

ShangmingCai merged 6 commits into
sgl-project:mainfrom
staugust:dcp_bug_fix_for_pp

Conversation

@staugust

@staugust staugust commented Jul 2, 2026

Copy link
Copy Markdown
Contributor

Motivation

Fix bugs for Deepseek V3.1/V3.2/Glm5/GLM5.1 and pp_size > 1, for pp ranks > 0, start _layer_id is not 0.
Fix #29877

Modifications

get_key_buffer from start_layer.

Accuracy Tests

Speed Tests and Profiling

Checklist

Review and Merge Process

  1. Ping Merge Oncalls to start the process. See the PR Merge Process.
  2. Get approvals from CODEOWNERS and other reviewers.
  3. Trigger CI tests with comments or contact authorized users to do so.
    • Common commands include /tag-and-rerun-ci, /tag-run-ci-label, /rerun-failed-ci
  4. After green CI and required approvals, ask Merge Oncalls or people with Write permission to merge the PR.

CI States

Latest PR Test (Base): ❌ Run #28799607482
Latest PR Test (Extra): ❌ Run #28799607823

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Warning

You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again!

@whybeyoung

Copy link
Copy Markdown
Collaborator

/tag-and-rerun-ci

@github-actions github-actions Bot added the run-ci label Jul 2, 2026
Comment thread python/sglang/srt/model_executor/runner/eager_runner.py Outdated

@ShangmingCai ShangmingCai left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@ShangmingCai

Copy link
Copy Markdown
Collaborator

/tag-and-rerun-ci

@whybeyoung whybeyoung changed the title fix bugs for dsv3.1 with pp in get kv_buffer_shape fix bugs with pp in get kv_buffer_shape Jul 2, 2026
@ShangmingCai

Copy link
Copy Markdown
Collaborator

/rerun-failed-ci

@ShangmingCai

Copy link
Copy Markdown
Collaborator

/rerun-test test/registered/models_e2e/test_mimo_v2.py

@github-actions

github-actions Bot commented Jul 6, 2026

Copy link
Copy Markdown
Contributor

Results for /rerun-test test/registered/models_e2e/test_mimo_v2.py:

🚀 8-gpu-h200 (1 test): ✅ View workflow run

cd test/ && python3 registered/models_e2e/test_mimo_v2.py

@ShangmingCai

Copy link
Copy Markdown
Collaborator

/rerun-failed-ci

@ShangmingCai ShangmingCai changed the title fix bugs with pp in get kv_buffer_shape [PP] Fix start_layer_id with pp in get kv_buffer_shape Jul 7, 2026
@ShangmingCai
ShangmingCai merged commit 669fd4b into sgl-project:main Jul 7, 2026
385 of 439 checks passed
hzwzwzw added a commit to hzwzwzw/sglang that referenced this pull request Jul 13, 2026
…last-leaf fix

Upstream sgl-project#29860 (a375e9f, merged). Adapted to our
fork's inline _evict_swa in schedule_batch.py (upstream lives in
mem_cache/common.py:free_swa_out_of_window_slots).

Before: _evict_swa evicted up to pre_len - sliding_window_size - page_size
under the assumption that "extra page keeps the frontier below the insert
boundary". The env flag SGLANG_OPT_SWA_EVICT_DROP_PAGE_MARGIN toggled
whether to apply the extra -page_size subtraction.

After: gate on tree_cache.is_chunk_cache() (replicated tree-cache
property, uniform across ranks) instead of the env flag:
- chunk-cache: no radix tree -> no tombstone-leaf concern; evict up to
  the window boundary (pre_len - sliding_window_size).
- radix: keep max(window, page). The trailing floor page-aligns the
  frontier, and subtracting at least one page keeps the frontier below
  the insert boundary (page_floor(seq_len)) so the last leaf is never
  all-tombstone. This is the case sgl-project#29860 was fixing.

The env var SGLANG_OPT_SWA_EVICT_DROP_PAGE_MARGIN is now unused but
left in environ.py for backward compatibility (harmless no-op).

P2 batch, other PRs skipped this round:
- sgl-project#27550 fix(hiradix): wait for extra pool IO - target code path (the
  completed+pool_transfers_done branch in can_terminate_prefetch) was
  simplified away by our PR sgl-project#27010 port; the pool_transfers_done
  invariant is now enforced through the ack-queue ordering instead.
- sgl-project#28422 decode-hicache _storage_hit_query - the "pre-query" feature
  sgl-project#28422 patches does not exist in our fork.
- sgl-project#29887 [PP] get kv_buffer_shape - target file (eager_runner.py) does
  not exist in our fork.
- sgl-project#29817 write_back policy refinement - refactor, not a bug fix; our
  fork's evict() has diverged from the upstream shape and porting
  cleanly is out of scope for this batch.
- sgl-project#28614 remove large host mem constraint - depends on sgl-project#29817's evict
  refactor for the bulk of its diff, and the standalone piece flips
  prefetch_capacity_limit from max(0, 0.8*(host-device)) to 0.5*host,
  a subtle memory-budget semantics change we won't ship silently.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Chronostasys pushed a commit to MindLab-Research/sglang that referenced this pull request Aug 24, 2026
…9887)

Co-authored-by: ybyang <10629930+whybeyoung@users.noreply.github.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug] prefill pp start error

3 participants