Repository navigation
Conversation
The final multimodal placement check silently kept a suffix when an embedding had too many rows, which can shift image rows onto different placeholder tokens, and counted only the first dimension of higher-rank encoder outputs. Require an exact flattened token-row count at final placement and report mismatches with both counts and a chunked-prefill hint. Model-owned padding/trimming and EVS placeholder redistribution keep their existing behavior. Relationship to upstream: sgl-project#36724 (open) requires exact token counts for fresh encoder output before cache insertion, as part of a larger request-isolation change, but leaves _adjust_embedding_length's suffix crop at final placement unchanged. Main still crops. Carried from #43. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
rodamani
marked this pull request as ready for review
September 29, 2026 21:15
rodamani
requested review from
Ying1123,
hnyls2002,
merrymercy and
xiezhq-hermann
as code owners
September 29, 2026 21:15
Cropping overlong multimodal embeddings with a warning stays the default. SGLANG_ENABLE_STRICT_MM_EMBEDDING_LENGTH=1 raises on any mismatch between flattened embedding rows and placeholder tokens instead. Restores main's crop tests and the test file's original CI suite. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
….devin.ai/proxy/github.com/modal-projects/sglang into rohan/up/mm-embedding-row-count
Contributor
Author
|
/tag-and-rerun-ci |
3 of 5 tasks
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
The final multimodal placement check in
mm_schedule.py(_adjust_embedding_length) keeps a suffix (with a warning) when an embedding has more rows than placeholder tokens, which can shift image rows onto the wrong placeholders, and it counts only the first dimension of higher-rank encoder outputs. Some deployments may rely on the crop, so this PR adds an opt-in strict check instead of changing the default.Modifications
SGLANG_ENABLE_STRICT_MM_EMBEDDING_LENGTH(defaultFalse). When set,_adjust_embedding_lengthrequires the flattened embedding row count (_embedding_token_count) to equal the placeholder count and raises with both counts and a chunked-prefill hint on any mismatch.test/registered/unit/managers/test_mm_embedding_length.py.Relationship to open PRs: #36724 requires exact token counts for fresh encoder output before cache insertion, but leaves the suffix crop at final placement unchanged.
Accuracy Tests
CPU:
test_mm_embedding_length.py45 passed.Compatibility
No behavior change unless the env var is set. Short-term by design: the underlying question (whether any model legitimately depends on the final-placement crop) needs a per-model audit; if none do, the strict check could become the default in a follow-up.
Speed Tests and Profiling
No hot-path change beyond the fix itself; not separately benchmarked.
Checklist
Review and Merge Process
/tag-and-rerun-ci,/tag-run-ci-label,/rerun-failed-ciCI States
Latest PR Test (Base): ❌ Run #36767693421
Latest PR Test (Extra): ❌ Run #36767693169
Latest PR Test (AMD ROCm 10): ❌ Run #36767693525