Repository navigation
dsv4.1: vision tower and image preprocessing - #39668
Merged
Merged
Conversation
hnyls2002
requested review from
HaiShaw,
JustinTong0323,
mickqian,
yctseng0211,
yhyang201 and
yuan-luo
September 15, 2026 22:42
hnyls2002
added this pull request to stack #39669
September 15, 2026 22:43
hnyls2002
removed this pull request from stack #39669
September 15, 2026 22:56
hnyls2002
force-pushed
the
dsv4.1-vision
branch
from
September 15, 2026 22:57
bf0fe9d to
9315d28
Compare
hnyls2002
force-pushed
the
dsv4.1-mhc
branch
from
September 15, 2026 22:57
e6a5980 to
124c47f
Compare
hnyls2002
added this pull request to stack #39672
September 15, 2026 22:58
hnyls2002
force-pushed
the
dsv4.1-mhc
branch
from
September 15, 2026 23:38
124c47f to
e996f4e
Compare
hnyls2002
force-pushed
the
dsv4.1-vision
branch
from
September 15, 2026 23:38
9315d28 to
de8b756
Compare
hnyls2002
force-pushed
the
dsv4.1-mhc
branch
from
September 16, 2026 00:27
e996f4e to
63f7d97
Compare
hnyls2002
force-pushed
the
dsv4.1-vision
branch
from
September 16, 2026 00:27
de8b756 to
3e293dd
Compare
hnyls2002
force-pushed
the
dsv4.1-mhc
branch
from
September 16, 2026 01:12
63f7d97 to
111265c
Compare
hnyls2002
force-pushed
the
dsv4.1-vision
branch
from
September 16, 2026 01:12
3e293dd to
64ca38d
Compare
hnyls2002
force-pushed
the
dsv4.1-mhc
branch
from
September 16, 2026 01:52
111265c to
6dcb649
Compare
hnyls2002
force-pushed
the
dsv4.1-vision
branch
from
September 16, 2026 01:52
64ca38d to
8a9a57b
Compare
hnyls2002
force-pushed
the
dsv4.1-mhc
branch
from
September 16, 2026 03:45
6dcb649 to
ca17422
Compare
hnyls2002
requested review from
BBuf,
Fridge003,
Qiaolin-Yu,
Ying1123,
ch-wan,
hebiao064,
ispobock and
merrymercy
as code owners
September 16, 2026 03:45
Collaborator
Author
|
/tag-and-rerun-ci |
Collaborator
Author
|
/rerun-test test/registered/unit/multimodal/test_processor_async_call_sites.py test/registered/unit/utils/test_hf_transformers.py test/registered/unit/utils/test_hf_transformers_loading.py test/registered/unit/configs/test_model_config_scaling.py |
Contributor
|
Results for 🚀 |
This was referenced Sep 18, 2026
This was referenced Sep 28, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
model_type = "deepseek_v41"andvision_n_layers > 0is loaded.multimodal/deepseek_v41_image_processing.py: grid planning (plan_image_grid,llm_grid), resize/pad/normalise and patchify on PIL, through the Rust extension (patchify_image_rust, from dsv4.1: Rust extension modules for image preprocessing, KV pool names, and PD bootstrap #39677) or on the GPU (prepare_image_gpu/materialize_image_gpu), and the per-image token-type layout (image_token_types).models/deepseek_v41_vit.py:ViT(patch embedding, 2D rotary, transformer blocks on the sharedVisionAttention,RMSNormon its native path with an fp32 weight) andAligner(3x3 downsample to the LLM width).multimodal/processors/deepseek_v41.py:DeepseekV41ImageProcessor, which keeps the raw token ids the model's n-gram hashing needs and therefore runs its own preprocessing instead of the shared worker-pool chain (declared intest_processor_async_call_sites.py). The backend (PIL / Rust / GPU) follows the existing image-processor settings.configs/model_config.py:has_dsv41_visionmarks such configs multimodal and image-understanding;utils/hf_transformers/processor.pyreturns the plain tokenizer for them, since the processor above does the image work.Changes to existing behavior
model_type == "deepseek_v41", and the processor registers forDeepseekV4ForCausalLMonly through the processor registry, which text-only V4 configs never consult (is_multimodalstays false).Verification
test_processor_async_call_sites.py,test_hf_transformers.py,test_hf_transformers_loading.pyandtest_model_config_scaling.pycover the touched shared modules.CI States
Latest PR Test (Base): 🚫 Run #35168832639
Latest PR Test (Extra): ❌ Run #35168832441
Latest PR Test (AMD ROCm 10): 🚫 Run #35168832637