Skip to content

[Pooling] Cap max-length padding for chunked embeddings - #56505

Merged
yewentao256 merged 5 commits into
vllm-project:mainfrom
taneem-ibrahim:fix/pooling-chunked-max-length-padding
Sep 16, 2026
Merged

yewentao256 merged 5 commits into
vllm-project:mainfrom
taneem-ibrahim:fix/pooling-chunked-max-length-padding

Conversation

@taneem-ibrahim

@taneem-ibrahim taneem-ibrahim commented Sep 11, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Fix padding="max_length" with chunked embeddings. Padding incorrectly used the aggregate max_embed_len, creating artificial pad-only chunks. The fix caps padding at the smaller of max_model_len and max_embed_len while preserving the original validation limit.

Reproducer

Start vLLM:

VLLM_USE_V2_MODEL_RUNNER=1 \
vllm serve intfloat/multilingual-e5-small \
  --runner pooling \
  --dtype bfloat16 \
  --enforce-eager \
  --max-model-len 512 \
  --pooler-config \
  '{"pooling_type":"MEAN","use_activation":true,"enable_chunked_processing":true,"max_embed_len":10000}' \
  --gpu-memory-utilization 0.8

In another terminal run:

curl http://localhost:8000/v1/embeddings \
  -H "Content-Type: application/json" \
  -d '{
    "model": "intfloat/multilingual-e5-small",
    "input": ["hello"],
    "encoding_format": "float",
    "padding": "max_length"
  }'

Output

Main

prompt_tokens=10000

Branch

prompt_tokens=512
short_reference_max_abs_diff=9.313225746154785e-10
long_input_tokens=1802
long_input_max_abs_diff=1.862645149230957e-09

Test Plan and Results

.venv/bin/python -m pytest \
  tests/entrypoints/pooling/embed/test_protocol.py \
  tests/entrypoints/pooling/embed/test_io_processor.py -q
# 66 passed

VLLM_USE_V2_MODEL_RUNNER=1 CUDA_VISIBLE_DEVICES=<GPU> \
  .venv/bin/python -m pytest \
  tests/entrypoints/pooling/embed/test_online_long_text.py -q
# 5 passed

AI assistance disclosure

OpenAI Codex (GPT-5) assisted with the implementation.

Signed-off-by: Taneem Ibrahim <taneem.ibrahim@gmail.com>

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

@mergify mergify Bot added the frontend label Sep 11, 2026

@yewentao256 yewentao256 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We now firstly padding, then chunk. So 513 -> 512 + 1.
Shall we chunk first then padding? So 513 -> 512 + 512(1 padding to 512)

@taneem-ibrahim

Copy link
Copy Markdown
Contributor Author

We now firstly padding, then chunk. So 513 -> 512 + 1. Shall we chunk first then padding? So 513 -> 512 + 512(1 padding to 512)

Correct - the current code does turn 513 tokens into chunks of 512 and 1.

However, padding after chunking would turn the last 1-token chunk into 512 tokens. That means processing 1,024 tokens instead of 513, and the mostly empty second chunk could have too much influence on the final embedding.

For this narrow fix, I think keeping the current order is safer. The current PR behavior of 513 tokens becoming 512 + 1 preserves the original long-input embedding.

Padding each chunk would need a separate and broader change that keeps track of how many real tokens each chunk contained. For example, it would require coordinated changes to padding, real-token tracking, aggregation weights, usage reporting, and accuracy validation.

@yewentao256 yewentao256 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I am not sure the 1 token request would generate correct output, could you test on this?

@taneem-ibrahim

Copy link
Copy Markdown
Contributor Author

I am not sure the 1 token request would generate correct output, could you test on this?

Using intfloat/multilingual-e5-small, using MRV2 on an H100 with bfloat16:

one_token_usage max_length=512 explicit=512
one_token_vs_explicit_511_pads max_abs=4.65661287308e-10 cosine=1

tail_usage max_length=513 do_not_pad=513 chunks=512+1
513_max_length_vs_do_not_pad max_abs=9.31322574615e-10 cosine=1
513_vs_manual_512_plus_1 max_abs=7.20731115067e-09 cosine=1

shape_and_finite one=(384,) tail=(384,) finite=True
batch_usage 1025
batch_one_vs_single max_abs=4.65661287308e-10 cosine=1
batch_513_vs_single max_abs=9.31322574615e-10 cosine=1

I tested both the one-token input and the 513-token boundary that produces a 512+1 split on MRV2 with an H100. The one-token padded output matched an explicitly constructed token-plus-511-pads reference (max_abs=4.66e-10, cosine 1.0). The 513-token result matched both do_not_pad (max_abs=9.31e-10) and a manual weighted aggregation of independent 512-token and 1-token chunks (max_abs=7.21e-09, cosine 1.0). Mixed batching also matched the individual results. Therefore, the one-token tail is handled correctly and I don’t think an additional code change is needed.

@taneem-ibrahim

taneem-ibrahim commented Sep 15, 2026 •

Copy link
Copy Markdown
Contributor Author

The current vLLM behavior is:

  • padding omitted/default: no padding, tokenized input is chunked directly.
  • padding="do_not_pad": same—chunk directly.
  • padding="max_length": padding happens first, then chunking.

This PR change only caps the padding target from max_embed_len to max_model_len. The existing padding before chunking remains as is. This gains:

  • Far less GPU computation, latency, and memory.
  • Correct token accounting: 512 instead of 10,000.
  • No artificial pad-only chunks contaminating the weighted embedding.
  • Real long inputs can still contain up to 10,000 tokens.

The cap is min(max_model_len, max_embed_len), preserving a smaller configured embedding limit.

However, chunking first and padding each chunk would introduce new semantics rather than preserve existing behavior. I think that might be a risky one to do.

@yewentao256 yewentao256 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, thanks for the work!

@yewentao256 yewentao256 added the ready ONLY add when PR is ready to merge/full CI is needed label Sep 15, 2026
@github-actions

Copy link
Copy Markdown

✅ @taneem-ibrahim, CI is now available for this PR.

  • /ci run starts upstream CI; /amd-ci run starts AMD CI only.
  • Your branch must contain every commit currently on its upstream target branch. Merge or rebase onto the latest target branch, then rerun the command. Append --allow-stale to a run command to test an outdated branch at your own risk.
  • /ci retry retries failed jobs in the CI build for the current PR head. If the current head has no CI build, it starts a new CI build for the current head containing only jobs that failed in the latest earlier CI build for this PR.
  • /amd-ci retry retries failed jobs in AMD CI for the current PR head. Use /amd-ci run when the current head has no AMD CI build.
  • /ci cancel cancels scheduled or running CI builds for this PR branch; /amd-ci cancel does the same for AMD CI only.

@taneem-ibrahim

Copy link
Copy Markdown
Contributor Author

/ci run

@github-actions

Copy link
Copy Markdown

✅ Triggered Buildkite CI #89170 for commit f60cb698c52a.

@yewentao256
yewentao256 merged commit 1fd119d into vllm-project:main Sep 16, 2026
91 checks passed
@taneem-ibrahim
taneem-ibrahim deleted the fix/pooling-chunked-max-length-padding branch September 16, 2026 00:34
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

frontend ready ONLY add when PR is ready to merge/full CI is needed

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants