Skip to content

[Core] Support both structured-output manager APIs - #7799

Draft
0z5a wants to merge 5 commits into
vllm-project:mainfrom
0z5a:fix/structured-output-manager-compat
Draft

0z5a wants to merge 5 commits into
vllm-project:mainfrom
0z5a:fix/structured-output-manager-compat

Conversation

@0z5a

@0z5a 0z5a commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Current A100 validation at 719bcf95b0f8 (2×A100 40GB; Torch 2.13.0+cu130, vLLM 0.30.0).

Check Baseline Current result Speedup
Native scoped regression suite N/A 7 passed, 3 skipped N/A
A100 full-model high-concurrency comparison Pending Pending Pending

The installed vLLM 0.30 manager API passed seven cases; three legacy-API cases were skipped. Full BAGEL JSON-schema requests and high concurrency remain pending.

The instance was externally stopped during the remaining test queue. All completed results above were saved off-instance; unfinished measurements remain pending.

Purpose

Keep structured-output advancement compatible with both vLLM manager APIs: vLLM 0.30.0 supplies accept_tokens; vLLM 0.29.0 uses should_advance, reasoning-prefix trimming and the request grammar. Forward the sampled token delta, accept an empty post-reasoning suffix and preserve AR grammar rejection.

This branch is rebased after merged #7820. Current main already calls the shared grammar-rejection method from both schedulers; this patch updates that method only. Scheduler test stubs with an instance-bound accept_tokens remain supported.

Validation

Check Before After Relative speed
vLLM 0.29.0 tests/core/sched 223 passed, 3 failed 233 passed, 3 failed N/A — correctness fix
New dispatch and reasoning-boundary cases Absent 10 passed N/A
Full Qwen3-VL-8B JSON E2E, warm median 566.49 ms 599.04 ms 0.946×
Historical MiniCPM-o-4_5 JSON E2E, warm median 292.33 ms 289.85 ms 1.009×

The same three timeout/log-capture tests fail on main and candidate in the vLLM 0.29.0 host environment. The full Qwen3-VL run used the same GPU and memory budget for both variants: seven greedy schema requests each, 7/7 valid identical JSON outputs per variant, first request excluded from timing. The installed engine is 0.29.0, so both runs used the same 0.29-compatible Omni source and the candidate received this PR's exact helper and scheduler dispatch. The complete rebased 0.30.0 source tree was not started in that environment.

The earlier PR head 92be0140 completed a real three-stage MiniCPM-o-4_5 run on two L20s, with 28/28 valid identical JSON outputs. Neither timing difference on these shared hosts establishes a useful speed improvement. See task-local evidence for the measurement details. This PR stays draft.

@hsliuustc0106 hsliuustc0106 added core related to core module: cache, scheduler, engine, worker, modelrunner enhancement New feature or request labels Sep 18, 2026
@0z5a
0z5a marked this pull request as ready for review September 19, 2026 07:16
@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/input_output_modality_contracts.md, docs/design/module/ar_runtime.md, docs/design/module/cache_management.md.

Module owners: @tzhouam @alex-jw-brooks @amy-why-3459

Routing: @tzhouam via module of the changed files, CODEOWNERS; @alex-jw-brooks via module of the changed files; @amy-why-3459 via module of the changed files

@0z5a, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@vllm-omni-review-bot

vllm-omni-review-bot commented Sep 19, 2026 •

Copy link
Copy Markdown

Omni ReviewBot triage note

Automated triage of commit 0cdb876c379b produced:

  • Priority: high. Prompt maintainer attention is suggested.

These are automated triage suggestions only — the final decision belongs to the maintainers.

@MrlixiangWE MrlixiangWE left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One blocking compatibility issue remains in the vLLM 0.29 fallback. Separately, #7820 rewrites these same two scheduler call sites for the vLLM 0.30 API and currently conflicts with this branch. Please state the intended landing order and rebase or supersede this compatibility layer accordingly so conflict resolution does not restore the broken 0.29 fallback.

"""Support both manager APIs without requiring scheduler stubs to bind a helper."""
if hasattr(type(manager), "accept_tokens"):
return bool(manager.accept_tokens(request, new_token_ids))
if not manager.should_advance(request):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The vLLM 0.29 fallback needs to preserve the reasoning-aware contract used by the released scheduler. That path calls should_advance(request, new_token_ids=...), then trim_reasoning_for_advance, and only feeds the non-empty post-reasoning suffix to the grammar. Here the no-argument gate can derive the wrong delta after async/spec rejection, and even when it detects the boundary, forwarding the full block includes the reasoning prefix and end marker. A speculative block ending with [..., </think>, {] can therefore be rejected, after which the AR scheduler marks the request FINISHED_ERROR. Please pass the sampled delta, trim before acceptance, treat an empty suffix as accepted, and add a vLLM 0.29 boundary regression through the production caller; the current Boolean mocks do not exercise this case.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 92be014: forward new_token_ids, trim the reasoning prefix before grammar acceptance, and accept an empty suffix. Production AR caller regressions cover mixed reasoning-end/JSON tokens and an end-marker-only delta. This should land before #7820; its 0.30 upgrade can then use native accept_tokens and remove the 0.29 fallback while retaining the boundary coverage.

@0z5a

0z5a commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor Author

Validation for 92be014 (before: ed8016c): vLLM 0.29.0, full MiniCPM-o-4_5 three-stage pipeline, two L20s (Stage 0 on one card; Stages 1/2 on the other). Fresh A/P/P/A processes, 7 JSON requests each, first request per process excluded: 12 measured requests per variant. All 28 outputs are valid JSON and match exactly (11 tokens).

E2E case Before median After median Speedup Latency reduction
Structured JSON 292.33 ms 289.85 ms 1.009× +0.85%

This is a correctness fix; the small timing difference on a shared host is not evidence of a meaningful speed improvement. All four processes completed. This replaces the earlier single-pair result; interrupted attempts are excluded.

Regression: the original helper fails both reasoning-boundary cases; the fix passes all three cases through the production AR caller. Core/worker CPU suites: 557 passed, 1 skipped on 0.29.0; 554 passed, 4 skipped on the development build. Other pre-commit checks passed; mypy matches the original head’s 57 existing diagnostics, with no new diagnostics.

0z5a added a commit to 0z5a/vllm-omni that referenced this pull request Sep 23, 2026
Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
0z5a added a commit to 0z5a/vllm-omni that referenced this pull request Sep 24, 2026
Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
@0z5a
0z5a force-pushed the fix/structured-output-manager-compat branch from 8e03b43 to f3cec17 Compare September 24, 2026 06:10
0z5a added a commit to 0z5a/vllm-omni that referenced this pull request Sep 24, 2026
Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
@0z5a
0z5a force-pushed the fix/structured-output-manager-compat branch from f3cec17 to 0cdb876 Compare September 24, 2026 07:10
@0z5a
0z5a marked this pull request as draft September 27, 2026 08:52
Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
@0z5a
0z5a force-pushed the fix/structured-output-manager-compat branch from 0cdb876 to 6c10baa Compare September 27, 2026 09:25
0z5a and others added 3 commits September 27, 2026 19:10
Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
Signed-off-by: 0z5a <192209249+0z5a@users.noreply.github.com>
Signed-off-by: 0z5a <192209249+0z5a@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

core related to core module: cache, scheduler, engine, worker, modelrunner enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants