Skip to content

add check_health in inline client - #3052

Merged
Gaohan123 merged 3 commits into
vllm-project:mainfrom
lengrongfu:fix/health-api
Apr 30, 2026
Merged

Gaohan123 merged 3 commits into
vllm-project:mainfrom
lengrongfu:fix/health-api

Conversation

@lengrongfu

@lengrongfu lengrongfu commented Apr 23, 2026 •

Copy link
Copy Markdown
Contributor

PLEASE FILL IN THE PR DESCRIPTION HERE ENSURING ALL CHECKLIST ITEMS (AT THE BOTTOM) HAVE BEEN CONSIDERED.

Purpose

Fix: #3050

Because when use inline diffusion client, but in this client not check_health method, so need add this method, to check if the inline diffusion engine and its workers are healthy.

Test Plan

  1. Start serve
$ vllm serve /home/jovyan/Wan2.2-T2V-A14B-Diffusers/ --omni --port 8091 
  1. kill work process, we can look this log
ERROR 04-23 03:02:52 [multiproc_executor.py:220] Diffusion worker(s) died unexpectedly: ['DiffusionWorker-0']
  1. Send health reqeust
$ curl -i http://localhost:8091/health
HTTP/1.1 503 Service Unavailable
date: Thu, 23 Apr 2026 03:02:58 GMT
server: uvicorn
content-length: 0

Test Result


Essential Elements of an Effective PR Description Checklist
  • The purpose of the PR, such as "Fix some issue (link existing issues this PR will resolve)".
  • The test plan. Please provide the test scripts & test commands. Please state the reasons if your codes don't require additional test scripts. For test file guidelines, please check the test style doc
  • The test results. Please paste the results comparison before and after, or the e2e results.
  • (Optional) The necessary documentation update, such as updating supported_models.md and examples for a new model. Please run mkdocs serve to sync the documentation editions to ./docs.
  • (Optional) Release notes update. If your change is user-facing, please update the release notes draft.

BEFORE SUBMITTING, PLEASE READ https://github.com/vllm-project/vllm-omni/blob/main/CONTRIBUTING.md (anything written below this line will be removed by GitHub Actions)

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Credits must be used to enable repository wide code reviews.

Signed-off-by: rongfu.leng <lenronfu@gmail.com>
Signed-off-by: rongfu.leng <lenronfu@gmail.com>
@lengrongfu

Copy link
Copy Markdown
Contributor Author

@hsliuustc0106 Hi, can take a look this pr.

Comment thread vllm_omni/diffusion/inline_stage_diffusion_client.py Outdated
Signed-off-by: rongfu.leng <lenronfu@gmail.com>
@Gaohan123 Gaohan123 added the ready label to trigger buildkite CI label Apr 30, 2026

@Gaohan123 Gaohan123 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. Thanks

@Gaohan123
Gaohan123 enabled auto-merge (squash) April 30, 2026 11:17
@Gaohan123
Gaohan123 merged commit da60663 into vllm-project:main Apr 30, 2026
8 checks passed
lengrongfu added a commit to lengrongfu/vllm-omni that referenced this pull request May 1, 2026
Signed-off-by: rongfu.leng <lenronfu@gmail.com>
BeatSeat pushed a commit to BeatSeat/vllm-omni that referenced this pull request May 2, 2026
Signed-off-by: rongfu.leng <lenronfu@gmail.com>
sphinxkkkbc pushed a commit to sphinxkkkbc/vllm-omni that referenced this pull request May 4, 2026
Signed-off-by: rongfu.leng <lenronfu@gmail.com>
Signed-off-by: sphinxkkkbc <binchengkang8@gmail.com>
clodaghwalsh17 pushed a commit to clodaghwalsh17/nm-vllm-omni-ent that referenced this pull request May 12, 2026
Signed-off-by: rongfu.leng <lenronfu@gmail.com>
quyifei23 pushed a commit to quyifei23/vllm-omni that referenced this pull request Jun 6, 2026
Signed-off-by: rongfu.leng <lenronfu@gmail.com>
tzhouam added a commit that referenced this pull request Sep 21, 2026
Brings in 65 main commits (825ab20..ae3880f). Motivation: the three
CI failures left on Buildkite #3060 are main drift, not rebase regressions:

- MiniCPM-o 4.5 seed-tts perf (ready + nightly): the identical request
  ("Tim Tebow ..." item, 30.2 s of silence, response never reaching
  response.done, audio_past_key_values reset at 1500) also fails on main's
  own builds #3049 and #3052, and passes on main from #3054 on, after
  [CI/Build][MiniCPM-o] Fix the Nightly tests (#7758, 6df62d5), which
  the rebase branch did not carry.
- Nemotron VoiceChat native-duplex e2e: main marked the test skip in the
  same #7758 (the pipeline declares no duplex_plugin, so /v1/realtime falls
  through to vLLM's speech-to-text handler and rejects session.update);
  main's build #3054 skipped it, ours ran it.

Conflict resolutions:

- tests/e2e/offline_inference/test_qwen3_omni.py: main's side. #7742/#7803
  compare the deploy-YAML literal (top_k -1) verbatim and normalize the
  runtime check through SamplingParams, which supersedes 3dee645's
  top_k=0 expectation.
- vllm_omni/diffusion/diffusion_kv/model_runner_backend.py: both sides
  dropped. Main (#7166) removed the scheduler_config.max_num_seqs override;
  the rebase removed the kv_caches list because f2aad6aa70's init_kv_cache
  returns the caches as a dict.
- vllm_omni/benchmarks/patch/patch.py: main's Video-MME helpers kept
  together with the rebase's get_samples(args, tokenizer, **kwargs)
  signature; both delegate paths forward **kwargs.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
khairulkabir1661 pushed a commit to khairulkabir1661/vllm-omni that referenced this pull request Sep 25, 2026
Signed-off-by: rongfu.leng <lenronfu@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ready label to trigger buildkite CI

Projects

None yet

3 participants