Skip to content

test(examples): cover the remote custom encoder topology end to end - #15200

Draft
furionw wants to merge 2 commits into
qiwa/external-encoder-orchestrator-examplefrom
qiwa/custom-encoder-remote-serve-test
Draft

furionw wants to merge 2 commits into
qiwa/external-encoder-orchestrator-examplefrom
qiwa/custom-encoder-remote-serve-test

Conversation

@furionw

@furionw furionw commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Based on #14859, which depends on utilities PR #15552. This remains a separate draft verification PR.

Full PR sequence

  1. #13293 — encoder handoff (merged)
  2. #15195 — unary endpoint helpers (merged)
  3. #15552 — model and client utilities
  4. #14859 — application-owned orchestrator example
  5. #15200 — draft end-to-end verification

The three open PRs form GitHub stack #15572. The first two remain in completed stack #15197.

Why

The in-process custom encoder has serve coverage, but the remote topology needs a real serve test proving that packed embeddings cross the request plane, drive generation, and remain available to the application classifier. Text-only conversation turns must also work.

What changes

  • Launch the application orchestrator and stock aggregated generator through the service-name configuration and LLMUnaryClient.connect() startup helper.
  • Publish the profile's model ID through the optional public-name override.
  • Check a plain text turn first, then assert 42 generation and a structured classifier label for an image request.
  • Keep dynamic worker ports for concurrent serve tests.

End-to-end verification

  • Exact stacked head: 7655a83052a9c721571a6a101715337b4cbc329d.
  • Image: nvcr.io/nvstaging/ai-dynamo/vllm-runtime:qiwa-dev-vllm-x86-09-27 (baked vLLM 0.29.0); checked-out code rebuilt in release mode and installed editable.
  • Hardware: dlcluster dl-a100, Slurm job 2359334, node 4u4g-0104, two NVIDIA A100 80GB PCIe GPUs.
  • Focused utility/orchestrator tests: 29 passed. The plain-text and image serve cases: 2 passed. The image case checked 42, the classifier label, and the logged encoder-result handoff.
  • Build, test, source-hash, and environment logs are saved locally under dynamo-tmp/logs/10-02/pr15200-7655a83052/. The source hashes match the exact test checkout. Pre-commit and pre-merge checks passed on all three stacked PR heads.
  • The previous stacked head's saved image request returned 42 by default and Paris. when only DYN_CUSTOM_PHRASE changed; both responses included classifier_label: class_a. Its requests, responses, and logs remain under dynamo-tmp/logs/10-01/pr15200-80069e8566/.

@copy-pr-bot

copy-pr-bot Bot commented Sep 23, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

@github-actions github-actions Bot added the test label Sep 23, 2026
@furionw
furionw force-pushed the qiwa/custom-encoder-remote-serve-test branch 2 times, most recently from 8e63edf to 448052e Compare September 23, 2026 01:42
@pull-request-size pull-request-size Bot added size/L and removed size/M labels Sep 23, 2026
@furionw
furionw force-pushed the qiwa/custom-encoder-remote-serve-test branch from 448052e to 93035e4 Compare September 23, 2026 02:11
@pull-request-size pull-request-size Bot added size/M and removed size/L labels Sep 23, 2026
@furionw
furionw force-pushed the qiwa/custom-encoder-remote-serve-test branch from 93035e4 to 04c38ff Compare September 23, 2026 02:12
@furionw
furionw added this pull request to stack #15197 September 23, 2026 02:55
@furionw
furionw force-pushed the qiwa/custom-encoder-remote-serve-test branch from 04c38ff to 8a3f45a Compare September 23, 2026 19:41
@furionw
furionw force-pushed the qiwa/custom-encoder-remote-serve-test branch 3 times, most recently from 48e8f11 to b536e41 Compare September 24, 2026 23:07
@pull-request-size pull-request-size Bot added size/L and removed size/M labels Sep 24, 2026
@furionw
furionw force-pushed the qiwa/custom-encoder-remote-serve-test branch from b536e41 to e7816bd Compare September 25, 2026 00:10
@furionw
furionw removed this pull request from stack #15197 October 2, 2026 06:07
@furionw
furionw force-pushed the qiwa/custom-encoder-remote-serve-test branch 2 times, most recently from 17df95c to 4aee13a Compare October 2, 2026 06:33
@furionw
furionw force-pushed the qiwa/custom-encoder-remote-serve-test branch from 4aee13a to 80069e8 Compare October 2, 2026 06:44
@devin-ai-integration

devin-ai-integration Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

❌ Dynamo PR CI failed — run 37495756206 (attempt 1) on 8ab611d4a5

Gate checks: ❌ backend-status-check · ✅ deploy-status-check · ❌ dynamo-status-check

Framework 1-GPU amd64 1-GPU arm64 Multi-GPU amd64
SGLang ❌ 1 ❌ 1 ❌ 1
TRT-LLM ❌ 1 ❌ 1 ❌ 1
Other Jobs
dynamo-runtime ❌ 5 ✅ 3
planner ⏹️ 1

Failure details

11 jobs failed, all for the same reason: tests/examples/custom_encoder/test_remote_orchestrator.py fails at collection because it imports dynamo.vllm.multimodal_utils, which imports vllm, and vllm is not available in the TRT-LLM, SGLang or dynamo-runtime images. Every other selected test passed. The cancelled planner / Build multi-arch cpu job was cancelled in Build and Push Image after about 60 minutes (16:28 to 17:29 UTC). That does not look like a fail-fast cancel and is consistent with the job time limit.

❌ trtllm-runtime / 2-GPU Test cuda13.1, amd64: collection error, ModuleNotFoundError: No module named 'vllm'

Job: trtllm-runtime / 2-GPU Test cuda13.1, amd64 · Failed step: Run GPU tests (sequential) · Logs: gh run view --job 112387853495 -R ai-dynamo/dynamo --log-failed

ERROR tests/examples/custom_encoder/test_remote_orchestrator.py
tests/examples/custom_encoder/test_remote_orchestrator.py:13: in <module>
    from dynamo.vllm.multimodal_utils.custom_encoder.handoff import (  # noqa: E402
dynamo/vllm/multimodal_utils/__init__.py:5: in <module>
    from dynamo.vllm.multimodal_utils.chat_message_utils import extract_user_text
dynamo/vllm/multimodal_utils/chat_message_utils.py:8: in <module>
    from dynamo.vllm.multimodal_utils.protocol import ChatMessage
dynamo/vllm/multimodal_utils/protocol.py:25: in <module>
    from vllm.inputs import TokensPrompt  # noqa: F401
E   ModuleNotFoundError: No module named 'vllm'

tests/examples/custom_encoder/test_remote_orchestrator.py errors during collection: its import of dynamo.vllm.multimodal_utils.custom_encoder.handoff pulls in dynamo/vllm/multimodal_utils/protocol.py, which runs from vllm.inputs import TokensPrompt, and vllm is not installed in this image. Pytest result: 4 passed, 3 skipped, 1 error. The step fails only because of this collection error.

❌ trtllm-runtime / Test cuda13.1, amd64: collection error, ModuleNotFoundError: No module named 'vllm'

Job: trtllm-runtime / Test cuda13.1, amd64 · Failed step: Run CPU-only tests (parallelized) · Logs: gh run view --job 112387853755 -R ai-dynamo/dynamo --log-failed

ERROR tests/examples/custom_encoder/test_remote_orchestrator.py
tests/examples/custom_encoder/test_remote_orchestrator.py:13: in <module>
    from dynamo.vllm.multimodal_utils.custom_encoder.handoff import (  # noqa: E402
dynamo/vllm/multimodal_utils/__init__.py:5: in <module>
    from dynamo.vllm.multimodal_utils.chat_message_utils import extract_user_text
dynamo/vllm/multimodal_utils/chat_message_utils.py:8: in <module>
    from dynamo.vllm.multimodal_utils.protocol import ChatMessage
dynamo/vllm/multimodal_utils/protocol.py:25: in <module>
    from vllm.inputs import TokensPrompt  # noqa: F401
E   ModuleNotFoundError: No module named 'vllm'

tests/examples/custom_encoder/test_remote_orchestrator.py errors during collection: its import of dynamo.vllm.multimodal_utils.custom_encoder.handoff pulls in dynamo/vllm/multimodal_utils/protocol.py, which runs from vllm.inputs import TokensPrompt, and vllm is not installed in this image. Pytest result: collection error in two pytest stages (313 and 217 tests, 0 failed, 1 error each). The step fails only because of this collection error.

❌ sglang-runtime / Test cuda13.0, amd64: collection error, ModuleNotFoundError: No module named 'vllm'

Job: sglang-runtime / Test cuda13.0, amd64 · Failed step: Run CPU-only tests (parallelized) · Logs: gh run view --job 112395943221 -R ai-dynamo/dynamo --log-failed

ERROR tests/examples/custom_encoder/test_remote_orchestrator.py
tests/examples/custom_encoder/test_remote_orchestrator.py:13: in <module>
    from dynamo.vllm.multimodal_utils.custom_encoder.handoff import (  # noqa: E402
dynamo/vllm/multimodal_utils/__init__.py:5: in <module>
    from dynamo.vllm.multimodal_utils.chat_message_utils import extract_user_text
dynamo/vllm/multimodal_utils/chat_message_utils.py:8: in <module>
    from dynamo.vllm.multimodal_utils.protocol import ChatMessage
dynamo/vllm/multimodal_utils/protocol.py:25: in <module>
    from vllm.inputs import TokensPrompt  # noqa: F401
E   ModuleNotFoundError: No module named 'vllm'

tests/examples/custom_encoder/test_remote_orchestrator.py errors during collection: its import of dynamo.vllm.multimodal_utils.custom_encoder.handoff pulls in dynamo/vllm/multimodal_utils/protocol.py, which runs from vllm.inputs import TokensPrompt, and vllm is not installed in this image. Pytest result: collection error in two pytest stages (1200 and 29 tests, 0 failed, 1 error each). The step fails only because of this collection error.

❌ sglang-runtime / 2-GPU Test cuda13.0, amd64: collection error, ModuleNotFoundError: No module named 'vllm'

Job: sglang-runtime / 2-GPU Test cuda13.0, amd64 · Failed step: Run GPU tests (sequential) · Logs: gh run view --job 112395943282 -R ai-dynamo/dynamo --log-failed

ERROR tests/examples/custom_encoder/test_remote_orchestrator.py
tests/examples/custom_encoder/test_remote_orchestrator.py:13: in <module>
    from dynamo.vllm.multimodal_utils.custom_encoder.handoff import (  # noqa: E402
dynamo/vllm/multimodal_utils/__init__.py:5: in <module>
    from dynamo.vllm.multimodal_utils.chat_message_utils import extract_user_text
dynamo/vllm/multimodal_utils/chat_message_utils.py:8: in <module>
    from dynamo.vllm.multimodal_utils.protocol import ChatMessage
dynamo/vllm/multimodal_utils/protocol.py:25: in <module>
    from vllm.inputs import TokensPrompt  # noqa: F401
E   ModuleNotFoundError: No module named 'vllm'

tests/examples/custom_encoder/test_remote_orchestrator.py errors during collection: its import of dynamo.vllm.multimodal_utils.custom_encoder.handoff pulls in dynamo/vllm/multimodal_utils/protocol.py, which runs from vllm.inputs import TokensPrompt, and vllm is not installed in this image. Pytest result: 5 passed, 4 skipped, 1 error. The step fails only because of this collection error.

❌ dynamo-runtime / test / gpu cuda13.0, amd64: collection error, ModuleNotFoundError: No module named 'vllm.inputs'

Job: dynamo-runtime / test / gpu cuda13.0, amd64 · Failed step: Run GPU tests (sequential) · Logs: gh run view --job 112394991814 -R ai-dynamo/dynamo --log-failed

ERROR tests/examples/custom_encoder/test_remote_orchestrator.py
tests/examples/custom_encoder/test_remote_orchestrator.py:13: in <module>
    from dynamo.vllm.multimodal_utils.custom_encoder.handoff import (  # noqa: E402
dynamo/vllm/multimodal_utils/__init__.py:5: in <module>
    from dynamo.vllm.multimodal_utils.chat_message_utils import extract_user_text
dynamo/vllm/multimodal_utils/chat_message_utils.py:8: in <module>
    from dynamo.vllm.multimodal_utils.protocol import ChatMessage
dynamo/vllm/multimodal_utils/protocol.py:25: in <module>
    from vllm.inputs import TokensPrompt  # noqa: F401
E   ModuleNotFoundError: No module named 'vllm.inputs'

tests/examples/custom_encoder/test_remote_orchestrator.py errors during collection: its import of dynamo.vllm.multimodal_utils.custom_encoder.handoff pulls in dynamo/vllm/multimodal_utils/protocol.py, which runs from vllm.inputs import TokensPrompt, and vllm.inputs is not installed in this image. Pytest result: 36 passed, 7 skipped, 1 error. The step fails only because of this collection error.

❌ dynamo-runtime / test / parallel cuda13.0, amd64: collection error, ModuleNotFoundError: No module named 'vllm.inputs'

Job: dynamo-runtime / test / parallel cuda13.0, amd64 · Failed step: Run CPU-only tests (parallelized) · Logs: gh run view --job 112394991927 -R ai-dynamo/dynamo --log-failed

ERROR tests/examples/custom_encoder/test_remote_orchestrator.py
tests/examples/custom_encoder/test_remote_orchestrator.py:13: in <module>
    from dynamo.vllm.multimodal_utils.custom_encoder.handoff import (  # noqa: E402
dynamo/vllm/multimodal_utils/__init__.py:5: in <module>
    from dynamo.vllm.multimodal_utils.chat_message_utils import extract_user_text
dynamo/vllm/multimodal_utils/chat_message_utils.py:8: in <module>
    from dynamo.vllm.multimodal_utils.protocol import ChatMessage
dynamo/vllm/multimodal_utils/protocol.py:25: in <module>
    from vllm.inputs import TokensPrompt  # noqa: F401
E   ModuleNotFoundError: No module named 'vllm.inputs'

tests/examples/custom_encoder/test_remote_orchestrator.py errors during collection: its import of dynamo.vllm.multimodal_utils.custom_encoder.handoff pulls in dynamo/vllm/multimodal_utils/protocol.py, which runs from vllm.inputs import TokensPrompt, and vllm.inputs is not installed in this image. Pytest result: 796 passed, 23 skipped, 1 error. The step fails only because of this collection error.

❌ dynamo-runtime / test / sequential cuda13.0, amd64: collection error, ModuleNotFoundError: No module named 'vllm.inputs'

Job: dynamo-runtime / test / sequential cuda13.0, amd64 · Failed step: Run CPU-only tests (parallelized) · Logs: gh run view --job 112394992049 -R ai-dynamo/dynamo --log-failed

ERROR tests/examples/custom_encoder/test_remote_orchestrator.py
tests/examples/custom_encoder/test_remote_orchestrator.py:13: in <module>
    from dynamo.vllm.multimodal_utils.custom_encoder.handoff import (  # noqa: E402
dynamo/vllm/multimodal_utils/__init__.py:5: in <module>
    from dynamo.vllm.multimodal_utils.chat_message_utils import extract_user_text
dynamo/vllm/multimodal_utils/chat_message_utils.py:8: in <module>
    from dynamo.vllm.multimodal_utils.protocol import ChatMessage
dynamo/vllm/multimodal_utils/protocol.py:25: in <module>
    from vllm.inputs import TokensPrompt  # noqa: F401
E   ModuleNotFoundError: No module named 'vllm.inputs'

tests/examples/custom_encoder/test_remote_orchestrator.py errors during collection: its import of dynamo.vllm.multimodal_utils.custom_encoder.handoff pulls in dynamo/vllm/multimodal_utils/protocol.py, which runs from vllm.inputs import TokensPrompt, and vllm.inputs is not installed in this image. Pytest result: 4291 passed, 151 skipped, 1 error. The step fails only because of this collection error.

❌ trtllm-runtime / Test cuda13.1, arm64: collection error, ModuleNotFoundError: No module named 'vllm'

Job: trtllm-runtime / Test cuda13.1, arm64 · Failed step: Run CPU-only tests (parallelized) · Logs: gh run view --job 112387853611 -R ai-dynamo/dynamo --log-failed

ERROR tests/examples/custom_encoder/test_remote_orchestrator.py
tests/examples/custom_encoder/test_remote_orchestrator.py:13: in <module>
    from dynamo.vllm.multimodal_utils.custom_encoder.handoff import (  # noqa: E402
dynamo/vllm/multimodal_utils/__init__.py:5: in <module>
    from dynamo.vllm.multimodal_utils.chat_message_utils import extract_user_text
dynamo/vllm/multimodal_utils/chat_message_utils.py:8: in <module>
    from dynamo.vllm.multimodal_utils.protocol import ChatMessage
dynamo/vllm/multimodal_utils/protocol.py:25: in <module>
    from vllm.inputs import TokensPrompt  # noqa: F401
E   ModuleNotFoundError: No module named 'vllm'

tests/examples/custom_encoder/test_remote_orchestrator.py errors during collection: its import of dynamo.vllm.multimodal_utils.custom_encoder.handoff pulls in dynamo/vllm/multimodal_utils/protocol.py, which runs from vllm.inputs import TokensPrompt, and vllm is not installed in this image. Pytest result: 268 passed, 24 skipped, 1 error. The step fails only because of this collection error.

Remaining failed jobs: same collection error, ModuleNotFoundError from tests/examples/custom_encoder/test_remote_orchestrator.py:

For agents
{"pr": 15200, "run_id": 37495756206, "run_attempt": 1, "head_sha": "8ab611d4a58912d181e9adf4422b80f5f5c6b704", "failures": [{"job": "trtllm-runtime / 2-GPU Test cuda13.1, amd64", "job_id": 112387853495, "failed_step": "Run GPU tests (sequential)", "signature": "collection error: ModuleNotFoundError: No module named 'vllm'", "tests": ["tests/examples/custom_encoder/test_remote_orchestrator.py"], "log_cmd": "gh run view --job 112387853495 -R ai-dynamo/dynamo --log-failed"}, {"job": "trtllm-runtime / Test cuda13.1, amd64", "job_id": 112387853755, "failed_step": "Run CPU-only tests (parallelized)", "signature": "collection error: ModuleNotFoundError: No module named 'vllm'", "tests": ["tests/examples/custom_encoder/test_remote_orchestrator.py"], "log_cmd": "gh run view --job 112387853755 -R ai-dynamo/dynamo --log-failed"}, {"job": "sglang-runtime / Test cuda13.0, amd64", "job_id": 112395943221, "failed_step": "Run CPU-only tests (parallelized)", "signature": "collection error: ModuleNotFoundError: No module named 'vllm'", "tests": ["tests/examples/custom_encoder/test_remote_orchestrator.py"], "log_cmd": "gh run view --job 112395943221 -R ai-dynamo/dynamo --log-failed"}, {"job": "sglang-runtime / 2-GPU Test cuda13.0, amd64", "job_id": 112395943282, "failed_step": "Run GPU tests (sequential)", "signature": "collection error: ModuleNotFoundError: No module named 'vllm'", "tests": ["tests/examples/custom_encoder/test_remote_orchestrator.py"], "log_cmd": "gh run view --job 112395943282 -R ai-dynamo/dynamo --log-failed"}, {"job": "dynamo-runtime / test / gpu cuda13.0, amd64", "job_id": 112394991814, "failed_step": "Run GPU tests (sequential)", "signature": "collection error: ModuleNotFoundError: No module named 'vllm.inputs'", "tests": ["tests/examples/custom_encoder/test_remote_orchestrator.py"], "log_cmd": "gh run view --job 112394991814 -R ai-dynamo/dynamo --log-failed"}, {"job": "dynamo-runtime / test / parallel cuda13.0, amd64", "job_id": 112394991927, "failed_step": "Run CPU-only tests (parallelized)", "signature": "collection error: ModuleNotFoundError: No module named 'vllm.inputs'", "tests": ["tests/examples/custom_encoder/test_remote_orchestrator.py"], "log_cmd": "gh run view --job 112394991927 -R ai-dynamo/dynamo --log-failed"}, {"job": "dynamo-runtime / test / sequential cuda13.0, amd64", "job_id": 112394992049, "failed_step": "Run CPU-only tests (parallelized)", "signature": "collection error: ModuleNotFoundError: No module named 'vllm.inputs'", "tests": ["tests/examples/custom_encoder/test_remote_orchestrator.py"], "log_cmd": "gh run view --job 112394992049 -R ai-dynamo/dynamo --log-failed"}, {"job": "trtllm-runtime / Test cuda13.1, arm64", "job_id": 112387853611, "failed_step": "Run CPU-only tests (parallelized)", "signature": "collection error: ModuleNotFoundError: No module named 'vllm'", "tests": ["tests/examples/custom_encoder/test_remote_orchestrator.py"], "log_cmd": "gh run view --job 112387853611 -R ai-dynamo/dynamo --log-failed"}, {"job": "sglang-runtime / Test cuda13.0, arm64", "job_id": 112395943283, "failed_step": "Run CPU-only tests (parallelized)", "signature": "collection error: ModuleNotFoundError: No module named 'vllm'", "tests": ["tests/examples/custom_encoder/test_remote_orchestrator.py"], "log_cmd": "gh run view --job 112395943283 -R ai-dynamo/dynamo --log-failed"}, {"job": "dynamo-runtime / test / sequential cuda13.0, arm64", "job_id": 112394992037, "failed_step": "Run CPU-only tests (parallelized)", "signature": "collection error: ModuleNotFoundError: No module named 'vllm.inputs'", "tests": ["tests/examples/custom_encoder/test_remote_orchestrator.py"], "log_cmd": "gh run view --job 112394992037 -R ai-dynamo/dynamo --log-failed"}, {"job": "dynamo-runtime / test / parallel cuda13.0, arm64", "job_id": 112394992172, "failed_step": "Run CPU-only tests (parallelized)", "signature": "collection error: ModuleNotFoundError: No module named 'vllm.inputs'", "tests": ["tests/examples/custom_encoder/test_remote_orchestrator.py"], "log_cmd": "gh run view --job 112394992172 -R ai-dynamo/dynamo --log-failed"}]}

Posted automatically by Devin for run 37495756206. Updated on every full-CI run of this PR.

@pull-request-size pull-request-size Bot added size/XL and removed size/L labels Oct 2, 2026
@furionw
furionw force-pushed the qiwa/custom-encoder-remote-serve-test branch from 80069e8 to 0efc556 Compare October 2, 2026 16:04
@pull-request-size pull-request-size Bot added size/L and removed size/XL labels Oct 2, 2026
@pull-request-size pull-request-size Bot added size/XL and removed size/L labels Oct 2, 2026
@furionw
furionw force-pushed the qiwa/custom-encoder-remote-serve-test branch from 0efc556 to 7655a83 Compare October 2, 2026 16:10
@pull-request-size pull-request-size Bot added size/L and removed size/XL labels Oct 2, 2026
@furionw
furionw added this pull request to stack #15572 October 2, 2026 17:19
furionw and others added 2 commits October 6, 2026 09:26
The in-process CustomEncoder topology has a serve test; the remote one had
none, so nothing exercised the encoder handoff over a real request plane:
the existing unit tests mock both the handoff and the generator.

Add an agg_custom_remote topology that launches the example's own script and
asserts the same phrase-splice semantics as agg_custom, plus a text-only turn
so a conversation's non-image messages stay served. The wrapper script maps
the harness's dynamic system ports onto the two the example names, so parallel
topologies do not collide on its 8081/8082 defaults.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: furionw <qiwa@nvidia.com>
@furionw
furionw force-pushed the qiwa/custom-encoder-remote-serve-test branch from 7655a83 to 8ab611d Compare October 6, 2026 16:26

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant