Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
26 commits
Select commit Hold shift + click to select a range
d95b086
fix(sglang): use embedding requests for worker canary checks
xianlubird Sep 11, 2026
c97491b
test(sglang): bound embedding handler unit tests
xianlubird Sep 11, 2026
f68dc2b
test(deploy): share chat and embedding API coverage
xianlubird Sep 11, 2026
108e129
test(deploy): fix shared response contracts and endpoint routing
xianlubird Sep 11, 2026
327737d
test(deploy): route shared helper tests through core CPU selection
xianlubird Sep 11, 2026
f9f8150
test(deploy): fix embedding mode and verify stop against baseline
xianlubird Sep 11, 2026
80418f0
test(deploy): remove redundant null stop case
xianlubird Sep 11, 2026
7e6abe8
test(sglang): deduplicate embedding health payload coverage
xianlubird Sep 11, 2026
8cb3ad3
Merge embedding canary test cleanup into deploy API coverage
xianlubird Sep 11, 2026
804ddad
test(deploy): remove redundant validator docstrings
xianlubird Sep 11, 2026
5d09b41
chore: merge main into deploy API coverage
xianlubird Sep 11, 2026
054dc25
chore: merge main into deploy API coverage
xianlubird Sep 15, 2026
2f80e40
test(deploy): tighten API contracts and simplify check wiring
xianlubird Sep 15, 2026
cf52116
chore: merge main into deploy API coverage
xianlubird Sep 16, 2026
368d5be
Merge branch 'main' into codex/deploy-api-coverage
xianlubird Sep 16, 2026
2654467
ci(deploy): avoid Kubernetes E2E fan-out on core changes
xianlubird Sep 17, 2026
9453730
test(deploy): tighten shared API checks
xianlubird Sep 18, 2026
827833a
Merge branch 'main' into codex/deploy-api-coverage
xianlubird Sep 18, 2026
6bb4ca1
fix(deploy): read typed KVCR chat response
xianlubird Sep 18, 2026
63b35df
chore: merge main into deploy API coverage
xianlubird Sep 20, 2026
d8fe62d
fix(deploy): retry dropped streaming responses
xianlubird Sep 22, 2026
b955b60
Merge branch 'main' into codex/deploy-api-coverage
xianlubird Sep 22, 2026
1df3b85
Merge branch 'main' into codex/deploy-api-coverage
xianlubird Sep 23, 2026
74dd78a
chore: merge main into deploy API coverage
xianlubird Oct 8, 2026
de3fb23
fix(deploy): reuse reconnected port forwards across API checks
xianlubird Oct 9, 2026
b780508
chore: merge main and update deploy output helper import
xianlubird Oct 9, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 9 additions & 1 deletion .github/actions/dynamo-deploy-test/action.yml
Original file line number Diff line number Diff line change
Expand Up @@ -41,6 +41,10 @@ inputs:
description: 'Full container image reference for the framework runtime'
required: false
default: ''
frontend_image:
description: 'Full container image reference for the standalone frontend'
required: false
default: ''
platform_arch:
description: 'Platform architecture (amd64, arm64)'
required: false
Expand All @@ -58,7 +62,7 @@ inputs:
required: false
default: ''
extra_pytest_args:
description: 'Additional pytest arguments (e.g., --frontend-image=...)'
description: 'Additional pytest arguments (e.g., -m framework_only)'
required: false
default: ''
cleanup_deployment:
Expand Down Expand Up @@ -105,6 +109,7 @@ runs:
FRAMEWORK: ${{ inputs.framework }}
PROFILE: ${{ inputs.profile }}
IMAGE: ${{ inputs.image }}
FRONTEND_IMAGE: ${{ inputs.frontend_image }}
DYN_TEST_OUTPUT_PATH: ${{ github.workspace }}/test-output
TEST_NAME: ${{ inputs.test_name }}
TEST_FILE: ${{ inputs.test_file }}
Expand All @@ -127,6 +132,9 @@ runs:
if [ -n "${IMAGE}" ]; then
PYTEST_ARGS+=" --image=${IMAGE}"
fi
if [ -n "${FRONTEND_IMAGE}" ]; then
PYTEST_ARGS+=" --frontend-image=${FRONTEND_IMAGE}"
fi
if [ -n "${MODEL_CACHE_PVC}" ]; then
PYTEST_ARGS+=" --model-cache-pvc=${MODEL_CACHE_PVC}"
fi
Expand Down
23 changes: 19 additions & 4 deletions .github/workflows/nightly-ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -536,6 +536,18 @@ jobs:
copy_timeout_minutes: 20
secrets: inherit

frontend-copy-to-acr:
name: frontend-nightly
needs: [compute-release-mode, frontend-build]
if: ${{ needs.compute-release-mode.outputs.run_tests == 'true' }}
uses: $/.github/workflows/shared-copy.yml
with:
target_tag_plain: ${{ format('{0}-nightly', needs.frontend-build.outputs.target_tag_plain) }}
cuda_version: '[""]'
override_arch: amd64
copy_timeout_minutes: 10
secrets: inherit

deepseek-v4-pro-agg-perf:
name: Recipe tests
needs: [compute-release-mode, resolve-source-sha, vllm-build, vllm-copy-to-acr]
Expand Down Expand Up @@ -632,41 +644,44 @@ jobs:

deploy-test-vllm:
name: vLLM Nightly Deploy Test
needs: [compute-release-mode, deploy-operator, vllm-build, vllm-copy-to-acr]
needs: [compute-release-mode, deploy-operator, vllm-build, vllm-copy-to-acr, frontend-copy-to-acr]
if: ${{ needs.compute-release-mode.outputs.run_tests == 'true' }}
uses: $/.github/workflows/shared-deploy-test.yml
with:
framework: vllm
profiles: '["agg", "agg_router_kv_approx"]'
image_tag: ${{ needs.vllm-copy-to-acr.outputs.image_tag }}
frontend_tag: ${{ needs.frontend-copy-to-acr.outputs.image_tag }}
namespace: ${{ needs.deploy-operator.outputs.namespace }}
vcluster_name: ${{ needs.deploy-operator.outputs.vcluster_name }}
operator_tag: ${{ needs.deploy-operator.outputs.operator_tag }}
secrets: inherit

deploy-test-sglang:
name: sglang Nightly Deploy Test
needs: [compute-release-mode, deploy-operator, sglang-build, sglang-copy-to-acr]
needs: [compute-release-mode, deploy-operator, sglang-build, sglang-copy-to-acr, frontend-copy-to-acr]
if: ${{ needs.compute-release-mode.outputs.run_tests == 'true' }}
uses: $/.github/workflows/shared-deploy-test.yml
with:
framework: sglang
profiles: '["agg", "agg_logging"]'
profiles: '["agg", "agg_logging", "agg_embed"]'
image_tag: ${{ needs.sglang-copy-to-acr.outputs.image_tag }}
frontend_tag: ${{ needs.frontend-copy-to-acr.outputs.image_tag }}
namespace: ${{ needs.deploy-operator.outputs.namespace }}
vcluster_name: ${{ needs.deploy-operator.outputs.vcluster_name }}
operator_tag: ${{ needs.deploy-operator.outputs.operator_tag }}
secrets: inherit

deploy-test-trtllm:
name: TensorRT-LLM Nightly Deploy Test
needs: [compute-release-mode, deploy-operator, trtllm-build, trtllm-copy-to-acr]
needs: [compute-release-mode, deploy-operator, trtllm-build, trtllm-copy-to-acr, frontend-copy-to-acr]
if: ${{ needs.compute-release-mode.outputs.run_tests == 'true' }}
uses: $/.github/workflows/shared-deploy-test.yml
with:
framework: trtllm
profiles: '["agg"]'
image_tag: ${{ needs.trtllm-copy-to-acr.outputs.image_tag }}
frontend_tag: ${{ needs.frontend-copy-to-acr.outputs.image_tag }}
namespace: ${{ needs.deploy-operator.outputs.namespace }}
vcluster_name: ${{ needs.deploy-operator.outputs.vcluster_name }}
operator_tag: ${{ needs.deploy-operator.outputs.operator_tag }}
Expand Down
22 changes: 17 additions & 5 deletions .github/workflows/pr.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -757,6 +757,9 @@ jobs:
name: frontend
needs: [changed-files]
if: |
needs.changed-files.outputs.vllm == 'true' ||
needs.changed-files.outputs.sglang == 'true' ||
needs.changed-files.outputs.trtllm == 'true' ||
needs.changed-files.outputs.frontend == 'true' ||
needs.changed-files.outputs.operator == 'true' ||
needs.changed-files.outputs.snapshot == 'true' ||
Expand Down Expand Up @@ -1191,7 +1194,13 @@ jobs:
needs: [changed-files, frontend-build]
if: |
always() && !cancelled() &&
(needs.changed-files.outputs.snapshot == 'true' ||
(needs.changed-files.outputs.frontend == 'true' ||
needs.changed-files.outputs.vllm == 'true' ||
needs.changed-files.outputs.sglang == 'true' ||
needs.changed-files.outputs.trtllm == 'true' ||
needs.changed-files.outputs.deploy == 'true' ||
needs.changed-files.outputs.operator == 'true' ||
needs.changed-files.outputs.snapshot == 'true' ||
needs.changed-files.outputs.snapshot_vllm == 'true' ||
needs.changed-files.outputs.snapshot_sglang == 'true' ||
needs.changed-files.outputs.snapshot_trtllm == 'true') &&
Expand Down Expand Up @@ -1311,7 +1320,7 @@ jobs:

deploy-test-vllm:
name: vllm Deploy Test
needs: [changed-files, deploy-operator, vllm-copy-to-acr]
needs: [changed-files, deploy-operator, vllm-copy-to-acr, frontend-copy-to-acr]
if: |
needs.changed-files.outputs.run_deploy_tests == 'true' &&
!cancelled() && !failure() &&
Expand All @@ -1325,14 +1334,15 @@ jobs:
framework: vllm
profiles: '["agg", "agg_router", "disagg", "disagg_router"]'
image_tag: ${{ needs.vllm-copy-to-acr.outputs.image_tag }}
frontend_tag: ${{ needs.frontend-copy-to-acr.outputs.image_tag }}
namespace: ${{ needs.deploy-operator.outputs.namespace }}
vcluster_name: ${{ needs.deploy-operator.outputs.vcluster_name }}
operator_tag: ${{ needs.deploy-operator.outputs.operator_tag }}
secrets: inherit

deploy-test-sglang:
name: sglang Deploy Test
needs: [changed-files, deploy-operator, sglang-copy-to-acr]
needs: [changed-files, deploy-operator, sglang-copy-to-acr, frontend-copy-to-acr]
if: |
needs.changed-files.outputs.run_deploy_tests == 'true' &&
!cancelled() && !failure() &&
Expand All @@ -1344,16 +1354,17 @@ jobs:
uses: $/.github/workflows/shared-deploy-test.yml
with:
framework: sglang
profiles: '["agg", "agg_router"]'
profiles: '["agg", "agg_router", "agg_embed"]'
image_tag: ${{ needs.sglang-copy-to-acr.outputs.image_tag }}
frontend_tag: ${{ needs.frontend-copy-to-acr.outputs.image_tag }}
namespace: ${{ needs.deploy-operator.outputs.namespace }}
vcluster_name: ${{ needs.deploy-operator.outputs.vcluster_name }}
operator_tag: ${{ needs.deploy-operator.outputs.operator_tag }}
secrets: inherit

deploy-test-trtllm:
name: trtllm Deploy Test
needs: [changed-files, deploy-operator, trtllm-copy-to-acr]
needs: [changed-files, deploy-operator, trtllm-copy-to-acr, frontend-copy-to-acr]
if: |
needs.changed-files.outputs.run_deploy_tests == 'true' &&
!cancelled() && !failure() &&
Expand All @@ -1367,6 +1378,7 @@ jobs:
framework: trtllm
profiles: '["agg", "agg_router"]'
image_tag: ${{ needs.trtllm-copy-to-acr.outputs.image_tag }}
frontend_tag: ${{ needs.frontend-copy-to-acr.outputs.image_tag }}
namespace: ${{ needs.deploy-operator.outputs.namespace }}
vcluster_name: ${{ needs.deploy-operator.outputs.vcluster_name }}
operator_tag: ${{ needs.deploy-operator.outputs.operator_tag }}
Expand Down
5 changes: 5 additions & 0 deletions .github/workflows/shared-deploy-test.yml
Original file line number Diff line number Diff line change
Expand Up @@ -18,6 +18,10 @@ on:
description: 'Runtime image tag'
type: string
required: true
frontend_tag:
description: 'Standalone frontend image tag'
type: string
default: ''
namespace:
description: 'Host namespace where the vCluster lives'
type: string
Expand Down Expand Up @@ -94,6 +98,7 @@ jobs:
profile: ${{ matrix.profile }}
image: ${{ secrets.AZURE_ACR_HOSTNAME }}/${{ vars.ACR_REPOSITORY }}:${{ inputs.image_tag }}
platform_arch: amd64
frontend_image: ${{ inputs.frontend_tag != '' && format('{0}/{1}:{2}', secrets.AZURE_ACR_HOSTNAME, vars.ACR_REPOSITORY, inputs.frontend_tag) || '' }}
extra_pytest_args: -m framework_only
# Mount the shared model cache only when both endpoint vars are set
# (matches the PV/PVC creation gate); otherwise workers download from HF.
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -13,6 +13,7 @@

from sglang.srt.managers.io_struct import EmbeddingReqInput # noqa: E402

from dynamo.sglang.health_check import SglangEmbeddingHealthCheckPayload # noqa: E402
from dynamo.sglang.request_handlers.embedding import ( # noqa: E402
embedding_handler as eh,
)
Expand All @@ -24,6 +25,7 @@
pytest.mark.gpu_0,
pytest.mark.profiled_vram_gib(0),
pytest.mark.pre_merge,
pytest.mark.timeout(30),
]


Expand Down Expand Up @@ -61,6 +63,29 @@ def _handler(*, enable_trace: bool = True) -> eh.EmbeddingWorkerHandler:
return handler


@pytest.mark.asyncio
@pytest.mark.parametrize("use_text_input", [True, False])
async def test_health_check_runs_embedding_inference(monkeypatch, use_text_input):
monkeypatch.delenv("DYN_HEALTH_CHECK_PAYLOAD", raising=False)
handler = _handler()
payload = SglangEmbeddingHealthCheckPayload(
"embedding-model", use_text_input=use_text_input
).to_dict()

outputs = [output async for output in handler.generate(payload, _Context())]

assert len(outputs) == 1
assert outputs[0]["model"] == "embedding-model"
assert len(outputs[0]["data"]) == 1
if use_text_input:
assert handler.engine.async_encode_calls[0]["prompt"] == "Test"
assert handler.engine.tokenizer_manager.requests == []
else:
[(request, _)] = handler.engine.tokenizer_manager.requests
assert request.input_ids == [1]
assert handler.engine.async_encode_calls == []


@pytest.mark.asyncio
@pytest.mark.parametrize(
("embedding_input", "expected_request_id"),
Expand Down
27 changes: 27 additions & 0 deletions components/src/dynamo/sglang/tests/test_sglang_health_check.py
Original file line number Diff line number Diff line change
Expand Up @@ -19,12 +19,15 @@
from dynamo.health_check import HEALTH_CHECK_KEY
from dynamo.sglang.health_check import (
SglangDisaggHealthCheckPayload,
SglangEmbeddingHealthCheckPayload,
SglangHealthCheckPayload,
)
from dynamo.sglang.protocol import EmbeddingRequest

pytestmark = [
pytest.mark.unit,
pytest.mark.sglang,
pytest.mark.fault_tolerance,
pytest.mark.gpu_0,
pytest.mark.pre_merge,
]
Expand All @@ -39,6 +42,30 @@ def test_decode_payload_has_no_marker():
assert HEALTH_CHECK_KEY not in SglangHealthCheckPayload().to_dict()


def test_embedding_payload_uses_engine_bos_token(monkeypatch):
monkeypatch.delenv("DYN_HEALTH_CHECK_PAYLOAD", raising=False)
engine = SimpleNamespace(
tokenizer_manager=SimpleNamespace(tokenizer=SimpleNamespace(bos_token_id=42))
)
payload = SglangEmbeddingHealthCheckPayload(
"embedding-model", engine, use_text_input=False
).to_dict()

request = EmbeddingRequest(**payload)
assert request.model == "embedding-model"
assert request.input == [42]


def test_embedding_env_override(monkeypatch):
override = {"model": "embedding-model", "input": "custom probe"}
monkeypatch.setenv("DYN_HEALTH_CHECK_PAYLOAD", json.dumps(override))

payload = SglangEmbeddingHealthCheckPayload("embedding-model").to_dict()

assert payload == override
assert EmbeddingRequest(**payload).input == "custom probe"


def test_disagg_env_override_preserves_marker(monkeypatch):
"""DYN_HEALTH_CHECK_PAYLOAD must not drop the canary marker."""
monkeypatch.setenv(
Expand Down
13 changes: 13 additions & 0 deletions docs/fern/pages/recipes/kubernetes-templates/dgd/sglang.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -22,6 +22,19 @@ kubectl apply -f agg.yaml
<Accordion title="agg.yaml · Baseline aggregated serving">
<Code src="../../../../../../examples/backends/sglang/deploy/agg.yaml" title="agg.yaml" language="yaml" maxLines={0} />
</Accordion>
<Accordion title="agg_embed.yaml · Aggregated embedding serving">
This Kubernetes template uses `Qwen/Qwen3-Embedding-0.6B` for the single-GPU CI profile; the
[local CLI example](../../cli-templates/sglang.mdx) defaults to `Qwen/Qwen3-Embedding-4B`.

Comment thread
xianlubird marked this conversation as resolved.
<Code src="../../../../../../examples/backends/sglang/deploy/agg_embed.yaml" title="agg_embed.yaml" language="yaml" maxLines={0} />

After forwarding the Frontend service to `localhost:8000`, check the `/v1/embeddings` endpoint:

```bash
curl http://localhost:8000/v1/embeddings -H 'Content-Type: application/json' \
-d '{"model": "Qwen/Qwen3-Embedding-0.6B", "input": "Hello world"}'
```
</Accordion>
<Accordion title="agg_router.yaml · Aggregated with KV-aware routing">
<Code src="../../../../../../examples/backends/sglang/deploy/agg_router.yaml" title="agg_router.yaml" language="yaml" maxLines={0} />
</Accordion>
Expand Down
51 changes: 51 additions & 0 deletions examples/backends/sglang/deploy/agg_embed.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,51 @@
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0

apiVersion: nvidia.com/v1beta1
kind: DynamoGraphDeployment
metadata:
name: sglang-agg-embed
spec:
components:
- name: Frontend
podTemplate:
spec:
containers:
- image: my-registry/sglang-runtime:my-tag
name: main
replicas: 1
type: frontend
- name: decode
Comment thread
xianlubird marked this conversation as resolved.
Comment thread
xianlubird marked this conversation as resolved.
podTemplate:
spec:
containers:
- args:
- --embedding-worker
- --use-sglang-tokenizer
- --model-path
- Qwen/Qwen3-Embedding-0.6B
Comment thread
xianlubird marked this conversation as resolved.
- --served-model-name
- Qwen/Qwen3-Embedding-0.6B
- --page-size
- "16"
- --tp
- "1"
- --trust-remote-code
command:
- python3
- -m
- dynamo.sglang
envFrom:
- secretRef:
name: hf-token-secret
image: my-registry/sglang-runtime:my-tag
name: main
resources:
limits:
nvidia.com/gpu: "1"
requests:
# Increase this value for larger models.
ephemeral-storage: 2Gi
workingDir: /workspace/examples/backends/sglang
replicas: 1
type: worker
Loading
Loading