Skip to content
Merged
45 changes: 31 additions & 14 deletions docs/design/metrics.md
Original file line number Diff line number Diff line change
Expand Up @@ -60,7 +60,7 @@ Upstream per-engine metrics retain the `vllm:` prefix but are now registered by

There are five independent paths for metric collection.

**Path 1: Pipeline-level metrics (`vllm_omni:*`)**
#### Path 1: Pipeline-level metrics (`vllm_omni:*`)

`OmniPrometheusMetrics` registers the Gauge / Counter / Histogram collectors at import time. It is instantiated once per entrypoint, labeled with the model name. The entrypoint calls its methods as requests progress:

Expand All @@ -69,7 +69,7 @@ There are five independent paths for metric collection.
- `request_failed()` — recorded by the cleanup path when a request exits without natural completion. Internally maps to `finished_reason="abort"` on the existing completion Counter.
- `inc_requests_failed(reason)` — increments the dedicated failure-attribution Counter once per request. Reasons are normalized to the bounded set `client_abort`, `client_disconnect`, `stage_error`, and `unknown`.

**Path 2: Image and diffusion metrics**
#### Path 2: Image and diffusion metrics

Image and diffusion metrics use two cooperating collectors. `OmniPrometheusMetrics` owns pipeline/stage workload families, while `OmniModalityMetrics` owns per-replica diffusion breakdowns. Finished stage messages are consumed once in `OmniBase._process_single_result`, which prevents replayed terminal messages from incrementing Counters or Histograms twice.

Expand All @@ -81,15 +81,15 @@ Image and diffusion metrics use two cooperating collectors. `OmniPrometheusMetri

Zero and missing values have different meanings. Valid zero-duration queue observations are recorded, while unavailable optional measurements are omitted instead of being exported as synthetic zeroes. `peak_memory_mb`, `diffusion_forward_s`, `diffusion_kv_load_s`, `vae_decode_s`, and `kv_wait_s` therefore appear only when their data source is active.

**Path 3: Audio modality metrics (`vllm_omni:audio_*`)**
#### Path 3: Audio modality metrics (`vllm_omni:audio_*`)

`OmniModalityMetrics` registers seven audio families with `{model_name, stage, replica}` (plus an extra `threshold_ms` / `reason` label on the two extra-cardinality Counters). Three observation sites:

- `observe_modality_at_finalize(...)` — called from `omni_base._process_single_result` inside the existing `e2e_done` finalize guard. For `output_type == "audio"` it emits `audio_frames_total`, `audio_duration_s`, `audio_rtf` (or `audio_skipped_requests_total{reason="no_audio_data"}` when no audio was produced). Sample rate is resolved from `engine_outputs.multimodal_output` via `definitions.resolve_audio_sample_rate(...)` (fallback chain mirrors `serving_chat.py`'s audio response path).
- `observe_audio_first_packet(...)` — called from the OpenAI SSE audio branch in `serving_chat.py` on the first audio packet for a request. The once-per-request guard is held by `ClientRequestState.first_audio_ts`. The `request_arrival_ts` anchor is stored in `ClientRequestState` by `async_omni.generate()`, computed at request entry.
- `observe_audio_streaming_finalize(...)` — called from `serving_chat.py` after the streaming chunk loop exhausts. It runs the per-chunk player simulation from `vllm_omni/benchmarks/audio_continuity.py` to compute the worst-case underrun and emits `audio_underrun_s` plus (when the request stayed below the threshold) `audio_continuity_ok_total{threshold_ms}`. Per-chunk PCM byte counts and arrival timestamps are recorded by the same audio branch that updates `first_audio_ts`.
- `observe_audio_first_packet(...)` — called from the OpenAI SSE audio branch in `serving_chat.py` and the streaming TTS output paths in `serving_speech.py` on the first non-empty audio payload for a request. Streaming Speech observes the first PCM payload rather than a header-only WAV chunk. HTTP carries the middleware request timestamp through response generation; WebSocket uses the start of each sentence synthesis request. Non-streaming Speech does not emit TTFP and uses `e2e_request_latency_s` for overall latency instead.
- `observe_audio_streaming_finalize(...)` — called from `serving_chat.py` and `serving_speech.py` after a streaming chunk loop exhausts normally. It runs the per-chunk player simulation from `vllm_omni/benchmarks/audio_continuity.py` to compute the worst-case underrun and emits `audio_underrun_s` plus (when the request stayed below the threshold) `audio_continuity_ok_total{threshold_ms}`. Per-chunk PCM byte counts and arrival timestamps are recorded by the same output paths that update TTFP. Speech records raw PCM payloads before raw/SSE/WebSocket encoding, excludes WAV headers and empty chunks, and does not emit continuity samples for failed, cancelled, or non-streaming requests.

**Path 4: Cross-stage transfer metrics (`vllm_omni:transfer_*`)**
#### Path 4: Cross-stage transfer metrics (`vllm_omni:transfer_*`)

`OmniTransferMetrics` registers four Histogram families with `{model_name, from_stage, from_replica, to_stage, to_replica}` labels. Each observation corresponds to one physical transfer hop (one chunk between adjacent stages), not the per-request accumulated total — so the histograms track per-transfer distribution.

Expand All @@ -100,7 +100,7 @@ The hook lives in `OrchestratorAggregator.record_transfer_tx` and `record_transf

Defensive fail-safe: if `transfer_emitter` or `replica_resolver` is missing, or the resolver returns `None` for either side, the emit is skipped silently (the underlying `TransferEdgeStats` accumulation is unaffected).

**Path 5: Per-engine metrics (`vllm:*`, stage/replica wrap)**
#### Path 5: Per-engine metrics (`vllm:*`, stage/replica wrap)

The Orchestrator instantiates `OmniPrometheusStatLogger` (a thin subclass of upstream `vllm.v1.metrics.loggers.PrometheusStatLogger`) and feeds it scheduler stats and iteration stats after processing each batch of engine outputs. This populates the standard ~37 vLLM metric families (TTFT, ITL, TPOT, KV cache usage, etc.) using the same upstream code path — but with the `engine` label reshaped into `stage` + `replica` so multi-replica deployments produce distinct series per replica. See the next section for the wrap mechanics.

Expand Down Expand Up @@ -174,7 +174,7 @@ second.
### Pipeline (4)

| Metric | Type | Labels | Description |
|--------|------|--------|-------------|
| -------- | ------ | -------- | ------------- |
| `vllm_omni:num_requests_running` | Gauge | `model_name` | Requests currently executing across all stages |
| `vllm_omni:num_requests_waiting` | Gauge | `model_name` | Requests queued but not yet scheduled |
| `vllm_omni:requests_success_total` | Counter | `model_name`, `finished_reason` | Total requests by completion reason ({stop, length, abort, ...}); aborts cover client-disconnect / cancellation paths in addition to upstream `FinishReason.ABORT` |
Expand All @@ -183,7 +183,7 @@ second.
### Image and diffusion service-level metrics

| Metric | Type | Labels | Description |
|--------|------|--------|-------------|
| -------- | ------ | -------- | ------------- |
| `vllm_omni:stage_gen_time_s` | Histogram | `model_name`, `stage`, `stage_type` | Stage submit to finished output; includes in-stage queueing |
| `vllm_omni:request_queue_wait_s` | Histogram | `model_name` | Orchestration-layer queue wait; a present zero is recorded |
| `vllm_omni:stage_waiting_requests` | Gauge | `model_name`, `stage` | Sum of the latest waiting snapshots across live replicas in the stage |
Expand All @@ -200,7 +200,7 @@ second.
Labels: `{model_name, stage, replica}`.

| Metric | Type | Description |
|--------|------|-------------|
| -------- | ------ | ------------- |
| `vllm_omni:diffusion_exec_s` | Histogram | Core diffusion-step execution time per request |
| `vllm_omni:diffusion_exec_per_step_s` | Histogram | Core diffusion execution divided by inference-step count |
| `vllm_omni:diffusion_preprocess_s` | Histogram | Diffusion input preprocessing time |
Expand All @@ -218,7 +218,7 @@ The optional breakdown and memory families are sparse by design: when an engine
Labels: `{model_name, stage, replica}` plus the listed extra label.

| Metric | Type | Extra label | Description |
|--------|------|-------------|-------------|
| -------- | ------ | ------------- | ------------- |
| `vllm_omni:audio_ttfp_s` | Histogram | — | Time from request arrival to first audio packet/frame |
| `vllm_omni:audio_duration_s` | Histogram | — | Audio content duration (`audio_frames / sample_rate`) |
| `vllm_omni:audio_rtf` | Histogram | — | Real-time factor `stage_gen_time_s / audio_duration_s` (SLO `< 1`); uses `RTF_BUCKETS` |
Expand All @@ -227,12 +227,29 @@ Labels: `{model_name, stage, replica}` plus the listed extra label.
| `vllm_omni:audio_continuity_ok_total` | Counter | `threshold_ms` | Incremented when the request's worst underrun stayed below `threshold_ms` |
| `vllm_omni:audio_skipped_requests_total` | Counter | `reason` | Silent-loss counter — code2wav rejected malformed codec input and returned `200 OK` with empty audio |

### Speech streaming (2)

Labels: `{model_name}` plus the listed extra label.

| Metric | Type | Extra label | Description |
| -------- | ------ | ------------- | ------------- |
| `vllm_omni:speech_stream_aborted_total` | Counter | `reason` | Interrupted audio generators, including before first PCM; reasons: `cancelled`, `closed`, `engine_dead`, `error` |
| `vllm_omni:speech_stream_completed_total` | Counter | — | Normally completed audio generators; does not confirm client receipt |

Scope: started Speech audio generators (raw/SSE/WebSocket), counted once per
generation; WebSocket counts each sentence. Chat, non-streaming Speech, and
failures before generator execution are excluded. Interrupted streams do not
emit continuity or underrun samples.

Interruption ratio: `aborted / (aborted + completed)`, using increments over the
same time window and summing all reasons per model.

### Cross-stage transfer (4)

Labels: `{model_name, from_stage, from_replica, to_stage, to_replica}`.

| Metric | Type | Description |
|--------|------|-------------|
| -------- | ------ | ------------- |
| `vllm_omni:transfer_size_bytes` | Histogram | Per-transfer payload size in bytes |
| `vllm_omni:transfer_tx_s` | Histogram | Sender-side time (serialize + submit to connector) |
| `vllm_omni:transfer_rx_s` | Histogram | Receiver-side time (recv + deserialize) |
Expand All @@ -245,8 +262,8 @@ After the wrap, every upstream `vllm:*` family — TTFT, ITL, TPOT, e2e latency,
## Naming Convention

- All time-bearing metrics use the `_s` suffix (values in seconds). Two bucket families are used:
- `SECONDS_BUCKETS` (0.05 s – 300 s) for e2e / generation / TTFP style values.
- `SECONDS_FAST_BUCKETS` (0.001 s – 60 s) for fine-grained cross-stage transfer and audio-underrun values that need millisecond-level resolution.
- `SECONDS_BUCKETS` (0.05 s – 300 s) for e2e / generation / TTFP style values.
- `SECONDS_FAST_BUCKETS` (0.001 s – 60 s) for fine-grained cross-stage transfer and audio-underrun values that need millisecond-level resolution.
- Counters use the `_total` suffix (auto-appended by `prometheus_client`).
- Sizes use the `_bytes` suffix.
- All omni-specific families are prefixed `vllm_omni:`. The upstream `unregister_vllm_metrics()` function is monkey-patched to a scoped version that still strips upstream `vllm:*` collectors (so multi-engine init within one process does not crash on duplicate registration) but preserves anything prefixed `vllm_omni:` / `vllm_omni`.
Expand Down
25 changes: 17 additions & 8 deletions docs/usage/metrics.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,7 +12,7 @@ curl http://localhost:8000/metrics
## Metric Namespaces

| Prefix | Source | Present when |
|--------|--------|--------------|
| -------- | -------- | -------------- |
| `vllm_omni:` | vLLM-Omni orchestrator / audio modality / cross-stage transfer | Pipeline-dependent |
| `vllm:` | Upstream vLLM engine, wrapped by `OmniPrometheusStatLogger` to expose `{stage, replica}` | Pipeline includes an LLM (AR) stage |
| `http_` / `process_` | Uvicorn / Python runtime | Always |
Expand All @@ -24,23 +24,23 @@ Defined in `vllm_omni/metrics/prometheus.py`. Track request lifecycle across the
### Request counts

| Metric | Type | Labels | Description |
|--------|------|--------|-------------|
| -------- | ------ | -------- | ------------- |
| `vllm_omni:num_requests_running` | Gauge | `model_name` | Pipeline-global in-flight requests (dispatched to engine, not yet finalized) |
| `vllm_omni:num_requests_waiting` | Gauge | `model_name` | Requests waiting in the Orchestrator queue |
| `vllm_omni:requests_success_total` | Counter | `model_name`, `finished_reason` | Total requests by completion reason. `finished_reason` ∈ {`stop`, `length`, `abort`, ...} mirroring upstream `vllm:request_success_total`; aborts cover client disconnect / cancellation paths in addition to upstream `FinishReason.ABORT` |

### Latency

| Metric | Type | Labels | Description |
|--------|------|--------|-------------|
| -------- | ------ | -------- | ------------- |
| `vllm_omni:e2e_request_latency_s` | Histogram | `model_name` | Pipeline-global end-to-end request latency in seconds |

## Audio Modality Metrics (`vllm_omni:`)

Emitted at request finalize, except for `audio_ttfp_s` (streaming-hook at the first audio packet) and `audio_underrun_s` / `audio_continuity_ok_total` (streaming finalize, after the chunk stream is exhausted). All carry `{model_name, stage, replica}` plus the listed extra label.
Emitted at request finalize, except for `audio_ttfp_s` (the first audio packet/frame from streaming Chat or Speech APIs) and `audio_underrun_s` / `audio_continuity_ok_total` (Chat or Speech streaming finalize, after the chunk stream is exhausted). Non-streaming Speech latency is represented by `e2e_request_latency_s`, not TTFP. All audio metrics carry `{model_name, stage, replica}` plus the listed extra label.

| Metric | Type | Extra label | Description |
|--------|------|-------------|-------------|
| -------- | ------ | ------------- | ------------- |
| `vllm_omni:audio_ttfp_s` | Histogram | — | Time from request arrival to first audio packet/frame |
| `vllm_omni:audio_duration_s` | Histogram | — | Audio content duration (`audio_frames / sample_rate`) |
| `vllm_omni:audio_rtf` | Histogram | — | Real-time factor (`stage_gen_time_s / audio_duration_s`); streaming TTS SLO red line `< 1`; uses `RTF_BUCKETS` |
Expand All @@ -50,13 +50,22 @@ Emitted at request finalize, except for `audio_ttfp_s` (streaming-hook at the fi
| `vllm_omni:audio_skipped_requests_total` | Counter | `reason` | Silent-loss counter — code2wav rejected malformed codec input and returned `200 OK` with empty audio |

The continuity math comes from `vllm_omni/benchmarks/audio_continuity.py::compute_continuity_stats` so the server-side observation aligns with the bench-side definition.
For Speech, continuity applies to raw audio, SSE, and WebSocket streaming. WAV headers and empty chunks are excluded; non-streaming Speech requests do not emit continuity samples.

Continuity rate for the default 100 ms threshold can be derived from the successful-continuity counter and the underrun Histogram's request count:

```promql
sum by (model_name, stage, replica) (rate(vllm_omni:audio_continuity_ok_total{threshold_ms="100"}[5m]))
/
sum by (model_name, stage, replica) (rate(vllm_omni:audio_underrun_s_count[5m]))
```

## Diffusion Metrics (`vllm_omni:`)

Per-request timing breakdowns for diffusion (image/video) stages. Emitted at request finalize when `stage_metrics.diffusion_metrics` is present. All carry `{model_name, stage, replica}`.

| Metric | Type | Description |
|--------|------|-------------|
| -------- | ------ | ------------- |
| `vllm_omni:diffusion_exec_s` | Histogram | DiT forward pass execution time per request in seconds |
| `vllm_omni:diffusion_exec_per_step_s` | Histogram | DiT forward pass execution time per denoising step in seconds (`exec_s / num_inference_steps`) |
| `vllm_omni:diffusion_preprocess_s` | Histogram | Diffusion input preprocessing time per request in seconds |
Expand All @@ -80,7 +89,7 @@ rate(vllm_omni:diffusion_preprocess_s_sum[5m])
Per-physical-transfer histograms tracking the data hop between adjacent stages. Labels `{model_name, from_stage, from_replica, to_stage, to_replica}` let dashboards attribute latency to specific replica edges. `from_replica` / `to_replica` are resolved from the orchestrator's sticky-routing binding (`stage_pool.get_bound_replica_id(request_id)`), so no extra plumbing through `TransferEdgeStats` is needed.

| Metric | Type | Description |
|--------|------|-------------|
| -------- | ------ | ------------- |
| `vllm_omni:transfer_size_bytes` | Histogram | Per-transfer payload size in bytes |
| `vllm_omni:transfer_tx_s` | Histogram | Sender-side time (serialize + submit to connector) |
| `vllm_omni:transfer_rx_s` | Histogram | Receiver-side time (recv + deserialize) |
Expand All @@ -106,7 +115,7 @@ For the full list of upstream metrics, see [the vLLM docs](https://github.com/vl
## Metric Availability by Pipeline Type

| Metric group | Multi-stage LLM (Qwen3-Omni) |
|---|---|
| --- | --- |
| `vllm_omni:` request tracking + latency | With `--log-stats` |
| `vllm_omni:` audio modality | With `--log-stats`, if pipeline has a talker stage |
| `vllm_omni:` transfer | With `--log-stats`, if pipeline has ≥ 2 stages |
Expand Down
3 changes: 2 additions & 1 deletion tests/entrypoints/openai_api/test_serving_speech.py
Original file line number Diff line number Diff line change
Expand Up @@ -3287,7 +3287,7 @@ def test_raw_stream_passes_tts_params_to_common_guard(self, streaming_app, monke
finalized_tts_params = {"_qwen3_tts_effective_max_tokens": [192]}
captured: dict = {}

async def prepare(_request, request_id=None):
async def prepare(_request, request_id=None, arrival_time=None):
return request_id, object(), finalized_tts_params

async def generate_chunks(
Expand All @@ -3296,6 +3296,7 @@ async def generate_chunks(
_response_format="pcm",
raw_request=None,
request_start_s=None,
request_arrival_ts=None,
include_sample_rate=False,
usage_acc=None,
tts_params=None,
Expand Down
Loading
Loading