Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
93 changes: 75 additions & 18 deletions docs/METRICS.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,11 +12,15 @@ The repository includes a ready-to-run monitoring stack:
- `prometheus.yaml` scrapes the vLLM server from the Docker host every five
seconds.
- `prometheus-rules.yaml` records one-minute averages for prompt, generation,
and total token throughput.
and total token throughput, plus speculative-decoding metrics (acceptance
rate, mean accepted length, draft and accepted token rates, and
per-position acceptance).
- `grafana/provisioning/` configures Prometheus as Grafana's default data source
and loads dashboards from disk.
- `grafana/dashboards/vllm-throughput.json` defines the default **vLLM
Throughput** dashboard.
- `grafana/dashboards/vllm-spec-decode.json` defines the **vLLM Speculative
Decoding** dashboard.

The default configuration assumes vLLM is listening on port `8000` on the
Docker host. Confirm the endpoint first:
Expand Down Expand Up @@ -88,16 +92,55 @@ imported into Grafana:

## Speculative decoding and MTP panels

The general vLLM dashboard may not include speculative-decoding panels. Add
Grafana panels with the following PromQL queries.
The provisioned **vLLM Speculative Decoding** dashboard
(`grafana/dashboards/vllm-spec-decode.json`) covers these metrics out of the
box. It appears in the **vLLM** folder and queries the `vllm:spec_decode_*`
recording rules from the `vllm-spec-decode` group in `prometheus-rules.yaml`:

- **Draft acceptance rate** — fraction of draft tokens accepted
(`vllm:spec_decode_acceptance_rate:rate1m`)
- **Mean accepted length** — tokens emitted per verification step, including
the bonus token (`vllm:spec_decode_mean_accepted_length:rate1m`); with N
speculative tokens it ranges from 1.0 to N+1
- **Draft tokens / second** (`vllm:spec_decode_draft_tokens_per_second:rate1m`)
- **Draft vs accepted tokens / second** — the gap between the two series is
wasted draft compute
- **Acceptance rate by draft position** — per-position quality of the MTP
cascade (`vllm:spec_decode_acceptance_rate_by_pos:rate1m`)

The stat panels query the recorded metrics as instant values. The ratio
rules (acceptance rate, mean accepted length, per-position) evaluate to no
value once traffic stops, so those panels show **No value** shortly after
the last request, when the last recorded sample leaves Prometheus'
five-minute lookback window; this avoids a misleading 100% acceptance rate
at idle. The token-rate rules have no such guard and keep recording zero
while vLLM is running, so the draft and accepted token panels show 0 at
idle.

If you add or replace dashboard files on disk, restart the Grafana container
for the file provider to reload them:

```bash
docker compose -f docker-compose.metrics.yaml restart grafana
```

If you prefer to build your own panels directly on the raw counters (for
example with a longer rate window), the following PromQL queries are
equivalent to the recording rules:

### Draft-token acceptance rate

```promql
100 *
sum(rate(vllm:spec_decode_num_accepted_tokens_total[5m]))
/
sum(rate(vllm:spec_decode_num_draft_tokens_total[5m]))
(
100 *
sum(rate(vllm:spec_decode_num_accepted_tokens_total[5m]))
/
sum(rate(vllm:spec_decode_num_draft_tokens_total[5m]))
)
and on()
(
sum(rate(vllm:spec_decode_num_draft_tokens_total[5m])) > 0
)
```

Use a Gauge or Time series visualization with the unit set to percent.
Expand All @@ -108,10 +151,16 @@ This convention includes the target/bonus token emitted by a verification
step:

```promql
1 +
sum(rate(vllm:spec_decode_num_accepted_tokens_total[5m]))
/
sum(rate(vllm:spec_decode_num_drafts_total[5m]))
(
1 +
sum(rate(vllm:spec_decode_num_accepted_tokens_total[5m]))
/
sum(rate(vllm:spec_decode_num_drafts_total[5m]))
)
and on()
(
sum(rate(vllm:spec_decode_num_drafts_total[5m])) > 0
)
```

### Draft and accepted tokens per second
Expand All @@ -131,18 +180,26 @@ sum(rate(vllm:spec_decode_num_accepted_tokens_total[5m]))
### Acceptance rate by draft position

```promql
100 *
sum by (position) (
rate(vllm:spec_decode_num_accepted_tokens_per_pos_total[5m])
(
100 *
sum by (position) (
rate(vllm:spec_decode_num_accepted_tokens_per_pos_total[5m])
)
/
scalar(
sum(rate(vllm:spec_decode_num_drafts_total[5m]))
)
)
/
scalar(
and on()
(
sum(rate(vllm:spec_decode_num_drafts_total[5m]))
> 0
)
```

Use a Bar chart or Time series visualization and set the legend to
`Position {{position}}`.
The provisioned dashboard renders this metric as a Bar gauge with one bar
per draft position. If you build your own panel, use a Bar gauge or Time
series visualization and set the legend to `Position {{position}}`.

## Generate test traffic

Expand Down
Loading