patches: bench-probe-errors — the /tokenize alignment probe names its failure and sends the key - #165
Conversation
… and sends the key
vllm bench serve's tokenizer-alignment probe collapsed every failure into
one wrong message ("/tokenize unavailable") and sent no Authorization header,
so a keyed server 401'd it even with a valid model name. The patch sends the
same Bearer the benchmark requests carry and names the real cause: 404
(no route, or the model name was rejected — /v1/models lists the served
names), 401 (key required), unreachable, timeout, or the raw exception.
Independent of every other patch in the series (benchmarks/serve.py is
untouched by them), so it rides at the end of patches/series. Cut from the
extended cpuchip/vllm qwen38/0.28 branch, topic commit [qwen38]
bench-probe-errors; kind: fix, retires when upstream takes it.
c512beb to
5c8755d
Compare
fetch_spec_decode_metrics and fetch_diffusion_metrics GET /metrics with no headers and return None on any non-200, so on a server bound with --api-key the 401 is indistinguishable from 'speculative decoding is off' and the benchmark's whole spec-decode block (acceptance length, accepted/drafted, per-position acceptance) silently disappears from every run. Both fetchers now take the headers the benchmark already built; the four call sites in benchmark() pass extra_headers, the same dict the request path uses. Unkeyed servers and the None-on-missing-metrics contract are unchanged. Observed on hyperqwen d6e094a3 with a 48-char VLLM_API_KEY: GET /metrics keyless 401, with the key 200, /health 200 keyless; the bench client's 401s share an ephemeral port with its own POST /v1/completions, which is what identifies it (rather than run_benchmarks.sh's curl) as the caller.
|
Amended: the patch now also fixes the /metrics scrapes.
Both fetchers take Runtime evidence (hyperqwen The 401s in the server log share an ephemeral port with the bench client's own Verification of the regenerated file: Disclosure: this repo has no |
… header Threading extra_headers into fetch_spec_decode_metrics and fetch_diffusion_metrics was not enough. That dict only ever carries the --header pairs, so on a server bound with --api-key and no --header it is empty, and all four /metrics scrapes still came back 401 -- which both fetchers return as None, indistinguishable from 'speculative decoding is off'. The spec-decode block then vanished from every run on a keyed server, which is the exact symptom the patch set out to fix. Both fetchers now build their scrape headers the way the /tokenize probe does: start from whatever the caller passed, then add the Bearer from OPENAI_API_KEY unless one is already set. Observed on a keyed server: repeated 'GET /metrics 401 Unauthorized' from a single long-lived bench-client session, while the harness's own keyed curl scrapes on neighbouring ports returned 200.
|
Merging this, with an independent check of the series rather than of the file alone. Applied the whole The premise checks out in the pinned source too, which is the part worth stating since neither of us has booted a rebuilt image: The amended shape is the right one. On the regeneration disclosure: noted, and it is the right thing to have said. Reverse-applying to recover pristine upstream and re-diffing is content-equivalent — The remaining gap is the same for all five: nobody has booted a server from a rebuilt image carrying them. The image rebuilds on this merge, so that is coming; if anything in the series misbehaves at boot I will revert rather than patch forward. |
…, the four serving and bench patches from syv-ai#165-syv-ai#168, the triton message fix)
What
Adds
patches/bench-probe-errors.patch, itspatches/seriesline and itsPATCHES.mdrow. The patch changesvllm bench serve's tokenizer-alignment probe: it sends theAuthorizationheader the benchmark requests already send, and it classifies its failure instead of printing one line for every cause.Upstream: vllm-project/vllm#58024.
Why
In 0.28.0, the probe posts
args.modelverbatim as the request model and swallows every exception into one line:bench/run_benchmarks.sh:25passes--model $MODEL(a checkpoint path) with--served-model-name qwen3.8-27b, so the probe asks for a name the server does not serve and gets a 404 from the model check. The server is healthy and the benchmark runs, but the warning says the endpoint is unavailable. In my results tree, 314 of 602 benchmark logs carry that line, split exactly alongdataset_name='random'versus'custom'.The probe also sends no key, so on a keyed server alignment is skipped in every run.
#170 fixes the caller in this repo. This patch fixes the message for everyone who reads it.
Runtime evidence
Patched function loaded from the fork commit, driven against a local
aiohttpserver (CPU only):Verification
patch integrity(the workflow'sgit-applyjob) applies the whole series to a pristinevllm-project/vllmcheckout at the pin: passes on this PR.patch -p1 --fuzz 0 --dry-runof this file against the installed tree inghcr.io/syv-ai/hyperqwen:latest(06150174): applies, no fuzz.verify.shneeds no new entry: its loop readspatches/series, andpatches/_check_applied.pyparses the patch file itself.Not done: I have not rebuilt the image and restarted a server on this patch, so the runtime evidence above comes from the code path, not from a rebuilt server.