Fix fast-model verification and align single-mode model selection - #126
Conversation
|
Half of this is landing as-is and half is superseded — details so you can rebase without guessing. The verify.sh hunk: #122 went in instead. Same false rejection, stricter fix. Both accept 4 and 8; #122 additionally checks the packed geometry the declared width implies against the shard header. On a fast dir whose config was edited to claim
A verifier that reads the declared width and never looks at what is on disk has given up the thing it is for, so I took the one that looks. This breaks your head fixture, concretely. An empty file has no safetensors header to read. The fixture needs a real one — an 8-byte little-endian length followed by a JSON header naming The rest is good and I want it. So: drop the Leaving #125 open until this lands, since the selection half of it is still live. |
750f4d4 to
f4a9d0c
Compare
|
Thanks for the detailed review and the concrete geometry example — agreed that #122 is the stronger verifier fix. I've rebased onto current main ( The head fixture now writes real safetensors headers rather than empty files. It covers matching int4/int8 geometry, both declared/stored-width mismatch directions, and missing/unsupported heads and scales. It also checks that rejection cases are reported failures rather than crashes. Validation: all 8 head cases and 4 selection cases pass under Linux/WSL. The unchanged upstream model-check block also passes against both installed model directories, including geometry and duplicate-shard checks. No model files were changed and no server was restarted; this was not a full install-verifier run. Pushed the rebased branch through |
# Conflicts: # single-user/start_qwen.sh
# Conflicts: # .github/workflows/docker-image.yml # bench/test_model_verification.py # docker-compose.yml # docker/prepare.sh # patches/dflash2-ngram-chains.patch # patches/dflash2-prewarm.patch # patches/dflash2-z-adaptive-emitted.patch # patches/hybrid-sw-block-promote.patch # patches/vllm-pr54282-draft-gumbel-salt.patch # prepare/quant_heads_stream.py # single-user/start_qwen.sh # verify.sh
f4a9d0c to
2203c00
Compare
|
Rebased onto
Head is now |
|
Merging. Everything I asked for is here, and I re-ran the evidence rather than reading it. On the box (CPU only — the 3090 is on a training job, which none of this needs):
The mismatch coverage in both directions is the part worth having in CI. That case — a config declaring one width over the other's tensors — is the one that separates a verifier that looks at the disk from one that reads the config and shrugs, and until now nothing tested it.
Two follow-ups, neither blocking:
Thanks for the rebase through the |
# Conflicts: # single-user/start_qwen.sh
# Conflicts: # .github/workflows/docker-image.yml # bench/test_model_verification.py # docker-compose.yml # docker/prepare.sh # patches/dflash2-ngram-chains.patch # patches/dflash2-prewarm.patch # patches/dflash2-z-adaptive-emitted.patch # patches/hybrid-sw-block-promote.patch # patches/vllm-pr54282-draft-gumbel-salt.patch # prepare/quant_heads_stream.py # single-user/start_qwen.sh # verify.sh
Accept supported packed int4 and int8 heads and select the same model before Docker verification and launch. Preserve explicit overrides and batch defaults. Add CPU regression checks for head formats and selection. Fixes syv-ai#125
Use header-only shards, reject declared/stored bit-width mismatches in both directions, and retain model-selection regression coverage. Keep upstream syv-ai#122 verifier unchanged.
2203c00 to
5fad43d
Compare
…#113) bench/warmup.sh re-derived the model dir (fast variant when present, else the base dir) with a comment saying it mirrors single-user/start_qwen.sh. syv-ai#126 extracts exactly that resolution into single-user/select_model.sh and points the launcher and the Docker gate at it, which leaves this copy as the only one left to drift. Source the same file here. Equivalence, not approximation: the sourced resolver resolves identically to the inline block in every case - fast present, base only, and MODEL unset/empty/explicit - checked against the old block on identical fixtures.
|
Rebased onto current Both conflicts were adjacent additions rather than competing edits:
No content changed. Re-run on the rebased tip rather than carried over: #136 is rebased on top and is now the single commit you asked for. |
|
Verified on the reference box tonight, merged onto today's And #125's first defect against the real files, with The only thing holding the merge is on my side: the branch edits |
|
Merging. Re-verified on top of the vLLM 0.29 pin (#148, merged just now) rather than on the old main: this branch plus #136 merged onto |
…yv-ai#136) into main Conflicts: docs/vllm-0.29.md takes syv's (a superset: the maintainer's greedy-divergence section); the image workflow keeps this fork's identity (Dockerfile.fork into ghcr.io/cpuchip/...); PATCHES.md keeps this fork's header and marlin-int8-asym-zp row, which describe the fork-exported file this main carries (hunks identical to syv's hand-cut file; syv's series replays to fork tip 291980422 with 0 differing files). Batch GPU_UTIL 0.95 comes in from syv-ai#148.
Closes #125
Scope after review
Rebased onto upstream main at
d1df6dc(the README/docs restructure), which includes #122. This PR no longer changesverify.sh: upstream's packed-geometry and shard-header validation is retained unchanged.Changes
Share single-user model selection between the launcher and Docker entrypoint, resolving after preparation and exporting MODEL before verification.
Preserve explicit MODEL overrides, base fallback, batch defaults, and standalone verifier defaults.
Keep the GPU-free model-verification CI job, entrypoint selection fixture, and Docker documentation.
Replace empty safetensors fixtures with header-only shards: 8-byte little-endian header length, padded JSON header, tensor shapes, dtypes, and offsets. No tensor payload is required by these header-only checks.
Cover valid int4/int8 geometry, declared int8 over int4 tensors and the reverse, unsupported width, missing group, missing packed head, and missing scale. Rejections must be verifier failures, not Python crashes.
Validation
wsl python3 /mnt/g/dev/qwen38-27b-rtx3090/bench/test_model_verification.py: PASS (2 test methods; 8 head cases and 4 selection cases).Exact current upstream model-check heredoc executed with Python against both installed base and fast model directories: exit 0 for both, including packed geometry and duplicate-shard checks.
bash -npassed for verify.sh, docker/entrypoint.sh, single-user/start_qwen.sh, and single-user/select_model.sh.Selection fixture executes the actual entrypoint in a temporary root with verifier/launcher path recorders. Covers base fallback, fast auto-selection, explicit path containing a space, and batch default.
Limits
No full verify.sh run, image rebuild, or GPU server restart. The prior container was already stopped, so installed-model headers were checked directly through its host bind mount without modifying model files. Run the fixture suite on Linux/WSL (as CI does); native Windows Python launching WSL bash does not reliably forward its environment.
Rebased onto current
main(2026-09-22)The branch was conflicting after #138 (
resolve_config.sh) landed a few lines from theMODELblock, and after #134 added adocs/docker.mdbullet next to the one this PR extends. Rebased ontomainat8b4dccd; both conflicts were adjacent additions, not competing edits:single-user/start_qwen.sh: launchers: one validated effective-configuration resolver (F13 / backlog 6) #138's validated-resolver block is kept as-is, and the three-lineMODELdefault below it becomessource "$REPO/single-user/select_model.sh".docs/docker.md: the selection paragraph stays attached to thedocker compose run --rm single verifybullet; prepare: serialise concurrent model preparation with flock #134's "Concurrent prepares are serialised" bullet follows it.No content changed in the rebase. Re-run on the rebased tip, not carried over:
bench/test_model_verification.py— 2 tests, 12 cases, pass (WSL, Python 3.12.3).git-applyandmodel-verificationboth green; the PR is 6 files,MERGEABLE / CLEAN.