diff --git a/.env.example b/.env.example index bdfea6cc..f8f1c2f1 100644 --- a/.env.example +++ b/.env.example @@ -16,13 +16,16 @@ PREFIX_CACHE=1 # Typical use: long docs, code edits, RAG with quoting, prompt-copy workloads. # CTX=long -# Use when you need about 150k context and are willing to trade speed for more tokens. -# Slower than the default speed mode; good for longer chats or long prompts, not the fastest profile. +# About 150k context, at ~95-100 tok/s instead of 127. Set SPEC=mtp with it: DFlash2 +# past 64k is worth it only for context reproduction and loses to SPEC=mtp CTX=long +# roughly 2:1 on everything else. # CTX=huge -# Use only after running: bash kvarn/install.sh in a Linux shell with bash available. -# This is the maximum-context KVarN profile for 200k+ context; it is slower and not the speed winner. -# Best for genuinely long prompts where the context window matters more than decode rate. +# The KVarN 4/2-bit KV cache: 240k context (268k tokens of pool), and it combines with +# the SPEC=dflash2 above -- keep PREFIX_CACHE=1, since the first turn over a long document +# is the expensive one. Slower than the default profile: 67 tok/s across mixed tasks, +# 164 while reproducing. The container already has KVarN built in; a venv install needs +# bash kvarn/install.sh first. # The launcher's own default is SPEC=mtp; this template sets dflash2 above because it is the # faster single-user profile. The line below is enabled for Docker Desktop/WSL2, which needs diff --git a/.github/FUNDING.yml b/.github/FUNDING.yml index 7f375ea2..040d9bfb 100644 --- a/.github/FUNDING.yml +++ b/.github/FUNDING.yml @@ -1,4 +1,4 @@ # Renders the "Sponsor" button at the top of the repo. # One RTX 3090 at 250 W is the whole hardware budget behind this repo's numbers; -# see "Help us test more hardware" in README.md for what funding actually buys. +# see "Field tests wanted" in README.md for what funding actually buys. ko_fi: mhenrichsen diff --git a/.github/ISSUE_TEMPLATE/config.yml b/.github/ISSUE_TEMPLATE/config.yml new file mode 100644 index 00000000..411f0fbb --- /dev/null +++ b/.github/ISSUE_TEMPLATE/config.yml @@ -0,0 +1,8 @@ +blank_issues_enabled: true +contact_links: + - name: Offer compute or hardware + url: https://github.com/syv-ai/HyperQwen/issues/new?title=compute%20offer + about: Credits, a box we can SSH into, or an architecture you can lend time on. + - name: Fund rented GPU hours + url: https://ko-fi.com/mhenrichsen + about: Goes to cards this project does not own; the runs get written up here. diff --git a/.github/ISSUE_TEMPLATE/field-report.yml b/.github/ISSUE_TEMPLATE/field-report.yml new file mode 100644 index 00000000..81e0b254 --- /dev/null +++ b/.github/ISSUE_TEMPLATE/field-report.yml @@ -0,0 +1,92 @@ +name: Field report (a run on your hardware) +description: Numbers from a card this repo cannot test on. Partial runs welcome. +title: "Field report: — " +labels: ["field-report"] +body: + - type: markdown + attributes: + value: | + The protocol is in the README under + [Field tests](https://github.com/syv-ai/HyperQwen#field-tests-wanted). + Short version: start the server, run the harness **twice**, keep the + second run, then check quality. Anything less than the full protocol is + still worth posting — fill in what you have and say what you skipped. + - type: input + id: card + attributes: + label: Card, and how many + description: Include the architecture if you know it (sm80/86/89/90/120). + placeholder: 1x RTX A4000 16 GB (sm86) + validations: + required: true + - type: input + id: power + attributes: + label: Power limit during the run + description: | + `nvidia-smi -q -d POWER | grep "Power Limit"`. This is the single most + common reason two runs disagree — a 3090 at 200 W measures its power cap, + not this stack (33% slower than at 250 W). + placeholder: 250 W (nvidia-smi -pl 250) + validations: + required: true + - type: input + id: stack + attributes: + label: Driver, CUDA, OS, and container or venv + placeholder: 580.65, CUDA 13.0, Ubuntu 24.04 bare metal / Docker / WSL2 + validations: + required: true + - type: input + id: rev + attributes: + label: Commit you ran + description: "`git rev-parse --short HEAD`, or the image tag." + placeholder: 9b47388 + validations: + required: true + - type: dropdown + id: setup + attributes: + label: Which setup + description: The letters from the README decision tree. + options: + - A — batch + - B — single, default (SPEC=dflash2 PREFIX_CACHE=1) + - C — reproduction (DFLASH_TOKENS=15) + - D — long context (SPEC=mtp CTX=long) + - E — huge context (CTX=huge) + - something else (describe below) + validations: + required: true + - type: textarea + id: rows + attributes: + label: Harness output — the second run + description: | + The `ROW` lines from `bash bench/run_benchmarks.sh `. The first run + after a start includes JIT warmup and reads 30-50% low, so post the + second. If you measured with your own client instead, say so here and + give the prompt shape, output length and how you define tok/s — it will + be listed as an independent report rather than a table row. + render: text + validations: + required: true + - type: textarea + id: quality + attributes: + label: Quality check + description: | + `python bench/quality_battery.py --gsm-only` (about 5 minutes). + A fast server that emits garbage is worth nothing, and a quantization + or kernel bug on new hardware shows up here first. Every shipped config + reads 95.0-96.5% on this harness. + render: text + - type: textarea + id: notes + attributes: + label: Anything that did not work + description: | + Build failures, wrong-looking output, things the docs got wrong on your + platform. This is often the more useful half of the report — several + entries in docs/gotchas.md started as one line in an issue like this. diff --git a/README.md b/README.md index 971fe651..d02ce01c 100644 --- a/README.md +++ b/README.md @@ -31,8 +31,8 @@ speculation that drafts out of the prompt. Most of that is not specific to one checkpoint or one card; it is just that one checkpoint and one card are what has been measured to death. See [Roadmap](#roadmap-more-qwen-models-more-cards) for what is and is not portable yet, and -[Help us test more hardware](#help-us-test-more-hardware) if you have credits to -spare on a GPU we have never touched. +[Field tests wanted](#field-tests-wanted) if you have half an hour on a GPU we +have never touched. ## Quick start @@ -55,7 +55,42 @@ into `./models`), then serves on `:18020`. One GPU runs one mode at a time. [docs/docker.md](docs/docker.md#plain-docker-no-compose) · [docs/install.md](docs/install.md) -### Which mode +## Which setup do I want? + +Three questions, and the answers are the whole of your `.env`. + +```mermaid +flowchart TD + Q1{"Who is sending
the requests?"} + Q1 -->|"an API backend, a pipeline,
8+ at once"| A["A — batch"] + Q1 -->|"one or a few
people chatting"| Q2{"Longest prompt
you will send?"} + Q2 -->|"under 64k tokens"| Q3{"Do the answers quote
the prompt back?"} + Q2 -->|"up to 150k"| D["D — long context"] + Q2 -->|"200k and beyond"| E["E — huge context"] + Q3 -->|"no — normal chat"| B["B — single, default"] + Q3 -->|"yes — code edits, RAG,
rewrites, translation"| C["C — reproduction"] +``` + +Starting from the `.env` you copied in Quick start, which already ships **B**: + +| | change in `.env` | start with | what you get | +|---|---|---|---| +| **A** | nothing | `--profile batch` | ~1,035 tok/s aggregate at 64 concurrent | +| **B** | nothing | `--profile single` | 127 tok/s single stream, 64k context | +| **C** | add `DFLASH_TOKENS=15` | `--profile single` | 381 tok/s while quoting, at 4 slots and 56k | +| **D** | `SPEC=mtp` and `CTX=long` | `--profile single` | 150k context, ~95-100 tok/s | +| **E** | add `CTX=huge` | `--profile single` | 240k context, 67 tok/s mixed and 164 while quoting | + +Torn between B and C: B is the safe default, and C only pays off when the output +really does repeat the input — it trades request slots and context for it. D drops +back to MTP speculation on purpose, because DFlash2 past 64k is worth it only for +reproduction and loses to `SPEC=mtp CTX=long` about 2:1 on everything else; at E +the KVarN cache buys the context instead, so DFlash2 stays on. A needs no edit at +all — batch ignores `SPEC` and says so on startup. Every knob, and what each costs: +[single-user/README.md](single-user/README.md#if-you-are-the-only-user-do-this) · +[batch/README.md](batch/README.md). + +### The numbers behind that | | **batch** → [batch/](batch/) | **single** → [single-user/](single-user/) | |---|---|---| @@ -72,24 +107,6 @@ few of ([measurement](docs/long-context.md)). Both modes share one install. [Full prefill matrix](batch/README.md#prefill) · [how each number was won](docs/optimizations.md). -## If you are the only user - -The shipped default is conservative — MTP speculation, 8 request slots, 64k -context. If the card is yours alone, two settings are worth more than every -other knob in this repo put together: - -```bash -printf 'SPEC=dflash2\nPREFIX_CACHE=1\n' >> .env -docker compose --profile single up -d -``` - -`SPEC=dflash2` proposes 7 tokens in one pass instead of 4 chained ones; -`PREFIX_CACHE=1` keeps the document you already sent. Add `DFLASH_TOKENS=15` if -your answers quote your prompts — that is the 381 tok/s row above. - -Full numbers, the venv equivalent, and what each setting costs: -[single-user/README.md](single-user/README.md#if-you-are-the-only-user-do-this). - ## Roadmap: more Qwen models, more cards The name changed because the scope did. This started as "Qwen3.8-27B on one RTX @@ -127,46 +144,60 @@ same treatment rather than a shrug; a larger one for 48 GB; the EXL3 route in path. If you want a specific model or card prioritized, say so in an issue — the order is mostly driven by who turns up with a reproduction. -## Help us test more hardware - -**We are looking for Runpod or Vast.ai credits.** One RTX 3090 at 250 W is the -entire hardware budget behind every number in this README, and it is shared with -training jobs. That constraint shows: the tables above lean on community -reproductions for every card that is not a 3090, an sm80 speculation bug has sat -open because nobody here owns an sm80 card to debug it on, and each architecture -sweep is scheduled around whatever else needs the GPU that week. - -Rented compute converts directly into things this repo does not currently have: - -- **An architecture matrix that is measured rather than collected.** sm80, sm89, - sm90, sm120 and multi-GPU under one harness, one power protocol, one set of - prompts — instead of a table where every row came from a different person's - client and cannot honestly be compared to the row above it. -- **Fixing the bugs we cannot reproduce.** - [#98](https://github.com/syv-ai/HyperQwen/issues/98) and - [#72](https://github.com/syv-ai/HyperQwen/issues/72) are sm80 faults diagnosed - entirely from other people's logs. A few hours on a rented A100 probably closes - both. -- **Porting faster.** The vLLM 0.29.0 port - ([#106](https://github.com/syv-ai/HyperQwen/issues/106)) is blocked on - packaging details that need a machine to bisect on, while the one card here is - the card serving production. -- **More models.** Every new checkpoint needs its draft vocabulary calibrated and - its drafter trained — GPU-hours, not cleverness. - -What you get: results published here as reproductions with the raw harness -output, credit in this section and on the runs themselves, and — if you would -rather have it than the credit — an honest writeup of what your hardware does -badly, which is usually the more useful half. - -If you can help, open an issue titled "compute offer" and we will take it from -there. Small amounts are genuinely useful: a single day on one unfamiliar card -has historically been worth more to this project than a month on a familiar one. - -Money works too, if that is easier than credits: -[ko-fi.com/mhenrichsen](https://ko-fi.com/mhenrichsen) — it goes to rented GPU -hours and the runs get written up here like any other reproduction. +## Field tests wanted + +Every number above is one RTX 3090 at 250 W, shared with training jobs. Every +row for any other card came from someone else running the harness — so the ask +here is a protocol rather than "try it and let us know". + +**About 30 minutes on one GPU**, most of it unattended. The harness lives in +the container, so it runs through `compose exec` — on a venv install drop that +prefix and use `python` instead of `venv/bin/python`: + +```bash +sudo nvidia-smi -pl 250 # your card's reference wattage; say which +docker compose --profile single up -d # first start also prepares the model + +docker compose exec single bash bench/run_benchmarks.sh single # discard: a first run after a start reads 30-50% low +docker compose exec single bash bench/run_benchmarks.sh single # keep this one — its ROW lines are the report + +docker compose exec single venv/bin/hf download openai/gsm8k --repo-type dataset \ + --include "main/test-*" --local-dir bench/quality-data/gsm8k +docker compose exec single venv/bin/python bench/quality_battery.py mycard --gsm-only +``` + +Then open a +[field report](https://github.com/syv-ai/HyperQwen/issues/new?template=field-report.yml). +The form asks for the six things that make two runs comparable — card, power +cap, driver, OS or container, commit, which setup letter — because without them +a number cannot honestly be put next to the row above it. Ran something +narrower, or with your own client? Post it anyway and say what you skipped; it +gets listed as an independent report instead of a table row, which is what +happened to most of +[docs/reproductions/](docs/reproductions/README.md#results-from-other-hardware). + +**Most wanted, in order.** Several of these already have a thread with someone's +hardware in it — check before you duplicate, and add to theirs if it matches: +| | why it is worth your GPU hour | +|---|---| +| **sm90** — H100, H200 | The only architecture here with no datapoint at all. | +| **sm80** — A100, A30, CMP 170HX | Two owners are mid-bisect on a speculation fault that only their cards produce ([#98](https://github.com/syv-ai/HyperQwen/issues/98), [#72](https://github.com/syv-ai/HyperQwen/issues/72)). A third sm80 box would separate the card from the build. | +| **Four Ampere cards** | Four sm120 cards are measured ([#105](https://github.com/syv-ai/HyperQwen/issues/105)); nobody has run four 3090s, and two of them already give *less* aggregate throughput than one ([#135](https://github.com/syv-ai/HyperQwen/issues/135)). | +| **12 GB cards, on the harness** | Two 3060s do serve this model ([#68](https://github.com/syv-ai/HyperQwen/issues/68)), reported with their owners' own clients — so the numbers cannot be set against the rows above. A harness run on that pair is most of what decides whether smaller Qwen checkpoints are worth preparing. | +| **vLLM 0.29.0** ([#106](https://github.com/syv-ai/HyperQwen/issues/106)) | Ready on a fork, blocked on packaging details that need a machine to bisect on. | + +**What you get:** the run published in +[docs/reproductions/](docs/reproductions/README.md) with your raw output and +credit, the gotchas your platform exposes written into +[docs/gotchas.md](docs/gotchas.md), and an honest account of what your hardware +does badly — historically the more useful half. + +**Rather give hardware or money than time?** A box we can reach, or credits: +open an issue titled "compute offer". Otherwise +[ko-fi.com/mhenrichsen](https://ko-fi.com/mhenrichsen) buys hours on cards this +project does not own, and those runs get written up here like any other +reproduction. ## Documentation diff --git a/docs/reproductions/README.md b/docs/reproductions/README.md index a588a3a9..4e6cc1d9 100644 --- a/docs/reproductions/README.md +++ b/docs/reproductions/README.md @@ -62,6 +62,37 @@ other only loosely, and not rows for the table above: attention adding +1.8% at 21k and +6.3% at 66k on top, against this repo's +2.7% at 16k and +5.3% at 51k. Plus the power-limit ladder quoted above — [#62](https://github.com/syv-ai/HyperQwen/issues/62). +- **4x RTX 5060 Ti 16 GB (sm120, TP4)**: the canonical harness at TP4, and the + cleanest multi-arm A/B this project has received --- four arms on one box, + `PREFIX_CACHE=1` throughout, one variable apart. C1 greedy decode 72.8 + (`SPEC=mtp` k=3, fp8/FlashInfer) -> 137.6 (`SPEC=mtp`, int8 KV on + `TRITON_ATTN`) -> 154.3 (`SPEC=dflash2` k=7, same KV). So the headline 2.37x + from the first round is **1.89x KV dtype and attention backend, 1.12x + drafter**, and the drafter's advantage is gone by C4. `DFLASH_TOKENS=15` at + TP4 is a clear loss (154.3 -> 89.4 at C1, TTFT 3.3x at C8), which is the TP4 + half the launcher's keep-it-at-7 warning was missing. Also the source of the + `curand` headers gotcha and the `NCCL_P2P_LEVEL=SYS` note --- + [#105](https://github.com/syv-ai/HyperQwen/issues/105). +- **2x RTX 3060 12 GB**: 24 GB of VRAM in two 12 GB cards, which forces a + tensor-parallel split the single-card path never takes. Two independent boxes + in the thread, both on the Docker path at `SPEC=mtp CTX=long`, one reporting + ~50 tok/s at 49k and ~35 at 80k context at a 130 W per-card cap. Measured with + their own clients, so not comparable to the table above --- a harness run on + this pair is the open ask in + [#68](https://github.com/syv-ai/HyperQwen/issues/68). +- **3x RTX 3090**: confirms TP=3 is refused by the checkpoint rather than by this + repo (4 KV heads, 64 layers: neither TP=3 nor an even PP=3 split exists), so + the third card idles under `--tensor-parallel-size 2` by construction. The + thread's workaround --- a second independent engine on the spare card, with + `VLLM_OFFLOAD_KEEP_SHM=1` on both containers --- is the shape a triple-card + owner wants: [#104](https://github.com/syv-ai/HyperQwen/issues/104). +- **2x RTX 3090 (TP=2), native Ubuntu**: batch harness at 917 decode / 861 e2e + at 64 concurrent against the single card's ~1,035 / 948, and 60.1 tok/s at C1 + against 46 --- two cards giving less aggregate throughput and ~30% more + single-stream, the same shape as [#40](https://github.com/syv-ai/HyperQwen/issues/40). + Run with `KV=kvarn VISION=1` rather than the reference batch profile, so + indicative rather than a controlled row: + [#135](https://github.com/syv-ai/HyperQwen/issues/135). - **Dual-GPU reports**: the controlled 1-vs-2×3090 A/B in [#40](https://github.com/syv-ai/HyperQwen/issues/40) (+16–35%, 161.6 C1 greedy at 275 W, PCIe x8 without NVLink; independently reproduced diff --git a/single-user/README.md b/single-user/README.md index 1236945b..000b8e49 100644 --- a/single-user/README.md +++ b/single-user/README.md @@ -128,8 +128,10 @@ happens when the streams are big. Where it is *not* the better choice: 5 sliding-window layers from padding the target's attention/GDN layers (105 → 78 KB of pool per token; without it this mode caps out at ~40k), and the V2 runner's profiled activation peak swings ~1 GiB between starts, which makes a - utilization-based setting non-deterministic. `CTX=huge` stays MTP (the script falls back - with a message). `CTX=long` doubles the context — 138,696 tokens at `DFLASH_TOKENS=7`, + utilization-based setting non-deterministic. `CTX=huge` combines with this drafter — + 268k tokens of pool at 245760 max-model-len over the KVarN cache + ([docs/long-context.md](../docs/long-context.md#dflash2-at-240k-ctxhuge-kvarn-also-combines-with-specdflash2)); + any other `CTX` falls back to MTP with a message. `CTX=long` doubles the context — 138,696 tokens at `DFLASH_TOKENS=7`, 114,224 at 15 — by moving to an `int8_per_token_head` cache on the Triton backend; it is worth it only for context reproduction, and `SPEC=mtp CTX=long` beats it about 2:1 on everything else. See [docs/long-context.md](../docs/long-context.md#dflash2-past-64k-specdflash2-ctxlong).