Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
13 changes: 8 additions & 5 deletions .env.example
Original file line number Diff line number Diff line change
Expand Up @@ -16,13 +16,16 @@ PREFIX_CACHE=1
# Typical use: long docs, code edits, RAG with quoting, prompt-copy workloads.

# CTX=long
# Use when you need about 150k context and are willing to trade speed for more tokens.
# Slower than the default speed mode; good for longer chats or long prompts, not the fastest profile.
# About 150k context, at ~95-100 tok/s instead of 127. Set SPEC=mtp with it: DFlash2
# past 64k is worth it only for context reproduction and loses to SPEC=mtp CTX=long
# roughly 2:1 on everything else.

# CTX=huge
# Use only after running: bash kvarn/install.sh in a Linux shell with bash available.
# This is the maximum-context KVarN profile for 200k+ context; it is slower and not the speed winner.
# Best for genuinely long prompts where the context window matters more than decode rate.
# The KVarN 4/2-bit KV cache: 240k context (268k tokens of pool), and it combines with
# the SPEC=dflash2 above -- keep PREFIX_CACHE=1, since the first turn over a long document
# is the expensive one. Slower than the default profile: 67 tok/s across mixed tasks,
# 164 while reproducing. The container already has KVarN built in; a venv install needs
# bash kvarn/install.sh first.

# The launcher's own default is SPEC=mtp; this template sets dflash2 above because it is the
# faster single-user profile. The line below is enabled for Docker Desktop/WSL2, which needs
Expand Down
2 changes: 1 addition & 1 deletion .github/FUNDING.yml
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
# Renders the "Sponsor" button at the top of the repo.
# One RTX 3090 at 250 W is the whole hardware budget behind this repo's numbers;
# see "Help us test more hardware" in README.md for what funding actually buys.
# see "Field tests wanted" in README.md for what funding actually buys.
ko_fi: mhenrichsen
8 changes: 8 additions & 0 deletions .github/ISSUE_TEMPLATE/config.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,8 @@
blank_issues_enabled: true
contact_links:
- name: Offer compute or hardware
url: https://github.com/syv-ai/HyperQwen/issues/new?title=compute%20offer
about: Credits, a box we can SSH into, or an architecture you can lend time on.
- name: Fund rented GPU hours
url: https://ko-fi.com/mhenrichsen
about: Goes to cards this project does not own; the runs get written up here.
92 changes: 92 additions & 0 deletions .github/ISSUE_TEMPLATE/field-report.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,92 @@
name: Field report (a run on your hardware)
description: Numbers from a card this repo cannot test on. Partial runs welcome.
title: "Field report: <card> — <what you ran>"
labels: ["field-report"]
body:
- type: markdown
attributes:
value: |
The protocol is in the README under
[Field tests](https://github.com/syv-ai/HyperQwen#field-tests-wanted).
Short version: start the server, run the harness **twice**, keep the
second run, then check quality. Anything less than the full protocol is
still worth posting — fill in what you have and say what you skipped.
- type: input
id: card
attributes:
label: Card, and how many
description: Include the architecture if you know it (sm80/86/89/90/120).
placeholder: 1x RTX A4000 16 GB (sm86)
validations:
required: true
- type: input
id: power
attributes:
label: Power limit during the run
description: |
`nvidia-smi -q -d POWER | grep "Power Limit"`. This is the single most
common reason two runs disagree — a 3090 at 200 W measures its power cap,
not this stack (33% slower than at 250 W).
placeholder: 250 W (nvidia-smi -pl 250)
validations:
required: true
- type: input
id: stack
attributes:
label: Driver, CUDA, OS, and container or venv
placeholder: 580.65, CUDA 13.0, Ubuntu 24.04 bare metal / Docker / WSL2
validations:
required: true
- type: input
id: rev
attributes:
label: Commit you ran
description: "`git rev-parse --short HEAD`, or the image tag."
placeholder: 9b47388
validations:
required: true
- type: dropdown
id: setup
attributes:
label: Which setup
description: The letters from the README decision tree.
options:
- A — batch
- B — single, default (SPEC=dflash2 PREFIX_CACHE=1)
- C — reproduction (DFLASH_TOKENS=15)
- D — long context (SPEC=mtp CTX=long)
- E — huge context (CTX=huge)
- something else (describe below)
validations:
required: true
- type: textarea
id: rows
attributes:
label: Harness output — the second run
description: |
The `ROW` lines from `bash bench/run_benchmarks.sh <mode>`. The first run
after a start includes JIT warmup and reads 30-50% low, so post the
second. If you measured with your own client instead, say so here and
give the prompt shape, output length and how you define tok/s — it will
be listed as an independent report rather than a table row.
render: text
validations:
required: true
- type: textarea
id: quality
attributes:
label: Quality check
description: |
`python bench/quality_battery.py <tag> --gsm-only` (about 5 minutes).
A fast server that emits garbage is worth nothing, and a quantization
or kernel bug on new hardware shows up here first. Every shipped config
reads 95.0-96.5% on this harness.
render: text
- type: textarea
id: notes
attributes:
label: Anything that did not work
description: |
Build failures, wrong-looking output, things the docs got wrong on your
platform. This is often the more useful half of the report — several
entries in docs/gotchas.md started as one line in an issue like this.
151 changes: 91 additions & 60 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -31,8 +31,8 @@ speculation that drafts out of the prompt. Most of that is not specific to one
checkpoint or one card; it is just that one checkpoint and one card are what has
been measured to death. See [Roadmap](#roadmap-more-qwen-models-more-cards) for
what is and is not portable yet, and
[Help us test more hardware](#help-us-test-more-hardware) if you have credits to
spare on a GPU we have never touched.
[Field tests wanted](#field-tests-wanted) if you have half an hour on a GPU we
have never touched.

## Quick start

Expand All @@ -55,7 +55,42 @@ into `./models`), then serves on `:18020`. One GPU runs one mode at a time.
[docs/docker.md](docs/docker.md#plain-docker-no-compose) ·
[docs/install.md](docs/install.md)

### Which mode
## Which setup do I want?

Three questions, and the answers are the whole of your `.env`.

```mermaid
flowchart TD
Q1{"Who is sending<br/>the requests?"}
Q1 -->|"an API backend, a pipeline,<br/>8+ at once"| A["<b>A — batch</b>"]
Q1 -->|"one or a few<br/>people chatting"| Q2{"Longest prompt<br/>you will send?"}
Q2 -->|"under 64k tokens"| Q3{"Do the answers quote<br/>the prompt back?"}
Q2 -->|"up to 150k"| D["<b>D — long context</b>"]
Q2 -->|"200k and beyond"| E["<b>E — huge context</b>"]
Q3 -->|"no — normal chat"| B["<b>B — single, default</b>"]
Q3 -->|"yes — code edits, RAG,<br/>rewrites, translation"| C["<b>C — reproduction</b>"]
```

Starting from the `.env` you copied in Quick start, which already ships **B**:

| | change in `.env` | start with | what you get |
|---|---|---|---|
| **A** | nothing | `--profile batch` | ~1,035 tok/s aggregate at 64 concurrent |
| **B** | nothing | `--profile single` | 127 tok/s single stream, 64k context |
| **C** | add `DFLASH_TOKENS=15` | `--profile single` | 381 tok/s while quoting, at 4 slots and 56k |
| **D** | `SPEC=mtp` and `CTX=long` | `--profile single` | 150k context, ~95-100 tok/s |
| **E** | add `CTX=huge` | `--profile single` | 240k context, 67 tok/s mixed and 164 while quoting |

Torn between B and C: B is the safe default, and C only pays off when the output
really does repeat the input — it trades request slots and context for it. D drops
back to MTP speculation on purpose, because DFlash2 past 64k is worth it only for
reproduction and loses to `SPEC=mtp CTX=long` about 2:1 on everything else; at E
the KVarN cache buys the context instead, so DFlash2 stays on. A needs no edit at
all — batch ignores `SPEC` and says so on startup. Every knob, and what each costs:
[single-user/README.md](single-user/README.md#if-you-are-the-only-user-do-this) ·
[batch/README.md](batch/README.md).

### The numbers behind that

| | **batch** → [batch/](batch/) | **single** → [single-user/](single-user/) |
|---|---|---|
Expand All @@ -72,24 +107,6 @@ few of ([measurement](docs/long-context.md)). Both modes share one install.
[Full prefill matrix](batch/README.md#prefill) ·
[how each number was won](docs/optimizations.md).

## If you are the only user

The shipped default is conservative — MTP speculation, 8 request slots, 64k
context. If the card is yours alone, two settings are worth more than every
other knob in this repo put together:

```bash
printf 'SPEC=dflash2\nPREFIX_CACHE=1\n' >> .env
docker compose --profile single up -d
```

`SPEC=dflash2` proposes 7 tokens in one pass instead of 4 chained ones;
`PREFIX_CACHE=1` keeps the document you already sent. Add `DFLASH_TOKENS=15` if
your answers quote your prompts — that is the 381 tok/s row above.

Full numbers, the venv equivalent, and what each setting costs:
[single-user/README.md](single-user/README.md#if-you-are-the-only-user-do-this).

## Roadmap: more Qwen models, more cards

The name changed because the scope did. This started as "Qwen3.8-27B on one RTX
Expand Down Expand Up @@ -127,46 +144,60 @@ same treatment rather than a shrug; a larger one for 48 GB; the EXL3 route in
path. If you want a specific model or card prioritized, say so in an issue — the
order is mostly driven by who turns up with a reproduction.

## Help us test more hardware

**We are looking for Runpod or Vast.ai credits.** One RTX 3090 at 250 W is the
entire hardware budget behind every number in this README, and it is shared with
training jobs. That constraint shows: the tables above lean on community
reproductions for every card that is not a 3090, an sm80 speculation bug has sat
open because nobody here owns an sm80 card to debug it on, and each architecture
sweep is scheduled around whatever else needs the GPU that week.

Rented compute converts directly into things this repo does not currently have:

- **An architecture matrix that is measured rather than collected.** sm80, sm89,
sm90, sm120 and multi-GPU under one harness, one power protocol, one set of
prompts — instead of a table where every row came from a different person's
client and cannot honestly be compared to the row above it.
- **Fixing the bugs we cannot reproduce.**
[#98](https://github.com/syv-ai/HyperQwen/issues/98) and
[#72](https://github.com/syv-ai/HyperQwen/issues/72) are sm80 faults diagnosed
entirely from other people's logs. A few hours on a rented A100 probably closes
both.
- **Porting faster.** The vLLM 0.29.0 port
([#106](https://github.com/syv-ai/HyperQwen/issues/106)) is blocked on
packaging details that need a machine to bisect on, while the one card here is
the card serving production.
- **More models.** Every new checkpoint needs its draft vocabulary calibrated and
its drafter trained — GPU-hours, not cleverness.

What you get: results published here as reproductions with the raw harness
output, credit in this section and on the runs themselves, and — if you would
rather have it than the credit — an honest writeup of what your hardware does
badly, which is usually the more useful half.

If you can help, open an issue titled "compute offer" and we will take it from
there. Small amounts are genuinely useful: a single day on one unfamiliar card
has historically been worth more to this project than a month on a familiar one.

Money works too, if that is easier than credits:
[ko-fi.com/mhenrichsen](https://ko-fi.com/mhenrichsen) — it goes to rented GPU
hours and the runs get written up here like any other reproduction.
## Field tests wanted

Every number above is one RTX 3090 at 250 W, shared with training jobs. Every
row for any other card came from someone else running the harness — so the ask
here is a protocol rather than "try it and let us know".

**About 30 minutes on one GPU**, most of it unattended. The harness lives in
the container, so it runs through `compose exec` — on a venv install drop that
prefix and use `python` instead of `venv/bin/python`:

```bash
sudo nvidia-smi -pl 250 # your card's reference wattage; say which
docker compose --profile single up -d # first start also prepares the model

docker compose exec single bash bench/run_benchmarks.sh single # discard: a first run after a start reads 30-50% low
docker compose exec single bash bench/run_benchmarks.sh single # keep this one — its ROW lines are the report

docker compose exec single venv/bin/hf download openai/gsm8k --repo-type dataset \
--include "main/test-*" --local-dir bench/quality-data/gsm8k
docker compose exec single venv/bin/python bench/quality_battery.py mycard --gsm-only
```

Then open a
[field report](https://github.com/syv-ai/HyperQwen/issues/new?template=field-report.yml).
The form asks for the six things that make two runs comparable — card, power
cap, driver, OS or container, commit, which setup letter — because without them
a number cannot honestly be put next to the row above it. Ran something
narrower, or with your own client? Post it anyway and say what you skipped; it
gets listed as an independent report instead of a table row, which is what
happened to most of
[docs/reproductions/](docs/reproductions/README.md#results-from-other-hardware).

**Most wanted, in order.** Several of these already have a thread with someone's
hardware in it — check before you duplicate, and add to theirs if it matches:

| | why it is worth your GPU hour |
|---|---|
| **sm90** — H100, H200 | The only architecture here with no datapoint at all. |
| **sm80** — A100, A30, CMP 170HX | Two owners are mid-bisect on a speculation fault that only their cards produce ([#98](https://github.com/syv-ai/HyperQwen/issues/98), [#72](https://github.com/syv-ai/HyperQwen/issues/72)). A third sm80 box would separate the card from the build. |
| **Four Ampere cards** | Four sm120 cards are measured ([#105](https://github.com/syv-ai/HyperQwen/issues/105)); nobody has run four 3090s, and two of them already give *less* aggregate throughput than one ([#135](https://github.com/syv-ai/HyperQwen/issues/135)). |
| **12 GB cards, on the harness** | Two 3060s do serve this model ([#68](https://github.com/syv-ai/HyperQwen/issues/68)), reported with their owners' own clients — so the numbers cannot be set against the rows above. A harness run on that pair is most of what decides whether smaller Qwen checkpoints are worth preparing. |
| **vLLM 0.29.0** ([#106](https://github.com/syv-ai/HyperQwen/issues/106)) | Ready on a fork, blocked on packaging details that need a machine to bisect on. |

**What you get:** the run published in
[docs/reproductions/](docs/reproductions/README.md) with your raw output and
credit, the gotchas your platform exposes written into
[docs/gotchas.md](docs/gotchas.md), and an honest account of what your hardware
does badly — historically the more useful half.

**Rather give hardware or money than time?** A box we can reach, or credits:
open an issue titled "compute offer". Otherwise
[ko-fi.com/mhenrichsen](https://ko-fi.com/mhenrichsen) buys hours on cards this
project does not own, and those runs get written up here like any other
reproduction.

## Documentation

Expand Down
31 changes: 31 additions & 0 deletions docs/reproductions/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -62,6 +62,37 @@ other only loosely, and not rows for the table above:
attention adding +1.8% at 21k and +6.3% at 66k on top, against this repo's
+2.7% at 16k and +5.3% at 51k. Plus the power-limit ladder quoted above —
[#62](https://github.com/syv-ai/HyperQwen/issues/62).
- **4x RTX 5060 Ti 16 GB (sm120, TP4)**: the canonical harness at TP4, and the
cleanest multi-arm A/B this project has received --- four arms on one box,
`PREFIX_CACHE=1` throughout, one variable apart. C1 greedy decode 72.8
(`SPEC=mtp` k=3, fp8/FlashInfer) -> 137.6 (`SPEC=mtp`, int8 KV on
`TRITON_ATTN`) -> 154.3 (`SPEC=dflash2` k=7, same KV). So the headline 2.37x
from the first round is **1.89x KV dtype and attention backend, 1.12x
drafter**, and the drafter's advantage is gone by C4. `DFLASH_TOKENS=15` at
TP4 is a clear loss (154.3 -> 89.4 at C1, TTFT 3.3x at C8), which is the TP4
half the launcher's keep-it-at-7 warning was missing. Also the source of the
`curand` headers gotcha and the `NCCL_P2P_LEVEL=SYS` note ---
[#105](https://github.com/syv-ai/HyperQwen/issues/105).
- **2x RTX 3060 12 GB**: 24 GB of VRAM in two 12 GB cards, which forces a
tensor-parallel split the single-card path never takes. Two independent boxes
in the thread, both on the Docker path at `SPEC=mtp CTX=long`, one reporting
~50 tok/s at 49k and ~35 at 80k context at a 130 W per-card cap. Measured with
their own clients, so not comparable to the table above --- a harness run on
this pair is the open ask in
[#68](https://github.com/syv-ai/HyperQwen/issues/68).
- **3x RTX 3090**: confirms TP=3 is refused by the checkpoint rather than by this
repo (4 KV heads, 64 layers: neither TP=3 nor an even PP=3 split exists), so
the third card idles under `--tensor-parallel-size 2` by construction. The
thread's workaround --- a second independent engine on the spare card, with
`VLLM_OFFLOAD_KEEP_SHM=1` on both containers --- is the shape a triple-card
owner wants: [#104](https://github.com/syv-ai/HyperQwen/issues/104).
- **2x RTX 3090 (TP=2), native Ubuntu**: batch harness at 917 decode / 861 e2e
at 64 concurrent against the single card's ~1,035 / 948, and 60.1 tok/s at C1
against 46 --- two cards giving less aggregate throughput and ~30% more
single-stream, the same shape as [#40](https://github.com/syv-ai/HyperQwen/issues/40).
Run with `KV=kvarn VISION=1` rather than the reference batch profile, so
indicative rather than a controlled row:
[#135](https://github.com/syv-ai/HyperQwen/issues/135).
- **Dual-GPU reports**: the controlled 1-vs-2×3090 A/B in
[#40](https://github.com/syv-ai/HyperQwen/issues/40) (+16–35%,
161.6 C1 greedy at 275 W, PCIe x8 without NVLink; independently reproduced
Expand Down
6 changes: 4 additions & 2 deletions single-user/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -128,8 +128,10 @@ happens when the streams are big. Where it is *not* the better choice:
5 sliding-window layers from padding the target's attention/GDN layers (105 →
78 KB of pool per token; without it this mode caps out at ~40k), and the V2
runner's profiled activation peak swings ~1 GiB between starts, which makes a
utilization-based setting non-deterministic. `CTX=huge` stays MTP (the script falls back
with a message). `CTX=long` doubles the context — 138,696 tokens at `DFLASH_TOKENS=7`,
utilization-based setting non-deterministic. `CTX=huge` combines with this drafter —
268k tokens of pool at 245760 max-model-len over the KVarN cache
([docs/long-context.md](../docs/long-context.md#dflash2-at-240k-ctxhuge-kvarn-also-combines-with-specdflash2));
any other `CTX` falls back to MTP with a message. `CTX=long` doubles the context — 138,696 tokens at `DFLASH_TOKENS=7`,
114,224 at 15 — by moving to an `int8_per_token_head` cache on the Triton backend; it is
worth it only for context reproduction, and `SPEC=mtp CTX=long` beats it about 2:1 on
everything else. See [docs/long-context.md](../docs/long-context.md#dflash2-past-64k-specdflash2-ctxlong).
Expand Down
Loading