Skip to content

llama.cpp backend: GPU offload and sequence-copy fan-out for shared mode - #34

Open
maximelefrancois86 wants to merge 7 commits into
TheoLeeCJ:masterfrom
maximelefrancois86:llamacpp-cuda-parallel
Open

maximelefrancois86 wants to merge 7 commits into
TheoLeeCJ:masterfrom
maximelefrancois86:llamacpp-cuda-parallel

Conversation

@maximelefrancois86

Copy link
Copy Markdown

The llama.cpp backend (#18) runs GGUF checkpoints on CPU only and branches from a prefilled state
by saving sequence 0 and restoring it once per decision. This PR keeps that path and adds two
options in separable commits. What it buys, measured on one laptop (RTX 3080, 16 GB): the GPU
offload turns 0.05 / 0.49 decisions per second (fresh / shared, CPU) into 1.45 / 9.86 — ×30 and
×20; the fan-out adds 13–18 % on top. It does not make the GGUF path faster than Torch BF16 on the same GPU (see the tables):
what the GGUF keeps is 3 GB of weights instead of 8.8, and no torch.

The two options

--llama-gpu-layers auto stops forcing n_gpu_layers = 0 and leaves llama.cpp's own default
(every layer, in current builds); 0 and N remain. The metadata records the effective value and
llama_supports_gpu_offload(), so a CPU-only wheel that ignores the flag shows in the output, not
in the timings.

--llama-parallel auto makes shared mode copy the prefilled sequence into branches with
llama_memory_seq_cp and decode all their suffixes in one llama_decode, one logit row per
branch. The number of branches per decode is not a setting: prefix + Σ suffixes ≤ n_ctx decides
it per state (cap 32). The context is sized once with kv_unified, so every sequence sees the whole
n_ctx and copies share the prefix cells. N fixes n_seq_max; 1 keeps state restore. This is
the llama.cpp counterpart of the Torch native-state-prefix-parallel-v1 path.

benchmarks/shape777.py gains --backend llamacpp with the same options.

One docstring corrected

The module said hybrid Qwen3.5 memories "support neither sequence copies nor partial tail removal".
Only the second holds: llama_memory_recurrent implements seq_cp — both branching modes use it —
and what a recurrent state cannot do is be partially erased (seq_rm refuses any range that
includes the last position, because the state exists only for the last token). Docstring and
docs/LLAMACPP.md now say that.

Validation

  • pytest -q: 71 passed, 3 skipped (new unit tests: argument validation, fan-out layout and
    branch removal against a fake library).
  • The real-GGUF test now also loads three sequences, checks copy-shared against restore-shared
    (same decisions; option_logits within 0.5 — they were bit-identical only because both were the
    same restore path). Passes on the PyPI CPU wheel (29 s) and on the
    cu124 wheel with all layers on an RTX 3080 Laptop (15 s).
  • GGUF: bartowski/Qwen_Qwen3.5-4B-GGUF @ 4168f45a, Q4_K_M, the artifact already pinned in
    manifests/models.json.

Results

Quality, authored144, direct mode:

backend GPU balanced accuracy ECE (own T, out-of-fold)
Torch BF16 (committed) RTX 3090 0.813 0.038
Torch BF16 RTX 3080 Laptop 0.813 0.050
llama.cpp Q4_K_M, this PR RTX 3080 Laptop 0.796 0.063

Speed, shape777, decisions per second. First the line this PR starts from — the backend as merged,
CPU only — on the first 5 states (105 decisions; backend functions timed directly, since
shape777.py takes only the full fixture; raw file in the PR):

llama.cpp Q4_K_M, laptop, 5 states fresh shared
before: CPU, state restore (#18) 0.05 0.49
after: GPU, state restore 1.45 9.86
after: GPU, fan-out 1.43 11.18

Then the full fixture (37 states × 21 questions), where the subset's GPU rows land within 7 %:

backend GPU fresh serial prefix parallel shared
Torch BF16 (committed) RTX 3090 2.33 10.75 20.03
Torch BF16, flash-linear-attention RTX 3080 Laptop 1.51 10.19 14.14
Torch BF16, reference PyTorch kernels RTX 3080 Laptop 1.10 7.78 10.24
llama.cpp Q4_K_M, this PR RTX 3080 Laptop 1.40 9.21 10.88 (8) · 10.51 (auto)

The laptop BF16 rows are the same GPU as the GGUF rows, so they are the fair comparison: BF16 with
the fast delta-rule kernel is 8 / 11 / 30 % faster than the 4-bit GGUF and 1.7 points more accurate,
for 8.8 GB of weights instead of 3. The GGUF is the deployment path, not the fast path. pip install -e . does not install that kernel — transformers says so at load
time and falls back to reference PyTorch, the second BF16 row. Neither kernel touches llama.cpp.

Fan-out gains 18 % over state restore by removing the save/restore round-trip; one decode per state
(auto) is no faster than three, because llama.cpp re-splits the batch into n_ubatch
micro-batches either way. Argmax flips between decode paths (16–19 / 777) are the 4-bit model's
noise floor: two sequential reads of one prompt already differ on ~3 % of decisions when only the
micro-batch size changes.

New raw files: results/raw/llamacpp-gguf-cuda-authored144.json,
results/raw/shape777-llamacpp-gguf-cuda{,-auto}.json (+ predictions), calibration report, lines
appended to SHA256SUMS. results/phase1-summary.json and the headline claims are untouched.

Relation to JevBench

JevBench's semif_direct adapter runs SemIf in Torch BF16. This backend is what lets the same
adapter run the pinned GGUF on a laptop GPU. On the 231 public tasks, same mapping: published BF16
on a 3090 187/231, BF16 on the laptop 186/231, this backend 183/231.

Known issue, not in this backend

The cu124 wheel's CPU path raises Illegal instruction on an Intel Xeon W-11955M (AVX-512 without
avx512_bf16/AMX); the PyPI CPU wheel is fine. Documented; belongs to the wheel.


Dropped on the way

An earlier version of this branch carried a marginal readout that folded one-token preambles
(**) back into the answer slots, motivated by Qwen3-8B answering in Markdown under another
prompt. Run on this repository's prompt and authored144, Qwen3-8B puts its first-token mass on a
letter in 144 of 144 rows (preamble_mass 0, identical accuracy 0.793), so the readout had nothing
to fold and was removed rather than shipped without a demonstrated use.

À vérifier avant d'envoyer

  1. Trois fonctionnalités, trois commits séparables — le dire ainsi vous convient ?
  2. La correction du docstring de l'auteur reste-t-elle courtoise ?
  3. Auteur des commits : « Maxime Lefrançois » avec votre adresse gmail — à confirmer.

The llama.cpp backend was CPU-only and branched from the prefilled state by
serializing and restoring sequence 0 once per decision. This keeps that path as
the default and adds three things.

--llama-gpu-layers N passes the offload count to llama.cpp. The metadata
records it together with llama_supports_gpu_offload(), so a CPU wheel that
silently ignores the flag is visible in the output.

--llama-parallel N sets n_seq_max. Shared mode then copies the prefilled
sequence 0 into up to N-1 branches with llama_memory_seq_cp and decodes all
their suffixes in one batched llama_decode, reading one logit row per branch.
Branches are removed whole afterwards, so sequence 0 needs no restore between
chunks. The old docstring said hybrid Qwen3.5 memories support neither sequence
copies nor partial tail removal; only the second half is true. seq_cp is
implemented for llama_memory_recurrent, and both branching modes rely on it.
What hybrid memories refuse is truncating a tail back to the prefix, because a
recurrent state cannot be partially erased. The docstring and docs now say so.

--llama-readout marginal (shared mode, parallel >= 2) folds one-token
preambles into the answer-slot masses: for each branch whose most likely next
tokens are not slots, the probes are appended in one extra batched decode and
the slot masses read after them are added, weighted by the preamble's
probability. option_logits become log-masses so softmax and temperature
scaling are unchanged; preamble_mass records the share that came through a
preamble. Qwen3-8B, which wants to write "**" before the letter, is the case
this exists for; on Qwen3.5-4B it changes nothing.

benchmarks/shape777.py accepts --backend llamacpp with the same options, so the
37x21 systems fixture runs on GGUF, and reads the GPU name without a CUDA
torch. Unit tests cover the new validation, the fan-out batch layout, branch
removal, and the marginal fold; the real-GGUF test now also loads with three
sequences and compares copy-shared to restore-shared. It passes on the PyPI CPU
wheel and on the cu124 wheel with full offload.
…sults on a laptop GPU

--llama-parallel and --llama-gpu-layers now default to "auto". For the layers
that means not overriding llama.cpp's own default (-1, every layer, in current
builds) and recording what the library did; the original backend forced 0.

For the branches it means that nothing is asked of the user: the rows are
already tokenized by the time shared mode runs, so the prefix length and every
suffix length are known, and each state's fan-out holds as many questions as
prefix + sum(suffixes) <= n_ctx allows, 32 at most. The context is sized once,
for the longest single prompt, not multiplied by the sequence count: with a
unified KV buffer every sequence sees the whole n_ctx (measured on this build,
16 sequences and n_ctx 4096 give 4096 per sequence against 256 without), and
copied branches share the prefix cells. branches_per_decode in the shared
timing records what was done. An integer still fixes n_seq_max; 1 keeps the
state-restore path.

Results on an RTX 3080 Laptop with the pinned Q4_K_M GGUF, all layers
offloaded: authored144 direct 0.796 mean family balanced accuracy (Torch BF16
on a 3090: 0.813), ECE 0.063 out of fold (0.038), median allowed_token_mass
0.9997; shape777 1.40 / 9.21 / 10.88 decisions per second fresh, serial and
parallel with 8 branches, 10.51 with auto sizing (Torch: 2.33 / 10.75 / 20.03),
16 to 19 argmax flips out of 777 between decode paths, which is this quantized
model's noise floor. One decode per state is no faster than three: llama.cpp
re-splits into n_ubatch micro-batches either way. Raw reports, predictions and
calibration are committed with their checksums; phase-1 claims are untouched
and verify_published.py passes.

The real-GGUF test compared serial and shared logits with an exact tolerance
that only held because both were the restore path; it now requires exactness
when n_seq_max == 1 and 0.5 logits otherwise, decisions still equal. It passes
on the PyPI CPU wheel and on the cu124 wheel with offload.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant