Repository navigation
llama.cpp backend: GPU offload and sequence-copy fan-out for shared mode - #34
Open
maximelefrancois86 wants to merge 7 commits into
Open
maximelefrancois86 wants to merge 7 commits into
maximelefrancois86 wants to merge 7 commits into
Conversation
The llama.cpp backend was CPU-only and branched from the prefilled state by serializing and restoring sequence 0 once per decision. This keeps that path as the default and adds three things. --llama-gpu-layers N passes the offload count to llama.cpp. The metadata records it together with llama_supports_gpu_offload(), so a CPU wheel that silently ignores the flag is visible in the output. --llama-parallel N sets n_seq_max. Shared mode then copies the prefilled sequence 0 into up to N-1 branches with llama_memory_seq_cp and decodes all their suffixes in one batched llama_decode, reading one logit row per branch. Branches are removed whole afterwards, so sequence 0 needs no restore between chunks. The old docstring said hybrid Qwen3.5 memories support neither sequence copies nor partial tail removal; only the second half is true. seq_cp is implemented for llama_memory_recurrent, and both branching modes rely on it. What hybrid memories refuse is truncating a tail back to the prefix, because a recurrent state cannot be partially erased. The docstring and docs now say so. --llama-readout marginal (shared mode, parallel >= 2) folds one-token preambles into the answer-slot masses: for each branch whose most likely next tokens are not slots, the probes are appended in one extra batched decode and the slot masses read after them are added, weighted by the preamble's probability. option_logits become log-masses so softmax and temperature scaling are unchanged; preamble_mass records the share that came through a preamble. Qwen3-8B, which wants to write "**" before the letter, is the case this exists for; on Qwen3.5-4B it changes nothing. benchmarks/shape777.py accepts --backend llamacpp with the same options, so the 37x21 systems fixture runs on GGUF, and reads the GPU name without a CUDA torch. Unit tests cover the new validation, the fan-out batch layout, branch removal, and the marginal fold; the real-GGUF test now also loads with three sequences and compares copy-shared to restore-shared. It passes on the PyPI CPU wheel and on the cu124 wheel with full offload.
…sults on a laptop GPU --llama-parallel and --llama-gpu-layers now default to "auto". For the layers that means not overriding llama.cpp's own default (-1, every layer, in current builds) and recording what the library did; the original backend forced 0. For the branches it means that nothing is asked of the user: the rows are already tokenized by the time shared mode runs, so the prefix length and every suffix length are known, and each state's fan-out holds as many questions as prefix + sum(suffixes) <= n_ctx allows, 32 at most. The context is sized once, for the longest single prompt, not multiplied by the sequence count: with a unified KV buffer every sequence sees the whole n_ctx (measured on this build, 16 sequences and n_ctx 4096 give 4096 per sequence against 256 without), and copied branches share the prefix cells. branches_per_decode in the shared timing records what was done. An integer still fixes n_seq_max; 1 keeps the state-restore path. Results on an RTX 3080 Laptop with the pinned Q4_K_M GGUF, all layers offloaded: authored144 direct 0.796 mean family balanced accuracy (Torch BF16 on a 3090: 0.813), ECE 0.063 out of fold (0.038), median allowed_token_mass 0.9997; shape777 1.40 / 9.21 / 10.88 decisions per second fresh, serial and parallel with 8 branches, 10.51 with auto sizing (Torch: 2.33 / 10.75 / 20.03), 16 to 19 argmax flips out of 777 between decode paths, which is this quantized model's noise floor. One decode per state is no faster than three: llama.cpp re-splits into n_ubatch micro-batches either way. Raw reports, predictions and calibration are committed with their checksums; phase-1 claims are untouched and verify_published.py passes. The real-GGUF test compared serial and shared logits with an exact tolerance that only held because both were the restore path; it now requires exactness when n_seq_max == 1 and 0.5 logits otherwise, decisions still equal. It passes on the PyPI CPU wheel and on the cu124 wheel with offload.
…the backend sits relative to JevBench
… DeltaNet kernel, next to the GGUF rows
…fore the GPU offload
…t has nothing to fold
This was referenced Sep 26, 2026
tvpavan
added a commit
to tvpavan/SemIf-OpenJev
that referenced
this pull request
Oct 2, 2026
Merge picks: upstream PRs TheoLeeCJ#34 TheoLeeCJ#41 TheoLeeCJ#42 TheoLeeCJ#47 TheoLeeCJ#55 + lazy option permutations
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The llama.cpp backend (#18) runs GGUF checkpoints on CPU only and branches from a prefilled state
by saving sequence 0 and restoring it once per decision. This PR keeps that path and adds two
options in separable commits. What it buys, measured on one laptop (RTX 3080, 16 GB): the GPU
offload turns 0.05 / 0.49 decisions per second (fresh / shared, CPU) into 1.45 / 9.86 — ×30 and
×20; the fan-out adds 13–18 % on top. It does not make the GGUF path faster than Torch BF16 on the same GPU (see the tables):
what the GGUF keeps is 3 GB of weights instead of 8.8, and no torch.
The two options
--llama-gpu-layers autostops forcingn_gpu_layers = 0and leaves llama.cpp's own default(every layer, in current builds);
0andNremain. The metadata records the effective value andllama_supports_gpu_offload(), so a CPU-only wheel that ignores the flag shows in the output, notin the timings.
--llama-parallel automakes shared mode copy the prefilled sequence into branches withllama_memory_seq_cpand decode all their suffixes in onellama_decode, one logit row perbranch. The number of branches per decode is not a setting:
prefix + Σ suffixes ≤ n_ctxdecidesit per state (cap 32). The context is sized once with
kv_unified, so every sequence sees the wholen_ctxand copies share the prefix cells.Nfixesn_seq_max;1keeps state restore. This isthe llama.cpp counterpart of the Torch
native-state-prefix-parallel-v1path.benchmarks/shape777.pygains--backend llamacppwith the same options.One docstring corrected
The module said hybrid Qwen3.5 memories "support neither sequence copies nor partial tail removal".
Only the second holds:
llama_memory_recurrentimplementsseq_cp— both branching modes use it —and what a recurrent state cannot do is be partially erased (
seq_rmrefuses any range thatincludes the last position, because the state exists only for the last token). Docstring and
docs/LLAMACPP.mdnow say that.Validation
pytest -q: 71 passed, 3 skipped (new unit tests: argument validation, fan-out layout andbranch removal against a fake library).
(same decisions;
option_logitswithin 0.5 — they were bit-identical only because both were thesame restore path). Passes on the PyPI CPU wheel (29 s) and on the
cu124 wheel with all layers on an RTX 3080 Laptop (15 s).
bartowski/Qwen_Qwen3.5-4B-GGUF@4168f45a, Q4_K_M, the artifact already pinned inmanifests/models.json.Results
Quality,
authored144, direct mode:Speed,
shape777, decisions per second. First the line this PR starts from — the backend as merged,CPU only — on the first 5 states (105 decisions; backend functions timed directly, since
shape777.pytakes only the full fixture; raw file in the PR):#18)Then the full fixture (37 states × 21 questions), where the subset's GPU rows land within 7 %:
flash-linear-attention8) · 10.51 (auto)The laptop BF16 rows are the same GPU as the GGUF rows, so they are the fair comparison: BF16 with
the fast delta-rule kernel is 8 / 11 / 30 % faster than the 4-bit GGUF and 1.7 points more accurate,
for 8.8 GB of weights instead of 3. The GGUF is the deployment path, not the fast path.
pip install -e .does not install that kernel —transformerssays so at loadtime and falls back to reference PyTorch, the second BF16 row. Neither kernel touches llama.cpp.
Fan-out gains 18 % over state restore by removing the save/restore round-trip; one decode per state
(
auto) is no faster than three, because llama.cpp re-splits the batch inton_ubatchmicro-batches either way. Argmax flips between decode paths (16–19 / 777) are the 4-bit model's
noise floor: two sequential reads of one prompt already differ on ~3 % of decisions when only the
micro-batch size changes.
New raw files:
results/raw/llamacpp-gguf-cuda-authored144.json,results/raw/shape777-llamacpp-gguf-cuda{,-auto}.json(+ predictions), calibration report, linesappended to
SHA256SUMS.results/phase1-summary.jsonand the headline claims are untouched.Relation to JevBench
JevBench's
semif_directadapter runs SemIf in Torch BF16. This backend is what lets the sameadapter run the pinned GGUF on a laptop GPU. On the 231 public tasks, same mapping: published BF16
on a 3090 187/231, BF16 on the laptop 186/231, this backend 183/231.
Known issue, not in this backend
The cu124 wheel's CPU path raises
Illegal instructionon an Intel Xeon W-11955M (AVX-512 withoutavx512_bf16/AMX); the PyPI CPU wheel is fine. Documented; belongs to the wheel.Dropped on the way
An earlier version of this branch carried a marginal readout that folded one-token preambles
(
**) back into the answer slots, motivated by Qwen3-8B answering in Markdown under anotherprompt. Run on this repository's prompt and
authored144, Qwen3-8B puts its first-token mass on aletter in 144 of 144 rows (
preamble_mass0, identical accuracy 0.793), so the readout had nothingto fold and was removed rather than shipped without a demonstrated use.
À vérifier avant d'envoyer