Repository navigation
Pluggable direct prompt: built-in en/fr wordings, prompt files, one prompt_version per wording - #35
maximelefrancois86 wants to merge 8 commits into
Conversation
The llama.cpp backend was CPU-only and branched from the prefilled state by serializing and restoring sequence 0 once per decision. This keeps that path as the default and adds three things. --llama-gpu-layers N passes the offload count to llama.cpp. The metadata records it together with llama_supports_gpu_offload(), so a CPU wheel that silently ignores the flag is visible in the output. --llama-parallel N sets n_seq_max. Shared mode then copies the prefilled sequence 0 into up to N-1 branches with llama_memory_seq_cp and decodes all their suffixes in one batched llama_decode, reading one logit row per branch. Branches are removed whole afterwards, so sequence 0 needs no restore between chunks. The old docstring said hybrid Qwen3.5 memories support neither sequence copies nor partial tail removal; only the second half is true. seq_cp is implemented for llama_memory_recurrent, and both branching modes rely on it. What hybrid memories refuse is truncating a tail back to the prefix, because a recurrent state cannot be partially erased. The docstring and docs now say so. --llama-readout marginal (shared mode, parallel >= 2) folds one-token preambles into the answer-slot masses: for each branch whose most likely next tokens are not slots, the probes are appended in one extra batched decode and the slot masses read after them are added, weighted by the preamble's probability. option_logits become log-masses so softmax and temperature scaling are unchanged; preamble_mass records the share that came through a preamble. Qwen3-8B, which wants to write "**" before the letter, is the case this exists for; on Qwen3.5-4B it changes nothing. benchmarks/shape777.py accepts --backend llamacpp with the same options, so the 37x21 systems fixture runs on GGUF, and reads the GPU name without a CUDA torch. Unit tests cover the new validation, the fan-out batch layout, branch removal, and the marginal fold; the real-GGUF test now also loads with three sequences and compares copy-shared to restore-shared. It passes on the PyPI CPU wheel and on the cu124 wheel with full offload.
…sults on a laptop GPU --llama-parallel and --llama-gpu-layers now default to "auto". For the layers that means not overriding llama.cpp's own default (-1, every layer, in current builds) and recording what the library did; the original backend forced 0. For the branches it means that nothing is asked of the user: the rows are already tokenized by the time shared mode runs, so the prefix length and every suffix length are known, and each state's fan-out holds as many questions as prefix + sum(suffixes) <= n_ctx allows, 32 at most. The context is sized once, for the longest single prompt, not multiplied by the sequence count: with a unified KV buffer every sequence sees the whole n_ctx (measured on this build, 16 sequences and n_ctx 4096 give 4096 per sequence against 256 without), and copied branches share the prefix cells. branches_per_decode in the shared timing records what was done. An integer still fixes n_seq_max; 1 keeps the state-restore path. Results on an RTX 3080 Laptop with the pinned Q4_K_M GGUF, all layers offloaded: authored144 direct 0.796 mean family balanced accuracy (Torch BF16 on a 3090: 0.813), ECE 0.063 out of fold (0.038), median allowed_token_mass 0.9997; shape777 1.40 / 9.21 / 10.88 decisions per second fresh, serial and parallel with 8 branches, 10.51 with auto sizing (Torch: 2.33 / 10.75 / 20.03), 16 to 19 argmax flips out of 777 between decode paths, which is this quantized model's noise floor. One decode per state is no faster than three: llama.cpp re-splits into n_ubatch micro-batches either way. Raw reports, predictions and calibration are committed with their checksums; phase-1 claims are untouched and verify_published.py passes. The real-GGUF test compared serial and shared logits with an exact tolerance that only held because both were the restore path; it now requires exactness when n_seq_max == 1 and 0.5 logits otherwise, decisions still equal. It passes on the PyPI CPU wheel and on the cu124 wheel with offload.
…the backend sits relative to JevBench
… DeltaNet kernel, next to the GGUF rows
…fore the GPU offload
…t has nothing to fold
…es, prompt_version per wording
|
Your "not done, on purpose" note here:
That is implemented and open as #47, cut against It is not only The object case is the one to keep an eye on. With the pinned The fix reads the separator from the payload you already have I have not touched |
The direct prompt — system instruction and the five JSON key names — becomes a value instead of two
module constants.
What changes
core.Prompt: frozen dataclass withversion,systemand the key names (evidence,criterion,options,letter,description).messages(row)renders the turns;evidence_text(state)is the prefix the serial and shared modes cut at.en(the current prompt,direct-options-v1) andfr(direct-options-fr-v1);resolve_prompt()also loads a JSON file.--prompt en|fr|path.json;prompt=onscore,SerialPrefixScorer,score_sharedand_state_prefixin the Torch, MLX and llama.cpp backends.prompt_versionin every result is thewording's own.
docs/PROMPTS.md; one paragraph under Input in the README.What does not change
--prompt en(default) renders byte for byte what the code rendered before — pinned bytest_default_prompt_is_the_published_wording_byte_for_byte;verify_published.pystill passes(69 claims).
and rewrites the instruction, nothing else — which keeps the state-prefix cut valid for every
wording (tested for both built-ins). Answer-slot verification runs unchanged.
Why
Evidence is JSON and answers are letters, so the interface is language-neutral in principle — but
the only wording is English. Measuring whether that costs anything on another language needs a
second wording labelled as such; hence one
prompt_versionper wording rather than a free-textoverride.
The measurement: it costs nothing, and it gains nothing
300 French e-mails × 4 yes/no criteria asked in French, hand labels, pinned Qwen3.5-4B GGUF, serial
mode,
--prompt enthenfr:Ranking metrics are identical; the French wording makes the model say yes more often, so
precision at the natural threshold drops. A threshold fitted per wording removes the gap. The honest
case for this option is traceability, not a quality gain. The data set is private.
Test changes
Fakes in
test_cli.py,test_llamacpp.py,test_shared.pygainedprompt=None; twoassert_called_withintest_cli.pynow expectprompt=DEFAULT_PROMPT. Newtests/test_prompt.py(8 tests, one of them importing the scoring modules with
torchmade unimportable). 85 pass, 3skipped.
Not done, on purpose
pyproject.tomluntouched: makingtorchan extra would breakpip install .for your defaultbackend; it would be one
[project.optional-dependencies]entry, and it is your call.serialising
{"evidence": state}, dropping the final}and the last token. Whenstateis aJSON object, its own closing
}merges with the wrapper's into one token; the prefix no longermatches and
SerialPrefixScorer.scoreraisesState prefix does not match the full prompt(10 of 300 structured states here). Cutting at the longest common token prefix would fix it.
Review points
Prompt,evidence_key, …).--promptis refused with--mode reranker, which has its own prompt.