Skip to content

Pluggable direct prompt: built-in en/fr wordings, prompt files, one prompt_version per wording - #35

Draft
maximelefrancois86 wants to merge 8 commits into
TheoLeeCJ:masterfrom
maximelefrancois86:pluggable-prompt
Draft

maximelefrancois86 wants to merge 8 commits into
TheoLeeCJ:masterfrom
maximelefrancois86:pluggable-prompt

Conversation

@maximelefrancois86

Copy link
Copy Markdown

Stacked on #34 (llamacpp-cuda-parallel): the diff below includes that branch's commits until it merges; the one commit of this PR is the last one. Opened as a draft for that reason.

The direct prompt — system instruction and the five JSON key names — becomes a value instead of two
module constants.

What changes

  • core.Prompt: frozen dataclass with version, system and the key names (evidence,
    criterion, options, letter, description). messages(row) renders the turns;
    evidence_text(state) is the prefix the serial and shared modes cut at.
  • Built-ins en (the current prompt, direct-options-v1) and fr (direct-options-fr-v1);
    resolve_prompt() also loads a JSON file.
  • --prompt en|fr|path.json; prompt= on score, SerialPrefixScorer, score_shared and
    _state_prefix in the Torch, MLX and llama.cpp backends. prompt_version in every result is the
    wording's own. docs/PROMPTS.md; one paragraph under Input in the README.

What does not change

  • --prompt en (default) renders byte for byte what the code rendered before — pinned by
    test_default_prompt_is_the_published_wording_byte_for_byte; verify_published.py still passes
    (69 claims).
  • The payload shape: evidence first, then criterion, then lettered options. A prompt renames keys
    and rewrites the instruction, nothing else — which keeps the state-prefix cut valid for every
    wording (tested for both built-ins). Answer-slot verification runs unchanged.

Why

Evidence is JSON and answers are letters, so the interface is language-neutral in principle — but
the only wording is English. Measuring whether that costs anything on another language needs a
second wording labelled as such; hence one prompt_version per wording rather than a free-text
override.

The measurement: it costs nothing, and it gains nothing

300 French e-mails × 4 yes/no criteria asked in French, hand labels, pinned Qwen3.5-4B GGUF, serial
mode, --prompt en then fr:

en fr
macro AP / AUC 0.79 / 0.984 0.78 / 0.984
macro F1 at the best threshold 0.82 0.83
macro F1 at 0.5 0.70 0.61
decisions that flip 5.0 %

Ranking metrics are identical; the French wording makes the model say yes more often, so
precision at the natural threshold drops. A threshold fitted per wording removes the gap. The honest
case for this option is traceability, not a quality gain. The data set is private.

Test changes

Fakes in test_cli.py, test_llamacpp.py, test_shared.py gained prompt=None; two
assert_called_with in test_cli.py now expect prompt=DEFAULT_PROMPT. New tests/test_prompt.py
(8 tests, one of them importing the scoring modules with torch made unimportable). 85 pass, 3
skipped.

Not done, on purpose

  • pyproject.toml untouched: making torch an extra would break pip install . for your default
    backend; it would be one [project.optional-dependencies] entry, and it is your call.
  • A bug found on the way, not fixed here: serial and shared modes cut the state prefix by
    serialising {"evidence": state}, dropping the final } and the last token. When state is a
    JSON object, its own closing } merges with the wrapper's into one token; the prefix no longer
    matches and SerialPrefixScorer.score raises State prefix does not match the full prompt
    (10 of 300 structured states here). Cutting at the longest common token prefix would fix it.

Review points

  1. Names (Prompt, evidence_key, …).
  2. Ship the French built-in, or built-ins English-only and everything else from files?
  3. --prompt is refused with --mode reranker, which has its own prompt.

The llama.cpp backend was CPU-only and branched from the prefilled state by
serializing and restoring sequence 0 once per decision. This keeps that path as
the default and adds three things.

--llama-gpu-layers N passes the offload count to llama.cpp. The metadata
records it together with llama_supports_gpu_offload(), so a CPU wheel that
silently ignores the flag is visible in the output.

--llama-parallel N sets n_seq_max. Shared mode then copies the prefilled
sequence 0 into up to N-1 branches with llama_memory_seq_cp and decodes all
their suffixes in one batched llama_decode, reading one logit row per branch.
Branches are removed whole afterwards, so sequence 0 needs no restore between
chunks. The old docstring said hybrid Qwen3.5 memories support neither sequence
copies nor partial tail removal; only the second half is true. seq_cp is
implemented for llama_memory_recurrent, and both branching modes rely on it.
What hybrid memories refuse is truncating a tail back to the prefix, because a
recurrent state cannot be partially erased. The docstring and docs now say so.

--llama-readout marginal (shared mode, parallel >= 2) folds one-token
preambles into the answer-slot masses: for each branch whose most likely next
tokens are not slots, the probes are appended in one extra batched decode and
the slot masses read after them are added, weighted by the preamble's
probability. option_logits become log-masses so softmax and temperature
scaling are unchanged; preamble_mass records the share that came through a
preamble. Qwen3-8B, which wants to write "**" before the letter, is the case
this exists for; on Qwen3.5-4B it changes nothing.

benchmarks/shape777.py accepts --backend llamacpp with the same options, so the
37x21 systems fixture runs on GGUF, and reads the GPU name without a CUDA
torch. Unit tests cover the new validation, the fan-out batch layout, branch
removal, and the marginal fold; the real-GGUF test now also loads with three
sequences and compares copy-shared to restore-shared. It passes on the PyPI CPU
wheel and on the cu124 wheel with full offload.
…sults on a laptop GPU

--llama-parallel and --llama-gpu-layers now default to "auto". For the layers
that means not overriding llama.cpp's own default (-1, every layer, in current
builds) and recording what the library did; the original backend forced 0.

For the branches it means that nothing is asked of the user: the rows are
already tokenized by the time shared mode runs, so the prefix length and every
suffix length are known, and each state's fan-out holds as many questions as
prefix + sum(suffixes) <= n_ctx allows, 32 at most. The context is sized once,
for the longest single prompt, not multiplied by the sequence count: with a
unified KV buffer every sequence sees the whole n_ctx (measured on this build,
16 sequences and n_ctx 4096 give 4096 per sequence against 256 without), and
copied branches share the prefix cells. branches_per_decode in the shared
timing records what was done. An integer still fixes n_seq_max; 1 keeps the
state-restore path.

Results on an RTX 3080 Laptop with the pinned Q4_K_M GGUF, all layers
offloaded: authored144 direct 0.796 mean family balanced accuracy (Torch BF16
on a 3090: 0.813), ECE 0.063 out of fold (0.038), median allowed_token_mass
0.9997; shape777 1.40 / 9.21 / 10.88 decisions per second fresh, serial and
parallel with 8 branches, 10.51 with auto sizing (Torch: 2.33 / 10.75 / 20.03),
16 to 19 argmax flips out of 777 between decode paths, which is this quantized
model's noise floor. One decode per state is no faster than three: llama.cpp
re-splits into n_ubatch micro-batches either way. Raw reports, predictions and
calibration are committed with their checksums; phase-1 claims are untouched
and verify_published.py passes.

The real-GGUF test compared serial and shared logits with an exact tolerance
that only held because both were the restore path; it now requires exactness
when n_seq_max == 1 and 0.5 logits otherwise, decisions still equal. It passes
on the PyPI CPU wheel and on the cu124 wheel with offload.
@pxtroniwnl

Copy link
Copy Markdown

Your "not done, on purpose" note here:

serial and shared modes cut the state prefix by serialising {"evidence": state},
dropping the final } and the last token. When state is a JSON object, its own
closing } merges with the wrapper's into one token; the prefix no longer matches
and SerialPrefixScorer.score raises ... (10 of 300 structured states here).
Cutting at the longest common token prefix would fix it.

That is implemented and open as #47, cut against master. It is not a
competing fix — it is the one you described, so you may prefer to just rebase on
top of it. Two notes that may save you some work:

It is not only serial.py. src/semif_phase1/shared.py had a duplicated copy
of _state_prefix with the same [:-1], so score_shared raises on the same
states. Both are fixed in #47.

The object case is the one to keep an eye on. With the pinned
Qwen/Qwen3.5-4B tokenizer at 851bf6e, over 70 structured states whose tail
ends in a merge-prone character, 11 produce an invalid prefix and 0 do after the
fix. Your 10 of 300 is the same order of magnitude, so this should clear it — but
if you still see it after rebasing, the interesting variable is the character the
state ends in, not whether the state is an object. String states are affected too
(5 of the 100 printable ASCII characters), which is easy to miss if you only test
with structured data.

The fix reads the separator from the payload you already have
(payload[len(evidence):]) rather than hardcoding it, specifically so it keeps
working with --prompt fr and prompt files. That is the one thing to preserve if
you resolve the shared hunks by hand instead of rebasing.

I have not touched core.py / Prompt, so there is no conflict with the
pluggable-prompt work itself — the overlap is confined to _state_prefix.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants