Skip to content

Merge picks: upstream PRs #34 #41 #42 #47 #55 + lazy option permutations - #1

Merged
tvpavan merged 17 commits into
masterfrom
picks
Oct 2, 2026
Merged

tvpavan merged 17 commits into
masterfrom
picks

Conversation

@tvpavan

@tvpavan tvpavan commented Oct 2, 2026

Copy link
Copy Markdown
Owner

Brings picks into master: upstream TheoLeeCJ/SemIf-OpenJev plus its open PRs and one fix.

From upstream PRs:

Fork fix (b7bcb30): order.select_permutations listed all n! option orders before picking K (1.45 s at 10 options, factorial beyond). It now samples lazily above 8 options. Outputs are byte-identical up to 8 options (checked n=2..8, K in 2,3,4,5,8,16, seeds 0,1,7); K=2 is always identity + reverse. For 9+ options with K>=3 the extra orders come from a different draw.

Tests: pytest tests 192 passed, 4 skipped on picks at b7bcb30 (macOS, llama-cpp-python 0.3.35).

Used by rsa-ecom/xb-decisions-agent PR TheoLeeCJ#14, which pins b7bcb30.

maximelefrancois86 and others added 17 commits September 23, 2026 06:55
The llama.cpp backend was CPU-only and branched from the prefilled state by
serializing and restoring sequence 0 once per decision. This keeps that path as
the default and adds three things.

--llama-gpu-layers N passes the offload count to llama.cpp. The metadata
records it together with llama_supports_gpu_offload(), so a CPU wheel that
silently ignores the flag is visible in the output.

--llama-parallel N sets n_seq_max. Shared mode then copies the prefilled
sequence 0 into up to N-1 branches with llama_memory_seq_cp and decodes all
their suffixes in one batched llama_decode, reading one logit row per branch.
Branches are removed whole afterwards, so sequence 0 needs no restore between
chunks. The old docstring said hybrid Qwen3.5 memories support neither sequence
copies nor partial tail removal; only the second half is true. seq_cp is
implemented for llama_memory_recurrent, and both branching modes rely on it.
What hybrid memories refuse is truncating a tail back to the prefix, because a
recurrent state cannot be partially erased. The docstring and docs now say so.

--llama-readout marginal (shared mode, parallel >= 2) folds one-token
preambles into the answer-slot masses: for each branch whose most likely next
tokens are not slots, the probes are appended in one extra batched decode and
the slot masses read after them are added, weighted by the preamble's
probability. option_logits become log-masses so softmax and temperature
scaling are unchanged; preamble_mass records the share that came through a
preamble. Qwen3-8B, which wants to write "**" before the letter, is the case
this exists for; on Qwen3.5-4B it changes nothing.

benchmarks/shape777.py accepts --backend llamacpp with the same options, so the
37x21 systems fixture runs on GGUF, and reads the GPU name without a CUDA
torch. Unit tests cover the new validation, the fan-out batch layout, branch
removal, and the marginal fold; the real-GGUF test now also loads with three
sequences and compares copy-shared to restore-shared. It passes on the PyPI CPU
wheel and on the cu124 wheel with full offload.
…sults on a laptop GPU

--llama-parallel and --llama-gpu-layers now default to "auto". For the layers
that means not overriding llama.cpp's own default (-1, every layer, in current
builds) and recording what the library did; the original backend forced 0.

For the branches it means that nothing is asked of the user: the rows are
already tokenized by the time shared mode runs, so the prefix length and every
suffix length are known, and each state's fan-out holds as many questions as
prefix + sum(suffixes) <= n_ctx allows, 32 at most. The context is sized once,
for the longest single prompt, not multiplied by the sequence count: with a
unified KV buffer every sequence sees the whole n_ctx (measured on this build,
16 sequences and n_ctx 4096 give 4096 per sequence against 256 without), and
copied branches share the prefix cells. branches_per_decode in the shared
timing records what was done. An integer still fixes n_seq_max; 1 keeps the
state-restore path.

Results on an RTX 3080 Laptop with the pinned Q4_K_M GGUF, all layers
offloaded: authored144 direct 0.796 mean family balanced accuracy (Torch BF16
on a 3090: 0.813), ECE 0.063 out of fold (0.038), median allowed_token_mass
0.9997; shape777 1.40 / 9.21 / 10.88 decisions per second fresh, serial and
parallel with 8 branches, 10.51 with auto sizing (Torch: 2.33 / 10.75 / 20.03),
16 to 19 argmax flips out of 777 between decode paths, which is this quantized
model's noise floor. One decode per state is no faster than three: llama.cpp
re-splits into n_ubatch micro-batches either way. Raw reports, predictions and
calibration are committed with their checksums; phase-1 claims are untouched
and verify_published.py passes.

The real-GGUF test compared serial and shared logits with an exact tolerance
that only held because both were the restore path; it now requires exactness
when n_seq_max == 1 and 0.5 logits otherwise, decisions still equal. It passes
on the PyPI CPU wheel and on the cu124 wheel with offload.
Expose the existing per-workload temperature scaling as an installable
command and wire evaluate's non-screen auto-decide/review operating point,
without changing the frozen 96-row screening gate or published tables.
Measure ID-aligned flips, total variation, and position bias from committed
predictions without changing default scores; optional --stabilize-order K
averages permutation logits under a distinct prompt_version at Kx cost.
validate_row declares state may be a string, object or array and score()
honours all three, but the prefix-sharing paths rejected a subset of valid
states with "The fixed state prefix does not match every full prompt". The same
row succeeded with one question and raised with two, and which states failed was
a property of the tokenizer vocabulary, so a caller could not predict or avoid it.

Both copies of _state_prefix ended with encode(text)[:-1], assuming that
appending the next field's punctuation always merges with exactly one final
token. That is a property of the state text, not a guarantee. Against the pinned
Qwen/Qwen3.5-4B tokenizer at revision 851bf6e, a state ending in one of ( - = ? }
produces a prefix one token too long, so it is not a token prefix of the full
prompt at all: 5 of the 100 printable ASCII characters, and the state from the
issue report, "Did the deploy use the same parameters)?", is one of them.

serial.py carried the same duplicated helper with the same defect, so
SerialPrefixScorer raised on exactly the same states even though the issue only
reports score_shared. Both are fixed here.

Replace the guess with the longest common token prefix between the truncated
evidence and that text plus the separator that always follows it. The separator
is taken from the payload already in hand rather than hardcoded, so the fix
survives a change to the prompt wording. A consequence worth stating: when the
boundary token merges forward, the shared prefix now stops just before it, which
is the longest prefix that can be shared at all.

tests/test_state_prefix.py drives both copies through a byte tokenizer that
merges only at the evidence boundary, so the regression is covered without torch
and without network. Four of its cases fail against the previous code.

The fix is a no-op on every committed fixture. Audited with the pinned tokenizer
over all 217 distinct states in shape777, authored144 and perturbations108: zero
prefix differences against the previous implementation, and both copies now
agree on all of them, so no committed evidence is affected and nothing needs
regenerating. verify_published.py still reproduces all 69 claims and every
results/raw checksum still matches.
The stub tokenizer only reproduced the string-state variant, so the test proved
less than it appeared to: it would still have passed if the object case were
broken. Objects and arrays are the shapes that actually fail in practice, because
the value's own closing brace or quote sits right in front of the comma.

Model both extremes instead of one hand-picked triple. PlainTokenizer never merges
so the boundary costs exactly one token, which is the case the old code handled
correctly; MergingTokenizer lets a token straddle the comma, which is the case it
did not. Six states across both tokenizers and both copies: 48 cases, and 10 of
them fail against the previous implementation.

Verified against the pinned Qwen tokenizer that this is not a stub artefact: of 70
structured states ending in a merge-prone character, 11 produce an invalid prefix
with the old code and 0 with the fix.
semif-score --backend sglang scores the direct, serial, shared, and
reranker modes through the /v1/score endpoint of a running SGLang
server that contains sgl-project/sglang#40826. SemIf still renders
every prompt, checks the answer slots, and validates the input on the
pinned reference tokenizer, so prompt_sha256, input_tokens,
prompt_version, option order, and answer_token_ids equal the Torch
values. The server only scores SemIf's token ids, with each row's
answer slots as that item's candidates and return_token_logprobs, and
the client uses the standard library, so there is no new dependency.

Before the first row, the backend refuses an unreachable server, a
server without per-item candidate scoring, another model or revision,
a Hub model whose server reports no revision, a non-generation model,
--load-format dummy, --enable-mis, --allow-auto-truncate, any
--preferred-sampling-params, a server tokenizer that disagrees with
the reference tokenizer, and a --max-tokens at or above the server's
max_req_input_len. Every response must carry one finite
log-probability per candidate and a usage.prompt_tokens equal to the
token ids sent. Requests go straight to the server with no proxy,
redirect, or retry, and SEMIF_SGLANG_API_KEY is sent only as a bearer
header and never written to a row.

option_logits hold the candidates' full-vocabulary log-probabilities,
one constant per row away from raw logits, as the readout string says.
cli.py now binds a reranker scorer per backend, so reranker mode
reaches the SGLang scorer, and --sglang-url and --sglang-timeout
require --backend sglang. docs/SGLANG.md covers the server launch and
the refused settings, and the real-server tests read SEMIF_SGLANG_URL
and SEMIF_SGLANG_RERANKER_URL.
# Conflicts:
#	README.md
# Conflicts:
#	src/semif_phase1/cli.py
#	tests/test_cli.py
select_permutations listed all n! permutations before picking K: 1.45 s at
10 options, factorial beyond. Up to 8 options the enumerate-and-shuffle path
is kept, so existing stabilized results reproduce exactly (checked for
n=2..8, K in 2,3,4,5,8,16, seeds 0,1,7). Above that, orders are drawn with
random.Random(seed).sample until K distinct ones are chosen; identity and
reverse are still first.
@tvpavan
tvpavan merged commit f964e73 into master Oct 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants