Repository navigation
Conversation
The llama.cpp backend was CPU-only and branched from the prefilled state by serializing and restoring sequence 0 once per decision. This keeps that path as the default and adds three things. --llama-gpu-layers N passes the offload count to llama.cpp. The metadata records it together with llama_supports_gpu_offload(), so a CPU wheel that silently ignores the flag is visible in the output. --llama-parallel N sets n_seq_max. Shared mode then copies the prefilled sequence 0 into up to N-1 branches with llama_memory_seq_cp and decodes all their suffixes in one batched llama_decode, reading one logit row per branch. Branches are removed whole afterwards, so sequence 0 needs no restore between chunks. The old docstring said hybrid Qwen3.5 memories support neither sequence copies nor partial tail removal; only the second half is true. seq_cp is implemented for llama_memory_recurrent, and both branching modes rely on it. What hybrid memories refuse is truncating a tail back to the prefix, because a recurrent state cannot be partially erased. The docstring and docs now say so. --llama-readout marginal (shared mode, parallel >= 2) folds one-token preambles into the answer-slot masses: for each branch whose most likely next tokens are not slots, the probes are appended in one extra batched decode and the slot masses read after them are added, weighted by the preamble's probability. option_logits become log-masses so softmax and temperature scaling are unchanged; preamble_mass records the share that came through a preamble. Qwen3-8B, which wants to write "**" before the letter, is the case this exists for; on Qwen3.5-4B it changes nothing. benchmarks/shape777.py accepts --backend llamacpp with the same options, so the 37x21 systems fixture runs on GGUF, and reads the GPU name without a CUDA torch. Unit tests cover the new validation, the fan-out batch layout, branch removal, and the marginal fold; the real-GGUF test now also loads with three sequences and compares copy-shared to restore-shared. It passes on the PyPI CPU wheel and on the cu124 wheel with full offload.
…sults on a laptop GPU --llama-parallel and --llama-gpu-layers now default to "auto". For the layers that means not overriding llama.cpp's own default (-1, every layer, in current builds) and recording what the library did; the original backend forced 0. For the branches it means that nothing is asked of the user: the rows are already tokenized by the time shared mode runs, so the prefix length and every suffix length are known, and each state's fan-out holds as many questions as prefix + sum(suffixes) <= n_ctx allows, 32 at most. The context is sized once, for the longest single prompt, not multiplied by the sequence count: with a unified KV buffer every sequence sees the whole n_ctx (measured on this build, 16 sequences and n_ctx 4096 give 4096 per sequence against 256 without), and copied branches share the prefix cells. branches_per_decode in the shared timing records what was done. An integer still fixes n_seq_max; 1 keeps the state-restore path. Results on an RTX 3080 Laptop with the pinned Q4_K_M GGUF, all layers offloaded: authored144 direct 0.796 mean family balanced accuracy (Torch BF16 on a 3090: 0.813), ECE 0.063 out of fold (0.038), median allowed_token_mass 0.9997; shape777 1.40 / 9.21 / 10.88 decisions per second fresh, serial and parallel with 8 branches, 10.51 with auto sizing (Torch: 2.33 / 10.75 / 20.03), 16 to 19 argmax flips out of 777 between decode paths, which is this quantized model's noise floor. One decode per state is no faster than three: llama.cpp re-splits into n_ubatch micro-batches either way. Raw reports, predictions and calibration are committed with their checksums; phase-1 claims are untouched and verify_published.py passes. The real-GGUF test compared serial and shared logits with an exact tolerance that only held because both were the restore path; it now requires exactness when n_seq_max == 1 and 0.5 logits otherwise, decisions still equal. It passes on the PyPI CPU wheel and on the cu124 wheel with offload.
…the backend sits relative to JevBench
… DeltaNet kernel, next to the GGUF rows
…fore the GPU offload
…t has nothing to fold
Expose the existing per-workload temperature scaling as an installable command and wire evaluate's non-screen auto-decide/review operating point, without changing the frozen 96-row screening gate or published tables.
Measure ID-aligned flips, total variation, and position bias from committed predictions without changing default scores; optional --stabilize-order K averages permutation logits under a distinct prompt_version at Kx cost.
validate_row declares state may be a string, object or array and score() honours all three, but the prefix-sharing paths rejected a subset of valid states with "The fixed state prefix does not match every full prompt". The same row succeeded with one question and raised with two, and which states failed was a property of the tokenizer vocabulary, so a caller could not predict or avoid it. Both copies of _state_prefix ended with encode(text)[:-1], assuming that appending the next field's punctuation always merges with exactly one final token. That is a property of the state text, not a guarantee. Against the pinned Qwen/Qwen3.5-4B tokenizer at revision 851bf6e, a state ending in one of ( - = ? } produces a prefix one token too long, so it is not a token prefix of the full prompt at all: 5 of the 100 printable ASCII characters, and the state from the issue report, "Did the deploy use the same parameters)?", is one of them. serial.py carried the same duplicated helper with the same defect, so SerialPrefixScorer raised on exactly the same states even though the issue only reports score_shared. Both are fixed here. Replace the guess with the longest common token prefix between the truncated evidence and that text plus the separator that always follows it. The separator is taken from the payload already in hand rather than hardcoded, so the fix survives a change to the prompt wording. A consequence worth stating: when the boundary token merges forward, the shared prefix now stops just before it, which is the longest prefix that can be shared at all. tests/test_state_prefix.py drives both copies through a byte tokenizer that merges only at the evidence boundary, so the regression is covered without torch and without network. Four of its cases fail against the previous code. The fix is a no-op on every committed fixture. Audited with the pinned tokenizer over all 217 distinct states in shape777, authored144 and perturbations108: zero prefix differences against the previous implementation, and both copies now agree on all of them, so no committed evidence is affected and nothing needs regenerating. verify_published.py still reproduces all 69 claims and every results/raw checksum still matches.
The stub tokenizer only reproduced the string-state variant, so the test proved less than it appeared to: it would still have passed if the object case were broken. Objects and arrays are the shapes that actually fail in practice, because the value's own closing brace or quote sits right in front of the comma. Model both extremes instead of one hand-picked triple. PlainTokenizer never merges so the boundary costs exactly one token, which is the case the old code handled correctly; MergingTokenizer lets a token straddle the comma, which is the case it did not. Six states across both tokenizers and both copies: 48 cases, and 10 of them fail against the previous implementation. Verified against the pinned Qwen tokenizer that this is not a stub artefact: of 70 structured states ending in a merge-prone character, 11 produce an invalid prefix with the old code and 0 with the fix.
semif-score --backend sglang scores the direct, serial, shared, and reranker modes through the /v1/score endpoint of a running SGLang server that contains sgl-project/sglang#40826. SemIf still renders every prompt, checks the answer slots, and validates the input on the pinned reference tokenizer, so prompt_sha256, input_tokens, prompt_version, option order, and answer_token_ids equal the Torch values. The server only scores SemIf's token ids, with each row's answer slots as that item's candidates and return_token_logprobs, and the client uses the standard library, so there is no new dependency. Before the first row, the backend refuses an unreachable server, a server without per-item candidate scoring, another model or revision, a Hub model whose server reports no revision, a non-generation model, --load-format dummy, --enable-mis, --allow-auto-truncate, any --preferred-sampling-params, a server tokenizer that disagrees with the reference tokenizer, and a --max-tokens at or above the server's max_req_input_len. Every response must carry one finite log-probability per candidate and a usage.prompt_tokens equal to the token ids sent. Requests go straight to the server with no proxy, redirect, or retry, and SEMIF_SGLANG_API_KEY is sent only as a bearer header and never written to a row. option_logits hold the candidates' full-vocabulary log-probabilities, one constant per row away from raw logits, as the readout string says. cli.py now binds a reranker scorer per backend, so reranker mode reaches the SGLang scorer, and --sglang-url and --sglang-timeout require --backend sglang. docs/SGLANG.md covers the server launch and the refused settings, and the real-server tests read SEMIF_SGLANG_URL and SEMIF_SGLANG_RERANKER_URL.
# Conflicts: # README.md
# Conflicts: # src/semif_phase1/cli.py # tests/test_cli.py
select_permutations listed all n! permutations before picking K: 1.45 s at 10 options, factorial beyond. Up to 8 options the enumerate-and-shuffle path is kept, so existing stabilized results reproduce exactly (checked for n=2..8, K in 2,3,4,5,8,16, seeds 0,1,7). Above that, orders are drawn with random.Random(seed).sample until K distinct ones are chosen; identity and reverse are still first.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Brings
picksintomaster: upstream TheoLeeCJ/SemIf-OpenJev plus its open PRs and one fix.From upstream PRs:
score_shared),load_model(gpu_layers=, sequences=)semif_phase1.order)semif-calibrateCLI and calibrated-threshold gateFork fix (b7bcb30):
order.select_permutationslisted all n! option orders before picking K (1.45 s at 10 options, factorial beyond). It now samples lazily above 8 options. Outputs are byte-identical up to 8 options (checked n=2..8, K in 2,3,4,5,8,16, seeds 0,1,7); K=2 is always identity + reverse. For 9+ options with K>=3 the extra orders come from a different draw.Tests:
pytest tests192 passed, 4 skipped on picks at b7bcb30 (macOS, llama-cpp-python 0.3.35).Used by rsa-ecom/xb-decisions-agent PR TheoLeeCJ#14, which pins b7bcb30.