Skip to content

Add an SGLang server backend for all four scoring modes - #55

Open
rwang5203 wants to merge 1 commit into
TheoLeeCJ:masterfrom
rwang5203:sglang-backend
Open

rwang5203 wants to merge 1 commit into
TheoLeeCJ:masterfrom
rwang5203:sglang-backend

Conversation

@rwang5203

Copy link
Copy Markdown

Motivation

sgl-project/sglang#40826 added per-item candidate scoring to SGLang's /v1/score. This adds an SGLang backend so SemIf can score rows on an SGLang server.

Modifications

  • Add --backend sglang for the direct, serial, shared, and reranker modes in src/semif_phase1/sglang_backend.py. SemIf still renders and validates prompts, and the server only scores SemIf's token ids.
  • Check the server before scoring and refuse unsupported settings with a message that names the fix.
  • Add --sglang-url, --sglang-timeout, docs/SGLANG.md, and tests.

Related Issues

Requires an SGLang build with sgl-project/sglang#40826, which is SGLang main or a nightly from 0.5.21.dev20260925 on.

Accuracy Test

python -m sglang.launch_server --model-path Qwen/Qwen3.5-4B --revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a
semif-score --backend sglang --mode direct --model Qwen/Qwen3.5-4B --revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a --input examples/decisions.jsonl --output results-sglang-direct.jsonl
pytest -q

semif-score --backend sglang scores the direct, serial, shared, and
reranker modes through the /v1/score endpoint of a running SGLang
server that contains sgl-project/sglang#40826. SemIf still renders
every prompt, checks the answer slots, and validates the input on the
pinned reference tokenizer, so prompt_sha256, input_tokens,
prompt_version, option order, and answer_token_ids equal the Torch
values. The server only scores SemIf's token ids, with each row's
answer slots as that item's candidates and return_token_logprobs, and
the client uses the standard library, so there is no new dependency.

Before the first row, the backend refuses an unreachable server, a
server without per-item candidate scoring, another model or revision,
a Hub model whose server reports no revision, a non-generation model,
--load-format dummy, --enable-mis, --allow-auto-truncate, any
--preferred-sampling-params, a server tokenizer that disagrees with
the reference tokenizer, and a --max-tokens at or above the server's
max_req_input_len. Every response must carry one finite
log-probability per candidate and a usage.prompt_tokens equal to the
token ids sent. Requests go straight to the server with no proxy,
redirect, or retry, and SEMIF_SGLANG_API_KEY is sent only as a bearer
header and never written to a row.

option_logits hold the candidates' full-vocabulary log-probabilities,
one constant per row away from raw logits, as the readout string says.
cli.py now binds a reranker scorer per backend, so reranker mode
reaches the SGLang scorer, and --sglang-url and --sglang-timeout
require --backend sglang. docs/SGLANG.md covers the server launch and
the refused settings, and the real-server tests read SEMIF_SGLANG_URL
and SEMIF_SGLANG_RERANKER_URL.
tvpavan added a commit to tvpavan/SemIf-OpenJev that referenced this pull request Oct 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant