Repository navigation
dsv4.1: Engram module and request history support - #39666
Merged
Merged
Conversation
hnyls2002
requested review from
BBuf,
Edwardf0t1,
Fridge003,
HaiShaw,
Ying1123,
ch-wan,
ispobock,
merrymercy and
xiezhq-hermann
September 15, 2026 22:31
hnyls2002
added this pull request to stack #39667
September 15, 2026 22:32
hnyls2002
removed this pull request from stack #39667
September 15, 2026 22:43
hnyls2002
force-pushed
the
dsv4.1-engram
branch
from
September 15, 2026 22:43
43bdd6c to
e21bf81
Compare
hnyls2002
force-pushed
the
dsv4.1-chat
branch
from
September 15, 2026 22:43
2f28471 to
72f3422
Compare
hnyls2002
added this pull request to stack #39669
September 15, 2026 22:43
hnyls2002
removed this pull request from stack #39669
September 15, 2026 22:56
hnyls2002
force-pushed
the
dsv4.1-chat
branch
from
September 15, 2026 22:57
72f3422 to
3e83da8
Compare
hnyls2002
force-pushed
the
dsv4.1-engram
branch
from
September 15, 2026 22:57
e21bf81 to
d388cfc
Compare
hnyls2002
added this pull request to stack #39672
September 15, 2026 22:58
hnyls2002
force-pushed
the
dsv4.1-engram
branch
from
September 15, 2026 23:21
d388cfc to
fa664c2
Compare
hnyls2002
force-pushed
the
dsv4.1-chat
branch
from
September 15, 2026 23:21
3e83da8 to
99c9a8f
Compare
hnyls2002
force-pushed
the
dsv4.1-engram
branch
from
September 15, 2026 23:46
fa664c2 to
a42a85b
Compare
hnyls2002
force-pushed
the
dsv4.1-chat
branch
2 times, most recently
from
September 16, 2026 00:27
6cbb581 to
e350d0b
Compare
hnyls2002
force-pushed
the
dsv4.1-engram
branch
from
September 16, 2026 00:27
a42a85b to
2b46619
Compare
hnyls2002
force-pushed
the
dsv4.1-chat
branch
from
September 16, 2026 01:12
e350d0b to
c730e0f
Compare
hnyls2002
force-pushed
the
dsv4.1-engram
branch
from
September 16, 2026 01:12
2b46619 to
4c12309
Compare
hnyls2002
force-pushed
the
dsv4.1-chat
branch
from
September 16, 2026 01:52
c730e0f to
096ca10
Compare
# Conflicts: # python/sglang/srt/entrypoints/openai/chat_encoding.py # python/sglang/srt/entrypoints/openai/encoding_dsv41.py # python/sglang/srt/entrypoints/openai/serving_chat.py # python/sglang/srt/environ.py # python/sglang/srt/function_call/deepseekv41_detector.py # python/sglang/srt/parser/template_detection.py
…yout factory; env section
…cuda gate; image_token_id from caller; host table lifetime note
Collaborator
Author
|
/rerun-test test_ngram_embedding_manager.py test_scheduler_hicache_events.py test_scheduler_chunked_req_gate.py test_auxiliary_output.py |
Contributor
|
Results for 🚀 |
Collaborator
Author
|
/tag-and-rerun-ci |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
layers/engram.py): hashed n-gram lookup into per-layer fp8 tables with e8m0 block scales, gated into the hc residual stream. Hashing, table gather and the gate run on the existingkernels/ops/embeddings/engram_*Triton kernels; the torch paths are the ROCm/CPU fallbacks.NgramEmbeddingManagerthe request-history bookkeeping the hasher needs: a per-request history row seeded on extend (ScheduleBatch.engram_history->ForwardBatch.engram_history), a PD-decode seed for prebuilt requests, andupdate_after_verifyfor speculative verify commits (the caller lands with the DSpark wiring).ModelConfig:use_engram(fromengram_layer_ids) andngram_embedding_nnext to the existinguse_ngram_embedding.Changes
Hash layout
EngramLayout.from_configdraws one ascending prime sequence in (layer, n-gram size, head) order;EngramHashermaps token ids through a tokenizer-normalized compressed vocabulary and keeps a[req_slots, n - 1]history table. Primality is a local deterministic Miller-Rabin, so no new dependency.Table placement
SGLANG_ENABLE_DSV41_ENGRAM_HOST_TABLE=1keeps the tables in pinned host memory and gathers rows over the CPU link.SGLANG_DSV41_ENGRAM_HOST_TABLE_LAYOUTselectsshared(default: one buffer per TP group, no lookup all-reduce; the ranks must share a PID namespace) orper_rank(one shard per rank, all-reduce kept, huge pages viaMADV_COLLAPSEafter dropping the checkpoint page cache). The mapping lives for the whole process.Shared detectors and config
NgramEmbeddingManagerkeeps its LongCat token-table behavior unchanged for models withoutengram_layer_ids; the unusedkfield is dropped (no readers).Verification
engram_compressed_vocab_sizewhen the hasher is built, so atokenizersbehavior change fails at startup instead of silently rehashing.ModelConfigfields.sharedandper_ranktables construct, register and report huge-page coverage; unknown layout values fail at startup.CI States
Latest PR Test (Base): 🚫 Run #35185092582
Latest PR Test (Extra): ❌ Run #35190590426
Latest PR Test (AMD ROCm 10): 🚫 Run #35185092572