Skip to content

dsv4.1: Engram module and request history support - #39666

Merged
hnyls2002 merged 105 commits into
mainfrom
dsv4.1-engram
Sep 17, 2026
Merged

hnyls2002 merged 105 commits into
mainfrom
dsv4.1-engram

Conversation

@hnyls2002

@hnyls2002 hnyls2002 commented Sep 15, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • Add the DeepSeek-V4.1 Engram module (layers/engram.py): hashed n-gram lookup into per-layer fp8 tables with e8m0 block scales, gated into the hc residual stream. Hashing, table gather and the gate run on the existing kernels/ops/embeddings/engram_* Triton kernels; the torch paths are the ROCm/CPU fallbacks.
  • Give NgramEmbeddingManager the request-history bookkeeping the hasher needs: a per-request history row seeded on extend (ScheduleBatch.engram_history -> ForwardBatch.engram_history), a PD-decode seed for prebuilt requests, and update_after_verify for speculative verify commits (the caller lands with the DSpark wiring).
  • ModelConfig: use_engram (from engram_layer_ids) and ngram_embedding_n next to the existing use_ngram_embedding.

Changes

Hash layout

  • EngramLayout.from_config draws one ascending prime sequence in (layer, n-gram size, head) order; EngramHasher maps token ids through a tokenizer-normalized compressed vocabulary and keeps a [req_slots, n - 1] history table. Primality is a local deterministic Miller-Rabin, so no new dependency.

Table placement

  • Default: rows sharded over the TP group in device memory, reassembled with an all-reduce; under DP attention the lookup gathers every DP rank's ids first.
  • SGLANG_ENABLE_DSV41_ENGRAM_HOST_TABLE=1 keeps the tables in pinned host memory and gathers rows over the CPU link. SGLANG_DSV41_ENGRAM_HOST_TABLE_LAYOUT selects shared (default: one buffer per TP group, no lookup all-reduce; the ranks must share a PID namespace) or per_rank (one shard per rank, all-reduce kept, huge pages via MADV_COLLAPSE after dropping the checkpoint page cache). The mapping lives for the whole process.

Shared detectors and config

  • NgramEmbeddingManager keeps its LongCat token-table behavior unchanged for models without engram_layer_ids; the unused k field is dropped (no readers).

Verification

  • The compressed vocabulary size the tokenizer normalizes to is asserted against engram_compressed_vocab_size when the hasher is built, so a tokenizers behavior change fails at startup instead of silently rehashing.
  • Existing manager, scheduler and config CPU suites pass unchanged; the MLX runner stubs gain the two new ModelConfig fields.
  • Host-table smoke: shared and per_rank tables construct, register and report huge-page coverage; unknown layout values fail at startup.

CI States

Latest PR Test (Base): 🚫 Run #35185092582
Latest PR Test (Extra): ❌ Run #35190590426
Latest PR Test (AMD ROCm 10): 🚫 Run #35185092572

@hnyls2002
hnyls2002 added this pull request to stack #39667 September 15, 2026 22:32
@hnyls2002
hnyls2002 removed this pull request from stack #39667 September 15, 2026 22:43
@hnyls2002
hnyls2002 added this pull request to stack #39669 September 15, 2026 22:43
@hnyls2002
hnyls2002 removed this pull request from stack #39669 September 15, 2026 22:56
@hnyls2002
hnyls2002 added this pull request to stack #39672 September 15, 2026 22:58
@hnyls2002
hnyls2002 force-pushed the dsv4.1-chat branch 2 times, most recently from 6cbb581 to e350d0b Compare September 16, 2026 00:27
Base automatically changed from dsv4.1-chat to main September 17, 2026 03:49
# Conflicts:
#	python/sglang/srt/entrypoints/openai/chat_encoding.py
#	python/sglang/srt/entrypoints/openai/encoding_dsv41.py
#	python/sglang/srt/entrypoints/openai/serving_chat.py
#	python/sglang/srt/environ.py
#	python/sglang/srt/function_call/deepseekv41_detector.py
#	python/sglang/srt/parser/template_detection.py
@hnyls2002

Copy link
Copy Markdown
Collaborator Author

/rerun-test test_ngram_embedding_manager.py test_scheduler_hicache_events.py test_scheduler_chunked_req_gate.py test_auxiliary_output.py

@github-actions

github-actions Bot commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

Results for /rerun-test test_ngram_embedding_manager.py test_scheduler_hicache_events.py test_scheduler_chunked_req_gate.py test_auxiliary_output.py:

🚀 ubuntu-latest (4 tests): ✅ View workflow run

cd test/ && python3 registered/unit/model_executor/model_runner_components/test_ngram_embedding_manager.py
cd test/ && python3 registered/unit/managers/test_scheduler_hicache_events.py
cd test/ && python3 registered/unit/managers/test_scheduler_chunked_req_gate.py
cd test/ && python3 registered/unit/managers/test_auxiliary_output.py

@hnyls2002

Copy link
Copy Markdown
Collaborator Author

/tag-and-rerun-ci

@hnyls2002 hnyls2002 added the run-ci CI: run the baseline test suite on this PR label Sep 17, 2026
@hnyls2002
hnyls2002 merged commit 3401b75 into main Sep 17, 2026
188 of 230 checks passed
@hnyls2002
hnyls2002 deleted the dsv4.1-engram branch September 17, 2026 06:49
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

apple-silicon run-ci CI: run the baseline test suite on this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants