Skip to content

spec : update speculative-simple - #26904

Merged
ggerganov merged 3 commits into
masterfrom
gg/examples-fix-spec-simple
Aug 11, 2026
Merged

spec : update speculative-simple#26904
ggerganov merged 3 commits into
masterfrom
gg/examples-fix-spec-simple

Conversation

@ggerganov

@ggerganov ggerganov commented Aug 11, 2026

Copy link
Copy Markdown
Member

Overview

  • Update the speculative-simple example to the latest common/speculative.
  • Remove common_speculative_need_embd
  • Remove common_speculative_need_embd_nextn

Requirements

@ggerganov
ggerganov requested review from a team as code owners August 11, 2026 12:58
@github-actions github-actions Bot added documentation Improvements or additions to documentation server labels Aug 11, 2026
@ggerganov
ggerganov merged commit f785fc9 into master Aug 11, 2026
23 of 27 checks passed
@ggerganov
ggerganov deleted the gg/examples-fix-spec-simple branch August 11, 2026 16:52
gabe-l-hart added a commit to gabe-l-hart/llama.cpp that referenced this pull request Aug 12, 2026
Signed-off-by: Gabe Goodhart <ghart@us.ibm.com>

* origin/master: (383 commits)
  cmake :  introduce semantic versioning  (ggml-org#26839)
  gguf : harden loader against malformed tensor dims and metadata types (ggml-org#25596)
  kleidiai: Add runtime feature detection mechanism for aarch64/kleidiai (ggml-org#26076)
  model : disallow integer dflash sliding_window_pattern (ggml-org#26900)
  sync : ggml
  cmake : add config version support (ggml/1582)
  server : support slot save/restore with media inputs (ggml-org#26640)
  ui: add read_media tool (ggml-org#25877)
  opencl: default FA c8 cluster width to 16 on X1E (ggml-org#26433)
  tests : update speculative params (ggml-org#26925)
  vulkan: add TQ2_0 (ternary) support (ggml-org#25850)
  wavtokenizer-dec : bound posnet/convnext block_count against n_layer_all (ggml-org#26892)
  convert : handle per_layer_config in Gemma4 (transformers 5.15) (ggml-org#26882)
  opencl: use flat mv q5_k when weight exceeds image1d_buffer_t limit (ggml-org#26880)
  chat : fix muse-glimmer detection of tool calls after EOM (ggml-org#26879)
  ci : add missing release check (ggml-org#26923)
  CUDA: only disable CUDA graphs when mul_mat_id actually needs a stream sync (ggml-org#26802)
  cuda : add warp-per-row wkv7 kernel for single-token decode (ggml-org#26111)
  spec : update speculative-simple (ggml-org#26904)
  chat : tighten bare function parsing for Qwen models (ggml-org#26793)
  ...
huaxel pushed a commit to huaxel/CachyLLama that referenced this pull request Aug 12, 2026
* spec : update speculative-simple

* cont : simplify

* cont : clean-up
brittlewis12 pushed a commit to brittlewis12/llama.cpp that referenced this pull request Aug 17, 2026
* spec : update speculative-simple

* cont : simplify

* cont : clean-up
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation server

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants