Skip to content

fix(container): update image ghcr.io/ggml-org/llama.cpp (server-cuda-b9843 ➔ server-cuda-b9853) - #6619

Merged
drag0n141 merged 1 commit into
masterfrom
renovate/ghcr.io-ggml-org-llama.cpp-0.x
Jul 1, 2026
Merged

drag0n141 merged 1 commit into
masterfrom
renovate/ghcr.io-ggml-org-llama.cpp-0.x

Conversation

@drag0n141-bot

@drag0n141-bot drag0n141-bot Bot commented Jul 1, 2026

Copy link
Copy Markdown
Contributor

This PR contains the following updates:

Package Update Change
ghcr.io/ggml-org/llama.cpp patch server-cuda-b9843 → server-cuda-b9853

Configuration

📅 Schedule: (in timezone Europe/Berlin)

  • Branch creation
    • At any time (no schedule defined)
  • Automerge
    • At any time (no schedule defined)

🚦 Automerge: Disabled by config. Please merge this manually once you are satisfied.

♻ Rebasing: Whenever PR becomes conflicted, or you tick the rebase/retry checkbox.

🔕 Ignore: Close this PR and you won't be reminded about these updates again.


  • If you want to rebase/retry this PR, check this box

This PR has been generated by Mend Renovate.

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ghcr.io/ggml-org/llama.cpp: server-cuda-b9843 → server-cuda-b9853

Verdict: Safe to merge

Small patch bump (10 build numbers) in llama.cpp's rolling b#### release scheme. Reviewed every intermediate commit between b9843 and b9853 via the GitHub compare API and per-tag release notes:

  • ggml-webgpu: add NVFP4 support (#25143) — WebGPU backend, not used (CUDA image).
  • HIP: hipBLAS tuning for gfx900 (#24588) — HIP backend, not used.
  • vulkan: matmul roll-bk fix for Asahi Linux (#24663) — Vulkan backend, not used.
  • CUDA: fix Gemma E4B MTP FlashAttention (#25148) — model-specific to Gemma E4B MTP; this deployment runs Qwen3.5-4B.
  • CUDA: fix get_rows_back grid-y clamp for >65535 rows (#25103) — backward/training-only op, not exercised by the server binary.
  • common,server: bracket IPv6 literals in URL authority/logs (#25140) — cosmetic/log-formatting fix; server binds 0.0.0.0, not IPv6.
  • model: register t_layer_inp for qwen3next, fixing a crash in speculative-decoding drafters (DFlash/EAGLE3) (#25141) — Qwen3.5's hybrid architecture is qwen3next-derived, but this compose file doesn't configure a draft model / speculative decoding, so the previously-crashing path isn't in use.
  • CUDA: fix integer truncation/overflow in flash_attn_mask_to_KV_max KQ mask strides, relevant only at very large context/ubatch sizes (#24945) — this deployment sets LLAMA_ARG_CTX_SIZE=8192, well below the range where the overflow could trigger.
  • opencl: initial q1_0 support (#25160) — OpenCL backend, not used.
  • ui: remove PWA navigate fallback (#25174) — the deployment sets LLAMA_ARG_NO_UI=1.

No breaking changes or deprecations in this range, and none of the fixes intersect with how docker/nas01/llama/docker-compose.yaml is configured (CUDA backend, no UI, no speculative decoding, 8192 context, non-Gemma model). Only file changed is the image tag bump.

Sources consulted:

@drag0n141
drag0n141 merged commit 6167dc2 into master Jul 1, 2026
2 checks passed
@drag0n141
drag0n141 deleted the renovate/ghcr.io-ggml-org-llama.cpp-0.x branch July 1, 2026 08:37
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant