fix(container): update image ghcr.io/ggml-org/llama.cpp (server-cuda-b9843 ➔ server-cuda-b9853) - #6619
Merged
Conversation
…b9843 ➔ server-cuda-b9853)
There was a problem hiding this comment.
ghcr.io/ggml-org/llama.cpp: server-cuda-b9843 → server-cuda-b9853
Verdict: Safe to merge
Small patch bump (10 build numbers) in llama.cpp's rolling b#### release scheme. Reviewed every intermediate commit between b9843 and b9853 via the GitHub compare API and per-tag release notes:
ggml-webgpu: add NVFP4 support (#25143) — WebGPU backend, not used (CUDA image).- HIP: hipBLAS tuning for gfx900 (#24588) — HIP backend, not used.
vulkan: matmul roll-bk fix for Asahi Linux (#24663) — Vulkan backend, not used.- CUDA: fix Gemma E4B MTP FlashAttention (#25148) — model-specific to Gemma E4B MTP; this deployment runs Qwen3.5-4B.
- CUDA: fix
get_rows_backgrid-y clamp for >65535 rows (#25103) — backward/training-only op, not exercised by theserverbinary. common,server: bracket IPv6 literals in URL authority/logs (#25140) — cosmetic/log-formatting fix; server binds0.0.0.0, not IPv6.model: registert_layer_inpfor qwen3next, fixing a crash in speculative-decoding drafters (DFlash/EAGLE3) (#25141) — Qwen3.5's hybrid architecture is qwen3next-derived, but this compose file doesn't configure a draft model / speculative decoding, so the previously-crashing path isn't in use.- CUDA: fix integer truncation/overflow in
flash_attn_mask_to_KV_maxKQ mask strides, relevant only at very large context/ubatch sizes (#24945) — this deployment setsLLAMA_ARG_CTX_SIZE=8192, well below the range where the overflow could trigger. opencl: initial q1_0 support (#25160) — OpenCL backend, not used.ui: remove PWA navigate fallback (#25174) — the deployment setsLLAMA_ARG_NO_UI=1.
No breaking changes or deprecations in this range, and none of the fixes intersect with how docker/nas01/llama/docker-compose.yaml is configured (CUDA backend, no UI, no speculative decoding, 8192 context, non-Gemma model). Only file changed is the image tag bump.
Sources consulted:
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This PR contains the following updates:
server-cuda-b9843→server-cuda-b9853Configuration
📅 Schedule: (in timezone Europe/Berlin)
🚦 Automerge: Disabled by config. Please merge this manually once you are satisfied.
♻ Rebasing: Whenever PR becomes conflicted, or you tick the rebase/retry checkbox.
🔕 Ignore: Close this PR and you won't be reminded about these updates again.
This PR has been generated by Mend Renovate.