Skip to content

llama: R-SWA reference sliding window attention for Unlimited-OCR - #24975

Open
sfallah wants to merge 1 commit into
ggml-org:masterfrom
sfallah:sf/unlimited-ocr-rswa
Open

llama: R-SWA reference sliding window attention for Unlimited-OCR#24975
sfallah wants to merge 1 commit into
ggml-org:masterfrom
sfallah:sf/unlimited-ocr-rswa

Conversation

@sfallah

@sfallah sfallah commented Jun 24, 2026

Copy link
Copy Markdown
Contributor

Overview

Adds baidu/Unlimited-OCR to mtmd. It is DeepSeek-OCR v1 with a different decoder attention, R-SWA (Reference Sliding Window Attention): every token sees the whole prompt (image + text) plus the last 128 generated tokens. With plain SWA the image would slide out of view.

Changes:

  • new LLAMA_SWA_TYPE_REFERENCE and its mask rule
  • the KV cache records where the prompt ends (n_ref) when the first generated token arrives, and keeps it correct across seq_rm / seq_cp / seq_add / seq_div and session save/load (LLAMA_SESSION_VERSION and LLAMA_STATE_SEQ_VERSION bumped)
  • one full KV cache, no eviction; llama_model_n_swa returns 0 for this arch so the server does not treat it as SWA
  • V cache defaults to F32 for R-SWA (F16 garbles dense tables); a quantized -ctv is still honored

Only single-page parity for now.

Validation

bf16, single page, HF reference in its release config (bf16 weights, R-SWA, CUDA):

CER
llama.cpp (Metal and CUDA) 0.1574
HF reference 0.1591

Re-validated 2026-08-19 on current master (incl. #26727), same numbers.

Design note

R-SWA sits in core (swa type, mask rule, n_ref in llama_kv_cache) and is only enabled for deepseek2-ocr. It could live in the model code instead; open to moving it there.
Anyhow my immediate aim is that Unlimited-OCR is supported in llama.cpp with its original R-SWA.

Limitations

The KV cache guesses the prompt end from the first single-token decode that requests output. Where the guess is off:

  • speculative decode (contexts allowing more than one output per sequence): no guess, plain causal attention
  • a prompt whose last ubatch is a single token: boundary one token early
  • context shift: boundary moves to the current position
  • seq_cp copies the boundary as-is

An explicit prompt-length signal from the caller (mtmd knows it) would remove the guessing. Open to that as well.

How to run

GGUF models: sabafallah/Unlimited-OCR-GGUF

build/bin/llama-mtmd-cli -hf sabafallah/Unlimited-OCR-GGUF:bf16 \
  --image tools/mtmd/test-1.jpeg -p "document parsing." \
  --chat-template deepseek-ocr \
  --temp 0 --flash-attn off --no-warmup \
  -n 4096 -c 16384 \
  --dry-multiplier 0.8 --dry-base 1.75 --dry-allowed-length 35 \
  --dry-penalty-last-n 128 --dry-sequence-breaker none

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES - this is largely an AI-written first draft, opened to start the discussion. I designed the R-SWA approach (from the model's paper), directed and reviewed the implementation; the code itself is mostly AI-generated. I take responsibility for the contents.

@github-actions github-actions Bot added model Model specific examples python python script changes labels Jun 24, 2026
@jason-ni

Copy link
Copy Markdown
Contributor

It seems the impl in this branch has numeric precision issue:
https://huggingface.co/sabafallah/Unlimited-OCR-GGUF/discussions/1

@o7si could you please have a test on image in that discussion? Thanks.

@sfallah

sfallah commented Jun 27, 2026

Copy link
Copy Markdown
Contributor Author

It seems the impl in this branch has numeric precision issue: https://huggingface.co/sabafallah/Unlimited-OCR-GGUF/discussions/1

@o7si could you please have a test on image in that discussion? Thanks.

@jason-ni
yes, I will.
Thanks for reporting this.

sfallah added a commit to sfallah/llama.cpp that referenced this pull request Jun 27, 2026
The DeepSeek-OCR / Unlimited-OCR decoder reads dense layout (e.g. tables)
by attending over the always-visible visual prefix. With the default F16
V-cache, those value vectors are truncated enough to garble the output:
table headers come out as """ / ">" (reported on ggml-org#24975), while the
official HF reference parses them correctly.

The HF reference accumulates attention in F32, so match it by promoting the
F16 V-cache default to F32 for LLM_ARCH_DEEPSEEK2OCR. An explicit
lower-precision -ctv (e.g. q8_0) is still honored. This is the in-graph
equivalent of running with --cache-type-v f32. It is not the cuBLAS compute
mode: the headers are emitted deep in autoregressive decode (a mat-vec path
that bypasses cuBLAS), so FORCE_CUBLAS_COMPUTE_32F has no effect; F16 V
storage/accumulation is what truncates.

Verified on the reported image (parses cleanly with no flags) and with
tools/mtmd/tests/test-deepseek-ocr.py: all cases pass and improve, no
regression (v1 0.2626, v2 0.6877, unlimited 0.1641 CER).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015dykwunMpwXWxHPVhbjhiK
@github-actions github-actions Bot added mtmd Related to multimodal functionality (video/image/audio) conversion labels Jun 27, 2026
@sfallah

sfallah commented Jul 12, 2026

Copy link
Copy Markdown
Contributor Author

@jason-ni
the root cause was found and fixed.
It was the DRY flags in my run example, not the model.
See details in the HF discussion: https://huggingface.co/sabafallah/Unlimited-OCR-GGUF/discussions/1

@sfallah
sfallah force-pushed the sf/unlimited-ocr-rswa branch from ac42aca to 8d14be3 Compare August 19, 2026 06:53
@sfallah
sfallah marked this pull request as ready for review August 25, 2026 07:11
@sfallah
sfallah requested review from CISC and ggerganov as code owners August 25, 2026 07:11
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

conversion examples model Model specific mtmd Related to multimodal functionality (video/image/audio) python python script changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants