Skip to content

Scheduler UMA ring buffer (+ sanitizer and fixes) - #27311

Open
pwilkin wants to merge 17 commits into
ggml-org:masterfrom
pwilkin:sched-uma-ring
Open

Scheduler UMA ring buffer (+ sanitizer and fixes)#27311
pwilkin wants to merge 17 commits into
ggml-org:masterfrom
pwilkin:sched-uma-ring

Conversation

@pwilkin

@pwilkin pwilkin commented Aug 18, 2026

Copy link
Copy Markdown
Member

Overview

As per discussion in #25863 , implement the ring buffer mechanism for input tensors, on top of @am17an 's sanitizer (#26167), plus additional hardening for the sanitizer and extra scheduler fixes (there was an error with duplicating pinned memory that was a view's base).

Additional information

Makes host buffers viable again, ping @ORippler for feedback / tests on CUDA integrated boxes. No measurable efficiency losses.

Supersedes #25863 , #26167 , #26225

Moved all scheduler documentation to a dedicated doc.

Requirements

@pwilkin
pwilkin requested review from a team, CISC and ggerganov as code owners August 18, 2026 09:24
@github-actions github-actions Bot added documentation Improvements or additions to documentation ggml changes relating to the ggml tensor library for machine learning CUDA Related to the CUDA backend labels Aug 18, 2026
@pwilkin

pwilkin commented Aug 18, 2026

Copy link
Copy Markdown
Member Author

@ggerganov this should enable proper host memory usage on UMA devices other than Metal (i.e. ones that actually pin host memory), once this is verified the props.integrated flag can be reenabled on CUDA devices.

@pwilkin

pwilkin commented Aug 19, 2026

Copy link
Copy Markdown
Member Author

As suggested by @am17an I've asked Claude to add a post-mortem doc about my chases of the various bugs during this PR here: https://pwilkin.github.io/llama-scheduler/ring.html

@ORippler

Copy link
Copy Markdown
Collaborator

@aendk Does the behavior here conform to #27258, where we try to add tests to formalize a backend's behavior?

@ORippler ORippler left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for taking a stab at formalizing this! I'll take a tour on DGX/RTX Spark later on and report back on perf

Comment thread docs/development/backend-scheduler.md Outdated
Comment on lines +28 to +31
2. **Op support and operand location.** Otherwise the highest priority backend that supports the
op is used, preferring the backend holding the operands. Ops reading tensors in a buffer marked
`GGML_BACKEND_BUFFER_USAGE_WEIGHTS` prefer that buffer's backend, so that weights are not
copied.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

How are these priorities determined?

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Augmented the docs to include the entire algorithm.

Comment thread docs/development/backend-scheduler.md Outdated
says the memory is dead - so a later split on the owning backend can be given the same memory and
overwrite it.

Such tensors are pinned for the lifetime of the graph with `ggml_gallocr_pin_tensor()`, keeping

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For my understanding. We are talking about pinning the memory on backend A for a node later on consumed by backend B. This will hold also for the path A -> C -> A -> B (which is why ggml_set_output is insufficient, as it garuantuees only within single cgraph execution).

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yeah.

Comment on lines +135 to +141
That move is only made when the target backend supports the op. Support can be conditional on
the tensor types - CUDA runs `GGML_OP_SET` only for F32 and I32 - so the backend owning the
aliased memory is not guaranteed to be able to run the op writing into it. There is no correct
placement in that case: the scheduler copies operands into a split, never results out of one, so
whichever backend runs the op, the write cannot reach the aliased memory. The node is left where
the earlier passes put it, which is what the scheduler did before this rule existed, and the
reason is logged under `GGML_SCHED_DEBUG`.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

so the backend owning the
aliased memory is not guaranteed to be able to run the op writing into it. There is no correct
placement in that case: the scheduler copies operands into a split, never results out of one, so
whichever backend runs the op, the write cannot reach the aliased memory.

Wouldn't the correct behavior be to check for this during node-placement + expansion time?

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It might be, but I didn't want to do a full scheduler refactor for this. One is probably due anyway since the work in this PR outlined quite a few issues with the scheduler and quite a few of the current solutions (esp. regarding the "special" input / output tensors) seem really hacky.

Comment on lines +149 to +152
It maintains a vector clock per actor - the host thread and each backend - and a shadow map of
which actor last read or wrote every byte of every buffer. Synchronization points (backend
synchronize, event record, event wait, event synchronize, async copies) advance those clocks. When
an access conflicts with a previous one and no happens-before edge connects them, it reports:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we make this respect inter-backend events? AFAIK it's not current POR, but CUDA offers cross-stream events like such

        CUDA_CHECK(cudaEventRecord(cuda_ctx_src->copy_event, cuda_ctx_src->stream()));

        // wait on dst stream for the copy to complete
        CUDA_CHECK(cudaStreamWaitEvent(cuda_ctx_dst->stream(), cuda_ctx_src->copy_event, 0));

Comment thread ggml/src/ggml-alloc.c Outdated
Comment thread docs/development/backend-scheduler.md Outdated
Comment thread ggml/src/ggml-cuda/ggml-cuda.cu
@github-actions github-actions Bot added the testing Everything test related label Aug 20, 2026
@DeMaulwurfn

Copy link
Copy Markdown

I tested #26225 on this box a few days ago and reported the numbers over there, so here is the same again
on its successor. Same machine, same model same method, so everything is directly comparable to what i
posted before. Also nobody seems to be covering the AMD side, the ping in your description only mentions
CUDA integrated boxes. Short version, correct, no speed loss, and the sanitizer stays quiet.

The box is unchanged from last time, a Framework Desktop with Strix Halo, gfx1151, 96 GB unified memory as
a BIOS carve out, Fedora 44 with kernel 7.1.8. Model is still Qwen3.5-122B-A10B as unsloth UD-Q4_K_XL with
-c 262144, flags are -ngl 99 -dio --no-mmap --jinja -fa on --cache-type-k q8_0 --cache-type-v q8_0. Built
from your branchs own .devops/rocm.Dockerfile with one change, AMDGPU_TARGETS=gfx1151 only because i have
exactly one card. Base image rocm/dev-ubuntu-24.04:7.2.1-complete, same as the official recipe. The commit
i tested is b16cc7d, so after the rebase and after the ring allocation hardening. All throughput numbers
come from the servers own timings object, not from wall clock.

For correctness i use the same needle test as before. A marker KANARIE-<8 hex> goes at a defined position
into a deterministic filler prompt, it has no relation to prompt length or position so it can not be
guessed, and the model has to either echo it or answer NICHTGEFUNDEN, which is german for "not found". Six
prompt lengths times five positions, so 30 cells. Result on your branch is 30 of 30, and it stays 30 of 30
with the sanitizer switched on.

For comparison, the 13 of 30 that i reported in #26225 was measured on build 10454, wich is an ancestor of
your base here. The failing cells were all the ones where the marker sits more then about 1024 tokens away
from the end of the prompt. Nothing, that fixes this, got merged between 10454 and your base, i checked the
log, so the base still has the bug. I did not build the base seperately this time, that is the one gap in
my chain. I did look, at what your base commit actually is though, 60eeeb6 "cuda : skip UMA override for
HIP builds", and it only touches memory reporting, two preprocessor lines, so it cant be what heals this.

The sanitizer is the part i find most interesting, that one is new compared to my last report. First
attempt gave me nothing at all, not even the summary line, and that turned out to be my own fault, the
proxy i normally run in front of llama-server swallows its stderr. So i started llama-server directly
instead and got this:

ggml-sched-sanitize: 0 race(s) reported

That was over four cells at 2000 and 4000 tokens, not the full grid, so its four requests and not thirty.
But zero is zero, the summary line proves the thing was actually running, wich matters because "no races"
and "sanitizer never ran" look exactly the same in a log otherwise, both empty.

Now speed. Prefill in t/s on the same machine and model, 9776 is the last build before #24233 and still
what i run in production because of the corruption.

prompt tokens 9776 (before #24233) 10454 (master, no fix) your branch b16cc7d
3373 274,5 380,0 363,1
26357 109,0 318,7 312,4
52592 64,2 265,3 261,6
104777 35,4 196,9 193,1

At 104777 tokens thats 193,1 against 196,9 with no fix at all, so 1,9 percent apart while my measurement
noise on prefill is also 1,9 percent. Your ring buffer keeps the whole scaling. In wall clock, reading a
100k token prompt takes 49 minutes on 9776 and 9 minutes here, so for long context work on an APU this is
the difference between usable and not. Fair warning that 10454 and your base are not the same commit, there
are a couple of weeks of master between them, so please dont read the small gaps at the short end as your
patch.

Decode is unchanged as well, 9,62 t/s at 104777 tokens against 9,61 on unpatched master

One thing that is not about your PR at all, but you might want to know since it showed up in the same runs.
The build i made for #26225, wich sits on master from end of july, does 11,67 t/s decode at that same
104777 token point. Thats about 21 percent faster then both current master and your branch. Prefill is fine
on all of them, its only decode and only at long context. So something in master between end of july and
now cost that, and your branch just inherits it. Happy to open a seperate issue with the numbers if thats
useful, i did not want to clutter this one.

Anyway from where i sit, this looks good, better then the previous approach because the sanitizer confirms
it structurally and not just by symptoms. If you want anything else run on gfx1151 just say so, the build
pipeline is allready set up so its about an hour turnaround. Test scripts are plain perl and curl if anyone
wants them.

am17an and others added 12 commits August 21, 2026 15:34
The sanitizer aborts on the first race, so enumerating the races in a
workload takes one run per race. Add GGML_SCHED_SANITIZE_NONFATAL=1 to
report each one and keep going, with a count printed at exit.

This is safe because access_range() updates the shadow state after
report(), so a range that has been reported does not keep re-firing.

Assisted-By: Claude Opus 5
llama.cpp writes several graph inputs straight through tensor->data after
checking the buffer is host-visible, bypassing ggml_backend_tensor_set.
The scheduler sanitizer only hooks the backend API, so it could not see
those writes and silently missed every race against them.

Add ggml_backend_tensor_set_direct(), a no-op unless GGML_SCHED_SANITIZE
is set, to announce such a write. Wrap it in llama_host_write(), which
also carries the is_host assertion, and use it in place of the bare
assertion at all 22 direct-write sites. Folding the announcement into the
idiom that already marked these sites keeps the two from drifting apart.

Measured on a multi-ubatch prefill on an integrated GPU (gfx1151): the
run reports 87 races, 69 of which were previously invisible. 40 are on
attn_inp_kq_mask alone, which was entirely unseen before.

Assisted-By: Claude Opus 5
Most graph inputs were never named, so they showed up as leaf_N in
GGML_SCHED_DEBUG output, in scheduler sanitizer reports and in any other
tooling that identifies tensors by name. Of the 24 inputs, 15 were
anonymous.

Name them, following the existing conventions: cb() where the builder is
an llm_graph_context method, ggml_set_name() in the free functions and
memory-module methods that have no callback to hand.

The kq_mask helper also stamped the same "attn_inp_kq_mask" on every
variant, so the base, SWA, MLA and LID masks were indistinguishable from
one another. Give it a name parameter and pass a distinct name at each
call site.

No functional change: a prefill that reports 87 sanitizer races reports
the same 87 before and after, with every tensor now identified.

Assisted-By: Claude Opus 5
ggml_cuda_graph_update_required() treats an unchanged non-zero cgraph->uid
as proof that nothing the captured graph baked in has changed, and returns
early without comparing node properties. That holds today only because the
uid is re-stamped in ggml_backend_sched_split_graph(), and splitting always
precedes re-allocation.

The failure mode if it ever stops holding is bad: there is no fallback, the
executable graph is simply replayed against stale addresses and the results
are wrong with no diagnostic. A change that re-points tensors without going
through a re-split - pipelining or ring-buffered graph inputs, say - would
hit exactly this.

Verify the promise rather than trusting it: on the fast path, recompute the
node properties and abort if they differ. On by default in debug builds and
available via GGML_CUDA_GRAPH_VERIFY_UID elsewhere, since release builds are
where such a change is actually exercised. Document the contract at both
ends under [TAG_CUDA_GRAPH_UID].

Checked by freezing the split uid to a constant: without the check a
multi-ubatch prefill completes silently with stale addresses (exit 0), with
it enabled the run aborts naming the offending node. No false positives -
the fast path is taken 38 times in a 40-token decode and stays quiet, with
decode speed unchanged.

Assisted-By: Claude Opus 5
…t memory

Graph inputs are pinned to the last (CPU) backend, and llama.cpp hands the
scheduler a device host buffer type for that slot. On an integrated GPU the
device also accepts that buffer type, so ggml_backend_sched_buffer_supported()
is true, no split input copy is made, and the device reads the very memory the
host thread writes. Nothing then stops the host from writing the next ubatch's
inputs while the device is still reading the previous one.

Detect that case and give the graph inputs a ring of 2 instead of a single
buffer. The predicate tests the exact condition that elides the copy - a host
buffer type in the CPU slot that some compute backend accepts - rather than
guessing at "is this an APU", so backends that deliberately refuse to compute
on pinned host memory are unaffected. Verified on a Strix Halo APU: ROCm gets
a ring, Vulkan on the same device does not, and neither does CPU-only.

Note the buffer type has to be resolved directly here, since sched->bufts[] is
not populated until further down.

This is detection and allocation only - rotating the ring on the graph reuse
path is a separate change, and until it lands the race is reduced but not
fixed (87 reports to 61 on a multi-ubatch prefill, and a deeper ring only
widens the window: 49 at N=3, 42 at N=4, never 0). Depth 2 is the default
because the extra input memory is not free: measured +8 MiB at n_ctx 8192 /
n_ubatch 512, and +128 MiB at 32768 / 2048, all of it in the host buffer.

GGML_SCHED_UMA_RING overrides the detection - 0 or 1 disables, larger sets
the depth - as an escape hatch and to A/B the cost.

Assisted-By: Claude Opus 5
Enabling the ring was not enough on its own. The addresses only move as a
side effect of re-splitting, so on the graph reuse path - which re-runs
neither the split nor the allocator - bumping cur_copy moved nothing, and
the host still overwrote inputs the device was reading.

Add ggml_backend_sched_prepare_inputs(), which llama.cpp calls before
writing the inputs of a reused graph. It steps onto the next ring slot,
waits for that slot's previous reader, and re-points each graph input by
hand from addresses captured when the allocator placed them. Nodes hold
ggml_tensor pointers and split_graph does not rewrite node->src[] for
graph inputs, so re-pointing the tensor moves every reader with it.

It replaces an unconditional ggml_backend_sched_synchronize() that was
already at that call site for pipeline parallelism, so multi GPU setups
now rotate instead of stalling on every reused graph.

Three things this needs to get right:

- An input reached only through a view of it - the recurrent state copy -
  was never registered, so it never rotated and kept racing. Resolve
  through view chains when registering, and re-derive the views after
  re-pointing, since a view's address is baked at allocation time.

- The wait belongs on the allocate path too. Rotating onto a slot is only
  safe once its previous reader is finished, however the addresses got
  there, so both paths share it.

- [TAG_CUDA_GRAPH_UID] re-stamp the split uids when re-pointing. Moving
  addresses without a re-split is exactly the case the uid fast path
  assumes cannot happen, and a captured graph would otherwise replay
  against the old ones.

Rotation is confined to where it is needed: anything that synchronizes
clears the flag, so single token decode inside a sampling loop holds its
slot, keeps its addresses stable and keeps its captured graphs.

Measured on a Strix Halo APU, multi ubatch prefill, sanitizer race reports
to zero:

  n_ubatch 512 (rebuild every ubatch)   87 -> 0
  n_ubatch 32  (11 graphs reused)       21 -> 0

Decode stays clean, and so do the CPU only and Vulkan controls. Over a 200
token generation the reuse path rotates zero times, the captured graph is
reused 198 times, and warmup completes once and never resets - so decode
addresses are as stable as before. GGML_CUDA_GRAPH_VERIFY_UID is quiet
across both, including the 9 reuse path rotations.

Assisted-By: Claude Opus 5
The ring buffer arrived as a running commentary spread over the scheduler,
the CUDA backend and llama.cpp, which is the wrong place for it: the parts
that matter to a caller or to a backend author were only discoverable by
reading the implementation, and the parts that only matter to the
implementation were repeated at every site that touched it.

Move the description into the "Backend scheduler" block in ggml-backend.h
as its own section, covering when the ring engages and why, what it costs,
when it rotates, and the two obligations it creates - callers must call
ggml_backend_sched_prepare_inputs() before writing the inputs of a reused
graph, and backends that cache work against a graph must key it on
ggml_cgraph::uid. Drop the commentary from the code.

Two one-line pointers are kept rather than removed outright, both at
places where the invariant is not local and violating it fails silently:
the cgraph->uid check in the CUDA backend, and the note that sched->bufts[]
is not populated yet where the ring is detected.

No functional change. Sanitizer race reports are unchanged on all three
models - Qwen3.5 87 -> 0 at n_ubatch 512 and 21 -> 0 at 32, Muse-Glimmer
135 -> 0 - and the reuse path still rotates 9 times on the small ubatch
workload.

Assisted-By: Claude Opus 5
A race is reported against the tensor that was accessed, but when that
tensor is a view the memory belongs to its root, and the two names can be
completely unrelated. Chasing one such report cost a wrong diagnosis and
two rebuilds: the report named a recurrent state tensor, while the memory
being recycled belonged to an unnamed intermediate several links up the
view chain.

Name both, as "accessed <- root", so a report points at the allocation
that is actually in contention.

Assisted-By: Claude Opus 5
… reuse pool

The graph allocator frees a tensor's memory once its last consumer in
graph order has run, and hands it to a later tensor. That is only sound if
the consumers have actually finished. When a backend computes directly on
a buffer another backend owns - as an integrated GPU does on host memory -
the reading split is asynchronous and is still reading well after graph
order says the memory is dead, so a later split on the owning backend
writes over it.

Give the allocator an explicit pin and apply it to exactly those tensors.
The pin is keyed on the view root, since that is the tensor that owns the
memory and the one ggml_gallocr_free_node() acts on - keying it on the
accessed tensor misses every access that goes through a view, which is
most of them here.

A dedicated pin rather than the existing GGML_TENSOR_FLAG_OUTPUT: the flag
also suppresses fusion in the CUDA and SYCL backends, and the tensors that
need pinning here are precisely gated delta net state, so reusing it would
have disabled ggml_cuda_try_gdn_cache_fusion() on the models this fixes.
It also propagates along view chains, which pins more than intended.

Measured on a Strix Halo APU with a recurrent model, sanitizer race
reports over partially offloaded layers:

  -ngl    0   10   20   30   99
  before 160  120   60   20    0
  after    0    0    0    0    0

Cost is +4.3 MiB of compute buffer, flat in context size, and nothing at
all when no split reads across backends in place. GGML_SCHED_PIN_ASYNC_READS=0
disables it.

Assisted-By: Claude Opus 5
An op whose result aliases one of its sources - ggml_set() and friends,
where the destination is a view of src0 - writes through to that source.
Pass 4 already says "views are always on the same backend as the source",
but only applies it to nodes that are still unassigned; passes 2 and 3 can
have moved an aliasing node to another backend before that. Pass 5 then
finds a source on a different backend than the split, substitutes a copy
for it, and the write lands in the copy. The copy is never written back,
so the update is discarded.

Nothing detects this. The graph computes, every op runs, and the result is
quietly wrong. It is only reachable when the placement puts such a node
across a backend boundary, which on a recurrent model means the state
carried between tokens stops being updated and generation degenerates into
a repeated token.

Enforce the invariant for nodes that alias their source and are not pure
view ops.

Found on a Strix Halo APU with a hybrid recurrent model at partial
offload, where every one of the 13 split input copies fed a SET:

  -ngl                  0     8    16    24    30    99
  before          garbage  ...  ...   ...   ...    ok
  after                ok    ok    ok    ok    ok    ok

The trigger is the backend accepting another's buffers, which changes
placement: forcing devices[].integrated = false, as the CUDA path already
does, also avoids it. That makes it reachable today on HIP integrated GPUs
and on any backend that computes on buffers it does not own.

The fix moves nothing at full offload, and at partial offload it moves
only the SET nodes - 13 at -ngl 16 - removing the copies that fed them.

Assisted-By: Claude Opus 5
The scheduler's documentation was a prose block at the top of the sched
section in ggml-backend.h, and the mechanisms added since were commented at
the sites that implement them. Neither is somewhere a reader looks to find
out how the scheduler works, and the header block had grown past what
belongs in a header.

Add docs/development/backend-scheduler.md covering backend assignment,
splits, allocation and graph reuse, and then the parts that exist because
the reuse path runs neither the split nor the allocator: the graph input
ring buffer and its two contracts, pinning memory that another backend
reads in place, and the placement rule for ops that alias their source.
Document the sanitizer with it, since every one of those is an ordering
rule and the sanitizer is how they are checked.

Move the header block and its usage example there and drop the commentary
from the code, leaving a pointer at the three places where the invariant is
not local: the cgraph uid check in the CUDA backend, llama_host_write(),
and the sched section of the header.

No functional change - race reports are unchanged at 87/0 with the ring off
and on, 0 at partial offload, and partial offload output is still correct.

Assisted-By: Claude Opus 5
Two problems, both reached by test-opt, which builds a scheduler with the
backend under test in front of the full backend list - so for the CPU
device that is a CPU backend in front of the CPU backend.

The detection asked whether any backend other than the last accepts the
last backend's buffer type. A CPU backend trivially accepts CPU memory, so
a CPU-only scheduler was treated as a device computing on host memory and
given a ring it has no use for. Skip CPU devices: the ring exists for a
non-CPU device reading memory the host writes.

That alone was a segfault rather than wasted memory. The tensor aliased by
the ring sat at slot cur_copy, so which tensor occupied a given leaf index
changed from one allocation to the next. ggml_gallocr_needs_realloc()
compares node and leaf counts and sizes but never identity, so it saw no
reason to reallocate, and a leaf_alloc recorded for a pre-allocated tensor
(buffer_id -1) was applied to a placeholder that still needed allocating,
indexing galloc->buffers out of bounds.

Keep the aliased tensor at slot 0 so leaf identity is stable, and rotate
purely by re-pointing the inputs, which both paths now do explicitly. This
is what the graph reuse path already relied on; the allocate path was
getting rotation as a side effect of where split_graph placed the alias.

Also stop GGML_SCHED_UMA_RING from switching the ring on where it was not
detected. It is an escape hatch for a scheduler that has one, not a way to
impose one on a caller that does not meet its contract - ggml_opt keeps its
inputs across separate allocations and does not.

Sanitizer race reports are unchanged: 94/0 with the ring off and on, 0 at
partial offload, reuse path still rotating, output correct at full and
partial offload.

Assisted-By: Claude Opus 5
The detection excluded CPU devices, but BLAS is an accelerator device that
computes on host memory, so a BLAS plus CPU scheduler was still given a
ring. ggml_opt keeps its inputs across separate allocations and does not
meet the ring's contract, so test-opt went to 73/118 with it enabled.

Excluding device types one at a time is the wrong shape. The hazard is a
reader that is still reading after graph_compute returns, so require the
backend to be asynchronous. That covers CPU and BLAS together, and anything
else synchronous: their work is finished before control comes back, and
nothing of theirs can be reading the inputs the host is about to overwrite.

Verified against a BLAS build of test-opt, which is the configuration that
failed: 118/118, ring never enabled. The ring is still enabled for ROCm and
its race reports are unchanged at 94/0 with it off and on, 0 at partial
offload.

Assisted-By: Claude Opus 5
Pass 4 assigns a node that aliases its source to the backend owning that
source, so the write reaches the aliased memory instead of a discarded copy.
It did so unconditionally.

Support for these ops can be conditional on the tensor types - CUDA runs
GGML_OP_SET only for F32 and I32, and OpenCL has no GGML_OP_ACC at all - so
the backend that owns the aliased memory is not guaranteed to be able to run
the op writing into it. Forcing the node there hands the backend an op it
rejects, which aborts in ggml_backend_graph_compute.

There is no placement that computes such a graph correctly: operands are
copied into a split, results are never copied out, so the write cannot reach
the aliased memory from any other backend. Check support before moving, and
leave the node where the earlier passes put it otherwise - the behaviour
before this rule existed. The declined move is logged under GGML_SCHED_DEBUG.

Measured on Qwen3.5-4B at -ngl 16, the case the rule was added for: 13 moves
before and after, 0 declined, output unchanged. Races still 0 with the ring
on and 1 with it off, and test-opt passes 46/46.

Assisted-By: Claude Opus 5
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CUDA Related to the CUDA backend documentation Improvements or additions to documentation ggml changes relating to the ggml tensor library for machine learning testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants