Skip to content

[Bugfix] Reclaim unconsumed SHM segments and sender state on request abort - #4349

Open
gagandhakrey wants to merge 14 commits into
vllm-project:mainfrom
gagandhakrey:fix/shm-segment-leak-on-abort
Open

gagandhakrey wants to merge 14 commits into
vllm-project:mainfrom
gagandhakrey:fix/shm-segment-leak-on-abort

Conversation

@gagandhakrey

@gagandhakrey gagandhakrey commented Jun 11, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Fixes a /dev/shm and host-memory leak on request abort in the SHM chunk-transfer path.

What breaks

SharedMemoryConnector.cleanup() has no callers on current main. SHM segments are only unlinked inside shm_read_bytes() during successful consumption, so a chunk that is written but never read — client disconnect, downstream early exit, consumer failure — stays allocated until process exit, along with its /dev/shm/shm_{key}_lockfile.lock.

/dev/shm is physical RAM. Once exhausted, every subsequent put() fails and chunk transfer breaks pipeline-wide, not per-request.

Aborted requests additionally retain sender-side state (put_req_chunk, request_payload, code_prompt_token_ids, requests_num_chunks_sent, ramp_chunk_count, _pending_streaming_prefills), which can pin CPU tensors held by TTS input processors. cleanup_sender() is reachable only from the natural-finish path today.

This is issue #6352, measured at ~45 KB/abort and 258 → 60 MiB/h under sustained load.

Repro

# Any 2-stage chunk-transfer pipeline (e.g. TTS talker -> vocoder).
ls /dev/shm | wc -l          # baseline

# Issue N streaming requests, disconnect each before the final chunk.

ls /dev/shm | wc -l          # grows by ~1 segment + 1 lock file per aborted
                             # request, never returns to baseline until exit

Fix

1. Abort-time connector sweep.
finish_requests() enqueues a connector_cleanup task onto the save queue; _send_single_request() dispatches it to _run_connector_cleanup(), which calls connector.cleanup(external_req_id) and then cleanup_sender().

Enqueued rather than called inline: both put() and the sweep run on the save-loop thread, so FIFO ordering guarantees no chunk is written after its own sweep.

2. Gated on finished_status.
Only FINISHED_ABORTED / FINISHED_ERROR sweep. Any other terminal status may still have an unconsumed terminal chunk downstream, and unlinking it would leave the next stage waiting on a finish marker that no longer exists.

The is_finished() guard alone is no longer sufficient — #6360's resumable-FINISHED_STOPPED carve-out lets a finished request fall through it.

3. Structural cleanup-key matching.
cleanup() previously matched any key with request_id as a bare _-delimited prefix or suffix. Both collide: abc swept abc_def's live segments, and cleanup("0") swept chunk 0 of every request in flight.

_key_belongs_to() now matches only the real key shapes:

  • {id}
  • {id}_{stage}_{chunk}
  • {id}_{stage}_{chunk}_{from_rank}_{to_rank}
  • omni_{from}_to_{to}_kv_cache_{id}

4. Lock-file leak on consume.
_get_data_with_lock() now removes the lock file whenever the segment was consumed, not only when deserialization succeeded.

Once shm_read_bytes() has unlinked the segment, no retry can succeed, so a deserialization failure must not strand the lock file.

Scope

  • Fixes the SHM connector path. Mooncake / MORI transfer-engine connectors key their buffers by _make_key(key, from_stage, to_stage), so a bare-id cleanup() is a no-op there; their equivalent leak is out of scope.
  • Each stage sweeps only segments it wrote; full coverage relies on abort propagating to every stage.
  • Producer-side _pending_keys growth on the successful path is not addressed here — put() adds to the producer's set while get() prunes the consumer's set in a different process. Abort-path growth is bounded by this PR; the natural-finish case needs a separate change.

Note for reviewers

Fix 3 required re-pointing an existing test. TestCleanup::test_cleanup_removes_unconsumed_segment asserted the generic _-suffix match (key cleanup_req_42, swept by cleanup("req_42")) — exactly the behaviour being removed.

It now uses the real KV key shape, with three added regressions covering the collisions. This is a deliberate narrowing of cleanup()'s contract, not an incidental edit.

Test Plan

pytest -sv tests/distributed/omni_connectors/test_shm_connector.py \
           tests/distributed/omni_connectors/test_chunk_transfer_adapter.py \
           -m 'core_model and cpu'

Adapter (test_chunk_transfer_adapter.py) — 6 tests, 7 cases:

  • Abort enqueues exactly one cleanup task; processing it sweeps connector + sender state.
  • Cleanup runs after saves already queued before the abort (FIFO).
  • Already-finished and unknown request IDs are never swept.
  • FINISHED_STOPPED does not enqueue cleanup (finished_status gate).
  • Sender-only stage_id=0 abort is swept, parametrized over FINISHED_ABORTED / FINISHED_ERROR.
  • Natural-finish terminal send never invokes connector.cleanup().

Connector (test_shm_connector.py) — 5 tests:

  • Chunk-style keys and their lock files are swept, unrelated requests untouched.
  • abc does not sweep abc_def.
  • cleanup("0") does not sweep other requests' chunk 0.
  • Rank-aware KV keys with non-numeric stage names are still swept.
  • Lock file is removed even for a falsy consumed payload.

Additional Updates

Also updated the title with the [Bugfix] prefix and refreshed the description. Four claims had gone stale against current main after the rebase, notably "producer _pending_keys remains bounded", which only holds on the abort path and is now explicitly scoped out.

Signed-off-by: Gagan Dhakrey <gagandhakrey@gmail.com>
@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Credits must be used to enable repository wide code reviews.

Comment thread vllm_omni/distributed/omni_connectors/transfer_adapter/chunk_transfer_adapter.py Outdated
@gagandhakrey

Copy link
Copy Markdown
Contributor Author

@princepride @yuanheng-zhao @yenuo26
could you please review this PR

@yuanheng-zhao yuanheng-zhao left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for contributing! Left some comments

Comment thread vllm_omni/distributed/omni_connectors/transfer_adapter/chunk_transfer_adapter.py Outdated
Comment on lines +819 to +822
"connector_cleanup": True,
"external_req_id": external_req_id,
# save_loop's error log reads task["request_id"].
"request_id": external_req_id,

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not sure if these three keys are necessary. Shall we shrink the code, for example, appending {"connector_cleanup": external_req_id} and update related code about sending request?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

could merge connector_cleanup/external_req_id into one key , keeping request_id though save_loop's error log relies on it for a useful id on failure, minor style choice, can shrink if you want

@yuanheng-zhao

Copy link
Copy Markdown
Collaborator

PTAL @Shirley125 , chunk_transfer_adapter.py related changes

…on-abort

# Conflicts:
#	vllm_omni/distributed/omni_connectors/transfer_adapter/chunk_transfer_adapter.py

Signed-off-by: Gagan Dhakrey <gagandhakrey@gmail.com>
…o fix/shm-segment-leak-on-abort

Signed-off-by: Gagan Dhakrey <gagandhakrey@gmail.com>
request is already known non-None and Request.external_req_id is
always a real attribute, so getattr(..., None) is dead defensiveness.

Signed-off-by: Gagan Dhakrey <gagandhakrey@gmail.com>
@gagandhakrey
gagandhakrey force-pushed the fix/shm-segment-leak-on-abort branch from 0dfcea1 to b9902ed Compare June 15, 2026 08:58
@ckc5800

ckc5800 commented Aug 19, 2026

Copy link
Copy Markdown

Production data point supporting this PR: we independently root-caused the sender-state leak on a 2-stage TTS pipeline (details in #6352). Constant ~45 KB per aborted request across 4 configs; sustained-load A/B shows 258 → 60 MiB/h with a minimal abort-time cleanup_sender() call. The residual ~13 KB/abort we still measure is exactly the re-creation race your save-queue FIFO design prevents — so this PR should close it completely. Would love to see this merged.

@amy-why-3459

Copy link
Copy Markdown
Collaborator

Thanks for this — the leak is real, and the save-queue FIFO design is the
right fix. SharedMemoryConnector.cleanup() still has no callers on current
main; #6352 independently measured ~45 KB/abort of sender state and a
258 → 60 MiB/h drop from cleanup_sender(), with a residual ~13 KB/abort
that is exactly the queued-put race this PR serializes away.

Request changes:

  1. Rebase onto current main (mergeable: dirty). Today's finish_requests
    already clears _active_streams / _held_non_active; landing the June
    diff as-is would regress that. Tests still use pooling_output and
    will KeyError against the current _send_single_request contract
    (multimodal_output). June CI is not merge evidence.

  2. Gate the connector sweep on finished_status
    (FINISHED_ABORTED / FINISHED_ERROR), not “any still-live request in
    finish_requests”. finished_status is unused today. If a later change
    (e.g. [Core][Benchmark]Omniinteract for Minicpm-o4.5 #5102 resumable FINISHED_STOPPED) feeds a live request through
    this path, this sweep would unlink an unread terminal chunk and stall
    the next stage. Please add a test that FINISHED_STOPPED does not
    enqueue cleanup.

  3. Add a stage_id=0 sender-abort test. [Bug]: Host memory leak — sender-side per-request state never freed for aborted requests (chunk_transfer_adapter) #6352's production leak is on
    the talker: process_pending_chunks returns early at stage 0, and
    abort never sends a terminal chunk. All four new adapter tests use
    stage_id=1.

Should-fix:

  • Match cleanup keys as {external_req_id}{stage}{chunk}; startswith
    request_id+"_" collides (abc vs abc_def).
  • Update cleanup_sender's docstring — it currently forbids the abort
    call this PR adds. Natural finish still waits for the terminal put;
    abort/error may reclaim immediately, but only after queued puts.
  • On rebase of _get_data_with_lock, keep this PR's consumed-after-
    shm_read_bytes semantics (unlink the lock even if deserialize
    fails). Do not fall back to main's deserialized flag or the old
    if obj check.
  • Coordinate with [Bug]: Host memory leak — sender-side per-request state never freed for aborted requests (chunk_transfer_adapter) #6352: this PR is the complete fix. Please don't
    land a scheduler-only cleanup_sender() alongside it.
  • Fine to shrink the task to {"connector_cleanup": external_req_id}
    and keep request_id for save_loop's error log (yuanheng's comment).

The FIFO enqueue, the “don't unlink on natural finish” invariant, and
the new tests for that invariant all look correct. Happy to re-review
after the rebase.

…on-abort

Conflicts resolved:

* shm_connector._get_data_with_lock: main independently fixed the same
  lock-file leak by gating removal on a `deserialized` flag. Kept this
  branch's `consumed` gate (the segment is already unlinked by
  shm_read_bytes, so the lock file must go even when deserialization
  fails and no retry can succeed), with main's type annotations and
  FileNotFoundError handling.

* chunk_transfer_adapter.finish_requests: took main's restructured skip
  logic (`request is None` plus the resumable-segment-stop carve-out from
  vllm-project#6360) and its cleanup_receiver()/_held_non_active handling, while
  keeping the aborted_external_ids collection and the connector-cleanup
  enqueue. A resumable segment stop is now fully reclaimed here, so it is
  swept like any other abort.

Signed-off-by: Gagan Dhakrey <gagandhakrey@gmail.com>
The branch tip already carried a merge of an older main (3c62c53), which
is a subset of the main merged in the previous commit. No content change.

Signed-off-by: Gagan Dhakrey <gagandhakrey@gmail.com>
@gagandhakrey gagandhakrey changed the title Critical Memory leak bug : Unlink unconsumed SHM segments and lock files for aborted requests [Bugfix] Reclaim unconsumed SHM segments and sender state on request abort Aug 20, 2026
Signed-off-by: Gagan Dhakrey <gagandhakrey@gmail.com>
@gagandhakrey

Copy link
Copy Markdown
Contributor Author

Thanks @amy-why-3459 for detailed review
Thanks @ckc5800 for validating this in production ,really appreciate it

All three blockers are done, plus the should-fixes.

Blockers (Resolved)

1. Rebased — now MERGEABLE.

2. Gated on finished_status — (FINISHED_ABORTED, FINISHED_ERROR).

3. stage_id=0 test added — test_finish_requests_sweeps_sender_only_stage_zero, parametrized over both abort statuses.

Should-fixes (done)

Key matching — now structural (_key_belongs_to).

  • The suffix half was the worse one: cleanup("0") would unlink chunk 0 of every request in flight.
  • Required changing an existing test. test_cleanup_removes_unconsumed_segment asserted a generic suffix match. Re-pointed it at the real KV key shape and added three collision regressions. Flagging it as a deliberate contract narrowing, not an incidental edit.

Task shape — one correction to your note.

  • request_id is already dead: save_loop reads getattr(task.get("request"), "request_id", None).
  • Task is now just {"connector_cleanup": external_req_id}.

cleanup_sender docstring — rewritten around the real invariant (save_loop thread only), not "after the terminal chunk".

_get_data_with_lock — kept consumed semantics.

On #6352

  • The residual ~13 KB/abort should be closed, not just serialized away.
  • A late put() is the only re-leak path, and it's unreachable: save_async is reached only from update_from_output, which skips aborted requests before that line.
  • @ckc5800 — a re-run against this branch should show zero.

Ready for re-review.

(this comment is AI polished)

@vllm-omni-review-bot

vllm-omni-review-bot commented Aug 31, 2026 •

Copy link
Copy Markdown

Omni ReviewBot triage note

Automated triage of commit f2c09b53affe produced:

  • Priority: high. Prompt maintainer attention is suggested.

These are automated triage suggestions only — the final decision belongs to the maintainers.

@vllm-omni-review-bot

Copy link
Copy Markdown

Omni ReviewBot: no human activity for 8 days

@gagandhakrey this pull request has had no human commit, comment or review since 2026-08-31. Per repository policy it may be closed if it stays inactive.

To keep it moving, any one of these is enough: push an update, reply to the open blocker, or post the current plan and timeline.

@ckc5800

ckc5800 commented Sep 18, 2026 •

Copy link
Copy Markdown

We carried this PR's approach into a v0.24 deployment for pre-production load testing
and ended up shipping a narrower variant. The reasons are all leak-mechanics, so here they
are in case they help land this.

Production-scale throughput and volume figures are omitted as a matter of policy. Everything load-bearing below is normalized — per abort, per request, per token — which is also the form that makes the invariance argument work. The load range quoted for the four configurations is the one already published in #6352.

Stage-0 sender-abort test (review point 3)

That path is the one our load testing reproduces. Conditions:

  • 2-stage pipeline, talker at stage 0 → code2wav at stage 1, streaming speech
  • client disconnect or server-side stream cancellation mid-utterance, so the engine request
    ends with finished_reason="abort" and no terminal chunk is ever sent
  • process_pending_chunks returns early for stage_id == 0, so the abort purge never
    sweeps the sender; cleanup_sender() has exactly one call site, on the terminal-send
    path, which an aborted request never reaches

Retained state per abort is one utterance's codec frame history plus the ref tensor. We
measured a constant ~45 KB/abort across four configurations whose raw leak rates differed
by 13x
(2–12 rps, response cache on/off, fixed/unfixed seed, glibc/jemalloc). That
invariance is what pins it to per-abort state rather than to load.

Two constraints we ran into

Sweep cost scales with the pending-key set. Sender-side _pending_keys grows on every
put() while the matching discard only runs in the receiving process, so on stage 0 the
set never shrinks — and an abort-time sweep then walks the whole thing. We had to abandon
the scanning form for that reason.

Reconstructing the keys avoids it entirely: keys are f"{external_req_id}_{stage}_{chunk_id}"
with chunk_id monotonic from 0, so the exact set for a request can be rebuilt from its
chunk count and discarded directly — no iteration, no lock, GIL-atomic. That also sidesteps
the should-fix note on key matching (prefix matching on request_id + "_" collides: abc vs abc_def).

To be fair to the scanning form, reconstruction is not race-free either: a put() already
in flight at abort time can re-add a key the reconstruction has just passed. We cover it by
reconstructing a small margin beyond the recorded chunk count, and the worst case we could
construct leaves one key (~130 B) per abort — about 1% of what it reclaims. This PR's
save-queue ordering is the cleaner answer to that race; the trade is that it also unlinks.

Reclaiming at sender-finish time is not safe for segments. sent != consumed — the
sender cannot tell whether the receiver has picked a segment up, so unlinking on the sender's
completion can remove a segment that is still to be read. Our shipped version therefore
touches no filesystem state at all: it only discards heap bookkeeping keys, which put/get
never read except to add/discard. Strictly less capable than this PR, but it cannot disturb
an in-flight transfer. If the unlink is retained here, it needs to be ordered by the
save-queue rather than inferred from sender completion.

What we shipped

A minimal abort-only cleanup_sender() from omni_ar_scheduler._free_request, gated on
RequestStatus.FINISHED_ABORTED. Sustained-load A/B: 258 → 60 MiB/h. With the
_pending_keys fix described above, plus a second connector-bookkeeping fix
(_cancelled_load_reqs, written up in #6352), a multi-day load run shows no leak at all.

The gating matches review point 2: normally-finished requests must keep their state until
the pending terminal chunk is sent, so the sweep has to key on FINISHED_ABORTED /
FINISHED_ERROR, and FINISHED_STOPPED must not enqueue cleanup.

I can contribute the stage_id=0 sender-abort test case as a PR against this branch if that
unblocks it.

@amy-why-3459

Copy link
Copy Markdown
Collaborator

@gagandhakrey this PR currently has merge conflicts with main (mergeable_state: dirty), so it cannot be merged as-is.

Please rebase onto the latest main, resolve the conflicts, and push:

git fetch origin main
git rebase origin/main
# resolve conflicts, then:
git add -A
git rebase --continue
git push --force-with-lease

Once it is mergeable again we can continue review. Thanks!

@natureofnature

Copy link
Copy Markdown
Collaborator

The unread-SHM cleanup is still useful, but I found a correctness issue that should be addressed before merging.

_key_belongs_to() still allows cross-request collisions. For example:

  • Request abc_1_2 writes chunk key abc_1_2_0_0.
  • cleanup("abc") accepts the four-field suffix as a rank-aware key.
  • It consequently deletes the other request’s segment and lock file.

An isolated reproduction using this PR’s unmodified cleanup methods and real POSIX shared memory confirms the deletion. Checking that all suffix fields are numeric would not resolve this ambiguity. Could we track exact request-to-key ownership instead, and add this case as a regression test?

Also, current main already calls cleanup_sender() during request cleanup, with in-flight send cancellation and deferred cleanup for queued terminal chunks. The PR still has merge conflicts, so the rebase should preserve those protections and focus on the remaining unread-SHM and lock-file cleanup.

Finally, the latest update in #6352 identifies the remaining memory growth as separate issues. The description should distinguish those from the abort-path leak rather than imply this PR resolves all host-memory growth.

@Gaohan123 Gaohan123 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. Thanks

@Gaohan123

Copy link
Copy Markdown
Collaborator

Please resolve conflicts

gagandhakrey and others added 2 commits September 24, 2026 09:25
…on-abort

Resolve conflicts with the sender-generation fencing from vllm-project#6670:

- shm_connector: keep main's non-blocking flock and BlockingIOError
  retry path alongside the `consumed` lock-file cleanup.
- chunk_transfer_adapter: dispatch the connector_cleanup task before
  main's generation-fenced send body, and take main's cleanup_sender
  docstring (it is now safe to call from the scheduler thread).
- _run_connector_cleanup no longer calls cleanup_sender: finish_requests
  already retires the generation inline, and a second call from the save
  thread would cancel a successor generation that reused the external id.
  Add a regression test for that case.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@gagandhakrey

gagandhakrey commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor Author

resolved @Gaohan123 _run_connector_cleanup no longer calls cleanup_sender(), since finish_requests now does it inline and a second call could cancel a request reusing the same external id. Test added

@vllm-omni-review-bot

Copy link
Copy Markdown

Omni ReviewBot: no human activity for 7 days

@gagandhakrey this pull request has had no human commit, comment or review since 2026-09-24. Please confirm the current plan and next step. The author or a maintainer decides whether to change the PR state.

To keep it moving, any one of these is enough: push an update, reply to the open blocker, or post the current plan and timeline.

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot routing record

Assigned Strict on cursor (cursor-grok-4.6-high) under experiment fleet-strict-cursor-grok46-zcode-glm53flash-5050-c5-z10-20261002.

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot attempt record

Review attempt ended as timeout (Strict attempt outlived its budget; falling back to direct/cursor/auto).

1 similar comment
@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot attempt record

Review attempt ended as timeout (Strict attempt outlived its budget; falling back to direct/cursor/auto).

@vllm-omni-review-bot vllm-omni-review-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Omni ReviewBot review

Changes since the previous review

  • 0 new inline finding(s); 1 finding(s) below.

CI at f2c09b53affe (2026-10-09T13:06:43.964448+00:00): required check(s) blocking: buildkite/vllm-omni (missing).

Note: The assigned review arm strict/cursor/cursor-grok-4.6-high could not complete this review, so it was produced by the fallback arm direct/cursor/auto. It is excluded from the routing experiment.

Full review analysis

PR description

On abort or error, finish_requests still reclaims sender state inline, then enqueues a connector_cleanup task so the save thread unlinks that request's unconsumed shared-memory segments only after chunks already queued for it. SharedMemoryConnector.cleanup now matches structural key shapes instead of a bare prefix or suffix, and a consumed read deletes the lock file even when deserialization fails. Natural finish and FINISHED_STOPPED do not unlink segments. The PR text still says the sweep also calls cleanup_sender(); at this head that second call is gone so a reused external id is not retired twice.

Change flow

flowchart LR
  Abort["[EXISTING] finish_requests abort or error"]:::existing
  Gate["[CHANGED] FINISHED_ABORTED or ERROR gate"]:::changed
  Task["[NEW] connector_cleanup save-queue task"]:::new
  Match["[CHANGED] cleanup key match and unlink"]:::changed
  Shm["[EXISTING] /dev/shm segment and lock file"]:::existing
  Abort --> Gate --> Task --> Match --> Shm
classDef existing fill:#e5e7eb,stroke:#6b7280,color:#111827
classDef changed fill:#fef3c7,stroke:#d97706,color:#451a03,stroke-width:2px
classDef new fill:#dcfce7,stroke:#16a34a,color:#052e16,stroke-width:2px
classDef removed fill:#fee2e2,stroke:#dc2626,color:#450a0a,stroke-width:2px
Loading

Findings

  • [P1] Abort cleanup still unlinks another request's chunk key — vllm_omni/distributed/omni_connectors/connectors/shm_connector.py:175
    Existing thread: #4349 (comment)
Evidence for Abort cleanup still unlinks another request's chunk key

_key_belongs_to treats a four-field remainder as a rank-aware key (len(fields) in (2, 4) and only the last field decimal). Chunk puts are {external_req_id}_{stage_id}_{chunk_id} (chunk_transfer_adapter.py around the connector_put_key format). For live ids abc and abc_1_2, key abc_1_2_0_0 starts with abc_, splits into 1,2,0,0, and cleanup("abc") unlinks that segment and its lock file. external_req_id is the client-facing id, so underscore ids are in contract. The new tests only cover abc vs abc_def_0_0 (three remainder fields) and cleanup("0"); they never put abc_1_2_0_0. That unlink drops a chunk the other request's consumer is still waiting on, which is worse than the leak. Record the exact keys at put and delete those. Requiring every suffix field to be numeric does not separate this key from the rank-aware shape.


🤖 This review was generated by InferMatrix Copilot, an open-source repo-maintenance agent for PR review, CI debugging and issue triage. Try it on your own repo, and ⭐ star it if it helped!

@vllm-omni-review-bot

Copy link
Copy Markdown

Omni ReviewBot: finding feedback

[p1] Abort cleanup still unlinks another request's chunk key — vllm_omni/distributed/omni_connectors/connectors/shm_connector.py:175

See the review for details. If you are the PR author and disagree, react 👎 here; the maintainer will see your disagreement.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working core related to core module: cache, scheduler, engine, worker, modelrunner

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: Host memory leak — sender-side per-request state never freed for aborted requests (chunk_transfer_adapter)

8 participants