Studio: make Stop interrupt a llama.cpp generation stalled mid-stream - #7117
Conversation
There was a problem hiding this comment.
Code Review
This pull request updates the llama_cpp backend to ensure that active HTTPX sockets are shut down first during cancellation, allowing reads blocked in recv() during a mid-stream stall to be interrupted immediately rather than hanging. It also adds a new integration test using a raw HTTP/1.1 server to verify this behavior. The feedback suggests improving the test's determinism by using a threading.Event instead of a hardcoded sleep to trigger the cancellation, preventing potential flakiness in slow CI environments.
Important
The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.
|
@codex review |
|
Codex Review: Didn't find any major issues. Delightful! Reviewed commit: ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |
|
@codex review |
|
Codex Review: Didn't find any major issues. 🚀 Reviewed commit: ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |
When a GGUF generation goes quiet mid-response, clicking Stop did nothing: the chat sat on "running" until the ~20 min first-token read deadline elapsed. Reported on Discord.
The stream reader blocks in
recv()waiting for the next token. Stop calledresponse.close(), which doesn't interrupt a read already blocked on the socket, so the read stayed parked until the read deadline fired. The socket shutdown that does interrupt it was only wired for the pre-header case, before any response arrives.Fix
Shut the socket down on cancel, not just close the response.
_shutdown_active_httpx_sockets(already used before a response arrives) unblocks the parkedrecv(); the read raises, and the existing handler turns it into a clean cancel.One ordering detail: shutdown has to run before
response.close(). On amax_keepalive_connections=0client, closing the response first drops the connection from the pool, so a later shutdown finds nothing to close.This doesn't add an auto-timeout, which is where it parts from #7002. A stall nobody stops still waits out the read deadline, on purpose: slow hardware (a big model mmap'd off SSD, minutes between tokens) looks like silence at the socket too, so a 2 min auto-kill would cut it off mid-run. Stop is the escape hatch.
Verification
Added a regression test: a real httpx/httpcore stream over a raw socket sends one chunk, goes silent, then cancels mid-read. The cancel lands in under a second instead of at the far-off deadline. It hangs the full deadline on the old code and passes now.
Ran it live too: loaded a GGUF, streamed, aborted mid-stream, and the worker tore down cleanly and served the next generation with no traceback in the log. Adjacent suites stay green:
test_llama_cpp_stream_cancel.py(3) plus the tool-loop, cursor-reset, orchestrator-unload-cancel, and route-timeout suites (124 together).