Skip to content

fix(deepseek): release the PoW worker slot when spawning fails - #13097

Merged
diegosouzapw merged 1 commit into
diegosouzapw:release/v3.8.51from
huuhungn:fix/deepseek-pow-slot-leak-13094
Sep 9, 2026
Merged

diegosouzapw merged 1 commit into
diegosouzapw:release/v3.8.51from
huuhungn:fix/deepseek-pow-slot-leak-13094

Conversation

@anhtahaylove

Copy link
Copy Markdown
Contributor

Fixes #13094.

Problem

solveInWorker takes the concurrency slot before the worker exists:

activeWorkerCount += 1;                       // slot taken
return new Promise<number>((resolve, reject) => {
  const worker = new Worker(resolveWorkerPath(), { workerData: validated });
  ...
  const cleanup = () => { activeWorkerCount -= 1; ... };   // only reachable
});                                                        // once construction succeeded

resolveWorkerPath() resolves the worker script against process.cwd() and throws when it isn't there. That throw escapes before any handler is wired up, so the counter is never decremented.

MAX_CONCURRENT_WORKERS is 2, so two such failures disable the solver for the lifetime of the process. Every later call then rejects with:

DeepSeek PoW worker capacity reached (2)

while no worker is actually running — and the real cause (missing worker script) is hidden behind a misleading capacity error.

Fix

Construct the Worker in a try/catch and hand the slot back before rejecting, so the rejection carries the real error.

Tests

tests/unit/deepseek-pow-slot-leak-13094.test.ts, two cases:

  1. Two consecutive spawn failures must not exhaust the budget — the third call still has to fail with worker script not found, not capacity reached.
  2. After a failure, a normal solve from a working directory must still succeed, proving the slot was genuinely returned rather than the counter just being clamped.

The failure is triggered the same way users hit it — process.chdir() into an empty temp dir so resolveWorkerPath() cannot find the script — rather than by mocking, and the cwd is restored in finally.

Verified as a real detector: both fail on release/v3.8.51 before the change, with the first showing actual: 'DeepSeek PoW worker capacity reached (2)'. Both pass after.

Scope

fail 0 across the full DeepSeek PoW suite (11 tests, existing ones included), and tsc -p tsconfig.json --noEmit is clean.

Same class of bug as the worker-slot leak in #12812, but on the logic counter rather than an OS thread, and reachable without any worker ever starting.

`solveInWorker` increments `activeWorkerCount` before constructing the Worker,
but the `cleanup()` that decrements it lives inside the promise executor and only
runs once the worker exists. Anything that throws first -- most obviously
`resolveWorkerPath()` when the worker script is missing, since it resolves
against `process.cwd()` -- leaves the counter permanently incremented.

With `MAX_CONCURRENT_WORKERS = 2`, two such failures wedge the solver for the
lifetime of the process: every later call rejects with "capacity reached (2)"
while no worker is actually running, and the real cause is hidden.

Construct the Worker inside a try/catch and release the slot before rejecting.

Fixes diegosouzapw#13094
@anhtahaylove

Copy link
Copy Markdown
Contributor Author

The checks shown here (Mergify, semgrep) are the ones that run without approval — the workflow runs for tests and typecheck are held in action_required because I am a first-time contributor, so the real signal has not run yet on this PR.

There are now 7 of these open, all small and independent, all from the same resource-leak audit:

PR Fixes Area
#13091 #12812 compression worker pool
#13092 #12819 plugin loader
#13093 #12822 LLMLingua worker spawn
#13096 #13095 ACP sendPrompt teardown
#13097 #13094 DeepSeek PoW worker slot
#13100 #13095 ACP output buffers
#13106 #13103 badge SSE stream

Approving the held workflow runs (once per branch) is all that is needed to get real CI on them. No rush on review itself — I would just rather you judge them on this repo's CI than on my local runs.

For what it is worth locally on Windows 11 / Node 24.19.0: tsc -p tsconfig.json --noEmit is clean, and each PR's touched suites pass. Every test in these PRs was confirmed to fail on the unpatched code first, so they are regression detectors rather than tests written to match the fix.

Happy to rebase, split, or drop any of them if the batch is too much at once.

@diegosouzapw
diegosouzapw merged commit 9492357 into diegosouzapw:release/v3.8.51 Sep 9, 2026
3 checks passed
diegosouzapw added a commit that referenced this pull request Sep 10, 2026
…agment (#13200)

Um caractere. O fragmento do #13097 subiu sem o `- ` inicial e derrubou o `Merge integrity` para todo mundo que veio depois.

Terceira ocorrência da mesma causa nesta release; a anterior foi o `reset-aware-model-family.md`, que o #12711 consertou de carona.
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…souzapw#13097)

`solveInWorker` increments `activeWorkerCount` before constructing the Worker,
but the `cleanup()` that decrements it lives inside the promise executor and only
runs once the worker exists. Anything that throws first -- most obviously
`resolveWorkerPath()` when the worker script is missing, since it resolves
against `process.cwd()` -- leaves the counter permanently incremented.

With `MAX_CONCURRENT_WORKERS = 2`, two such failures wedge the solver for the
lifetime of the process: every later call rejects with "capacity reached (2)"
while no worker is actually running, and the real cause is hidden.

Construct the Worker inside a try/catch and release the slot before rejecting.

Fixes diegosouzapw#13094
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…agment (diegosouzapw#13200)

Um caractere. O fragmento do diegosouzapw#13097 subiu sem o `- ` inicial e derrubou o `Merge integrity` para todo mundo que veio depois.

Terceira ocorrência da mesma causa nesta release; a anterior foi o `reset-aware-model-family.md`, que o diegosouzapw#12711 consertou de carona.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(deepseek): PoW worker slot counter leaks on any throw before cleanup, permanently disabling the solver after 2 failures

2 participants