fix(db): cap per-request batch sweep and guard shared files (#13680, #13681) - #13805
Merged
diegosouzapw merged 2 commits intoSep 16, 2026
Merged
Conversation
…13681) deleteCompletedBatches ran an unbounded synchronous for(;;) loop over INSTANCE_SWEEP_CHUNK-sized chunks, so one request could hold the event loop for as long as it took to sweep every completed batch on the instance, with no way for the caller to detect or bound the work. Separately, the sweep (and deleteBatch) nulled a batch's input/output/ error file unconditionally, even when another batch — in progress, or completed but outside the swept chunk — still referenced the same file. Adds MAX_CHUNKS_PER_REQUEST (25 * INSTANCE_SWEEP_CHUNK = 5000 batches) to sweepLoop with a hasMore continuation flag threaded through the DELETE /v1/batches/delete-completed response (resumption is natural via rowid ordering, no cursor needed); and isFileReferencedByOtherBatch(), applied before every file soft-delete in deleteCompletedBatches, deleteBatch, and cleanupExpiredBatches. Regression tests: tests/unit/issue-13680-batches-delete-completed-unbounded-work.test.ts, tests/unit/issue-13681-shared-file-across-batches.test.ts
diegosouzapw
force-pushed
the
fix/13680-batches-sweep-cap-shared-file
branch
from
September 15, 2026 22:56
f221057 to
0de1660
Compare
…ile (base-red fix #13747)
muhamadgalihsaputra
pushed a commit
to niyatna/NiyatnaRoute
that referenced
this pull request
Sep 27, 2026
…zapw#13680, diegosouzapw#13681) (diegosouzapw#13805) Merged in the 2026-09-16 sweep of the maintainer's own open PRs, at the owner's explicit instruction. No push was made to the PR branch: the merge took the head as the owning session left it (verified OPEN, non-draft and MERGEABLE against the release tip immediately before merging).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #13680
Closes #13681
Root cause (short)
deleteCompletedBatches'ssweepLoop(src/lib/db/batches.ts) ran an unbounded synchronousfor (;;)overINSTANCE_SWEEP_CHUNK-sized (200) chunks.better-sqlite3is synchronous, so oneDELETE /v1/batches/delete-completedrequest could hold the Node.js event loop for as long as it took to sweep every completed batch on the instance, with no way for the caller to detect or bound the work (nohasMore).deleteCompletedBatches(insidesweepIds),deleteBatch, andcleanupExpiredBatches(open-sse/services/batchProcessor.ts) all soft-deleted a batch'sinput_file_id/output_file_id/error_file_idunconditionally, with no check for whether some OTHER batch — in progress, queued, or completed but outside the current sweep chunk — still referenced the same file id. A tenant reusing one input file across two batches could lose that file out from under a still-live batch.Both residual after PR #13684 (merged same day), which fixed different findings (chunked write-lock duration + forward-progress guard) from the same omni-code-sec battery but did not cap chunks or guard shared files.
Fix
src/lib/db/batches.ts: addedMAX_CHUNKS_PER_REQUEST = 25(5000 batches/request).sweepLoopnow peeks the next chunk once the cap is hit — without deleting or counting anything — and returnshasMore: trueif more remain; the caller (the DELETE route) simply calls again. Resumption is natural (no cursor): swept rows are gone,rowidonly increases, so the next call'sORDER BY rowid LIMIT ?picks up exactly where the last one left off.isFileReferencedByOtherBatch(fileId, excludeBatchIds)— a smallSELECT 1 … WHERE (input_file_id = ? OR output_file_id = ? OR error_file_id = ?) AND id NOT IN (…) LIMIT 1guard (drops theNOT INclause entirely when the exclude list is empty, since an emptyNOT IN ()is invalid SQL). Applied before every file soft-delete indeleteCompletedBatches'ssweepIds,deleteBatch, andcleanupExpiredBatches's terminal-batch expiry loop.deleteCompletedBatches's return type and theDELETE /v1/batches/delete-completedroute response both gainedhasMore: boolean(additive field; existingdeletedBatches/deletedFilescallers are unaffected).deleteCompletedBatchesJSDoc to document the cap/hasMorecontract and to narrow the pre-existing "shared file can be nulled across chunks" caveat to the one race the guard genuinely cannot see (a new batch created between chunk commits, reusing the just-nulled file id) — the guard now covers every case of a concurrently existing sibling batch, regardless of its status or chunk.Regression tests
tests/unit/issue-13680-batches-delete-completed-unbounded-work.test.ts— RED on unfixed code:MAX_CHUNKS_PER_REQUESTdid not exist (assertion evaluated toNaN/undefined); GREEN after the fix (2/2 passing), including a resumption test that calls the sweep repeatedly untilhasMoreisfalseand verifies every seeded batch is eventually swept.RED excerpt:
GREEN excerpt:
tests/unit/issue-13681-shared-file-across-batches.test.ts— RED on unfixed code (2/2 failing): a file shared by a completed batch and a survivingin_progresssibling was nulled by bothdeleteCompletedBatchesanddeleteBatch; GREEN after the fix (3/3 passing), including a case proving the file IS still deleted once the last referencing batch is gone (no regression toward "never delete").RED excerpt:
GREEN excerpt:
Gates run
npx eslint --suppressions-location config/quality/eslint-suppressions.json <changed files>→ 0 findingsnpm run typecheck:core→ exit 0node scripts/check/check-file-size.mjs→ no ✗ on touched files (the one pre-existing ✗,open-sse/utils/stream.ts, is unrelated/untouched)node scripts/check/check-complexity.mjsandnode scripts/check/check-cognitive-complexity.mjs→ 0 findings on touched filesnode scripts/check/check-test-discovery.mjs→ OK, both new test files discoveredbatch-deletion.test.ts(8/8),batches-delete-completed-ownership-wvxc.test.ts(15/15),batch-delete-completed-ownership-wvxc.test.ts(5/5),batches-delete-completed-route-scope.test.ts(11/11),batch-deletion-route-logic.test.ts(11/11), plus the two new regression files (2/2, 3/3)Existing tests aligned
tests/unit/batches-delete-completed-route-scope.test.ts: extended the shared response-body type withhasMore?: booleanand added one assertion (body.hasMore === false) to the existing "an inference key sweeps only its own batches" case — additive, no existing assertion weakened or removed.No other existing assertion needed changing:
batch-deletion.test.ts's "shared file IDs across multiple completed batches" test seeds two completed batches (both inside the same sweep chunk), which the new guard does not affect — a file is only preserved when a batch outside the current chunk still references it.