fix(oracle): unblock Oracle CI — free runner disk space + fix the retain deadlock - #2948
Merged
Conversation
The three Oracle jobs run the Oracle 23ai `free` service image, which together with the Python ML deps (torch) exhausts the runner's ~14 GB root disk. Two symptoms, one cause: - uv fails to extract a wheel with "No space left on device (os error 28)" (fast ~2 min failure), and - a near-full disk starves I/O badly enough to trip the 30-minute job timeout. test-python-client-oracle and test-typescript-client-oracle have been red on every open PR (#2941, #2942, #2943) from this, independent of the code under test. Reclaim ~20 GB of preinstalled tooling (the same jlumbroso action the Docker build job already uses) before the Oracle setup step. docker-images stays false here: unlike the Docker build job, the Oracle service container is already running by the time steps execute, so pruning images could disrupt it. The savings come from the tool cache, Android SDK, .NET, Haskell, large apt packages and swap.
The first pass enabled every reclaim, which cost ~4 minutes of job time — counterproductive on jobs that are already fighting a 30-minute limit. android + dotnet + haskell + swap are a few rm -rf's worth ~16-21 GB, which is ample headroom for the Oracle image plus torch. Dropped: - large-packages: apt-get remove, costs minutes for little extra space; - tool-cache: deletes the preinstalled Python that actions/setup-python then re-downloads, making the job slower rather than faster.
…e hang)
Retain hung forever on the Oracle backend: every retain test burned its 120s
client timeout while the server sat idle, so test-python-client-oracle and
test-typescript-client-oracle only ever reached ~5% of the suite before the
30-minute job limit.
The server was not slow — it was deadlocked. flush_pending_stats() acquires
its own connection, but it was being called while the enclosing
acquire_with_retry(...) block still held one:
async with acquire_with_retry(pool) as conn: # conn checked out
async with conn.transaction(): # SAVEPOINT only
...write facts/entities...
await entity_resolver.flush_pending_stats() # takes a 2nd connection
oracledb does not autocommit and OracleConnection.transaction() is only a
SAVEPOINT, so the write is committed by OracleBackend.acquire() when its block
exits. Connection #2's `UPDATE entities ...` therefore waits on row locks held
by the still-open connection #1, which cannot commit until the call returns —
a circular wait. Oracle never reports ORA-00060 because session #1 is blocked
in Python, not on the database, so it hangs indefinitely instead of erroring.
Move the flush after the acquire block in all three call sites (streaming
retain, delta retain, transfer importer), which is what its own docstring
already required ("must be called AFTER the retain transaction commits") and
which PostgreSQL satisfied only by accident via asyncpg autocommit.
Guarded with an AST lint test rather than a behavioural one: the deadlock
cannot be reproduced against PostgreSQL, which is what the suite runs on.
test_dry_run_creates_nothing still flaked in test-api shard 3. CONCURRENTLY avoids ACCESS EXCLUSIVE but still takes ShareUpdateExclusive, which conflicts with the ShareLock a fresh bank's plain CREATE INDEX holds — and that one cannot be made concurrent, since it runs inside the bank-create transaction. So _drop_bank_indexes can still be picked as the deadlock victim while another xdist worker seeds a bank: Process A waits for ShareUpdateExclusiveLock on memory_units; blocked by B. Process B waits for ShareLock on virtual transaction; blocked by A. The bank-create side already retries (#2943); give the drop the same treatment. The drop is idempotent, so retrying is safe.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
test-python-client-oracleandtest-typescript-client-oraclehave been red on every open PR (#2941, #2942, #2943), independent of the code under test. There were two independent causes; this PR fixes both.1. Runner disk exhaustion (CI)
The Oracle 23ai
freeservice image plus the Python ML deps (torch) exhaust the hosted runner's ~14 GB root disk, souvcan't extract a wheel.Fix: add the same
jlumbroso/free-disk-spacestep the Docker build job already uses to the three Oracle jobs. Only the cheap, high-yield reclaims are enabled (Android/.NET/Haskell/swap ≈ 16-21 GB, a fewrm -rfs):large-packagesrunsapt-get removeand cost ~4 min for little gain;tool-cachewould delete the preinstalled Python thatactions/setup-pythonthen re-downloads;docker-imagesis pointless — the Oracle service container is already running, so its image is in use and can't be pruned.Verified:
No space left on deviceis gone (0 occurrences).2. Retain deadlocks on Oracle (the real blocker)
With disk fixed, the jobs still timed out. The server was not slow — it was deadlocked. Server logs showed the retain work finishing in under a second and then the process sitting completely idle for 7 minutes while the client waited out its 120s timeout:
[streaming] Consumer batch … total— the very next log line — never appeared. The only statement in between isawait entity_resolver.flush_pending_stats().Cause.
flush_pending_stats()acquires its own connection, but was called while the enclosingacquire_with_retry(...)block still held one:oracledbdoes not autocommit andOracleConnection.transaction()is only aSAVEPOINT, so the write is committed byOracleBackend.acquire()when its block exits. Connection #2'sUPDATE entities …waits on row locks held by the still-open connection #1, which cannot commit until the call returns — a circular wait. Oracle never raisesORA-00060, because session #1 is blocked in Python rather than on the database, so it hangs indefinitely instead of erroring. (3×Phase 1but only 1×Phase 2in the logs: once retain #1 wedged, later retains blocked in their own write txn.)Fix. Move the flush after the acquire block in all three call sites — streaming retain, delta retain, and the transfer importer. This is what its own docstring already required ("Must be called AFTER the retain transaction commits"); PostgreSQL satisfied it only by accident, via asyncpg autocommit.
Test. Guarded with an AST lint test (in the spirit of
test_migration_shape.py) rather than a behavioural one — the deadlock cannot be reproduced against PostgreSQL, which is what the suite runs on. It includes a self-check so it can't pass vacuously.Validation
This PR touches
.github/workflows/**, so both Oracle client jobs run here and the fixes validate themselves.lint.sh+tyclean; retain suites pass locally (25/25).Not addressed: the pre-existing
ORA-00903: invalid table namewarning when writingllm_traceson Oracle. It's caught and logged best-effort on every retain and does not block — worth a separate fix.