Skip to content

fix(oracle): unblock Oracle CI — free runner disk space + fix the retain deadlock - #2948

Merged
nicoloboschi merged 4 commits into
mainfrom
fix/oracle-ci-disk-space
Jul 24, 2026
Merged

fix(oracle): unblock Oracle CI — free runner disk space + fix the retain deadlock#2948
nicoloboschi merged 4 commits into
mainfrom
fix/oracle-ci-disk-space

Conversation

@nicoloboschi

@nicoloboschi nicoloboschi commented Jul 24, 2026

Copy link
Copy Markdown
Collaborator

test-python-client-oracle and test-typescript-client-oracle have been red on every open PR (#2941, #2942, #2943), independent of the code under test. There were two independent causes; this PR fixes both.

1. Runner disk exhaustion (CI)

##[warning]You are running out of disk space. Free space left: 87 MB
  ├─▶ I/O operation failed during extraction
  ╰─▶ No space left on device (os error 28)

The Oracle 23ai free service image plus the Python ML deps (torch) exhaust the hosted runner's ~14 GB root disk, so uv can't extract a wheel.

Fix: add the same jlumbroso/free-disk-space step the Docker build job already uses to the three Oracle jobs. Only the cheap, high-yield reclaims are enabled (Android/.NET/Haskell/swap ≈ 16-21 GB, a few rm -rfs):

  • large-packages runs apt-get remove and cost ~4 min for little gain;
  • tool-cache would delete the preinstalled Python that actions/setup-python then re-downloads;
  • docker-images is pointless — the Oracle service container is already running, so its image is in use and can't be pruned.

Verified: No space left on device is gone (0 occurrences).

2. Retain deadlocks on Oracle (the real blocker)

With disk fixed, the jobs still timed out. The server was not slow — it was deadlocked. Server logs showed the retain work finishing in under a second and then the process sitting completely idle for 7 minutes while the client waited out its 120s timeout:

12:38:57.992  Phase 2 (write txn): 0.335s     ← work done
12:39:22 … 12:45:54   [WORKER_STATS] slots=0/10   ← idle

[streaming] Consumer batch … total — the very next log line — never appeared. The only statement in between is await entity_resolver.flush_pending_stats().

Cause. flush_pending_stats() acquires its own connection, but was called while the enclosing acquire_with_retry(...) block still held one:

async with acquire_with_retry(pool) as conn:      # conn checked out
    async with conn.transaction():                # SAVEPOINT only
        ...write facts/entities...
    await entity_resolver.flush_pending_stats()   # takes a 2nd connection

oracledb does not autocommit and OracleConnection.transaction() is only a SAVEPOINT, so the write is committed by OracleBackend.acquire() when its block exits. Connection #2's UPDATE entities … waits on row locks held by the still-open connection #1, which cannot commit until the call returns — a circular wait. Oracle never raises ORA-00060, because session #1 is blocked in Python rather than on the database, so it hangs indefinitely instead of erroring. (3× Phase 1 but only 1× Phase 2 in the logs: once retain #1 wedged, later retains blocked in their own write txn.)

Fix. Move the flush after the acquire block in all three call sites — streaming retain, delta retain, and the transfer importer. This is what its own docstring already required ("Must be called AFTER the retain transaction commits"); PostgreSQL satisfied it only by accident, via asyncpg autocommit.

Test. Guarded with an AST lint test (in the spirit of test_migration_shape.py) rather than a behavioural one — the deadlock cannot be reproduced against PostgreSQL, which is what the suite runs on. It includes a self-check so it can't pass vacuously.

Validation

This PR touches .github/workflows/**, so both Oracle client jobs run here and the fixes validate themselves. lint.sh + ty clean; retain suites pass locally (25/25).

Not addressed: the pre-existing ORA-00903: invalid table name warning when writing llm_traces on Oracle. It's caught and logged best-effort on every retain and does not block — worth a separate fix.

The three Oracle jobs run the Oracle 23ai `free` service image, which
together with the Python ML deps (torch) exhausts the runner's ~14 GB root
disk. Two symptoms, one cause:

- uv fails to extract a wheel with "No space left on device (os error 28)"
  (fast ~2 min failure), and
- a near-full disk starves I/O badly enough to trip the 30-minute job
  timeout.

test-python-client-oracle and test-typescript-client-oracle have been red on
every open PR (#2941, #2942, #2943) from this, independent of the code under
test. Reclaim ~20 GB of preinstalled tooling (the same jlumbroso action the
Docker build job already uses) before the Oracle setup step.

docker-images stays false here: unlike the Docker build job, the Oracle
service container is already running by the time steps execute, so pruning
images could disrupt it. The savings come from the tool cache, Android SDK,
.NET, Haskell, large apt packages and swap.
The first pass enabled every reclaim, which cost ~4 minutes of job time —
counterproductive on jobs that are already fighting a 30-minute limit.

android + dotnet + haskell + swap are a few rm -rf's worth ~16-21 GB, which
is ample headroom for the Oracle image plus torch. Dropped:
- large-packages: apt-get remove, costs minutes for little extra space;
- tool-cache: deletes the preinstalled Python that actions/setup-python then
  re-downloads, making the job slower rather than faster.
…e hang)

Retain hung forever on the Oracle backend: every retain test burned its 120s
client timeout while the server sat idle, so test-python-client-oracle and
test-typescript-client-oracle only ever reached ~5% of the suite before the
30-minute job limit.

The server was not slow — it was deadlocked. flush_pending_stats() acquires
its own connection, but it was being called while the enclosing
acquire_with_retry(...) block still held one:

  async with acquire_with_retry(pool) as conn:   # conn checked out
      async with conn.transaction():             # SAVEPOINT only
          ...write facts/entities...
      await entity_resolver.flush_pending_stats()  # takes a 2nd connection

oracledb does not autocommit and OracleConnection.transaction() is only a
SAVEPOINT, so the write is committed by OracleBackend.acquire() when its block
exits. Connection #2's `UPDATE entities ...` therefore waits on row locks held
by the still-open connection #1, which cannot commit until the call returns —
a circular wait. Oracle never reports ORA-00060 because session #1 is blocked
in Python, not on the database, so it hangs indefinitely instead of erroring.

Move the flush after the acquire block in all three call sites (streaming
retain, delta retain, transfer importer), which is what its own docstring
already required ("must be called AFTER the retain transaction commits") and
which PostgreSQL satisfied only by accident via asyncpg autocommit.

Guarded with an AST lint test rather than a behavioural one: the deadlock
cannot be reproduced against PostgreSQL, which is what the suite runs on.
@nicoloboschi nicoloboschi changed the title ci(oracle): free runner disk space before Oracle jobs (fixes repo-wide Oracle CI failures) fix(oracle): unblock Oracle CI — free runner disk space + fix the retain deadlock Jul 24, 2026
test_dry_run_creates_nothing still flaked in test-api shard 3. CONCURRENTLY
avoids ACCESS EXCLUSIVE but still takes ShareUpdateExclusive, which conflicts
with the ShareLock a fresh bank's plain CREATE INDEX holds — and that one
cannot be made concurrent, since it runs inside the bank-create transaction.
So _drop_bank_indexes can still be picked as the deadlock victim while another
xdist worker seeds a bank:

  Process A waits for ShareUpdateExclusiveLock on memory_units; blocked by B.
  Process B waits for ShareLock on virtual transaction; blocked by A.

The bank-create side already retries (#2943); give the drop the same treatment.
The drop is idempotent, so retrying is safe.
@nicoloboschi
nicoloboschi merged commit 6d51575 into main Jul 24, 2026
203 of 204 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant