Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
9 changes: 6 additions & 3 deletions .env.example
Original file line number Diff line number Diff line change
Expand Up @@ -1163,15 +1163,18 @@ PROVIDER_LIMITS_SYNC_SPACING_MS=1500
# Used by: src/lib/db/core.ts::getDbHealthCheckIntervalMs().
#OMNIROUTE_DB_HEALTHCHECK_INTERVAL_MS=21600000

# WAL truncate cadence override (ms). Set to 0 to disable. Default: 21600000 (6h).
# Used by: src/lib/db/core.ts::getWalTruncateIntervalMs().
# Removed: periodic live wal_checkpoint(TRUNCATE) could SIGBUS the process (issue
# #13973). The variable is inert: a positive value logs a one-time deprecation warning,
# while 0 or unset stays silent. The WAL is kept small
# by the PASSIVE scheduler below and truncated by the shutdown checkpoint.
#OMNIROUTE_WAL_TRUNCATE_INTERVAL_MS=21600000

# Frequent wal_checkpoint(PASSIVE) cadence (ms). Set to 0 to disable. Default: 300000 (5m).
# Used by: src/lib/db/walMaintenance.ts.
#OMNIROUTE_WAL_PASSIVE_INTERVAL_MS=300000

# WAL size (MB) above which a PASSIVE tick escalates to wal_checkpoint(TRUNCATE). Default: 256.
# WAL size (MB) above which a PASSIVE tick runs wal_checkpoint(RESTART) so the
# WAL starts over without rewriting the mapped wal-index. Default: 256.
# Used by: src/lib/db/walMaintenance.ts.
#OMNIROUTE_WAL_GUARD_MAX_MB=256

Expand Down
1 change: 1 addition & 0 deletions changelog.d/fixes/wal-truncate-sigbus.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
- **fix(db):** remove the periodic `wal_checkpoint(TRUNCATE)` scheduler. Truncating the WAL rewrites the shared wal-index (`storage.sqlite-shm`) while other connections and in-flight statements still hold it mapped, which crashed long-running servers with SIGBUS roughly every six hours (#13973). WAL hygiene is unchanged: PASSIVE checkpoints still run every five minutes, now count busy contention and retry after 60s instead of silently logging, and a WAL above `OMNIROUTE_WAL_GUARD_MAX_MB` runs `wal_checkpoint(RESTART)` so the file stays bounded without rewriting the mapped index. The shutdown checkpoint still truncates. `OMNIROUTE_WAL_TRUNCATE_INTERVAL_MS` is ignored and logs a deprecation warning.
7 changes: 3 additions & 4 deletions docs/reference/ENVIRONMENT.md
Original file line number Diff line number Diff line change
Expand Up @@ -101,13 +101,12 @@ OmniRoute uses **SQLite** (via `better-sqlite3`) for all persistence. These vari
| `OMNIROUTE_CRYPT_KEY` | _(unset)_ | `src/lib/db/encryption.ts` | **Legacy alias** for `STORAGE_ENCRYPTION_KEY`. Accepted as a fallback when the primary variable is absent. |
| `OMNIROUTE_API_KEY_BASE64` | _(unset)_ | `src/lib/db/encryption.ts` | **Legacy alias** (Base64-encoded form) accepted as a fallback. Decoded automatically before use. |
| `OMNIROUTE_DB_HEALTHCHECK_INTERVAL_MS` | _(unset)_ | `src/lib/db/core.ts` | Override the periodic SQLite healthcheck interval (ms). When unset, defaults are derived from `NODE_ENV`. |
| `OMNIROUTE_WAL_TRUNCATE_INTERVAL_MS` | `21600000` (6h) | `src/lib/db/walMaintenance.ts` | Override the periodic `wal_checkpoint(TRUNCATE)` interval (ms). Auto-checkpoint never shrinks the WAL file itself, and a long-running server never closes its DB. `0` disables. |
| `OMNIROUTE_WAL_PASSIVE_INTERVAL_MS` | `300000` (5m) | `src/lib/db/walMaintenance.ts` | Override the frequent `wal_checkpoint(PASSIVE)` interval (ms). Keeps pending WAL frames small so the periodic TRUNCATE never copies a multi-GB backlog on the main thread. `0` disables. |
| `OMNIROUTE_WAL_GUARD_MAX_MB` | `256` | `src/lib/db/walMaintenance.ts` | When a PASSIVE tick finds the WAL file above this size, escalate to `wal_checkpoint(TRUNCATE)` immediately instead of waiting for the slow tick. |
| `OMNIROUTE_WAL_TRUNCATE_INTERVAL_MS` | _(removed)_ | `src/lib/db/walMaintenance.ts` | **Removed.** A periodic live `wal_checkpoint(TRUNCATE)` can invalidate the shared wal-index mapping and crash the process with SIGBUS (#13973), so the scheduler no longer exists. The variable is inert: a positive value logs a one-time deprecation warning, while `0` or unset stays silent. The WAL is maintained by PASSIVE checkpoints (below) and truncated by the shutdown checkpoint. |
| `OMNIROUTE_WAL_PASSIVE_INTERVAL_MS` | `300000` (5m) | `src/lib/db/walMaintenance.ts` | Override the frequent `wal_checkpoint(PASSIVE)` interval (ms). Keeps pending WAL frames small so checkpoints stay fast and the WAL file stays bounded between shutdown truncations. `0` disables. |
| `OMNIROUTE_WAL_GUARD_MAX_MB` | `256` | `src/lib/db/walMaintenance.ts` | When a PASSIVE tick finds the WAL file above this size, run `wal_checkpoint(RESTART)` so the WAL starts over without rewriting the mapped wal-index. Live truncate-mode checkpoints were removed (see the `OMNIROUTE_WAL_TRUNCATE_INTERVAL_MS` row). |
| `OMNIROUTE_PRESSURE_SELF_RESTART` | `false` | `open-sse/utils/resourcePressure.ts` | Set to `1`/`true`/`yes`/`on` to exit the process after critical resource pressure is sustained for `OMNIROUTE_PRESSURE_SELF_RESTART_AFTER_MS`, letting a supervisor (systemd `Restart=always`, Docker restart policy) bring back a clean process instead of serving 503s indefinitely. |
| `OMNIROUTE_PRESSURE_SELF_RESTART_AFTER_MS` | `120000` (2m) | `open-sse/utils/resourcePressure.ts` | How long critical pressure must persist before the self-restart exit fires. |
| `OMNIROUTE_SQLJS_WASM_PATH` | _(auto-detect)_ | `src/lib/db/adapters/sqljsAdapter.ts` | Explicit path (absolute or relative to cwd) to `sql-wasm.wasm` when using the `sql.js` WASM fallback adapter. Auto-detected via package dependencies and candidate layouts when unset. |
| `OMNIROUTE_WAL_TRUNCATE_INTERVAL_MS` | `21600000` (6h) | `src/lib/db/core.ts` | Override the periodic `wal_checkpoint(TRUNCATE)` interval (ms). Auto-checkpoint never shrinks the WAL file itself, and a long-running server never closes its DB. `0` disables. |
| `OMNIROUTE_BATCH_RETENTION_DAYS` | `30` | `src/lib/db/cleanup.ts` | Days a terminal (completed/failed/cancelled/expired) Batch API job's checkpoints, referenced input/output/error files, and row are kept by the automatic cleanup sweep before deletion. Only takes effect once `BATCH_AND_FILE_AUTO_CLEANUP_ENABLED` is turned on; matches OpenAI's own Batch API output retention window. Does not affect the operator-triggered `DELETE /api/v1/batches/delete-completed` route, which stays unconditional (no age filter) by design. |
| `BATCH_AND_FILE_AUTO_CLEANUP_ENABLED` | `false` | `src/lib/db/cleanup.ts` | When `true`, let the automatic cleanup sweep delete terminal Batch API jobs (and their checkpoints) past `OMNIROUTE_BATCH_RETENTION_DAYS`, and clear the BLOB content of uploaded files past their own `expires_at`. Off by default: every existing install keeps this data exactly as before until an operator opts in. Also a dashboard-editable feature flag — see `docs/reference/FEATURE_FLAGS.md` → Runtime. |
| `OMNIROUTE_SKIP_DB_HEALTHCHECK` | `0` | `src/lib/db/core.ts`, `src/lib/db/healthCheck.ts` | Set to `1` to skip the DB healthcheck entirely on startup. Useful for short-lived tasks and integration tests. |
Expand Down
7 changes: 5 additions & 2 deletions src/lib/db/core.ts
Original file line number Diff line number Diff line change
Expand Up @@ -939,8 +939,11 @@ function startDbHealthCheckScheduler(db: SqliteDatabase) {
}

// Auto-checkpoint moves WAL pages back into the main DB file but never shrinks the WAL
// file itself; only wal_checkpoint(TRUNCATE) does, and a long-running server never closes its DB.
// The scheduler lives in ./walMaintenance (periodic TRUNCATE + busy warn + PASSIVE retry).
// file itself; only wal_checkpoint(TRUNCATE) does. TRUNCATE runs at shutdown
// (closeDbInstance) and never on a live timer: truncating a live process's WAL rewrites
// the shared wal-index under handles that hold it mapped and can SIGBUS the event loop
// (issue #13973). Runtime maintenance lives in ./walMaintenance (PASSIVE + busy warn +
// retry).

const healthShutdown = new AbortController();
const managedHealth = createDbHealthCoordinator(async (autoRepair, skipIntegrity) => {
Expand Down
137 changes: 70 additions & 67 deletions src/lib/db/walMaintenance.ts
Original file line number Diff line number Diff line change
Expand Up @@ -5,9 +5,16 @@ import type { SqliteAdapter } from "./adapters/types";
import { registerDbStateResetter } from "./stateReset";

/**
* WAL maintenance owns the periodic `wal_checkpoint(TRUNCATE)` lifecycle that
* used to live inside `core.ts`: interval parsing, the scheduler, and reading
* the pragma result so a busy checkpoint warns instead of logging success.
* WAL maintenance owns the periodic `wal_checkpoint(PASSIVE)` lifecycle that
* used to live inside `core.ts`: interval parsing, the scheduler, busy
* accounting, and reading the pragma result so a busy checkpoint warns
* instead of logging success.
*
* There is deliberately no periodic TRUNCATE: truncating the WAL of a live
* process rewrites the shared wal-index (storage.sqlite-shm) while other
* handles and in-flight statements hold it mapped, which can crash the event
* loop with SIGBUS (issue #13973). A WAL above the size guard uses RESTART
* instead, which starts a new WAL file without rewriting the mapped index.
*/
export type WalCheckpointMode = "PASSIVE" | "FULL" | "RESTART" | "TRUNCATE";

Expand Down Expand Up @@ -36,15 +43,13 @@ export interface WalMaintenanceState {

const isCloud = typeof globalThis.caches === "object" && globalThis.caches !== null;

const DEFAULT_WAL_TRUNCATE_INTERVAL_MS = 6 * 60 * 60 * 1000;
const DEFAULT_WAL_PASSIVE_INTERVAL_MS = 5 * 60 * 1000;
const DEFAULT_WAL_GUARD_MAX_BYTES = 256 * 1024 * 1024;
const RETRY_DELAY_MS = 60_000;

export const WAL_BUSY_NAMESPACE = "walMaintenance";
export const WAL_BUSY_KEY = "busyTotal";

let walTimer: NodeJS.Timeout | null = null;
let walPassiveTimer: NodeJS.Timeout | null = null;
let retryTimer: NodeJS.Timeout | null = null;
let ticks = 0;
Expand All @@ -56,6 +61,29 @@ let lastOkAt: string | null = null;
let pendingBusyDelta = 0;
// The handle the running scheduler was started with; used for the shutdown flush.
let activeDb: SqliteAdapter | null = null;
let truncateDeprecationWarned = false;

/**
* Operators who set OMNIROUTE_WAL_TRUNCATE_INTERVAL_MS to a positive value believe a
* periodic TRUNCATE is reclaiming their WAL on a timer. It is not: that scheduler was
* removed because a live TRUNCATE can SIGBUS the process (issue #13973). Warn once per
* process so the stale setting is visible instead of silently ignored.
*/
function warnPeriodicTruncateRemoved(env: NodeJS.ProcessEnv): void {
if (truncateDeprecationWarned) return;
const rawValue = env.OMNIROUTE_WAL_TRUNCATE_INTERVAL_MS;
if (typeof rawValue !== "string" || rawValue.trim().length === 0) return;
const parsed = Number(rawValue);
if (!Number.isFinite(parsed) || parsed <= 0) return;
truncateDeprecationWarned = true;
console.warn(
"[DB] OMNIROUTE_WAL_TRUNCATE_INTERVAL_MS is no longer used and has no effect: periodic " +
"live wal_checkpoint(TRUNCATE) was removed because truncating the WAL of a live " +
"process can invalidate the shared wal-index mapping and crash the server with " +
"SIGBUS (issue #13973). Runtime checkpoints are PASSIVE; the shutdown checkpoint " +
"truncates the WAL."
);
}

/**
* Count one busy checkpoint. Memory only, on purpose: a busy checkpoint means the
Expand Down Expand Up @@ -160,17 +188,6 @@ export function runCheckpointNow(
}
}

export function getWalMaintenanceIntervalMs(env: NodeJS.ProcessEnv = process.env): number {
const rawValue = env.OMNIROUTE_WAL_TRUNCATE_INTERVAL_MS;
if (typeof rawValue === "string" && rawValue.trim().length > 0) {
const parsed = Number(rawValue);
if (Number.isFinite(parsed) && parsed >= 0) {
return parsed;
}
}
return DEFAULT_WAL_TRUNCATE_INTERVAL_MS;
}

export function getWalPassiveIntervalMs(env: NodeJS.ProcessEnv = process.env): number {
const rawValue = env.OMNIROUTE_WAL_PASSIVE_INTERVAL_MS;
if (typeof rawValue === "string" && rawValue.trim().length > 0) {
Expand Down Expand Up @@ -214,7 +231,8 @@ export function logCheckpointOutcome(
if (outcome.skipped) return;
if (outcome.busy) {
console.warn(
`[DB] SQLite WAL checkpoint busy — ${outcome.logFrames} frames pending (streak ${streak})`
`[DB] SQLite WAL checkpoint busy - ${outcome.logFrames} frames pending, ` +
`${outcome.checkpointedFrames ?? 0} checkpointed this attempt (streak ${streak})`
);
return;
}
Expand Down Expand Up @@ -274,26 +292,46 @@ function startWalPassiveScheduler(
isBuildPhase: isNextBuildPhase(),
});
if (stats.skipped) return;
if (!stats.busy && stats.ok) flushBusyTotal(db);
if (stats.busy || (stats.checkpointedFrames ?? 0) > 0) {
console.log(
`[DB] WAL passive checkpoint (busy=${stats.busy ? 1 : 0} logFrames=${stats.logFrames} ` +
`checkpointedFrames=${stats.checkpointedFrames} walMb=${formatWalMb(walBeforeBytes)})`
);
}
ticks++;
// Check the size guard on every tick, busy or not: a WAL that stays above the
// guard while readers hold the database is exactly the case the operator must
// hear about, and a busy tick must not hide it.
const guardMaxBytes = getWalGuardMaxBytes(env);
if (walBeforeBytes != null && walBeforeBytes > guardMaxBytes) {
const startedAtMs = Date.now();
const truncateStats = runCheckpointNow(db, "TRUNCATE", {
// Never TRUNCATE a live WAL: rewriting the shared wal-index under handles that
// hold it mapped can SIGBUS the process (issue #13973). RESTART checkpoints
// the WAL and starts a new one without changing the mapped file geometry.
const restart = runCheckpointNow(db, "RESTART", {
sqliteFile,
isCloud,
isBuildPhase: isNextBuildPhase(),
});
console.log(
console.warn(
`[DB] WAL above guard (${formatWalMb(walBeforeBytes)}MB > ${Math.floor(guardMaxBytes / (1024 * 1024))}MB); ` +
`ran TRUNCATE in ${Date.now() - startedAtMs}ms ` +
`(walMbAfter=${formatWalMb(getWalFileSizeBytes(sqliteFile))} busy=${truncateStats.busy ? 1 : 0} ` +
`checkpointedFrames=${truncateStats.checkpointedFrames})`
`ran wal_checkpoint(RESTART) ok=${restart.ok} busy=${restart.busy}` +
` checkpointedFrames=${restart.checkpointedFrames}` +
(restart.error ? ` error=${restart.error}` : "")
);
}
if (stats.busy) {
// Passive ticks carry the busy telemetry now that no TRUNCATE tick is left to
// do it: a contended PASSIVE is the same "readers never let go" signal, and the
// 60s retry gives a busy WAL a second chance long before the next 5m tick.
recordBusy();
logCheckpointOutcome(stats, "PASSIVE", busyStreak);
schedulePassiveRetry(db);
return;
}
if (stats.ok) {
recordOk();
flushBusyTotal(db);
} else {
logCheckpointOutcome(stats, "PASSIVE", busyStreak);
}
if ((stats.checkpointedFrames ?? 0) > 0) {
console.log(
`[DB] WAL passive checkpoint (logFrames=${stats.logFrames} ` +
`checkpointedFrames=${stats.checkpointedFrames} walMb=${formatWalMb(walBeforeBytes)})`
);
}
} catch (error: unknown) {
Expand All @@ -309,46 +347,14 @@ export function startWalMaintenance(
sqliteFile: string | null,
env: NodeJS.ProcessEnv = process.env
): void {
warnPeriodicTruncateRemoved(env);
// stopWalMaintenance() flushes what it can and zeroes session state, so capture the
// in-memory total first; the gate stays before any DB touch.
const priorBusyTotal = busyTotal;
stopWalMaintenance();
if (sqliteFile === null || isCloud || isNextBuildPhase() || isAutomatedTestProcess()) return;
activeDb = db;
busyTotal = mergeBusyTotal(priorBusyTotal, loadPersistedBusyTotal(db));
const intervalMs = getWalMaintenanceIntervalMs(env);
if (intervalMs <= 0) {
startWalPassiveScheduler(db, sqliteFile, env);
return;
}
walTimer = setInterval(() => {
try {
if (!db.open) return;
const walBeforeBytes = getWalFileSizeBytes(sqliteFile);
const startedAtMs = Date.now();
const outcome = runCheckpointNow(db, "TRUNCATE");
if (outcome.skipped) return;
ticks++;
if (outcome.busy) {
recordBusy();
logCheckpointOutcome(outcome, "TRUNCATE", busyStreak);
schedulePassiveRetry(db);
} else if (outcome.ok) {
recordOk();
flushBusyTotal(db);
console.log(
`[DB] Periodic SQLite WAL checkpoint completed (TRUNCATE) in ${Date.now() - startedAtMs}ms ` +
`(walMbBefore=${formatWalMb(walBeforeBytes)} walMbAfter=${formatWalMb(getWalFileSizeBytes(sqliteFile))} ` +
`busy=${outcome.busy ? 1 : 0} logFrames=${outcome.logFrames} checkpointedFrames=${outcome.checkpointedFrames})`
);
} else {
logCheckpointOutcome(outcome, "TRUNCATE", busyStreak);
}
} catch {
// A periodic scheduler must never throw into the event loop.
}
}, intervalMs);
walTimer.unref?.();
startWalPassiveScheduler(db, sqliteFile, env);
}

Expand All @@ -357,10 +363,6 @@ export function stopWalMaintenance(): void {
flushBusyTotal(activeDb);
activeDb = null;
pendingBusyDelta = 0;
if (walTimer) {
clearInterval(walTimer);
walTimer = null;
}
if (walPassiveTimer) {
clearInterval(walPassiveTimer);
walPassiveTimer = null;
Expand Down Expand Up @@ -403,6 +405,7 @@ export function mergeBusyTotal(prior: number, loaded: number): number {

export function __resetForTests(): void {
stopWalMaintenance();
truncateDeprecationWarned = false;
}

registerDbStateResetter(stopWalMaintenance);
Loading
Loading