Skip to content

fix(db): never TRUNCATE-checkpoint a live WAL (SIGBUS under traffic) - #14005

Merged
diegosouzapw merged 4 commits into
diegosouzapw:release/v3.8.51from
HouMinXi:fix/wal-truncate-under-traffic
Sep 18, 2026
Merged

diegosouzapw merged 4 commits into
diegosouzapw:release/v3.8.51from
HouMinXi:fix/wal-truncate-under-traffic

Conversation

@HouMinXi

@HouMinXi HouMinXi commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

Refs #13973.

Long-running servers died about every six hours with SIGBUS (exit 135), no JS stack. A coredump put the fault on the mapped storage.sqlite-shm (WAL-index). The journal showed wal_checkpoint(TRUNCATE) one or two seconds before each death.

Cause

walMaintenance.ts ran wal_checkpoint(TRUNCATE) on a six-hour timer and from the 256 MB size guard. TRUNCATE rewrites the wal-index while other handles still have it mapped. A later read of that mapping is SIGBUS; JS cannot catch it.

No interval makes a live TRUNCATE safe.

Change

  • Drop the periodic TRUNCATE scheduler. PASSIVE checkpoints still run every five minutes; a busy tick records telemetry and retries after 60s.
  • Size guard (OMNIROUTE_WAL_GUARD_MAX_MB, default 256) now runs wal_checkpoint(RESTART) instead of TRUNCATE or a warning-only log. RESTART starts a new WAL without rewriting the mapped index, so the file stays bounded.
  • Shutdown still truncates (closeDbInstance).
  • OMNIROUTE_WAL_TRUNCATE_INTERVAL_MS is ignored; a positive value logs a one-time deprecation warning.

This implements one of the three items on #13973 (stop live TRUNCATE). It does not stagger the other 6h timers and does not add the conversation_turn_nodes(last_seen_at) index, so the issue stays open.

Tests

WAL suites: 38 pass, 0 fail, 1 pre-existing bun:sqlite skip. Injecting the old warn-only guard fails the RESTART assertion; restoring it passes.

Note

OMNIROUTE_STRIP_SYSTEM_PREAMBLE / COMBO_LOOP_SAFETY_TIMEOUT_MS env-doc drift is on the base tip (#14004 / #14054), not this PR.


⚠️ base-red inherited: #13866 — the 16 unit-test failures reproduce on the clean tip 7cc454d9 (6/6 sampled) and are identical across PRs with disjoint content; none touches what this PR changes.

@diegosouzapw

Copy link
Copy Markdown
Owner

Consertei o que travava, falta uma decisão sua.

O diagnóstico aqui é o melhor do lote — coredump, si_addr mapeando em storage.sqlite-shm, correlação temporal. A premissa confere: o tip roda wal_checkpoint(TRUNCATE) num timer de 6h (walMaintenance.ts:319-352) e um segundo TRUNCATE ao vivo no guard de 256 MB dentro do tick PASSIVE (:287).

O que eu consertei: a PR estava vermelha no próprio head — 9 testes falhando, embora o corpo afirme "38 pass / 0 fail". A causa era o helper fnBody(), duplicado em db-wal-passive-scheduler.test.ts:19 e db-wal-scheduler-wiring.test.ts:26: ele foi reescrito para fatiar até a próxima linha em coluna 0, mas o Prettier deixa o ) de fechamento da assinatura multi-linha exatamente na coluna 0, então o corpo nunca entrava na fatia. Substituí por um balanceador de parênteses e chaves. Agora: tests 39 | pass 38 | fail 0 | skipped 1 (o skip é o caso exclusivo de bun:sqlite). typecheck:core limpo, eslint e prettier limpos. Não toquei em walMaintenance.ts.

Decisão pendente — RESTART em vez de só avisar. A PR troca o guard de 256 MB por um console.warn puro, então o WAL passa a encolher apenas no restart. O enum WalCheckpointMode já inclui "RESTART", que reinicia o WAL do zero sem truncar o arquivo — não invalida mapeamento nenhum, logo não tem o risco de SIGBUS, e mantém o WAL limitado. Seria o melhor dos dois. Vale também notar que o warn dispara a cada tick de 5 min enquanto o WAL estiver acima do limite.

Segunda pendência: o corpo diz Fixes #13973, mas a PR implementa 1 das 3 sugestões da issue (ficam de fora escalonar os três timers de 6h que disparam no mesmo segundo, e o índice em conversation_turn_nodes(last_seen_at)). Como está, o merge auto-fecharia a issue indevidamente — precisa virar Refs #13973.

@HouMinXi
HouMinXi force-pushed the fix/wal-truncate-under-traffic branch from a6a9d07 to 2f253ca Compare September 18, 2026 02:12
A live TRUNCATE checkpoint rewrites the shared wal-index while other
processes hold it mapped; dereferencing the stale mapping SIGBUSes the
process. Two production crashes six hours apart, coredump stack in
better-sqlite3 native memcpy (issue diegosouzapw#13973).

Remove the periodic TRUNCATE scheduler. Runtime checkpoints are
PASSIVE-only, which move pages without changing the wal-index geometry,
while TRUNCATE stays on the shutdown path where reclaiming the file is
safe. The 256MB size guard now warns instead of escalating to a live
TRUNCATE, busy PASSIVE ticks feed the persisted busy telemetry that the
TRUNCATE tick used to carry, and a positive
OMNIROUTE_WAL_TRUNCATE_INTERVAL_MS logs a one-time deprecation warning.

Signed-off-by: Minxi Hou <houminxi@gmail.com>
The 256MB size guard only warned, so a live WAL could keep growing until
the next restart. wal_checkpoint(RESTART) starts a new WAL file without
rewriting the mapped wal-index, which is what SIGBUS'd the process when
we used TRUNCATE under traffic.

Related to diegosouzapw#13973.

Signed-off-by: Minxi Hou <houminxi@gmail.com>
@HouMinXi
HouMinXi force-pushed the fix/wal-truncate-under-traffic branch from 2f253ca to 5c219e7 Compare September 18, 2026 06:52
@HouMinXi

Copy link
Copy Markdown
Contributor Author

Rebased onto current release/v3.8.51 and addressed the two review notes:

Tests: 38 pass / 1 skip, matching the numbers from the earlier test-only commit. Injection: swapping RESTART for PASSIVE turns the new source-contract assertion red; restoring it turns it green.

Head is 5c219e770a.

HouMinXi and others added 2 commits September 18, 2026 03:08
Backticks around TRUNCATE made the env/docs checker treat it as a
variable. The VACUUM rows next to it were never part of this change
and are not in the base docs.

Signed-off-by: Minxi Hou <houminxi@gmail.com>
@diegosouzapw
diegosouzapw merged commit 9956f13 into diegosouzapw:release/v3.8.51 Sep 18, 2026
12 of 16 checks passed
@HouMinXi
HouMinXi deleted the fix/wal-truncate-under-traffic branch September 19, 2026 13:08
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…iegosouzapw#14005)

* fix(db): never TRUNCATE-checkpoint a live WAL

A live TRUNCATE checkpoint rewrites the shared wal-index while other
processes hold it mapped; dereferencing the stale mapping SIGBUSes the
process. Two production crashes six hours apart, coredump stack in
better-sqlite3 native memcpy (issue diegosouzapw#13973).

Remove the periodic TRUNCATE scheduler. Runtime checkpoints are
PASSIVE-only, which move pages without changing the wal-index geometry,
while TRUNCATE stays on the shutdown path where reclaiming the file is
safe. The 256MB size guard now warns instead of escalating to a live
TRUNCATE, busy PASSIVE ticks feed the persisted busy telemetry that the
TRUNCATE tick used to carry, and a positive
OMNIROUTE_WAL_TRUNCATE_INTERVAL_MS logs a one-time deprecation warning.

Signed-off-by: Minxi Hou <houminxi@gmail.com>

* fix(db): RESTART the WAL when it exceeds the size guard

The 256MB size guard only warned, so a live WAL could keep growing until
the next restart. wal_checkpoint(RESTART) starts a new WAL file without
rewriting the mapped wal-index, which is what SIGBUS'd the process when
we used TRUNCATE under traffic.

Related to diegosouzapw#13973.

Signed-off-by: Minxi Hou <houminxi@gmail.com>

* docs: drop a fake TRUNCATE env name from the WAL guard row

Backticks around TRUNCATE made the env/docs checker treat it as a
variable. The VACUUM rows next to it were never part of this change
and are not in the base docs.

Signed-off-by: Minxi Hou <houminxi@gmail.com>

---------

Signed-off-by: Minxi Hou <houminxi@gmail.com>
Co-authored-by: Diego Rodrigues de Sa e Souza <diegosouza.pw@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants