fix(db): never TRUNCATE-checkpoint a live WAL (SIGBUS under traffic) - #14005
diegosouzapw merged 4 commits into
Conversation
|
Consertei o que travava, falta uma decisão sua. O diagnóstico aqui é o melhor do lote — coredump, O que eu consertei: a PR estava vermelha no próprio head — 9 testes falhando, embora o corpo afirme "38 pass / 0 fail". A causa era o helper Decisão pendente — Segunda pendência: o corpo diz |
a6a9d07 to
2f253ca
Compare
A live TRUNCATE checkpoint rewrites the shared wal-index while other processes hold it mapped; dereferencing the stale mapping SIGBUSes the process. Two production crashes six hours apart, coredump stack in better-sqlite3 native memcpy (issue diegosouzapw#13973). Remove the periodic TRUNCATE scheduler. Runtime checkpoints are PASSIVE-only, which move pages without changing the wal-index geometry, while TRUNCATE stays on the shutdown path where reclaiming the file is safe. The 256MB size guard now warns instead of escalating to a live TRUNCATE, busy PASSIVE ticks feed the persisted busy telemetry that the TRUNCATE tick used to carry, and a positive OMNIROUTE_WAL_TRUNCATE_INTERVAL_MS logs a one-time deprecation warning. Signed-off-by: Minxi Hou <houminxi@gmail.com>
The 256MB size guard only warned, so a live WAL could keep growing until the next restart. wal_checkpoint(RESTART) starts a new WAL file without rewriting the mapped wal-index, which is what SIGBUS'd the process when we used TRUNCATE under traffic. Related to diegosouzapw#13973. Signed-off-by: Minxi Hou <houminxi@gmail.com>
2f253ca to
5c219e7
Compare
|
Rebased onto current
Tests: 38 pass / 1 skip, matching the numbers from the earlier test-only commit. Injection: swapping RESTART for PASSIVE turns the new source-contract assertion red; restoring it turns it green. Head is |
Backticks around TRUNCATE made the env/docs checker treat it as a variable. The VACUUM rows next to it were never part of this change and are not in the base docs. Signed-off-by: Minxi Hou <houminxi@gmail.com>
9956f13
into
diegosouzapw:release/v3.8.51
…iegosouzapw#14005) * fix(db): never TRUNCATE-checkpoint a live WAL A live TRUNCATE checkpoint rewrites the shared wal-index while other processes hold it mapped; dereferencing the stale mapping SIGBUSes the process. Two production crashes six hours apart, coredump stack in better-sqlite3 native memcpy (issue diegosouzapw#13973). Remove the periodic TRUNCATE scheduler. Runtime checkpoints are PASSIVE-only, which move pages without changing the wal-index geometry, while TRUNCATE stays on the shutdown path where reclaiming the file is safe. The 256MB size guard now warns instead of escalating to a live TRUNCATE, busy PASSIVE ticks feed the persisted busy telemetry that the TRUNCATE tick used to carry, and a positive OMNIROUTE_WAL_TRUNCATE_INTERVAL_MS logs a one-time deprecation warning. Signed-off-by: Minxi Hou <houminxi@gmail.com> * fix(db): RESTART the WAL when it exceeds the size guard The 256MB size guard only warned, so a live WAL could keep growing until the next restart. wal_checkpoint(RESTART) starts a new WAL file without rewriting the mapped wal-index, which is what SIGBUS'd the process when we used TRUNCATE under traffic. Related to diegosouzapw#13973. Signed-off-by: Minxi Hou <houminxi@gmail.com> * docs: drop a fake TRUNCATE env name from the WAL guard row Backticks around TRUNCATE made the env/docs checker treat it as a variable. The VACUUM rows next to it were never part of this change and are not in the base docs. Signed-off-by: Minxi Hou <houminxi@gmail.com> --------- Signed-off-by: Minxi Hou <houminxi@gmail.com> Co-authored-by: Diego Rodrigues de Sa e Souza <diegosouza.pw@gmail.com>
Refs #13973.
Long-running servers died about every six hours with SIGBUS (exit 135), no JS stack. A coredump put the fault on the mapped
storage.sqlite-shm(WAL-index). The journal showedwal_checkpoint(TRUNCATE)one or two seconds before each death.Cause
walMaintenance.tsranwal_checkpoint(TRUNCATE)on a six-hour timer and from the 256 MB size guard. TRUNCATE rewrites the wal-index while other handles still have it mapped. A later read of that mapping is SIGBUS; JS cannot catch it.No interval makes a live TRUNCATE safe.
Change
OMNIROUTE_WAL_GUARD_MAX_MB, default 256) now runswal_checkpoint(RESTART)instead of TRUNCATE or a warning-only log. RESTART starts a new WAL without rewriting the mapped index, so the file stays bounded.closeDbInstance).OMNIROUTE_WAL_TRUNCATE_INTERVAL_MSis ignored; a positive value logs a one-time deprecation warning.This implements one of the three items on #13973 (stop live TRUNCATE). It does not stagger the other 6h timers and does not add the
conversation_turn_nodes(last_seen_at)index, so the issue stays open.Tests
WAL suites: 38 pass, 0 fail, 1 pre-existing bun:sqlite skip. Injecting the old warn-only guard fails the RESTART assertion; restoring it passes.
Note
OMNIROUTE_STRIP_SYSTEM_PREAMBLE/COMBO_LOOP_SAFETY_TIMEOUT_MSenv-doc drift is on the base tip (#14004 / #14054), not this PR.7cc454d9(6/6 sampled) and are identical across PRs with disjoint content; none touches what this PR changes.