Skip to content

fix(state): enforce write patience during lock acquisition - #79140

Open
RonLi79 wants to merge 2 commits into
NousResearch:mainfrom
RonLi79:fix/sqlite-write-patience-budget
Open

RonLi79 wants to merge 2 commits into
NousResearch:mainfrom
RonLi79:fix/sqlite-write-patience-budget

Conversation

@RonLi79

@RonLi79 RonLi79 commented Aug 5, 2026 •

Copy link
Copy Markdown

Summary

Make _execute_write()'s application-level patience_s budget authoritative for SQLite write-lock acquisition.

SessionDB keeps the existing 1-second SQLite timeout for direct connection operations. During BEGIN IMMEDIATE, it now snapshots PRAGMA busy_timeout, temporarily sets it to 0, and restores the original value in finally. The randomized application retry loop therefore owns the complete wall-clock budget without leaking a fail-fast timeout to later connection users.

Root cause

The writer connection is opened with a native SQLite busy timeout, while some application write classes use a shorter patience budget (notably 0.5 seconds for S1 activity writes). BEGIN IMMEDIATE could block inside SQLite's busy handler before _execute_write() regained control to check its deadline. The configured application budget was therefore not end-to-end enforceable.

Changes

  • Snapshot and restore the connection's existing PRAGMA busy_timeout around write-lock acquisition.
  • Set busy_timeout=0 only for BEGIN IMMEDIATE, so lock contention fails fast into the existing deadline-bounded jitter loop.
  • Use a local checked connection reference for the full transaction.
  • Add regression tests proving timeout restoration after both a successful write and exhausted lock patience.

Verification

Fresh evidence on rebased commit c22367383:

  • 15 passed — tests/state/test_write_lock_patience.py plus tests/gateway/test_watchdog_review_76354.py
  • 40 passed — complete tests/state plus the gateway watchdog regression
  • ruff check hermes_state.py tests/state/test_write_lock_patience.py — clean
  • py_compile for implementation and test — clean
  • git diff --check upstream/main...HEAD — clean
  • Manual probe: a non-default busy_timeout=73ms is restored after both success and lock exhaustion

The pre-fix release candidate reproduced the short-budget failures at roughly 3.7–4.1 seconds despite a 0.5-second budget. The patched snapshot keeps the existing retry/error semantics while enforcing the configured wall-clock deadline.

Independent review

Read-only diff review with Codex gpt-5.3-codex-spark: APPROVE, no blockers. Claude Fable was attempted first but its OAuth session was expired; no Fable verdict is claimed.

@RonLi79
RonLi79 force-pushed the fix/sqlite-write-patience-budget branch from 8734c16 to c223673 Compare August 5, 2026 07:37
@RonLi79

RonLi79 commented Aug 5, 2026

Copy link
Copy Markdown
Author

Rebased onto current upstream/main and added explicit busy_timeout restoration regressions. Fresh local gates: 15/15 focused and 40/40 complete state/watchdog tests, plus Ruff, py_compile, and diff-check. The PR currently has no instantiated check suite; please approve the fork workflow and review when convenient.

@alt-glitch alt-glitch added type/bug Something isn't working P2 Medium — degraded but workaround exists comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint area/sessions Session lifecycle, resume, persistence, history sweeper:risk-session-state Sweeper risk: may lose/corrupt/mis-associate session or context state labels Aug 5, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/sessions Session lifecycle, resume, persistence, history comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint P2 Medium — degraded but workaround exists sweeper:risk-session-state Sweeper risk: may lose/corrupt/mis-associate session or context state type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants