Repository navigation
fix(memory): run nightly consolidation off the px-mind tick (#291) - #293
Closed
adrianwedd wants to merge 1 commit into
Closed
adrianwedd wants to merge 1 commit into
adrianwedd wants to merge 1 commit into
Conversation
Consolidation could not succeed under load, and the three reasons were mutually reinforcing — fixing any one alone would not have helped. It ran inline on px-mind's ~60s awareness tick while its declared deadline is 600s, twice px-mind's own 300s staleness window. Honouring the budget and keeping the mind loop alive were mutually exclusive, so the call carried an ad-hoc timeout=180 that silently overrode the declared 600s: every live failure timed out at exactly 180.1s while every success took 30-65s. And attempt 2 was unspendable — memory.MAX_ATTEMPTS_PER_DAY promised two tries while _TYPE_QUOTAS["consolidate"] was 1 and its cooldown 20h. - mind: `_consolidation_tick` now only supervises. It reaps a finished worker, heartbeats the in-flight marker, clears what a previous process left behind, and starts at most one `px-mind-consolidation` daemon thread. It always returns promptly; awareness, reflection and the battery check run behind it. - memory: `state/consolidation_job.json`, keyed on pid + heartbeat. The worker dies with its process, so a marker outliving its owner is always a lie — a restart cannot leave a false in-progress claim, and the cleanup is recorded as a failure rather than silently reset. An unfinished run past JOB_OVERRUN_AFTER_S is reported once, not every 60s. - memory: `consolidate()` passes no `timeout`, and `run_claude_session`'s default is now None, so brain._DEADLINE_S is the single deadline source and the declared 600s is what reaches `ask_brain`. - claude_session: consolidate quota 1 -> 2 (== MAX_ATTEMPTS_PER_DAY) and cooldown 72000 -> 2400 (== RETRY_SPACING_S). 40min spaces the retry *past* the 30-min global cooldown rather than exempting it from one — nobody is waiting on a 3am retry. - px-motd renders an in-flight run additively; no new health status value. Resident-only Claude is untouched: same ask_brain, same single-flight lock, same mailbox, no second process and no cold fallback. Only which thread blocks on it changed. Tests are inert and every duration is synthetic: the ">300s run" is a threading.Event plus a backdated marker, and the 600s deadline is asserted as a plumbed number read out of the request file. No live brain call. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ThqC6Gq4mZyXWnnvGC2a57
Owner
Author
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #291.
#289 made the nightly consolidation failure visible. This makes it
survivable: three defects were reinforcing each other so that success was
structurally impossible under load, and fixing any one alone would not have
helped.
deadline is 600s — twice px-mind's own 300s staleness window. Honouring the
budget and keeping the mind loop alive were mutually exclusive.
timeout=180that silently overrode the declared600s. The tighter number always wins, so the declared budget was never once
reachable: every live failure measured exactly 180.1s, every success 30–65s.
memory.MAX_ATTEMPTS_PER_DAYpromisedtwo tries a night while
_TYPE_QUOTAS["consolidate"]was 1 and_TYPE_COOLDOWNS["consolidate"]was 20 hours. Attempt 1 consumed the onlyslot attempt 2 could ever have used.
Scheduling semantics: before → after
px-mind-consolidation) inside px-mindbrain._DEADLINE_S["consolidate"], the declared valuerun_claude_session(timeout=)int = 300defaultint | None = None— None means "use the kind's declared deadline"_TYPE_QUOTAS["consolidate"]== memory.MAX_ATTEMPTS_PER_DAY)_TYPE_COOLDOWNS["consolidate"]== memory.RETRY_SPACING_S)memory.consolidation_due()consolidatestays out of_GLOBAL_COOLDOWN_EXEMPTstate/consolidation_job.json, pid + heartbeat keyedoverran,abandoned), no new health status enum valuestate/consolidation_meta.jsonlast_attempt_tsfield)Why a daemon thread + a state-file job record
The brief asked for the smallest existing repo pattern. This is two of them
composed, no new machinery:
wander._ExploringRefresher— the in-daemonthreading.Thread(daemon=True)idiom, for work that must not block its owner's loop.
/proc/{pid}single-instance guard — the liveness idiom,for deciding whether a marker on disk still has an owner.
Alternatives considered and rejected:
api.py's async-job pattern (_jobsdict + worker thread, 202 +job_id).Right shape, wrong scope:
_jobsis in-memory and process-local, and px-mindhas no HTTP surface for anyone to poll. The part that is actually needed —
cross-process visibility — is exactly the part that pattern does not provide.
bin/px-consolidate. A second Python interpreter anda second import of
pxhon a memory-tight Pi (ops(kernel): restore usable memory accounting and pressure instrumentation #218/perf(voice): px-wake-listen holds ~689 MiB anonymous, 446 MiB in one unreclaimable heap (900 MiB peak) #219/fix(brain,mic,ops): contain operator Claude memory, extend PSI/cgroup observability #286), to do workthat is 99% waiting on a mailbox file. It would also need its own
single-instance guard, its own health reporting and its own systemd unit — more
moving parts than the problem has.
of the cognitive loop that owns it and duplicates the window/attempt logic in
unit files. Worth revisiting only if the worker ever needs to outlive px-mind.
The thread dying with its process is a feature, not a limitation: it is what
makes "a marker whose pid is gone" unambiguously a lie rather than a maybe.
Proof that px-mind cannot be stalled
tests/test_consolidation_background_job.py::test_tick_stays_responsive_while_consolidation_exceeds_300sThe worker blocks on a
threading.Eventthe test owns and never sets duringthe loop. The tick is then called five more times with the job marker
backdated past 600s, and each iteration asserts:
Reaching the end of that loop is the proof: a tick that waited on the worker
could not return at all, because the only thing that can release the worker is
the test itself, after the loop. It is deliberately structural rather than a
stopwatch — this Pi is the live robot and routinely sits above a load average
of 10 (26 while writing this), where a single
fsynchas been measured takingtens of seconds. A wall-clock budget would have proved nothing extra and failed
for the wrong reason; an earlier draft of this test did exactly that (a tick
measured at 35.5s that was not blocked on anything).
Supporting assertions in the same file:
test_no_duplicate_concurrent_consolidation— three further ticks while theworker runs start nothing new (
len(calls) == 1, same thread object), andclaim_consolidation_job()independently returnsFalse, so the guard doesnot rest on a process-local variable alone.
test_the_tick_heartbeats_the_marker_while_the_worker_runs— the heartbeat iswritten by the tick, not the worker (the worker is blocked in
ask_brainand could not beat if it wanted to).
test_an_overrunning_worker_is_reported_exactly_once— four ticks pastJOB_OVERRUN_AFTER_Sproduceconsecutive_failures == 1, not four.test_a_marker_from_a_dead_owner_is_not_a_running_claim— a marker with adead pid is cleared and recorded as
abandoned, and the next tick is free torun; no persisted "running" marker survives a restart as a false claim.
test_a_marker_with_a_silent_heartbeat_is_stale— a live pid is not enough;a process that exists and stopped ticking is stale, and so is a marker with no
heartbeat at all (absent-is-stale, so an unreadable marker can never block a
night forever).
Proof that the declared 600s is what reaches the brain
Three layers, no live call and no real wait anywhere:
test_consolidate_passes_no_ad_hoc_timeout—memory.consolidate()callsrun_claude_sessionwith notimeoutkwarg at all.test_run_claude_session_defaults_to_the_declared_deadline— the default isNone, andNoneis what arrives atask_brain;brain.deadline_for_kind ("consolidate") == 600.test_the_declared_600s_is_what_reaches_the_request—ask_brainis drivenwith
session_statestubbed validated andtmux_claude.injectstubbed tocapture-then-fail, so it writes the real request into conftest's tmp mailbox
and returns immediately. The assertion is on the number in that file:
580 <= request["deadline"] - before <= 601. A window rather than anequality because
ask_braindeducts validation/lock wait from the budget.Other callers passing ad-hoc timeouts for classified kinds (listed, not fixed)
src/pxh/mind.py:3498timeout=600self_debug= 900bin/px-blog:739timeout=300blog= 300bin/tool-blog:59timeout=300blog= 300bin/tool-research:48timeout=300research= 300bin/tool-compose:49timeout=300compose= 300src/pxh/vision.pytimeout_s=CLAUDE_TIMEOUT→ask_brain("describe_scene")DESCRIBE_SCENE_TIMEOUTbudget bytests/test_wander.pybin/px-evolve×4timeout=…evolveis not brain-routed and raisesColdStartForbiddenOnly the
self_debugone is a live defect. Out of scope here; left for afollow-up so this PR stays about #291.
Invariants preserved
reaches the same
ask_brain→spark-brainmailbox as before, through thesame single-flight
FileLock; the only change is which thread blocks on it.python tools/check_resident_claude.py→ clean.ask_brainconcurrently — there is oneworker at a time, enforced twice over (process-local thread handle and the
pid-keyed marker under a
FileLock).health.pystill reports only what finished.bin/px-motdrenders thein-flight hint additively (
long-term memory (consolidating now for 3m) last formed 20h ago) and ignores a marker whose heartbeat has gone quiet.state/consolidation_meta.json, success-or-correct-skip marks the date done —all still pinned, in
tests/test_memory.pyandtest_the_window_and_meta_shape_are_unchanged.Tests
All inert: no service touched, no tmux session reached, no Claude call, no
restart. Every duration is synthetic — the "600s" is asserted as a plumbed
number and the ">300s run" is an
Eventplus a backdated marker. There is nomanufactured live 600s brain call anywhere.
One unrelated red flag seen along the way, for the record:
tests/test_health.py::test_overall_is_the_worst_componentfailed once inside a915s combined run (
assert 'stale' == 'ok') and passes in isolation — thatrun was starved enough for
px-mind's 300s window to elapse betweenrecord_successandread_health. Nothing here toucheshealth.py's statusderivation; it is the repo's known "flaky suite under load" class, and CI is the
gate.
Run targeted (
-m "not live") rather than as a full suite — this checkout is onthe robot itself and CI is the gate.
Filed separately
#292 — 82 orphaned
state/health/tmp*.tmpfiles (2026-08-06 → 2026-08-23).Investigated while here; the cause is
atomic_write's temp file surviving aSIGKILL (all are mode
0600, i.e. pre-chmod; complete-but-unrenamed at169/338 bytes or empty at 0), which no code inside the dying process can clean
up — the remedy is a sweeper, which is a design decision rather than a tiny
proven fix. None have appeared since 2026-08-23, i.e. since the #286/#219
restart fixes landed. The live files were not deleted — they are the
evidence.