Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
53 changes: 51 additions & 2 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -153,6 +153,39 @@ threshold while the store has not grown for days. That is why `overdue`
(>48h) is a separate axis from `status`, and why `summarize()` returns a
non-empty string for overdue memory on an otherwise clean bill.

**Consolidation runs on a background daemon thread, never on the tick (#291).**
`_consolidation_tick` only *supervises*: it heartbeats the in-flight marker,
reaps what a previous process left behind, and decides whether to start a
worker. It must always return promptly, because awareness, reflection and the
battery check run behind it. The reason is arithmetic, not taste — the kind's
declared deadline is **600s** and px-mind's own staleness window is **300s**, so
an inline call that honoured its budget guaranteed px-mind read `stale`. That
made the three defects mutually reinforcing: the pass could not be given enough
time, so it was given an ad-hoc `timeout=180` that silently overrode the
declared deadline, so every failure timed out at exactly 180.1s.

**`state/consolidation_job.json` is the in-flight marker**, and it is keyed on
**pid plus a heartbeat** — `{"status","pid","attempt","started_ts","heartbeat_ts"}`.
A health record says what the last *finished* attempt did; nothing else answers
"is one running right now". The worker thread dies with its process, so a marker
outliving its owner is always a lie: `consolidation_job_is_stale()` calls it
stale when `/proc/<pid>` is gone (or is not px-mind) **or** the heartbeat has
been quiet past `JOB_HEARTBEAT_STALE_S`. A restart therefore cannot leave a
false in-progress claim behind, and the tick records that cleanup as a
*failure* — no memory formed that night — rather than silently resetting. The
heartbeat is written by the **tick**, not the worker: the worker spends its
whole life blocked in `ask_brain` and could not beat if it wanted to. Past
`JOB_OVERRUN_AFTER_S` an unfinished run is reported **once**, not every 60s.

**Two attempts a night, spaced 40 min apart** (`memory.RETRY_SPACING_S`). All
three numbers used to disagree: `MAX_ATTEMPTS_PER_DAY` promised 2 while the
`consolidate` quota was 1 and its type cooldown 20h, so attempt 2 was
structurally unreachable. The spacing deliberately **clears** the 30-min global
cooldown rather than adding `consolidate` to `_GLOBAL_COOLDOWN_EXEMPT` — nobody
is waiting on a 3am retry, and each exemption is one more way for a background
job to crowd a session someone *is* waiting on. Success **or a correct skip**
marks the date done; only a failure leaves the later retry open.

**Claude spend visibility:** `token_log.log_usage()` takes a `backend` argument and splits totals under `by_backend` in `state/token_usage.json`. The top-level totals mix free Ollama with paid Claude and cannot answer "what am I spending". `call_llm()` also sets `result["backend"]` to the tier that actually served — the `backend=` reflection log line shows the *configured* primary, not the one that answered.

### Idle-Alive Daemon
Expand Down Expand Up @@ -193,7 +226,7 @@ Three-layer architecture:
Ollama. Writes to `state/thoughts.jsonl` only after a valid M5 response.
- **Layer 3 — Expression** (30min cooldown; `greet_arrival` bypasses it on a real arrival, 120s anti-flap): dispatches to tool-voice/tool-look/tool-remember and cognitive tools. Valid actions include (wait, greet, greet_arrival, comment, remember, look_at, weather_comment, scan, play_sound, photograph, emote, look_around, time_check, calendar_check, introspect, evolve, morning_fact, research, compose, self_debug, blog_essay, message_obi, set_goal, update_goal, complete_goal). Suppressed during school, quiet time, bedtime (all calendar-driven). **Hardcoded night silence: 19:00–07:00 Hobart time — no speech/audio/motion. Silent cognitive actions (`NIGHT_ALLOWED_ACTIONS`: wait, remember, research, compose, introspect, self_debug, set_goal, update_goal, complete_goal) are exempt and run overnight.**
- **`message_obi` action**: SPARK initiates a direct message to Obi via the dashboard. Exponential backoff: starts at 10min, doubles on unanswered nudge, caps at 4h, resets when Obi replies. Respects all suppressors. Thoughts with `action=message_obi` are **redacted** in `thoughts-spark.jsonl` (written as `[private message to Obi]`) so the private DM content never reaches the public `/api/v1/public/thoughts` endpoint.
- **Memory consolidation**: nightly Haiku pass (02:00–06:00 Hobart, ≤2 attempts/day, state/consolidation_meta.json) distills the last 24h of thoughts into state/memories-spark.jsonl; reflection retrieves the top-3 relevant memories by keyword/tag overlap. Goal persistence in state/intention-spark.json (7-day expiry, one active at a time).
- **Memory consolidation**: nightly Haiku pass (02:00–06:00 Hobart, ≤2 attempts/day ≥40min apart, state/consolidation_meta.json) distills the last 24h of thoughts into state/memories-spark.jsonl; reflection retrieves the top-3 relevant memories by keyword/tag overlap. **Runs on a background daemon thread with a pid-keyed job marker (`state/consolidation_job.json`), never inline on the tick** — see the health section. Goal persistence in state/intention-spark.json (7-day expiry, one active at a time).

**Critical gotchas:**
- All time-of-day logic uses `ZoneInfo("Australia/Hobart")` — never hardcoded UTC offsets
Expand Down Expand Up @@ -244,7 +277,23 @@ Watches `state/thoughts-spark.jsonl` (salience ≥0.7 or spoken action), runs Cl
| `compose` | Haiku | 4h | 2/day |
| `conversation` | Sonnet | 15min | 4/day |
| `blog` | Haiku | 30min | 5/day |
| `consolidate` | Haiku | 20h | 1/day |
| `consolidate` | Haiku | 40min | 2/day |

`consolidate` is 40min/2 rather than 20h/1 so `memory.MAX_ATTEMPTS_PER_DAY`'s
second nightly attempt can actually be spent (#291). 40min also clears the
30-min global cooldown, so the retry is *spaced past* it rather than exempted
from it.

**Deadlines are declared once, in `brain._DEADLINE_S`, and `timeout=` on
`run_claude_session` defaults to `None` so that table is what reaches
`ask_brain`.** Pass a number only when you mean to override the kind's declared
budget — an override that is *tighter* silently wins and makes the declared
value unreachable, which is exactly what an ad-hoc `timeout=180` did to
`consolidate`'s 600s: every live failure timed out at 180.1s while every
success took 30–65s. Callers still passing ad-hoc values for classified kinds
(`mind.py` self_debug 600 vs. declared 900; `bin/px-blog`, `bin/tool-research`,
`bin/tool-compose`, `bin/tool-blog` passing 300, which matches) are redundant at
best and drift at worst.

Global: 30min cooldown between sessions (except `self_debug`/`blog`), 8/day cap. When ≤2 remaining: only `self_debug`/`evolve` allowed. Bypass: `PX_CLAUDE_BUDGET_DISABLED=1`. Session log: `state/claude_sessions.jsonl`.

Expand Down
20 changes: 18 additions & 2 deletions bin/px-motd
Original file line number Diff line number Diff line change
Expand Up @@ -141,19 +141,34 @@ def _trunc(s: str, maxlen: int = 72) -> str:
# because px-motd runs as root from PAM and deliberately imports no pxh module.
MEMORY_OVERDUE_S = 48 * 3600

# Mirrors pxh.memory.JOB_HEARTBEAT_STALE_S, duplicated for the same reason and
# pinned by test_motd_job_heartbeat_threshold_matches_memory_module.
JOB_HEARTBEAT_STALE_S = 300

def _memory_formation_line(rec: dict) -> str:

def _memory_formation_line(rec: dict, job: dict | None = None) -> str:
"""Render "long-term memory" from the px-mind-consolidation health record.

Keyed on ``last_success_ts``, never ``updated_ts``: a failed nightly pass
refreshes the record without distilling anything, so the age of the last
*success* is the only number that answers "when did SPARK last form a
long-term memory". `rec` is the raw record so this stays pure and testable;
px-motd's own path clamping decides which file it comes from.

`job` is `state/consolidation_job.json`, the background worker's in-flight
marker (#291). A run in progress is appended as a hint rather than
replacing the line: "last formed 26h ago" is still the honest answer while
tonight's pass is mid-flight, and a marker whose heartbeat has gone quiet
is not evidence of anything, so it is ignored.
"""
formed = (rec or {}).get("last_success_ts")
age = _age_secs(formed)
label = f" \ud83e\udde0 {DIM}long-term memory{RESET} "
beat = _age_secs((job or {}).get("heartbeat_ts"))
if beat is not None and beat <= JOB_HEARTBEAT_STALE_S:
started = _age_secs((job or {}).get("started_ts"))
for_s = f" for {started // 60}m" if started and started >= 60 else ""
label += f"{DIM}(consolidating now{for_s}){RESET} "
if age is None:
err = _trunc(str((rec or {}).get("last_error") or ""), 44)
tail = f" {DIM}{err}{RESET}" if err else ""
Expand Down Expand Up @@ -636,7 +651,8 @@ def section_spark() -> list[str]:
# all until now, so it could fail every night while this banner showed a
# busy, healthy-looking robot.
out.append(_memory_formation_line(
_json(STATE / "health" / "px-mind-consolidation.json")))
_json(STATE / "health" / "px-mind-consolidation.json"),
_json(STATE / "consolidation_job.json")))

# Weather
weather = aw.get("weather", {}) if aw else {}
Expand Down
28 changes: 24 additions & 4 deletions src/pxh/claude_session.py
Original file line number Diff line number Diff line change
Expand Up @@ -75,7 +75,13 @@ def _model_for_type(session_type: str) -> str:
"compose": 14400, # 4 hours
"conversation": 900, # 15 min
"blog": 1800, # 30 min
"consolidate": 72000, # 20 hours
# 40 min, matching memory.RETRY_SPACING_S (#291). It was 20 hours, which
# meant the *first* attempt of a night consumed the only slot the second
# one could ever have used: `memory.MAX_ATTEMPTS_PER_DAY` promised two
# tries between 02:00 and 06:00 and this made the second structurally
# unreachable. 40 min also clears the 30-min global cooldown, so attempt 2
# is spaced past it rather than exempted from it.
"consolidate": 2400,
}

_TYPE_QUOTAS: dict[str, int] = {
Expand All @@ -85,7 +91,9 @@ def _model_for_type(session_type: str) -> str:
"compose": 2,
"conversation": 4,
"blog": 5,
"consolidate": 1,
# Two, to match memory.MAX_ATTEMPTS_PER_DAY (#291). A quota of 1 made the
# retry that module offers impossible to spend.
"consolidate": 2,
}

# Higher number = higher priority. Used for budget-tight gating.
Expand Down Expand Up @@ -325,7 +333,7 @@ def brain_kinds() -> frozenset[str]:
def _run_via_brain(
session_type: str,
prompt: str,
timeout: int,
timeout: int | None,
model: str,
) -> RunResult:
"""Serve a session from the resident Claude session instead of a subprocess.
Expand All @@ -339,6 +347,14 @@ def _run_via_brain(
resident session's envelope is fixed when it launches and cannot be
widened for one request. That is a security property, not a limitation to
work around.

`timeout=None` means "use the deadline this kind declares" —
`brain._DEADLINE_S`, which is the single source of truth for how long a
classified kind may take. A caller that passes a number overrides it, and
#291 is what that costs when the number is wrong: `consolidate` declares
600s, memory.py passed an ad-hoc 180, and because the tighter value always
wins the declared budget was unreachable — every live failure timed out at
exactly 180.1s.
"""
from . import brain # local import keeps the tmux dependency off the hot path

Expand Down Expand Up @@ -376,7 +392,7 @@ def _run_via_brain(
def run_claude_session(
session_type: str,
prompt: str,
timeout: int = 300,
timeout: int | None = None,
allowed_tools: str = "",
skip_permissions: bool = False,
cwd: str | Path | None = None,
Expand All @@ -385,6 +401,10 @@ def run_claude_session(
) -> RunResult:
"""Run a Claude session with budget checking, model routing, and logging.

timeout: seconds, or None (the default) to use the deadline the kind
declares in `brain._DEADLINE_S`. Prefer None — the declared per-kind
deadline is the one source of truth, and an ad-hoc override that is
tighter silently replaces it (see #291).
model_override: use this model instead of the session-type default.
skip_budget_check: skip rate-limit check (use for sub-phases of an already-checked session).
Raises SessionBudgetExhausted if rate-limited (unless skip_budget_check=True).
Expand Down
Loading
Loading