fix(lcm): ingested_at backfill = second full-table scan on every boot; gate one-time backfills behind lcm_migration_state + lint the init path - #902
Merged
Conversation
…every boot — gate one-time backfills behind lcm_migration_state Same class as #887, next line down in _init_db. `UPDATE messages SET ingested_at = timestamp WHERE ingested_at IS NULL` cannot use any index (the column is NULL for zero rows on a healthy DB) so it scanned 10.9 GB / 2.5 M rows on every gateway start, ~21 min, inside the engine-load lock, with all of Apollo's turns parked behind it. Surfaced 2026-09-22 20:5x, one boot after #887 deployed — Ace: "Apollo's not responding in Discord again". Landed 2026-08-06 (27b6178) with no done-marker. Fix: the codebase already has the marker table (lcm_migration_state, mark_migration_step_complete); the backfills just never used it. New is_migration_step_complete(); the ingested_at UPDATE runs once under step `messages_ingested_at_backfill_v1`, then never again. ENFORCEMENT (the part that stops #3): test_init_path_has_no_unmarked_full_table_writes walks every method _init_db calls and fails on any UPDATE/DELETE against `messages` that is neither row-bounded (WHERE store_id/session_id, LIMIT ?) nor preceded by an is_migration_step_complete() gate. Proven: injecting `UPDATE messages SET source='x' WHERE source IS NULL` into _ensure_source_column -> red, naming the method and statement. Removing the new gate -> 2 red. tests/context_engine 389 passed.
Collaborator
Author
|
🤖 merged-by: aegis · lane: breakglass · gate: BYPASS: FR paused by Ace ruling 2026-09-22; Aegis-authored; second lcm boot-scan class fix + init-path lint; 389 lcm tests, 2 mutations red; Apollo frozen again 20:5x · why: (no reason given) |
Kyzcreig
enabled auto-merge
September 23, 2026 04:14
Collaborator
Author
|
🤖 merged-by: aegis · lane: breakglass · gate: BYPASS: FR paused by Ace ruling 2026-09-22; Aegis-authored; second lcm boot-scan class fix + init-path lint; runner flake rerun green · why: (no reason given) |
This was referenced Sep 23, 2026
github-merge-queue Bot
pushed a commit
that referenced
this pull request
Sep 24, 2026
…sed replays reported lost (t_43e058b7) (#961) * fix: retain session prompt across route metadata writes (#958) Verified targeted session-state, restore, accounting, and model-resume tests: 236 passed, 1 pre-existing dashboard-auth fixture warning deselected. Reproduced NULL from billing route before fix. Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com> * fix(gateway): restart follow-ups keep adapter-granted admission; refused replays reported lost (t_43e058b7) Argus r8 N1 (t_e253d9d5): SessionSource.to_dict drops is_bot / role_authorized / delivered_via_upstream_relay / profile_route_rejected, so a spooled follow-up admitted only by ALLOW_BOTS, ALLOWED_ROLES or the relay was refused as "Unauthorized user" on boot replay, its spool file acked, and restart_followup_lost logged 0 lines. Trust model: to_dict stays wire-safe (unchanged). The spool record carries the flags in a separate `admission` block and the whole record is HMAC-SHA256'd with a per-home 0600 key (<home>/gateway/restart_followups.key). On load the flags are restored only if the MAC verifies; otherwise no trust flag is restored (only fail-closed profile_route_rejected is honoured) and PHASE=restart_followup_untrusted is logged. Live policy is still re-evaluated by the normal intake. A replay the intake refuses (unauthorized / profile_route_rejected) now logs PHASE=restart_followup_lost with reason. MF (same review): AST contract that the post-turn draining site spools pending_event itself, not None. Verified: new real stop->boot e2e (human/bot/role/relay, forged, tampered, gate-closed-during-restart, to_dict class guard) 8/8; on base 3 admission arms fail, human control passes. Focused restart suites 49/49. Mutants: MAC unchecked, refusal unreported, admission unrestored, MF pending_event=None all KILLED. Argus probe_r8_source_authz_real_intake: B/R PRESERVED, CONTROL ok. Session/authz/startup-restore suites 445 passed. * fix(kanban): judge goal deliverables before completion; fail open for operator errors (#960) * fix(kanban): grade goal deliverables before completion and isolate judge errors Verified 81 targeted tests pass (one ACP-dependent test excluded). Mutating the completion rubric makes the first-completion regression fail as expected. * refactor(kanban): one shared goal-mode handoff gate for CLI and tool surfaces Argus r1 (t_c4e23682): _goal_mode_handoff_rejection was byte-identical in tools/kanban_tools.py and hermes_cli/kanban.py; only the tool copy was test-gated, so 4 CLI mutants survived (Issue NousResearch#38367 two-copies class). - goals.kanban_handoff_rejection is now the single predicate (judge with completion_handoff=True; owned worker retries then blocks transient on judge error; operator fails open with a judge_error event; caller's conn). - Both surfaces' complete + request-review delegate to it, injecting only their run-id resolver and judge-availability probe. - CLI tests drive the real `kanban complete` / `request-review` argv path (build_parser -> kanban_command): completion_handoff reaches the judge and the card closes; real judge prompt accepts first completion; owned-worker 500 -> 2 calls, blocked transient, error on stderr, rc!=0; review gated. - AST contract: exactly one function in the tree calls the judge with completion_handoff, and both surface wrappers delegate to it. Verified: 86 passed, 1 deselected (inherited ModuleNotFoundError: acp, also red on ce0c9d3). Mutation matrix: baseline green precondition, 16/16 KILLED by named failing tests (M01-M13 re-targeted at the shared helper + W1-W9 wiring/duplicate-predicate mutants). --------- Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com> * fix(kanban): resume dependency-wait PR and page stranded ready cards (#952) * fix(kanban): respawn guard honors worker dependency_wait->promoted resume; add requeue_task + stuck-guard probe t_7d7ff489. Rule 4 active_pr no longer strands a card whose own dependency block (kind=dependency) postdates the newest PR comment and whose promotion has not yet spawned. requeue_task emits operator-intent 'requeued' for READY cards. respawn_guard_stuck_tasks lists cards held by active_pr >= 30 min. * fix(kanban): surface guarded ready cards and provide requeue verb Verify dependency_wait promotion dispatches once and subsequent crash is guarded; CLI requeue and one-shot alert tests pass (168 passed, 1 skipped in focused suites). t_7d7ff489. * test(kanban): keep corruption probe independent of watcher call count Verified: corrupt-board regression 2 passed, 22 deselected; ruff and diff check pass. Original two failures reproduced on clean fork base. * fix(kanban): keep PR continuation through status comments and unobserved ticks Verified 213 passed, 2 skipped across focused DB/CLI/watcher suites; subprocess stdin guard passed. * fix(kanban): consume event-ordered PR requeue intent Verified 183 passed, 1 skipped across DB/CLI/watcher; ruff and subprocess stdin guard pass. Same-second event-order mutant fails the regression test. * docs(kanban): state one-shot PR intent ordering * fix(kanban): bind PR comments to events and consume dependency intent Verified real dispatcher regressions RED before fix, then 210 passed, 2 skipped in focused DB/CLI/watcher suite; ruff and subprocess guard passed. * fix(kanban): pin READY requeue to PR comment identity Legacy same-second inline audit comments can share author and length; requeue snapshots the PR row id and remains one-shot. Verified 211 passed, 2 skipped; ruff clean. * fix(kanban): snapshot comment identity for every resume intent Verified focused DB/CLI/watcher/core suite: 217 passed, 2 skipped. Legacy equal-second strict mutant fails the intended arm. * fix(kanban): guard-stuck age ignores data events; board-scoped recovery command - respawn_guard_stuck_tasks: only kinds that can change the active_pr answer (_RESPAWN_GUARD_FAILURE_RESET_KINDS + dependency_wait + spawned, or a guard decline for another reason) restart the continuous-guard age; comments, heartbeats, attachments are data. - render_operator_command(board, verb, *args): single renderer, always emits --board <slug>; clear_verb uses it; watcher passes the probed board. - test_triage_resolve_records_who_and_why: expect after_comment_id == max(task_comments.id) at emit time (CI slice 15/16 red). Verified: new tests RED on 581299d (4 failed), GREEN here; recovery command executed via real CLI on default and secondary boards (rc0); no-board renderer mutant fails; 244 passed/1 skipped on db/triage/ watchers/cli/boards; ruff clean. --------- Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com> * fix(lcm): FTS parity COUNT(*) ran under _LOAD_LOCK on every engine load (freeze #3) (#966) * fix(lcm): FTS parity COUNT(*) ran under _LOAD_LOCK on every engine load (freeze #3) Card t_d3963974. Third Apollo boot-cost freeze. The cause: _fts_needs_rebuild_structural ran `SELECT COUNT(*) FROM messages` (SCAN messages USING COVERING INDEX, 2.5M rows) plus `COUNT(*) FROM messages_fts_docsize` on EVERY MessageStore/SummaryDAG construction, under plugins.context_engine._LOAD_LOCK. Measured on an APFS clone of the fleet DB: 13.4 s of a 17.9 s cold MessageStore() init. The two autocommit COUNTs could also straddle a concurrent ingest. They gave a false mismatch 15/300 times, and each mismatch triggered a full inline FTS drop+rebuild. That happened twice on 2026-09-24: held 779.8 s, 12 turns queued. Provenance: the count came in with the original vendor import 8b86963 (2026-06-16) and was carried unchanged through re-vendor 27b6178. #887/#902/#903 did not touch it. #903's plan check exempted USING COVERING INDEX and skipped messages_fts*, so it passed on this code. Fix: - The parity check moves to _fts_count_parity_mismatch. It never runs on the throttle=True (load) path. A metadata marker (fts_parity_checked_at:<fts>) throttles it to once per LCM_FTS_PARITY_CHECK_INTERVAL_HOURS (default 6 h). When due, it runs on the existing background integrity thread (own connection). A mismatch sets the /lcm doctor integrity flag instead of rebuilding inline. Both counts are read in one snapshot. - Explicit repair (throttle=False) and /lcm doctor still run parity synchronously. - A no-op load writes nothing: - _clear_integrity_failed runs only after a real repair. Before, it also erased the background corruption flag on every load. - The messages_dedup_v1 and schema_version upserts are marker-gated. - The integrity claim uses a 1 s busy timeout, not 30 s. - Lifecycle GC (on_session_start, every agent init): - no longer holds BEGIN IMMEDIATE across two SELECT DISTINCT session_id full scans; - uses indexed per-session probes; - runs at most once per 6 h per process. - _backfill_search_content no longer rewrites NULL over NULL for undecryptable rows. That rewrite fired msg_fts_update on every boot. Test: test_lcm_init_cost_regression now traces engine construction + on_session_start on the loading thread. It fails on ANY SCAN of messages, messages_fts*, summary_nodes and nodes_fts*, covering index included (only LIMIT-bounded statements are exempt). It also asserts that a steady-state load needs no write lock and preserves the corruption flag. RED on df43599 (3 failed); GREEN with this change; tests/context_engine: 393 passed. * test(lcm): lock parity-race repro and document restart meltdown --------- Co-authored-by: Apollo <apollo@angventures.io> * fix(gateway): restart follow-ups keep adapter-granted admission; refused replays reported lost (t_43e058b7) Argus r8 N1 (t_e253d9d5): SessionSource.to_dict drops is_bot / role_authorized / delivered_via_upstream_relay / profile_route_rejected, so a spooled follow-up admitted only by ALLOW_BOTS, ALLOWED_ROLES or the relay was refused as "Unauthorized user" on boot replay, its spool file acked, and restart_followup_lost logged 0 lines. Trust model: to_dict stays wire-safe (unchanged). The spool record carries the flags in a separate `admission` block and the whole record is HMAC-SHA256'd with a per-home 0600 key (<home>/gateway/restart_followups.key). On load the flags are restored only if the MAC verifies; otherwise no trust flag is restored (only fail-closed profile_route_rejected is honoured) and PHASE=restart_followup_untrusted is logged. Live policy is still re-evaluated by the normal intake. A replay the intake refuses (unauthorized / profile_route_rejected) now logs PHASE=restart_followup_lost with reason. MF (same review): AST contract that the post-turn draining site spools pending_event itself, not None. Verified: new real stop->boot e2e (human/bot/role/relay, forged, tampered, gate-closed-during-restart, to_dict class guard) 8/8; on base 3 admission arms fail, human control passes. Focused restart suites 49/49. Mutants: MAC unchecked, refusal unreported, admission unrestored, MF pending_event=None all KILLED. Argus probe_r8_source_authz_real_intake: B/R PRESERVED, CONTROL ok. Session/authz/startup-restore suites 445 passed. * fix(gateway): reject torn spool keys and gate replay refusals Verified: 57 focused restart tests passed; forged invalid-key, tampered-admission, and two refusal-site arms exercised via real restart. --------- Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com> Co-authored-by: Apollo <apollo@angventures.io>
Kyzcreig
added a commit
that referenced
this pull request
Sep 25, 2026
…25 rows (t_caabb1fa) 49/84 rows. LCM rows graded against stephenschoettler/hermes-lcm @ 8d1b1e6 (v1.0.0-rc.1) per D1a and reconciled with the lcm-upstream-divergence-ledger (DD-1 -> #168 SUPERSEDED, DD-2 -> #107 SUPERSEDED, DC-1/DC-2/DC-3 KEEP). Three generic fixes upstream still lacks by read: #412, #902, dc4245a -> UPSTREAM. Route disclosure corrected in plugins.md header (judgment on claude-fable-5-1).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
fix(lcm): the ingested_at backfill was the SECOND full-table scan on every boot — gate one-time backfills behind lcm_migration_state
Same class as #887, next line down in _init_db.
UPDATE messages SET ingested_at = timestamp WHERE ingested_at IS NULLcannot use any index (the column is NULL for zero rows on a healthy DB) so itscanned 10.9 GB / 2.5 M rows on every gateway start, ~21 min, inside the engine-load lock, with all
of Apollo's turns parked behind it. Surfaced 2026-09-22 20:5x, one boot after #887 deployed — Ace:
"Apollo's not responding in Discord again". Landed 2026-08-06 (27b6178) with no done-marker.
Fix: the codebase already has the marker table (lcm_migration_state, mark_migration_step_complete);
the backfills just never used it. New is_migration_step_complete(); the ingested_at UPDATE runs once
under step
messages_ingested_at_backfill_v1, then never again.ENFORCEMENT (the part that stops #3): test_init_path_has_no_unmarked_full_table_writes walks every
method _init_db calls and fails on any UPDATE/DELETE against
messagesthat is neither row-bounded(WHERE store_id/session_id, LIMIT ?) nor preceded by an is_migration_step_complete() gate. Proven:
injecting
UPDATE messages SET source='x' WHERE source IS NULLinto _ensure_source_column -> red,naming the method and statement. Removing the new gate -> 2 red. tests/context_engine 389 passed.
Need help on this PR? Tag
@codesmith-botwith what you need. Autofix is disabled.