Skip to content

fix(gateway): restart follow-ups keep adapter-granted admission; refused replays reported lost (t_43e058b7) - #961

Merged
Kyzcreig merged 8 commits into
mainfrom
card/t_43e058b7
Sep 24, 2026
Merged

Kyzcreig merged 8 commits into
mainfrom
card/t_43e058b7

Conversation

@Kyzcreig

@Kyzcreig Kyzcreig commented Sep 24, 2026 •

Copy link
Copy Markdown
Collaborator

Argus r8 N1 (t_e253d9d5): SessionSource.to_dict drops is_bot / role_authorized /
delivered_via_upstream_relay / profile_route_rejected, so a spooled follow-up
admitted only by ALLOW_BOTS, ALLOWED_ROLES or the relay was refused as
"Unauthorized user" on boot replay, its spool file acked, and
restart_followup_lost logged 0 lines.

Trust model: to_dict stays wire-safe (unchanged). The spool record carries the
flags in a separate admission block and the whole record is HMAC-SHA256'd
with a per-home 0600 key (/gateway/restart_followups.key). On load the
flags are restored only if the MAC verifies; otherwise no trust flag is
restored (only fail-closed profile_route_rejected is honoured) and
PHASE=restart_followup_untrusted is logged. Live policy is still re-evaluated
by the normal intake. A replay the intake refuses (unauthorized /
profile_route_rejected) now logs PHASE=restart_followup_lost with reason.

MF (same review): AST contract that the post-turn draining site spools
pending_event itself, not None.

Verified: new real stop->boot e2e (human/bot/role/relay, forged, tampered,
gate-closed-during-restart, to_dict class guard) 8/8; on base 3 admission arms
fail, human control passes. Focused restart suites 49/49. Mutants: MAC
unchecked, refusal unreported, admission unrestored, MF pending_event=None all
KILLED. Argus probe_r8_source_authz_real_intake: B/R PRESERVED, CONTROL ok.
Session/authz/startup-restore suites 445 passed.

Card: t_43e058b7 (from Argus t_e253d9d5 r8 N1 + MF). Live reachability today is 0 (no profile enables ALLOW_BOTS with a token, ALLOWED_ROLES, or the relay).


View with [code]smith Autofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.

Kyzcreig and others added 8 commits September 24, 2026 10:28
Verified targeted session-state, restore, accounting, and model-resume tests: 236 passed, 1 pre-existing dashboard-auth fixture warning deselected. Reproduced NULL from billing route before fix.

Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
…sed replays reported lost (t_43e058b7)

Argus r8 N1 (t_e253d9d5): SessionSource.to_dict drops is_bot / role_authorized /
delivered_via_upstream_relay / profile_route_rejected, so a spooled follow-up
admitted only by ALLOW_BOTS, ALLOWED_ROLES or the relay was refused as
"Unauthorized user" on boot replay, its spool file acked, and
restart_followup_lost logged 0 lines.

Trust model: to_dict stays wire-safe (unchanged). The spool record carries the
flags in a separate `admission` block and the whole record is HMAC-SHA256'd
with a per-home 0600 key (<home>/gateway/restart_followups.key). On load the
flags are restored only if the MAC verifies; otherwise no trust flag is
restored (only fail-closed profile_route_rejected is honoured) and
PHASE=restart_followup_untrusted is logged. Live policy is still re-evaluated
by the normal intake. A replay the intake refuses (unauthorized /
profile_route_rejected) now logs PHASE=restart_followup_lost with reason.

MF (same review): AST contract that the post-turn draining site spools
pending_event itself, not None.

Verified: new real stop->boot e2e (human/bot/role/relay, forged, tampered,
gate-closed-during-restart, to_dict class guard) 8/8; on base 3 admission arms
fail, human control passes. Focused restart suites 49/49. Mutants: MAC
unchecked, refusal unreported, admission unrestored, MF pending_event=None all
KILLED. Argus probe_r8_source_authz_real_intake: B/R PRESERVED, CONTROL ok.
Session/authz/startup-restore suites 445 passed.
… operator errors (#960)

* fix(kanban): grade goal deliverables before completion and isolate judge errors

Verified 81 targeted tests pass (one ACP-dependent test excluded). Mutating the completion rubric makes the first-completion regression fail as expected.

* refactor(kanban): one shared goal-mode handoff gate for CLI and tool surfaces

Argus r1 (t_c4e23682): _goal_mode_handoff_rejection was byte-identical in
tools/kanban_tools.py and hermes_cli/kanban.py; only the tool copy was
test-gated, so 4 CLI mutants survived (Issue NousResearch#38367 two-copies class).

- goals.kanban_handoff_rejection is now the single predicate (judge with
  completion_handoff=True; owned worker retries then blocks transient on
  judge error; operator fails open with a judge_error event; caller's conn).
- Both surfaces' complete + request-review delegate to it, injecting only
  their run-id resolver and judge-availability probe.
- CLI tests drive the real `kanban complete` / `request-review` argv path
  (build_parser -> kanban_command): completion_handoff reaches the judge and
  the card closes; real judge prompt accepts first completion; owned-worker
  500 -> 2 calls, blocked transient, error on stderr, rc!=0; review gated.
- AST contract: exactly one function in the tree calls the judge with
  completion_handoff, and both surface wrappers delegate to it.

Verified: 86 passed, 1 deselected (inherited ModuleNotFoundError: acp, also
red on ce0c9d3). Mutation matrix: baseline green precondition, 16/16 KILLED
by named failing tests (M01-M13 re-targeted at the shared helper + W1-W9
wiring/duplicate-predicate mutants).

---------

Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
…952)

* fix(kanban): respawn guard honors worker dependency_wait->promoted resume; add requeue_task + stuck-guard probe

t_7d7ff489. Rule 4 active_pr no longer strands a card whose own dependency block (kind=dependency) postdates the newest PR comment and whose promotion has not yet spawned. requeue_task emits operator-intent 'requeued' for READY cards. respawn_guard_stuck_tasks lists cards held by active_pr >= 30 min.

* fix(kanban): surface guarded ready cards and provide requeue verb

Verify dependency_wait promotion dispatches once and subsequent crash is guarded; CLI requeue and one-shot alert tests pass (168 passed, 1 skipped in focused suites). t_7d7ff489.

* test(kanban): keep corruption probe independent of watcher call count

Verified: corrupt-board regression 2 passed, 22 deselected; ruff and diff check pass. Original two failures reproduced on clean fork base.

* fix(kanban): keep PR continuation through status comments and unobserved ticks

Verified 213 passed, 2 skipped across focused DB/CLI/watcher suites; subprocess stdin guard passed.

* fix(kanban): consume event-ordered PR requeue intent

Verified 183 passed, 1 skipped across DB/CLI/watcher; ruff and subprocess stdin guard pass. Same-second event-order mutant fails the regression test.

* docs(kanban): state one-shot PR intent ordering

* fix(kanban): bind PR comments to events and consume dependency intent

Verified real dispatcher regressions RED before fix, then 210 passed, 2 skipped in focused DB/CLI/watcher suite; ruff and subprocess guard passed.

* fix(kanban): pin READY requeue to PR comment identity

Legacy same-second inline audit comments can share author and length; requeue snapshots the PR row id and remains one-shot. Verified 211 passed, 2 skipped; ruff clean.

* fix(kanban): snapshot comment identity for every resume intent

Verified focused DB/CLI/watcher/core suite: 217 passed, 2 skipped. Legacy equal-second strict mutant fails the intended arm.

* fix(kanban): guard-stuck age ignores data events; board-scoped recovery command

- respawn_guard_stuck_tasks: only kinds that can change the active_pr answer
  (_RESPAWN_GUARD_FAILURE_RESET_KINDS + dependency_wait + spawned, or a guard
  decline for another reason) restart the continuous-guard age; comments,
  heartbeats, attachments are data.
- render_operator_command(board, verb, *args): single renderer, always emits
  --board <slug>; clear_verb uses it; watcher passes the probed board.
- test_triage_resolve_records_who_and_why: expect after_comment_id ==
  max(task_comments.id) at emit time (CI slice 15/16 red).

Verified: new tests RED on 581299d (4 failed), GREEN here; recovery
command executed via real CLI on default and secondary boards (rc0);
no-board renderer mutant fails; 244 passed/1 skipped on db/triage/
watchers/cli/boards; ruff clean.

---------

Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
…ad (freeze #3) (#966)

* fix(lcm): FTS parity COUNT(*) ran under _LOAD_LOCK on every engine load (freeze #3)

Card t_d3963974. Third Apollo boot-cost freeze. The cause:
_fts_needs_rebuild_structural ran `SELECT COUNT(*) FROM messages`
(SCAN messages USING COVERING INDEX, 2.5M rows) plus
`COUNT(*) FROM messages_fts_docsize` on EVERY MessageStore/SummaryDAG
construction, under plugins.context_engine._LOAD_LOCK. Measured on an APFS
clone of the fleet DB: 13.4 s of a 17.9 s cold MessageStore() init. The
two autocommit COUNTs could also straddle a concurrent ingest. They gave a
false mismatch 15/300 times, and each mismatch triggered a full inline
FTS drop+rebuild. That happened twice on 2026-09-24: held 779.8 s, 12
turns queued.

Provenance: the count came in with the original vendor import 8b86963
(2026-06-16) and was carried unchanged through re-vendor 27b6178.
#887/#902/#903 did not touch it. #903's plan check exempted
USING COVERING INDEX and skipped messages_fts*, so it passed on this code.

Fix:
- The parity check moves to _fts_count_parity_mismatch. It never runs on
  the throttle=True (load) path. A metadata marker
  (fts_parity_checked_at:<fts>) throttles it to once per
  LCM_FTS_PARITY_CHECK_INTERVAL_HOURS (default 6 h). When due, it runs
  on the existing background integrity thread (own connection). A
  mismatch sets the /lcm doctor integrity flag instead of rebuilding
  inline. Both counts are read in one snapshot.
- Explicit repair (throttle=False) and /lcm doctor still run parity
  synchronously.
- A no-op load writes nothing:
  - _clear_integrity_failed runs only after a real repair. Before, it
    also erased the background corruption flag on every load.
  - The messages_dedup_v1 and schema_version upserts are marker-gated.
  - The integrity claim uses a 1 s busy timeout, not 30 s.
- Lifecycle GC (on_session_start, every agent init):
  - no longer holds BEGIN IMMEDIATE across two SELECT DISTINCT
    session_id full scans;
  - uses indexed per-session probes;
  - runs at most once per 6 h per process.
- _backfill_search_content no longer rewrites NULL over NULL for
  undecryptable rows. That rewrite fired msg_fts_update on every boot.

Test: test_lcm_init_cost_regression now traces engine construction +
on_session_start on the loading thread. It fails on ANY SCAN of messages,
messages_fts*, summary_nodes and nodes_fts*, covering index included
(only LIMIT-bounded statements are exempt). It also asserts that a
steady-state load needs no write lock and preserves the corruption flag.
RED on df43599 (3 failed); GREEN with this change;
tests/context_engine: 393 passed.

* test(lcm): lock parity-race repro and document restart meltdown

---------

Co-authored-by: Apollo <apollo@angventures.io>
…sed replays reported lost (t_43e058b7)

Argus r8 N1 (t_e253d9d5): SessionSource.to_dict drops is_bot / role_authorized /
delivered_via_upstream_relay / profile_route_rejected, so a spooled follow-up
admitted only by ALLOW_BOTS, ALLOWED_ROLES or the relay was refused as
"Unauthorized user" on boot replay, its spool file acked, and
restart_followup_lost logged 0 lines.

Trust model: to_dict stays wire-safe (unchanged). The spool record carries the
flags in a separate `admission` block and the whole record is HMAC-SHA256'd
with a per-home 0600 key (<home>/gateway/restart_followups.key). On load the
flags are restored only if the MAC verifies; otherwise no trust flag is
restored (only fail-closed profile_route_rejected is honoured) and
PHASE=restart_followup_untrusted is logged. Live policy is still re-evaluated
by the normal intake. A replay the intake refuses (unauthorized /
profile_route_rejected) now logs PHASE=restart_followup_lost with reason.

MF (same review): AST contract that the post-turn draining site spools
pending_event itself, not None.

Verified: new real stop->boot e2e (human/bot/role/relay, forged, tampered,
gate-closed-during-restart, to_dict class guard) 8/8; on base 3 admission arms
fail, human control passes. Focused restart suites 49/49. Mutants: MAC
unchecked, refusal unreported, admission unrestored, MF pending_event=None all
KILLED. Argus probe_r8_source_authz_real_intake: B/R PRESERVED, CONTROL ok.
Session/authz/startup-restore suites 445 passed.
Verified: 57 focused restart tests passed; forged invalid-key, tampered-admission, and two refusal-site arms exercised via real restart.
@Kyzcreig

Copy link
Copy Markdown
Collaborator Author

🤖 merged-by: apollo · lane: discord · gate: BYPASS: FR paused by Ace ruling 2026-09-22; gate = kanban Argus PASS r2 t_43e058b7 run 8841 · why: restart spool: follow-ups from bot/role-admitted senders preserved across restart (adapter-granted admission kept; refused replays); spool key validation + atomic publish; Argus r2 PASS-WITH-CAVEATS @f323654f run 8841 (t_43e058b7)

@Kyzcreig
Kyzcreig added this pull request to the merge queue Sep 24, 2026
Merged via the queue into main with commit 60d226b Sep 24, 2026
55 checks passed
@Kyzcreig
Kyzcreig deleted the card/t_43e058b7 branch September 24, 2026 17:24
@Kyzcreig Kyzcreig added the fleetreview:post-merge Ask FleetReview to review this MERGED pull (merge commit vs first parent) label Sep 24, 2026
@ang-prism

ang-prism Bot commented Sep 27, 2026

Copy link
Copy Markdown

FleetReview

Post-merge review (fleetreview:post-merge override): this reviewed the merge commit against its first parent — the bytes that already shipped. It is not a pre-merge gate pass.

Confidence: 3/5

Findings

  • P1 gateway/fork_ext/restart_followups.py:151 — Key replacement race
  • P2 tests/gateway/test_restart_followups_admission_e2e.py:136 — Slow refusals
  • P1 gateway/fork_ext/restart_followups.py:76 — Restored role_authorized / delivered_via_upstream_relay skip the live policy check, so a gate closed during the restart does not refuse the replay
  • P1 gateway/run.py:15951 — Replayed role grants bypass revoked Discord role policy
  • P1 tests/gateway/test_restart_followups_admission_e2e.py:148 — Role replay is authorized even after the role allowlist is removed

FleetReview provenance · models: D=grok-4.6 · cost: $0.00 · duration: 33m 02s · rounds: 1 · files examined: 5

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

fleetreview:post-merge Ask FleetReview to review this MERGED pull (merge commit vs first parent)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant