fix(kanban): judge goal deliverables before completion; fail open for operator errors - #960
Merged
Merged
Conversation
…dge errors Verified 81 targeted tests pass (one ACP-dependent test excluded). Mutating the completion rubric makes the first-completion regression fail as expected.
…surfaces Argus r1 (t_c4e23682): _goal_mode_handoff_rejection was byte-identical in tools/kanban_tools.py and hermes_cli/kanban.py; only the tool copy was test-gated, so 4 CLI mutants survived (Issue NousResearch#38367 two-copies class). - goals.kanban_handoff_rejection is now the single predicate (judge with completion_handoff=True; owned worker retries then blocks transient on judge error; operator fails open with a judge_error event; caller's conn). - Both surfaces' complete + request-review delegate to it, injecting only their run-id resolver and judge-availability probe. - CLI tests drive the real `kanban complete` / `request-review` argv path (build_parser -> kanban_command): completion_handoff reaches the judge and the card closes; real judge prompt accepts first completion; owned-worker 500 -> 2 calls, blocked transient, error on stderr, rc!=0; review gated. - AST contract: exactly one function in the tree calls the judge with completion_handoff, and both surface wrappers delegate to it. Verified: 86 passed, 1 deselected (inherited ModuleNotFoundError: acp, also red on ce0c9d3). Mutation matrix: baseline green precondition, 16/16 KILLED by named failing tests (M01-M13 re-targeted at the shared helper + W1-W9 wiring/duplicate-predicate mutants).
Collaborator
Author
|
🤖 merged-by: apollo · lane: discord · gate: BYPASS: FR paused by Ace ruling 2026-09-22; gate = kanban Argus PASS run 8602 · why: goal-mode judge: judge proposed deliverables, not a prior completion receipt (t_c4e23682): Argus PASS-WITH-CAVEATS r2 run 8602 at ddf8108 |
github-merge-queue Bot
pushed a commit
that referenced
this pull request
Sep 24, 2026
…sed replays reported lost (t_43e058b7) (#961) * fix: retain session prompt across route metadata writes (#958) Verified targeted session-state, restore, accounting, and model-resume tests: 236 passed, 1 pre-existing dashboard-auth fixture warning deselected. Reproduced NULL from billing route before fix. Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com> * fix(gateway): restart follow-ups keep adapter-granted admission; refused replays reported lost (t_43e058b7) Argus r8 N1 (t_e253d9d5): SessionSource.to_dict drops is_bot / role_authorized / delivered_via_upstream_relay / profile_route_rejected, so a spooled follow-up admitted only by ALLOW_BOTS, ALLOWED_ROLES or the relay was refused as "Unauthorized user" on boot replay, its spool file acked, and restart_followup_lost logged 0 lines. Trust model: to_dict stays wire-safe (unchanged). The spool record carries the flags in a separate `admission` block and the whole record is HMAC-SHA256'd with a per-home 0600 key (<home>/gateway/restart_followups.key). On load the flags are restored only if the MAC verifies; otherwise no trust flag is restored (only fail-closed profile_route_rejected is honoured) and PHASE=restart_followup_untrusted is logged. Live policy is still re-evaluated by the normal intake. A replay the intake refuses (unauthorized / profile_route_rejected) now logs PHASE=restart_followup_lost with reason. MF (same review): AST contract that the post-turn draining site spools pending_event itself, not None. Verified: new real stop->boot e2e (human/bot/role/relay, forged, tampered, gate-closed-during-restart, to_dict class guard) 8/8; on base 3 admission arms fail, human control passes. Focused restart suites 49/49. Mutants: MAC unchecked, refusal unreported, admission unrestored, MF pending_event=None all KILLED. Argus probe_r8_source_authz_real_intake: B/R PRESERVED, CONTROL ok. Session/authz/startup-restore suites 445 passed. * fix(kanban): judge goal deliverables before completion; fail open for operator errors (#960) * fix(kanban): grade goal deliverables before completion and isolate judge errors Verified 81 targeted tests pass (one ACP-dependent test excluded). Mutating the completion rubric makes the first-completion regression fail as expected. * refactor(kanban): one shared goal-mode handoff gate for CLI and tool surfaces Argus r1 (t_c4e23682): _goal_mode_handoff_rejection was byte-identical in tools/kanban_tools.py and hermes_cli/kanban.py; only the tool copy was test-gated, so 4 CLI mutants survived (Issue NousResearch#38367 two-copies class). - goals.kanban_handoff_rejection is now the single predicate (judge with completion_handoff=True; owned worker retries then blocks transient on judge error; operator fails open with a judge_error event; caller's conn). - Both surfaces' complete + request-review delegate to it, injecting only their run-id resolver and judge-availability probe. - CLI tests drive the real `kanban complete` / `request-review` argv path (build_parser -> kanban_command): completion_handoff reaches the judge and the card closes; real judge prompt accepts first completion; owned-worker 500 -> 2 calls, blocked transient, error on stderr, rc!=0; review gated. - AST contract: exactly one function in the tree calls the judge with completion_handoff, and both surface wrappers delegate to it. Verified: 86 passed, 1 deselected (inherited ModuleNotFoundError: acp, also red on ce0c9d3). Mutation matrix: baseline green precondition, 16/16 KILLED by named failing tests (M01-M13 re-targeted at the shared helper + W1-W9 wiring/duplicate-predicate mutants). --------- Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com> * fix(kanban): resume dependency-wait PR and page stranded ready cards (#952) * fix(kanban): respawn guard honors worker dependency_wait->promoted resume; add requeue_task + stuck-guard probe t_7d7ff489. Rule 4 active_pr no longer strands a card whose own dependency block (kind=dependency) postdates the newest PR comment and whose promotion has not yet spawned. requeue_task emits operator-intent 'requeued' for READY cards. respawn_guard_stuck_tasks lists cards held by active_pr >= 30 min. * fix(kanban): surface guarded ready cards and provide requeue verb Verify dependency_wait promotion dispatches once and subsequent crash is guarded; CLI requeue and one-shot alert tests pass (168 passed, 1 skipped in focused suites). t_7d7ff489. * test(kanban): keep corruption probe independent of watcher call count Verified: corrupt-board regression 2 passed, 22 deselected; ruff and diff check pass. Original two failures reproduced on clean fork base. * fix(kanban): keep PR continuation through status comments and unobserved ticks Verified 213 passed, 2 skipped across focused DB/CLI/watcher suites; subprocess stdin guard passed. * fix(kanban): consume event-ordered PR requeue intent Verified 183 passed, 1 skipped across DB/CLI/watcher; ruff and subprocess stdin guard pass. Same-second event-order mutant fails the regression test. * docs(kanban): state one-shot PR intent ordering * fix(kanban): bind PR comments to events and consume dependency intent Verified real dispatcher regressions RED before fix, then 210 passed, 2 skipped in focused DB/CLI/watcher suite; ruff and subprocess guard passed. * fix(kanban): pin READY requeue to PR comment identity Legacy same-second inline audit comments can share author and length; requeue snapshots the PR row id and remains one-shot. Verified 211 passed, 2 skipped; ruff clean. * fix(kanban): snapshot comment identity for every resume intent Verified focused DB/CLI/watcher/core suite: 217 passed, 2 skipped. Legacy equal-second strict mutant fails the intended arm. * fix(kanban): guard-stuck age ignores data events; board-scoped recovery command - respawn_guard_stuck_tasks: only kinds that can change the active_pr answer (_RESPAWN_GUARD_FAILURE_RESET_KINDS + dependency_wait + spawned, or a guard decline for another reason) restart the continuous-guard age; comments, heartbeats, attachments are data. - render_operator_command(board, verb, *args): single renderer, always emits --board <slug>; clear_verb uses it; watcher passes the probed board. - test_triage_resolve_records_who_and_why: expect after_comment_id == max(task_comments.id) at emit time (CI slice 15/16 red). Verified: new tests RED on 581299d (4 failed), GREEN here; recovery command executed via real CLI on default and secondary boards (rc0); no-board renderer mutant fails; 244 passed/1 skipped on db/triage/ watchers/cli/boards; ruff clean. --------- Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com> * fix(lcm): FTS parity COUNT(*) ran under _LOAD_LOCK on every engine load (freeze #3) (#966) * fix(lcm): FTS parity COUNT(*) ran under _LOAD_LOCK on every engine load (freeze #3) Card t_d3963974. Third Apollo boot-cost freeze. The cause: _fts_needs_rebuild_structural ran `SELECT COUNT(*) FROM messages` (SCAN messages USING COVERING INDEX, 2.5M rows) plus `COUNT(*) FROM messages_fts_docsize` on EVERY MessageStore/SummaryDAG construction, under plugins.context_engine._LOAD_LOCK. Measured on an APFS clone of the fleet DB: 13.4 s of a 17.9 s cold MessageStore() init. The two autocommit COUNTs could also straddle a concurrent ingest. They gave a false mismatch 15/300 times, and each mismatch triggered a full inline FTS drop+rebuild. That happened twice on 2026-09-24: held 779.8 s, 12 turns queued. Provenance: the count came in with the original vendor import 8b86963 (2026-06-16) and was carried unchanged through re-vendor 27b6178. #887/#902/#903 did not touch it. #903's plan check exempted USING COVERING INDEX and skipped messages_fts*, so it passed on this code. Fix: - The parity check moves to _fts_count_parity_mismatch. It never runs on the throttle=True (load) path. A metadata marker (fts_parity_checked_at:<fts>) throttles it to once per LCM_FTS_PARITY_CHECK_INTERVAL_HOURS (default 6 h). When due, it runs on the existing background integrity thread (own connection). A mismatch sets the /lcm doctor integrity flag instead of rebuilding inline. Both counts are read in one snapshot. - Explicit repair (throttle=False) and /lcm doctor still run parity synchronously. - A no-op load writes nothing: - _clear_integrity_failed runs only after a real repair. Before, it also erased the background corruption flag on every load. - The messages_dedup_v1 and schema_version upserts are marker-gated. - The integrity claim uses a 1 s busy timeout, not 30 s. - Lifecycle GC (on_session_start, every agent init): - no longer holds BEGIN IMMEDIATE across two SELECT DISTINCT session_id full scans; - uses indexed per-session probes; - runs at most once per 6 h per process. - _backfill_search_content no longer rewrites NULL over NULL for undecryptable rows. That rewrite fired msg_fts_update on every boot. Test: test_lcm_init_cost_regression now traces engine construction + on_session_start on the loading thread. It fails on ANY SCAN of messages, messages_fts*, summary_nodes and nodes_fts*, covering index included (only LIMIT-bounded statements are exempt). It also asserts that a steady-state load needs no write lock and preserves the corruption flag. RED on df43599 (3 failed); GREEN with this change; tests/context_engine: 393 passed. * test(lcm): lock parity-race repro and document restart meltdown --------- Co-authored-by: Apollo <apollo@angventures.io> * fix(gateway): restart follow-ups keep adapter-granted admission; refused replays reported lost (t_43e058b7) Argus r8 N1 (t_e253d9d5): SessionSource.to_dict drops is_bot / role_authorized / delivered_via_upstream_relay / profile_route_rejected, so a spooled follow-up admitted only by ALLOW_BOTS, ALLOWED_ROLES or the relay was refused as "Unauthorized user" on boot replay, its spool file acked, and restart_followup_lost logged 0 lines. Trust model: to_dict stays wire-safe (unchanged). The spool record carries the flags in a separate `admission` block and the whole record is HMAC-SHA256'd with a per-home 0600 key (<home>/gateway/restart_followups.key). On load the flags are restored only if the MAC verifies; otherwise no trust flag is restored (only fail-closed profile_route_rejected is honoured) and PHASE=restart_followup_untrusted is logged. Live policy is still re-evaluated by the normal intake. A replay the intake refuses (unauthorized / profile_route_rejected) now logs PHASE=restart_followup_lost with reason. MF (same review): AST contract that the post-turn draining site spools pending_event itself, not None. Verified: new real stop->boot e2e (human/bot/role/relay, forged, tampered, gate-closed-during-restart, to_dict class guard) 8/8; on base 3 admission arms fail, human control passes. Focused restart suites 49/49. Mutants: MAC unchecked, refusal unreported, admission unrestored, MF pending_event=None all KILLED. Argus probe_r8_source_authz_real_intake: B/R PRESERVED, CONTROL ok. Session/authz/startup-restore suites 445 passed. * fix(gateway): reject torn spool keys and gate replay refusals Verified: 57 focused restart tests passed; forged invalid-key, tampered-admission, and two refusal-site arms exercised via real restart. --------- Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com> Co-authored-by: Apollo <apollo@angventures.io>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Goal-mode completion now grades the proposed deliverables without requiring a prior completion receipt. Worker judge errors retry once then block transient; operator completions fail open with a judge_error event. Both tool and CLI handoffs covered. Verification: 81 targeted tests passed; one unrelated ACP-dependent test excluded (acp not installed). Mutant disabling the completion rubric makes the first-completion regression fail.
Need help on this PR? Tag
@codesmith-botwith what you need. Autofix is disabled.