Skip to content

fix: speed up session startup and escalate home-summary failures - #51

Merged
MrGTV-love merged 13 commits into
mainfrom
fm/fm-session-start-fast
Oct 8, 2026
Merged

MrGTV-love merged 13 commits into
mainfrom
fm/fm-session-start-fast

Conversation

@MrGTV-love

@MrGTV-love MrGTV-love commented Oct 8, 2026 •

Copy link
Copy Markdown
Owner

Intent

we also need to fix an issue with the session start hooks. you increased the time to 180s, but even 120s was too long to get a session started. here is an analysis. HOME_SUMMARY failed 295x and never did anythign about it. another exaple of why ignoring and working around problems instead of solving them is the wrong thing to do.

What Changed

  • Detach home-summary refreshes from session start, spawn, and teardown, with single-flight execution and pending-trigger follow-ups.
  • Escalate three consecutive owned refresh failures through a diagnostic wake, reset the streak after successful publication, and fence publication and failure accounting against stale lock owners.
  • Reduce startup and fleet-scan overhead with bounded bulk backlog-state reads, validated open-decision checkpoint reuse in readers and snapshot copies, and shell-native parsing and path handling.

Risk Assessment

⚠️ Medium: The change modifies concurrency-sensitive refresh ownership and persisted checkpoint reuse, but the complete source review found no substantiated merge-blocking defect, unauthorized scope expansion, or intent contradiction.

Testing

Live validation passed across startup latency, checkpoint migration, detached publication, lifecycle triggers, failure escalation, recovery, and concurrency. Focused regression runs exposed two fixture assumptions; the corrected scenarios, ownership checks, and checkpoint regressions passed. Evidence consists of real CLI transcripts and persisted ledgers; this change has no visual-layout surface requiring screenshots.

  • Live validation: ✅ go - 13 of 13 scenarios driven live against the product
Scenario Result Live Evidence
Start an upgraded real Claude primary with an obsolete checkpoint and long status history ✅ pass live Real Claude upgraded-startup digest; Upgraded-startup timing
Re-emit startup context while retaining a buried unresolved decision ✅ pass live Real Claude context re-emission digest; claude-live-warm-time.txt
Open a real Claude primary without waiting for stalled summary validation ✅ pass live Startup completed while detached publication was blocked; Blocked-summary startup timing; Detached ledger published after releasing validation
Reject a directory ledger destination without false publication ✅ pass live Live refresh outcomes, streaks, wakes, and published ledgers
Escalate three real acquired failures once without repeated wake spam ✅ pass live Live refresh outcomes, streaks, wakes, and published ledgers
Recover with a valid ledger and escalate a subsequent fresh failure streak ✅ pass live Live refresh outcomes, streaks, wakes, and published ledgers
Escalate and release refresh ownership even when diagnostic logging blocks ✅ pass live Live refresh outcomes, streaks, wakes, and published ledgers
Return detached triggers immediately and drain more than three pending follow-ups ✅ pass live Live refresh outcomes, streaks, wakes, and published ledgers
Cancel a refresh parent and prevent its orphan from overwriting newer state ✅ pass live Live refresh outcomes, streaks, wakes, and published ledgers
Spawn and tear down a real worker without waiting for blocked summary publication ✅ pass live Real isolated Herdr worker lifecycle
Reconcile quoted-looking task IDs in a comma-containing home and release lifecycle locks ✅ pass live Batched backlog reconciliation and concurrent lock acquisition
Complete a status line after a partial checkpoint and preserve earlier decisions in publication ✅ pass live Partial-checkpoint drain and publication behavior
Bound stalled refreshes, preserve the last good ledger, and report a changed failure reason ✅ pass live Live refresh outcomes, streaks, wakes, and published ledgers
Evidence: Live scenario manifest and validation setup

Source: Live scenario manifest and validation setup

{
  "runtime": "real Firstmate CLI scripts; real Claude 2.1.294 primary; isolated Herdr 0.9.1 session; real tasks-axi 0.2.6",
  "primary_transport": "tmux private fm-lab socket, 120 columns x 40 rows; explicit lab-bound command because this Claude launcher did not forward FM_HOME to its agent shell",
  "fault_injection": "Real jq execution delayed at validation barriers; real FIFO diagnostic log. No Firstmate script or summary document was substituted.",
  "scenarios": [
    {
      "name": "Reject a directory ledger destination without false publication",
      "result": "pass",
      "live": true,
      "evidence": "live-refresh-results.json",
      "reason": "fm-home-summary-refresh: atomic ledger replacement failed: destination is a directory: ~/.no-mistakes/worktrees/32d18ed9638d/01M4CVJREG2ZWYYTNDVWK2DHWT/.validation-tmp/live/escalation/state/home-summary.json"
    },
    {
      "name": "Escalate three real acquired failures once, without repeated wake spam",
      "result": "pass",
      "live": true,
      "evidence": "live-refresh-results.json",
      "reason": "[{\"attempt\": 1, \"exit\": 0, \"elapsed_seconds\": 0.73, \"streak\": \"count=1\\nfirst=2026-10-08T14:26:57Z\\nclass=atomic ledger replacement failed: destination is a directory: ~\\nescalated=\\n\", \"wake_count\": 0}, {\"attempt\": 2, \"exit\": 0, \"elapsed_seconds\": 0.729, \"streak\": \"count=2\\nfirst=2026-10-08T14:26:57Z\\nclass=atomic ledger replacement failed: destination is a directory: ~\\nescalated=\\n\", \"wake_count\": 0}, {\"attempt\": 3, \"exit\": 0, \"elapsed_seconds\": 0.947, \"streak\": \"count=3\\nfirst=2026-10-08T14:26:57Z\\nclass=atomic ledger replacement failed: destination is a directory: ~\\nescalated=atomic ledger replacement failed: destination is a directory: ~\\n\", \"wake_count\": 1}, {\"attempt\": 4, \"exit\": 0, \"elapsed_seconds\": 1.255, \"streak\": \"count=4\\nfirst=2026-10-08T14:26:57Z\\nclass=atomic ledger replacement failed: destination is a directory: ~\\nescalated=atomic ledger replacement failed: destination is a directory: ~\\n\", \"wake_count\": 1}]"
    },
    {
      "name": "Recover with a valid ledger and escalate a subsequent fresh failure streak",
      "result": "pass",
      "live": true,
      "evidence": "live-refresh-results.json",
      "reason": "Success in 1.74s; schema=fm-secondmate-home-summary.v1; streak removed=True; subsequent wake count=2"
    },
    {
      "name": "Escalate and release refresh ownership even when diagnostic logging blocks",
      "result": "pass",
      "live": true,
      "evidence": "live-refresh-results.json",
      "reason": "[{\"attempt\": 1, \"elapsed_seconds\": 5.396, \"streak\": \"count=1\\nfirst=2026-10-08T14:27:07Z\\nclass=atomic ledger replacement failed: destination is a directory: ~\\nescalated=\\n\", \"wake_count\": 0, \"lock_remaining\": false}, {\"attempt\": 2, \"elapsed_seconds\": 5.272, \"streak\": \"count=2\\nfirst=2026-10-08T14:27:07Z\\nclass=atomic ledger replacement failed: destination is a directory: ~\\nescalated=\\n\", \"wake_count\": 0, \"lock_remaining\": false}, {\"attempt\": 3, \"elapsed_seconds\": 5.593, \"streak\": \"count=3\\nfirst=2026-10-08T14:27:07Z\\nclass=atomic ledger replacement failed: destination is a directory: ~\\nescalated=atomic ledger replacement failed: destination is a directory: ~\\n\", \"wake_count\": 1, \"lock_remaining\": false}]"
    },
    {
      "name": "Return detached triggers immediately and drain more than three pending follow-ups",
      "result": "pass",
      "live": true,
      "evidence": "live-refresh-results.json",
      "reason": "Initial detach=0.05s; queued=['trigger-1', 'trigger-2', 'trigger-3', 'trigger-4', 'trigger-5']; followups=[{'round': 1, 'contender_return_seconds': 0.018, 'refresh_owner_pid': '53359', 'validation_process': '65057 63100'}, {'round': 2, 'contender_return_seconds': 0.033, 'refresh_owner_pid': '53359', 'validation_process': '69647 67902'}, {'round': 3, 'contender_return_seconds': 0.019, 'refresh_owner_pid': '53359', 'validation_process': '78190 71918'}, {'round': 4, 'contender_return_seconds': 0.029, 'refresh_owner_pid': '53359', 'validation_process': '85326 82760'}, {'round': 5, 'contender_return_seconds': 0.023, 'refresh_owner_pid': '53359', 'validation_process': '90211 88240'}]"
    },
    {
      "name": "Cancel a refresh parent and prevent its orphan from overwriting newer state",
      "result": "pass",
      "live": true,
      "evidence": "live-refresh-results.json",
      "reason": "Dead owner=50222; newer refresh exit=0, duration=1.081s; orphan did not publish into directory or erase streak"
    },
    {
      "name": "Start an upgraded real Claude primary with a version-9 checkpoint and 24,000 routine status lines",
      "result": "pass",
      "live": true,
      "evidence": "live-refresh-results.json",
      "reason": "elapsed_seconds=10.494; status_bytes=936060; buried-live retained; obsolete phantom rejected; NEXT STEP reached"
    },
    {
      "name": "Re-emit startup context without losing a buried unresolved decision",
      "result": "pass",
      "live": true,
      "evidence": "live-refresh-results.json",
      "reason": "elapsed_seconds=8.281; buried-live retained; NEXT STEP reached"
    },
    {
      "name": "Spawn and tear down a real worker without waiting for blocked summary publication",
      "result": "pass",
      "live": true,
      "evidence": "live-refresh-results.json",
      "reason": "Spawn=13.846s; teardown=21.565s with summary validation blocked; final active_children=[]; pool inside worktree"
    },
    {
      "name": "Reconcile quoted-looking task IDs in a comma-containing home and release lifecycle locks",
      "result": "pass",
      "live": true,
      "evidence": "live-refresh-results.json",
      "reason": "bootstrap-batch-transcript.txt: numeric/boolean IDs avoided per-record reads; queued owned row healed; both worker locks acquired before bootstrap exited"
    },
    {
      "name": "Complete a status line after a partial checkpoint and preserve pre-checkpoint decisions in publication",
      "result": "pass",
      "live": true,
      "evidence": "live-refresh-results.json",
      "reason": "Completed split decision surfaced in the real drain; after resolving it, published keys=['late', 'still-open']; cold partial checkpoint offset=offset=85"
    },
    {
      "name": "Bound stalled refreshes, preserve the last good ledger, and report a changed failure reason",
      "result": "pass",
      "live": true,
      "evidence": "live-refresh-results.json",
      "reason": "Three acquired validation timeouts retained identical ledger bytes and released ownership; wake on third only; directory-destination failure then raised a second wake"
    },
    {
      "name": "Open a real Claude primary without waiting for stalled summary validation",
      "result": "pass",
      "live": true,
      "evidence": "live-refresh-results.json",
      "reason": "Startup completed in 8.563s while the detached validation barrier was held and no ledger existed; releasing the barrier subsequently produced a valid ledger."
    }
  ],
  "cleanup": "Private Claude socket stopped and lab home removed; every provisioned Herdr lab torn down with the helper tripwire. Disposable worktree pool returned; accidental default-pool fixture slot destroyed through Treehouse and subsequent pools configured in-project.",
  "test_setup_fixes": [
    "Stopped the short-deadline competing watcher before direct stale-owner recovery, preserving the original single-flight and publication assertions.",
    "Repeating-failure regression now drives real directory-destination publication failures under the unchanged default deadline and checks every acquired streak count. Separate timeout and serialization checks remain intact; the initial 1s fixture could fail before acquisition, and a 5s validation-entry trial remained host-dependent."
  ],
  "targeted_regressions": {
    "ownership": "tests/fm-home-summary-refresh-ownership.test.sh passed",
    "checkpoint_migration": "tests/fm-wake-drain-open-decisions-cursor.test.sh passed",
    "refresh": "Full focused refresh runs exercised watcher cadence, serialization, initialization and logging deadlines, directory publication, accumulated-history publication, and detached follow-ups; fixture races were corrected. The final verbatim escalation/recovery/large-history scenario driver passed.",
    "production_changes": "None; only tests/fm-home-summary-refresh.test.sh modified."
  }
}
Evidence: Live refresh outcomes, streaks, wakes, and published ledgers

Source: Live refresh outcomes, streaks, wakes, and published ledgers

{
  "results": [
    {
      "name": "Reject a directory ledger destination without false publication",
      "result": "pass",
      "live": true,
      "evidence": "live-refresh-results.json",
      "reason": "fm-home-summary-refresh: atomic ledger replacement failed: destination is a directory: ~/.no-mistakes/worktrees/32d18ed9638d/01M4CVJREG2ZWYYTNDVWK2DHWT/.validation-tmp/live/escalation/state/home-summary.json"
    },
    {
      "name": "Escalate three real acquired failures once, without repeated wake spam",
      "result": "pass",
      "live": true,
      "evidence": "live-refresh-results.json",
      "reason": "[{\"attempt\": 1, \"exit\": 0, \"elapsed_seconds\": 0.73, \"streak\": \"count=1\\nfirst=2026-10-08T14:26:57Z\\nclass=atomic ledger replacement failed: destination is a directory: ~\\nescalated=\\n\", \"wake_count\": 0}, {\"attempt\": 2, \"exit\": 0, \"elapsed_seconds\": 0.729, \"streak\": \"count=2\\nfirst=2026-10-08T14:26:57Z\\nclass=atomic ledger replacement failed: destination is a directory: ~\\nescalated=\\n\", \"wake_count\": 0}, {\"attempt\": 3, \"exit\": 0, \"elapsed_seconds\": 0.947, \"streak\": \"count=3\\nfirst=2026-10-08T14:26:57Z\\nclass=atomic ledger replacement failed: destination is a directory: ~\\nescalated=atomic ledger replacement failed: destination is a directory: ~\\n\", \"wake_count\": 1}, {\"attempt\": 4, \"exit\": 0, \"elapsed_seconds\": 1.255, \"streak\": \"count=4\\nfirst=2026-10-08T14:26:57Z\\nclass=atomic ledger replacement failed: destination is a directory: ~\\nescalated=atomic ledger replacement failed: destination is a directory: ~\\n\", \"wake_count\": 1}]"
    },
    {
      "name": "Recover with a valid ledger and escalate a subsequent fresh failure streak",
      "result": "pass",
      "live": true,
      "evidence": "live-refresh-results.json",
      "reason": "Success in 1.74s; schema=fm-secondmate-home-summary.v1; streak removed=True; subsequent wake count=2"
    },
    {
      "name": "Escalate and release refresh ownership even when diagnostic logging blocks",
      "result": "pass",
      "live": true,
      "evidence": "live-refresh-results.json",
      "reason": "[{\"attempt\": 1, \"elapsed_seconds\": 5.396, \"streak\": \"count=1\\nfirst=2026-10-08T14:27:07Z\\nclass=atomic ledger replacement failed: destination is a directory: ~\\nescalated=\\n\", \"wake_count\": 0, \"lock_remaining\": false}, {\"attempt\": 2, \"elapsed_seconds\": 5.272, \"streak\": \"count=2\\nfirst=2026-10-08T14:27:07Z\\nclass=atomic ledger replacement failed: destination is a directory: ~\\nescalated=\\n\", \"wake_count\": 0, \"lock_remaining\": false}, {\"attempt\": 3, \"elapsed_seconds\": 5.593, \"streak\": \"count=3\\nfirst=2026-10-08T14:27:07Z\\nclass=atomic ledger replacement failed: destination is a directory: ~\\nescalated=atomic ledger replacement failed: destination is a directory: ~\\n\", \"wake_count\": 1, \"lock_remaining\": false}]"
    },
    {
      "name": "Return detached triggers immediately and drain more than three pending follow-ups",
      "result": "pass",
      "live": true,
      "evidence": "live-refresh-results.json",
      "reason": "Initial detach=0.05s; queued=['trigger-1', 'trigger-2', 'trigger-3', 'trigger-4', 'trigger-5']; followups=[{'round': 1, 'contender_return_seconds': 0.018, 'refresh_owner_pid': '53359', 'validation_process': '65057 63100'}, {'round': 2, 'contender_return_seconds': 0.033, 'refresh_owner_pid': '53359', 'validation_process': '69647 67902'}, {'round': 3, 'contender_return_seconds': 0.019, 'refresh_owner_pid': '53359', 'validation_process': '78190 71918'}, {'round': 4, 'contender_return_seconds': 0.029, 'refresh_owner_pid': '53359', 'validation_process': '85326 82760'}, {'round': 5, 'contender_return_seconds': 0.023, 'refresh_owner_pid': '53359', 'validation_process': '90211 88240'}]"
    },
    {
      "name": "Cancel a refresh parent and prevent its orphan from overwriting newer state",
      "result": "pass",
      "live": true,
      "evidence": "live-refresh-results.json",
      "reason": "Dead owner=50222; newer refresh exit=0, duration=1.081s; orphan did not publish into directory or erase streak"
    },
    {
      "name": "Start an upgraded real Claude primary with a version-9 checkpoint and 24,000 routine status lines",
      "result": "pass",
      "live": true,
      "evidence": "live-refresh-results.json",
      "reason": "elapsed_seconds=10.494; status_bytes=936060; buried-live retained; obsolete phantom rejected; NEXT STEP reached"
    },
    {
      "name": "Re-emit startup context without losing a buried unresolved decision",
      "result": "pass",
      "live": true,
      "evidence": "live-refresh-results.json",
      "reason": "elapsed_seconds=8.281; buried-live retained; NEXT STEP reached"
    },
    {
      "name": "Spawn and tear down a real worker without waiting for blocked summary publication",
      "result": "pass",
      "live": true,
      "evidence": "live-refresh-results.json",
      "reason": "Spawn=13.846s; teardown=21.565s with summary validation blocked; final active_children=[]; pool inside worktree"
    },
    {
      "name": "Reconcile quoted-looking task IDs in a comma-containing home and release lifecycle locks",
      "result": "pass",
      "live": true,
      "evidence": "live-refresh-results.json",
      "reason": "bootstrap-batch-transcript.txt: numeric/boolean IDs avoided per-record reads; queued owned row healed; both worker locks acquired before bootstrap exited"
    },
    {
      "name": "Complete a status line after a partial checkpoint and preserve pre-checkpoint decisions in publication",
      "result": "pass",
      "live": true,
      "evidence": "live-refresh-results.json",
      "reason": "Completed split decision surfaced in the real drain; after resolving it, published keys=['late', 'still-open']; cold partial checkpoint offset=offset=85"
    },
    {
      "name": "Bound stalled refreshes, preserve the last good ledger, and report a changed failure reason",
      "result": "pass",
      "live": true,
      "evidence": "live-refresh-results.json",
      "reason": "Three acquired validation timeouts retained identical ledger bytes and released ownership; wake on third only; directory-destination failure then raised a second wake"
    },
    {
      "name": "Open a real Claude primary without waiting for stalled summary validation",
      "result": "pass",
      "live": true,
      "evidence": "live-refresh-results.json",
      "reason": "Startup completed in 8.563s while the detached validation barrier was held and no ledger existed; releasing the barrier subsequently produced a valid ledger."
    }
  ],
  "product_outputs": [
    {
      "scenario": "directory publication and escalation",
      "direct": {
        "exit": 1,
        "elapsed_seconds": 0.533,
        "stderr": "fm-home-summary-refresh: atomic ledger replacement failed: destination is a directory: ~/.no-mistakes/worktrees/32d18ed9638d/01M4CVJREG2ZWYYTNDVWK2DHWT/.validation-tmp/live/escalation/state/home-summary.json\n"
      },
      "steps": [
        {
          "attempt": 1,
          "exit": 0,
          "elapsed_seconds": 0.73,
          "streak": "count=1\nfirst=2026-10-08T14:26:57Z\nclass=atomic ledger replacement failed: destination is a directory: ~\nescalated=\n",
          "wake_count": 0
        },
        {
          "attempt": 2,
          "exit": 0,
          "elapsed_seconds": 0.729,
          "streak": "count=2\nfirst=2026-10-08T14:26:57Z\nclass=atomic ledger replacement failed: destination is a directory: ~\nescalated=\n",
          "wake_count": 0
        },
        {
          "attempt": 3,
          "exit": 0,
          "elapsed_seconds": 0.947,
          "streak": "count=3\nfirst=2026-10-08T14:26:57Z\nclass=atomic ledger replacement failed: destination is a directory: ~\nescalated=atomic ledger replacement failed: destination is a directory: /User

... [12208 bytes truncated] ...

,
          "valid_until": 900,
          "captain": [],
          "captain_omitted": 0
        },
        "generated": "2026-10-08T14:32:39Z",
        "generated_epoch": 1791469959,
        "home": "~/.no-mistakes/worktrees/32d18ed9638d/01M4CVJREG2ZWYYTNDVWK2DHWT/.validation-tmp/live/lifecycle",
        "valid": false,
        "reason": "live child state has no in-flight backlog item: live-lifecycle=unknown",
        "invalidity": {
          "kind": "unowned_current",
          "ids": [
            "live-lifecycle"
          ]
        },
        "state": "unknown",
        "active_children": [],
        "decisions_open": [],
        "holds": [],
        "queued": [
          {
            "id": "live-lifecycle",
            "title": "Exercise detached lifecycle refresh",
            "blocked_by": null,
            "blocked_by_ids": [],
            "unresolved_blocker_ids": [],
            "blocked_reason": null,
            "hold_reason": null,
            "hold_kind": null,
            "hold_until": null,
            "hold_bucket": null,
            "hold_age_days": 0,
            "captain_actionable": false,
            "repo": "fixture",
            "kind": "ship",
            "since": "2026-10-08"
          }
        ],
        "landed": [],
        "endpoints": [
          {
            "id": "live-lifecycle",
            "state": "unknown",
            "source": "pane",
            "endpoint": {
              "target": "fm-lab-start-fast-16920-17434:w1:p2",
              "exists": true,
              "agent_alive": "dead",
              "status": "dead",
              "observed_at": "2026-10-08T14:32:39Z",
              "freshness": "fresh"
            }
          }
        ],
        "counts": {
          "active_children": 0,
          "decisions_open": 0,
          "holds": 0,
          "queued": 1,
          "landed": 0,
          "endpoints": 1
        },
        "omitted": []
      }
    },
    {
      "scenario": "verified-parent in-project lifecycle",
      "metadata": {
        "window": "fm-lab-start-fast-72061-4735:w2:p2",
        "endpoint_task_id": "live-local-pool",
        "worktree": "~/.no-mistakes/worktrees/32d18ed9638d/01M4CVJREG2ZWYYTNDVWK2DHWT/.validation-tmp/live/lifecycle-project/.treehouse/lifecycle-project-5fdf64/1/lifecycle-project",
        "project": "~/.no-mistakes/worktrees/32d18ed9638d/01M4CVJREG2ZWYYTNDVWK2DHWT/.validation-tmp/live/lifecycle-project",
        "harness": "sh",
        "kind": "ship",
        "mode": "local-only",
        "yolo": "off",
        "branch": "fm/live-local-pool",
        "tasktmp": "/tmp/fm-live-local-pool",
        "model": "default",
        "effort": "default",
        "skill_selection": "undelivered",
        "skill_selection_reason": "raw launch command has no supported brief transport",
        "spawn_gen": "s1791470221.40113.32462",
        "backend": "herdr",
        "herdr_session": "fm-lab-start-fast-72061-4735",
        "herdr_workspace_id": "w2",
        "herdr_tab_id": "w2:t2",
        "herdr_pane_id": "w2:p2"
      },
      "spawn_seconds": 13.846,
      "teardown_seconds": 21.565,
      "final_ledger": {
        "schema": "fm-secondmate-home-summary.v1",
        "hold_classifier_schema": "fm-captain-hold-buckets.v1",
        "contributions": {
          "known": 0,
          "checked": 0,
          "counts": {
            "captain": 0,
            "fleet": 0,
            "maintainer": 0,
            "nobody": 0
          },
          "unmeasured": 0,
          "complete": true,
          "proven_clear": true,
          "stale_verdicts": 0,
          "missing_verdicts": 0,
          "unreadable_records": 0,
          "valid_until": 900,
          "captain": [],
          "captain_omitted": 0
        },
        "generated": "2026-10-08T14:37:02Z",
        "generated_epoch": 1791470222,
        "home": "~/.no-mistakes/worktrees/32d18ed9638d/01M4CVJREG2ZWYYTNDVWK2DHWT/.validation-tmp/live/lifecycle-local-pool",
        "valid": false,
        "reason": "live child state has no in-flight backlog item: live-local-pool=unknown",
        "invalidity": {
          "kind": "unowned_current",
          "ids": [
            "live-local-pool"
          ]
        },
        "state": "unknown",
        "active_children": [],
        "decisions_open": [],
        "holds": [],
        "queued": [
          {
            "id": "live-local-pool",
            "title": "Exercise detached lifecycle refresh",
            "blocked_by": null,
            "blocked_by_ids": [],
            "unresolved_blocker_ids": [],
            "blocked_reason": null,
            "hold_reason": null,
            "hold_kind": null,
            "hold_until": null,
            "hold_bucket": null,
            "hold_age_days": 0,
            "captain_actionable": false,
            "repo": "lifecycle-project",
            "kind": "ship",
            "since": "2026-10-08"
          }
        ],
        "landed": [],
        "endpoints": [
          {
            "id": "live-local-pool",
            "state": "unknown",
            "source": "pane",
            "endpoint": {
              "target": "fm-lab-start-fast-72061-4735:w2:p2",
              "exists": true,
              "agent_alive": "dead",
              "status": "dead",
              "observed_at": "2026-10-08T14:37:02Z",
              "freshness": "fresh"
            }
          }
        ],
        "counts": {
          "active_children": 0,
          "decisions_open": 0,
          "holds": 0,
          "queued": 1,
          "landed": 0,
          "endpoints": 1
        },
        "omitted": []
      }
    },
    {
      "scenario": "timeouts preserve atomic ledger and failure class change",
      "attempts": [
        {
          "attempt": 1,
          "exit": 0,
          "elapsed_seconds": 10.094,
          "streak": "count=1\nfirst=2026-10-08T14:44:03Z\nclass=refresh exceeded its -second deadline\nescalated=\n",
          "wake_count": 0,
          "last_good_ledger_preserved": true,
          "lock_remaining": false
        },
        {
          "attempt": 2,
          "exit": 0,
          "elapsed_seconds": 7.301,
          "streak": "count=2\nfirst=2026-10-08T14:44:03Z\nclass=refresh exceeded its -second deadline\nescalated=\n",
          "wake_count": 0,
          "last_good_ledger_preserved": true,
          "lock_remaining": false
        },
        {
          "attempt": 3,
          "exit": 0,
          "elapsed_seconds": 8.358,
          "streak": "count=3\nfirst=2026-10-08T14:44:03Z\nclass=refresh exceeded its -second deadline\nescalated=refresh exceeded its -second deadline\n",
          "wake_count": 1,
          "last_good_ledger_preserved": true,
          "lock_remaining": false
        }
      ],
      "wake_queue": "1791470667\t1\tcheck\thome-summary-refresh\tcheck: home-summary-refresh: 3 consecutive refresh failures since 2026-10-08T14:44:03Z; last: refresh exceeded its 5-second deadline (the failed attempt ran 6s). state/home-summary.json is not being republished; reproduce with bin/fm-home-summary-refresh.sh and read state/.home-summary-refresh.log\n1791470671\t2\tcheck\thome-summary-refresh\tcheck: home-summary-refresh: 4 consecutive refresh failures since 2026-10-08T14:44:03Z; last: atomic ledger replacement failed: destination is a directory: ~/.no-mistakes/worktrees/32d18ed9638d/01M4CVJREG2ZWYYTNDVWK2DHWT/.validation-tmp/live/deadline-preserves-ledger/state/home-summary.json (the failed attempt ran 1s). state/home-summary.json is not being republished; reproduce with bin/fm-home-summary-refresh.sh and read state/.home-summary-refresh.log\n",
      "diagnostic_log": "[2026-10-08T14:44:03Z] refresh exceeded its 5-second deadline\n[2026-10-08T14:44:13Z] refresh exceeded its 5-second deadline\n[2026-10-08T14:44:20Z] refresh exceeded its 5-second deadline\n[2026-10-08T14:44:29Z] atomic ledger replacement failed: destination is a directory: ~/.no-mistakes/worktrees/32d18ed9638d/01M4CVJREG2ZWYYTNDVWK2DHWT/.validation-tmp/live/deadline-preserves-ledger/state/home-summary.json\n"
    }
  ]
}
Evidence: Upgraded-startup timing

Source: Upgraded-startup timing

elapsed_seconds=10.494
Evidence: Blocked-summary startup timing

Source: Blocked-summary startup timing

PATH= FM_LIVE_GATE=/tmp/fm-lab.XrOZ3E/state/validation-gate FM_HOME= TMUX=     1.33s user 2.81s system 48% cpu 8.563 total
Evidence: Detached ledger published after releasing validation

Source: Detached ledger published after releasing validation

{
  "schema": "fm-secondmate-home-summary.v1",
  "hold_classifier_schema": "fm-captain-hold-buckets.v1",
  "contributions": {
    "known": 0,
    "checked": 0,
    "counts": {
      "captain": 0,
      "fleet": 0,
      "maintainer": 0,
      "nobody": 0
    },
    "unmeasured": 0,
    "complete": true,
    "proven_clear": true,
    "stale_verdicts": 0,
    "missing_verdicts": 0,
    "unreadable_records": 0,
    "valid_until": 900,
    "captain": [],
    "captain_omitted": 0
  },
  "generated": "2026-10-08T14:46:22Z",
  "generated_epoch": 1791470782,
  "home": "/tmp/fm-lab.XrOZ3E",
  "valid": true,
  "reason": null,
  "invalidity": {
    "kind": null,
    "ids": []
  },
  "state": "no_active_work",
  "active_children": [],
  "decisions_open": [],
  "holds": [],
  "queued": [],
  "landed": [],
  "endpoints": [],
  "counts": {
    "active_children": 0,
    "decisions_open": 0,
    "holds": 0,
    "queued": 0,
    "landed": 0,
    "endpoints": 0
  },
  "omitted": []
}
Evidence: Batched backlog reconciliation and concurrent lock acquisition

Source: Batched backlog reconciliation and concurrent lock acquisition

BOOTSTRAP OUTPUT
BOOTSTRAP_INFO: marked queued-owned in flight to match the worker this home already owns

OBSERVED REAL TASKS-AXI CALLS
--version 
update --help
mv --help
list --file
show queued-owned
show queued-owned
start queued-owned

CONCURRENT LOCK ACQUISITION
all worker lifecycle locks available while bootstrap is still running

FINAL BACKLOG
count: 5
tasks[5]{id,state,kind,repo,title}:
  queued-owned,in_flight,ship,fixture,Exercise batch reconciliation
  "null",in_flight,ship,fixture,Exercise batch reconciliation
  "false",in_flight,ship,fixture,Exercise batch reconciliation
  "true",in_flight,ship,fixture,Exercise batch reconciliation
  "123",in_flight,ship,fixture,Exercise batch reconciliation
help[2]:
  - "Run `tasks-axi show <id> --file=~/.no-mistakes/worktrees/32d18ed9638d/01M4CVJREG2ZWYYTNDVWK2DHWT/.validation-tmp/live/batch,a,b/data/backlog.md` for full notes on a task"
  - "Run `tasks-axi ready --file=~/.no-mistakes/worktrees/32d18ed9638d/01M4CVJREG2ZWYYTNDVWK2DHWT/.validation-tmp/live/batch,a,b/data/backlog.md` to see unblocked queued work"
Evidence: Partial-checkpoint drain and publication behavior

Source: Partial-checkpoint drain and publication behavior

FIRST DRAIN (PARTIAL APPEND)
OPEN DECISIONS (still open, folded from the durable status logs - not just the latest line):
edge [key=still-open] needs-decision: retain the buried decision
OPEN DECISIONS: close one by answering it: bin/fm-send.sh <task> --resolve-key <key> '<answer>'
●━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
●  WATCHER DOWN - SUPERVISION IS OFF
●  1 task(s) in flight, but no watcher has a fresh beacon (last beat: never, grace 300s).
●  Trust the emitted supervision protocol for this harness; do not use shell & for watcher repair.
●  This is a supervision warning only; the guarded operation WILL still run.
●  repair a missing or failed watcher cycle with the omp tool fm_watch_arm_omp, or restart omp inside this home so ~/.no-mistakes/worktrees/32d18ed9638d/01M4CVJREG2ZWYYTNDVWK2DHWT/.omp/extensions/fm-primary-turnend-guard.ts and ~/.no-mistakes/worktrees/32d18ed9638d/01M4CVJREG2ZWYYTNDVWK2DHWT/.omp/extensions/fm-primary-omp-watch.ts auto-load from .omp/extensions/ (use -e with both paths only when starting omp from another directory).
●━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

PARTIAL CHECKPOINT
version=10:ship
offset=85
ident=strong:16777232:134357519:1791470480.442324101
still-open	needs-decision	retain the buried decision
SECOND DRAIN (COMPLETED APPEND)
OPEN DECISIONS (still open, folded from the durable status logs - not just the latest line):
edge [key=still-open] needs-decision: retain the buried decision
edge [key=split] needs-decision: choose the endpoint
OPEN DECISIONS: close one by answering it: bin/fm-send.sh <task> --resolve-key <key> '<answer>'
WARNING: watcher still down (same stale episode; last beat: never, grace 300s) - full banner already printed this episode.

PUBLISHED LEDGER AFTER RESOLUTION AND LATE APPEND
{
  "schema": "fm-secondmate-home-summary.v1",
  "hold_classifier_schema": "fm-captain-hold-buckets.v1",
  "contributions": {
    "known": 0,
    "checked": 0,
    "counts": {
      "captain": 0,
      "fleet": 0,
      "maintainer": 0,
      "nobody": 0
    },
    "unmeasured": 0,
    "complete": true,
    "proven_clear": true,
    "stale_verdicts": 0,
    "missing_verdicts": 0,
    "unreadable_records": 0,
    "valid_until": 900,
    "captain": [],
    "captain_omitted": 0
  },
  "generated": "2026-10-08T14:41:36Z",
  "generated_epoch": 1791470496,
  "home": "~/.no-mistakes/worktrees/32d18ed9638d/01M4CVJREG2ZWYYTNDVWK2DHWT/.validation-tmp/live/partial-checkpoint",
  "valid": false,
  "reason": "child current state unavailable: edge",
  "invalidity": {
    "kind": "child_current_unavailable",
    "ids": [
      "edge"
    ]
  },
  "state": "unknown",
  "active_children": [],
  "decisions_open": [
    {
      "id": "edge",
      "key": "still-open",
      "verb": "needs-decision",
      "summary": "retain the buried decision",
      "reason": null,
      "source": "status"
    },
    {
      "id": "edge",
      "key": "late",
      "verb": "needs-decision",
      "summary": "added after checkpoint",
      "reason": null,
      "source": "status"
    }
  ],
  "holds": [],
  "queued": [],
  "landed": [],
  "endpoints": [
    {
      "id": "edge",
      "state": "unknown",
      "source": "none",
      "endpoint": {
        "target": "absent:fm-edge",
        "exists": false,
        "agent_alive": "missing",
        "status": "absent",
        "observed_at": "2026-10-08T14:41:36Z",
        "freshness": "fresh"
      }
    }
  ],
  "counts": {
    "active_children": 0,
    "decisions_open": 2,
    "holds": 0,
    "queued": 0,
    "landed": 0,
    "endpoints": 1
  },
  "omitted": []
}
- Outcome: 🔧 4 issues found → auto-fixed ✅ across 4 runs (3h9m42s)

Pipeline

Updates from git push no-mistakes

... (12 earlier update rounds omitted to keep the PR body within GitHub's 65536-char limit; full history is in the run log.)

🔧 **Review** - 6 issues found → auto-fixed (5) ✅

🔧 Fix applied.
✅ Re-checked - no issues remain.

✅ No issues found.

🔧 **Test** - 4 issues found → auto-fixed ✅
  • 🚨 bin/fm-classify-lib.sh:358 - First-upgrade startup still reproduces the reported latency problem. A real Claude primary with a valid version-9 checkpoint and an 840,058-byte status log containing 24,000 routine lines hit the 120-second startup bound at 120.784s in wake-queue. The buried decision and remaining digest sections were not delivered. Version 10 correctly rejects the old checkpoint, but the incremental cold fold still invokes a subprocess for every status line. Optimize cold/refused-checkpoint folding without trusting obsolete checkpoints or increasing the startup timeout, and retain an executable large-history migration regression.
  • ⚠️ tests/fm-home-summary-refresh.test.sh:708 - The focused refresh script remains failing at its accumulated-home publication scenario: 300 wide working lines plus the cost-gate decision exceeded the configured 30-second deadline and produced no ledger. Fixture-root isolation was corrected, and the preceding modified timeout, initialization, and logging checks passed. A separate real no-endpoint variant published the same history shape in 12.006s, so that narrower live result does not resolve the fuller owned-worker case. Diagnose the classification/fixture difference rather than increasing the deadline.
  • ⚠️ bin/fm-home-summary-refresh.sh:217 - A directory at state/home-summary.json produces false publication success. The real writer moved its temporary JSON inside that directory, returned zero, and left no failure streak while the public ledger path remained unreadable as JSON. This bypasses the new failure accounting. Attribution to this change has not been established, but the exercised publication boundary should reject a directory destination and account for the acquired failure.
  • 🚨 live validation verdict: no-go (10 of 10 scenarios were driven live against the product); failed: Start the first upgraded session with an older checkpoint and long status history, Surface an acquired publication failure when the ledger destination is a directory
  • Live validation: ❌ no-go - 10 of 10 scenarios driven live against the product
Scenario Result Live Evidence
Start a cached session while another summary refresh holds the lock ✅ pass live live-busy-summary-start-digest.txt; live-startup-observations.json
Start the first upgraded session with an older checkpoint and long status history ❌ fail live live-upgrade-session-start-digest.txt; live-upgrade-startup-results.json
Heal queued worker records without per-worker backlog reads ✅ pass live live-bootstrap-bulk-results.json
Preserve decisions across cached tails, partial appends, invalid checkpoints, and log replacement ✅ pass live live-fold-scenarios.json; live-fold-copied-snapshot.json
Publish accumulated wide status history for a worker with no recorded endpoint ✅ pass live live-accumulated-wide-results.json; live-accumulated-wide-summary.json
Escalate three real publication failures, deduplicate repeats, and reset after recovery ✅ pass live summary/live-results.json; summary/live-failure-3/.wake-queue
Keep failure accounting and lock release working when diagnostic logging blocks ✅ pass live summary/live-blocked-logger-results.json
Consume pending transitions arriving beyond the fourth successful refresh ✅ pass live live-five-followups-results.json; live-five-followups-ledger.json
Cancel an old refresh and prevent its resumed worker from overwriting a newer ledger ✅ pass live live-cancelled-owner-result.json; live-cancelled-owner-ledger.json
Surface an acquired publication failure when the ledger destination is a directory ❌ fail live summary/directory-obstacle-observation.json; summary/live-results.json
  • Scoped the change with git diff --stat 90172034c69544cbdf8e37ece135976a6e8f16b9 04b38188cc1dc715a5e07ad71829d70b8a0c8b13; read the trusted Herdr lab contracts and checked required tools on PATH.
  • TMPDIR=$WORKTREE/.validation-summary/tmp bash tests/fm-home-summary-refresh.test.sh and bash tests/fm-home-summary-refresh-ownership.test.sh; initial runs exposed fixture-root isolation and incidental timing assumptions.
  • TMPDIR=$WORKTREE/.validation-primary/tmp bash tests/fm-wake-drain-open-decisions-cursor.test.sh — passed.
  • Ran the focused summary scripts from an unchanged disposable code copy with sibling temporary homes inside the worktree, avoiding the secondmate-inside-code-root refusal.
  • bash .validation-primary/focused-test-world/code/tests/fm-home-summary-refresh-ownership.test.sh — passed after the test-only deadline correction.
  • bash .validation-primary/focused-test-world/code/tests/fm-home-summary-refresh.test.sh — modified failure-handling checks passed; the final run failed accumulated-home publication at its 30-second deadline.
  • Provisioned and tore down fm-lab-start-fast-77979-1475 and fm-lab-upgrade-cold-53699-6156 only through bin/fm-herdr-lab.sh; both teardown tripwires verified the default session remained unchanged.
  • From real Claude primaries, executed bash -c &#39;TIMEFORMAT=&#34;STARTUP_RUNTIME_SECONDS=%R&#34;; time ./bin/fm-session-start.sh&#39; with a held summary lock and with a previous-version long-history checkpoint.
  • FM_BOOTSTRAP_NETWORK=skip bin/fm-bootstrap.sh against a marked comma-containing home, using a transparent tracer delegating to installed tasks-axi; observed one list and one show while healing the queued record.
  • Executed real bin/fm-wake-drain.sh and bin/fm-fleet-snapshot.sh --secondmate-home-summary across cached tails, partial appends, invalid boundaries, log replacement, terminal closure, and reopening.
  • Executed real bin/fm-home-summary-refresh.sh --best-effort with an immutable ledger for three-failure escalation, fourth-failure deduplication, recovery, and re-escalation; also exercised an unavailable logger with a full stderr pipe.
  • Executed five real --detach triggers during successive refreshes; verified the same owner drained all pending transitions.
  • Paused an actual refresh worker, killed its diagnostic parent, published a newer ledger, then resumed the stale worker; verified the newer ledger remained byte-identical.
  • Executed the real writer against a directory ledger destination and a separate 300-wide-line marked-home fixture without backend mocks; preserved resulting product state and removed all disposable fixtures.

🔧 Fix applied.
✅ Re-checked - no issues remain.

  • Live validation: ✅ go - 15 of 15 scenarios driven live against the product
Scenario Result Live Evidence
Start a real Claude primary with an old checkpoint and 24,000 status lines ✅ pass live claude-startup-digest.txt and claude-lab-digest-stream.jsonl: completed in 26.18 seconds, surfaced upgrade-gate, and emitted the final NEXT STEP section without truncation.
Reopen startup context using the migrated checkpoint ✅ pass live claude-warm-digest.txt: retained upgrade-gate and emitted the complete digest; the real Claude turn completed in 5.62 seconds.
Publish accumulated history within a 30-second deadline ✅ pass live accumulated-home-ledger.json and summary-writer-command-transcript.json: publication completed in 3.94 seconds with cost-gate preserved.
Publish accumulated history for an owned Claude endpoint ✅ pass live owned-worker-history-ledger.json and owned-worker-pane.txt: endpoint exists=true, cost-gate preserved, and the real Claude hook plus model turn completed in 7.89 seconds.
Reject a directory ledger destination and escalate the third acquired failure once ✅ pass live publication-failure-state.json: directory contents remained unchanged, streak counts advanced, one wake appeared on failure three, and failure four did not duplicate it.
Recover publication and allow a later failure streak to escalate again ✅ pass live publication-failure-state.json and recovered-home-ledger.json: successful publication cleared the streak; three subsequent failures produced a new wake.
Escalate acquired failures despite blocked diagnostic logging ✅ pass live blocked-logger-state.json: a directory log destination and full undrained stderr pipe did not prevent accounting, third-failure escalation, or lock release.
Bound a genuinely slow refresh and escalate its third acquired timeout ✅ pass live real-timeout-state.json: three real large-history refreshes exceeded their one-second deadlines, released ownership, and produced one threshold wake.
Raise a new wake for a changed failure class without repeating identical timeout wakes ✅ pass live changed-failure-class-state.json: changing directory-publication failures to real timeouts produced one additional wake; the next timeout did not duplicate it.
Return immediately under a held refresh lock and converge concurrent triggers ✅ pass live concurrent-trigger-ledger.json: initial detach returned in 0.073 seconds; the trigger burst converged on the latest decision with no remaining pending marker or lock.
Cancel a refresh parent without allowing its orphaned worker to overwrite newer state ✅ pass live parent-cancellation-state.json: resuming the real orphaned worker preserved both the newer ledger and newer failure streak.
Preserve decision transitions across append, terminal completion, replacement, and partial-line boundaries ✅ pass live decision-transition-transcript.json: public wake drains closed terminal decisions, surfaced reopened and replacement decisions, and correctly folded a completed partial line.
Recover a queued owned task with numeric-looking IDs and a comma-containing home ✅ pass live backlog-batching-transcript.txt: real tasks-axi output decoded 123, true, false, and null; bootstrap healed heal-queued to in_flight.
Read a large child publication through the parent fleet snapshot ✅ pass live large-parent-snapshot.json: the supported isolated topology consumed the local ledger and preserved all 600 orphan IDs.
Spawn and tear down a real long-running Herdr task without waiting for summary publication ✅ pass live herdr-lifecycle-transcript.txt: spawn completed in 16.51 seconds and teardown in 12.63 seconds while the refresh lock remained held; metadata and the endpoint were retired.
  • bash tests/fm-home-summary-refresh-ownership.test.sh — passed.
  • bash tests/fm-wake-drain-open-decisions-cursor.test.sh — passed.
  • bash tests/fm-home-summary-refresh.test.sh — initially failed because workspace-local temporary homes were nested under the code root; the unchanged test passed as bash .live-validation-emr992ci/code-root/tests/fm-home-summary-refresh.test.sh using a disposable sibling code root.
  • Launched real Claude primaries on private, non-zero-sized tmux sessions; disposable SessionStart settings invoked bin/fm-session-start.sh --source startup and bin/fm-session-start.sh --reemit --source clear.
  • Drove bin/fm-home-summary-refresh.sh, --best-effort, and --detach against accumulated histories, directory destinations, held refresh locks, blocked logging, and actual one-second deadlines.
  • Killed a real refresh parent while its worker was suspended, published newer state, recorded a newer failure, and resumed the orphaned worker.
  • Drove bin/fm-wake-drain.sh through decision opening, terminal completion, reopening, authoritative-log replacement, and partial-line completion.
  • Used real tasks-axi add, list, and show commands plus bin/fm-bootstrap.sh to recover an owned queued task with numeric-looking IDs and a comma-containing home path.
  • Provisioned a named Herdr lab through bin/fm-herdr-lab.sh; scaffolded with bin/fm-brief.sh lifecycle batch-project --mode local-only --herdr-lab; drove real long-running task spawn and teardown while summary publication was blocked.
  • Ran the real summary writer from Claude SessionStart with the pane's actual private-socket TMUX environment, an owned endpoint, 300 wide status lines, and real no-mistakes state reads.
  • Tore down Herdr with its default-fleet tripwire intact, stopped private tmux servers, removed their helper-owned socket directories, destroyed the exact test-created Treehouse worktree, and removed disposable worktree fixtures.

✅ No issues found.

  • Live validation: ✅ go - 12 of 12 scenarios driven live against the product
Scenario Result Live Evidence
Run the first upgraded startup with an old checkpoint and long history; receive the buried decision and complete digest ✅ pass live Upgraded startup digest; Validation observations
Start while another summary refresh owns the lock; finish startup without waiting or recording a contention failure ✅ pass live Upgraded startup digest; Validation observations
Append a decision after a large-history checkpoint; publish both the earlier and new decisions ✅ pass live Warm publication retains earlier and appended decisions
Publish an owned worker's accumulated wide status history within the existing deadline ✅ pass live Accumulated wide-history publication
Refresh when the ledger destination is a directory; reject false publication and preserve its contents ✅ pass live Real publication failures, escalation, and recovery
Repeat an acquired publication failure; wake on the third attempt, suppress duplicates, and clear the streak after recovery ✅ pass live Real publication failures, escalation, and recovery
Repeat acquired failures while logging is blocked; still escalate and release refresh ownership ✅ pass live Escalation and lock release despite blocked logging
Append multiple transitions during overlapping detached refreshes; converge to the latest state ✅ pass live Detached refresh burst convergence
Cancel a refresh after calculation while publication is fenced; preserve the newer publication ✅ pass live Cancelled-owner publication fencing
Complete a partial status line or change fold overrides; refuse incompatible checkpoints without losing the decision ✅ pass live Partial-line and override checkpoint drains
Bootstrap numeric and boolean-looking task IDs from a comma-containing home; heal queued ownership and release lifecycle locks ✅ pass live Real batched bootstrap and concurrent lifecycle locks
Spawn and tear down a real scout while summary publication is blocked; complete lifecycle without waiting for the refresh ✅ pass live Real scout spawn; Real scout teardown; Spawn and teardown do not wait for summary ownership
  • Inspected the runtime changes against base commit cb7205a65d14a1d45338c26a030ae551ed6ec5db and read the trusted Herdr lab contracts.
  • Provisioned two named non-default sessions through bin/fm-herdr-lab.sh provision; drove their real Claude processes through the helper and removed both through guarded teardown.
  • A real Claude primary ran FM_HOME=&lt;marked lab&gt; bin/fm-session-start.sh --source startup with a version-9 checkpoint, 24,000 routine status lines, and a held summary-refresh lock.
  • Drove bin/fm-home-summary-refresh.sh, --best-effort, and --detach against real isolated producers for accumulated history, appended decisions, directory destinations, failure escalation, recovery, overlapping triggers, blocked logging, and parent cancellation.
  • Drove bin/fm-wake-drain.sh across an incomplete status line, its completion, resolution, and a changed resolution-verb override.
  • Used the installed tasks-axi CLI and real bin/fm-bootstrap.sh with numeric/boolean-looking IDs and a comma-containing home path; acquired lifecycle locks concurrently while bootstrap remained running.
  • Ran real bin/fm-spawn.sh live-scout &lt;disposable project&gt; --scout --harness claude --backend herdr, then bin/fm-captain-hold.sh complete live-scout --none and bin/fm-teardown.sh live-scout, while summary-refresh ownership remained held.
  • bash tests/fm-home-summary-refresh-ownership.test.sh passed.
  • bash tests/fm-home-summary-refresh.test.sh initially failed because repository-local temporary secondmate homes violated the supported isolation guard; after collecting the actual failed-state diagnostic, TMPDIR=/tmp/ bash tests/fm-home-summary-refresh.test.sh passed.
  • bash tests/fm-wake-drain-open-decisions-cursor.test.sh initially exceeded an external 300-second runner limit; the unchanged focused script subsequently completed successfully in 469.16 seconds with an 1800-second external limit. The product startup deadline was not increased.
  • Preserved product digests, JSON ledgers, wake/streak state, lifecycle transcripts, and focused regression logs in the evidence directory; removed transient fixtures and generated startup state.

✅ No issues found.

  • Live validation: ✅ go - 13 of 13 scenarios driven live against the product
Scenario Result Live Evidence
Start an upgraded real Claude primary with an obsolete checkpoint and long status history ✅ pass live Real Claude upgraded-startup digest; Upgraded-startup timing
Re-emit startup context while retaining a buried unresolved decision ✅ pass live Real Claude context re-emission digest; claude-live-warm-time.txt
Open a real Claude primary without waiting for stalled summary validation ✅ pass live Startup completed while detached publication was blocked; Blocked-summary startup timing; Detached ledger published after releasing validation
Reject a directory ledger destination without false publication ✅ pass live Live refresh outcomes, streaks, wakes, and published ledgers
Escalate three real acquired failures once without repeated wake spam ✅ pass live Live refresh outcomes, streaks, wakes, and published ledgers
Recover with a valid ledger and escalate a subsequent fresh failure streak ✅ pass live Live refresh outcomes, streaks, wakes, and published ledgers
Escalate and release refresh ownership even when diagnostic logging blocks ✅ pass live Live refresh outcomes, streaks, wakes, and published ledgers
Return detached triggers immediately and drain more than three pending follow-ups ✅ pass live Live refresh outcomes, streaks, wakes, and published ledgers
Cancel a refresh parent and prevent its orphan from overwriting newer state ✅ pass live Live refresh outcomes, streaks, wakes, and published ledgers
Spawn and tear down a real worker without waiting for blocked summary publication ✅ pass live Real isolated Herdr worker lifecycle
Reconcile quoted-looking task IDs in a comma-containing home and release lifecycle locks ✅ pass live Batched backlog reconciliation and concurrent lock acquisition
Complete a status line after a partial checkpoint and preserve earlier decisions in publication ✅ pass live Partial-checkpoint drain and publication behavior
Bound stalled refreshes, preserve the last good ledger, and report a changed failure reason ✅ pass live Live refresh outcomes, streaks, wakes, and published ledgers
  • Launched real Claude primaries with tmux -L fm-lab new-session on disposable private sockets, with 120×40 terminal dimensions.
  • Drove bin/fm-session-start.sh --source startup through Claude against a version-9 checkpoint and a 936,060-byte status history containing 24,000 routine lines; startup completed in 10.494s.
  • Drove bin/fm-session-start.sh --reemit --source clear through Claude; completed in 8.281s with the buried decision retained.
  • Drove startup through another real Claude primary while detached summary validation was blocked; startup completed in 8.563s, and publication succeeded after releasing the barrier.
  • Executed bin/fm-home-summary-refresh.sh, --best-effort, and --detach against disposable homes to exercise directory-destination failures, escalation, recovery, blocked FIFO logging, five pending follow-ups, parent cancellation, and acquired validation timeouts.
  • Used bin/fm-herdr-lab.sh to provision and tear down named non-default labs; exercised real bin/fm-spawn.sh and bin/fm-teardown.sh calls while summary publication was blocked.
  • Executed bin/fm-bootstrap.sh with real tasks-axi records named 123, true, false, and null in a comma-containing home; verified queued-row healing and concurrent lifecycle-lock availability.
  • Executed bin/fm-wake-drain.sh and the real summary writer across a partial-line checkpoint, completed append, resolution, and later decision append.
  • Ran bash tests/fm-home-summary-refresh-ownership.test.sh and bash tests/fm-wake-drain-open-decisions-cursor.test.sh; both passed.
  • Ran focused bash tests/fm-home-summary-refresh.test.sh checks. Diagnosed and corrected competing-watcher and short-deadline fixture assumptions, then executed the corrected escalation, recovery, and large-history sections verbatim through a disposable targeted driver; those sections passed.
  • Stopped private primary sockets, tore down Herdr labs, returned disposable worker slots, and removed repository-local fixtures and diagnostic drivers. No full repository suite, linters, formatters, or static analysis ran.
✅ **Document** - passed

✅ No issues found.

✅ No issues found.

✅ No issues found.

🔧 **Lint** - 1 issue found → auto-fixed (2) ✅
  • ⚠️ linter found issues (exit code 1)

🔧 Fix applied.
✅ Re-checked - no issues remain.

  • ⚠️ linter found issues (exit code 1)

🔧 Fix applied.
✅ Re-checked - no issues remain.

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

✅ No issues found.

✅ No issues found.

@MrGTV-love
MrGTV-love force-pushed the fm/fm-session-start-fast branch from ea05f11 to 38fe9c2 Compare October 8, 2026 13:28
@MrGTV-love MrGTV-love changed the title fix: speed up session start and surface home-summary failures fix: keep session starts responsive and escalate home-summary failures Oct 8, 2026
…escalate repeating failures

A locked session start spent its whole 60-second refresh deadline on a
publication that needed 117 seconds, so every start paid the deadline, the
refresh never published, and 425 identical failures reached nobody.

- Session start, spawn, and teardown run the refresh with the new --detach:
  single-flight, detached, never waited on. A trigger that finds a refresh in
  flight leaves a pending marker and the in-flight run refreshes once more, so
  bursts coalesce into one follow-up and none is lost.
- Three consecutive refresh failures append one check: home-summary-refresh
  wake naming the reason and measured duration; it is raised again only when
  the reason changes or after a success ends the streak.
- The refresh itself folds each task's status log from the checkpoint the
  wake drain keeps beside it instead of from line 1 (hunks from the lifecycle
  checkpoint work, reused as written).
- Bootstrap reads backlog row states once instead of once per owned record,
  canonicalizes paths and checks control bytes in the shell, and the herdr
  journal reader no longer forks per field; the wake drain drops per-line tr
  and dirname/basename forks.
- FM_SESSION_START_STAGE_TIMES_FILE times a session start stage by stage.
…cklog-atomicity.test.sh. Its aggregate show-call limit incorrectly rejected two legitimate reads of the queued task: an eligibility probe and a full-body read for drop-provenance handling. The test now records command and task ID and asserts that none of the five nonqueued tasks is read individually, in both ordinary and comma-containing home paths. Existing one-list, queued-healing, state, and metadata assertions remain intact; production behavior is unchanged. Verification: the complete backlog atomicity suite passed locally with tasks-axi 0.2.6, including the previously failing scenario. An initial run passed that scenario but hit the final secondmate fixture's repository-containment guard; rerunning with the harness's normal ephemeral temp location passed the full suite. Temporary verification fixtures were removed. No other pipeline phase or push was invoked
@MrGTV-love
MrGTV-love force-pushed the fm/fm-session-start-fast branch from ecf4abd to a3dea5c Compare October 8, 2026 16:14
@MrGTV-love MrGTV-love changed the title fix: keep session starts responsive and escalate home-summary failures fix: speed up session startup and escalate home-summary failures Oct 8, 2026
@MrGTV-love
MrGTV-love merged commit 1776681 into main Oct 8, 2026
22 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant