Skip to content

chore(integration): keep reviewed fleet fixes live across handoffs - #102730

Open
jayleaton wants to merge 12 commits into
NousResearch:mainfrom
jayleaton:wt/t_fe169b0e
Open

jayleaton wants to merge 12 commits into
NousResearch:mainfrom
jayleaton:wt/t_fe169b0e

Conversation

@jayleaton

Copy link
Copy Markdown

Integration scope

This is a single cumulative operational candidate; it does not replace or duplicate the four feature PRs. It merge-preserves the exact reviewed heads of:

The only integration conflict was cron/jobs.py::mark_job_run: the resolution retains #102073’s provider_backoff argument and forwarding while adopting #102627’s Optional[bool] tri-state return contract.

Verification

scripts/run_tests.sh over every changed Python test file: 17 files, 473 passed, 0 failed. This one matrix covers provider backoff, Chronos/provider/manual-run paths, worktree decomposition and real child spawn resolution, worker lifecycle finalization/PRODUCT_SIGNOFF, fire-fence contention/heartbeat/shutdown races, and the associated agent/provider compatibility paths.

git merge-base --is-ancestor passes for all four exact reviewed heads; git diff --check 7b72fd124..HEAD is clean.

No deployment is included in this PR, and no upstream merge is requested as part of the operational handoff.

jayleaton and others added 12 commits September 3, 2026 03:59
A decomposed child with workspace_kind=worktree and no explicit path got
workspace_path=NULL. Dispatch then tried to anchor it on the board's
default_workdir and, when the board had none, failed the spawn with
'no default_workdir set' — twice, tripping the failure-limit circuit
breaker (live: t_ab2a7ce8, t_2dd0f5b7, 2026-09-02).

The root of a decomposed callback graph is usually itself a
dispatcher-materialized worktree <repo>/.worktrees/<root-id>, so the
repo is recoverable from the root row: when the root path is a linked
worktree checkout, stamp each pathless worktree child with the root's
git common-dir parent (the main repo) as its explicit anchor. Children
still get per-task worktrees (never the root's literal checkout), so
sibling isolation is unchanged; a repo that can't be recovered keeps
the anchorless row and the board-default dispatch path; scratch roots
are untouched.

Regression tests cover anchor inheritance + end-to-end child spawn
resolution with no board default_workdir, the unrecoverable-root
fallback, and scratch-root non-leakage.
A task whose title/body carries the durable PRODUCT_SIGNOFF marker defines
its job as reporting back for human sign-off rather than driving to a
terminal lifecycle transition. Repeated clean exits without one are a
property of that workflow, so _account_protocol_violation now classifies
such tasks at accounting time (read from the task row, so the exemption
survives retries, reclaims, and dispatcher restarts) and never auto-blocks
them — the violation is still recorded and the task retried, but the
gave_up breaker never fires across any number of limit-reaching runs.

Review round 1 correction for t_4067fdf1: the one-run PRODUCT_SIGNOFF test
could not reach the breaker; it is replaced by a regression that drives the
task through the full violation limit and proves it stays ready with no
gave_up event, while ordinary coder/QA omissions keep their bounded
retry/block behavior (covered by the existing limit test).
…wnership loss (c-027)

heartbeat_fire_claim() collapsed fire-fence acquisition failure into
False. During slow fenced delivery the worker's own heartbeat thread
cannot acquire the same process-local RLock, so a successfully
delivering run was misclassified as ownership loss and terminalled as
'Interrupted by shutdown before terminal completion.' (TrustMRR exec
4219b48d while RSI recovery exec f502e929 was settling).

Tri-state contract:
- True  = renewed/confirmed owner
- False = authoritatively inspected: owner mismatch / absent claim
- None  = fence unavailable/unconfirmed (contention or store error)

- initial validation stays fail-closed: run starts only on exactly True
- heartbeat loop: None rides the existing grace window (180s prod vs
  30s fence wait); immediate cancel only on explicit False
- post-run probes (_fire_claim_ownership_lost, both interrupted blocks,
  terminal owner-CAS read): None can never adjudicate a confirmed loss;
  uncertain outcomes get distinct ledger errors
- mark_job_run also tri-state: None = fence unavailable (CAS never ran),
  False stays reserved for confirmed owner mismatch
- _side_effect_fence unchanged: still the exactly-once save/delivery
  barrier

Tests (all written red-first, verified failing on unfixed code):
slow_owned_delivery (parent handoff), sustained fence-contention grace,
pre/post-delivery None probes, grace-exhausted None probe, terminal
write None, mark_job_run fence-unavailable None, RSI recovery dispatch
fails closed during original's fenced delivery, TrustMRR slow-owned
delivery terminal success. Managed-gateway restart E2E (systemd scope)
passes on host: one gateway, single side effect, single delivery.
@alt-glitch alt-glitch added invalid This doesn't seem right P3 Low — cosmetic, nice to have comp/cron Cron scheduler and job management comp/cli CLI entry point, hermes_cli/, setup wizard labels Sep 4, 2026
@alt-glitch

Copy link
Copy Markdown

This was generated by AI during triage.

This appears to be an integration/handoff artifact that bundles the reviewed heads of #102073, #101629, #102620 and #102627 and states that no upstream merge is requested. Upstream review happens on those four PRs individually; if this PR is meant only as an operational checkpoint, consider keeping it on the fork.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp/cli CLI entry point, hermes_cli/, setup wizard comp/cron Cron scheduler and job management invalid This doesn't seem right P3 Low — cosmetic, nice to have

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants