Skip to content

fix(kanban): inherit parent priority in create_task and decompose_triage_task (P10) - #80427

Open
MarkWin91 wants to merge 5 commits into
NousResearch:mainfrom
MarkWin91:feat/kanban-retriage-on-timeout
Open

MarkWin91 wants to merge 5 commits into
NousResearch:mainfrom
MarkWin91:feat/kanban-retriage-on-timeout

Conversation

@MarkWin91

Copy link
Copy Markdown

Restores P10 (t_5023c583): children now inherit the parent's priority exactly in the specifier.

  • create_task: when priority is None and the task has parents, the child inherits the HIGHEST parent priority. Explicit priority always wins. Parents with NULL priority contribute nothing. Falls back to 0.
  • decompose_triage_task: children inherit the root task's priority exactly (was hardcoded to 0). NULL root -> schema default 0.
  • Transport: kanban_tools.py no longer forces priority to 0; CLI --priority default changed from 0 to None.

Tests: 9 new (7 create_task + 2 decompose), 251 total pass, 0 regression.

Commit: 9199560 (fix) + 4f638aa (test).

MarkWin91 and others added 5 commits July 16, 2026 19:56
…nstead of giving up

A timeout is a deterministic failure: a task that needs more than
max_runtime_seconds fails identically on every blind retry, so the
dispatcher burned failure_limit x max_runtime_seconds of compute before
parking the task in blocked with gave_up — the least productive outcome.

New opt-in config kanban.retriage_on_timeout (default false): when the
circuit breaker trips with trigger outcome timed_out, the task is sent
back to triage with a failure-context block appended to its body and a
fresh failure counter, where the existing auto-decompose pipeline
(kanban.auto_decompose) splits it into smaller children on its next
tick. At most one retriage per task — guarded by the new retriaged
event — so subdivision cannot recurse unbounded; the second trip blocks
normally. Crashes and spawn failures (plausibly transient) and
force_trip callers (own retry policy) keep the existing semantics.

Related: NousResearch#54156 (structured failure annotation + decomposition
suggestion on auto-block — this goes one step further, opt-in).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ry, config warning

Four review findings on the initial implementation:

- The retriage UPDATE now matches status=ready ONLY and reports back:
  if a concurrent dispatcher re-claimed the task (ready -> running) in
  the window between the timeout txn and the failure-recording txn,
  retriage loses the race and falls through to the plain block
  semantics instead of yanking a freshly-spawned live worker off the
  task (the window pre-exists; the block path keeps its behavior).
- retriaged is delivered by the kanban notifier (TERMINAL_KINDS) so a
  subscriber who saw the timed_out event also sees the recovery.
- The gateway warns when retriage_on_timeout is enabled while
  auto_decompose is disabled (retriaged tasks would sit in Triage
  waiting for a manual decompose).
- DispatchResult.retriaged documents that its ids also appear in
  timed_out, so failure-counting consumers do not double-count.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Root cause: _cleanup_workspace (called by complete_task) only cleaned
scratch workspaces — worktree/dir workspaces were preserved but never
auto-committed. Workers that finished their job (build, test, all green)
left their worktree dirty, requiring manual Claude rescue (git add -A && commit)
every time.

Fix:
1. New _cleanup_auto_commit_worktree() helper — resolves the real worktree
   checkout dir from the DB-stored anchor repo path (three strategies:
   linked checkout, <repo>/.worktrees/<task_id> convention, fallback path).
2. New _git_auto_commit_worktree() — checks git status and commits dirty files
   with a descriptive wt/<task_id> message.
3. _cleanup_workspace now calls _cleanup_auto_commit_worktree for worktree tasks.
4. enforce_max_runtime also auto-commits worktree changes before SIGTERM,
   and now includes workspace_kind/workspace_path in its SELECT (was missing).
5. Regression test: test_complete_task_auto_commits_worktree — creates a task,
   dirties a file, completes, verifies commit exists and tree is clean.

Fixes: #t_37ac9048
…age_task

Restores P10 (t_5023c583): children now inherit the parent's priority
exactly in the specifier — a parent created with a priority passes that
priority down to every child.

create_task: when priority is None and the task has parents, the child
inherits the HIGHEST parent priority. Explicit priority always wins.
Parents with NULL priority contribute nothing. Falls back to 0.

decompose_triage_task: children now inherit the root task's priority
exactly (was hardcoded to 0). NULL root → schema default 0.

Transport: kanban_tools.py no longer forces priority to 0 (None reaches
create_task → inheritance), CLI --priority default changed from 0 to None.

Tests: 9 new (7 create_task + 2 decompose), 251 total pass, 0 regression.

Closes t_5023c583.
The fix f02a3c98e made create_task inherit parent priority, but the
transport layer (tools/kanban_tools.py _handle_create) had an 'else 0'
coercion that could zero it out before reaching the DB layer. This test
proves the inheritance survives the full tool-call path: omitted
priority inherits the parent's, explicit priority still wins, parent
untouched.
@alt-glitch alt-glitch added type/bug Something isn't working comp/cron Cron scheduler and job management comp/gateway Gateway runner, session dispatch, delivery P3 Low — cosmetic, nice to have sweeper:risk-compatibility Sweeper risk: may break existing users, config, migrations, defaults, or upgrades labels Aug 6, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp/cron Cron scheduler and job management comp/gateway Gateway runner, session dispatch, delivery P3 Low — cosmetic, nice to have sweeper:risk-compatibility Sweeper risk: may break existing users, config, migrations, defaults, or upgrades type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants