Skip to content

feat(agent): background server, management CLI, and mode-aware goal prompts - #134

Merged
w0wl0lxd merged 34 commits into
mainfrom
feat/agent-cli-simplify
Jul 29, 2026
Merged

w0wl0lxd merged 34 commits into
mainfrom
feat/agent-cli-simplify

Conversation

@w0wl0lxd

@w0wl0lxd w0wl0lxd commented Jul 27, 2026 •

Copy link
Copy Markdown
Owner

Summary

  • Adds n00n agent run --background Unix-socket server for long-lived background agents.
  • Implements n00n agent message, stop, list, status, pause, and resume client commands.
  • Background agent state is persisted under state_dir/agents/<id>/agent.json with a control.sock.
  • Server streams agent events (TextDelta, ToolOutput, Done, Error) as NDJSON and supports pause/resume.
  • CLI AgentMode enum supports General, Research, Task, Team, and Workflow.
  • Adds --goal, --team-mode, --max-agents, --waves, --workflow-inputs, and --task-description to n00n agent run so team/workflow/task modes build the matching tool call automatically.
  • Refactors plugins/team/init.lua and plugins/team/roles.lua to launch subagents through the shared n00n.subagent helper, reducing duplication with plugins/task and plugins/workflow.
  • Introduces shared plugins/lib/n00n/structured_output.lua and plugins/lib/n00n/subagent.lua modules used by task, workflow, and team plugins.
  • Hardens background agent CLI: validates agent_id to prevent path traversal, sets state directory permissions to 0o700, sets the Unix control socket to 0o600, and waits for the agent task to finish before the stop command exits.
  • Threads the thinking option through task and workflow subagent sessions.

Test plan

  • cargo fmt --all
  • cargo check --all --tests
  • cargo clippy --all --tests -- -D warnings
  • cargo nextest run --workspace
  • cargo test -p n00n
  • cargo test -p n00n-lua --test spec

Note

Medium Risk
Research mode and read-only tool filtering change which tools run in production sessions; background agent fork/cancel/stop paths touch long-lived process control and persisted state.

Overview
Introduces AgentMode::Research end-to-end: read-only tool filtering (same path as plan mode via is_readonly()), PromptId::Research for system prompts, session/UI persistence as StoredMode::Research, and CLI research mode that excludes write/edit tools when spawning headless or interactive agents.

plugins/task and plugins/workflow now call n00n.subagent.launch when an output_schema is set; task still uses a direct n00n.agent.session path for the legacy done tool. Task previews move to ActivityPreview. Team sets ctx:set_deadline from timeout_secs.

Background n00n agent run server behavior is tightened: child process fork/setsid, mode stored in agent.json and restored for client messages, pause sends cancel, stop drops input and cancels the interactive task, and error events end the message loop. Headless CLI runs fail with Err when the agent emits an error event.

Reviewed by Cursor Bugbot for commit cf2092b. Bugbot is set up for automated code reviews on this repo. Configure here.

w0wl0lxd added 3 commits July 26, 2026 23:18
Phase 2 implementation:

- Create plugins/lib/n00n/structured_output.lua with constants and helper functions for schema validation
- Create plugins/lib/n00n/subagent.lua with unified subagent launch API
- Migrate plugins/task/init.lua to use n00n.structured_output helper
- Migrate plugins/workflow/init.lua to use n00n.structured_output helper
- Add n00n agent run CLI command with stubs for future daemon commands
- Add tests for structured_output helper in plugins/lib/tests/spec.lua
- Keep team/roles.lua unchanged due to custom patterns (hardcoded audience, budget handling, preview support)

All tests pass. Lua syntax validated with luac.
Implement `n00n agent run --background` Unix-socket server plus
`message`, `stop`, and `list` client commands. Background agents keep
state in `state_dir/agents/<id>/agent.json` and a `control.sock`.

Server streams `TextDelta`/`ToolOutput`/`Done` events as NDJSON, and
supports an initial `--prompt` before entering the command loop. The
`--mode` flag selects the same workflow/team/etc modes used by the
one-shot runner.

- Add `AgentMode` CLI enum (General, Research, Task, Team, Workflow)
- Add `async-lock` for per-agent message serialization
- Add unit tests for CLI parsing, state serde, and state listing

Phase 3 stubs (status, pause, resume, policy) remain unimplemented.
@coderabbitai

coderabbitai Bot commented Jul 27, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

📝 Walkthrough

Walkthrough

Task and workflow plugins now route structured-output execution through n00n.subagent, while task execution without a schema retains its manual session path. Workflow semaphore, progress, and run-guard handling remain in place.

Changes

Subagent routing

Layer / File(s) Summary
Task execution paths
plugins/task/init.lua
Structured-output tasks use subagent.launch; tasks without output_schema retain manual sessions, progress polling, cost attachment, and captured results.
Workflow agent orchestration
plugins/workflow/init.lua, changelog.d/134.changed.md
Workflow agents use subagent.launch while preserving concurrency limits, progress events, run-guard recording, and updated changelog wording.

Estimated code review effort: 4 (Complex) | ~45 minutes

Possibly related PRs

  • w0wl0lxd/n00n#5: Adds progress plumbing used by task execution and previews.
  • w0wl0lxd/n00n#12: Updates task model-tier routing used by subagent launches.
  • w0wl0lxd/n00n#20: Changes structured-output execution and validation behavior in task and workflow plugins.

Poem

A bunny hops where subagents run,
Structured outputs neatly spun.
Tasks keep their manual trail,
Workflows guard each launch without fail.
Progress blooms, and errors clear—
“Hop onward!” whispers the engineer.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 0.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title is concise and captures the broader agent background and mode-aware work described in the PR.
Description check ✅ Passed The description is related to the changeset and includes the task/workflow subagent and thinking-option updates.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/agent-cli-simplify

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions

Copy link
Copy Markdown

Dependency Review

✅ No vulnerabilities or license issues or OpenSSF Scorecard issues found.

Scanned Files

None

w0wl0lxd added 2 commits July 26, 2026 23:49
Add `ClientCommand::Pause` and `Resume` to the background agent
protocol. The server stores a shared `paused` flag and rejects new
messages while paused, updating `agent.json` status accordingly.

Implement `n00n agent status`, `pause`, and `resume` client commands.
`status` reads the persisted agent state; `pause`/`resume` send a
control command over the Unix socket. The `policy` subcommand remains
a stub.
The text-mode client was streaming TextDelta chunks and then printing
the full Done.text again when the run ended with an error. Only print
the error in that case; the streamed text is already on the terminal.

@github-actions github-actions Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Criterion

Details
Benchmark suite Current: cf2092b Previous: d464812 Ratio
fib/jit_mlua_hook 7671510 ns/iter (± 229105) 7284823 ns/iter (± 75675) 1.05
fib/jit_watchdog 1480787 ns/iter (± 36716) 2603040 ns/iter (± 41118) 0.57
fib/jit_none 1477877 ns/iter (± 69862) 2604751 ns/iter (± 36917) 0.57
fib/interp_mlua_hook 8105136 ns/iter (± 53510) 7927475 ns/iter (± 19132) 1.02
fib/interp_watchdog 2898223 ns/iter (± 6957) 3874557 ns/iter (± 81135) 0.75
fib/interp_none 2895160 ns/iter (± 10123) 3816445 ns/iter (± 47350) 0.76
buffer_rw/jit_mlua_hook 735240 ns/iter (± 5791) 572889 ns/iter (± 1758) 1.28
buffer_rw/jit_watchdog 89487 ns/iter (± 229) 167764 ns/iter (± 597) 0.53
buffer_rw/jit_none 89519 ns/iter (± 1443) 167883 ns/iter (± 194) 0.53
buffer_rw/interp_mlua_hook 1086904 ns/iter (± 6387) 1062203 ns/iter (± 20350) 1.02
buffer_rw/interp_watchdog 456111 ns/iter (± 1465) 642735 ns/iter (± 6550) 0.71
buffer_rw/interp_none 456154 ns/iter (± 6815) 640283 ns/iter (± 7979) 0.71
splash_render_120x40 51201 ns/iter (± 3296) 65472 ns/iter (± 6958) 0.78
splash_render_200x60 128923 ns/iter (± 4796) 154383 ns/iter (± 12329) 0.84

This comment was automatically generated by workflow using github-action-benchmark.

@github-actions github-actions Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Performance Alert ⚠️

Possible performance regression was detected for benchmark 'Criterion'.
Benchmark result of this commit is worse than the previous benchmark result exceeding threshold 2.

Benchmark suite Current: 698a7b3 Previous: 5883bbf Ratio
buffer_rw/jit_watchdog 191486 ns/iter (± 313) 75535 ns/iter (± 1409) 2.54
buffer_rw/jit_none 191954 ns/iter (± 473) 76858 ns/iter (± 905) 2.50
splash_render_120x40 80553 ns/iter (± 806) 32557 ns/iter (± 1739) 2.47

This comment was automatically generated by workflow using github-action-benchmark.

@codecov

codecov Bot commented Jul 27, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 20.00000% with 68 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
src/cmd/agent.rs 12.85% 61 Missing ⚠️
n00n-ui/src/app/session.rs 0.00% 2 Missing ⚠️
n00n-acp/src/server.rs 0.00% 1 Missing ⚠️
n00n-agent/src/agent/instructions.rs 75.00% 1 Missing ⚠️
n00n-agent/src/headless.rs 0.00% 1 Missing ⚠️
src/print.rs 0.00% 1 Missing ⚠️
src/sdk_mode.rs 0.00% 1 Missing ⚠️

📢 Thoughts on this report? Let us know!

w0wl0lxd added 10 commits July 27, 2026 00:24
Extend `plugins/lib/n00n/subagent.lua` with options the team plugin needs:
- `system` override for role-specific prompts
- `preview` and `activity_label` for ActivityPreview integration
- `budget` with `:consume()` for agent-call budgets
- `fail_on_pricing_error` to keep the existing team semantics
- return the resolved `model_spec` as a fifth value

Migrate `plugins/team/roles.lua` to launch subagents through the shared
helper, preserving its `{ok, text, cost, model, usage, error}` return
shape and all existing tests.

Simplify `plugins/team/init.lua` `run_supervisor` by delegating the
structured-output plan generation to `n00n.subagent.launch` with
`output_schema = PLANNER_OUTPUT`.
Surface failed socket/directory removals in the background agent stop
handler instead of silently ignoring them. The process still exits, but
the warning gives operators a signal that stale state may remain.
Prevent path traversal from user-supplied agent ids by validating them
before any filesystem operation, and reject ids that contain path
separators, control characters, or other unsafe content. Limit ids to
64 ASCII alphanumeric/hyphen/underscore characters.

Set the agent state directory to 0o700 and the Unix control socket to
0o600 so only the owning user can read or connect to background agent
state.
Extract `prepare_agent_env` and `PreparedEnv` to remove the duplicated
plugin/model/MCP initialization between `run` and `server`. Both paths now
build their `HeadlessParams`/`InteractiveParams` from the same prepared
environment, making the two entry points easier to keep in sync.
Record the user-facing additions, refactor, and security hardening from
PR #134 in changelog.d.
The background agent Stop command now waits for the agent task to finish
after sending the cancel signal, instead of exiting immediately. This
prevents dropping active tool calls and unfinished MCP sessions and
ensures state is persisted before cleanup.
The `goal` flag on `n00n agent run` was parsed but never used, and the
`policy` subcommands were stubs. Remove both until they are wired to
actual behavior.
Thread the `thinking` option through `n00n.agent.session` in both the
task and workflow plugins, and include it in the workflow journal key so
cached results honor the thinking setting.
…tive

`headless::spawn_interactive` was calling `permissions.toggle_yolo()` when
`params.yolo` was true, which flipped an already-true yolo state back to
false. This broke `--yolo` in background agent mode (and TUI/ACP when
always_yolo was set). Add `PermissionManager::set_yolo` and use it to
reliably enable yolo when requested, leaving the config-derived default
otherwise.
Re-add --goal to n00n agent run and wire it into mode-specific prompts.
Team mode now emits a team tool call with goal, mode, max_agents,
waves, and auto-tier flags. Workflow mode emits a workflow tool call
with the prompt as the script and goal merged into inputs. Task mode
emits a task tool call with description/prompt. General/Research
modes prepend the goal to the prompt.

Also introduce AgentRunOptions to keep run/server signatures tidy and
avoid excessive boolean parameters.
@w0wl0lxd w0wl0lxd changed the title feat(agent): background server and management CLI feat(agent): background server, management CLI, and mode-aware goal prompts Jul 27, 2026
w0wl0lxd added 2 commits July 27, 2026 02:10
Remove the hard-coded 24/32 agent-call ceilings in team and workflow.
- team `max_agents` is now configurable with no hard maximum and an
  optional `timeout_secs`.
- workflow `max_agents_per_run` is configurable with no hard maximum.
- Add `plugins/lib/n00n/guard.lua` to enforce call budgets, wall-clock
  timeouts, repeated-prompt loops, and consecutive subagent errors.
- Wire the guard through `n00n.subagent.launch`, `team`, and `workflow`.
…tive errors

- Move the used counter increment from record() into check() so the
  budget is reserved before any async yield; this prevents concurrent
  subagent calls from overshooting max_calls.
- Check consecutive_errors in check() and use > in record() so the
  guard blocks the next call after the threshold instead of wasting
  a token on the failing call.
- Add lib spec coverage for budget, repeated-prompt, consecutive-error,
  and legacy consume() behavior.
@w0wl0lxd
w0wl0lxd marked this pull request as ready for review July 27, 2026 06:48

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 22d4ff09f0

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/cmd/agent.rs Outdated
Comment thread src/cmd/mod.rs
Comment thread src/cmd/agent.rs
Comment thread src/cmd/agent.rs
Comment thread src/cmd/agent.rs
Comment thread src/cmd/agent.rs Outdated
Comment thread src/cmd/agent.rs Outdated
Comment thread src/cmd/agent.rs
Comment thread plugins/team/init.lua
w0wl0lxd added a commit that referenced this pull request Jul 27, 2026
Wire PluginHost ui_action_tx into an in-process daemon listener so CLI
agent list unions TUI sessions while the UI is up. Document stacked
follow-ups for worker absorb (#134) and steer/control (#129).
w0wl0lxd added a commit that referenced this pull request Jul 27, 2026
Unify background worker socks under n00n agent run --background with
daemon-first list/status/message/pause/resume/stop, plus agent daemon
subcommand for headless worker planes.
@w0wl0lxd

Copy link
Copy Markdown
Owner Author

Absorbed into #149 (unified n00n agent CLI + daemon-first control). Prefer closing this once #149 lands, or leave open only if you want an independent review trail.

@w0wl0lxd

Copy link
Copy Markdown
Owner Author

cursor review

@cursor

cursor Bot commented Jul 28, 2026

Copy link
Copy Markdown

Bugbot couldn't run - usage limit reached

Bugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit.

A user or team admin can review and increase usage limits in the Cursor dashboard.

(requestId: serverGenReqId_f8e60ee7-46ac-4537-9ea3-0f5f814aa589)

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 5

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
plugins/task/init.lua (1)

157-177: 📐 Maintainability & Code Quality | 🟠 Major | ⚡ Quick win

Dead local_tools table — remove it.

local_tools is assigned here but never read; the legacy path builds its own done_tool at Line 252, and the structured path delegates to subagent.launch. The handler at Line 172-174 also drops value entirely, making this block a misleading duplicate. Luacheck flags it (W231).

♻️ Remove the unused block
   local preview = make_preview(ctx, input.description or "task")
-
-  -- Build local tools: either structured_output (with schema) or done tool
-  local local_tools
-  if not input.output_schema then
-    local_tools = {
-      [DONE_NAME] = {
-        description = DONE_DESCRIPTION,
-        input_schema = {
-          type = "object",
-          properties = {
-            answer = { type = "string", description = "Final answer to return to the parent agent." },
-          },
-          required = { "answer" },
-        },
-        handler = function(value)
-          return "Done."
-        end,
-      },
-    }
-  end
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@plugins/task/init.lua` around lines 157 - 177, Remove the unused local_tools
declaration and its conditional DONE_NAME handler block from the task
initialization flow after make_preview; retain the existing legacy done_tool
construction and structured subagent.launch path unchanged.

Source: Linters/SAST tools

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@plugins/task/init.lua`:
- Around line 307-318: Update the local assignment in do_poll to discard the
unused second return value from sess:get_progress, using a single progress
variable or an explicit underscore while preserving the existing polling and
completion logic.
- Around line 217-222: Update the return path that builds the `llm_output`
result to capture and validate the error returned by
`n00n.json.encode(captured)`, matching the checked encoding patterns elsewhere
in the handler. Handle encoding failure before returning so `ctx:finish` never
receives a nil `llm_output`.

In `@plugins/workflow/init.lua`:
- Around line 413-421: Update the subagent.launch result handling in the
workflow initialization flow to avoid silently capturing unused cost and
usage_val values: either remove them from the destructuring or propagate them
through existing progress/telemetry, preferably the agent_done event, so
sub-agent token spend remains observable.
- Around line 430-439: Update the completion handling around launch_err so
progress.agent_done and the logger.log("agent_done", ...) telemetry execute only
when launch_err is absent; preserve the existing error("sub-agent error: " ..
launch_err, 0) path for failed launches.
- Around line 441-450: Update the output handling around subagent.launch so
captured string results are stored unchanged, while table results are
JSON-encoded with the existing encode error handling. Use the type of captured
to select the path, preserving the empty fallback for nil values and assigning
the encoded value only for structured table output.

---

Outside diff comments:
In `@plugins/task/init.lua`:
- Around line 157-177: Remove the unused local_tools declaration and its
conditional DONE_NAME handler block from the task initialization flow after
make_preview; retain the existing legacy done_tool construction and structured
subagent.launch path unchanged.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 0dc01059-d475-4f57-82a8-e59938efc332

📥 Commits

Reviewing files that changed from the base of the PR and between effc288 and 483bdcd.

📒 Files selected for processing (3)
  • changelog.d/134.changed.md
  • plugins/task/init.lua
  • plugins/workflow/init.lua
📜 Review details
⏰ Context from checks skipped due to timeout. (12)
  • GitHub Check: Test (macOS)
  • GitHub Check: Test (Windows)
  • GitHub Check: Build (Windows)
  • GitHub Check: MSRV (1.97)
  • GitHub Check: Coverage
  • GitHub Check: Lint (macOS)
  • GitHub Check: Lint
  • GitHub Check: Lint (Windows)
  • GitHub Check: Test
  • GitHub Check: Build
  • GitHub Check: Docs
  • GitHub Check: Analyze (rust)
🧰 Additional context used
📓 Path-based instructions (2)
**/*

📄 CodeRabbit inference engine (AGENTS.md)

**/*: Non-trivial or multi-file changes should use a dedicated git worktree and new branch; do not modify unrelated user changes, force-push, or push to main.
Ship finished work with a clear Conventional Commit message, a pushed branch, and a draft pull request; never add AI-agent attribution to authored content.
Before investigating unfamiliar failures or third-party behavior, research documented behavior first and distinguish unrelated baseline failures from regressions in the touched surface.
Use structural tools before broad searches: prefer codegraph or arbor for cross-file relationships, index before reading files, targeted reads, parallel calls, and code_execution for filtering large outputs.

Files:

  • changelog.d/134.changed.md
  • plugins/workflow/init.lua
  • plugins/task/init.lua
plugins/**/*.lua

📄 CodeRabbit inference engine (AGENTS.md)

Built-in Lua plugins belong under ./plugins and should use the repository's plugin tooling and conventions.

Files:

  • plugins/workflow/init.lua
  • plugins/task/init.lua
🪛 Luacheck (1.2.0)
plugins/workflow/init.lua

[warning] 413-413: unused variable 'cost'

(W211)


[warning] 413-413: unused variable 'usage_val'

(W211)

plugins/task/init.lua

[warning] 160-160: variable 'local_tools' is never accessed

(W231)


[warning] 309-309: unused variable 'err'

(W211)

🪛 markdownlint-cli2 (0.23.0)
changelog.d/134.changed.md

[warning] 1-1: First line in a file should be a top-level heading

(MD041, first-line-heading, first-line-h1)

🔇 Additional comments (4)
plugins/task/init.lua (2)

12-12: LGTM!


223-306: LGTM!

Also applies to: 319-327

plugins/workflow/init.lua (1)

24-24: LGTM!

changelog.d/134.changed.md (1)

1-1: LGTM!

Comment thread plugins/task/init.lua
Comment thread plugins/task/init.lua
Comment thread plugins/workflow/init.lua Outdated
Comment thread plugins/workflow/init.lua
Comment thread plugins/workflow/init.lua
w0wl0lxd added 2 commits July 28, 2026 01:32
- Fix background detach: fork and setsid to daemonize server process
- Fix stop command: close input_tx and await outer task for proper cleanup
- Fix error handling: break event loop on Error to prevent state hang
- Fix task plugin: remove unused error from get_progress, fix encode error handling
- Fix workflow plugin: emit agent_done only on success, remove unused cost/usage returns
@cursor

cursor Bot commented Jul 28, 2026

Copy link
Copy Markdown

Bugbot couldn't run - usage limit reached

Bugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit.

A user or team admin can review and increase usage limits in the Cursor dashboard.

(requestId: serverGenReqId_78dfef73-cbf8-494b-ae25-b3342b839a63)

@w0wl0lxd

Copy link
Copy Markdown
Owner Author

cursor review

@cursor

cursor Bot commented Jul 28, 2026

Copy link
Copy Markdown

Bugbot couldn't run - usage limit reached

Bugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit.

A user or team admin can review and increase usage limits in the Cursor dashboard.

(requestId: serverGenReqId_c081034c-8f4f-45de-bfbb-f8217c1523a9)

w0wl0lxd added 2 commits July 28, 2026 01:45
- Return error on AgentEvent::Error in one-shot run for proper exit code
- Preserve agent state on read failures (only delete on not found)
- Avoid overwriting paused status when active run completes
- UTF-8 truncation already uses character-aware .chars().take()
@cursor

cursor Bot commented Jul 28, 2026

Copy link
Copy Markdown

Bugbot couldn't run - usage limit reached

Bugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit.

A user or team admin can review and increase usage limits in the Cursor dashboard.

(requestId: serverGenReqId_41e21cfe-671c-42ee-839c-e8ae3d69b1bc)

@w0wl0lxd

Copy link
Copy Markdown
Owner Author

cursor review

@cursor

cursor Bot commented Jul 28, 2026

Copy link
Copy Markdown

Bugbot couldn't run - usage limit reached

Bugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit.

A user or team admin can review and increase usage limits in the Cursor dashboard.

(requestId: serverGenReqId_c2b66b2f-d547-4fab-8a69-8226271b2d75)

@cursor

cursor Bot commented Jul 28, 2026

Copy link
Copy Markdown

Bugbot couldn't run - usage limit reached

Bugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit.

A user or team admin can review and increase usage limits in the Cursor dashboard.

(requestId: serverGenReqId_6f4f499a-0e4f-4e75-8caf-f4f725a5a07f)

@w0wl0lxd

Copy link
Copy Markdown
Owner Author

cursor review

@cursor

cursor Bot commented Jul 28, 2026

Copy link
Copy Markdown

Bugbot couldn't run - usage limit reached

Bugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit.

A user or team admin can review and increase usage limits in the Cursor dashboard.

(requestId: serverGenReqId_4e14ddea-1325-4b7e-8e38-036d7bc38f57)

w0wl0lxd added 2 commits July 28, 2026 04:53
…ctions

- Add AgentMode::Research and is_readonly() so build_system_prompt picks
the research prompt and filter_tools_for_mode strips execute-kind tools.
- Plumb runtime AgentMode through HeadlessParams/InteractiveParams and
persist it in background AgentState so follow-up messages keep research
mode.
- Exclude write/edit/multiedit/edit_lines/insert_lines for research runs.
- Fix task plugin preview to use ActivityPreview (required by n00n.subagent)
and return raw structured-output validation errors for tests.
@cursor

cursor Bot commented Jul 28, 2026

Copy link
Copy Markdown

Bugbot couldn't run - usage limit reached

Bugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit.

A user or team admin can review and increase usage limits in the Cursor dashboard.

(requestId: serverGenReqId_81fc75de-b42a-4ec6-844b-03c9cb07562d)

…tables

Subagent.launch may return either a plain string or a structured table.
Store strings unchanged, JSON-encode tables with error handling, and use an
empty string for nil output instead of encoding nil.
@cursor

cursor Bot commented Jul 28, 2026

Copy link
Copy Markdown

Bugbot couldn't run - usage limit reached

Bugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit.

A user or team admin can review and increase usage limits in the Cursor dashboard.

(requestId: serverGenReqId_726566c2-a243-4022-b30d-fdbecb0c86b0)

@w0wl0lxd
w0wl0lxd merged commit 6ebcd45 into main Jul 29, 2026
32 checks passed
@linear-code

linear-code Bot commented Jul 29, 2026

Copy link
Copy Markdown

N00N-92

w0wl0lxd added a commit that referenced this pull request Jul 29, 2026
…rompts (#134)

* feat(agent): add structured output helper and agent CLI command

Phase 2 implementation:

- Create plugins/lib/n00n/structured_output.lua with constants and helper functions for schema validation
- Create plugins/lib/n00n/subagent.lua with unified subagent launch API
- Migrate plugins/task/init.lua to use n00n.structured_output helper
- Migrate plugins/workflow/init.lua to use n00n.structured_output helper
- Add n00n agent run CLI command with stubs for future daemon commands
- Add tests for structured_output helper in plugins/lib/tests/spec.lua
- Keep team/roles.lua unchanged due to custom patterns (hardcoded audience, budget handling, preview support)

All tests pass. Lua syntax validated with luac.

* fix(agent-orchestration): repair route_tier and structured_output helpers

* feat(agent): add background agent server and management CLI

Implement `n00n agent run --background` Unix-socket server plus
`message`, `stop`, and `list` client commands. Background agents keep
state in `state_dir/agents/<id>/agent.json` and a `control.sock`.

Server streams `TextDelta`/`ToolOutput`/`Done` events as NDJSON, and
supports an initial `--prompt` before entering the command loop. The
`--mode` flag selects the same workflow/team/etc modes used by the
one-shot runner.

- Add `AgentMode` CLI enum (General, Research, Task, Team, Workflow)
- Add `async-lock` for per-agent message serialization
- Add unit tests for CLI parsing, state serde, and state listing

Phase 3 stubs (status, pause, resume, policy) remain unimplemented.

* feat(agent): implement status, pause, and resume commands

Add `ClientCommand::Pause` and `Resume` to the background agent
protocol. The server stores a shared `paused` flag and rejects new
messages while paused, updating `agent.json` status accordingly.

Implement `n00n agent status`, `pause`, and `resume` client commands.
`status` reads the persisted agent state; `pause`/`resume` send a
control command over the Unix socket. The `policy` subcommand remains
a stub.

* fix(agent): avoid duplicate text when message run ends with error

The text-mode client was streaming TextDelta chunks and then printing
the full Done.text again when the run ended with an error. Only print
the error in that case; the streamed text is already on the terminal.

* refactor(team): migrate roles and supervisor to n00n.subagent

Extend `plugins/lib/n00n/subagent.lua` with options the team plugin needs:
- `system` override for role-specific prompts
- `preview` and `activity_label` for ActivityPreview integration
- `budget` with `:consume()` for agent-call budgets
- `fail_on_pricing_error` to keep the existing team semantics
- return the resolved `model_spec` as a fifth value

Migrate `plugins/team/roles.lua` to launch subagents through the shared
helper, preserving its `{ok, text, cost, model, usage, error}` return
shape and all existing tests.

Simplify `plugins/team/init.lua` `run_supervisor` by delegating the
structured-output plan generation to `n00n.subagent.launch` with
`output_schema = PLANNER_OUTPUT`.

* fix(agent): log cleanup warnings in stop handler

Surface failed socket/directory removals in the background agent stop
handler instead of silently ignoring them. The process still exits, but
the warning gives operators a signal that stale state may remain.

* fix(agent): validate agent ids and lock down socket permissions

Prevent path traversal from user-supplied agent ids by validating them
before any filesystem operation, and reject ids that contain path
separators, control characters, or other unsafe content. Limit ids to
64 ASCII alphanumeric/hyphen/underscore characters.

Set the agent state directory to 0o700 and the Unix control socket to
0o600 so only the owning user can read or connect to background agent
state.

* refactor(agent): share one-shot and background agent setup

Extract `prepare_agent_env` and `PreparedEnv` to remove the duplicated
plugin/model/MCP initialization between `run` and `server`. Both paths now
build their `HeadlessParams`/`InteractiveParams` from the same prepared
environment, making the two entry points easier to keep in sync.

* docs(changelog): add fragments for agent CLI and team refactor

Record the user-facing additions, refactor, and security hardening from
PR #134 in changelog.d.

* fix(agent): await task completion before stop exits

The background agent Stop command now waits for the agent task to finish
after sending the cancel signal, instead of exiting immediately. This
prevents dropping active tool calls and unfinished MCP sessions and
ensures state is persisted before cleanup.

* refactor(cli): remove unused goal arg and stub policy commands

The `goal` flag on `n00n agent run` was parsed but never used, and the
`policy` subcommands were stubs. Remove both until they are wired to
actual behavior.

* fix(task,workflow): pass thinking config to subagent sessions

Thread the `thinking` option through `n00n.agent.session` in both the
task and workflow plugins, and include it in the workflow journal key so
cached results honor the thinking setting.

* fix(n00n-agent): make yolo enable rather than toggle in spawn_interactive

`headless::spawn_interactive` was calling `permissions.toggle_yolo()` when
`params.yolo` was true, which flipped an already-true yolo state back to
false. This broke `--yolo` in background agent mode (and TUI/ACP when
always_yolo was set). Add `PermissionManager::set_yolo` and use it to
reliably enable yolo when requested, leaving the config-derived default
otherwise.

* feat(agent): add --goal and mode-aware team/workflow/task prompts

Re-add --goal to n00n agent run and wire it into mode-specific prompts.
Team mode now emits a team tool call with goal, mode, max_agents,
waves, and auto-tier flags. Workflow mode emits a workflow tool call
with the prompt as the script and goal merged into inputs. Task mode
emits a task tool call with description/prompt. General/Research
modes prepend the goal to the prompt.

Also introduce AgentRunOptions to keep run/server signatures tidy and
avoid excessive boolean parameters.

* feat(agent): configurable agent-call limits and runaway guard

Remove the hard-coded 24/32 agent-call ceilings in team and workflow.
- team `max_agents` is now configurable with no hard maximum and an
  optional `timeout_secs`.
- workflow `max_agents_per_run` is configurable with no hard maximum.
- Add `plugins/lib/n00n/guard.lua` to enforce call budgets, wall-clock
  timeouts, repeated-prompt loops, and consecutive subagent errors.
- Wire the guard through `n00n.subagent.launch`, `team`, and `workflow`.

* fix(guard): reserve call budget atomically in check and guard consecutive errors

- Move the used counter increment from record() into check() so the
  budget is reserved before any async yield; this prevents concurrent
  subagent calls from overshooting max_calls.
- Check consecutive_errors in check() and use > in record() so the
  guard blocks the next call after the threshold instead of wasting
  a token on the failing call.
- Add lib spec coverage for budget, repeated-prompt, consecutive-error,
  and legacy consume() behavior.

* fix(agent): gate background agent IPC behind Unix-only cfg

Background agents use Unix domain sockets and permission modes that are
not available on Windows. Return a clear error on unsupported platforms
so the crate builds on all CI targets.

* fix(nix): use platform-correct loader env var in wrappers

Use DYLD_LIBRARY_PATH on Darwin and LD_LIBRARY_PATH elsewhere, and wrap the computed package binary path to avoid runtime loader failures on non-Linux builds.

* refactor(plugins): migrate task and workflow to n00n.subagent.launch for structured output

* fix(agent): resolve P1 issues in background server and error handling

- Fix background detach: fork and setsid to daemonize server process
- Fix stop command: close input_tx and await outer task for proper cleanup
- Fix error handling: break event loop on Error to prevent state hang
- Fix task plugin: remove unused error from get_progress, fix encode error handling
- Fix workflow plugin: emit agent_done only on success, remove unused cost/usage returns

* fix(agent): address P2 review feedback in agent.rs

- Return error on AgentEvent::Error in one-shot run for proper exit code
- Preserve agent state on read failures (only delete on not found)
- Avoid overwriting paused status when active run completes
- UTF-8 truncation already uses character-aware .chars().take()

* fix(team): enforce wall-clock timeout during subagent calls

* feat(agent): wire research mode through system prompt and tool restrictions

- Add AgentMode::Research and is_readonly() so build_system_prompt picks
the research prompt and filter_tools_for_mode strips execute-kind tools.
- Plumb runtime AgentMode through HeadlessParams/InteractiveParams and
persist it in background AgentState so follow-up messages keep research
mode.
- Exclude write/edit/multiedit/edit_lines/insert_lines for research runs.
- Fix task plugin preview to use ActivityPreview (required by n00n.subagent)
and return raw structured-output validation errors for tests.

* fix(workflow): preserve string subagent output and encode structured tables

Subagent.launch may return either a plain string or a structured table.
Store strings unchanged, JSON-encode tables with error handling, and use an
empty string for nil output instead of encoding nil.
w0wl0lxd added a commit that referenced this pull request Jul 29, 2026
…rompts (#134)

* feat(agent): add structured output helper and agent CLI command

Phase 2 implementation:

- Create plugins/lib/n00n/structured_output.lua with constants and helper functions for schema validation
- Create plugins/lib/n00n/subagent.lua with unified subagent launch API
- Migrate plugins/task/init.lua to use n00n.structured_output helper
- Migrate plugins/workflow/init.lua to use n00n.structured_output helper
- Add n00n agent run CLI command with stubs for future daemon commands
- Add tests for structured_output helper in plugins/lib/tests/spec.lua
- Keep team/roles.lua unchanged due to custom patterns (hardcoded audience, budget handling, preview support)

All tests pass. Lua syntax validated with luac.

* fix(agent-orchestration): repair route_tier and structured_output helpers

* feat(agent): add background agent server and management CLI

Implement `n00n agent run --background` Unix-socket server plus
`message`, `stop`, and `list` client commands. Background agents keep
state in `state_dir/agents/<id>/agent.json` and a `control.sock`.

Server streams `TextDelta`/`ToolOutput`/`Done` events as NDJSON, and
supports an initial `--prompt` before entering the command loop. The
`--mode` flag selects the same workflow/team/etc modes used by the
one-shot runner.

- Add `AgentMode` CLI enum (General, Research, Task, Team, Workflow)
- Add `async-lock` for per-agent message serialization
- Add unit tests for CLI parsing, state serde, and state listing

Phase 3 stubs (status, pause, resume, policy) remain unimplemented.

* feat(agent): implement status, pause, and resume commands

Add `ClientCommand::Pause` and `Resume` to the background agent
protocol. The server stores a shared `paused` flag and rejects new
messages while paused, updating `agent.json` status accordingly.

Implement `n00n agent status`, `pause`, and `resume` client commands.
`status` reads the persisted agent state; `pause`/`resume` send a
control command over the Unix socket. The `policy` subcommand remains
a stub.

* fix(agent): avoid duplicate text when message run ends with error

The text-mode client was streaming TextDelta chunks and then printing
the full Done.text again when the run ended with an error. Only print
the error in that case; the streamed text is already on the terminal.

* refactor(team): migrate roles and supervisor to n00n.subagent

Extend `plugins/lib/n00n/subagent.lua` with options the team plugin needs:
- `system` override for role-specific prompts
- `preview` and `activity_label` for ActivityPreview integration
- `budget` with `:consume()` for agent-call budgets
- `fail_on_pricing_error` to keep the existing team semantics
- return the resolved `model_spec` as a fifth value

Migrate `plugins/team/roles.lua` to launch subagents through the shared
helper, preserving its `{ok, text, cost, model, usage, error}` return
shape and all existing tests.

Simplify `plugins/team/init.lua` `run_supervisor` by delegating the
structured-output plan generation to `n00n.subagent.launch` with
`output_schema = PLANNER_OUTPUT`.

* fix(agent): log cleanup warnings in stop handler

Surface failed socket/directory removals in the background agent stop
handler instead of silently ignoring them. The process still exits, but
the warning gives operators a signal that stale state may remain.

* fix(agent): validate agent ids and lock down socket permissions

Prevent path traversal from user-supplied agent ids by validating them
before any filesystem operation, and reject ids that contain path
separators, control characters, or other unsafe content. Limit ids to
64 ASCII alphanumeric/hyphen/underscore characters.

Set the agent state directory to 0o700 and the Unix control socket to
0o600 so only the owning user can read or connect to background agent
state.

* refactor(agent): share one-shot and background agent setup

Extract `prepare_agent_env` and `PreparedEnv` to remove the duplicated
plugin/model/MCP initialization between `run` and `server`. Both paths now
build their `HeadlessParams`/`InteractiveParams` from the same prepared
environment, making the two entry points easier to keep in sync.

* docs(changelog): add fragments for agent CLI and team refactor

Record the user-facing additions, refactor, and security hardening from
PR #134 in changelog.d.

* fix(agent): await task completion before stop exits

The background agent Stop command now waits for the agent task to finish
after sending the cancel signal, instead of exiting immediately. This
prevents dropping active tool calls and unfinished MCP sessions and
ensures state is persisted before cleanup.

* refactor(cli): remove unused goal arg and stub policy commands

The `goal` flag on `n00n agent run` was parsed but never used, and the
`policy` subcommands were stubs. Remove both until they are wired to
actual behavior.

* fix(task,workflow): pass thinking config to subagent sessions

Thread the `thinking` option through `n00n.agent.session` in both the
task and workflow plugins, and include it in the workflow journal key so
cached results honor the thinking setting.

* fix(n00n-agent): make yolo enable rather than toggle in spawn_interactive

`headless::spawn_interactive` was calling `permissions.toggle_yolo()` when
`params.yolo` was true, which flipped an already-true yolo state back to
false. This broke `--yolo` in background agent mode (and TUI/ACP when
always_yolo was set). Add `PermissionManager::set_yolo` and use it to
reliably enable yolo when requested, leaving the config-derived default
otherwise.

* feat(agent): add --goal and mode-aware team/workflow/task prompts

Re-add --goal to n00n agent run and wire it into mode-specific prompts.
Team mode now emits a team tool call with goal, mode, max_agents,
waves, and auto-tier flags. Workflow mode emits a workflow tool call
with the prompt as the script and goal merged into inputs. Task mode
emits a task tool call with description/prompt. General/Research
modes prepend the goal to the prompt.

Also introduce AgentRunOptions to keep run/server signatures tidy and
avoid excessive boolean parameters.

* feat(agent): configurable agent-call limits and runaway guard

Remove the hard-coded 24/32 agent-call ceilings in team and workflow.
- team `max_agents` is now configurable with no hard maximum and an
  optional `timeout_secs`.
- workflow `max_agents_per_run` is configurable with no hard maximum.
- Add `plugins/lib/n00n/guard.lua` to enforce call budgets, wall-clock
  timeouts, repeated-prompt loops, and consecutive subagent errors.
- Wire the guard through `n00n.subagent.launch`, `team`, and `workflow`.

* fix(guard): reserve call budget atomically in check and guard consecutive errors

- Move the used counter increment from record() into check() so the
  budget is reserved before any async yield; this prevents concurrent
  subagent calls from overshooting max_calls.
- Check consecutive_errors in check() and use > in record() so the
  guard blocks the next call after the threshold instead of wasting
  a token on the failing call.
- Add lib spec coverage for budget, repeated-prompt, consecutive-error,
  and legacy consume() behavior.

* fix(agent): gate background agent IPC behind Unix-only cfg

Background agents use Unix domain sockets and permission modes that are
not available on Windows. Return a clear error on unsupported platforms
so the crate builds on all CI targets.

* fix(nix): use platform-correct loader env var in wrappers

Use DYLD_LIBRARY_PATH on Darwin and LD_LIBRARY_PATH elsewhere, and wrap the computed package binary path to avoid runtime loader failures on non-Linux builds.

* refactor(plugins): migrate task and workflow to n00n.subagent.launch for structured output

* fix(agent): resolve P1 issues in background server and error handling

- Fix background detach: fork and setsid to daemonize server process
- Fix stop command: close input_tx and await outer task for proper cleanup
- Fix error handling: break event loop on Error to prevent state hang
- Fix task plugin: remove unused error from get_progress, fix encode error handling
- Fix workflow plugin: emit agent_done only on success, remove unused cost/usage returns

* fix(agent): address P2 review feedback in agent.rs

- Return error on AgentEvent::Error in one-shot run for proper exit code
- Preserve agent state on read failures (only delete on not found)
- Avoid overwriting paused status when active run completes
- UTF-8 truncation already uses character-aware .chars().take()

* fix(team): enforce wall-clock timeout during subagent calls

* feat(agent): wire research mode through system prompt and tool restrictions

- Add AgentMode::Research and is_readonly() so build_system_prompt picks
the research prompt and filter_tools_for_mode strips execute-kind tools.
- Plumb runtime AgentMode through HeadlessParams/InteractiveParams and
persist it in background AgentState so follow-up messages keep research
mode.
- Exclude write/edit/multiedit/edit_lines/insert_lines for research runs.
- Fix task plugin preview to use ActivityPreview (required by n00n.subagent)
and return raw structured-output validation errors for tests.

* fix(workflow): preserve string subagent output and encode structured tables

Subagent.launch may return either a plain string or a structured table.
Store strings unchanged, JSON-encode tables with error handling, and use an
empty string for nil output instead of encoding nil.
w0wl0lxd added a commit that referenced this pull request Jul 29, 2026
…rompts (#134)

* feat(agent): add structured output helper and agent CLI command

Phase 2 implementation:

- Create plugins/lib/n00n/structured_output.lua with constants and helper functions for schema validation
- Create plugins/lib/n00n/subagent.lua with unified subagent launch API
- Migrate plugins/task/init.lua to use n00n.structured_output helper
- Migrate plugins/workflow/init.lua to use n00n.structured_output helper
- Add n00n agent run CLI command with stubs for future daemon commands
- Add tests for structured_output helper in plugins/lib/tests/spec.lua
- Keep team/roles.lua unchanged due to custom patterns (hardcoded audience, budget handling, preview support)

All tests pass. Lua syntax validated with luac.

* fix(agent-orchestration): repair route_tier and structured_output helpers

* feat(agent): add background agent server and management CLI

Implement `n00n agent run --background` Unix-socket server plus
`message`, `stop`, and `list` client commands. Background agents keep
state in `state_dir/agents/<id>/agent.json` and a `control.sock`.

Server streams `TextDelta`/`ToolOutput`/`Done` events as NDJSON, and
supports an initial `--prompt` before entering the command loop. The
`--mode` flag selects the same workflow/team/etc modes used by the
one-shot runner.

- Add `AgentMode` CLI enum (General, Research, Task, Team, Workflow)
- Add `async-lock` for per-agent message serialization
- Add unit tests for CLI parsing, state serde, and state listing

Phase 3 stubs (status, pause, resume, policy) remain unimplemented.

* feat(agent): implement status, pause, and resume commands

Add `ClientCommand::Pause` and `Resume` to the background agent
protocol. The server stores a shared `paused` flag and rejects new
messages while paused, updating `agent.json` status accordingly.

Implement `n00n agent status`, `pause`, and `resume` client commands.
`status` reads the persisted agent state; `pause`/`resume` send a
control command over the Unix socket. The `policy` subcommand remains
a stub.

* fix(agent): avoid duplicate text when message run ends with error

The text-mode client was streaming TextDelta chunks and then printing
the full Done.text again when the run ended with an error. Only print
the error in that case; the streamed text is already on the terminal.

* refactor(team): migrate roles and supervisor to n00n.subagent

Extend `plugins/lib/n00n/subagent.lua` with options the team plugin needs:
- `system` override for role-specific prompts
- `preview` and `activity_label` for ActivityPreview integration
- `budget` with `:consume()` for agent-call budgets
- `fail_on_pricing_error` to keep the existing team semantics
- return the resolved `model_spec` as a fifth value

Migrate `plugins/team/roles.lua` to launch subagents through the shared
helper, preserving its `{ok, text, cost, model, usage, error}` return
shape and all existing tests.

Simplify `plugins/team/init.lua` `run_supervisor` by delegating the
structured-output plan generation to `n00n.subagent.launch` with
`output_schema = PLANNER_OUTPUT`.

* fix(agent): log cleanup warnings in stop handler

Surface failed socket/directory removals in the background agent stop
handler instead of silently ignoring them. The process still exits, but
the warning gives operators a signal that stale state may remain.

* fix(agent): validate agent ids and lock down socket permissions

Prevent path traversal from user-supplied agent ids by validating them
before any filesystem operation, and reject ids that contain path
separators, control characters, or other unsafe content. Limit ids to
64 ASCII alphanumeric/hyphen/underscore characters.

Set the agent state directory to 0o700 and the Unix control socket to
0o600 so only the owning user can read or connect to background agent
state.

* refactor(agent): share one-shot and background agent setup

Extract `prepare_agent_env` and `PreparedEnv` to remove the duplicated
plugin/model/MCP initialization between `run` and `server`. Both paths now
build their `HeadlessParams`/`InteractiveParams` from the same prepared
environment, making the two entry points easier to keep in sync.

* docs(changelog): add fragments for agent CLI and team refactor

Record the user-facing additions, refactor, and security hardening from
PR #134 in changelog.d.

* fix(agent): await task completion before stop exits

The background agent Stop command now waits for the agent task to finish
after sending the cancel signal, instead of exiting immediately. This
prevents dropping active tool calls and unfinished MCP sessions and
ensures state is persisted before cleanup.

* refactor(cli): remove unused goal arg and stub policy commands

The `goal` flag on `n00n agent run` was parsed but never used, and the
`policy` subcommands were stubs. Remove both until they are wired to
actual behavior.

* fix(task,workflow): pass thinking config to subagent sessions

Thread the `thinking` option through `n00n.agent.session` in both the
task and workflow plugins, and include it in the workflow journal key so
cached results honor the thinking setting.

* fix(n00n-agent): make yolo enable rather than toggle in spawn_interactive

`headless::spawn_interactive` was calling `permissions.toggle_yolo()` when
`params.yolo` was true, which flipped an already-true yolo state back to
false. This broke `--yolo` in background agent mode (and TUI/ACP when
always_yolo was set). Add `PermissionManager::set_yolo` and use it to
reliably enable yolo when requested, leaving the config-derived default
otherwise.

* feat(agent): add --goal and mode-aware team/workflow/task prompts

Re-add --goal to n00n agent run and wire it into mode-specific prompts.
Team mode now emits a team tool call with goal, mode, max_agents,
waves, and auto-tier flags. Workflow mode emits a workflow tool call
with the prompt as the script and goal merged into inputs. Task mode
emits a task tool call with description/prompt. General/Research
modes prepend the goal to the prompt.

Also introduce AgentRunOptions to keep run/server signatures tidy and
avoid excessive boolean parameters.

* feat(agent): configurable agent-call limits and runaway guard

Remove the hard-coded 24/32 agent-call ceilings in team and workflow.
- team `max_agents` is now configurable with no hard maximum and an
  optional `timeout_secs`.
- workflow `max_agents_per_run` is configurable with no hard maximum.
- Add `plugins/lib/n00n/guard.lua` to enforce call budgets, wall-clock
  timeouts, repeated-prompt loops, and consecutive subagent errors.
- Wire the guard through `n00n.subagent.launch`, `team`, and `workflow`.

* fix(guard): reserve call budget atomically in check and guard consecutive errors

- Move the used counter increment from record() into check() so the
  budget is reserved before any async yield; this prevents concurrent
  subagent calls from overshooting max_calls.
- Check consecutive_errors in check() and use > in record() so the
  guard blocks the next call after the threshold instead of wasting
  a token on the failing call.
- Add lib spec coverage for budget, repeated-prompt, consecutive-error,
  and legacy consume() behavior.

* fix(agent): gate background agent IPC behind Unix-only cfg

Background agents use Unix domain sockets and permission modes that are
not available on Windows. Return a clear error on unsupported platforms
so the crate builds on all CI targets.

* fix(nix): use platform-correct loader env var in wrappers

Use DYLD_LIBRARY_PATH on Darwin and LD_LIBRARY_PATH elsewhere, and wrap the computed package binary path to avoid runtime loader failures on non-Linux builds.

* refactor(plugins): migrate task and workflow to n00n.subagent.launch for structured output

* fix(agent): resolve P1 issues in background server and error handling

- Fix background detach: fork and setsid to daemonize server process
- Fix stop command: close input_tx and await outer task for proper cleanup
- Fix error handling: break event loop on Error to prevent state hang
- Fix task plugin: remove unused error from get_progress, fix encode error handling
- Fix workflow plugin: emit agent_done only on success, remove unused cost/usage returns

* fix(agent): address P2 review feedback in agent.rs

- Return error on AgentEvent::Error in one-shot run for proper exit code
- Preserve agent state on read failures (only delete on not found)
- Avoid overwriting paused status when active run completes
- UTF-8 truncation already uses character-aware .chars().take()

* fix(team): enforce wall-clock timeout during subagent calls

* feat(agent): wire research mode through system prompt and tool restrictions

- Add AgentMode::Research and is_readonly() so build_system_prompt picks
the research prompt and filter_tools_for_mode strips execute-kind tools.
- Plumb runtime AgentMode through HeadlessParams/InteractiveParams and
persist it in background AgentState so follow-up messages keep research
mode.
- Exclude write/edit/multiedit/edit_lines/insert_lines for research runs.
- Fix task plugin preview to use ActivityPreview (required by n00n.subagent)
and return raw structured-output validation errors for tests.

* fix(workflow): preserve string subagent output and encode structured tables

Subagent.launch may return either a plain string or a structured table.
Store strings unchanged, JSON-encode tables with error handling, and use an
empty string for nil output instead of encoding nil.
w0wl0lxd added a commit that referenced this pull request Jul 29, 2026
…rompts (#134)

* feat(agent): add structured output helper and agent CLI command

Phase 2 implementation:

- Create plugins/lib/n00n/structured_output.lua with constants and helper functions for schema validation
- Create plugins/lib/n00n/subagent.lua with unified subagent launch API
- Migrate plugins/task/init.lua to use n00n.structured_output helper
- Migrate plugins/workflow/init.lua to use n00n.structured_output helper
- Add n00n agent run CLI command with stubs for future daemon commands
- Add tests for structured_output helper in plugins/lib/tests/spec.lua
- Keep team/roles.lua unchanged due to custom patterns (hardcoded audience, budget handling, preview support)

All tests pass. Lua syntax validated with luac.

* fix(agent-orchestration): repair route_tier and structured_output helpers

* feat(agent): add background agent server and management CLI

Implement `n00n agent run --background` Unix-socket server plus
`message`, `stop`, and `list` client commands. Background agents keep
state in `state_dir/agents/<id>/agent.json` and a `control.sock`.

Server streams `TextDelta`/`ToolOutput`/`Done` events as NDJSON, and
supports an initial `--prompt` before entering the command loop. The
`--mode` flag selects the same workflow/team/etc modes used by the
one-shot runner.

- Add `AgentMode` CLI enum (General, Research, Task, Team, Workflow)
- Add `async-lock` for per-agent message serialization
- Add unit tests for CLI parsing, state serde, and state listing

Phase 3 stubs (status, pause, resume, policy) remain unimplemented.

* feat(agent): implement status, pause, and resume commands

Add `ClientCommand::Pause` and `Resume` to the background agent
protocol. The server stores a shared `paused` flag and rejects new
messages while paused, updating `agent.json` status accordingly.

Implement `n00n agent status`, `pause`, and `resume` client commands.
`status` reads the persisted agent state; `pause`/`resume` send a
control command over the Unix socket. The `policy` subcommand remains
a stub.

* fix(agent): avoid duplicate text when message run ends with error

The text-mode client was streaming TextDelta chunks and then printing
the full Done.text again when the run ended with an error. Only print
the error in that case; the streamed text is already on the terminal.

* refactor(team): migrate roles and supervisor to n00n.subagent

Extend `plugins/lib/n00n/subagent.lua` with options the team plugin needs:
- `system` override for role-specific prompts
- `preview` and `activity_label` for ActivityPreview integration
- `budget` with `:consume()` for agent-call budgets
- `fail_on_pricing_error` to keep the existing team semantics
- return the resolved `model_spec` as a fifth value

Migrate `plugins/team/roles.lua` to launch subagents through the shared
helper, preserving its `{ok, text, cost, model, usage, error}` return
shape and all existing tests.

Simplify `plugins/team/init.lua` `run_supervisor` by delegating the
structured-output plan generation to `n00n.subagent.launch` with
`output_schema = PLANNER_OUTPUT`.

* fix(agent): log cleanup warnings in stop handler

Surface failed socket/directory removals in the background agent stop
handler instead of silently ignoring them. The process still exits, but
the warning gives operators a signal that stale state may remain.

* fix(agent): validate agent ids and lock down socket permissions

Prevent path traversal from user-supplied agent ids by validating them
before any filesystem operation, and reject ids that contain path
separators, control characters, or other unsafe content. Limit ids to
64 ASCII alphanumeric/hyphen/underscore characters.

Set the agent state directory to 0o700 and the Unix control socket to
0o600 so only the owning user can read or connect to background agent
state.

* refactor(agent): share one-shot and background agent setup

Extract `prepare_agent_env` and `PreparedEnv` to remove the duplicated
plugin/model/MCP initialization between `run` and `server`. Both paths now
build their `HeadlessParams`/`InteractiveParams` from the same prepared
environment, making the two entry points easier to keep in sync.

* docs(changelog): add fragments for agent CLI and team refactor

Record the user-facing additions, refactor, and security hardening from
PR #134 in changelog.d.

* fix(agent): await task completion before stop exits

The background agent Stop command now waits for the agent task to finish
after sending the cancel signal, instead of exiting immediately. This
prevents dropping active tool calls and unfinished MCP sessions and
ensures state is persisted before cleanup.

* refactor(cli): remove unused goal arg and stub policy commands

The `goal` flag on `n00n agent run` was parsed but never used, and the
`policy` subcommands were stubs. Remove both until they are wired to
actual behavior.

* fix(task,workflow): pass thinking config to subagent sessions

Thread the `thinking` option through `n00n.agent.session` in both the
task and workflow plugins, and include it in the workflow journal key so
cached results honor the thinking setting.

* fix(n00n-agent): make yolo enable rather than toggle in spawn_interactive

`headless::spawn_interactive` was calling `permissions.toggle_yolo()` when
`params.yolo` was true, which flipped an already-true yolo state back to
false. This broke `--yolo` in background agent mode (and TUI/ACP when
always_yolo was set). Add `PermissionManager::set_yolo` and use it to
reliably enable yolo when requested, leaving the config-derived default
otherwise.

* feat(agent): add --goal and mode-aware team/workflow/task prompts

Re-add --goal to n00n agent run and wire it into mode-specific prompts.
Team mode now emits a team tool call with goal, mode, max_agents,
waves, and auto-tier flags. Workflow mode emits a workflow tool call
with the prompt as the script and goal merged into inputs. Task mode
emits a task tool call with description/prompt. General/Research
modes prepend the goal to the prompt.

Also introduce AgentRunOptions to keep run/server signatures tidy and
avoid excessive boolean parameters.

* feat(agent): configurable agent-call limits and runaway guard

Remove the hard-coded 24/32 agent-call ceilings in team and workflow.
- team `max_agents` is now configurable with no hard maximum and an
  optional `timeout_secs`.
- workflow `max_agents_per_run` is configurable with no hard maximum.
- Add `plugins/lib/n00n/guard.lua` to enforce call budgets, wall-clock
  timeouts, repeated-prompt loops, and consecutive subagent errors.
- Wire the guard through `n00n.subagent.launch`, `team`, and `workflow`.

* fix(guard): reserve call budget atomically in check and guard consecutive errors

- Move the used counter increment from record() into check() so the
  budget is reserved before any async yield; this prevents concurrent
  subagent calls from overshooting max_calls.
- Check consecutive_errors in check() and use > in record() so the
  guard blocks the next call after the threshold instead of wasting
  a token on the failing call.
- Add lib spec coverage for budget, repeated-prompt, consecutive-error,
  and legacy consume() behavior.

* fix(agent): gate background agent IPC behind Unix-only cfg

Background agents use Unix domain sockets and permission modes that are
not available on Windows. Return a clear error on unsupported platforms
so the crate builds on all CI targets.

* fix(nix): use platform-correct loader env var in wrappers

Use DYLD_LIBRARY_PATH on Darwin and LD_LIBRARY_PATH elsewhere, and wrap the computed package binary path to avoid runtime loader failures on non-Linux builds.

* refactor(plugins): migrate task and workflow to n00n.subagent.launch for structured output

* fix(agent): resolve P1 issues in background server and error handling

- Fix background detach: fork and setsid to daemonize server process
- Fix stop command: close input_tx and await outer task for proper cleanup
- Fix error handling: break event loop on Error to prevent state hang
- Fix task plugin: remove unused error from get_progress, fix encode error handling
- Fix workflow plugin: emit agent_done only on success, remove unused cost/usage returns

* fix(agent): address P2 review feedback in agent.rs

- Return error on AgentEvent::Error in one-shot run for proper exit code
- Preserve agent state on read failures (only delete on not found)
- Avoid overwriting paused status when active run completes
- UTF-8 truncation already uses character-aware .chars().take()

* fix(team): enforce wall-clock timeout during subagent calls

* feat(agent): wire research mode through system prompt and tool restrictions

- Add AgentMode::Research and is_readonly() so build_system_prompt picks
the research prompt and filter_tools_for_mode strips execute-kind tools.
- Plumb runtime AgentMode through HeadlessParams/InteractiveParams and
persist it in background AgentState so follow-up messages keep research
mode.
- Exclude write/edit/multiedit/edit_lines/insert_lines for research runs.
- Fix task plugin preview to use ActivityPreview (required by n00n.subagent)
and return raw structured-output validation errors for tests.

* fix(workflow): preserve string subagent output and encode structured tables

Subagent.launch may return either a plain string or a structured table.
Store strings unchanged, JSON-encode tables with error handling, and use an
empty string for nil output instead of encoding nil.
@w0wl0lxd
w0wl0lxd deleted the feat/agent-cli-simplify branch July 29, 2026 22:29
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant