Skip to content

fix: remove tool-name restrictions on primary helper agents - #6

Merged
MrGTV-love merged 5 commits into
mainfrom
fm/fm-subagent-guard-fix
Oct 6, 2026
Merged

MrGTV-love merged 5 commits into
mainfrom
fm/fm-subagent-guard-fix

Conversation

@MrGTV-love

@MrGTV-love MrGTV-love commented Sep 30, 2026 •

Copy link
Copy Markdown
Owner

Intent

The firstmate primary could not use its own helper agents: bin/fm-subagent-pretool-check.sh denied the Agent tool when firstmate tried to delegate a read-only job (reading five reap-inventory reports and consolidating them into one table). The same stem list also denies Monitor, TaskCreate, ScheduleWakeup, SendMessage and Workflow.

The captain's words, in order:
"if firstmate cannot use its own helper agents, you should assume that is not desired"
"Are you suggested I want firstmate to block its helper agents?? I do not. I do not want any non-valuable friction whatsoever"
"I don't want any non-valuable friction. handicapping processes is almost never a good idea. I prefer fixing processes, not handicapping them"

Standing captain rule on guards (data/captain-shared.md, "Guards: fix, remove or consolidate by judgement"): judge each guard on harm prevented, reversibility, read vs write, who acts, cost, evidence, overlap and design; prefer better engineering over more engineering, and consider the whole system.

What Changed

  • Remove the subagent PreToolUse guard and its Claude hook registration, allowing primary helper-agent and session-tool calls without FM_ALLOW_SUBAGENT while retaining both Bash command protections.
  • Replace guard documentation with the work-based delegation boundary: non-project helper work is permitted, project-specific work still follows fleet requirements, and replacement home-local helper-tool deny lists are discouraged.
  • Remove obsolete guard tests and runner entries; add coverage for helper-tool availability and update Claude/Grok compatibility checks and session-start fixtures.

Risk Assessment

✅ Low: The change is a bounded removal of the explicitly unwanted tool-name restriction, preserves the separate Bash protections and project-work delegation contract, and introduces no substantiated correctness, authorization, or simplification issues.

Testing

Three targeted guard scripts passed, including their bundled lint checks. Real Claude sessions demonstrated five-report helper consolidation, native session-tool execution, Workflow completion, unsafe Bash denial, and correct read-only project routing. Baseline executable checks reproduced the former denials, and the current directory guard retained its deny/allow behavior. Runtime evidence was saved; all disposable environments were removed.

  • Live validation: ✅ go - 5 of 5 scenarios driven live against the product
Scenario Result Live Evidence
Primary delegates five private inventory reports to a helper and receives the correct consolidated table ✅ pass live claude-helper-consolidation.jsonl records the foreground helper, all five report reads, completion, and totals of 17 active and 10 reapable. The baseline executable separately denied Agent with exit 2…
Primary uses monitoring, task tracking, scheduling, and messaging without tool-name interception ✅ pass live claude-session-tools.jsonl records Monitor completion, task creation and completion, a scheduled wake followed by cancellation, and SendMessage's native missing-recipient response.
Primary runs a harmless inline Workflow without the former delegation-shaped denial ✅ pass live claude-workflow-and-bash-guard.jsonl records Workflow launch, completed task status, completion notification, and an empty background-task list afterward.
Removing the helper guard preserves unsafe watcher-launch and persistent-directory-change protections ✅ pass live The real Claude Bash call bin/fm-watch-arm.sh & was denied with watcher-background. baseline-denials-and-current-cwd-guard.json records the current executable denying cd projects/demo with exit 2…
Primary rejects read-only project work as a justification for bypassing fleet delegation ✅ pass live claude-readonly-project-boundary.jsonl records the primary answering that project-specific security audits and reproduction planning require a fleet worker, while private report consolidation may use…
Evidence: Real helper delegation and five-report consolidation

Source: Real helper delegation and five-report consolidation

{"type":"system","subtype":"hook_started","hook_id":"a52d87ef-6631-4add-996c-0e91cdefe5cd","hook_name":"SessionStart:startup","hook_event":"SessionStart","uuid":"b331b3ea-7dcc-4071-8b36-61637fce1bb7","session_id":"0f4410cc-8823-431f-8591-87d258b90703"}
{"type":"system","subtype":"hook_response","hook_id":"a52d87ef-6631-4add-996c-0e91cdefe5cd","hook_name":"SessionStart:startup","hook_event":"SessionStart","output":"","stdout":"","stderr":"","exit_code":0,"outcome":"success","uuid":"52ba9b02-5965-4717-8401-17fc8415aa41","session_id":"0f4410cc-8823-431f-8591-87d258b90703"}
{"type":"system","subtype":"hook_started","hook_id":"ccb59280-7791-4ed5-9b14-922e911a4e51","hook_name":"UserPromptSubmit","hook_event":"UserPromptSubmit","uuid":"ee759c10-c572-4b41-aa16-41da841e5d36","session_id":"0f4410cc-8823-431f-8591-87d258b90703"}
{"type":"system","subtype":"hook_response","hook_id":"ccb59280-7791-4ed5-9b14-922e911a4e51","hook_name":"UserPromptSubmit","hook_event":"UserPromptSubmit","output":"","stdout":"","stderr":"","exit_code":0,"outcome":"success","uuid":"964ce738-a0e1-4811-aa63-9f6507208ef5","session_id":"0f4410cc-8823-431f-8591-87d258b90703"}
{"type":"system","subtype":"init","cwd":"~/.no-mistakes/worktrees/32d18ed9638d/01M49CDRY666DRGYW0C04YJPQ1","session_id":"0f4410cc-8823-431f-8591-87d258b90703","tools":["Task","Bash","CronCreate","CronDelete","CronList","DesignSync","Edit","EnterWorktree","ExitWorktree","ListAgents","Monitor","NotebookEdit","PushNotification","Read","RemoteTrigger","ReportFindings","ScheduleWakeup","SendMessage","Skill","TaskStop","ToolSearch","WebFetch","WebSearch","Workflow","Write"],"mcp_servers":[],"model":"claude-sonnet-5-5","permissionMode":"default","slash_commands":["afk","ahoy","bearings","quiet","stow","updatefirstmate","deep-research","design","design-sync","dataviz","update-config","verify","debug","code-review","simplify","batch","fewer-permission-prompts","doctor","loop","schedule","claude-api","workflow-authoring","run","run-skill-generator","plugin-authoring","advisor","agents","auto-mode-setup","autocompact","clear","color","compact","config","output-style","context","effort","fast","focus","heapdump","init","mcp","import","model","__remote-workflow","workflow-launch-exec","reload-plugins","reload-skills","rename","ultrareview","security-review","usage-credits","extra-usage","usage","insights","recap","skill-doctor","goal","design-consent","design-revoke","list-agents","team-onboarding"],"terminal_slash_commands":["doctor","color","focus","reload-plugins"],"apiKeySource":"none","claude_code_version":"2.1.292","output_style":"default","agents":["claude","Explore","general-purpose","Plan","statusline-setup"],"skills":["afk","ahoy","bearings","quiet","stow","updatefirstmate","deep-research","design","design-sync","dataviz","update-config","verify","debug","code-review","simplify","batch","fewer-permission-prompts","doctor","loop","schedule","claude-api","workflow-authoring","run","run-skill-generator","plugin-authoring"],"plugins":[{"name":"firstmate-calm","path":"~/.no-mistakes/worktrees/32d18ed9638d/01M49CDRY666DRGYW0C04YJPQ1/.claude/skills/firstmate-calm","source":"firstmate-calm@skills-dir","version":"1.0.0"},{"name":"cc-plugin-agents-md","path":"builtin","source":"cc-plugin-agents-md@builtin"},{"name":"cc-plugin-telemetry","path":"builtin","source":"cc-plugin-telemetry@builtin"},{"name":"cc-plugin-plugin-authoring","path":"builtin","source":"cc-plugin-plugin-authoring@builtin"}],"capabilities":["interrupt_receipt_v1","interrupt_cancel_queued_v1","interrupt_send_now_v1","msg_lifecycle_v1","sdk_mcp_tools_list_changed","sdk_mcp_manifests","mcp_read_resource_v1","mcp_tool_ui_meta_v1","ui_surface_v1"],"analytics_disabled":false,"product_feedback_disabled":false,"uuid":"c31ecea5-743c-4af1-b01f-671d761319e1","messaging_socket_path":"/tmp/cc-socks/38416.sock","fast_mode_state":"off","fast_mode_disabled_reason":"sdk_opt_in_required","per_turn_effort_active":true,"view_mode":"default"}
{"type":"assistant","message":{"model":"claude-sonnet-5-5","id":"msg_011CfmaaaH3fa8gjwE4Ei3Xb","type":"message","role":"assistant","content":[{"type":"tool_use","id":"toolu_01VAdFpcxpH83LBetBe861uN","name":"Agent","input":{"description":"Consolidate reap-inventory reports","subagent_type":"general-purpose","run_in_background":false,"prompt":"Read-only task. List the files in ~/.no-mistakes/worktrees/32d18ed9638d/01M49CDRY666DRGYW0C04YJPQ1/.fm-live-lab-n0qygemt/data/reap-inventory (expect five report files) and read each one in full. Do not modify anything, and do not read any other files. For each report, extract its 'active' count and its 'reapable' count. Then produce a consolidated markdown table with columns: report, active, reapable, plus a final totals row summing both numeric columns. Verify the arithmetic. If a report lacks a clear active or reapable figure, say so explicitly rather than guessing. Return the table and any caveats."},"caller":{"type":"direct"}}],"container":null,"stop_reason":null,"stop_sequence":null,"stop_details":null,"usage":{"input_tokens":2,"cache_creation_input_tokens":39208,"cache_read_input_tokens":0,"cache_creation":{"ephemeral_5m_input_tokens":0,"ephemeral_1h_input_tokens":39208},"output_tokens":16,"service_tier":"standard","inference_geo":"not_available"},"input_transformations":[],"diagnostics":null,"context_management":null},"parent_tool_use_id":null,"session_id":"0f4410cc-8823-431f-8591-87d258b90703","uuid":"0ee1a337-5c9d-4c6d-ab0e-083795871973","timestamp":"2026-10-06T20:01:58.109Z","request_id":"req_011CfmaaZyRzYbZWaDRJzjLN","wire_tool_inputs":{"toolu_01VAdFpcxpH83LBetBe861uN":{"description":"Consolidate reap-inventory reports","subagent_type":"general-purpose","run_in_background":false,"prompt":"Read-only task. List the files in ~/.no-mistakes/worktrees/32d18ed9638d/01M49CDRY666DRGYW0C04YJPQ1/.fm-live-lab-n0qygemt/data/reap-inventory (expect five report files) and read each one in full. Do not modify anything, and do not read any other files. For each report, extract its 'active' count and its 'reapable' count. Then produce a consolidated markdown table with columns: report, active, reapable, plus a final totals row summing both numeric columns. Verify the arithmetic. If a report lacks a clear active or reapable figure, say so explicitly rather than guessing. Return the table and any caveats."}}}
{"type":"system","subtype":"task_started","task_id":"ac26f7bc60207c847","run_id":"0mux3uc14-688c10f4","tool_use_id":"toolu_01VAdFpcxpH83LBetBe861uN","description":"Consolidate reap-inventory reports","subagent_type":"general-purpose","is_backgrounded":false,"spawn_depth":1,"task_type":"local_agent","prompt":"Read-only task. List the files in ~/.no-mistakes/worktrees/32d18ed9638d/01M49CDRY666DRGYW0C04YJPQ1/.fm-live-lab-n0qygemt/data/reap-inventory (expect five report files) and read each one in full. Do not modify anything, and do not read any other files. For each report, extract its 'active' count and its 'reapable' count. Then produce a consolidated markdown table with columns: report, active, reapable, plus a final totals row summing both numeric columns. Verify the arithmetic. If a report lacks a clear active or reapable figure, say so explicitly rather than guessing. Return the table and any caveats.","uuid":"2b3b9ccf-4338-42e0-8d71-87798c0a107a","session_id":"0f4410cc-8823-431f-8591-87d258b90703"}
{"type":"user","message":{"role":"user","content":[{"type":"text","text":"Read-only task. List the files in ~/.no-mistakes/worktrees/32d18ed9638d/01M49CDRY666DRGYW0C04YJPQ1/.fm-live-lab-n0qygemt/data/reap-inventory (expect five report files) and read each one in full. Do not modify anything, and do not read any other files. For each report, extract its 'active' count and its 'reapable' count. Then produce a consolidated markdown table with columns: report, active, reapable, plus a final totals row summing both numeric columns. Verify the arithmetic. If a report lacks a clear active or reapable figure, say so explicitly rather than guessing. Return the table and any caveats."}]

... [19446 bytes truncated] ...

ous, so nothing was guessed.\n  - Charlie's active count and delta's reapable count are an explicit 0, not a missing value.\nagentId: ac26f7bc60207c847 (use SendMessage with to: 'ac26f7bc60207c847', summary: '<5-10 word recap>' to continue this agent)\n<usage>subagent_tokens: 36436\ntool_uses: 6\nduration_ms: 10128</usage>"}]}]},"parent_tool_use_id":null,"session_id":"0f4410cc-8823-431f-8591-87d258b90703","uuid":"2707b731-24c2-4518-ae8b-12f69bf8f501","timestamp":"2026-10-06T20:02:08.247Z","tool_use_result":{"status":"completed","prompt":"Read-only task. List the files in ~/.no-mistakes/worktrees/32d18ed9638d/01M49CDRY666DRGYW0C04YJPQ1/.fm-live-lab-n0qygemt/data/reap-inventory (expect five report files) and read each one in full. Do not modify anything, and do not read any other files. For each report, extract its 'active' count and its 'reapable' count. Then produce a consolidated markdown table with columns: report, active, reapable, plus a final totals row summing both numeric columns. Verify the arithmetic. If a report lacks a clear active or reapable figure, say so explicitly rather than guessing. Return the table and any caveats.","agentId":"ac26f7bc60207c847","agentType":"general-purpose","harnessNoteCount":0,"harnessTailCount":0,"harnessSectionHash":"6cd38f9a54c12874","content":[{"type":"text","text":"Captain, all five reports are in `.fm-live-lab-n0qygemt/data/reap-inventory/`, and each has a clear active and reapable figure.\n\n| report | active | reapable |\n|---|---|---|\n| alpha | 3 | 2 |\n| bravo | 7 | 1 |\n| charlie | 0 | 4 |\n| delta | 5 | 0 |\n| echo | 2 | 3 |\n| **Total** | **17** | **10** |\n\n**Arithmetic check:**\n- Active: 3 + 7 + 0 + 5 + 2 = 17.\n- Reapable: 2 + 1 + 4 + 0 + 3 = 10.\n- Overall: 17 + 10 = 27 items.\n\n**Caveats:**\n- I changed nothing and read only the five reports.\n- Each report is three lines: a heading, `active: N`, and `reapable: N`. No figure was missing or ambiguous, so nothing was guessed.\n- Charlie's active count and delta's reapable count are an explicit 0, not a missing value."}],"resolvedModel":"claude-sonnet-5-5","totalDurationMs":10128,"totalTokens":36436,"totalToolUseCount":6,"usage":{"output_tokens_details":{"thinking_tokens":0},"input_tokens":2,"cache_creation_input_tokens":935,"cache_read_input_tokens":35150,"output_tokens":349,"server_tool_use":{"web_search_requests":0,"web_fetch_requests":0},"service_tier":"standard","cache_creation":{"ephemeral_1h_input_tokens":0,"ephemeral_5m_input_tokens":935},"inference_geo":"not_available","iterations":[{"input_tokens":2,"output_tokens":349,"cache_read_input_tokens":35150,"cache_creation_input_tokens":935,"cache_creation":{"ephemeral_5m_input_tokens":935,"ephemeral_1h_input_tokens":0},"type":"message"}],"speed":"standard","fallback_credit":null},"toolStats":{"readCount":5,"searchCount":0,"bashCount":1,"editFileCount":0,"linesAdded":0,"linesRemoved":0,"otherToolCount":0}}}
{"type":"assistant","message":{"model":"claude-sonnet-5-5","id":"msg_011CfmabbaT9Mx2fNbC5bN1V","type":"message","role":"assistant","content":[{"type":"text","text":"Captain, the helper read all five reports and consolidated them. It changed nothing, and each report had clear active and reapable figures.\n\n| report | active | reapable |\n|---|---|---|\n| alpha | 3 | 2 |\n| bravo | 7 | 1 |\n| charlie | 0 | 4 |\n| delta | 5 | 0 |\n| echo | 2 | 3 |\n| **Total** | **17** | **10** |\n\nI didn't read the reports myself, so these figures rest on the helper's arithmetic. It cross-checked the sums: active 3+7+0+5+2 = 17, and reapable 2+1+4+0+3 = 10.\n\nHELPER_CONSOLIDATION_DONE"}],"container":null,"stop_reason":null,"stop_sequence":null,"stop_details":null,"usage":{"input_tokens":2,"cache_creation_input_tokens":979,"cache_read_input_tokens":39208,"cache_creation":{"ephemeral_5m_input_tokens":0,"ephemeral_1h_input_tokens":979},"output_tokens":236,"service_tier":"standard","inference_geo":"not_available"},"input_transformations":[],"diagnostics":null,"context_management":null},"parent_tool_use_id":null,"session_id":"0f4410cc-8823-431f-8591-87d258b90703","uuid":"80eb4098-4b7b-4812-a085-949867fc9100","timestamp":"2026-10-06T20:02:11.601Z","request_id":"req_011CfmabbJLYGPvMdmp1ahP6"}
{"type":"system","subtype":"hook_started","hook_id":"8430b6f9-8691-48b9-b387-c66f0ca7a1b3","hook_name":"Stop","hook_event":"Stop","uuid":"44b22ac3-7905-4460-8c57-a7d46fed0ef8","session_id":"0f4410cc-8823-431f-8591-87d258b90703"}
{"type":"system","subtype":"hook_started","hook_id":"4f3522c5-0b84-499a-8849-511b6ad3ccfe","hook_name":"Stop","hook_event":"Stop","uuid":"2f9f3c1a-4c83-477b-96d1-70fb17131f03","session_id":"0f4410cc-8823-431f-8591-87d258b90703"}
{"type":"system","subtype":"hook_started","hook_id":"47ff8fc4-e57d-4fb8-990a-a6329d8230ec","hook_name":"Stop","hook_event":"Stop","uuid":"a2fed889-6e98-4fb6-a5f8-84a8442e4674","session_id":"0f4410cc-8823-431f-8591-87d258b90703"}
{"type":"system","subtype":"hook_response","hook_id":"8430b6f9-8691-48b9-b387-c66f0ca7a1b3","hook_name":"Stop","hook_event":"Stop","output":"","stdout":"","stderr":"","exit_code":0,"outcome":"success","uuid":"18a9054f-61b9-4f69-b969-a01e3fbbc83b","session_id":"0f4410cc-8823-431f-8591-87d258b90703"}
{"type":"system","subtype":"hook_response","hook_id":"47ff8fc4-e57d-4fb8-990a-a6329d8230ec","hook_name":"Stop","hook_event":"Stop","output":"","stdout":"","stderr":"","exit_code":0,"outcome":"success","uuid":"42974ef1-8fd1-4cfb-ab15-5e9c1c5dbe4f","session_id":"0f4410cc-8823-431f-8591-87d258b90703"}
{"type":"system","subtype":"hook_response","hook_id":"4f3522c5-0b84-499a-8849-511b6ad3ccfe","hook_name":"Stop","hook_event":"Stop","output":"","stdout":"","stderr":"","exit_code":0,"outcome":"success","uuid":"bba7405f-7785-4717-945b-1887bab7ae11","session_id":"0f4410cc-8823-431f-8591-87d258b90703"}
{"duration_api_ms":15478,"stop_reason":"end_turn","session_id":"0f4410cc-8823-431f-8591-87d258b90703","total_cost_usd":0.29009690000000005,"usage":{"input_tokens":4,"cache_creation_input_tokens":40187,"cache_read_input_tokens":39208,"output_tokens":605,"output_tokens_details":{"thinking_tokens":0},"server_tool_use":{"web_search_requests":0,"web_fetch_requests":0},"service_tier":"standard","cache_creation":{"ephemeral_1h_input_tokens":40187,"ephemeral_5m_input_tokens":0},"inference_geo":"not_available","iterations":[{"input_tokens":2,"output_tokens":236,"cache_read_input_tokens":39208,"cache_creation_input_tokens":979,"cache_creation":{"ephemeral_5m_input_tokens":0,"ephemeral_1h_input_tokens":979},"type":"message"}],"speed":"standard","fallback_credit":null},"modelUsage":{"claude-sonnet-5-5":{"inputTokens":10,"outputTokens":1741,"cacheReadInputTokens":108532,"cacheCreationInputTokens":76272,"webSearchRequests":0,"costUSD":0.29009690000000005,"contextWindow":1000000,"maxOutputTokens":128000,"thinkingTokens":0,"canonicalModel":"claude-sonnet-5-5","provider":"firstParty","costBasis":"list"}},"permission_denials":[],"terminal_reason":"completed","fast_mode_state":"off","fast_mode_disabled_reason":"sdk_opt_in_required","subagent_stats":{"spawned":1,"requested":{"background":0,"foreground":1,"unset":0},"started_in_background":0,"max_depth":1,"spawned_by_subagents":0,"completed":1,"failed":0,"killed":{"parent":0,"user":0,"system":0},"refused":{"depth_limit":0,"concurrency_limit":0,"budget":0},"by_type":{"general-purpose":1}},"safety_stops":0,"is_error":false,"num_turns":2,"subtype":"success","api_error_status":null,"result":"Captain, the helper read all five reports and consolidated them. It changed nothing, and each report had clear active and reapable figures.\n\n| report | active | reapable |\n|---|---|---|\n| alpha | 3 | 2 |\n| bravo | 7 | 1 |\n| charlie | 0 | 4 |\n| delta | 5 | 0 |\n| echo | 2 | 3 |\n| **Total** | **17** | **10** |\n\nI didn't read the reports myself, so these figures rest on the helper's arithmetic. It cross-checked the sums: active 3+7+0+5+2 = 17, and reapable 2+1+4+0+3 = 10.\n\nHELPER_CONSOLIDATION_DONE","ttft_ms":3989,"type":"result","duration_ms":17599,"uuid":"28335ec0-35e2-4398-9263-85793f63bd6d","ttft_stream_ms":1740,"time_to_request_ms":238,"first_content_frame_ms":1740,"queued_turn_count":0,"result_index":0}
Evidence: Baseline tool denials and current directory-guard responses

Source: Baseline tool denials and current directory-guard responses

[
  {
    "product": "base guard executable",
    "tool": "Agent",
    "exit": 2,
    "stdout": "",
    "stderr": "{\"hookSpecificOutput\":{\"hookEventName\":\"PreToolUse\",\"permissionDecision\":\"deny\"},\"systemMessage\":\"[subagent-dispatch] the firstmate primary dispatches through the fleet, not the harness's own delegation tools: work started that way has no durable fleet record, leaves every firstmate guard inert, and dies with this session. Instead, first classify the work under the AGENTS.md intake contract, then use bin/fm-brief.sh followed by bin/fm-spawn.sh for dispatched work (blocked tool: Agent, delegation-shaped on \\\"agent\\\"). Launch the session with FM_ALLOW_SUBAGENT=1 for a deliberate exception.\"}\n"
  },
  {
    "product": "base guard executable",
    "tool": "Monitor",
    "exit": 2,
    "stdout": "",
    "stderr": "{\"hookSpecificOutput\":{\"hookEventName\":\"PreToolUse\",\"permissionDecision\":\"deny\"},\"systemMessage\":\"[subagent-dispatch] the firstmate primary dispatches through the fleet, not the harness's own delegation tools: work started that way has no durable fleet record, leaves every firstmate guard inert, and dies with this session. Instead, first classify the work under the AGENTS.md intake contract, then use bin/fm-brief.sh followed by bin/fm-spawn.sh for dispatched work (blocked tool: Monitor, delegation-shaped on \\\"monitor\\\"). Launch the session with FM_ALLOW_SUBAGENT=1 for a deliberate exception.\"}\n"
  },
  {
    "product": "base guard executable",
    "tool": "TaskCreate",
    "exit": 0,
    "stdout": "",
    "stderr": ""
  },
  {
    "product": "base guard executable",
    "tool": "ScheduleWakeup",
    "exit": 2,
    "stdout": "",
    "stderr": "{\"hookSpecificOutput\":{\"hookEventName\":\"PreToolUse\",\"permissionDecision\":\"deny\"},\"systemMessage\":\"[subagent-dispatch] the firstmate primary dispatches through the fleet, not the harness's own delegation tools: work started that way has no durable fleet record, leaves every firstmate guard inert, and dies with this session. Instead, first classify the work under the AGENTS.md intake contract, then use bin/fm-brief.sh followed by bin/fm-spawn.sh for dispatched work (blocked tool: ScheduleWakeup, delegation-shaped on \\\"schedul\\\"). Launch the session with FM_ALLOW_SUBAGENT=1 for a deliberate exception.\"}\n"
  },
  {
    "product": "base guard executable",
    "tool": "SendMessage",
    "exit": 2,
    "stdout": "",
    "stderr": "{\"hookSpecificOutput\":{\"hookEventName\":\"PreToolUse\",\"permissionDecision\":\"deny\"},\"systemMessage\":\"[subagent-dispatch] the firstmate primary dispatches through the fleet, not the harness's own delegation tools: work started that way has no durable fleet record, leaves every firstmate guard inert, and dies with this session. Instead, first classify the work under the AGENTS.md intake contract, then use bin/fm-brief.sh followed by bin/fm-spawn.sh for dispatched work (blocked tool: SendMessage, delegation-shaped on \\\"sendmessage\\\"). Launch the session with FM_ALLOW_SUBAGENT=1 for a deliberate exception.\"}\n"
  },
  {
    "product": "base guard executable",
    "tool": "Workflow",
    "exit": 2,
    "stdout": "",
    "stderr": "{\"hookSpecificOutput\":{\"hookEventName\":\"PreToolUse\",\"permissionDecision\":\"deny\"},\"systemMessage\":\"[subagent-dispatch] the firstmate primary dispatches through the fleet, not the harness's own delegation tools: work started that way has no durable fleet record, leaves every firstmate guard inert, and dies with this session. Instead, first classify the work under the AGENTS.md intake contract, then use bin/fm-brief.sh followed by bin/fm-spawn.sh for dispatched work (blocked tool: Workflow, delegation-shaped on \\\"workflow\\\"). Launch the session with FM_ALLOW_SUBAGENT=1 for a deliberate exception.\"}\n"
  },
  {
    "product": "target cwd guard executable",
    "command": "cd projects/demo",
    "exit": 2,
    "stdout": "",
    "stderr": "{\"hookSpecificOutput\":{\"hookEventName\":\"PreToolUse\",\"permissionDecision\":\"deny\"},\"systemMessage\":\"[persistent-cd] a persistent top-level directory change in the primary firstmate checkout is blocked; it would move the shell out of the home so a later firstmate-owned command runs inside a project clone. Reach the target without moving the shell - use git -C <dir> or an absolute path on the command itself - or scope the cd to a subshell like (cd <dir> && ...).\"}\n"
  },
  {
    "product": "target cwd guard executable",
    "command": "(cd projects/demo && pwd)",
    "exit": 0,
    "stdout": "",
    "stderr": ""
  },
  {
    "product": "target cwd guard executable",
    "command": "printf FM_CWD_SAFE",
    "exit": 0,
    "stdout": "",
    "stderr": ""
  }
]

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

✅ **Review** - passed

✅ No issues found.

✅ **Test** - passed

✅ No issues found.

  • Live validation: ✅ go - 5 of 5 scenarios driven live against the product
Scenario Result Live Evidence
Primary delegates five private inventory reports to a helper and receives the correct consolidated table ✅ pass live claude-helper-consolidation.jsonl records the foreground helper, all five report reads, completion, and totals of 17 active and 10 reapable. The baseline executable separately denied Agent with exit 2…
Primary uses monitoring, task tracking, scheduling, and messaging without tool-name interception ✅ pass live claude-session-tools.jsonl records Monitor completion, task creation and completion, a scheduled wake followed by cancellation, and SendMessage's native missing-recipient response.
Primary runs a harmless inline Workflow without the former delegation-shaped denial ✅ pass live claude-workflow-and-bash-guard.jsonl records Workflow launch, completed task status, completion notification, and an empty background-task list afterward.
Removing the helper guard preserves unsafe watcher-launch and persistent-directory-change protections ✅ pass live The real Claude Bash call bin/fm-watch-arm.sh &amp; was denied with watcher-background. baseline-denials-and-current-cwd-guard.json records the current executable denying cd projects/demo with exit 2…
Primary rejects read-only project work as a justification for bypassing fleet delegation ✅ pass live claude-readonly-project-boundary.jsonl records the primary answering that project-specific security audits and reproduction planning require a fleet worker, while private report consolidation may use…
  • bash tests/fm-turnend-guard.test.sh — passed, including normalized helper-tool matcher coverage and retained Bash-hook registration.
  • bash tests/fm-arm-pretool-check.test.sh and bash tests/fm-cd-pretool-check.test.sh — passed. These existing scripts also invoked their bundled fm-lint.sh checks; no standalone lint or formatter command was run.
  • Launched four real Claude 2.1.292 print-mode sessions from the worktree using private, 120×40 tmux sessions, marked disposable FM_HOME directories, project settings, existing authentication, and stream-json hook events.
  • Delegated five disposable reap-inventory reports to a real helper; exercised Monitor, TaskCreate/TaskUpdate, ScheduleWakeup cancellation, SendMessage's missing-recipient response, and an inline Workflow.
  • Attempted bin/fm-watch-arm.sh &amp; through the real Claude Bash tool and observed the retained PreToolUse denial.
  • Executed the base commit's removed guard from a disposable source snapshot, then exercised the current directory guard's Claude stdin interface against a plain-checkout fixture.
  • Asked the real primary whether a read-only project audit could use a helper instead of a fleet worker; observed its routing decision.
  • Collected runtime transcripts and executable-interface responses; removed all disposable labs, source fixtures, and private tmux servers.
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

Guard judgement

  • Harm prevented and evidence: The evidence reviewed does not establish an incident in which this guard prevented unrecorded project-changing work. It does establish the reported harm: the base executable denied Agent for read-only helper work, whereas the changed primary successfully delegated five private inventory reports and returned their consolidated table. Tool-name classification is not evidence that an operation changes project files or that a worker has a durable record.
  • Read versus write and who acts: The catch-all hook restricted primary helpers and session tools without distinguishing private report aggregation, communication, or session bookkeeping from project-changing work. The affected surface includes Agent, Monitor, TaskCreate, ScheduleWakeup, SendMessage, and Workflow. Earlier baseline payload replay actually allowed TaskCreate while denying the other five; regression coverage includes all six rather than assuming identical baseline behavior.
  • Reversibility and cost: The report aggregation changed no project files. Session tasks and wakeups can be completed or cancelled, and the live validation exercised those transitions. Blocking these operations imposed interruption and unnecessary dispatch overhead without an observed safety benefit. Actual project writes remain subject to the existing authorization boundary; hook removal grants no authority to discard work, merge, or bypass that boundary.
  • Overlap and whole-system design: AGENTS.md sections 1 and 7 already require project-specific coding, investigation, planning, reproduction, and audits to use the durable fleet process, including read-only project investigation. Hard rule 1 restricts primary project writes. fm-spawn owns isolated launch and durable task/runtime records, with existing supervision covering the resulting worker. A harness helper call does not create those records and is not a replacement for that process. The removed filter neither supplied those records nor enforced the semantic boundary against unclassified shell commands.
  • Decision: Remove the guard outright and document the work boundary at its existing owner. Do not replace it with an approval step, mandatory environment switch, local deny list, or another tool-name restriction. The separate Bash protections remain unchanged and passed targeted and live validation. GROK_BOT.md remains unchanged by firstmate's explicit decision because it describes a separate persistent-bot deployment.

Current verification

The final run exercised real Claude 2.1.292 helper consolidation, Monitor completion, TaskCreate/TaskUpdate transitions, ScheduleWakeup registration and cancellation, SendMessage reaching its native missing-recipient response, and Workflow completion. The missing-recipient response proves messaging reached its native consumer, not successful delivery; an earlier run separately recorded a helper-message acknowledgement. Targeted guard regressions passed, and the real primary correctly distinguished private reports from read-only project-specific audit work. No new source fixes were needed in this run. Review, test, and document all completed; the earlier document skip was corrected by full revalidation rather than waiving the required CI check.

@MrGTV-love
MrGTV-love force-pushed the fm/fm-subagent-guard-fix branch from ee4d5fb to 45bd16b Compare October 6, 2026 16:25
@MrGTV-love MrGTV-love changed the title fix(bin): stop blocking helper agents and session tools in the primary fix: restore firstmate helper-agent access Oct 6, 2026
@MrGTV-love
MrGTV-love force-pushed the fm/fm-subagent-guard-fix branch from 0ea119a to 80c0131 Compare October 6, 2026 20:09
@MrGTV-love MrGTV-love changed the title fix: restore firstmate helper-agent access fix: remove tool-name restrictions on primary helper agents Oct 6, 2026
@MrGTV-love
MrGTV-love merged commit 00eff61 into main Oct 6, 2026
23 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant