Repository navigation
Conversation
Port from nearai/ironclaw#5307 ("discourage disabled tool workarounds"). When a user disables a tool via `hermes tools` (or runs a restricted-toolset session), the runtime already enforces that the tool can't be invoked — but the model can still route around it by using a general-purpose tool (e.g. shelling out via terminal) to do what the disabled dedicated tool would have done, or by treating a never-enabled capability as something to work around silently. Adds a short, universal DISABLED_TOOL_GUIDANCE block to the cached system prompt telling the model: if the user names a capability with no available tool, report it as unavailable/disabled rather than substituting another tool. General-purpose tools remain fine for their own legitimate tasks. Follows the existing universal-guidance pattern (TASK_COMPLETION_GUIDANCE, PARALLEL_TOOL_CALL_GUIDANCE): constant in prompt_builder, injected in system_prompt gated on agent.valid_tool_names + config flag agent.disabled_tool_guidance (default True), wired in agent_init and config DEFAULT_CONFIG. Costs ~80 tokens once in the cached prefix.
🔎 Lint report:
|
| Rule | Count |
|---|---|
invalid-assignment |
1 |
First entries
tests/run_agent/test_credits_notices_toggle.py:76: [invalid-assignment] invalid-assignment: Object of type `None` is not assignable to attribute `_credits_session_start_micros` of type `int`
✅ Fixed issues (2):
| Rule | Count |
|---|---|
unresolved-attribute |
2 |
First entries
tests/run_agent/test_credits_notices_toggle.py:76: [unresolved-attribute] unresolved-attribute: Unresolved attribute `_credits_session_start_micros` on type `AIAgent`
run_agent.py:3040: [unresolved-attribute] unresolved-attribute: Object of type `Self@get_credits_spent_micros` has no attribute `_credits_session_start_micros`
Unchanged: 6139 pre-existing issues carried over.
Diagnostics are surfaced as warnings — this check never fails the build.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
The model now reports a named-but-unavailable capability as disabled instead of silently routing around it with a different tool.
When a user disables a tool via
hermes tools(or runs a restricted-toolset session — subagent, kanban worker, curated gateway), the runtime already enforces that the tool can't be invoked, including through the tool_search bridge. What was missing is the behavioral steer: a model whose dedicated email tool is disabled would happily shell out viaterminalto send the mail anyway, or curl an API to stand in for a disabled integration tool — silently working around the user's explicit decision to turn that capability off. This adds a short universal system-prompt block telling the model not to do that.Ported from nearai/ironclaw#5307 ("discourage disabled tool workarounds"), which adds a model-visible capability-surface usage policy. Adapted from IronClaw's Rust prompt-bundle mechanism to hermes-agent's Python prompt-assembly architecture, following the existing
TASK_COMPLETION_GUIDANCE/PARALLEL_TOOL_CALL_GUIDANCEuniversal-guidance pattern.Changes
agent/prompt_builder.py: newDISABLED_TOOL_GUIDANCEconstant (~80 tokens). Tells the model: if the user names a capability with no available tool, say it's unavailable/disabled; do NOT substitute another tool as a workaround. General-purpose tools stay fine for their own legitimate tasks.agent/system_prompt.py: inject the block, gated onagent.valid_tool_names(only when tools are loaded) and the config flag.agent/agent_init.py: readagent.disabled_tool_guidance(defaultTrue) intoagent._disabled_tool_guidance.hermes_cli/config.py: adddisabled_tool_guidance: TruetoDEFAULT_CONFIG(deep-merge picks it up for existing installs, no version bump needed).tests/agent/test_prompt_builder.py: behavior-contract tests for the new block (reports unavailable, forbids substitution, preserves general-purpose tools, stays short, has a heading).Validation
tests/agent/test_prompt_builder.py -k "DisabledTool or ParallelToolCall or TaskCompletion"→ 12 passed.HERMES_HOME(realAIAgent._build_system_prompt()):disabled_tool_guidance: falseCache safety
The block is part of the byte-stable cached system prefix (shipped once, amortised across the conversation) — no mid-conversation mutation, no cache invalidation. Same lifecycle as the sibling universal-guidance blocks.
Infographic