fix(skills): don't report persisted skill bodies as loaded - #98736
Closed
mira-solari wants to merge 1 commit into
Closed
mira-solari wants to merge 1 commit into
mira-solari wants to merge 1 commit into
Conversation
`skill_view` returns one JSON payload whose `content` IS the skill. Two
independent context-protection decisions in tools/tool_result_storage.py can
delete that body on its way to the model and leave a generic
`<persisted-output>` preview that still opens `{"success": true, …}` followed
by a fragment of the body. The model reads a truncated instruction set as a
loaded skill, and the repeat-view dedup cache then tells a retry the earlier
load "is still current and complete" about a body nobody ever saw.
Reproduced on main at 4f22543 through the registered handler plus layer 2:
a 353,176-char skill_view result comes back as a `<persisted-output>` block
whose preview is `{"success": true, "name": …, "content": "# Linked acceptance
reference\n\n## Overview…` — the skill, cut off mid-sentence, labelled success.
No size escapes this. The per-result threshold is 100,000 chars, but the
aggregate turn budget decides AFTER skill_view has returned, inside
enforce_turn_budget, which already had the tool name on the message and threw
it away; with threshold=0 and enough siblings it spills a 1,550-char
skill_view result, measured on the same commit. So the repair goes where the
truncation is decided.
tools/oversized_result_formatters.py is a per-tool replacement-formatter
registry consulted ONLY after the storage layer has already decided to persist.
skill_view registers the only formatter. When the body is gone it returns an
index-only receipt — one typed `[SKILL_INCOMPLETE:` marker saying the skill is
NOT loaded, `load_status: "incomplete"`, `content_returned: false`, a bounded
heading index, the availability metadata, and no fragment of the body — revokes
the premature dedup record, and puts back the `use` bump, because an
undelivered skill was viewed, not used. `skill_view(name, section=…)` is the
route back, and content comes back only when the whole of what was asked for
fits: sections are hierarchical, every heading occurrence gets a stable `#n`,
an oversized section becomes `#n.partK` selectors that concatenate back to the
original span byte-for-byte, and an overflowing index groups into `#a-b` rather
than truncating, so no entry is ever left without a route to it.
Index-only rather than a bounded head is deliberate: a labelled partial
instruction body is still persuasive enough to act on, and removing it entirely
makes the invariant structural.
A linked file needs one thing more. `skill_view(name=X, file_path="refs/big.md")`
produced a receipt honest about WHICH document was missing but whose only
actionable instruction named no file:
skill_view(name="X", section="<heading>")
Following that literally drops file_path, so the call takes the main-SKILL.md
branch and the selectors resolve against a different document with its own
heading index. On a skill whose SKILL.md and linked file both have an
`## Overview`, `section="Overview"` answered out of SKILL.md and `section="NousResearch#3"`
returned SKILL.md's third section advertised as the linked file's — both with
`section_found: true`, no marker, and nothing saying the file had changed. The
resolver was already correctly scoped to the linked file's own body; the
instruction was wrong. So every surface that tells the model how to continue
now keeps the path: the marker, the two notices that carry no marker, the
`section` schema description and the Skill Safety Rule all emit
`skill_view(name="X", file_path="Y", section="<heading>")`, and the marker
regex gained the matching optional segment so emit and detect cannot drift.
Thresholds, previews, persist decisions, aggregate candidate order and the
storage writes are untouched. resolve_threshold("skill_view") is still 100,000
with no override or pin; the real tool name travels in a separate
`formatter_name` kwarg used for the lookup and nothing else, so layer 3 keeps
forcing threshold=0 under its synthetic tool name. With an empty registry —
the shipping default for every tool but one — each receipt is byte-identical to
what it was, and that is asserted rather than argued: every generic path is
compared with and without the registry.
The four skill_view callers with no tool-result budget still receive the full
body: preloaded/slash commands, both cron injection sites, and the MCP surface,
which dispatches through model_tools.handle_function_call and never persists.
Putting the completeness decision inside skill_view() — the obvious design —
would strip all four while only one had a problem;
tests/tools/test_skill_incomplete_direct_consumers.py is the guard.
Tested on macOS 26.6.2 (arm64), Python 3.11. Four focused modules: 50 failed +
1 collection error before, 78 passed after. Twenty affected skill/persistence/
prompt modules: 460 passed, 2 skipped on both the base and this commit.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
mira-solari
force-pushed
the
fix/skill-view-incomplete-linked-file
branch
from
August 30, 2026 17:41
52b8fe5 to
2fce577
Compare
Author
|
Closing because I’m no longer pursuing upstream contributions. |
ppazosp
added a commit
to useomnia/hermes-agent
that referenced
this pull request
Sep 18, 2026
* fix(skills): make omitted instructions explicit and recoverable Adapt NousResearch#98736 (2fce577) to the fork without its upstream-only repeat-view cache. Preserve linked-file selectors and recover complete sections through both per-result and aggregate budgets. Co-authored-by: Mira Solari <268252643+mira-solari@users.noreply.github.com> * fix(delegation): preserve worker context and deliver complete artifacts Adapt the current-prompt budget correction from upstream NousResearch#103486, cache-path mapping from NousResearch#103667, and tasks-only schema from NousResearch#96424. Retain the legacy call interface, preserve shared batch context, and transfer only active-profile delegation artifacts into the paired Toolbox using existing file APIs. * fix(execute-code): deliver large RPC results without replaying tools Use the existing file transport or bounded shell chunks, publish atomically, and retain dispatched results through delivery retries. Fail explicitly after exhausted delivery instead of executing the same request again. * docs(delegation): explain remote transcript refresh behavior * test(execute-code): assert transferred bytes instead of shell command order --------- Co-authored-by: Mira Solari <268252643+mira-solari@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does this PR do?
skill_viewcan currently report a mandatory skill as loaded even when Hermes removed most of its body to protect the context window.The failure has two parts:
<persisted-output>preview that still begins with{"success": true, ...}and a partial instruction body.skill_viewrecords the load before that later replacement, so a retry can say the earlier body is “still current and complete” even though it never reached the model.There is a second correctness issue for linked files: the recovery guidance omitted
file_path. Following the suggested call therefore resolves headings and#nselectors againstSKILL.md, not the linked document that produced the index.This PR makes incomplete delivery explicit and gives the model a complete route back without increasing thresholds or weakening persistence.
Related Issue
No dedicated issue currently tracks this exact persistence-ordering and linked-file continuation bug. I searched open and closed issues/PRs before submitting.
Type of Change
Changes Made
skill_viewis its only consumer.[SKILL_INCOMPLETE: ...]index instead of a partial instruction preview.skill_viewdedup records and restored use accounting when the body did not reach the model.threshold=0, candidate order, persistence decisions, previews, and writes unchanged.file_pathin every recovery surface so linked-file selectors remain scoped to the document that produced them.skill_viewschema for incomplete loads and linked-file continuation.How to Test
SKILL.mdandreferences/big.mdboth contain## Overviewand## Omega, with different content, and make the linked file exceed the result threshold.skill_view(name="...", file_path="references/big.md"). It should return an incomplete marker and a heading index, with no partial body.file_path. It should return the linked file’s complete section, never the same-named section fromSKILL.md.Local result: 538 passed, 2 skipped, 0 failed across 24 files. The focused new modules are 78/78 green.
ruff checkpasses on all changed files, andscripts/check-windows-footguns.py --diff HEAD^reports no findings.The same real two-call sequence fails on current
main: the first result is a partial<persisted-output>preview and the second claims that omitted body is complete. It passes on this branch with all 11 acceptance checks true.Checklist
Code
fix(scope):,feat(scope):, etc.)pytest tests/ -qand all tests passDocumentation & Housekeeping
docs/, docstrings) — N/A; user-facing behavior is described in the tool schema and system guidance changed herecli-config.yaml.exampleif I added/changed config keys — N/A; no config changeCONTRIBUTING.mdorAGENTS.mdif I changed architecture or workflows — N/A; no contributor workflow changeskill_viewschema updatedCompatibility / non-goals
section=; that pre-existing limitation is not widened into this fix. An incomplete plugin skill now fails visibly instead of being falsely reported as complete.