Skip to content

fix(skills): resolve same-root duplicate names instead of refusing them - #112180

Draft
ankinow wants to merge 2 commits into
NousResearch:mainfrom
ankinow:fix/skills-same-root-duplicate-resolution
Draft

ankinow wants to merge 2 commits into
NousResearch:mainfrom
ankinow:fix/skills-same-root-duplicate-resolution

Conversation

@ankinow

@ankinow ankinow commented Sep 15, 2026 •

Copy link
Copy Markdown

What does this PR do?

Two fixes in the skill/bundle layer. Both were reproduced before being changed, and each one has a regression test that fails on main.

1. Same-root duplicate names resolve instead of refusing (tools/skills_tool.py)

_locate_skill() refused every name with more than one candidate — correct when the candidates come from two different tiers (silent shadowing), wrong when both live under the same root, which is the normal layout on installs whose skills root is a symlink view of another directory:

~/.hermes/skills/review                        -> symlink into the library
~/.hermes/skills/software-development/review/  -> nested copy of the same skill

Both candidates are the same skill in the same root, yet the bare name was refused, so skill_view('<name>') failed and bundle members declared by bare name were printed under Skills missing (skipped) while sitting on disk.

The approach: rank candidates within one root (real SKILL.md beats legacy <name>.md, then shallower path), keep refusing cross-tier ambiguity, and decide ownership lexically so a symlinked entry is not reclassified into the root it points at. The trust check additionally accepts the resolved target of every configured root entry, so a symlinked root no longer warns on every load.

2. A non-UTF-8 manifest no longer takes down bundle discovery (agent/skill_bundles.py)

_load_bundle_file() documents "None (logged) on any error so a broken bundle can't break discovery", but only caught OSError and yaml.YAMLError. UnicodeDecodeError is a ValueError, not an OSError, so a single manifest containing an invalid byte propagated out of scan_bundles() and broke every bundle surface for that install — the TUI /bundles listing, slash-command bundle loading, and cron prompts built from bundles.

Reproduced before the fix:

(tmp / "latin1.yaml").write_bytes("name: latin1\nskills: [skill-a]\n".encode("latin-1") + b"\xff\xfe\n")
scan_bundles()   # -> UnicodeDecodeError: 'utf-8' codec can't decode byte 0xff in position 31

Now that file is logged and skipped while valid bundles still load, which is what the docstring already promised.

Related Issue

Fixes #112179

Type of Change

  • 🐛 Bug fix (non-breaking change that fixes an issue)

Changes Made

  • tools/skills_tool.py — _owning_search_dir() (which configured root owns a candidate, decided lexically) and _rank_same_root_candidate() (tie-break inside one root); _collect_skill_candidates()/_locate_skill() use them so same-root duplication resolves while cross-tier ambiguity still refuses.
  • tools/skills_tool.py — _log_security_warnings() accepts a symlinked entry when the root that exposed it vouches for the resolved target.
  • agent/skill_bundles.py — _load_bundle_file() gains a UnicodeError branch (log + None) so an unreadable-encoding manifest cannot abort discovery of the other bundles.
  • tests/tools/test_skills_tool.py — TestSameRootDuplicationResolves (resolution) and TestTrustWarningSymlinkAware (warn/do-not-warn). Two of the new tests fail against main; test_genuinely_outside_file_still_warns passes on both and guards the invariant.
  • tests/agent/test_skill_bundles.py — test_skips_invalid_utf8_without_breaking_discovery (fails on main with the UnicodeDecodeError above, passes here).

How to Test

  1. Same-root duplication — reproduce on main: create a skills root containing both a top-level symlink and a nested copy of one skill, then skill_view('<name>') → Skill name collision for '<name>': 2 candidates. On this branch the name resolves to the real SKILL.md.
  2. Confirm the anti-shadowing contract is untouched:
    pytest tests/tools/test_skills_tool.py -q -k "Collision or SameRoot" → the cross-tier refusal test still passes.
  3. Trust warning, both ways:
    pytest tests/tools/test_skills_tool.py -q -k TrustWarningSymlinkAware → 2 passed here, 1 failed on main.
  4. Non-UTF-8 manifest:
    pytest tests/agent/test_skill_bundles.py -q -k utf8 → fails on main (UnicodeDecodeError), passes here.

Verification

Run with the repo's canonical runner (scripts/run_tests.sh, per-file isolation), not a bare pytest:

tests/agent/test_skill_bundles.py                                  19 passed  (18 + the regression test)
skills set (9 files: tests/tools/test_skills_tool.py,
  tests/agent/test_skill_bundles.py, test_external_skills.py,
  test_ghost_skill_pruning.py, test_skill_commands.py,
  test_skill_commands_reload.py, test_org_skill_namespace.py,
  test_skill_invocation_description.py, test_tool_dispatch_helpers.py)  184 passed

The full tests/ sweep is still running locally on a 4-core/4 GB host; I am not ticking the full-suite box until it reports, and the PR stays a draft until then.

Checklist

Code

  • My code follows the style of this project
  • I have performed a self-review of my own code
  • I have commented my code where it isn't self-explanatory
  • My changes generate no new warnings
  • I have added tests that prove my fix is effective
  • New and existing unit tests pass locally with my changes

Documentation

  • I have updated the relevant documentation (docstrings cover both behaviours)
  • No user-facing behaviour changed beyond the fixes above

Testing

  • I have tested this on Linux (Arch)
  • I have run the full test suite — pending, see Verification (still running)

General

  • I have read the contributing guidelines
  • My changes are backwards compatible

LERMF and others added 2 commits September 15, 2026 15:51
_locate_skill refused every name with more than one candidate. That is right for
one skill reachable through two tiers (silent shadowing) but wrong for
duplication inside a single search dir: the runtime root is a symlink view of the
library, so a top-level symlink and a nested category copy of the same skill
collided and made the bare name unusable — bundle members were then reported as
missing and skipped.

Candidates are now ranked within one root: a real SKILL.md wins over a legacy
<name>.md, then the shallower path. Cross-tier ambiguity keeps refusing. Ownership
is decided lexically so a symlinked entry is not reclassified into another root.

Trust warnings accept the resolved target of every root entry as well as the
lexical path, so a root that exposes skills through symlinks no longer warns on
every load and the injection-pattern signal stays readable.

Tests: 183 passed (tests/tools/test_skills_tool.py + tests/agent/test_skill_*.py).
The two new trust tests fail against the previous code; the outside-file test
passes on both and guards the invariant.
`_load_bundle_file` documents "None (logged) on any error so a broken bundle
can't break discovery", but only caught `OSError` and `YAMLError`.
`UnicodeDecodeError` is a `ValueError`, so one manifest with an invalid byte
propagated out of `scan_bundles` and took down discovery of EVERY bundle
(TUI `/bundles`, slash-command bundles, cron prompts built from bundles).

Adds a `UnicodeError` branch (log + skip) and a regression test asserting a
latin-1 manifest is skipped while valid bundles still load.

Verified with the canonical runner (scripts/run_tests.sh):
  tests/agent/test_skill_bundles.py - 19 passed (18 + 1 regression test)
  skills set (9 files)              - 184 passed
teknium1 added a commit that referenced this pull request Sep 17, 2026
…shape

Slim redo of the cherry-picked #112180 hunks without changing behaviour:
`_owning_search_dir` / `_rank_same_root_candidate` lose their unreachable
fallbacks (the owning root is chosen by lexical containment, so
`relative_to(root)` cannot fail), the trust check accepts the lexical path
inline instead of through a nested helper, and the WHY-only comments
replace the install-specific narration. Tests trimmed to two invariants
per fix: same-root nested copy resolves to the shallower path, equal-rank
tie still refuses; symlinked entry is trusted by the root that exposes
it, a genuinely outside file still warns. Dropped the legacy flat
`<name>.md` test (the rank still covers it; re-probed live).
@alt-glitch alt-glitch added type/bug Something isn't working P2 Medium — degraded but workaround exists comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint tool/skills Skills system (list, view, manage) labels Sep 17, 2026
@alt-glitch

Copy link
Copy Markdown

This was generated by AI during triage.

Related: #113126 merged the maintainer salvage of these skill and bundle-discovery fixes.

beardthelion added a commit to beardthelion/hermes-agent that referenced this pull request Sep 21, 2026
read_text(encoding="utf-8") raises UnicodeDecodeError - a ValueError,
not an OSError - so readers guarded only by except OSError crash instead
of degrading:

- _inject_context_from: one corrupt .md in a source job's output dir
  crashed _build_job_prompt (which runs outside run_job's try), so the
  downstream job failed on every fire while the file remained. The read
  now skips the file like a silent/blank archive, so an older usable
  archive still gets used instead of masking the whole source.
- read_active_org_id: a non-UTF-8 .active_org marker escaped its
  fail-safe contract ("None = no org skills load") into
  _build_skills_manifest during system-prompt assembly - every turn.
  Now returns None as documented.

The sibling arm in skill_bundles._load_bundle_file is already covered
by open PR NousResearch#112180, which adds the same UnicodeError branch there.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint P2 Medium — degraded but workaround exists tool/skills Skills system (list, view, manage) type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Skill names fail to resolve when duplicated inside a single search root (bundles report them as missing)

3 participants