Skip to content

fix(context_breakdown): search the volatile tier for the skills index - #78186

Open
pierrenode wants to merge 1 commit into
NousResearch:mainfrom
pierrenode:fix/context-breakdown-skills-volatile-band
Open

pierrenode wants to merge 1 commit into
NousResearch:mainfrom
pierrenode:fix/context-breakdown-skills-volatile-band

Conversation

@pierrenode

Copy link
Copy Markdown

Summary

#37117 (salvaged and merged today as #77696) moved the <available_skills> block from the stable band of the system prompt to the volatile band, so skill edits don't invalidate the cached identity prefix. A same-day CI-caught follow-up fixed hermes_cli/prompt_size.py's compute_prompt_breakdown() for this move.

agent/context_breakdown.py — which powers the /context and /context all commands (shipped via #72242) — has its own independent copy of the same _SKILLS_BLOCK_RE regex, used in two places (compute_session_context_breakdown(), compute_context_details()), and both were missed by the follow-up fix: they still search only the stable tier.

Empirically confirmed before writing the fix: with a skill loaded, /context's "Skills" category is silently absent from the breakdown, and /context all lists zero skills.

Fix

  • Both functions now search the volatile tier first, falling back to stable for sessions whose cached prompt predates fix(system_prompt): move the skills index out of the stable band to preserve the cached prefix #37117 (mirrors prompt_size.py's fix).
  • compute_session_context_breakdown() also now strips the found skills block out of the volatile tail before folding it into the system_prompt text, so skills tokens are attributed only to their own category rather than double-counted once they're found (a bug that would otherwise have been introduced by fixing the lookup alone).

Test plan

  • Updated test_breakdown_includes_major_categories's fixture to place <available_skills> in the volatile tier (matching current production reality) rather than stable — this is why the bug wasn't caught: the existing test's mock reflected the pre-fix(system_prompt): move the skills index out of the stable band to preserve the cached prefix #37117 layout.
  • Added test_skills_index_falls_back_to_stable_for_legacy_sessions for the legacy fallback path.
  • Added test_skills_index_not_double_counted_in_system_prompt proving the volatile-tail stripping fix.
  • Added test_context_details_finds_skills_in_volatile_tier — compute_context_details() had no prior direct behavioral test coverage at all.
  • Mutation-verified: reverting the fix makes all three new/updated assertions fail with the exact symptom described above.
  • Full neighboring test sweep (context breakdown, prompt_size, system_prompt, status command — 61 tests) passes.
  • ruff check clean on all changed files.

NousResearch#37117 moved the <available_skills> block from the stable band to the
volatile band so skill edits don't invalidate the cached identity
prefix. hermes_cli/prompt_size.py's compute_prompt_breakdown() was
updated for this today, but agent/context_breakdown.py — which
powers the /context and /context all commands (NousResearch#72242) — has its own
copy of the same regex and was missed: both
compute_session_context_breakdown() and compute_context_details()
still search only the stable tier, so the "Skills" category silently
disappears from /context and /context all reports zero skills for
any session with skills loaded.

Search volatile first, fall back to stable for sessions whose cached
prompt predates NousResearch#37117 (mirroring prompt_size.py's fix). Also strip
the found skills block from the volatile tail before folding it into
"system_prompt", so skills tokens aren't double-counted once they're
correctly attributed to their own category.
@alt-glitch alt-glitch added type/bug Something isn't working P3 Low — cosmetic, nice to have comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint tool/skills Skills system (list, view, manage) labels Aug 4, 2026
@86523553

Copy link
Copy Markdown

Independent reproduction on 0.21.5+2446.g9fc7f17906 (Windows 11, Python 3.14.7, Desktop app, platform=cli toolset resolution, 272 installed skills) — confirms the diagnosis in this PR and adds the user-visible numbers.

from hermes_cli.prompt_size import _build_inspection_agent
from agent.system_prompt import build_system_prompt_parts
from agent.context_breakdown import compute_session_context_breakdown

a = _build_inspection_agent("cli")
p = build_system_prompt_parts(a)
print(len(p["stable"]), len(p["volatile"]))                                      # 9942 29260
print("<available_skills>" in p["stable"], "<available_skills>" in p["volatile"]) # False True
print([(c["id"], c["tokens"]) for c in compute_session_context_breakdown(a, [])["categories"]])

Output:

9942 29260
False True
[('system_prompt', 12229), ('tool_definitions', 17477), ('subagent_definitions', 1165)]

So on a fresh session: there is no skills category at all, and system_prompt is inflated to 12,229 tokens while the stable instruction tier is only 9,942 chars (~2.5K tokens) — the missing ~9.7K is the <available_skills> index living in the volatile tier. hermes prompt-size independently reports the skills index at 31,336 B for the same config, and it never surfaces as its own category.

Two user-visible symptoms in the Desktop Context usage popover, both from this root cause:

  1. A Skills row never appears (zero-token categories are filtered out of categories in compute_session_context_breakdown), so the panel looks like skills are free.
  2. System prompt absorbs the entire skills index, so it reads ~12K instead of ~2.5K — i.e. the single biggest fixed cost in the prompt is mislabeled as instructions.

Same payload feeds the Desktop popover, /context, and the gateway slash-command renderer, all of which inherit the mislabel (and /context all lists zero skills, as noted in the PR). +1 to landing this.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint P3 Low — cosmetic, nice to have tool/skills Skills system (list, view, manage) type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants