spike(session-search-modes): three-mode session_search investigation - #9
Merged
Merged
Conversation
Three-mode investigation of hermes-agent's session_search tool: - Profiled fast vs summary vs guided against a real DB snapshot (280 sessions, 18k messages) with a fixed 4-query set - Quantified the 'summary fetches full sessions' gap that fast/summary share the same FTS5 retrieval set, but summary's auxiliary LLM ingests ~57-62× more raw transcript than fast surfaces to the calling agent - Reframed mode='guided' from a speed play into a steering surface — the end-to-end flow (fast → agent thinks → guided → agent answers) takes ~30s on Opus, same order as summary's 60s tool wall, but the user can intervene between retrieval and synthesis - Captured measured headlines: Q1 PII final design, Q2 session_search_ investigation root cause, Q3 latest planning, Q4 atlas commons (retrieval miss — the cleanest evidence that improving FTS5 ranking is the deeper lever) - Documented the H2 truncation hit on Q2/Q3 (2/3 sessions over the 100k per-session cap) Implementation lives in a separate hermes-agent PR (link in PR description); this page is the design + measurements artefact. Reproducibility: the harness pair (profile_session_search.py + replay_session_search_consumer.py) and raw outputs live in the investigation workspace, not republished here per the no-data rule for public commons / spikes content. First entry under spikes/. A spikes/index.html + skill / template will follow once a second spike lands.
Contributor
The <style> block used var(--surface-2) for background and var(--text) for text. In light mode --surface-2 is unset (falls back to #0f172a, dark) and --text resolves to #0f172a (slate-950) — bg and fg were the same colour, making JSON state blocks invisible. Fix: hardcode dark IDE colours so code blocks render the same in light and dark mode. Cleaner than fighting the design system for an artefact that wants IDE aesthetics regardless of page theme.
Two unclosed <small> tags in the stage-3 mermaid diagram source (aux and ret nodes) caused mermaid's HTML-label parser to keep an inline-formatting context open across nodes. foreignObject widths got measured wrong and node text spilled past every box. Confirmed by audit: 0 mismatched <small> across 19 diagrams in data-model.html / sequence-diagrams.html; 2 mismatched in this spike. Fix: close both tags. No font-load race, no theme bug — just bad markup on our side. Working as intended in upstream atlas pages.
…sion
Root cause of the truncated flowchart text. Mermaid emits <g class="label">
elements for every node and edge during its offscreen measurement pass.
Our inline <style> defined ".label { font-size: 11px }" for SHARED/FAST/
SUMMARY/GUIDED tags in the consumer-replay section. That global rule
matched mermaid's own .label group elements inside the offscreen sandbox,
shrinking measured label width to 11px-font dimensions.
After mermaid injected the SVG into .mermaid-canvas, base/styles.css
applied the correct sizes (.mermaid .nodeLabel = 15px, .mermaid .edgeLabel
= 12px), but by then the foreignObject widths were already locked at the
smaller measurement. Result: real text overflowed every node box.
Confirmed with same diagram source rendered on data-model.html — 0
overflow there, many overflows here. The only difference is our spike's
unscoped .label rule.
Fix: rename to .mode-label everywhere (5 CSS rules + 9 HTML use sites).
Earlier hypotheses (font-load race, unclosed <small> tags) were incorrect
or incidental — the <small> closures are kept since they were genuinely
malformed.
Earlier framing in §4, §5, §6, §6.5 implied guided was fast because the
DB read alone is sub-ms. That number is meaningless in isolation: guided
is never used standalone — it always rides on a prior fast call PLUS a
main-agent inference turn that picks the anchors (~15 s on Opus).
Honest e2e cost of fast → think → guided is ~15 s, dominated by the
agent inference turn, in the same order of magnitude as summary's tool
wall (42–66 s). Fast is the only mode whose tool wall is meaningful as
a latency comparison.
Changes:
- §4 table: rename 'fast+guided' column to 'fast → think → guided
(e2e)', show ~15 s instead of ~13 ms; lead with caption explaining
why before the table; add row showing how the e2e number is
computed (5–11 ms tool + ~15 s inference + 2–3 ms tool).
- §4 byte numbers refreshed from latest profile run (20260512-155222).
- §5 GUIDED PATH header rewritten from '~11 ms total' to
'fast → agent-think (~15 s) → guided (~2 ms)'.
- §6 fast-as-default card: 'partially supported' → 'rejected, with
caveat'. Body explicitly names the inference turn and concludes
'real win is steering, not speed'.
- §6.5 composed-payload card: refresh numbers to multi-anchor
(49 kB → 131 kB on Q1; 53 kB → 138 kB composed); replace the
misleading '~11 ms total' lesson with '~15 s end-to-end, dominated
by the agent-inference turn between the two tool calls'.
The spike under-specified how the three modes are actually meant to be
used at runtime. §6 framed guided as a 'steering surface' but didn't
walk through the mechanics; §7 specified the schema but not the loop.
Reader has to guess whether there's a multi-step state machine, a
'request guidance' API, or just standard tool-calling. Spoiler: it's
just standard tool-calling.
New §6.7 (Interaction Model) lays out:
- Activation surface: one tool, mode is just a string argument.
Table maps every possible call shape to its server-side path,
incl. multi-anchor.
- Default mechanics: SESSION_SEARCH_SCHEMA default='summary' AND
server-side mode = (mode or 'summary') normalisation. Both are
needed because not all providers honour schema defaults.
- The drill-down in concrete turns: TURN 1 user → TURN 2 fast tool
call → TURN 3 agent reasoning + guided tool call → TURN 4 final
answer. Each tool call is its own inference turn. Explains where
the ~15 s end-to-end actually goes (inference, not tool walls).
- 'What the tool does NOT do' list: no enforced stop, no
'ask-for-guidance' API, no fast→guided coupling enforcement, no
max-depth, no multi-turn state. All flow is driven by the schema
description plus model judgement.
- Two valid steering postures: agent-as-picker (typical) vs
user-as-picker (steering). Same code path, different actor doing
the anchor selection.
- Closes with 'what governs the loop': the schema description is the
only steering logic. Tuning the prompt + re-running consumer-replay
is the iteration unit, not adding code.
TOC entry added. Section numbered 6.7 so §7 (light spec) doesn't have to
shuffle.
…ummary's prose feel Live-test observation from May 12 multi-anchor drill: guided output reads as transactional and narrow compared to summary's briefing-prose recap. That's a real ergonomic difference, but it's not intrinsic to the mode — it's about where the synthesis work lands. Summary pre-digests through the aux LLM, so the main agent gets prose and passes it through. Guided returns raw messages, so the main agent does the synthesis from scratch — and apparently reaches for a 'msg X says…, then msg Y…' walk-through unless told otherwise. The fix is one sentence in the user prompt: ask for 'briefing-style recap, not a transcript walk' and guided lands at summary's reading feel while keeping its precision (no aux-LLM telephone, raw messages visible, anchor flagged). Added a new card to §6.7 with the synthesis-locus table (summary → aux LLM, fast → main agent from snippets, guided → main agent from raw messages) and the two phrasings side by side. Also added a bullet to the §6.7 closer card linking back to it — the real tradeoff isn't speed or precision, it's whose job synthesis is. This is the kind of observation that wouldn't have been spottable without the live test. Worth baking into the schema-tuning candidates for the next iteration of session_search's tool description.
…othesis findings Original card (yesterday's commit 967cb7b) predicted guided would default to transcript-walk style and asking for 'briefing-style recap' would fix it. Today's live A/B run (May 13, Opus, same multi-anchor drill on 'session-search-investigation') falsified the strong form of that claim. What actually happened: - Both prompts produced briefing prose. Neither was a transcript walk. - Prompt 1 (default) → chronological state-walk with named stages, analyst-notes feel, action items surfaced. - Prompt 2 (explicit briefing) → polished narrative arc with phase headings, story-shape closer, verbatim quote pulling, fewer action items. Refined finding: the lever isn't transcript-vs-prose, it's how much structural opinion the agent invents vs how much you specify. Default lands at analyst notes; explicit-briefing lands at executive narrative. Both have value; there's no 'right' default — but there IS real per-prompt steering room on guided you don't have on summary. Rewrote the card to: - Open with explicit acknowledgement that the earlier claim was partially wrong (recording falsification honestly, not silently retconning). - Keep what's still true: where synthesis happens (server-side aux LLM for summary, main agent for guided) is the underlying mechanism. - Replace the speculative 'default feel / fixable by prompt' columns with a tunability column that names the actual steering surface. - Show the two A/B prompts and their outputs side by side, not as hypothesis but as observed result. - Surface three schema-tuning candidates worth A/B'ing next (verbatim quote-pulling cue, action-item surfacing under briefing posture, honest-about-gaps cue). All prompt iterations — none need code. Updated the §6.7 closer bullet to match. This is the kind of cycle §6.7 itself names as the iteration unit: tune the prompt, observe the result, only promote to code when the evidence holds. Partial falsification is still evidence.
Replaces the §6.7 synthesis card with the full three-run progression:
bare drill instruction → briefing cue → full stack (briefing + tightness
+ active retrieval + dual closer). Each run gets its own bulleted
findings list. Run 3 is the keeper.
Key things Run 3 surfaced that Runs 1 and 2 didn't:
- Active retrieval: agent ran a supplemental fast search unprompted
to plug a Stage-4 shipping-state gap (initial fast → multi-anchor
guided → window refetch → 'guided mode session_search implementation
OR merged OR PR' supplemental search → synthesis). This was the
target behaviour for the 'play the surface' cue — it landed.
- Briefing prose survived the workflow instructions — Scope + 4
Stages + Handoff + Gaps, bulleted where appropriate, paragraphs
where appropriate. Tightness cue ('each section should earn its
place') held.
- Both closers present and substantive: 6 actionable handoff bullets,
5 retrieval-honesty bullets. Run 2 had lost the handoff entirely
when it gained the briefing polish; Run 3 has both.
- Verbatim quote pulling continued unprompted — appears to be what
Opus does once it has raw guided messages with user dialogue.
- The self-aware gaps bullet: agent caught itself instantiating the
very bug being investigated ('this whole briefing is itself an
example of the pathology — I'm reconstructing the Stage 4 shipping
state from cron/sibling references...'). Without the gaps cue, that
epistemic-surface finding would have been invisible.
Card now names the four shape-knobs explicitly: briefing framing,
tightness, active retrieval, dual closer. All four landed in Run 3.
Will promote to the tool's schema description once a few days of normal
use confirm the pattern holds — not before.
Also added the full Run 3 prompt verbatim to §10 (reproducibility) as a
new ve-card. Cross-linked from §6.7's Run 3 description. Lets readers
paste the prompt straight into their own CLI / TUI against their own DB
to reproduce the behaviour shape. The candidate phrasing is now part of
the public record, not just a workspace artifact.
Same evidence loop: tune the prompt, observe what comes back, only
promote to code when the evidence holds. §6.7 closer principle in
action.
…w-up' shape Live-test conversation pushed back on the 'three modes' framing — guided can't be a starting move (it needs anchors), so calling fast / summary / guided 'three peers' overstates the symmetry. Honest shape: • Two STARTING moves: fast, summary (called with query) • One FOLLOW-UP move: guided (called with anchors from a prior fast) Two cards updated: §6.7 'Activation surface' renamed to 'Two starting moves + one follow-up move'. New three-column table (Move | What the agent emits | What runs) replaces the previous flat list, making the Browse vs Start vs Follow-up distinction visible at a glance. Notes the user-config default-mode override (config.yaml) and explains why guided isn't acceptable as a default (rejected by the resolver). §7 'Surface area' updated to current LLM-facing schema. The old signature listed session_id / around_message_id as guided-only parameters; those were just removed from the schema in the matching hermes-agent commit. The schema now exposes only anchors (which handles single and multi cases identically) + window. The single-anchor params still exist as Python kwargs for direct callers and test fixtures — covered in the new 'Evolution note' caption so the history isn't lost. Pairs with hermes-agent commit 74fdfe6b5 (schema cleanup).
Companion to index.html (the investigation record). results.html is the
front-door page for someone who's never used the new modes: explains why
they exist, what they buy you, and how to try them — without the
investigation history.
Five sections, four visuals, each earning its place:
1. Three-track timeline showing wall time and cost per mode pattern
stacked on a single 90-second clock. Agent reasoning (grey), tool
calls (coloured), aux LLM (red). Makes it obvious where the dollars
and seconds go.
2. Question-shapes menu: 3 rows, each one binds a typical query shape
to the mode the agent picks, the wall time, and the cost. Scannable
in 10s.
3. Search-fidelity 2x2 (confidence x retrieval-strength). Names the
'confident-stale' failure mode the old default lived on. Plus one
real abstracted example pair (Telepath research from smoke-test
corpus) showing the old vs new behaviour on the same prompt.
4. Cost bar chart: stacked main-loop vs aux-LLM cost by mode pattern.
Summary is the only expensive mode; aux dominates it.
5. Bottom line: 3 takeaways + one config line (default_mode: fast in
~/.hermes/config.yaml) + a deep-dive link to index.html.
Deliberately NOT included on this page (lives in index.html for anyone
who wants depth):
- Test methodology / scenario IDs
- Pair-resolution regression details (branch-merge concern)
- FTS5 footgun investigation
- Spec-author bugs caught at gates during dry-run
- Run 3 prompt scaffolding details
TOC navigation: results.html sits ABOVE index.html in the shared TOC so
new readers land on the value-prop page first. The 'Read the deep-dive
→' footer is the explicit forward link to index.html for power users.
Audience: Nous internal (per user spec). Tone follows the user's
post-test framing — matter-of-fact, sparse copy, visuals do the work.
No paragraph longer than 3 lines per design-system rules.
Page-local CSS only (timeline, costbar, shapes, fidelity-grid, takeaway).
No mermaid on this page so the .label class collision doesn't apply, but
custom classes are scoped anyway (timeline-mode, costbar-label, etc.)
for forward-compat with future diagram additions.
Vision-audit feedback from rendered page identified three issues:
1. Timeline segments were too squished. Reference clock was 90s but
most rows only used 11-28s of that, so segments rendered as tiny
illegible boxes. Fixed by rescaling to an 80s reference where the
longest mode (summary, 60s) reaches 75% of track width. Short rows
now show readable segment labels ('agent thinks', 'fast', 'drill')
instead of cramped 1-char abbreviations. Also corrected the
active-retrieval row to show 2 guided drills (matching the actual
smoke-test scenario S14 which was fast×6 → guided×2).
2. Fidelity 2x2 didn't highlight the dangerous diagonal — the whole
point of the visual. Added background tint + corner labels:
- 'OLD DEFAULT' (red, top-left cell, confident+weak retrieval)
- 'NEW DEFAULT' (green, bottom-left cell, hedged+weak retrieval)
Reader can now see at a glance which cell the system used to live
in and which it's been moved to.
3. Page leaned cluttered — too much prose between visuals. Cut the
third fidelity card ('Tool priority' — repeated the same point in
different words; merged as a one-line note inside the example card).
Trimmed long captions on the timeline, fidelity, and cost cards.
Net effect: same information, less scrolling, the four visuals do
more of the work and the prose connecting them is reduced to one
line each.
1. 'recent' is not a mode. The schema enum is [fast, summary, guided] only. What I'd been calling 'recent mode' is the no-query browse path — when you call session_search with no query, the tool short- circuits before mode-checking and returns metadata for recent sessions. The response labels itself 'mode: recent' internally for clarity but you can't pass mode='recent'. Renamed the row label from 'recent only' to 'browse' on both the timeline and cost-bar, and updated the legend accordingly. 2. Several timeline rows were missing the final 'agent replies' segment. The agent always thinks AFTER the last tool call to write the user-facing reply — without that segment the agent would just hand back raw tool output. Added explicit 'agent replies' (or 'briefing' for the long active-retrieval row) closer to every row. Also renamed the inter-call think segments to be more descriptive: 'sharpen' on the active-retrieval row (where the agent is refining the query between fast calls), 'picks anchor(s)' on the rows that transition fast→guided (where the agent is deciding which result to drill into). The grey segments now tell a story instead of just 'think think think'.
…ession_search.default_mode The tools.session_search.default_mode path advertised in the prior revisions doesn't exist in the config schema. The actual knob lives under auxiliary.session_search alongside max_concurrency. Update the config snippet and the inline callout to match the shipped path.
User-flagged gap: the page documented modes' cost and timing but never
showed what each mode actually returns. Added a new §3 between
'shapes' and 'fidelity' (downstream sections renumbered 4-6).
Three visuals, no novel:
1. Per-call response anatomy — three side-by-side payload cards
showing real bytes from the smoke test. Fast: ~4KB of snippets +
metadata. Summary: ~12KB of LLM-rewritten prose. Guided: ~12-130KB
of raw conversation messages around a chosen anchor. Each card has
a labelled colour-coded mock-up of the JSON shape so readers can
see at a glance what kind of data lands in the calling agent's
context window.
2. Breadth-depth axis — single horizontal axis with three dots placed
by their character:
- fast (left, breadth): N sessions × snippet
- summary (middle): N sessions × LLM recap
- guided (right, depth): 1 anchor × full window
Caption frames the typical flow: fast scans horizontally,
guided drills vertically, summary collapses both into an
LLM-narrated middle at LLM-narrated cost.
3. The tunable parameters — six knobs in a 2-column grid showing
param name, which modes use it, and what it controls (query,
limit, anchors, window, role_filter, mode). Explains how the
agent composes them without making the user think about it —
mode driven by question shape, limit by topic breadth, anchors
and window by drill depth.
Closing line ties it back to user control: the agent picks defaults,
power users can pin any knob, explicit overrides win.
CSS additions are page-local and scoped (.payload-card, .bd-axis,
.lever-grid — no bare class names that could collide with mermaid).
TOC updated with new §3 entry.
…initions + interplay
User flag: the page jumped from KPI strip straight into facts, never
introduced the concepts. New readers needed (a) what was wrong before,
(b) what fast/summary/guided actually are, and (c) how they connect —
without that scaffolding the visuals later are floating in space.
New section sits between Overview (KPI strip) and §1 (timeline). Not
numbered — it's preamble to the numbered analysis sections.
Two side-by-side cards:
LEFT (slate) — The problem
Two short paragraphs naming: one mode previously, applied to every
question, paid synthesis cost (~30-60s, ~$1.00) on everything.
Confident-stale failure mode called out explicitly.
RIGHT (mode definitions + interplay diagram)
Three colour-coded mode rows (fast/summary/guided) with role tag
(starting move · discovery / starting move · synthesis /
follow-up move · drill-down) and one-line definition each.
Below the rows: a small interplay diagram showing the three valid
flows:
question → fast
question → summary
question → fast → guided
Caption: 'two ways to start, one way to follow up. Guided is the
deep-dive after fast or summary points to where to look.'
This is the framing readers need before the cost/timing/data
visualisations make sense. After this section, the rest of the page
(now numbered 1–6, plus this preamble) reads as a deepening of
concepts already introduced instead of a parade of facts about
unnamed modes.
Two fixes from re-render review:
1. Problem card trimmed from 4 long sentences to 2 short paragraphs.
The 'laundering thin hits into authoritative-sounding answers'
line is the killer point — now in <strong>, last position, lands
harder. Old version buried it after a cost paragraph.
2. Interplay diagram restructured. Previous version showed three
parallel 'question → X' rows, which read as three equal options —
not as 'two starts, one of which has a follow-up'. New version:
┌─ TWO STARTING MOVES ──────────┐
│ question → fast │
│ question → summary │
└───────────────────────────────┘
┌─ ONE FOLLOW-UP MOVE ──────────┐ (blue tint, accent border)
│ fast → guided │
│ summary → guided │
└───────────────────────────────┘
Two visually distinct groups. Follow-up group has a left accent
border in guided's blue, tying back to the mode tag. Also fixes
an inconsistency the audit caught — caption said guided follows
'fast OR summary' but only fast→guided was shown. Both paths now
visible.
User flag: the intro said WHAT each mode returns but not WHAT WORK each
mode does to produce it. That's the load-bearing distinction —
explains why summary costs ~6× more than fast even though both run the
same FTS5 query.
New 3-column diagram between mode definitions and the interplay
diagram. Each column shows the pipeline steps for one mode, colour-
coded by step type:
FAST SUMMARY GUIDED
step 1: FTS5 step 1: FTS5 (skip: no FTS5)
(skip: fetch) step 2: FETCH all step 1: FETCH ±N
messages from messages around
every match each anchor
(skip: LLM) step 3: AUX LLM (skip: LLM)
per matched session
return: snippets return: recaps return: raw window
Skipped steps are visually de-emphasised (opacity, italic, leading
em-dash) so the reader sees what each mode DOESN'T do as well as what
it does. The 'STEP N' labels are mode-coloured by accent border:
green for FTS5 steps, amber for FETCH, red for AUX LLM, slate for
RETURN.
Closing caption pins the lesson:
'The cost gap isn't about the search — fast and summary run the
same FTS5 query. The gap is what happens AFTER: summary fetches
every message from every match and runs an aux LLM over each one.'
This makes the 6x cost differential obvious and earned by the page's
own diagrams, instead of being a number you have to take on faith.
…erarchy
Two vision-audit fixes:
1. SKIPPED steps now visually de-emphasised at a glance, not by
inspection:
- ✗ marker prefix (instead of em-dash)
- dashed left border (instead of solid)
- line-through text decoration
- lower opacity, slightly darker background
The 'this mode does N steps' pattern now reads instantly: fast=1
live step + 2 skipped, summary=3 live, guided=1 live + 2 skipped.
2. Punchline elevated to a proper takeaway callout:
- Amber accent border + tinted background (same warning palette
the page uses for cost emphasis)
- 'The cost gap isn't about the search.' bolded mono-caps on its
own line as the lead
- 'after' (the load-bearing word) amber-highlighted in the body
No longer reads as a footnote; reads as the lesson the diagram is
teaching.
Same 3-column layout, same content — just better visual hierarchy so
the eye finds the right beats: pipeline → SKIPPED→ACTIVE contrast →
punchline.
Previous styling (line-through, opacity, dashed border) was too subtle — the rows still occupied the same height as live steps so visual scanning didn't make the 'this mode does N steps' pattern obvious. Structural change: SKIPPED rows are now visually shorter and inline: - 4px padding (vs 8px on live steps) - single line: 'skipped · no message fetching' - dashed left border, italic, dimmed grey - no full-card background Live steps stand out at full height with coloured accent borders (green for FTS5, amber for FETCH, red for AUX LLM, slate for RETURN). Skipped steps recede to thin grey lines you can scan past. Result: column heights now visibly differ. Fast has 1 tall live step + 2 thin skips + return. Summary has 3 tall live steps + return. Guided has 1 thin skip + 1 tall live step + 1 thin skip + return. The pattern reads at a glance.
…anism diagram height contrast Root cause: CSS grid auto-stretches columns to equal heights by default. So even though fast/guided had shorter live-step rows + thin skipped rows, the COLUMNS were padded out to match summary's height (all three were exactly 348px). Empty space at the bottom of fast and guided made the diagram look uniform even though the rows inside it were correctly sized. One-line fix: align-items: start on .mechanism-grid. Columns now size to their content. Summary stays tall (~348px with its 3 live steps), fast/guided collapse to ~210px with their 1 active step + 2 thin skipped rows. The visual 'this mode does more work' read is now in column height, not just row styling. JS check confirms live-step row heights are 64px vs skipped rows at 21px — the contrast was always there, just invisible because the columns were stretched.
1. Subtitle: drop 'Nous internal' (sub-header should be neutral).
2. Restructure intro section. Previous layout had the problem card on
left and the modes card on right in a 2-col grid — the modes card
was tall (3 mode rows + mechanism diagram + interplay diagram) and
forced the problem card to stretch with empty space below the
2-paragraph problem statement. New layout is single-column vertical:
section header
↓
problem statement (one paragraph, full width)
↓
'three modes' card (the 3 mode rows)
↓
'what each mode actually does' card (mechanism diagram + punchline)
↓
'how they connect' card (interplay diagram)
Each card stands on its own. No more empty space.
3. TOC titles shortened and renumbered. 'The problem, and what we
built' → 'The problem'. '1. The three modes, on one clock' → 'On
one clock'. Etc. Section header dots in the body also drop their
numbers. The TOC is for navigation, not narrative.
4. Bottom-line takeaways refreshed. Previous version led with the
4-6× cost improvement claim — accurate but understated the bigger
shift. New version leads with the schema reframe (fast is now the
default for 'catch me up on X', via the 1a00d730e tool-description
change), then guided's bookends + tool-noise filter, then summary
as opt-in synthesis, then the honesty improvement. Matches the
actual landed state at HEAD.
5. Bookends + tool-noise filter (commit b54b24607) now reflected in
three places:
- 'three modes' card: guided definition updated to mention
bookends.
- mechanism diagram for guided: added STEP 2 BOOKENDS row, added
'tool noise filtered' note to STEP 1 FETCH. fast STEP 1 FTS5
row now notes 'user + assistant roles by default'.
- punchline updated to mention bookends.
- §3 'what comes back' guided payload anatomy: bookend_start and
bookend_end rows added at top/bottom of the example, marked
'session opener'/'session closer'. Tool-noise filter note added.
- parameter table: role_filter description updated to reflect new
user,assistant default for fast/summary.
All page-internal changes. CSS unchanged (existing classes reused).
HTML balance verified.
1. Cost-bar green segments were off-vertically because the page used
a .main modifier class that collided with Atlas base/styles.css's
global '.main { padding-bottom: 40px; grid-column: 2 }' rule
(intended for the page's main content area). Same failure pattern
as the mermaid .label collision documented in atlas-lifecycle:
never use bare/common class names that overlap with base styles.
Renamed .costbar-seg.main → .costbar-seg.main-cost across the CSS
and 5 HTML usage sites. Also added explicit padding:0 and
min-width:0 to .costbar-seg as a belt-and-braces reset against
future global collisions. The green main-cost segment now sits
vertically centred matching the red aux segment.
2. Timeline tool-call segments had truncated labels ('f', 'drill',
etc.) that didn't fit the narrow widths and looked bad. Removed
the text content from .timeline-seg.tool-recent / .tool-fast /
.tool-guided / .aux — the colour-coded legend below the chart
already names them. Kept the 'think' segment labels ('think',
'sharpen', 'picks anchor', 'picks anchors', 'agent replies',
'agent replies (briefing)') because those describe reasoning
states, not modes, and aren't repeated in the legend.
… Optimisations' Updates <title>, h1, and TOC entry. No other changes.
Three of four cards were denominator-fishing from smoke-test methodology
the reader hasn't seen ('14 / 14' — of what?), and the fourth ('summary
is now opt-in') was a sentence pretending to be a metric.
Better signal already lives downstream:
- cost-bar chart proves the ~4-6× cost gap with stacked main+aux
- fidelity 2x2 proves the honesty improvement with worked example
- mode-shapes table tells the reader which mode runs when
- bottom-line callout summarises the change
The lead paragraph ('You used to have one way... You now have three')
is the right opener on its own. Reader scrolls straight from the lead
into 'THE PROBLEM' framing — no artificial KPI scaffolding pretending
to be the headline.
…l review Addresses 7 items from the May 2026 technical review: 1. Cost-section callout — fix mismatched ratio (was '4-6x cheaper / 3x faster', own numbers say ~6x and ~4x). Now reads '~6x cheaper and ~4x faster ($0.17 vs $1.00, 14s vs 60s on these scenarios)'. 2. Cost-section callout — added provider caveat: ratios are robust, absolute dollars assume Opus public rate card May 2026. Prevents the table being misread as a universal claim. 3. Bottom-line takeaways — softened 'zero hallucinations across the smoke-test set' to 'no hallucinations observed across the 14-scenario smoke set ... single sample size — but the failure mode is mechanically prevented, not just unobserved'. Same evidence, claim survives an adversarial reader. 4. Bottom-line takeaways — promoted 'memory first, external sources second' from the Telepath anecdote in the fidelity section to a top-level bullet. The behavioural shift was the real product win and was buried. 5. Cost section caption — added one-sentence note that browse is not one of the three modes (separate code path: session_search() with no query) so readers don't infer four-mode design. Also flagged single-run-per-scenario methodology + p90s not measured. 6. NEW card in §2: 'How mode selection actually happens' (~5 paras). The reviewer's headline ask. Explicitly names the three signals (schema teaching, user wording, configured default), points at the four schema commits (1a00d730e, 54d817f88, 74fdfe6b5, 659af123c) as the load-bearing surface, frames it as a prompt-engineering lever not a code lever. 7. NEW card in §3: 'When anchor selection goes wrong' (~5 paras). The reviewer's gap call-out. Three mitigations spelled out: bookends partial compensate, multi-anchor fan-out for ambiguous topics, re-fast with sharper query. Explicitly names what's NOT built in (automated re-anchoring, quality scoring) and flags adversarial-query worst case as worth a measurement follow-up.
…e scenarios
Addresses reviewer item 3.2: 'The Telepath example is compelling but
it is one anecdote. The fidelity case would be much stronger as a
matrix over the 14-scenario set ... surfacing this as a table would
convert the qualitative claim into the same level of evidence the
cost section enjoys.'
New card in §4 (search fidelity) right after the Telepath example:
'The whole smoke set, classified — would the old default have
laundered?'. Real data extracted from /tmp/hermes-test/results/
S{01-14}.json, classified per row across four columns:
- Scenario (S01-S14)
- Question shape (one-line summary)
- Mode picked (real choice the agent made — colour-coded fast /
summary / guided / browse, same tag style as elsewhere)
- Retrieval (strong / thin / n/a with the actual KB count)
- Old-default risk (HIGH / MEDIUM / NONE / SKIP, with one-line
explanation each)
Then a 4-up rollup at the bottom:
5 HIGH-risk — would have paid 6x, missed a behavioural fix, or
laundered confident-stale prose over zero hits
3 MEDIUM-risk — same answer, 6x the bill
4 No change — summary was the correct call for cross-session
synthesis questions
2 N/A — regression / A-B test infrastructure
Closing line: ~64% of scenarios (9/14 ex-test-infrastructure) would
have run measurably worse under the old default.
Visual treatment: amber-themed (matches §4's existing palette), risk
column uses left-border colour coding (red/amber/green/slate) so the
distribution is scannable down the column. Mode column uses the same
fast=green / summary=red / guided=blue / browse=purple tags the rest
of the page uses.
Three sets of changes for the v2 data:
1. Fidelity matrix refreshed with v2 rubric verdicts + S15-S18 added
(18 scenarios total). Risk distribution shifts: 10 HIGH / 4 MED /
1 LOW / 3 N-A. The 'memory-first / temporal cue / multi-anchor'
improvements are now visible at-a-glance.
2. Cost chart updated with median+p90 from v2 (5 iterations × 18
scenarios). Provider caveat made explicit ('ratios robust; absolute
dollars assume Opus pricing'). Summary row preserved as v1.8
reference with explanation of why v2 didn't measure it.
3. Bottom-line takeaways quote v2 measurements (90 iterations,
17/18 mode stability, 11/18 perfect rubric). 'How to try it'
section gets an honest caveat: setting default_mode: summary is
currently soft; the agent's training overrides it on synthesis-
shaped questions. Documented in red so users know what they're
buying.
Per-section commit messages document where each number comes from
(v2 aggregate.json files in /tmp/hermes-test/results/*/).
Three new visualisations inside §4 (Search Fidelity): 1. Wall variance: per-scenario min→max range bars with median (teal) and p90 (yellow) ticks. Reader spots tight scenarios at a glance, plus the outliers (S05 wide range, S14 highest absolute). 2. Cost variance: same shape in dollars. Shows one-outlier sensitivity for cheap scenarios (S06/11/12) and the tight cluster of synthesis scenarios around $0.30. 3. Rubric heatmap: 18 scenarios × 4 axes coloured by score (green=5, amber=4, orange=3, red=2). Predominantly green; 11/18 scenarios perfect 5/5/5/5. 7 docked scenarios get inline rubric-grader annotations. Where it sits: between the fidelity matrix card and the section's closing paragraph. Together with the matrix this gives: - matrix: WHICH MODE for each scenario, plus old-default risk - variance: HOW STABLE each scenario was across 5 iterations - rubric: WAS THE ANSWER GOOD by a 4-axis independent grader Addresses the review's 2.3 (p90/variance) and the implicit 'show the rubric data' ask.
Splits index.html and results.html along their natural axes: index
covers problem scoping, hypotheses, measurements, and design rationale
(the investigation record); results covers what shipped and what it
does (the value-prop and measured outcomes page).
This trim removes implementation/decision/verdict content from index
that had drifted out of sync with the as-shipped PR (default flipped
from summary to fast, schema description rewritten, session-recall
skill shipped) and over-duplicated the contract that now lives in the
PR itself, the schema, and the skill.
Sections removed:
- §6 Decision — verdicts about default and 'guided as steering'
- §6.7 Interaction model — implementation behaviour, not investigation
- §7 Guided mode spec — implementation contract; lives in code now
- §8 Downstream effects — pre-merge verification checklist
- §9.5 Guided mode results — verdicts reversed by what shipped
Two genuinely useful pieces from the removed sections were promoted
into §6 Consumer simulation rather than deleted:
- the three-run prompt-tuning arc (Run 1 → Run 3), reframed as a
follow-up measurement of how prompt shape changes synthesis quality
on the same retrieved messages
- the 'tool-side speedup isn't free at the agent-context level'
insight from the composed payload comparison
Other changes:
- Rename §3 'System today' → 'System as it was' so it stays locked to
the pre-change picture (would otherwise drift every time the system
evolves)
- Rename §1 stat 'Modes proposed: 3' → 'Hypotheses tested: 3';
drop 'Tests passing' (was investigation-state, would mix with PR
state)
- Sanitise personal references: literal session IDs replaced with
[sid_A]/[sid_B]/[work_session_sid] placeholders that preserve the
structural contrast without exposing local hashes; 'Yoni' references
removed; teammate names ('Jake', 'Teknium') and personal artefact
references trimmed where flavor, kept where load-bearing as
investigation context
- Rewrite intro to frame the page as the investigation record
informing the eventually-shipped PR (links the PR explicitly)
- Replace the 4 KB illustrative summary blob in §5 State walk with a
much shorter generic recap that still shows summary mode's
structured-recap shape
- Renumber §6.5/§9/§10 → §6/§7/§8 to match the new TOC
Net: 1650 → 1131 lines.
…k + repro; compress §6 §4 Findings table previously synthesised a 'fast → think → guided e2e' column from a ~15s inference estimate. v2 measured the real end-to-end interaction timing across 5 iterations × 18 scenarios on the as-shipped toolkit; that's the canonical number now. Index keeps the tool's intrinsic latency profile (the contract of each mode) and points at results.html for the full-interaction measurement. §7 Future work removed: - The 'retrieval is the deeper lever' synthesis-of-investigation insight was the only piece worth keeping; absorbed into §6's closing structural insight where it belongs (the section that earned it). - Other follow-ups (SQLite WAL/JSONL, H2 truncation, H3 silent FTS5 errors, FTS5 AND footgun, profile harness extensions) overlap with the PR's Out-of-scope section and don't belong here. §8 Reproducibility removed entirely — verbose harness setup that the PR companion-artefacts note already covers. §6 Consumer simulation compressed: - 'What §6.5 measures' framing callout removed (redundant with the renamed section header) - 'We'll measure once it ships' line removed (guided shipped, was measured in v2) - Per-query Q1-Q4 cards condensed: Q1 kept (it's the genuinely surprising 'summary's conclusion changed with wider hit list' finding); Q2 + Q3 merged into one card on multi-anchor's cross-session structure; Q4 trimmed to the retrieval-is-the-lever point. - 'Notable improvements from the wider hit list' paragraph removed (redundant with the Q1 card). Other cleanups: - Steering-improvements callout rewritten in past tense (limit cap raised, multi-anchor shipped) - Strip-tool-results follow-up reference moved inline (was pointing at deleted §7) - §6.5 → §6, §9/§10 references rewritten to point at where the content actually lives now Net: 1131 → 881 lines (about 22% reduction in this pass).
…shipped session-recall skill The 'How mode selection actually happens' section had a 'schema teaching' bullet listing four schema-iteration commits, framing the schema as the load-bearing surface for the mode picker. After the schema description was rewritten as a tight specification (and the playbook moved to the shipped session-recall skill), that framing was inaccurate. Now: - 'Schema description' bullet describes what the schema is today (a compact spec — what each mode does/returns/costs, anchor contract, FTS5 syntax, when-to-use). - 'session-recall skill' bullet describes the playbook layer that ships with the agent and gets loaded on demand. - 'User wording' and 'Configured default' bullets unchanged. - Closing paragraph reframes the prompt-engineering surface as two layers (schema = what, skill = when/why) that can be sharpened independently. - Bottom Line takeaways gains a fifth bullet on the shipped skill — surfaces it at the level of the headline outcomes.
The design system convention (set by base/atlas pages — index.html, data-model.html, diataxis.html, sequence-diagrams.html) is that ve-card elements live inside a grid container (.card-grid, .grid-2, .grid-3) which provides margin-bottom and gap. Plain .ve-card has no margin in the base CSS by design; the grid wrapper is what supplies spacing between conceptual groups. The spike pages had drifted from this — 31 ve-cards (13 in index.html, 18 in results.html) were sitting as direct children of <section> with no grid wrapper, so they butted up against each other with 0px spacing. Some had ad-hoc inline ``margin-top:NNpx`` patches as workarounds. This commit wraps each standalone ve-card with ``<div class="card-grid"> … </div>`` and strips the now-redundant inline margin-top values. The card-grid wrapper provides the standard 20px margin-bottom and gap behaviour, matching the base atlas pages. Changes: - index.html: 13 ve-cards wrapped - results.html: 18 ve-cards wrapped - Total: 31 ve-cards now follow the design system convention - All inline margin-top:Npx workarounds removed - HTML structure remains balanced - Animation index style="--i:N" preserved on each card
The <title> tag still carried the pre-trim framing ('optimisation —
investigation, findings, spec'); the H1 had already been updated to
'session_search investigation' during the trim pass. Aligning.
…sults.html results.html uses 'Session Search — Optimisations' (display-style title case with em-dash). index.html was still using the code-style 'session_search investigation' lowercase. Aligning so the two pages read as a coherent pair — the formal page title now reads as 'Session Search — Investigation' (browser tab, H1, and TOC self-title).
…I/redaction topic The Q1 worked example in §5 State walk was traced through both modes to demonstrate the cron-transcripts-vs-work-session retrieval pattern. The structural finding is the load-bearing insight; the subject of the example was PII / redaction / multimodal work, which is unnecessary detail for a public spike page. Replacements (all lowercase variants): - 'pii redaction multimodal' (query string) → '[q1 topic keywords]' - 'pii-multimodal-redaction' (session ID) → '[q1-topic-sid]' - 'PII work' (prose) → 'Q1 work' / 'the Q1 topic' - 'multimodal' → '[topic]' - 'redaction' / 'redact' → '[keywords]' - 'feat/redaction-passes' branch ref → 'feat/[impl-branch]' - 'feat/multimodal-redaction' branch ref → 'feat/[wip-branch]' - '1,304 base64 blobs, 293.6 MB, MIME breakdown, planned redaction approach' → 'a measured count, dataset size, classification breakdown, and planned approach' - 'Strip base64 encoded images and other multimodal payloads' → 'apply a targeted transform to [topic] data' All structural findings preserved: FTS5 matched cron transcripts not the work session, two queries needed (one for the named topic, one for 'session-search-investigation' meta), the chain of session_search calls inside the cron transcript that the anchor lands on. HTML balanced (depth 0). 19 PII/redaction/multimodal mentions before, 0 after.
The 7-step investigation procedural list and the salvage attribution paragraph weren't earning their space — readers can see the PR provenance in the linked PR itself, and the procedural list is narration that adds no signal beyond what §4 Findings and §5 State walk already show via their evidence. Removed §2 entirely; renumbered §3-§6 → §2-§5; TOC updated; HTML balanced (depth 0). Section count: 6 → 5.
21 tasks
yoniebans
pushed a commit
that referenced
this pull request
May 18, 2026
v11 (per hermes_state.py:36, migration at line 606) re-indexed FTS5 tables to cover content || tool_name || tool_calls and switched from external- content mode to inline mode. This enables session_search to find tool names and tool-call arguments, not just message content. Fixes: - data-model.html: KPI 10→11, section subtitle, ER diagram label, trigger SQL block (was external-content idiom, now inline-mode + tool_name/calls), trigram diagram label, new v11 row in migration history - index.html: SessionDB card 'Schema v10'→'Schema v11' Findings from 2026-05-18 atlas audit: items #9, #17.
yoniebans
pushed a commit
that referenced
this pull request
May 18, 2026
Five sites annotated to resolve to GitHub paths via refs.js: - index.html#state: agent/memory_manager.py (ref already in refs.js) - diataxis.html#mental-model: MemoryManager chip (ref already in refs.js) - diataxis.html#design-decisions: agent/memory_manager.py in Decision #9 (ref already in refs.js) - diataxis.html#patterns: get_hermes_home() in Profile Isolation (new hermes-constants ref → hermes_constants.py) - diataxis.html#patterns: tirith binary in Defense in Depth (new tirith-security ref → tools/tirith_security.py) Audit narrowing notes: - iteration-budget, auxiliary-client, acp-events were flagged in the audit but in each case either (a) the chip already had a working ref, (b) the reference lives inside a Mermaid diagram source where data-ref doesn't apply, or (c) no chip exists in the prose to annotate. Adding such chips would have been content additions, not ref expansions, so they're scoped out. Findings from 2026-05-18 atlas audit: item #27 (narrowed scope).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
The first spike under
spikes/<slug>/— a measured investigation ofhermes-agent'ssession_searchtool that informed the design of a three-mode toolkit (fast,summary,guided).Two pages:
results.html— value-prop + as-shipped behaviour, measured across 90 iterations × 18 scenarios on the as-shipped toolkit. Start here.index.html— the investigation record. Problem scoping, hypotheses, measurements, design rationale.Headline outcomes
fast → guidedinstead of paying for a synthesis LLM call by default. Same DB, same retrieval, but recall starts cheap.guidedmode returns raw messages around an anchor + session bookends (first/last user+assistant turns) — composes withfastfor the discover→drill pattern, no aux LLM, no truncation gamble.summaryis opt-in for genuine cross-session prose synthesis. The pre-existing default path; reserved now for the questions that actually need it.Companion implementation
The implementation ships from a separate branch on
hermes-agent: NousResearch/hermes-agent#26419. The spike'sresults.htmlmatches that PR's as-shipped state.Convention being introduced
spikes/<slug>/as the layout for measured investigations published underhermes-architecture— question → methodology → data → decision → reproducibility.<slug>is kebab-case, descriptive of the topic (this one issession-search-modes). Aspikes/index + spike-authoring skill will follow once a second spike lands.Reproducibility
No measurement data is committed. Harness scripts, raw profiling JSON, and consumer-replay reports stay in the investigation workspace (they consume real LLM tokens and reference private session content). The pages document reproducibility commands for anyone with their own
~/.hermes/state.db.