Skip to content

fix(agent): align empty hidden assistant placeholders on wire to eliminate sanitizer spam (#88955) - #88996

Closed
JoaoMarcos44 wants to merge 2 commits into
NousResearch:mainfrom
JoaoMarcos44:fix/group-chat-hidden-placeholder-sanitizer-88955
Closed

JoaoMarcos44 wants to merge 2 commits into
NousResearch:mainfrom
JoaoMarcos44:fix/group-chat-hidden-placeholder-sanitizer-88955

Conversation

@JoaoMarcos44

@JoaoMarcos44 JoaoMarcos44 commented Aug 18, 2026 •

Copy link
Copy Markdown

Summary

Fixes #88955.

In Bot Mode group chats (and active turn redirects / steering), when a member turn is interrupted before any assistant text has been streamed to the UI, _apply_active_turn_redirect() persists an empty role placeholder with display_kind="hidden":

{"role": "assistant", "content": "", "display_kind": "hidden"}

This placeholder is intentional and load-bearing in the session history to maintain proper role alternation (user -> assistant -> user) without leaking unstreamed reasoning or unfinished content into user-visible transcripts.

However, because this empty placeholder persisted in the session history, the pre-call sanitizer repair_empty_non_final_messages() in agent/agent_runtime_helpers.py detected it on every subsequent turn of the session and repeatedly emitted the following warning:

[WARNING] Pre-call sanitizer: healed 1 empty non-final message(s) by substituting placeholder content — an empty-content turn was in the transcript and would 400 the request ('messages must have non-empty content' / INVALID_REQUEST_BODY). Self-recovering the poisoned transcript in memory; no restart needed.

In long-running group chat sessions (e.g. 98+ API turns), this warning fired perpetually on every turn, causing log pollution, unnecessary repeated memory-copy healing passes, and false-positive alarm noise.


Root Cause Analysis

  1. Persisted History vs. Wire Messages:
    The placeholder {"role": "assistant", "content": "", "display_kind": "hidden"} is kept in messages so that SessionDB, transcripts, and timeline UI renderers can identify the turn as an interrupted redirect without rendering blank speech bubbles.
  2. Wire Stripping Order:
    When constructing outgoing api_messages in agent/conversation_loop.py, api_msg.pop("display_kind", None) stripped the display_kind metadata before sanitize_api_messages() was invoked.
  3. Indiscriminate Warning in Sanitizer:
    repair_empty_non_final_messages() received an empty non-final assistant row without the display_kind="hidden" context and treated it as unexpected data corruption rather than an intentional alternation placeholder, logging the [WARNING] every single turn.

Solution

We implement a clean, two-layer wire alignment:

  1. Wire Substitution in agent/conversation_loop.py:
    When constructing outgoing api_messages, before display_kind is stripped, empty hidden assistant placeholders are proactively populated with _INTERRUPTED_PLACEHOLDER ("[response interrupted]").
    • This ensures strict OpenAI/Anthropic/Bedrock API compliance without 400 errors.
    • The persisted transcript in SessionDB remains byte-stable with content="" and display_kind="hidden" for clean UI rendering.
  2. Defensive Discrimination in agent/agent_runtime_helpers.py:
    repair_empty_non_final_messages() distinguishes between intentional hidden placeholders and true accidental transcript corruption. It substitutes _INTERRUPTED_PLACEHOLDER cleanly while only logging warnings for unexpected non-hidden empty turns.

Verification Plan

Automated Tests

  • Added tests/agent/test_interrupted_hidden_placeholder.py covering:
    • Active turn redirects creating display_kind="hidden" placeholders when no text is visible.
    • Silent healing of hidden placeholders without warning logs.
    • Warning logs preserved for true non-hidden corruptions.
    • Multi-turn conversation regression verification (0 recurring warnings).
    • Preserving streamed visible assistant text during interruptions.
  • Ran all existing steer, conversation loop, and runtime helper suites:
    • pytest tests/agent/test_interrupted_hidden_placeholder.py -v (5/5 PASSED)
    • pytest tests/run_agent/test_steer.py -v (29/29 PASSED)
    • pytest tests/run_agent/test_run_agent.py -k "scaffold or redirect or display_kind" -v (5/5 PASSED)

@alt-glitch alt-glitch added type/bug Something isn't working comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint area/sessions Session lifecycle, resume, persistence, history P2 Medium — degraded but workaround exists sweeper:risk-session-state Sweeper risk: may lose/corrupt/mis-associate session or context state labels Aug 18, 2026
…eam_treats_pre_existing_poisoned_row_as_no_model
teknium1 added a commit that referenced this pull request Aug 18, 2026
…payload at projection time (#88955)

The salvaged writer-side fix stamps api_content on NEW hidden redirect
placeholders, but rows persisted before it (content="" + display_kind=hidden,
no sidecar) would keep re-triggering repair_empty_non_final_messages on every
call forever. Substitute [response interrupted] on the wire copy at the
api_content/display_kind projection stage so legacy sessions converge too.
Never the interrupt scaffold (#81841). Durable transcript untouched.

Regression tests drive run_conversation end-to-end with a spied sanitizer:
the projection must leave the sanitizer nothing to heal (its per-turn warning
spam is the bug), verified failing via sabotage run against the writer-only
fix.

Projection-side approach credit: @JoaoMarcos44 (PR #88996).
teknium1 added a commit that referenced this pull request Aug 18, 2026
…payload at projection time (#88955)

The salvaged writer-side fix stamps api_content on NEW hidden redirect
placeholders, but rows persisted before it (content="" + display_kind=hidden,
no sidecar) would keep re-triggering repair_empty_non_final_messages on every
call forever. Substitute [response interrupted] on the wire copy at the
api_content/display_kind projection stage so legacy sessions converge too.
Never the interrupt scaffold (#81841). Durable transcript untouched.

Regression tests drive run_conversation end-to-end with a spied sanitizer:
the projection must leave the sanitizer nothing to heal (its per-turn warning
spam is the bug), verified failing via sabotage run against the writer-only
fix.

Projection-side approach credit: @JoaoMarcos44 (PR #88996).
@teknium1

Copy link
Copy Markdown
Collaborator

Closing in favor of PR #89544, which just merged. Your PR was submitted first on this issue and your projection-side approach is what we shipped for legacy rows — the follow-up commit in #89544 credits you directly. The writer-side half came from #89525 (stamping api_content on the placeholder at creation), which avoids the sanitizer-warning bookkeeping split. Thanks for the fix and the fast turnaround!

@teknium1 teknium1 closed this Aug 18, 2026
lisajlau pushed a commit to lisajlau/hermes-agent that referenced this pull request Aug 20, 2026
…payload at projection time (NousResearch#88955)

The salvaged writer-side fix stamps api_content on NEW hidden redirect
placeholders, but rows persisted before it (content="" + display_kind=hidden,
no sidecar) would keep re-triggering repair_empty_non_final_messages on every
call forever. Substitute [response interrupted] on the wire copy at the
api_content/display_kind projection stage so legacy sessions converge too.
Never the interrupt scaffold (NousResearch#81841). Durable transcript untouched.

Regression tests drive run_conversation end-to-end with a spied sanitizer:
the projection must leave the sanitizer nothing to heal (its per-turn warning
spam is the bug), verified failing via sabotage run against the writer-only
fix.

Projection-side approach credit: @JoaoMarcos44 (PR NousResearch#88996).
bobaba76 pushed a commit to bobaba76/hermes-agent that referenced this pull request Aug 27, 2026
…payload at projection time (NousResearch#88955)

The salvaged writer-side fix stamps api_content on NEW hidden redirect
placeholders, but rows persisted before it (content="" + display_kind=hidden,
no sidecar) would keep re-triggering repair_empty_non_final_messages on every
call forever. Substitute [response interrupted] on the wire copy at the
api_content/display_kind projection stage so legacy sessions converge too.
Never the interrupt scaffold (NousResearch#81841). Durable transcript untouched.

Regression tests drive run_conversation end-to-end with a spied sanitizer:
the projection must leave the sanitizer nothing to heal (its per-turn warning
spam is the bug), verified failing via sabotage run against the writer-only
fix.

Projection-side approach credit: @JoaoMarcos44 (PR NousResearch#88996).
melon-xf added a commit to melon-xf/hermes-agent that referenced this pull request Sep 3, 2026
…payload at projection time (NousResearch#88955)

The salvaged writer-side fix stamps api_content on NEW hidden redirect
placeholders, but rows persisted before it (content="" + display_kind=hidden,
no sidecar) would keep re-triggering repair_empty_non_final_messages on every
call forever. Substitute [response interrupted] on the wire copy at the
api_content/display_kind projection stage so legacy sessions converge too.
Never the interrupt scaffold (NousResearch#81841). Durable transcript untouched.

Regression tests drive run_conversation end-to-end with a spied sanitizer:
the projection must leave the sanitizer nothing to heal (its per-turn warning
spam is the bug), verified failing via sabotage run against the writer-only
fix.

Projection-side approach credit: @JoaoMarcos44 (PR NousResearch#88996).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/sessions Session lifecycle, resume, persistence, history comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint P2 Medium — degraded but workaround exists sweeper:risk-session-state Sweeper risk: may lose/corrupt/mis-associate session or context state type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Bot Mode group chat: interrupted member turns persist empty display_kind=hidden assistant messages that re-trigger the pre-call sanitizer every turn

3 participants