Skip to content

Hide process wakeup prompts from transcripts - #5245

Closed
santastabber wants to merge 1 commit into
nesquena:masterfrom
santastabber:fix-process-wakeup-transcript-pollution
Closed

santastabber wants to merge 1 commit into
nesquena:masterfrom
santastabber:fix-process-wakeup-transcript-pollution

Conversation

@santastabber

Copy link
Copy Markdown
Contributor

Summary

  • keep background process wakeup prompts internal instead of materializing them as visible user transcript rows
  • filter legacy process_wakeup user rows from display windows, message summaries, and renderable-message counts
  • add focused regressions for recovery, eager checkpointing, display merge, pagination, and legacy hidden-row filtering

Tests

  • git diff --check origin/master...HEAD
  • python3 -m py_compile api/models.py api/routes.py api/streaming.py
  • ./scripts/test.sh tests/test_process_wakeup_synthetic.py -q
  • trufflehog filesystem --no-update --fail --only-verified api/models.py api/routes.py api/streaming.py tests/test_process_wakeup_synthetic.py
  • gitleaks detect --no-git --redact --verbose --source /tmp/webui-process-wakeup-committed.diff --report-path /tmp/webui-process-wakeup-gitleaks-committed.json --report-format json

@greptile-apps

greptile-apps Bot commented Jun 30, 2026 •

Copy link
Copy Markdown
Contributor

Greptile Summary

This PR changes process-wakeup prompts from materialized visible user rows to fully internal events — they are now suppressed at every persistence, recovery, display-merge, and metadata-count boundary instead of being stamped and filtered after the fact.

  • Shared predicates is_hidden_process_wakeup_message / filter_hidden_internal_messages are added to models.py and imported everywhere; _append_recovered_pending_turn, eager checkpointing, _materialize_pending_user_turn_before_error, and the full-load handle_get path all return/skip early for process_wakeup sources.
  • get_state_db_session_summary gains a Python-loop counting path when the _source column is present, and _metadata_only_message_summary gains a file-scan heuristic to correct inflated sidecar counts for legacy sessions that already persisted wakeup rows.
  • _merge_display_messages_after_agent_result and the new _merged_transcript_lacks_final_assistant_answer branch filter wakeup rows from display lists and from the settlement check, with a replay-dedup guard that avoids over-dropping on stable-id assistant messages.

Confidence Score: 3/5

The new process_wakeup settlement branch in _merged_transcript_lacks_final_assistant_answer can declare a session missing its final answer when the session also contains duplicate _partial messages from the pre-existing cancel_stream accumulation bug.

The tail slice uses len(previous_visible) as offset into merged_messages, but the merge function also deduplicates stale _partial entries not reflected in previous_visible. Sessions with cancel_stream duplicate partials receiving a process_wakeup will have tail_messages overshoot the new answer, causing false settlement failures and potential retry loops.

api/streaming.py — the _merged_transcript_lacks_final_assistant_answer process_wakeup branch where tail_messages = merged_messages[len(previous_visible):] needs to account for _partial deduplication.

Important Files Changed

Filename Overview
api/models.py Adds shared is_hidden_process_wakeup_message / filter_hidden_internal_messages predicates; gates _append_recovered_pending_turn against process_wakeup sources; adds _source to the column selection and a Python-loop counting path in get_state_db_session_summary so state.db metadata correctly excludes hidden rows.
api/routes.py Imports shared helpers; gates eager checkpointing against process_wakeup; adds full-sidecar scan in _metadata_only_message_summary to correct inflated counts; filters hidden rows in _message_summary and before _message_window_for_display. Sidecar text scan has a false-positive risk on literal user text.
api/streaming.py Filters wakeup rows throughout _merge_display_messages_after_agent_result; adds a dedicated process_wakeup branch in _merged_transcript_lacks_final_assistant_answer. The tail-slice offset len(previous_visible) is inconsistent with the _partial deduplication that the merge function applies, which can cause settlement to misreport a missing answer.
tests/test_process_wakeup_synthetic.py Inverts or renames pre-existing tests to match the new hide-not-stamp contract; adds thorough new regressions for renderable-count, filter, windowing, settlement, replay-drop, and backfill suppression.
tests/test_webui_state_db_reconciliation.py Adds _source column to test schema helpers and two new metadata-only summary tests covering state.db and sidecar legacy-wakeup-row scenarios.

Reviews (3): Last reviewed commit: "Hide process wakeup prompts from transcr..." | Re-trigger Greptile

Comment thread api/routes.py Outdated
@nesquena-hermes

Copy link
Copy Markdown
Collaborator

Reading the diff at HEAD (dfe6dd56) against origin/master, this is a clean, comprehensive suppression pass — six guard points across api/models.py:506, api/streaming.py (5 sites), and the eager checkpoint in api/routes.py:17167, plus a defense-in-depth display filter. The write-path coverage looks correct and the flipped test semantics in tests/test_process_wakeup_synthetic.py match the new "skip entirely" behavior.

I want to confirm Greptile's P1 because I think it's a genuine, persistent divergence (not just a transient poll artifact), and ground it in the code.

Code reference

The full-load summary now filters:

# api/routes.py:7007-7008
def _message_summary(messages) -> dict:
    messages = _filter_hidden_internal_messages(messages)

But the metadata-only path derives its count from stored JSON, not from a filtered message list:

# api/routes.py:7037-7040  (_metadata_only_message_summary)
sidecar_count = _numeric_count(getattr(sidecar_session, "_metadata_message_count", None))
if sidecar_count <= 0:
    sidecar_count = _numeric_count(sidecar_session.compact().get("message_count"))

And the stored count is written from the raw, unpruned message list:

# api/models.py:882
meta['message_count'] = len(self.messages or [])

Diagnosis

The key point: _filter_hidden_internal_messages (routes.py:17151) is a read/display filter — it never prunes the legacy process_wakeup row out of self.messages. So for any session that materialized a wakeup row under the old behavior:

  • _message_summary / the _all_msgs window (routes.py:7008, 10451) report the filtered count (N).
  • compact() at models.py:882 persists message_count = len(self.messages) = N+1, which then feeds _metadata_only_message_summary and the _metadata_message_count fallback in handle_get (routes.py:10442).

On origin/master both paths counted the wakeup row, so they agreed (both inflated). This PR makes the full-load path filter while leaving the metadata path unfiltered, so it introduces a persistent +1 mismatch between the sidebar tile and the opened transcript for legacy-polluted sessions. It won't self-heal on save either, since the row stays in self.messages and gets re-counted at models.py:882.

Recommendation

Two viable approaches:

  1. Symmetry (cheap, no migration): apply the same hidden-row discount in _metadata_only_message_summary. Since that path intentionally avoids a full transcript read, the sidecar JSON would need a stored "renderable count" or a hidden-row count to subtract — otherwise you can't know how many wakeup rows are in there without loading. That makes option 2 cleaner.

  2. Prune-on-load (self-healing): strip hidden wakeup rows from self.messages once, on load/save, so len(self.messages) at models.py:882 is already correct and both paths agree without a stored side count. A guarded prune in the load path (or a one-time migration) means new saves persist the filtered count and the divergence disappears permanently.

I'd lean toward option 2 — it keeps a single source of truth for the count and avoids teaching the cheap metadata path about hidden rows.

Test plan

A regression for the metadata path would catch this: build a session whose stored messages/_metadata_message_count include a _source='process_wakeup' user row, then assert _metadata_only_message_summary(sid)["message_count"] equals _message_summary(all_msgs)["message_count"]. The current new tests exercise _message_summary, _filter_hidden_internal_messages, and _message_window_for_display, but none cross-check the metadata-count path against the filtered full-load count, which is exactly where the divergence lives.

@santastabber
santastabber force-pushed the fix-process-wakeup-transcript-pollution branch from dfe6dd5 to a525c3b Compare June 30, 2026 01:12
@santastabber

Copy link
Copy Markdown
Contributor Author

Thanks — addressed in the updated head a525c3b2.

What changed:

  • moved the hidden process_wakeup predicate/filter into api.models so routes.py and streaming.py share the same source of truth;
  • aligned metadata-only summaries with the display filter for legacy process_wakeup rows, including sidecar and state-db coverage;
  • added regressions for both legacy sidecar rows and state-db rows.

Verification:

  • git diff --check
  • python3 -m py_compile api/models.py api/routes.py api/streaming.py tests/test_process_wakeup_synthetic.py tests/test_webui_state_db_reconciliation.py
  • ./scripts/test.sh tests/test_process_wakeup_synthetic.py tests/test_webui_state_db_reconciliation.py -q (56 passed)
  • TruffleHog changed files: no verified/unverified secrets
  • Gitleaks follow-up diff: no leaks found

@nesquena-hermes

Copy link
Copy Markdown
Collaborator

Gated this fresh (Codex reproduce + full suite) — the display-hiding is on the right track, but it introduces a CORE settlement regression I reproduced directly, so one fix is needed before it ships.

The bug (reproduced in-process): hiding the process_wakeup user prompt breaks the turn-settlement check. With the prompt hidden, _merged_transcript_lacks_final_assistant_answer() (api/streaming.py:~5162) can no longer find the current user turn, so it appends a synthetic wakeup user row after the kept assistant response (lines ~5167-5174); _session_lacks_final_assistant_answer() then sees the transcript ending in a user request and reports "lacks final answer". That feeds _terminal_failure (api/streaming.py:~8143-8149) — so a successful process-wakeup turn goes down the app-error path instead of normal done.

Direct repro against your branch:

_merged_transcript_lacks_final_assistant_answer(
    [], [], [{"role":"assistant","content":"done"}], "wakeup prompt", source="process_wakeup"
)  # returns True  -> should be False (a kept assistant answer IS a final answer)

Fix-spec (Codex, verified):

  1. api/streaming.py _merged_transcript_lacks_final_assistant_answer() — when _is_process_wakeup_source(source), do NOT append the synthetic pending user row after a kept assistant response. Judge the post-previous_display merged visible tail directly: return "lacks final answer" only when that visible tail has no final assistant answer.
  2. tests/test_process_wakeup_synthetic.py — add a regression asserting:
    _merged_transcript_lacks_final_assistant_answer([], [], [{"role":"assistant","content":"done"}], wakeup_prompt, source="process_wakeup") is False.

Note this is exactly the streaming-settlement-contract area where the 3-state error sentinel lives — keep the existing empty-string-sentinel tests green (a successful wakeup must settle as done, not apperror/no_response). Everything else (the display-window/summary/renderable-count filtering, legacy state.db reconciliation) gated clean. Re-gate on re-push. Thanks @santastabber.

@nesquena-hermes nesquena-hermes added the changes-requested Maintainer left detailed feedback requesting changes; PR is waiting on author to address label Jun 30, 2026
@nesquena-hermes

Copy link
Copy Markdown
Collaborator

🔬 Gate certification — RED ⛔ (wakeup-hiding mostly right, but 2 reproduced streaming-settlement edge bugs)

Certified head: sha:a525c3b2feca (rebased 1ee36c8a) · PR: #5245 · santastabber — "Hide process wakeup prompts from transcripts" · crit=3
Verdict: The core idea is sound and most of it is right — role=user + _source=process_wakeup rows are filtered from the visible transcript/count (incl. legacy state.db rows), and it does NOT disturb the error 3-state settlement sentinel (I verified the issue5121/non_auth_silent_failure tests pass). But Codex reproduced 2 real edge bugs in the streaming-settlement path: a successful wakeup turn can be settled as a failure (apperror), and one backfill branch can still re-leak a legacy hidden row. Both have precise fixes.

What I ran (isolated worktree /tmp/wt-gate-5245, rebased onto current master)

Gate Result
Codex (reproduce) SHIP ONLY WITH FIXES — 1 CORE (success→apperror) + 1 SILENT (tail-backfill re-leak), both verified
Full pytest suite (-p no:xdist, rebased) 11155 passed, 0 failed (green — but no test covers the assistant-only-wakeup-completion settlement or the tail-leak; that's why they slip)
PR's tests + reconciliation 56 passed
error-sentinel/settlement-contract tests (issue5121 / non_auth_silent_failure) 12 passed — change does NOT disturb the 3-state error sentinel ✓

⛔ Blocking #1 (CORE, Codex-reproduced + I confirmed in source) — successful wakeup turn settled as a terminal FAILURE

_merged_transcript_lacks_final_assistant_answer() (streaming.py:~5162) unconditionally builds a pending_user row and appends it to merged_messages (pending_user['_source']=source, ~lines 23-29). For a hidden process_wakeup turn, the synthetic user row was correctly removed by the merge, so there's no current-user boundary — and this re-appends the hidden wakeup prompt after the real assistant reply. Then _session_lacks_final_assistant_answer() sees the transcript ending in a user row → _terminal_failure (streaming.py:8143) → the settlement path can emit an apperror despite a real assistant reply. Codex verified the helper returns True for an assistant-only process-wakeup completion. (I confirmed the unconditional append in source.)
Fix (Codex): in _merged_transcript_lacks_final_assistant_answer(), when _is_process_wakeup_source(source) and there's no visible current user row, do NOT append pending_user — evaluate the merged transcript as-is (e.g. return _session_lacks_final_assistant_answer(merged_messages)). Completed process-wakeup turns must not trip _terminal_failure. Add a regression for an assistant-only wakeup completion.

⛔ Blocking #2 (SILENT, Codex-reproduced) — legacy hidden wakeup row re-persisted via the tail-backfill loop

The new not is_hidden_process_wakeup_message(_cmsg) guard was added to two context-backfill branches but NOT the final tail-backfill loop (streaming.py:~4927). So a legacy hidden wakeup row that is the final context-only tail row gets appended into merged → persisted back into s.messages (re-leaks into the durable visible transcript). Codex verified with a direct _merge_display_messages_after_agent_result() call.
Fix (Codex): add and not is_hidden_process_wakeup_message(_cmsg) to the condition at streaming.py:~4927; add a regression where the hidden row is the final context-only tail.

✅ Right (keep it)

  • Identification (role==user && _source==process_wakeup) is precise — a real user turn is never mis-hidden. Filtering applied at the merge, candidate-dedup, 2 backfill branches, pending-turn materialization guard, message-window/count exclusion, and legacy state.db rows. Does NOT touch the error 3-state sentinel (verified). 8 new tests + reconciliation tests pass.

Recommendation to the next agent

Bounce for the 2 streaming-settlement fixes (both Codex-reproduced, precise), then re-gate. The wakeup-hiding concept + identification + the bulk of the filtering are correct, and it correctly leaves the settlement error sentinel alone — but (#1) a successful wakeup turn can be mis-settled as an apperror (re-appended pending row trips _terminal_failure), and (#2) the tail-backfill loop misses the hidden-row guard so a legacy row can re-leak. Both fixes are one-liners + a regression each. crit=3 transcript/settlement correctness → fix before merge. Cert valid only at sha:a525c3b2feca.


Gate-certifier layer (warm-up → gate → release). I do not merge/tag/deploy/close. Cert valid only at sha:a525c3b2feca; a new push invalidates it → re-gate.

@nesquena-hermes nesquena-hermes added the gate-fail Gate found blocking issue(s); fix-spec in comment; awaiting fix/re-push label Jun 30, 2026
@santastabber

Copy link
Copy Markdown
Contributor Author

Addressed the process-wakeup settlement regression in the updated head.

What changed:

  • process_wakeup settlement no longer appends a synthetic hidden user row after a kept assistant response.
  • The settlement check now evaluates the visible merged tail for wakeup turns, so a real assistant response satisfies the final-answer check while an empty/no-answer wakeup still fails.
  • Added regressions for the direct repro, empty result, existing-history tail handling, hidden/marker boundary alignment, and replayed-assistant cases.

Verification:

  • git diff --check
  • python3 -m py_compile api/models.py api/routes.py api/streaming.py tests/test_process_wakeup_synthetic.py tests/test_webui_state_db_reconciliation.py tests/test_auto_compression_terminal_failure.py
  • ./scripts/test.sh tests/test_process_wakeup_synthetic.py tests/test_webui_state_db_reconciliation.py tests/test_auto_compression_terminal_failure.py -q (77 passed)
  • TruffleHog changed files: no verified/unverified secrets
  • Gitleaks final follow-up diff: no leaks found

@santastabber
santastabber force-pushed the fix-process-wakeup-transcript-pollution branch from a525c3b to 1f0c0ca Compare June 30, 2026 02:10
@santastabber

Copy link
Copy Markdown
Contributor Author

Follow-up pushed for the second gate finding as well.

What changed:

  • Added the missing is_hidden_process_wakeup_message(...) guard to the final context tail-backfill loop so legacy hidden wakeup rows cannot be reintroduced from previous_context.
  • Added a direct regression where a hidden process-wakeup row is the final context-only tail row and verified it is not backfilled while the fresh assistant answer remains.

This is in addition to the settlement fix from the previous push.

Verification:

  • git diff --check
  • python3 -m py_compile api/models.py api/routes.py api/streaming.py tests/test_process_wakeup_synthetic.py tests/test_webui_state_db_reconciliation.py tests/test_auto_compression_terminal_failure.py
  • ./scripts/test.sh tests/test_process_wakeup_synthetic.py tests/test_webui_state_db_reconciliation.py tests/test_auto_compression_terminal_failure.py -q (78 passed)
  • TruffleHog changed files: no verified/unverified secrets
  • Gitleaks final follow-up diff: no leaks found

@santastabber
santastabber force-pushed the fix-process-wakeup-transcript-pollution branch from 1f0c0ca to 1dfb762 Compare June 30, 2026 02:15
@nesquena-hermes nesquena-hermes added the size:L Large PR (>10 files or >250 LOC) label Jun 30, 2026
@nesquena-hermes nesquena-hermes removed the changes-requested Maintainer left detailed feedback requesting changes; PR is waiting on author to address label Jun 30, 2026
@nesquena-hermes

Copy link
Copy Markdown
Collaborator

🔬 Gate certification — RED ⛔ (converged on my 2 prior blockers, but a new defect + a structural rebase remain)

Certified head: sha:1dfb7622 · PR: #5245 · santastabber, Hide process wakeup prompts from transcripts
Verdict: The re-push fixed both blockers from my prior bounce (verified empirically) — good convergence. But the gate reproduced a NEW real defect (hidden wakeup rows still sent to the provider), and the PR no longer applies to current master because #5272's settlement refactor landed. Fix the leak + re-integrate onto the refactored structure, then re-gate.

What I ran (worktree on the PR head /tmp/wt-gate-5245, gated on the PR's own base since it doesn't rebase clean)

Gate Result
Prior-blocker verification (mine, in-process) ✅ BOTH fixed (see below)
Codex (reproduce) SHIP-WITH-FIXES — 1 reproduced SILENT (provider leak, confirmed by me) + 1 CORE + 1 SILENT (see below)
Full pytest suite 11161 passed, 0 failed
process_wakeup + reconciliation tests ✅ 64/64
error-sentinel / issue5121 contract ✅ 46/46 (the error:"" 3-state sentinel is untouched)

Findings

✅ CONVERGED — both my prior blockers are fixed (verified):

  1. (was CORE) apperror-on-success: the _is_process_wakeup_source(source) branch now returns early (return _session_lacks_final_assistant_answer(tail_messages)) before the pending_user append, so a successful wakeup turn no longer re-appends the hidden row → no false _terminal_failure. Verified in-process (lacks_final=True correctly for a no-fresh-answer wakeup).
  2. (was SILENT) tail-backfill: is_hidden_process_wakeup_message guard is now present at all three backfill sites (streaming.py:4893/4914/4927). Verified.
    The error-sentinel contract (issue5121) is intact (46/46). Nice work closing both.

⛔ SILENT (I REPRODUCED) — hidden wakeup rows are still sent to the provider (api/streaming.py _sanitize_messages_for_api): the row is hidden from display but _sanitize_messages_for_api() strips _source and never DROPS is_hidden_process_wakeup_message(msg), so a legacy hidden {'role':'user','content':'[IMPORTANT: hidden wakeup ping]','_source':'process_wakeup'} row is passed to the model as normal user history. Confirmed by direct call — the hidden row survives in the sanitized output. Fix: drop is_hidden_process_wakeup_message(msg) rows in _sanitize_messages_for_api() before sanitizing; add a regression covering legacy hidden sidecar/state.db rows. (The point of hiding the wakeup prompt is that the model shouldn't see it either.)

⚠️ Codex also flagged (verify on re-integration, not independently confirmed by me):

  • CORE (5168): the wakeup tail slice merged_messages[len(previous_visible):] may include a backfilled older context-only assistant row if _merge_display_messages_after_agent_result backfills before that boundary → could settle success without a fresh answer. My repro of the simplest case returned the correct True, so I could not confirm it as stated — but the tail-boundary math is worth hardening (compute the tail after the same backfill boundary / exclude pre-existing previous_context rows).
  • SILENT (9852): cancel_stream() recovery path may still persist the hidden wakeup prompt as a user row (bypasses the new skip guards). Verify when re-integrating.

⛔ BLOCKED — structural rebase onto current master (this is the contributor's, not a self-rebase): master's #5272 refactor (which I gate-certified GREEN earlier) split the evaluator into _turn_transcript_lacks_final_assistant_answer(merged_messages, previous_display, msg_text, ...) [receives merged_messages PRE-merged] + a _merged_transcript_... wrapper. This PR was written against the OLD single function — its process-wakeup branch calls _merge_display_messages_after_agent_result(previous_display, previous_context, result_messages, ...) using previous_context/result_messages which are no longer parameters of the extracted function. A git apply onto master conflicts in api/streaming.py. I did NOT self-resolve — placing the wakeup branch in the wrong layer risks a subtle settlement bug on a path that already had edge defects. The contributor should re-integrate the process-wakeup logic into the post-#5272 two-function shape (the wakeup-aware merge + tail-filter belongs in the wrapper that still has those vars).

Recommendation to the next agent

RED — gate-fail/changes-requested. Credit that both original blockers are genuinely fixed + the sentinel contract holds. Remaining: (1) drop hidden wakeup rows in _sanitize_messages_for_api (reproduced provider leak) + regression; (2) re-integrate onto master's #5272-refactored _turn_/_merged_ structure (the structural rebase); (3) while there, harden the tail-boundary (Codex CORE) and the cancel-path persistence (Codex SILENT). Converges fast — the hard settlement logic is mostly right. Author @santastabber. crit=3.


Gate-certifier layer (warm-up → gate → release). I do not merge/tag/deploy. Gated on the PR's own head (1dfb7622) because it structurally collides with post-#5272 master; the rebase is the contributor's (not a self-resolvable CHANGELOG/i18n rebase).

@nesquena-hermes nesquena-hermes added the changes-requested Maintainer left detailed feedback requesting changes; PR is waiting on author to address label Jul 1, 2026
@santastabber santastabber closed this by deleting the head repository Sep 16, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

changes-requested Maintainer left detailed feedback requesting changes; PR is waiting on author to address gate-fail Gate found blocking issue(s); fix-spec in comment; awaiting fix/re-push size:L Large PR (>10 files or >250 LOC)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants