Skip to content

refactor(streaming): split terminal-failure transcript evaluator (#5141) - #5272

Merged
1 commit merged into
nesquena:masterfrom
nankingjing:fix/5141-terminal-failure-transcript-evaluator
Jun 30, 2026
Merged

1 commit merged into
nesquena:masterfrom
nankingjing:fix/5141-terminal-failure-transcript-evaluator

Conversation

@nankingjing

Copy link
Copy Markdown
Contributor

Summary

  • Extract _turn_transcript_lacks_final_assistant_answer() to evaluate an already-merged transcript (pending-user materialization, replayed-assistant filtering, final-answer check).
  • Keep _merged_transcript_lacks_final_assistant_answer() as a thin merge → evaluate wrapper so #5129 / #5121 settlement semantics stay unchanged.

Fixes #5141

Why

The terminal-failure path still needs its own merge inputs (_all_result_messages, replay filtering), but the evaluation logic is now reusable without repeating the full merge body.

Test plan

  • tests/test_issue5141_terminal_failure_transcript_evaluator.py (4 cases: wrapper parity, final-answer path, pending-user materialization, delegation)
  • CI: existing #5121 regressions (tests/test_issue5121_provider_auth_terminal_error.py)

…quena#5141)

Extract _turn_transcript_lacks_final_assistant_answer() so settlement
logic can inspect an already-merged transcript. The merge wrapper now
delegates to the pure evaluator without changing behavior.

Fixes nesquena#5141
@greptile-apps

greptile-apps Bot commented Jun 30, 2026

Copy link
Copy Markdown
Contributor

Greptile Summary

This PR splits the monolithic _merged_transcript_lacks_final_assistant_answer into two layers: a thin merge-and-delegate wrapper and a new _turn_transcript_lacks_final_assistant_answer function that evaluates an already-merged transcript, making the evaluation logic reusable by the terminal-failure path without repeating the merge plumbing.

  • api/streaming.py: Extracted evaluation body (user-turn lookup, pending-user materialization, replay filtering, final-answer check) into _turn_transcript_lacks_final_assistant_answer; the original wrapper now calls _merge_display_messages_after_agent_result then delegates to the new function with the same arguments, preserving all existing semantics.
  • tests/test_issue5141_terminal_failure_transcript_evaluator.py: Four new tests covering wrapper/direct parity, the final-answer path, pending-user materialization, and delegation; the drop_replayed_assistant=True filtering branch inside the extracted function is not exercised directly.

Confidence Score: 4/5

The production code changes are a straight extraction of existing logic with no behavioral difference; the wrapper still delegates through the same merge pipeline to the same evaluation steps.

The streaming module change is a clean cut-and-paste refactor — no logic was altered, arguments flow identically, and the call site at line 8135 is untouched. The only real gaps are in the test file: the drop_replayed_assistant filtering branch of the extracted function has no direct test case, and the delegation test merged_len assertion is too loose to catch a broken merge. Neither affects runtime behavior of the current PR.

tests/test_issue5141_terminal_failure_transcript_evaluator.py — the new test suite would benefit from a direct case exercising the replay-filtering branch of _turn_transcript_lacks_final_assistant_answer.

Important Files Changed

Filename Overview
api/streaming.py Evaluation body extracted from _merged_transcript_lacks_final_assistant_answer into _turn_transcript_lacks_final_assistant_answer with no logic changes; wrapper now delegates cleanly.
tests/test_issue5141_terminal_failure_transcript_evaluator.py Four tests cover wrapper parity, final-answer path, pending-user materialization, and delegation. The drop_replayed_assistant=True filtering path in the extracted function is not directly tested, and the delegation test's merged_len assertion is overly loose.

Sequence Diagram

%%{init: {'theme': 'neutral'}}%%
sequenceDiagram
    participant Caller as Terminal Failure Path
    participant Wrapper as _merged_transcript_lacks_final_assistant_answer
    participant Merge as _merge_display_messages_after_agent_result
    participant Restore as _restore_reasoning_metadata
    participant Eval as _turn_transcript_lacks_final_assistant_answer
    participant Session as _session_lacks_final_assistant_answer

    Caller->>Wrapper: (previous_display, previous_context, result_messages, msg_text, source, drop_replayed_assistant)
    Wrapper->>Restore: (previous_display, result_messages)
    Restore-->>Wrapper: enriched result_messages
    Wrapper->>Merge: (previous_display, previous_context, enriched_result_messages, msg_text, source)
    Merge-->>Wrapper: merged_messages
    Wrapper->>Eval: (merged_messages, previous_display, msg_text, source, drop_replayed_assistant)
    Note over Eval: Finds current user turn index, materializes pending user if needed, optionally filters replayed assistant msgs
    Eval->>Session: (filtered_messages)
    Session-->>Eval: bool
    Eval-->>Wrapper: bool
    Wrapper-->>Caller: bool
Loading
%%{init: {'theme': 'base', 'themeVariables': {"darkMode": true, "background": "#0d1117", "primaryColor": "#21262d", "primaryTextColor": "#e6edf3", "primaryBorderColor": "#8b949e", "lineColor": "#8b949e", "textColor": "#e6edf3", "edgeLabelBackground": "#161b22", "actorBkg": "#21262d", "actorBorder": "#8b949e", "actorTextColor": "#e6edf3", "actorLineColor": "#8b949e", "signalColor": "#8b949e", "signalTextColor": "#e6edf3", "noteBkgColor": "#373320", "noteBorderColor": "#d4a72c", "noteTextColor": "#f0e6c0", "labelBoxBkgColor": "#21262d", "labelBoxBorderColor": "#8b949e", "labelTextColor": "#e6edf3", "loopTextColor": "#e6edf3", "activationBkgColor": "#30363d", "activationBorderColor": "#8b949e"}}}%%
sequenceDiagram
    participant Caller as Terminal Failure Path
    participant Wrapper as _merged_transcript_lacks_final_assistant_answer
    participant Merge as _merge_display_messages_after_agent_result
    participant Restore as _restore_reasoning_metadata
    participant Eval as _turn_transcript_lacks_final_assistant_answer
    participant Session as _session_lacks_final_assistant_answer

    Caller->>Wrapper: (previous_display, previous_context, result_messages, msg_text, source, drop_replayed_assistant)
    Wrapper->>Restore: (previous_display, result_messages)
    Restore-->>Wrapper: enriched result_messages
    Wrapper->>Merge: (previous_display, previous_context, enriched_result_messages, msg_text, source)
    Merge-->>Wrapper: merged_messages
    Wrapper->>Eval: (merged_messages, previous_display, msg_text, source, drop_replayed_assistant)
    Note over Eval: Finds current user turn index, materializes pending user if needed, optionally filters replayed assistant msgs
    Eval->>Session: (filtered_messages)
    Session-->>Eval: bool
    Eval-->>Wrapper: bool
    Wrapper-->>Caller: bool
Loading

Reviews (1): Last reviewed commit: "refactor(streaming): split terminal-fail..." | Re-trigger Greptile

Comment on lines +99 to +128
"merged_len": len(list(merged_messages or [])),
"previous_len": len(list(previous_display or [])),
"msg_text": msg_text,
"source": source,
"drop_replayed_assistant": drop_replayed_assistant,
}
)
return True

previous_display = [{"role": "user", "content": "hello"}]
with mock.patch.object(
streaming,
"_turn_transcript_lacks_final_assistant_answer",
side_effect=_fake_evaluator,
):
result = streaming._merged_transcript_lacks_final_assistant_answer(
previous_display,
previous_display,
previous_display,
"hello",
source="cli",
drop_replayed_assistant=True,
)
assert result is True
assert len(calls) == 1
assert calls[0]["previous_len"] == 1
assert calls[0]["msg_text"] == "hello"
assert calls[0]["source"] == "cli"
assert calls[0]["drop_replayed_assistant"] is True
assert calls[0]["merged_len"] >= 1

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Delegation test makes no coverage claim on drop_replayed_assistant=True logic

The delegation test (test_merged_wrapper_delegates_to_turn_evaluator) mocks out _turn_transcript_lacks_final_assistant_answer entirely, so it only verifies that the flag is forwarded — not that the replay-filtering branch inside the extracted function behaves correctly. The three wrapper-parity tests all call with drop_replayed_assistant=False. A case where a replayed assistant message would spuriously satisfy the "final answer" check without the filter is the exact failure mode the terminal-failure path guards against, and it has no direct coverage in this new test file. The existing test_issue5121_provider_auth_terminal_error.py suite presumably covers the regression path, but a direct case exercising the filtering logic in _turn_transcript_lacks_final_assistant_answer would close the gap opened by the extraction.

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

assert calls[0]["msg_text"] == "hello"
assert calls[0]["source"] == "cli"
assert calls[0]["drop_replayed_assistant"] is True
assert calls[0]["merged_len"] >= 1

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Loose assertion on merged_len — >= 1 accepts an empty merge result that happens to contain just the input message. Consider asserting the exact expected post-merge length so a regression in _merge_display_messages_after_agent_result is detected by this test rather than silently passing.

Suggested change
assert calls[0]["merged_len"] >= 1
assert calls[0]["merged_len"] == len(previous_display) # merge of identical inputs yields previous_display length

@nesquena-hermes

Copy link
Copy Markdown
Collaborator

🔬 Gate certification — GREEN ✅

Certified head: sha:c2bb4c65 · PR: #5272 · nankingjing — "refactor(streaming): split terminal-failure transcript evaluator (#5141)" · crit=3
Verdict: Clean behavior-preserving refactor of the settlement path — extracts the merged-transcript evaluator into a reusable _turn_transcript_lacks_final_assistant_answer, with the original name becoming a thin delegating wrapper. Crucially, it does NOT disturb the result['error'] 3-state sentinel or the _terminal_failure decision. Codex SAFE, suite green. Ready to merge.

What I ran (isolated worktree /tmp/wt-gate-5272, fresh-based on current master)

Gate Result
Codex (reproduce) SAFE TO SHIP (no findings — extraction truly behavior-preserving, sentinel/_terminal_failure unaltered)
Full pytest suite (-p no:xdist) 11188 passed, 0 failed
PR's tests (test_issue5141, +4) 4 passed (wrapper matches evaluator: no-replay-filter, with-final-answer, pending-user-after-boundary, delegation)
Sentinel/settlement-contract tests (non_auth_silent_failure, issue5121, 5224 terminal-transcript) 17 passed — the error:"" silent-failure sentinel convention is preserved

Findings (clean)

Concept: 4/5 — a maintainability refactor (#5141) that makes the terminal-failure evaluator reusable/testable without behavior change; reduces future-bug surface in a notoriously subtle area (good).

Recommendation to the next agent

Ready to merge — cert fresh for sha:c2bb4c65. Behavior-preserving settlement-path refactor; Codex SAFE, full suite green, 4 new evaluator tests + 17 sentinel/settlement-contract tests pass (the error:"" silent-failure sentinel preserved — the exact tripwire in this area). No regression to #5224/#5245/#5121. Concept 4/5. Cert valid only at sha:c2bb4c65.


Gate-certifier layer (warm-up → gate → release). I do not merge/tag/deploy/close. Cert valid only at sha:c2bb4c65.

@nesquena-hermes nesquena-hermes added gate-pass Full gate passed (Codex+Opus+suite+browser); queued Tier 1 for release agent size:M Medium PR (≤10 files, ≤250 LOC) labels Jun 30, 2026
@nesquena-hermes nesquena-hermes closed this pull request by merging all changes into nesquena:master in 8d5ff4c Jun 30, 2026
@nesquena-hermes

Copy link
Copy Markdown
Collaborator

Shipped in v0.51.765 (Wave 1 batch, via release PR #5277). Thanks @nankingjing! Combined-stage Codex SAFE + full pytest suite green + ESLint/ruff gates clean. 🎉

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

gate-pass Full gate passed (Codex+Opus+suite+browser); queued Tier 1 for release agent size:M Medium PR (≤10 files, ≤250 LOC)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Follow up terminal failure settlement to avoid duplicate transcript merge

2 participants