Skip to content

fix(worker): treat empty LLM response after text output as completion - #1677

Merged
serrrfirat merged 3 commits into
nearai:stagingfrom
j-bloggs:fix/empty-response-completion
Mar 28, 2026
Merged

serrrfirat merged 3 commits into
nearai:stagingfrom
j-bloggs:fix/empty-response-completion

Conversation

@j-bloggs

@j-bloggs j-bloggs commented Mar 26, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • Add LlmError::EmptyResponse variant for when LLM returns no content
  • Update nearai_chat and github_copilot providers to emit EmptyResponse instead of InvalidResponse for empty/no-choice responses
  • try_complete_on_error now only swallows EmptyResponse (not AuthFailed, ContextLengthExceeded, Http, Io, etc.)
  • Extract is_completion_eligible_error as testable pure function
  • Extract mark_completed_or_warn helper to DRY completion logic
  • Log mark_completed errors at warn level instead of silently dropping
  • Add EmptyResponse to retry and circuit breaker transient classifications

Change Type

  • Bug fix (non-breaking change which fixes an issue)

Linked Issue

Addresses review feedback from zmanian and gemini-code-assist on this PR.

Validation

  • cargo fmt --all -- --check: pass
  • cargo clippy (all-features, default, libsql): pass (zero warnings)
  • cargo test --lib -- is_completion_eligible: 1 test pass
  • cargo test --lib -- retry::tests circuit_breaker::tests: 34 tests pass

Security Impact

None - error handling change only.

Database Impact

None

Blast Radius

  • LLM providers now emit EmptyResponse instead of InvalidResponse for no-choice responses
  • Jobs with text output that hit non-empty errors (AuthFailed, Http, etc.) will now propagate the error instead of silently completing
  • retry and circuit_breaker treat EmptyResponse as transient (retryable)

Rollback Plan

Revert to staging behavior where all LLM errors propagate (re-introduces spurious failure bug but eliminates silent error swallowing).

Review Track

Track C - runtime changes in src/worker/ and src/llm/

Feature Parity

No FEATURE_PARITY.md changes needed.

@github-actions github-actions Bot added scope: worker Container worker size: M 50-199 changed lines labels Mar 26, 2026
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request addresses a critical bug in the worker's agentic loop where jobs would erroneously be marked as failed. The core problem stemmed from the system's inability to differentiate between a genuine LLM failure and a natural conclusion of the LLM's output (e.g., an empty response after providing all necessary information). By introducing a state flag to track prior substantive text output, the system can now correctly interpret subsequent empty responses or errors as successful completion, thereby ensuring accurate job status reporting and preventing unnecessary retries.

Highlights

  • Bug Fix: Incorrect Job Failure Reporting: Resolved an issue where routine and background jobs using LLMs would incorrectly report as 'failed' even after successfully producing results, particularly when the LLM subsequently returned an empty response or an error.
  • Improved LLM Response Handling: Implemented logic to treat subsequent empty LLM responses or provider errors as job completion, rather than triggering retries leading to failure, once a non-empty text response has already been produced.
  • New has_text_response Flag: Introduced a has_text_response AtomicBool flag within the JobDelegate to track whether substantive output has been generated by the LLM, enabling the new completion logic.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@github-actions github-actions Bot added risk: medium Business logic, config, or moderate-risk modules contributor: new First-time contributor labels Mar 26, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a has_text_response flag to the JobDelegate to refine job completion logic. This flag ensures that once a substantive text response has been generated, subsequent empty responses or LLM errors are treated as successful job completion rather than fatal failures or retry signals. The error handling in select_tools and respond_with_tools methods, as well as the handle_text_response method, have been updated to incorporate this new logic. A new test case has been added to validate this behavior. The review suggests refactoring duplicated error handling logic into a helper function to improve maintainability.

Comment thread src/worker/job.rs
@j-bloggs
j-bloggs force-pushed the fix/empty-response-completion branch from 22e5dbf to be47024 Compare March 26, 2026 11:27

@zmanian zmanian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review: fix(worker): treat empty LLM response after text output as completion

The intent is valid but the implementation is too aggressive.

High

  1. try_complete_on_error swallows ALL error variants: AuthFailed, ContextLengthExceeded, ModelNotAvailable, Io -- all treated as "job done" after any text output. Configuration and infrastructure errors should never be silently swallowed. Restrict to empty-response-like errors only.

  2. Premature termination of multi-step jobs: A job that produces intermediate text ("I'll fetch the data now") then hits a transient Http error on the next tool-selection call gets marked complete instead of retrying.

Medium

  1. Test doesn't exercise actual code: empty_response_after_text_signals_completion reimplements the decision logic with string comparisons rather than testing JobDelegate::handle_text_response. Would pass even if the real implementation were broken.

  2. mark_completed errors silently dropped: let _ = self.worker.mark_completed().await -- existing code at line 1315 logs the error; new code drops it.

Recommendations

  • Restrict error swallowing to specific variants (e.g., InvalidResponse with empty body)
  • Raise the completion threshold (e.g., text + no pending tool calls, or two consecutive empty responses)
  • Test the actual JobDelegate methods, not a reimplemented copy

@github-actions github-actions Bot added the scope: llm LLM integration label Mar 27, 2026
j-bloggs and others added 3 commits March 27, 2026 23:03
When a job's LLM produces a substantive text response (e.g., formatted
results from a routine) and the next LLM call returns empty or errors,
the worker now treats this as successful completion instead of
continuing the loop until failure.

Previously, empty responses always triggered TextAction::Continue,
causing the loop to re-call the LLM. The LLM had nothing more to say,
so the provider returned "Response contained no message or tool call
(empty)". This made routine jobs that successfully produced results
report as "failed".

The fix adds a `has_text_response` flag to JobDelegate:
- After any non-empty text response: flag is set
- Empty text after flag is set: treated as completion
- LLM errors (select_tools/respond_with_tools) after flag: treated
  as completion instead of propagating
- Empty text before any output: still retries (rate-limit backoff)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- Add LlmError::EmptyResponse variant for when LLM returns no content
- Update nearai_chat and github_copilot providers to emit EmptyResponse
  instead of InvalidResponse for empty/no-choice responses
- try_complete_on_error now only swallows EmptyResponse (not AuthFailed,
  ContextLengthExceeded, Http, Io, etc.)
- Extract is_completion_eligible_error as testable pure function
- Log mark_completed errors at warn level instead of silently dropping
- Add EmptyResponse to retry and circuit breaker transient classifications
- Rewrite test to exercise real classification logic against all variants

Addresses review feedback from zmanian and gemini-code-assist.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…tion logic

Extract shared mark-completed + warn-on-failure pattern into a single
helper method used by both try_complete_on_error and handle_text_response.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
@j-bloggs
j-bloggs force-pushed the fix/empty-response-completion branch from aa99a06 to 67d941f Compare March 27, 2026 12:07

@zmanian zmanian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-review: empty LLM response completion

Significant improvement. Core safety concern resolved.

Finding Previous Current
Error swallowing scope HIGH Fixed -- only EmptyResponse, not all errors
Premature termination HIGH Medium -- edge case in interleaved text+tool jobs
Test coverage MEDIUM Fixed -- tests actual is_completion_eligible_error()

The EmptyResponse-only gating is the right approach. The remaining edge case (premature completion during multi-step jobs) is acceptable for the routine/heartbeat use case. Approve.

@serrrfirat
serrrfirat merged commit fd41bdf into nearai:staging Mar 28, 2026
14 checks passed
DougAnderson444 pushed a commit to DougAnderson444/ironclaw that referenced this pull request Mar 29, 2026
…nearai#1677)

* fix(worker): treat empty LLM response after text output as completion

When a job's LLM produces a substantive text response (e.g., formatted
results from a routine) and the next LLM call returns empty or errors,
the worker now treats this as successful completion instead of
continuing the loop until failure.

Previously, empty responses always triggered TextAction::Continue,
causing the loop to re-call the LLM. The LLM had nothing more to say,
so the provider returned "Response contained no message or tool call
(empty)". This made routine jobs that successfully produced results
report as "failed".

The fix adds a `has_text_response` flag to JobDelegate:
- After any non-empty text response: flag is set
- Empty text after flag is set: treated as completion
- LLM errors (select_tools/respond_with_tools) after flag: treated
  as completion instead of propagating
- Empty text before any output: still retries (rate-limit backoff)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(worker): restrict error swallowing to EmptyResponse variant only

- Add LlmError::EmptyResponse variant for when LLM returns no content
- Update nearai_chat and github_copilot providers to emit EmptyResponse
  instead of InvalidResponse for empty/no-choice responses
- try_complete_on_error now only swallows EmptyResponse (not AuthFailed,
  ContextLengthExceeded, Http, Io, etc.)
- Extract is_completion_eligible_error as testable pure function
- Log mark_completed errors at warn level instead of silently dropping
- Add EmptyResponse to retry and circuit breaker transient classifications
- Rewrite test to exercise real classification logic against all variants

Addresses review feedback from zmanian and gemini-code-assist.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* refactor(worker): extract mark_completed_or_warn helper to DRY completion logic

Extract shared mark-completed + warn-on-failure pattern into a single
helper method used by both try_complete_on_error and handle_text_response.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: j-bloggs <j-bloggs@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
drchirag1991 pushed a commit to drchirag1991/ironclaw that referenced this pull request Apr 8, 2026
…nearai#1677)

* fix(worker): treat empty LLM response after text output as completion

When a job's LLM produces a substantive text response (e.g., formatted
results from a routine) and the next LLM call returns empty or errors,
the worker now treats this as successful completion instead of
continuing the loop until failure.

Previously, empty responses always triggered TextAction::Continue,
causing the loop to re-call the LLM. The LLM had nothing more to say,
so the provider returned "Response contained no message or tool call
(empty)". This made routine jobs that successfully produced results
report as "failed".

The fix adds a `has_text_response` flag to JobDelegate:
- After any non-empty text response: flag is set
- Empty text after flag is set: treated as completion
- LLM errors (select_tools/respond_with_tools) after flag: treated
  as completion instead of propagating
- Empty text before any output: still retries (rate-limit backoff)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(worker): restrict error swallowing to EmptyResponse variant only

- Add LlmError::EmptyResponse variant for when LLM returns no content
- Update nearai_chat and github_copilot providers to emit EmptyResponse
  instead of InvalidResponse for empty/no-choice responses
- try_complete_on_error now only swallows EmptyResponse (not AuthFailed,
  ContextLengthExceeded, Http, Io, etc.)
- Extract is_completion_eligible_error as testable pure function
- Log mark_completed errors at warn level instead of silently dropping
- Add EmptyResponse to retry and circuit breaker transient classifications
- Rewrite test to exercise real classification logic against all variants

Addresses review feedback from zmanian and gemini-code-assist.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* refactor(worker): extract mark_completed_or_warn helper to DRY completion logic

Extract shared mark-completed + warn-on-failure pattern into a single
helper method used by both try_complete_on_error and handle_text_response.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: j-bloggs <j-bloggs@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: new First-time contributor risk: medium Business logic, config, or moderate-risk modules scope: llm LLM integration scope: worker Container worker size: M 50-199 changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants