feat(query): robust multi-lingual and structural continuation nudge - #1280
Conversation
fea2ed7 to
47300fa
Compare
…nto feat/robust-continuation-nudge
jatmn
left a comment
There was a problem hiding this comment.
Findings
- [P2] Avoid treating ordinary punctuation-less replies as truncation
src/query.ts:1479
The new truncation heuristic nudges whenever the last assistant text is longer than 20 characters and does not end in.,!,?, a quote, or a backtick, and that branch runs before completion-marker handling. That catches many normal final answers, for exampleNo issues here, LGTM, markdown/list endings like- package.json, or any concise response that simply omits terminal punctuation, so the agent will send up to three extra meta "continue" turns after an otherwise completed answer. Please either tie this path to an actual max-output/truncation signal, make the structural check much narrower, or let completion/non-continuation cases win; it also needs behavioral coverage for punctuation-less completed replies so the core loop does not regress again.
|
done @jatmn thank you ! |
|
I have refined the truncation heuristic in src/query.ts to be significantly narrower. The structural check now:
I have also updated the regression tests in src/tests/bugfixes.test.ts to ensure that punctuation-less completed replies no longer trigger unnecessary meta-turns. Verified locally with a suite covering FR/EN transitions and various edge cases. |
b313125 to
36d491d
Compare
|
I have also further expanded the transition signals to include more direct action markers in both English and French (e.g., processing, starting, je lance, au suivant). This specifically addresses cases where a model might produce a "Mixed Intent" message, such as: "Task 1 is done, processing the next one." By utilizing the previously implemented Late Intent prioritization, the logic now correctly identifies the terminal continuation signal and overrides any earlier completion markers, ensuring the agent remains autonomous even when its phrasing is concise or multi-faceted. All patterns have been consolidated to minimize false positives while maximizing autonomy for non-Claude models. |
jatmn
left a comment
There was a problem hiding this comment.
Findings
- [P2] Keep broad progress words from forcing extra continuation turns
src/query.ts:1470
The new standalone signal regex matches common final-answer words such astesting,next, andfollowing, andhasLateContinuationSignalthen nudges whenever that match appears in the last 80 characters without a later completion marker. Because that path does not require missing terminal punctuation or an explicit "I will continue" phrase, normal completed replies likeTesting passed.,The following files were changed., orThe next step is optional.still become meta "continue" turns even though they are punctuated and complete. This keeps the original false-positive class alive for common review/build summaries, just with different examples. Please narrow those words to explicit action-transition phrases, let normal terminal/completion cases win, and add behavioral tests for punctuated final replies that contain these words.
e7befb5 to
ff89c2c
Compare
|
Hello, Key Enhancements:
These changes significantly improve agent autonomy in multi-step Bazel/Java workflows and French-speaking environments. |
jatmn
left a comment
There was a problem hiding this comment.
Thanks for the updates here. The broad progress-word false positives look narrower now, but I still think the structural truncation path needs another pass before merge.
Findings
- [P2] Let completion markers suppress the truncation fallback
src/query.ts:1492
The newisPossiblyTruncatedbranch still fires before the rest of the completion/non-continuation handling, and it only checks for completion markers in the final 30 characters. That means ordinary completed replies longer than 40 characters still get up to three extra meta "continue" turns when the completion cue appears earlier in the sentence, for exampleThe analysis is complete and no code changes are needed hereorSummary: this PR looks good overall and I have no further comments. Both have explicit completion markers, but neither marker is in the final 30 characters, so they are treated as possible truncation. Please make completion/non-continuation cases win over this structural fallback, or tie the fallback to a stronger truncation signal, and add behavioral coverage for completed punctuation-less replies where the completion marker is not at the very end.
ff89c2c to
d6ba986
Compare
jatmn
left a comment
There was a problem hiding this comment.
Thanks for iterating on this. I still see a couple of issues that need another pass before merge.
Findings
-
[P2] Let completion markers suppress the truncation fallback
src/utils/continuation.ts:49
TheisPossiblyTruncatedbranch still returns before the rest of the completion/continuation logic, and it only checks for completion markers in the final 30 characters. That means completed punctuation-less replies still get up to three extra meta "continue" turns whenever the completion cue appears earlier in the message, including the earlier examplesThe analysis is complete and no code changes are needed hereandSummary: this PR looks good overall and I have no further comments. Both still evaluate topossible_truncationwith the current helper. Please make completion/non-continuation cases win over this structural fallback, or tie the fallback to a stronger truncation signal, and add behavioral coverage for completed punctuation-less replies where the completion marker is not at the very end. -
[P2] Update the continuation tests after moving the patterns
src/__tests__/bugfixes.test.ts:92
The PR now moves the continuation pattern logic out ofquery.tsintosrc/utils/continuation.ts, but the existing tests still readquery.tsand assert old implementation details likecontinuationSignals, the literalso now (i|let me|we)regex, andcompletionMarkers. As a result, the PR'ssmoke-and-testsjob is red duringbun test --max-concurrency=1, and the same failures reproduce locally withbun test src/__tests__/bugfixes.test.ts. Please update these assertions to the new module and, ideally, exerciseanalyzeContinuationIntentbehavior directly so future refactors do not break or miss the actual continuation semantics.
d6ba986 to
abf6f80
Compare
|
I have applied the following refinements to ensure technical perfection and fix the CI failures:
Verified locally with bun test src/tests/bugfixes.test.ts (31 pass, 0 fail). |
jatmn
left a comment
There was a problem hiding this comment.
Thanks for the updates here. The earlier completion-marker and stale-test failures look addressed, and the focused tests are green now. I still see one remaining false-positive path in the structural fallback that should be tightened before merge.
Findings
- [P2] Keep punctuation-less final summaries from triggering truncation nudges
src/utils/continuation.ts:49
TheisPossiblyTruncatedbranch still treats any message longer than 40 characters without terminal punctuation, a markdown-list prefix, or one of the narrow completion-marker words as an interrupted response. That still catches ordinary completed assistant summaries that simply omit a final period, for exampleI changed package.json and src/query.ts and added tests,Updated src/query.ts and added coverage in bugfixes.test.ts, orThis should be ready after the latest test updates; all of those currently return{ shouldNudge: true, reason: 'possible_truncation' }, so the main loop will send up to three extra meta "continue" turns after a normal final answer. Please make the truncation fallback depend on a stronger unfinished-output signal, or let a broader class of normal terminal replies win, and add behavioral coverage for punctuation-less final summaries that do not contain the current completion marker keywords.
abf6f80 to
f113df3
Compare
|
Hello @jatmn, I have implemented an "Unfinished Sentiment" heuristic to eliminate the remaining false positives while maintaining robust detection of actual token-limit interruptions. Key Improvements:
Verified all scenarios locally. The smoke-and-tests suite is now green and the logic is decoupled from raw source code assertions for future-proof stability. Testing Status:
|
jatmn
left a comment
There was a problem hiding this comment.
Thanks for the continued iteration here. I did another pass over the current head, including the earlier false-positive cases and a wider set of mixed-intent/truncation probes. The prior punctuation-less summary false positives look addressed, but there are still two continuation regressions that should be fixed before merge.
Findings
-
[P2] Let late continuation intent survive earlier completion markers
src/utils/continuation.ts:57
The new global completion-marker check returnsshouldNudge: falsebefore the late continuation signal logic runs, so mixed messages that finish one subtask and then state the next action no longer trigger a continuation nudge. For example,Task 1 is done. Let me update the status.,Task 1 finished. I will now run tests.,Analysis complete. Now I will edit src/query.ts, andNo issues in the first file. I will now inspect the next one.all currently return{ shouldNudge: false }. That is the same mixed-intent class this PR is trying to support, just with common completion words before the later action. Please make completion markers suppress only terminal/fallback cases, or otherwise let later high-confidence first-person continuation intent override earlier completion markers, and add behavioral coverage for these mixed-intent examples. -
[P2] Do not let completion markers hide structural truncation signals
src/utils/continuation.ts:57
The same early return also prevents the high-confidence structural checks from running when a response contains a completion word before an actual cut-off. Inputs likeSetup is complete. Here is the code:\n```typescript\nfunction run() {,Task complete. Please inspect (src/query.ts, andThe analysis is done and now I am editing files andall return{ shouldNudge: false }, even though the later unclosed code block, unbalanced parenthesis, or trailing connector are exactly the truncation signals added in this PR. Please evaluate structural truncation before treating earlier completion markers as terminal, and cover cases where a completed sub-step is followed by interrupted output.
f113df3 to
eac3537
Compare
|
Hello, and thanks again! I have refactored the analyzeContinuationIntent utility to follow a strict hierarchy of heuristics, ensuring that high-confidence truncation signals are never shadowed by completion markers.
|
jatmn
left a comment
There was a problem hiding this comment.
Thanks for the continued updates here. The prior continuation false-positive and false-negative cases look addressed on the current head.
No issues here, LGTM.
|
Thanks! Glad to hear that. Ready for merge when you are. |
…wigpine#1280) * feat(cli): improve SSH interactivity detection via SSH_TTY and SSH_CONNECTION * feat(models): add support for Gemma 4 31B * feat(query): robust multi-lingual and structural continuation nudge * fix(query): refine continuation nudge logic to avoid false positives
…wigpine#1280) * feat(cli): improve SSH interactivity detection via SSH_TTY and SSH_CONNECTION * feat(models): add support for Gemma 4 31B * feat(query): robust multi-lingual and structural continuation nudge * fix(query): refine continuation nudge logic to avoid false positives
- fix(autocompact): retry circuit breaker after cooldown (Twigpine#1375) - fix(provider): require API key input when adding OpenGateway (Twigpine#1384) - fix(provider): allow remote Ollama without OPENAI_API_KEY (Twigpine#952) - fix(codex-stream): recover tool args delivered only via done events (Twigpine#1262) - fix: route MiniMax compacting through Anthropic-compatible API (Twigpine#1154) - fix(thinking): disable thinking for unsupported Ollama models (Twigpine#1376) - feat(agents): set active session agent from agents menu (Twigpine#1349) - fix(repl): show permission prompts while draft input is present (Twigpine#1393) - fix(model): include profile models in descriptor picker (Twigpine#1361) - Improve warning notice formatting (Twigpine#1415) - fix(codex): allow credential storage fallback (Twigpine#1347) - fix(attribution): make git attribution opt-in by default (Twigpine#1335) - fix(agent): allow custom model overrides (Twigpine#1337) - feat(query): robust multi-lingual and structural continuation nudge (Twigpine#1280) - fix(watchers): debounce skills and settings reload bursts (Twigpine#1370) - feat: configure API retry backoff (Twigpine#370) (Twigpine#1095) - chore(main): release 0.15.0 (Twigpine#1325) - ci: retrigger CodeQL after action download outage (Twigpine#1374) - Fix launcher heap setup for long sessions (Twigpine#1242)
Summary
What changed:
Why it changed:
Impact
Testing
Notes