Conversation
… annotation
replaceTextContent() prepended a {type:"text"} block to messages that had no
text block of their own, producing ["text","tool_result"] for a user message
whose only content was a tool_result. The Anthropic Messages API requires
tool_result blocks to come first in the message following a tool_use, so
upstream rejected it with 400 "tool_use ids were found without tool_result
blocks immediately after".
When the message carries a tool_result block, append the annotation after the
existing content instead of prepending it; messages without a tool_result keep
the prior prepend behavior.
Closes diegosouzapw#12890
(#12925) Rebased onto the tip and completed, per the maintainer's call to finish the wiring rather than merge the capability alone. What changed since your version: The tip had already cleared the TS2554 by deleting the 16th argument, leaving a comment that the highWaterMark stays at the helper default. So the base-red you found is gone, but the 64 KB #12179 asked for was still not applied and your new parameter had no caller. glm.ts now passes it, which is what turns the capability into the fix. Your test file also hung the runner: every stream createSSEStream builds arms a 10s idle watchdog via setInterval in start, and nothing cancelled them, so node:test waited on a non-empty event loop long after the assertions passed. Cancelling each readable in an after hook runs the cancel handler that clears the timer — the file now reports in about 7 seconds. Worth knowing for future stream tests. Your five assertions are unchanged and all pass. Reading the writable's desiredSize to measure the queue budget the stream was actually built with, rather than standing in for it, is the detail that makes this testable at all — and the 0-budget case pinning `??` against `||` is the kind of thing that silently rots otherwise. Thank you also for separating your own red checks from the base's and reporting what you found there. That is how #12919's identical failures got explained instead of chased.
|
Thanks for the clear root-cause write-up — the diagnosis and fix here are correct, but the Triage note: this is the review recommendation — the close itself happens only after the maintainer's per-PR sign-off (and, where a superseding PR is named, after it has landed). Nothing is being closed by this comment. |
|
Obrigado — o diagnóstico estava certo, mas o fix já entrou por outro caminho. Verifiquei contra o tip atual: A #12920 mergeou em 11/09 com o comportamento idêntico (append quando é Fechando como subsumida. O crédito pelo diagnóstico do ordenamento é seu. |
…gosouzapw#12179 (diegosouzapw#12925) Rebased onto the tip and completed, per the maintainer's call to finish the wiring rather than merge the capability alone. What changed since your version: The tip had already cleared the TS2554 by deleting the 16th argument, leaving a comment that the highWaterMark stays at the helper default. So the base-red you found is gone, but the 64 KB diegosouzapw#12179 asked for was still not applied and your new parameter had no caller. glm.ts now passes it, which is what turns the capability into the fix. Your test file also hung the runner: every stream createSSEStream builds arms a 10s idle watchdog via setInterval in start, and nothing cancelled them, so node:test waited on a non-empty event loop long after the assertions passed. Cancelling each readable in an after hook runs the cancel handler that clears the timer — the file now reports in about 7 seconds. Worth knowing for future stream tests. Your five assertions are unchanged and all pass. Reading the writable's desiredSize to measure the queue budget the stream was actually built with, rather than standing in for it, is the detail that makes this testable at all — and the 0-budget case pinning `??` against `||` is the kind of thing that silently rots otherwise. Thank you also for separating your own red checks from the base's and reporting what you found there. That is how diegosouzapw#12919's identical failures got explained instead of chased.
Closes #12890
Root cause
When the compression pipeline ages (or aggressively rewrites) a user message whose content is a lone
tool_resultblock and has no text block of its own,replaceTextContent()fell into itsif (!replaced)branch and prepended a fresh{type:"text"}block:For a message like
[{type:"tool_result", ...}]this produces["text", "tool_result"]. The Anthropic Messages API requirestool_resultblocks to come first in the message that immediately follows atool_use, so upstream rejects the request:Both callers hit this branch —
progressiveAging.ts(fullSummarytier tagging for user messages) andaggressive.ts(replaceContent) — so once a conversation crosses the aging threshold it fails on every subsequent turn (the reporter saw ~110 occurrences in a day).!replacedprepend when…progressiveAging.ts→tagAged("fullSummary", …)tool_result(no text block)aggressive.ts→replaceContent()Fix
Only the empty-of-text branch changes. When the message carries a
tool_resultblock, the annotation is appended after the existing content; messages without atool_resultkeep the previous prepend behavior (the annotation still leads, unchanged):Scoping the change to messages that actually contain a
tool_resultkeeps the existing ordering for every other shape (plain text, image-only, multi-text), so no other compression path shifts.Tests
New regression file
tests/unit/compression/tool-result-order-12890.test.ts(4 cases), next to the siblingtests/unit/compression/aggressive-fidelity.test.tswhich already covers Anthropictool_resultcompression:tool_result→ text block appended,tool_resultstays at index 0tool_resultblocks → both survive, no text block precedes themtool_result(image-only) → prepend behavior unchangedtool_resultuser message across thefullSummarythreshold keepstool_resultbefore any text blockBase-revert proof — with the source change stashed out and only the test present on the base commit:
The 3 ordering cases fail on base (the "unchanged behavior" case passes). With the fix:
Commands run
Not verified
I did not reproduce the original upstream 400 against a live Anthropic endpoint — the fix and its proof are at the⚠️ base-red inherited: #12732 — so compare CI failure sets against the base rather than reading a red check as this PR's fault.
replaceTextContent()unit boundary plus theapplyAging()integration path, matching the block ordering the issue traced. I did not run the full suite / coverage gate ortest:vitest; I ran the changed test plus its immediate siblings. The base branch is currently red independent of this change —