Conversation
…Research#94146) After a rate-limit episode, Weixin live replies were silently lost: the adapter reported success while nothing reached the client. Multiple false-success/silent-drop defects in the send path: - Rate-limited send results carried no retryable/retry_after/error_kind, so the gateway retry layer misclassified them as formatting failures and re-sent a mangled '(plain text:)' duplicate instead of honoring cooldown. - iLink sendmessage responses in the media path (caption + item sends) were never inspected: ret=-2/-14 came back and send_document/send_video/ send_voice still returned success=True with no error logged. - send_typing responses were never inspected either; throttled typing signals (a suspected trigger of server-side suppression) were swallowed at debug level - completely invisible. - Non-dict responses and empty/whitespace replies fell through as SendResult(success=True, message_id=None). - Successful sends logged nothing, so an accepted-but-dropped send was indistinguishable from a healthy one. Fixes: - Classify send failures via classify_send_error(); rate_limited results carry retryable=True and retry_after (cooldown remaining) so _send_with_retry retries after the cooldown instead of mangling. - Inspect + surface all sendmessage responses (text, caption, media) and typing responses; log every sendmessage response sanitized (ret, errcode, errmsg truncated, client_id prefix, context_token used). - Invalidate the cached context_token when the rate-limit circuit opens so the post-cooldown retry falls back to the tokenless degraded path. - Return an explicit failure for no-deliverable-content sends instead of success=True with message_id=None (live and direct paths). Regression tests (8, fake transport) were red pre-fix and green post-fix. Closes NousResearch#94146
This addresses #94146's real defects precisely: classified/retryable SendResults instead of bare failures, inspected responses on every path that previously faked success (media captions, non-dict payloads, blank replies), sanitized per-send logging, and token invalidation on circuit-open — all pinned by solid tests. Concerns:
Minor: |
Summary
This PR hardens the Weixin adapter's outbound path against client-side false success: five defects in
gateway/platforms/weixin.pythat let refused, throttled, or empty sends be reported as delivered (or swallowed silently) are fixed and covered by regression tests.It does not claim to fix the reported outage itself. The #94146 report describes replies that were dropped server-side with no error code at all (no
send failed,rate limited,session expired,-14, or-2log lines). That suppression is not reproduced or root-caused here — see "What this PR does not fix".What this PR fixes (client-side defects, all confirmed by code inspection; red→green tests)
send_document/send_video/send_voice(and the caption send) returnedSendResult(success=True, ...)even forret=-2/-14responses — a literal false success with no log line (the issue's "no-14or-2log")._send_typing()never inspected its response, and throttled typing bursts (a suspected suppression trigger in [Bug]: Weixin live replies are silently dropped after rate-limit incidents, even after fresh QR login and on latest main #94146) left no trace — now they surface at WARNING instead of a debug-level swallow.SendResult(success=True, message_id=None);send_weixin_directskipped the send and still returned success.sendmessageresponse was treated as delivered; now an empty response is a WARNING + classified failure and non-dict responses raise.send()returnedSendResult(success=False, error=...)with noretryable/retry_after/error_kind, soBasePlatformAdapter._send_with_retry()matched it against connection-error patterns, treated a throttle as a formatting failure, and re-sent a mangled"(Response formatting failed, plain text:)... "duplicate instead of honoring the cooldown. Nowclassify_send_error()marksrate_limitedwithretryable=True+retry_after(remaining cooldown, 5 s fallback).Plus: every
sendmessageresponse (text, caption, media) is now logged sanitized at INFO/WARNING (ret,errcode,errmsgtruncated,client_idprefix,context_tokenused), so an accepted-but-dropped send is at least distinguishable in production logs.What this PR does not fix (the reported case)
The #94146 repro shows a reply that produced no refusal code at all: the gateway logged the send as completed and no
-2/-14/ rate-limit / session-expired line appeared, while the message never arrived. The client-side defects above do not cover that: aret=0/errcode=0(or absent) response is still treated as accepted and the send still reports success. If iLink suppresses delivery without an error code, this PR cannot detect it. The server-side suppression is not reproduced (needs-repro) — the new per-send log lines are the instrumentation to capture it, and an evidence request has been left on #94146. The protocol-level gaps from #24989 also remain open.Hypothesized recovery — unverified, not a fix
The circuit-open path in this PR (invalidate the peer's cached
context_tokenso the post-cooldown retry starts from a tokenless send) is a hypothesis that a stale/poisoned context token suppresses server-side delivery. The report directly weakens it: the reporter states that a fresh QR login that produced a new bot account/token did not recover outbound delivery (#94146). A fresh login cannot carry a stale cached token, so a token-invalidation recovery cannot explain that case. This path is therefore labeled unverified: it may mitigate future incidents where a genuinely stale token matters, but it is not claimed to fix #94146.Rate-limit retry: double-send uncertainty
ret=-2semantics are ambiguous. In-repo,RATE_LIMIT_ERRCODEis documented as "iLink frequency limit — backoff and retry", while_is_stale_session_ret()treatsret=-2witherrmsg == "unknown error"as a stale-session signal (same as-14). If-2instead means "message accepted but delivery throttled", a post-cooldown retry double-sends. The handling here is conservative but cannot rule that out:WEIXIN_SEND_CHUNK_RETRIES(default 4) times with 3x backoff, then stops;_record_rate_limit_event()) aborts the in-flight send and surfaces a classified failure instead of hammering the endpoint;_send_with_retry()re-invokes the whole send at most 2 further times, honoringretry_afterwhen present.The double-send possibility is documented rather than asserted away; distinguishing "refused-and-retryable" from "accepted-but-throttled" needs a wire-level answer from iLink (see the evidence request on #94146).
Tests
8 new regression tests in
tests/gateway/test_weixin.py(TestWeixinRateLimitRecovery) using the fake-transport/stub-session pattern: circuit-open send returnsrate_limitedkind +retryable+retry_after; per-chunk loop exhaustion keeps rate_limited semantics; circuit open invalidates the cached context token; successful send logs a sanitized response line (accepted-but-dropped distinguishable); non-dict response is not a silent success; blank reply is a visible failure with no_send_messagecall;send_documentsurfaces a-2media response (wassuccess=True); typing rate limit logs at WARNING. All 8 were red pre-fix and green post-fix.Test results
tests/gateway/test_weixin.py+test_weixin_secret_scope.py+test_weixin_typing.py+test_poller_fd_lifecycle.py: 57 passed; 2 failures are pre-existing onmainin this environment (test_qr_login_timeout_uses_monotonic_clock,test_recycle_closes_old_and_installs_fresh_session— both fail identically withweixin.aiohttpunpatched/None in the test process, unrelated to this change).test_send_error_classification.py,test_send_retry.py,test_retry_response.py,test_completion_delivery.py: pass.tests/tools/test_send_message_tool.py: same result asmainin this environment (12 pre-existing Discord-forum failures, 94 passed).Status
Related: #94146 (stays open pending the evidence; the server-side root cause is not resolved by this PR).