fix(weixin): treat iLink "prepare failed" as a stale session, not a rate limit - #100815
raymondyan-zhijie wants to merge 2 commits into
Conversation
Duplicate of #91647. Both add |
`_is_stale_session_ret` already routes iLink's ret=-2 "prepare failed" / "unknown error" to the tokenless retry, but that retry is guarded by `not retried_without_token and context_token`. Once it is spent — already attempted, or never available because the send carried no context_token, which is exactly the cron / proactive-push case the retry exists for — the response falls through to the rate-limit arm purely because ret == -2. That arm is the wrong medicine for a dead session: it burns a 3x backoff per chunk and feeds `_record_rate_limit_event`, which can open the adapter-wide cooldown circuit and then fast-fail unrelated sends for the whole window. Surface it as a stale-session error instead, and break rather than raise so the generic retry arm does not re-send against the same dead session. Refs NousResearch#100815
`_is_stale_session_ret` already routes iLink's ret=-2 "prepare failed" / "unknown error" to the tokenless retry, but that retry is guarded by `not retried_without_token and context_token`. Once it is spent — already attempted, or never available because the send carried no context_token, which is exactly the cron / proactive-push case the retry exists for — the response falls through to the rate-limit arm purely because ret == -2. That arm is the wrong medicine for a dead session: it burns a 3x backoff per chunk and feeds `_record_rate_limit_event`, which can open the adapter-wide cooldown circuit and then fast-fail unrelated sends for the whole window. Surface it as a stale-session error instead, and break rather than raise so the generic retry arm does not re-send against the same dead session. Refs NousResearch#100815
|
Follow-up pushed to this branch: the original fix routed iLink's Two ways to hit that in production:
In both cases the rate-limit arm is the wrong medicine for a dead session: it burns a 3x backoff per chunk and feeds The new commit surfaces it as a stale-session error instead, and uses Observed on a production gateway: five |
|
现场补充一点:只把 实测(Hermes 0.20.6 / Windows 11,现场证据见 #101039 (comment) ):cron 日报投递失败时 job 的 建议随这次分类一起补两点:
另附一个实测的返回形状坑(与本 PR 的判定逻辑相关):iLink 发送成功时返回 |
|
Field data point on how far this fix reaches — from the deployment that filed the corroborating evidence on #82502 (issuecomment-5594240732). We applied exactly this widening ( But in our field case the widened classifier alone did not restore delivery. With the peer dormant, the tokenless retry was refused identically: Implication for this PR: the premise "the failure misses the existing tokenless retry, so widening the classifier fixes cron pushes" holds only while the peer session is expired-but-refreshable. For a dormant peer the adapter has no send-side recovery — the fix converts a silent swallow into a correct classification, not into a delivery. Suggest covering the second half in the same change:
Status note (2026-09-11): all three PRs for this errmsg variant are still open — #100815, #96437, #91647 — none merged, and |
`_is_stale_session_ret` already routes iLink's ret=-2 "prepare failed" / "unknown error" to the tokenless retry, but that retry is guarded by `not retried_without_token and context_token`. Once it is spent — already attempted, or never available because the send carried no context_token, which is exactly the cron / proactive-push case the retry exists for — the response falls through to the rate-limit arm purely because ret == -2. That arm is the wrong medicine for a dead session: it burns a 3x backoff per chunk and feeds `_record_rate_limit_event`, which can open the adapter-wide cooldown circuit and then fast-fail unrelated sends for the whole window. Surface it as a stale-session error instead, and break rather than raise so the generic retry arm does not re-send against the same dead session. Refs NousResearch#100815
…ate limit iLink answers ret=-2 for both a genuine frequency limit and a stale session; only ``errmsg`` disambiguates them. ``_is_stale_session_ret`` recognised ``unknown error`` (NousResearch#17228) but not ``prepare failed``, which is what the API returns once the stored ``context_token`` for a peer has gone stale. That token is refreshed only by an *inbound* message, so the failure is specific to unprompted sends: an interactive reply always carries a fresh token, while a cron push after a quiet period does not. Because the response was filed as a rate limit, ``_send_text_chunk_locked`` never reached the tokenless retry directly above -- the retry whose docstring states it exists to "keep cron-initiated push messages working even when no user message has refreshed the session recently". Instead the send burned six 30s backoffs and failed, while the cron fire fence expired underneath it. Observed in production as 16 consecutive cron delivery failures to a WeChat target whose interactive replies were succeeding normally. Widen the helper to a frozenset of stale-session errmsgs so both spellings reach the tokenless retry. Genuine rate limits (``freq limit``) and the separate errcode -14 path are unaffected.
`_is_stale_session_ret` already routes iLink's ret=-2 "prepare failed" / "unknown error" to the tokenless retry, but that retry is guarded by `not retried_without_token and context_token`. Once it is spent — already attempted, or never available because the send carried no context_token, which is exactly the cron / proactive-push case the retry exists for — the response falls through to the rate-limit arm purely because ret == -2. That arm is the wrong medicine for a dead session: it burns a 3x backoff per chunk and feeds `_record_rate_limit_event`, which can open the adapter-wide cooldown circuit and then fast-fail unrelated sends for the whole window. Surface it as a stale-session error instead, and break rather than raise so the generic retry arm does not re-send against the same dead session. Refs NousResearch#100815
04a9ed1 to
1776009
Compare
|
Field data from the same deployment as the Observed section above — window extended to 2026-08-29 → 2026-09-15 — corroborating @LohasGuy's result, with a count. The classifier widening works. The tokenless retry has never once delivered. I paired every Every tokenless attempt came back A measurement trap worth recordingBefore After it, the same condition logs one line: So a naive count of Compounding it: the adapter logs failures but nothing on success. Two suggestionsBoth overlap with @Lsy5515's comment above; this field data supports them:
Taken together, that means the failure mode is silent (no success log, and the failure is only distinguishable by |
Re-triage: no longer marked as a duplicate of #91647. The second commit on this branch (spent-session arm in |
|
Keeping this open — it is the minimal cut of the fix and still applies cleanly to One relationship worth flagging for triage. The earlier re-triage compared this against #91647 only. #96437 targets the same delta more broadly — it adds For the record, |
|
|
Problem
iLink answers
ret=-2for two different conditions: a genuine frequency limit, and a stale session. Onlyerrmsgtells them apart, which_is_stale_session_retalready accounts for — but it only recognisesunknown error(#17228).The API also returns
prepare failedonce the storedcontext_tokenfor a peer has gone stale. That token is refreshed only by an inbound message, so this failure is specific to unprompted sends:ret=-2 errmsg=prepare failedBecause
prepare failedmisses theunknown errorcheck,is_session_expiredis false and the response falls through to theis_rate_limited = (ret == RATE_LIMIT_ERRCODE)branch._send_text_chunk_lockedtherefore never reaches the tokenless retry directly above it — the retry whose own docstring says it exists toInstead the send burns six 30s backoffs and fails.
Observed
On a production deployment, 16 consecutive cron deliveries to a WeChat target failed this way while that same target's interactive replies were succeeding normally. Timeline of one occurrence:
The cron fire fence expires while the adapter is still backing off, so the job is marked failed and the cause is misattributed to rate limiting. The tokenless retry does not rescue this send either (below) — what the patch changes is that the failure is reported as itself, ~90s sooner, instead of as a frequency limit.
Fix
Widen the helper to a frozenset of stale-session
errmsgvalues so both spellings reach the existing tokenless retry. Genuine rate limits (freq limit) and the separateerrcode -14path are untouched.The branch also carries a follow-up commit — a spent stale session is not a rate limit. Without it the widening alone still ends in a 3x backoff loop and feeds
_record_rate_limit_event, because once the retry is spent (retried_without_tokennow true, orcontext_tokenfalsy to begin with) the response falls through to the rate-limit arm purely onret == -2.What this does not do
It does not restore delivery to a dormant peer. Pairing every
session expired for <peer>; retrying without context_tokenline in this deployment's gateway logs against subsequent rate-limit-backoff or send-failure lines (180s window), 2026-08-29 → 2026-09-15:Every tokenless attempt came back
ret=-2 errmsg="prepare failed"— byte-for-byte the same response as the attempt that carried the real-but-stale token. Delivery resumed only after the peer sent an inbound message. On iLink the tokenless path is not a degraded fallback; it is a second way to fail.Corroboration from other deployments is in the thread: @LohasGuy reports the same result from the field ("the widened classifier alone did not restore delivery … delivery resumed only after an inbound message from the peer"), and @Lsy5515 documents the silent-failure shape on a different deployment. Our own quantified run, and the measurement trap below, are here.
Verification
Added four cases to
TestIsStaleSessionRet, and confirmed the existing negative cases still hold:(-2, None, "prepare failed")(None, -2, "prepare failed")(-2, None, "Prepare Failed")(-2, None, "unknown error")(-2, None, "freq limit")(-14, None, "session expired")After deploying the patch, the previously-unreachable branch fires in production:
Firing is not delivering: that line now appears on the way to the same failure, ~90s earlier than it used to.
Measuring this is a trap — the same behaviour logs differently on either side of the follow-up commit
Before a spent stale session is not a rate limit:
After it:
So a naive count of
send failed— or of anything containingrate limited— gives different numbers for identical behaviour depending on which side of that commit the window falls. Compounded by the adapter logging nothing on success (send_textreturnsSendResult(success=True, ...)with no info line), so success is only ever inferable from the absence of an error — the exact signal whose wording this changes. Worth knowing before anyone A/Bs this patch on log counts.