fix(gateway): classify iLink stale sessions correctly - #96437
fangliquanflq wants to merge 5 commits into
Conversation
iLink reports a stale session as ret=-2 errmsg="prepare failed" on the sendmessage endpoint. The adapter previously classified this as a rate limit, opening the rate-limit circuit and locking out all outbound Weixin delivery (including cron-initiated pushes) until the cooldown expires. Classify "prepare failed" as an outbound stale-context signal, delete the cached context_token only when it still matches the failed token, load the token after acquiring the outbound gate, and allow one tokenless recovery send outside the transient retry budget. If the tokenless recovery also reports a stale session, surface a clear error instead of tripping the rate-limit circuit; genuine rate limits (e.g. "freq limit") still open the breaker. Follow-up to NousResearch#17228; complements NousResearch#74572.
|
Thanks for this — the direction matches what we've seen in production, and the actionable-error branch (tokenless retry still stale → surface a refresh instruction, no circuit) is the right shape. Three evidence-based notes from our incident forensics, one race consideration, and a convergence thought. 1. The dominant stale variant in our capture isn't in the keyword set. Across three failing windows captured end-to-end (gateway up 8 days, zero inbound for 3+ days — evidence in #80125), the stale responses were: {"ret": -2, "errcode": null, "errmsg": "prepare failed"}
2. The tokenless recovery retry consumes the retry budget here. The recovery 3. The token cleanup this PR inherits is a blind delete. The stale path still clears the store via 4. Wide substring matching cuts both ways. Convergence: #80426 is the same-domain candidate carrying the production evidence (three independent sites now: ours, Akie-Star's, EricCai's — #96416 makes a fourth). Your |
|
Thanks for the detailed production evidence. I addressed each point on this branch:
For convergence and attribution, I merged the implementation commit from #80426 into this branch and resolved it while retaining this PR's Verification:
|
|
Verified at
Three non-blocking notes for the record:
One correction to my own review: I wrote that "@ericcaiwx-star's Aug 24 reply confirms the race" — that Aug 24 comment in #80125 was mine; EricCai's production report was Aug 9. Misattribution on my part, apologies. From my side this closes the review. cc @teknium1 for visibility only: with #80426's commit folded in here verbatim, this PR settles the cluster in one diff, and #80426 can close as superseded when it lands — authorship preserved in the chain. |
|
Thanks for the thorough verification and the correction. I rechecked the current
Current required CI is green, including Python tests, E2E, Windows/macOS checks, and Python lint. A fresh local invocation of the focused command in this checkout could not start because no pytest-enabled virtualenv is installed, so I am not claiming an additional local pass from this run. |
|
Rebase needed to land. As of 2026-09-13 this PR is still Fresh evidence for this cluster is now on #96416: a 10-day, 22/22 delivery-failure streak on v0.21.2, plus a direct-endpoint capture returning |
|
Thanks for flagging the stale branch. I merged current upstream The new production capture is consistent with the existing exact Verification:
|
|
|
Summary
-14and session-semantic-2responses as stale sessionsTesting
scripts/run_tests.sh tests/gateway/test_weixin.py tests/gateway/test_weixin_typing.py tests/gateway/test_weixin_secret_scope.py -q(58 passed)Related Issue
Closes #96416