Skip to content

fix(routing): earliest-reset-first 4-fix bundle (post-deploy hotfix) - #23

Merged
i1hwan merged 2 commits into
mainfrom
fix/earliest-reset-first-v6
Apr 26, 2026
Merged

i1hwan merged 2 commits into
mainfrom
fix/earliest-reset-first-v6

Conversation

@i1hwan

@i1hwan i1hwan commented Apr 26, 2026

Copy link
Copy Markdown
Owner

Summary

Regressions fixed

# Symptom (production) Fix
F1 openai-compatible (self-hosted llama.cpp) → "all candidates excluded" 503 trackScores=0 branches: isAccountQuotaExhausted → excluded; otherwise score=0 fallback (paid still wins, self-hosted usable)
F2 opus/haiku routed to different accounts at the same instant (per-model weekly window split) Drop modelWindowMapping.ts entirely. Routing reads only overall weekly window. 429s on per-model quota retry via accountFallback
F3 100% fresh Anthropic accounts (Anthropic omits resetAt until first request) auto-excluded Q === 100 && resetAt === null → max-urgency T=100, score=10000; other null resetAt cases stay missing
F4 APEX 17%/2h outranked GNUMAX 87%/4h (additive formula privileges urgency over abundance) Multiplicative S = T_pts × Q_remain, score range [0, 10000]. Penalties rescaled 100×. Q_SATURATION_CAP deleted

Hand verification (user's actual scenario)

Account Track Q Reset T_pts Score = T·Q
APEX session 17 2h 23m 40 680
APEX weekly 35 3d 10h 20 700
APEX avg 690
GNUMAX session 87 4h 43m 20 1740
GNUMAX weekly 59 4d 18h 20 1180
GNUMAX avg 1460

→ GNUMAX wins by 770 points. Asserted in test F4-1.

Behaviour matrix (post-fix)

Scenario Result
openai-compatible (cache entry absent) score=0, eligible; lex tie-break across self-hosted only
Anthropic 429-marked empty cache excluded: "quota_exhausted_unknown_reset"
paid (score>0) + self-hosted (score=0) mixed paid wins
paid backoffLevel=4 (finalScore=-8400) + self-hosted self-hosted wins (intentional)
Q=100, resetAt=null (fresh) score=10000 (max urgency)
Q=80, resetAt=null (ambiguous) kind: "missing"
Per-model weekly Omelette=0% (would-have-excluded in v4) not excluded; routing uses overall weekly only
Q < 5% either track hard excluded
Ranking tie on T × Q product earliest-reset asc → lex id asc

API change

  • selectByEarliestResetFirst(candidates, sessionId) — modelHint parameter removed.
  • scoreAccount(conn), scoreWeeklyTrack(connId), isAffinityValid(conn, sessionId) — modelHint removed.
  • auth.ts:706-714 updated; only one caller site, no external API surface change.

Files

  • src/sse/services/strategies/earliestResetFirst.ts — rewritten (4-fix bundle)
  • src/sse/services/strategies/modelWindowMapping.ts — deleted
  • src/sse/services/auth.ts — caller site modelHint removed + comment updated
  • tests/unit/auth-strategy-earliest-reset-first.test.mjs — fixture simplified, 18 new tests added
  • package.json 3.7.0 → 3.7.1 (patch bump for hotfix)
  • docs/openapi.yaml info.version: 3.7.1
  • CHANGELOG.md — new [3.7.1] section
  • .sisyphus/plans/routing-strategy-v5.md — marked superseded
  • .sisyphus/plans/routing-strategy-v6.md — canonical SoT (separate workspace dir)

Out of scope

  • Active Sessions fingerprint algorithm (intentional per-agent fingerprinting confirmed).
  • cacheTelemetry follow-up fields (still v4 §6 follow-up).
  • per-provider strategy override.
  • auth.ts.orig cleanup.

Tag plan

Merge → tag apex-v2.0.1 (patch bump on top of apex-v2.0.0).

PR #21 (apex-v2.0.0) 머지 후 production 에서 다음 4가지 회귀 관찰됨:
  (1) openai-compatible llama.cpp → "all candidates excluded" 503
  (2) opus/haiku 가 같은 시점 다른 계정으로 분기 (modelWindowMapping)
  (3) 100% fresh Anthropic 계정이 missing 처리되어 자동 제외
  (4) APEX 17%/2h 가 GNUMAX 87%/4h 보다 우선 (burn-down 의도 위반)

Fix:
  - F1: self-hosted/openai-compatible (no cache) → score=0 fallback
        429-marked empty cache 는 excluded 유지 (부활 X)
  - F2: drop modelWindowMapping (Omelette/Sonnet hard guard 제거)
  - F3: resetAt=null AND Q=100% → fresh max-urgency (nothing else)
  - F4: additive S=0.85T+0.15Q → multiplicative S=T×Q
        penalty 100x rescale, Q_SATURATION_CAP 삭제

Verification:
  - 48/48 strategy unit tests PASS (기존 + F1×4, F2×2, F3×6, F4×6 신규)
  - 2791/2791 full unit suite PASS
  - prettier / eslint / typecheck:core / typecheck:noimplicit:core / docs-sync clean
  - APEX 17%/2h vs GNUMAX 87%/4h scenario test asserts GNUMAX selection
  - 429-marked cache excluded test (Oracle B1 patch verified)

Plan: .sisyphus/plans/routing-strategy-v6.md (Oracle 4-blocker + Momus 9-issue resolved)
Independent of PR #22 (mojibake hotfix).
Copilot AI review requested due to automatic review settings April 26, 2026 17:54

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Hotfix update to the earliest-reset-first routing strategy after production regressions in v3.7.0, shifting the strategy to a v6 model that removes per-model weekly window logic and changes scoring to a multiplicative urgency×remaining formulation.

Changes:

  • Rewrite earliest-reset-first scoring to T_pts × Q_remain, add fresh-quota handling (Q===100 && resetAt===null), and add a score=0 fallback for no-cache self-hosted providers.
  • Remove model-specific weekly-window mapping and drop the modelHint parameter from strategy APIs and the auth.ts call site.
  • Add/adjust unit tests for the four production regressions and bump version/docs/changelog to 3.7.1.

Reviewed changes

Copilot reviewed 7 out of 7 changed files in this pull request and generated 1 comment.

Show a summary per file
File Description
src/sse/services/strategies/earliestResetFirst.ts v6 rewrite: multiplicative scoring, fresh quota handling, self-hosted fallback, API signature changes.
src/sse/services/strategies/modelWindowMapping.ts Deleted; per-model weekly window mapping removed.
src/sse/services/auth.ts Updates caller to new selectByEarliestResetFirst(candidates, sessionId) signature and updates strategy comment.
tests/unit/auth-strategy-earliest-reset-first.test.mjs Updates fixtures and adds regression coverage for F1–F4 behaviors.
package.json Patch version bump to 3.7.1.
docs/openapi.yaml Updates OpenAPI info.version to 3.7.1.
CHANGELOG.md Adds 3.7.1 release notes describing the hotfix bundle.

Comment on lines 394 to 397
const s = scoreSessionTrack(conn.id);
if (s.kind === "excluded") return { valid: false, reason: s.reason };
const w = scoreWeeklyTrack(conn.id, modelHint);
const w = scoreWeeklyTrack(conn.id);
if (w.kind === "excluded") return { valid: false, reason: w.reason };

Copilot AI Apr 26, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

isAffinityValid can incorrectly keep affinity pinned to a connection that was marked quota-exhausted via markAccountExhaustedFrom429(). In that case both tracks are missing (empty quotas) so this function returns valid: true, while scoreAccount() would exclude it via isAccountQuotaExhausted(). Consider checking isAccountQuotaExhausted(conn.id) here (and returning invalid) so affinity breaks and selection can fall back to another account; add a regression test for this scenario.

Copilot uses AI. Check for mistakes.
Copilot review on PR #23 caught an inconsistency between scoreAccount
and isAffinityValid:

- scoreAccount's F1 branch correctly distinguishes self-hosted (no cache
  entry → score=0 fallback) from 429-marked exhausted (cache exists with
  empty quotas + exhausted=true → excluded).
- isAffinityValid had no equivalent guard. A connection marked exhausted
  via markAccountExhaustedFrom429() has empty quotas, so scoreSessionTrack
  and scoreWeeklyTrack both return kind:"missing" (not "excluded"). With
  no excluded gate, isAffinityValid returned valid:true and pinned the
  next request to the same 429-burning account.

Add the same isAccountQuotaExhausted guard in isAffinityValid (after
the static rate-limit/terminal checks, before quota-track scoring) so
affinity breaks and selection falls back to another candidate via the
normal scoring path.

Regression test added: 429-marked + bound session must produce
valid:false with reason "quota_exhausted_unknown_reset".

Strategy tests: 49/49 PASS. Full unit: 2792/2792 PASS.
@i1hwan
i1hwan merged commit 4f526b0 into main Apr 26, 2026
47 checks passed
@i1hwan
i1hwan deleted the fix/earliest-reset-first-v6 branch April 26, 2026 18:15
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants