Skip to content

fix(proxy): rewrite Claude Code compatible requests - #5

Merged
i1hwan merged 1 commit into
mainfrom
fix/native-claude-forwarding
Apr 11, 2026
Merged

i1hwan merged 1 commit into
mainfrom
fix/native-claude-forwarding

Conversation

@i1hwan

@i1hwan i1hwan commented Apr 11, 2026

Copy link
Copy Markdown
Owner

Summary

  • apply the lexical trigger rewrite inside the Claude Code-compatible request builder so the final Anthropic-bound request no longer bypasses the OAuth workaround
  • preserve and rewrite native Claude tool choice values in the same builder path, and carry the tool-name map for response-side restoration
  • add regression coverage proving system text, tool descriptions, tool names, tool choice, and assistant tool_use blocks are rewritten in Claude Code-compatible requests

Verification

  • node --import tsx/esm --test tests/unit/claude-code-compatible-request.test.mjs tests/unit/translator-openai-to-claude.test.mjs
  • npx eslint open-sse/services/claudeCodeCompatible.ts tests/unit/claude-code-compatible-request.test.mjs open-sse/translator/helpers/claudeHelper.ts open-sse/handlers/chatCore.ts
  • npm run typecheck:core
  • manual QA confirmed final built Claude Code-compatible request contains background_stop, background_result, directories:\nsrc/, and rewritten tool_choice

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
Copilot AI review requested due to automatic review settings April 11, 2026 12:21
@i1hwan
i1hwan merged commit 2b377e6 into main Apr 11, 2026
47 of 48 checks passed

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Ensures Claude Code-compatible request building applies the Claude OAuth lexical rewrite (text + tool names/tool_choice) so the final Anthropic Messages payload doesn’t bypass the OAuth workaround, and adds regression tests around the rewrite behavior.

Changes:

  • Apply applyClaudeOAuthLexicalRewrite() to the final Claude Code-compatible request body and propagate the resulting tool-name restoration map (_toolNameMap).
  • Ensure Claude-native tool_choice is carried through the Claude Code-compatible body preparation path.
  • Extend unit coverage to verify rewrite behavior across system text, message text, tool descriptions/names, tool_choice, and assistant tool_use blocks.

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated no comments.

File Description
open-sse/services/claudeCodeCompatible.ts Applies lexical rewrite to the constructed CC-compatible Messages payload; threads tool_choice through the Claude-native preparation path; attaches _toolNameMap for response restoration.
tests/unit/claude-code-compatible-request.test.mjs Adds regression assertions that the CC-compatible builder rewrites relevant text/tool fields and preserves the tool-name map.

i1hwan added a commit that referenced this pull request May 7, 2026
Security: /api/sessions enrichment now uses listProviderConnectionMetadata()
which selects only id/provider/name/display_name/email. Decryption of apiKey,
accessToken, refreshToken, idToken via getProviderConnections() is no longer
on this hot path. Fixes Copilot review #11.

Routing memory: heuristicBreakHistory map now bounded by BOTH age (10min) AND
hard size cap (1000). Re-set on insert moves the entry to the end of the
insertion-ordered Map; size-cap eviction loop removes oldest entries when
size exceeds 1000. >1000 distinct sessions inside the cooldown window can no
longer grow the map without bound. Fixes Copilot review #1/#6.

Routing perf: isAffinityValid heuristic-break #2 now finds the bound entry
in scoredAlternatives and reuses its score instead of calling scoreAccount()
again. Matches the 'score upfront once' contract of selectByEarliestResetFirst.
Fixes Copilot review #2/#7.

UI tooltip: RoutingBadge now flips below the badge when there isn't enough
space above (TOOLTIP_HEIGHT_ESTIMATE_PX + gap + viewport padding). top anchors
to r.bottom + gap on flip; transform switches from translateY(-100%) to
translateY(0). Fixes Copilot review #3/#8.

UI width: TOOLTIP_WIDTH_ESTIMATE_PX raised 240 -> 244 to match min-w-[220px]
plus px-3 padding. Eliminates 4px overflow on narrow viewports.
Fixes Copilot review #9.

Comments: SSR mounted comment rewritten to describe the actual typeof window
inline check and the eslint react-hooks/set-state-in-effect rule that blocks
the useState+useEffect mounted pattern. Fixes Copilot review #10.

Tests: SA-5/6 description corrected (re-test in same tick is strictly inside
the 60s cooldown window; '30s later' wording removed). Fixes Copilot review
#4. New tests/unit/api-sessions-route.test.mjs covers /api/sessions success
enrichment with explicit secret-leak guard, orphan connectionId fallback,
and simulated DB failure fallback. Fixes Copilot review #5.

Version 3.8.2 -> 3.8.3. test:unit 2833/2833 PASS.
i1hwan added a commit that referenced this pull request May 7, 2026
* feat(dashboard): limits page polish + smart session affinity (4 issues)

PR #25 (v3.1 + v7) merge follow-up. User feedback identified 4 polish
items observed on /dashboard/limits after Docker deploy:

1. AutoRefreshControl raw <input type="checkbox"> + raw <select> rendered
   like unstyled HTML controls — replaced checkbox with Toggle from the
   shared design system, restyled the select wrapper to match the rest of
   the dashboard (rounded bg-surface border, focus ring, custom chevron).

2. RoutingBadge hover tooltip was clipped by the Account Model Quotas
   card's overflow-hidden ancestor (PR #25 introduced the clipping). Now
   uses createPortal(..., document.body) with position:fixed and viewport
   edge clamping, matching the codebase's existing portal pattern in
   providers/[id]/page.tsx.

3. Smart session affinity continuation viability (v8). 5min
   SESSION_AFFINITY_WINDOW_MS unchanged (matches Anthropic's prompt cache
   TTL — librarian-confirmed via docs.anthropic.com prompt-caching). Two
   new heuristic break rules layered on top of existing hard exclusions:
     - affinity_break_low_quota: bound's min known remaining < 15% AND a
       usable alt exists (Oracle: 5% hard-floor × 3 cushion).
     - affinity_break_p1_too_urgent: alt score >= bound × 3 AND >= bound
       + 250 absolute delta. The absolute delta neutralizes near-zero
       misfire (e.g. bound=20, alt=61 trips 3× alone but 41-point gap
       doesn't justify a cache write).
   60s per-session cooldown gates oscillation; hard exclusions bypass
   cooldown. selectByEarliestResetFirst now scores upfront and passes the
   list to isAffinityValid (avoids double scoring).

4. SessionsTab account name. /api/sessions now joins active sessions
   against provider connections via getAccountDisplayName (graceful
   degradation: connection lookup failure returns sessions with
   accountName:null without breaking the API). UI replaces raw
   connectionId.slice(0,10) with the resolved account name + provider
   tag, falling back to "Account #xxxxxx" when name is unavailable. New
   i18n key 'usage.connectionFallback' added across 32 locales (en + ko
   hand-translated, 30 placeholder).

Tests: 10 new SA-1..SA-10 RED tests cover smart affinity (low-Q break,
heavy-gap break, no-alt edge, missing-track edge, cooldown, hard-
exclusion bypass, score<=0 special case, backwards-compat third arg).
Test cooldown reset hook (__resetAffinityHeuristicCooldownForTesting)
prevents cross-test pollution. Full unit suite: 2830/2830 PASS (was
2821; +9 net).

Verify: prettier ✓ eslint (errors 0, warnings unchanged from baseline) ✓
typecheck:core ✓ typecheck:noimplicit:core ✓ test:unit 2830/2830 ✓
docs:sync ✓.

Plan: .sisyphus/plans/limits-dashboard-polish.md
Oracle review: APPROVED with 7 revisions, all applied (bg_79458ad3)
Momus reviews: v1 REJECT (QA executability) → v2 OKAY (bg_999bd379, bg_ebbedf5b)

* fix(dashboard): restore reset countdown on session/weekly mini-bars + scoreAccount terminal guard

QuotaVisualization MiniBar now renders the reset countdown next to the label
(`⏱ 0h 34m` / `6d 11h`), mirroring the per-model bar format. PR #25
introduced the dual mini-bars but accidentally dropped the countdown for the
overall Session/Weekly windows, leaving the row visually flat.

scoreAccount() now excludes connections with terminal testStatus (expired /
banned / credits_exhausted) at scoring time, not only inside isAffinityValid.
Without this guard the fall-through path of selectByEarliestResetFirst could
re-select a terminal connection if its cached quotas still looked healthy.
Mirrors the auth.ts contract via the shared isTerminalConnectionStatus helper.
Fixes a SA-7 CI flake on Linux runners (selected.id === 'bound' instead of
'alt') without changing local behavior.

CHANGELOG documents the OAuth-lane prompt-cache reality: clients can request
ttl: '1h' but the server downgrades to 5m on the OAuth path (observed in
production responses; tracked upstream as anthropics/claude-code#46829), so
SESSION_AFFINITY_WINDOW_MS = 5min is the maximum we can rely on, not a default.

* fix(security,routing,ui): address Copilot PR #26 review (8 issues)

Security: /api/sessions enrichment now uses listProviderConnectionMetadata()
which selects only id/provider/name/display_name/email. Decryption of apiKey,
accessToken, refreshToken, idToken via getProviderConnections() is no longer
on this hot path. Fixes Copilot review #11.

Routing memory: heuristicBreakHistory map now bounded by BOTH age (10min) AND
hard size cap (1000). Re-set on insert moves the entry to the end of the
insertion-ordered Map; size-cap eviction loop removes oldest entries when
size exceeds 1000. >1000 distinct sessions inside the cooldown window can no
longer grow the map without bound. Fixes Copilot review #1/#6.

Routing perf: isAffinityValid heuristic-break #2 now finds the bound entry
in scoredAlternatives and reuses its score instead of calling scoreAccount()
again. Matches the 'score upfront once' contract of selectByEarliestResetFirst.
Fixes Copilot review #2/#7.

UI tooltip: RoutingBadge now flips below the badge when there isn't enough
space above (TOOLTIP_HEIGHT_ESTIMATE_PX + gap + viewport padding). top anchors
to r.bottom + gap on flip; transform switches from translateY(-100%) to
translateY(0). Fixes Copilot review #3/#8.

UI width: TOOLTIP_WIDTH_ESTIMATE_PX raised 240 -> 244 to match min-w-[220px]
plus px-3 padding. Eliminates 4px overflow on narrow viewports.
Fixes Copilot review #9.

Comments: SSR mounted comment rewritten to describe the actual typeof window
inline check and the eslint react-hooks/set-state-in-effect rule that blocks
the useState+useEffect mounted pattern. Fixes Copilot review #10.

Tests: SA-5/6 description corrected (re-test in same tick is strictly inside
the 60s cooldown window; '30s later' wording removed). Fixes Copilot review
#4. New tests/unit/api-sessions-route.test.mjs covers /api/sessions success
enrichment with explicit secret-leak guard, orphan connectionId fallback,
and simulated DB failure fallback. Fixes Copilot review #5.

Version 3.8.2 -> 3.8.3. test:unit 2833/2833 PASS.

* fix(routing,ui,api): address Copilot PR #26 round 3 review (3 issues, Oracle-verified)

/api/sessions now fetches metadata only for active session connectionIds:
listProviderConnectionMetadata(ids?) accepts an optional id filter, the route
collects distinct connectionIds from getActiveSessions(), short-circuits to
no DB call on empty, and otherwise issues a single WHERE id IN (?, ?, ...)
with bound parameters. Endpoint work is now proportional to active sessions,
not total connections. Fixes Copilot review #NEW-3.

RoutingBadge useLayoutEffect cleanup no longer calls setCoords(null). The
tooltip is already gated by 'open && coords' so the cleanup state update
was unnecessary and risked extra renders / strict-mode noise. Oracle
verified no stale-frame race (useLayoutEffect runs synchronously before
paint, updateCoords sets fresh coords on reopen). Fixes Copilot review
#NEW-1.

formatCountdown extracted to ProviderLimits/utils.tsx and reused from both
index.tsx (per-model bars) and QuotaVisualization.tsx (Session/Weekly mini
bars). Behavior preserved: <24h => h h m m, >=24h => d d h h, invalid =>
null. Eliminates the drift risk between the two countdown call sites.
Fixes Copilot review #NEW-2.

Tests: +2 in tests/unit/api-sessions-route.test.mjs covering 'no active
sessions skips DB query' and 'distinct connectionIds collapse to single
bound parameter, unrelated secrets never leak'. 2835/2835 unit tests pass.

Oracle pre-commit verification (ses_1fc62b147ffeUtgRRb2uvBKS4x): APPROVED
all 3 fixes with high confidence.

Version 3.8.3 -> 3.8.4.

* fix(routing,ui): address Oracle audit + Copilot R4 review (2 defects + 2 nits)

QuotaVisualization.pickWindow no longer absorbs per-model quotas (D1).
The previous Pass 2 fallback matched any name starting with 'weekly ' or
'session ', so a connection that only had 'weekly Sonnet (7d)' (no
canonical 'weekly' row) populated the overall mini-bar AND rendered as
its own per-model bar simultaneously. Pass 2 removed; only canonical
'session'/'weekly' (with or without parenthesised window) match. New
regression suite tests/unit/quota-visualization-pickwindow.test.mjs.

RoutingBadge tooltip is now visible on viewports shorter than the
tooltip estimate (D2). Previous fix flipped vertically when space-above
ran out but never clamped into the viewport, so very short viewports or
tooltips taller than viewH-2*pad still rendered partially off-screen.
The new placement (1) prefers the side with more room (mirrors
providers/[id]/page.tsx overlay rule), (2) clamps on-screen top with
transform-aware math (above-anchored uses translateY(-100%), so we
require top >= tooltipHeight + padding), (3) caps the rendered element
with maxHeight: calc(100vh - 16px) + overflow: hidden so tooltips
larger than the viewport degrade to a clipped frame, never off-screen.
Fixes Copilot R4-1.

isAffinityValid now guards against scoredAlternatives missing the bound
entry (N1). Previously fell through to boundScore=0, letting any
positive alt trip the urgent-break rule. Now returns { valid: true }
on missing-bound. Production selectByEarliestResetFirst always includes
bound, so this is a future-caller footgun guard. New SA-11 test.

/api/sessions metadata-lookup catch now logs a sanitized warning instead
of swallowing the error silently (N2). Format: '[sessions] connection
metadata lookup failed for N ids: <message>'. Includes only the count
and the error message text — no connectionIds, no SQL, no secrets.

Oracle pre-commit verification (ses_1fc4cce1bffePsARTIZ6x3AxlY)
returned NEEDS_REVISION on D1 and D2; both fixes applied per the
audit's concrete revision instructions, plus the two NITs.

Test count: 2841/2841 PASS (was 2835; +6 net). Version 3.8.4 -> 3.8.5.
i1hwan added a commit that referenced this pull request May 11, 2026
…stics, LowQuota bypass (#30)

* feat(settings): add Claude Compatibility tab skeleton (step 1/7)

Step 1 of the Claude Compatibility tab work — UI-only skeleton, no
behavior change. Default routing/tool-arg/streaming paths byte-identical
to pre-PR.

What this adds
- New 'Compatibility' tab in /dashboard/settings between Resilience and
  Advanced, registered in tabs array with material-symbol icon 'tune'.
- CompatibilityTab.tsx mounting 4 subsection cards in a vertical column:
  - ToolArgumentModeSection — default mode radio + per-provider/per-lane
    override tables. Local React state only; persistence wired in step 2.
  - LowQuotaBypassSection — default policy radio + per-provider/per-lane
    override tables. Local React state only.
  - SSEDiagnosticsSection — 3 capture checkboxes + 3 numeric retention
    inputs + Clear button (disabled until step 4).
  - TerminalRecoverySection — UI-only placeholder per plan §13. Marked
    'Coming soon'; no settings key, no API route.
- Shared OverrideTable<V> generic component used by both ToolArgs and
  LowQuotaBypass sections: editable rows with key dropdown sourced from
  USAGE_SUPPORTED_PROVIDERS (or claude-oauth-prefixed lane), value
  dropdown, remove button, '+ Add override' control. Duplicate keys
  prevented by remainingKeys filter.
- 39 new settings.* i18n keys in en.json + ko.json (hand-translated).
  Other 30 locales fall back to English per the PR #29 pattern.
- compatibility-tab-skeleton.test.mjs — verifies tab registration in
  page.tsx, 4 subsection mounts, and i18n key presence in both locales.

Plan references
- .sisyphus/plans/tool-arg-buffered-final-mode.md (rev3.2.2) §14 step 1
- reviewer final approval (rev3.2.1) — no scope expansion vs. plan

Verification
- prettier --write: all 10 touched paths unchanged after format
- npm run lint on touched paths: 0 errors
- npm run typecheck:core: clean
- node --import tsx/esm --test on the new + 5 related regression suites:
  49/49 pass

Files
  M src/app/(dashboard)/dashboard/settings/page.tsx
  A src/app/(dashboard)/dashboard/settings/components/CompatibilityTab.tsx
  A .../CompatibilityTab/ToolArgumentModeSection.tsx
  A .../CompatibilityTab/LowQuotaBypassSection.tsx
  A .../CompatibilityTab/SSEDiagnosticsSection.tsx
  A .../CompatibilityTab/TerminalRecoverySection.tsx
  A .../CompatibilityTab/OverrideTable.tsx
  M src/i18n/messages/en.json (+39 keys)
  M src/i18n/messages/ko.json (+39 keys)
  A tests/unit/compatibility-tab-skeleton.test.mjs (6 cases)

* feat(settings): wire Claude Compatibility settings layer (step 2/7)

Step 2 of the Claude Compatibility tab work — backend settings layer +
WebUI API connection. Routing/streaming behavior still byte-identical
to pre-PR; only operator preferences are now persisted.

What this adds
- Migration 021_compatibility_settings.sql with 3 settings keys
  (INSERT OR IGNORE — safe on re-run / existing installs):
    toolArgumentMode   = {default:'stream-normalized', byProvider:{}, byLane:{}}
    lowQuotaBypass     = {default:false, byProvider:{}, byLane:{}}
    sseDiagnostics     = {captureProviderRawSSELines:false, ...,
                          keepLastNDebugRequests:20, maxDebugBundleSizeMB:100,
                          maxActiveDebugBundles:5}
  Per plan rev3.2 fix #5, no terminalRecovery key — UI placeholder only.
- Zod schemas in src/shared/validation/settingsSchemas.ts with canonical
  provider-id whitelist (USAGE_SUPPORTED_PROVIDERS) per plan rev3.1 fix #4:
    toolArgumentModeSettingsSchema / lowQuotaBypassSettingsSchema /
    sseDiagnosticsSettingsSchema, plus matching DEFAULT exports.
- src/lib/db/settings.ts getSettings() now seeds the 3 new keys with
  imported defaults so older databases (pre-021 migration) fall back
  consistently.
- 4 API routes:
    GET/PUT /api/settings/tool-argument-mode
    GET/PUT /api/settings/low-quota-bypass
    GET/PUT /api/settings/sse-diagnostics
    POST    /api/settings/sse-diagnostics/clear
  All PUT handlers validate via validateBody(zod) and persist via
  updateSettings. The clear handler removes files from
  getSseDiagnosticsDir() — ENOENT is treated as success (idempotent).
- src/lib/logEnv.ts exports resolveDefaultDataDir + getSseDiagnosticsDir
  (shared between the clear route and the future step 4 capture write
  path so DATA_DIR resolution stays in one place).
- Subsection components now hydrate via GET on mount and persist via
  PUT on every change (optimistic update with rollback on failure,
  matching RoutingTab.tsx pattern). Local-state placeholders from
  step 1 are gone.
- compat-settings-api.test.mjs — static-analysis tests verifying
  migration shape, schema exports, route exports, validateBody usage,
  clear-route data-dir resolution, and section endpoint wiring.

Plan references
- .sisyphus/plans/tool-arg-buffered-final-mode.md (rev3.2.2) §14 step 2
- reviewer final approval (rev3.2.1) — no scope expansion vs. plan
- Verification commands deferred to a single end-of-PR pass per
  user policy (no per-step npm test / typecheck runs).

Files
  A src/lib/db/migrations/021_compatibility_settings.sql
  M src/shared/validation/settingsSchemas.ts (+3 schemas, +3 defaults)
  M src/lib/db/settings.ts (+3 default keys)
  M src/lib/logEnv.ts (export 2 path helpers)
  A src/app/api/settings/tool-argument-mode/route.ts
  A src/app/api/settings/low-quota-bypass/route.ts
  A src/app/api/settings/sse-diagnostics/route.ts
  A src/app/api/settings/sse-diagnostics/clear/route.ts
  M src/app/(dashboard)/.../CompatibilityTab/ToolArgumentModeSection.tsx
  M src/app/(dashboard)/.../CompatibilityTab/LowQuotaBypassSection.tsx
  M src/app/(dashboard)/.../CompatibilityTab/SSEDiagnosticsSection.tsx
  A tests/unit/compat-settings-api.test.mjs

* feat(translator): add buffered-final tool argument mode (step 3/7)

Step 3 of the Claude Compatibility tab work — implements the
buffered-final tool argument streaming mode at the Claude→OpenAI
translator layer. Default behavior (stream-normalized) byte-identical
to pre-PR; new mode kicks in only when the operator opts in via
WebUI (step 2 already wired).

What this adds
- open-sse/translator/helpers/toolArgumentMode.ts:
  - resolveToolArgumentMode(settings, provider, forwardingLane) →
    'stream-normalized' | 'buffered-final'
  - Precedence: byLane > byProvider > default > 'stream-normalized'
  - Resolver is total (never throws); malformed/missing → safe default
  - Whitelisted mode values via VALID_MODES set — unknown strings reject
- open-sse/translator/response/claude-to-openai.ts (3 case branches):
  - content_block_start (tool_use): buffered-final defers the placeholder
    chunk emit and skips normalizer init; stream-normalized unchanged.
  - content_block_delta (input_json_delta): buffered-final concatenates
    raw partial_json onto state.toolCalls[idx].function.arguments
    without emitting; stream-normalized unchanged (still routes through
    jsonUnicodeNormalizer).
  - content_block_stop (tool_use slot): buffered-final emits exactly
    one chunk with id + type + name + the entire accumulated arguments;
    stream-normalized continues to flush the normalizer tail.
- open-sse/handlers/chatCore.ts:
  - applyClaudeOAuthLexicalRewrite call sites (L1027 + L1072) now also
    stash _forwardingLane = 'claude-oauth-prefixed' alongside _toolNameMap
    (reviewer rev3.1 fix #1 — explicit lane marker, never inferred from
    toolNameMap presence).
  - Extraction block at L1234 also unstashes _forwardingLane and forwards
    it to the stream factory.
  - On the translate path, resolveToolArgumentMode is called once per
    request via dynamic import of @/lib/localDb (same cross-package
    pattern used by volumeDetector + rateLimitManager) — settings read
    is cached so cost is minimal.
- open-sse/utils/stream.ts:
  - StreamOptions and TranslateState extended with toolArgumentMode +
    forwardingLane fields.
  - State object construction propagates both into translator scope.
  - createSSETransformStreamWithLogger wrapper accepts two new positional
    args (defaults preserve old behavior; non-translate callers untouched).
- tests/unit/translator-resp-claude-to-openai-buffered-final.test.mjs:
  - 10 cases covering §7.2.1-7.2.8 from plan rev3.2:
    - single chunk emit at stop with full args
    - stream-normalized regression guard (mode absent / default)
    - multiple tool_use blocks independent
    - mid-escape stop preserves verbatim partial buffer
    - mixed text+tool flow
    - no internal-state key leak
    - T_FINAL < T_FINISH in both modes (reviewer rev3.1 fix #3)
    - client-style accumulator parity buffered vs normalized
      (reviewer rev3.1 fix #4 carry-over)
    - resolver precedence (lane > provider > default + malformed → safe)

Plan references
- .sisyphus/plans/tool-arg-buffered-final-mode.md (rev3.2.2) §5.2 + §5B
- reviewer final approval (rev3.2.1)
- Verification commands deferred to a single end-of-PR pass per
  user policy (no per-step npm test / typecheck runs).

Files
  A open-sse/translator/helpers/toolArgumentMode.ts
  M open-sse/translator/response/claude-to-openai.ts (mode branch in 3 cases)
  M open-sse/utils/stream.ts (options/state + wrapper)
  M open-sse/handlers/chatCore.ts (_forwardingLane stash + extract + resolve)
  A tests/unit/translator-resp-claude-to-openai-buffered-final.test.mjs (10 cases)

* feat(stream): add SSE diagnostics capture (step 4/7)

Step 4 of the Claude Compatibility tab work — captures provider SSE
streams into per-request debug bundles so the next incident yields
the wire-level evidence we currently lack. All capture toggles default
OFF; zero overhead when disabled.

What this adds
- open-sse/utils/sseDiagnosticsBundle.ts:
  - tryCreateBundle: returns null when no capture toggle is on OR when
    maxActiveDebugBundles concurrent cap is reached.
  - appendRawLine / appendParsedEvent / appendTranslatedChunk: gated by
    per-capture booleans; overflow guard short-circuits further appends
    when bundle._bytes would exceed maxDebugBundleSizeMB.
  - finalizeBundle: writes <DATA_DIR>/logs/sse-diagnostics/<TS>_<UUID>.json,
    then prunes to keepLastNDebugRequests by mtime. Termination marker
    is one of 'flush' | 'upstream_error' | 'client_abort' with optional
    detail string.
  - Module-level activeBundleCount + _testOnlyResetActiveCount helper
    for the test harness.
- open-sse/utils/stream.ts:
  - StreamOptions extended with sseDiagnosticsConfig.
  - createSSEStream wires capture hooks at 3 points:
    - raw line: in the line-split loop, before parseSSELine
    - parsed event: after parseSSELine succeeds, in translate path
    - translated chunk: after formatSSE + sanitize, before
      controller.enqueue
  - finalizeBundle called in 3 termination paths:
    - flush()'s finally block (normal completion or streamTimedOut)
    - upstream-error controller.error() in the translate fast-fail block
    - idle-timeout controller.error()
  - createSSETransformStreamWithLogger wrapper accepts 13th positional
    arg (sseDiagnosticsConfig).
- open-sse/handlers/chatCore.ts:
  - The existing single getSettings() read on the translate path now
    extracts sseDiagnostics alongside toolArgumentMode and forwards the
    config to createSSETransformStreamWithLogger.
- tests/unit/sse-diagnostics-capture.test.mjs:
  - 11 cases covering plan §5.3 + §7.3:
    - toggles OFF → no bundle file written
    - each capture type independently stores its data
    - parsed events stored without truncation
    - keepLastNDebugRequests prunes oldest by mtime
    - maxDebugBundleSizeMB overflow guard sets _capture_overflow
    - maxActiveDebugBundles concurrent cap (reviewer rev3.1 fix #7)
    - upstream_error termination marker + detail preserved
    - client_abort termination marker
    - null-bundle safety (all append* are no-ops)
    - clear API route deletes all bundle files

Plan references
- .sisyphus/plans/tool-arg-buffered-final-mode.md (rev3.2.2) §5.3
- reviewer final approval (rev3.2.1) — diagnostic boundary + abort path
  + concurrent cap all reflected.
- Verification commands deferred to a single end-of-PR pass per
  user policy (no per-step npm test / typecheck runs).

Files
  A open-sse/utils/sseDiagnosticsBundle.ts (~190 lines)
  M open-sse/utils/stream.ts (3 capture hooks + 3 termination paths + wrapper)
  M open-sse/handlers/chatCore.ts (forward sseDiagnostics config)
  A tests/unit/sse-diagnostics-capture.test.mjs (11 cases)

* feat(routing): add LowQuota bypass + Q<=0 hard-exclude (step 5/7)

Step 5 of the Claude Compatibility tab work — implements operator-toggleable
LowQuota routing bypass at the earliestResetFirst strategy. Bypass scope is
provider-only (rev3.3 page directive); the Q<=0 case is hard-excluded
regardless of bypass (page rev3.2 blocker #2). Default behavior
byte-identical to pre-PR.

Two-tier exclusion guard
- Tier 1 (rev3.2): Q<=0 → excluded with reason 'session<=0%' / 'weekly<=0%'.
  Hard-excluded regardless of bypass — Q=0 means actually exhausted, NOT
  'low quota'. Aligns with quotaCache.ts:65 isExhausted semantics.
  UI maps the new reasons to existing 'routingPriorityExcludedExhausted'
  label.
- Tier 2 (existing): 0<Q<5 → excluded with reason 'session<5%' / 'weekly<5%'
  unless lowQuotaBypass=true. When bypass is on, score falls through to
  pressure() with bypassMinUsable:true so 1%, 3%, 4.99% all produce
  distinct finite scores.

What this adds (page directive in rev3.3)
- open-sse/services/routing/lowQuotaBypass.ts:
  - resolveLowQuotaBypass(settings, provider) → boolean.
  - NO byLane parameter. auth.ts:auth() runs BEFORE chatCore +
    applyClaudeOAuthLexicalRewrite, so forwardingLane is not known at
    connection select time. byLane on LowQuota would persist to storage
    with no runtime effect — exactly the UX trap rev2.1 warned about.
    'byProvider:{claude:true}' covers every Claude connection (OAuth +
    API-key) which matches the operator goal '5% 미만이어도 계속 써라'.
  - toolArgumentMode.byLane is unaffected — its consumer (stream factory)
    IS lane-aware, so its byLane stays.

- src/sse/services/strategies/earliestResetFirst.ts:
  - pressure() gains optional PressureOptions { bypassMinUsable } so the
    Q<5 guard can be skipped without changing the default contract.
  - scoreSessionTrack(connId, lowQuotaBypass=false) — Tier 1 Q<=0 guard
    + Tier 2 Q<5 guard gated by bypass.
  - scoreWeeklyTrack(connId, lowQuotaBypass=false) — same.
  - scoreAccount(conn, lowQuotaBypass=false) — forwards to both tracks.
  - isAffinityValid(conn, sessionId, scored, provider, lowQuotaBypass=false)
    — same bypass decision the per-attempt scoring uses.
  - selectByEarliestResetFirst(candidates, sessionId, provider,
    lowQuotaBypass=false) — single boolean (NOT a Map<connId, bool>) since
    auth.ts always passes provider-filtered orderedConnections.

- src/sse/services/auth.ts (auth() at L709):
  - Resolves bypass once via resolveLowQuotaBypass(settings.lowQuotaBypass,
    provider) using the existing getSettings() read at L570 and forwards
    to selectByEarliestResetFirst.

- src/app/api/usage/provider-limits/route.ts (dashboard preview):
  - buildResponseBody resolves bypass per connection (Map<connId, bool>)
    because computeRouting mixes providers across the response map.
    computeRouting forwards each connection's resolved bypass into
    scoreAccount.

- src/shared/validation/settingsSchemas.ts:
  - lowQuotaBypassSettingsSchema is now z.strict() with { default, byProvider }
    only. PUT with a byLane field returns 400.
  - LOW_QUOTA_BYPASS_DEFAULT drops byLane.

- src/lib/db/migrations/021_compatibility_settings.sql:
  - lowQuotaBypass default value is '{"default":false,"byProvider":{}}'
    (no byLane field).

- src/app/(dashboard)/dashboard/settings/components/CompatibilityTab/
  LowQuotaBypassSection.tsx:
  - Per-lane override table REMOVED.
  - New hint text under per-provider table explains why lane scope is not
    applicable and instructs operator to use provider='claude'=true to
    cover the OAuth lane.

- src/i18n/messages/en.json + ko.json:
  - New key compatibilityLowQuotaProviderScopeOnlyHint with the explanation
    above (hand-translated en/ko).

Tests added
- tests/unit/low-quota-bypass-resolver.test.mjs (9 cases):
  - empty / null / undefined → false
  - default + byProvider precedence
  - per-provider override beats default both directions
  - provider=null → only default applies
  - malformed settings → safe fallback
  - rev3.3 — resolver signature is 2-arg (no byLane param)

- tests/unit/score-track-low-quota-bypass.test.mjs (10 cases):
  - pressure() default vs bypassMinUsable opts
  - pressure() Q=0 with bypass returns 0 (finite, lowest)
  - pressure() preserves ordering (1% < 3% < 4.99%)
  - scoreSessionTrack Q=3 bypass=false → excluded session<5%
  - scoreSessionTrack Q=3 bypass=true → known finite
  - scoreSessionTrack Q=0 → excluded session<=0% REGARDLESS of bypass
  - scoreWeeklyTrack Q=0 → excluded weekly<=0% REGARDLESS of bypass
  - scoreSessionTrack signature accepts optional bypass param

- tests/unit/compat-settings-api.test.mjs:
  - New case: lowQuotaBypassSettingsSchema PUT with byLane field rejects 400.

- tests/unit/compatibility-tab-skeleton.test.mjs:
  - New i18n key compatibilityLowQuotaProviderScopeOnlyHint added to
    REQUIRED_KEYS so en/ko coverage is enforced.

Plan references
- .sisyphus/plans/tool-arg-buffered-final-mode.md (rev3.3) §5B
- page rev3.3 mid-step-5 directive — option B (remove byLane on LowQuota)
- Verification commands deferred to a single end-of-PR pass per user policy.

Files
  A open-sse/services/routing/lowQuotaBypass.ts (~17 lines)
  M src/sse/services/strategies/earliestResetFirst.ts (Q<=0 + bypass plumbing)
  M src/sse/services/auth.ts (resolveLowQuotaBypass once per request)
  M src/app/api/usage/provider-limits/route.ts (dashboard map builder)
  M src/shared/validation/settingsSchemas.ts (strict schema, no byLane)
  M src/lib/db/migrations/021_compatibility_settings.sql (default shape)
  M src/app/(dashboard)/.../LowQuotaBypassSection.tsx (no lane override)
  M src/i18n/messages/en.json (+1 hint key)
  M src/i18n/messages/ko.json (+1 hint key)
  M tests/unit/compatibility-tab-skeleton.test.mjs
  M tests/unit/compat-settings-api.test.mjs
  A tests/unit/low-quota-bypass-resolver.test.mjs
  A tests/unit/score-track-low-quota-bypass.test.mjs

* feat(ui): RoutingBadge composite + Compatibility deep-link (step 6/7)

Step 6 of the Claude Compatibility tab work — UI polish. Adds the
'Rank N · LowQuota' composite badge for near-depletion accounts whose
LowQuota bypass is active, the matching near-depletion subtitle in the
RoutingBadge tooltip, and a cross-link from ProviderLimits to the new
Compatibility settings tab.

What this adds
- src/app/(dashboard)/.../ProviderLimits/RoutingBadge.tsx:
  - getExcludedI18nKey: added 'session<=0%' and 'weekly<=0%' cases
    (mapped to existing 'routingPriorityExcludedExhausted' label per
    plan rev3.2 — Q<=0 is functionally Exhausted from the UI side).
  - New nearDepletion derived: entry.excluded === false AND
    min(sessionRemainingPct, weeklyRemainingPct) > 0 AND < 5. This
    matches the bypass-active runtime state (LowQuota bypass kept the
    candidate in the routing pool; UI now surfaces the warning).
  - Composite label: 'Rank N · LowQuota' via new i18n key
    routingPriorityRankLowQuota. Falls back to plain 'Rank N' /
    'Next' when nearDepletion is false.
  - variant flips from 'default' to 'warning' (amber styling) when
    nearDepletion is true.
  - Tooltip: appends a near-depletion subtitle with both session/weekly
    percentages, separated from the breakdown rows by a faint divider.

- src/app/(dashboard)/.../ProviderLimits/index.tsx:
  - Header toolbar (right-side action group) gains a Compatibility
    deep-link button that goes to /dashboard/settings?tab=compatibility.
    Sits between 'Refresh All' and AutoRefreshControl, matching the
    visual weight of the existing buttons.

- src/i18n/messages/en.json + ko.json (usage namespace):
  - compatibilitySettingsLink           — button label
  - compatibilitySettingsLinkTooltip    — title= tooltip
  - routingPriorityRankLowQuota         — '{n} · LowQuota'
  - routingPriorityNearDepletionTooltip — '⚠ Near depletion: session={s}%, weekly={w}% (bypass active)'

Tests added
- tests/unit/routing-badge-composite.test.mjs (8 cases):
  - session<=0% / weekly<=0% map to existing Exhausted label
  - session<5% / weekly<5% mapping unchanged (regression guard)
  - nearDepletion branch + i18n keys present in source
  - condition uses min(s,w) < 5 with positive floor (rev3.3 §5B.4)
  - variant flips to 'warning' on nearDepletion
  - ProviderLimits/index has the Compatibility deep-link

- tests/unit/compatibility-tab-skeleton.test.mjs:
  - USAGE_NS_KEYS array + 2 new assertions confirming en/ko have the
    4 new usage namespace keys (deep-link + composite badge i18n).

Plan references
- .sisyphus/plans/tool-arg-buffered-final-mode.md (rev3.3) §4A.5 + §5B.4
- reviewer final approval (rev3.2.1) — composite badge policy locked in
- Verification commands deferred to a single end-of-PR pass per user policy.

Files
  M src/app/(dashboard)/.../ProviderLimits/RoutingBadge.tsx
  M src/app/(dashboard)/.../ProviderLimits/index.tsx (Compatibility link)
  M src/i18n/messages/en.json (+4 usage keys)
  M src/i18n/messages/ko.json (+4 usage keys)
  M tests/unit/compatibility-tab-skeleton.test.mjs
  A tests/unit/routing-badge-composite.test.mjs

* fix(tests,schemas,diagnostics): verification-pass cleanup (step 7/7)

Step 7 of the Claude Compatibility tab work — final verification pass
plus reviewer-driven hardening pulled into the same commit.

Verification fixes (3008/3008 → 3009/3009 pass)

1) settingsSchemas.ts — providerOverrideRecord / laneOverrideRecord
   were using z.record(z.enum([...]), ...) which under Zod v4 requires
   *every* enum value to appear in the record (an empty {} is rejected
   as 'expected boolean, received undefined' for each missing key).
   Switched to z.record(z.string(), valueSchema).refine(...) with an
   explicit whitelist check + a Zod v4 compatible refine() options
   shape ({ message: '...' }). Existing tests for valid-keys + strict()
   unknown field rejection continue to pass.

2) translator-resp-claude-to-openai-buffered-final.test.mjs — the
   §7.2.8 client-style accumulator parity test was reading the first
   emitted tool_calls chunk for the stream-normalized mode by reference;
   subsequent input_json_delta events mutated the same object via
   state.toolCalls.get(idx).function.arguments += delta. By the time
   the test ran, every captured chunk reflected the final-state
   arguments, not the emit-time snapshot the OpenAI client would
   actually see on the wire. Added a feedAndSnapshot() wrapper that
   JSON.stringify+parse-clones each feed result immediately so the
   parity comparison sees honest emit-time captures.

Reviewer-driven hardening (post-self-verify)

3) sseDiagnosticsBundle.finalizeBundle now sets bundle._finalized=true
   on first call and short-circuits any subsequent call. Defends
   against the (currently unreachable but theoretically possible)
   double-finalize path where flush() runs after controller.error()
   already fired the bundle write. Added a regression test that
   finalizes the same bundle twice with different termination
   reasons and asserts (a) only one bundle file is written, (b)
   activeBundleCount returns to 0 exactly once, (c) the first
   termination marker is preserved.

4) sseDiagnosticsBundle.approxByteLen now delegates to
   Buffer.byteLength(value, 'utf8') instead of estimating per-char
   UTF-8 size manually. Same result for ASCII/BMP/3-byte CJK; lifts
   the maxDebugBundleSizeMB cap to the actual disk byte cost without
   any per-call overhead worth mentioning.

Verification (full end-of-PR sweep — after all 4 fixes)
- npm run test:unit                → 3009/3009 pass (+1 idempotency)
- npm run typecheck:core           → clean
- npm run lint                     → 0 errors, 75 baseline warnings
                                     (identical to pre-PR baseline)
- prettier --write on changed files → unchanged after format
- LSP diagnostics on changed files  → clean
- docs:sync                         → script not present in this repo;
                                      no version bump emitted here.
                                      Operator may bump 3.8.8 → 3.9.0
                                      at PR merge time if desired.

Files
  M src/shared/validation/settingsSchemas.ts
  M tests/unit/translator-resp-claude-to-openai-buffered-final.test.mjs
  M open-sse/utils/sseDiagnosticsBundle.ts
  M tests/unit/sse-diagnostics-capture.test.mjs

* fix(diagnostics): connect client_abort path + flush remaining buffer capture

Addresses two review change requests on PR #30 SSE diagnostics.

Issue 1 — client_abort was not actually wired to the stream cancel path

Before this commit, finalizeBundle(..., 'client_abort') only fired when
helper code called it directly (e.g. tests). The real client-disconnect
flow is:
  - createDisconnectAwareStream.cancel(reason)
  - → streamController.handleDisconnect()
  - → reader.cancel()
  - → writer.abort()
The diagnostics bundle lived in a closure inside createSSEStream() that
this cancel path could not reach. reader.cancel() also short-circuits
TransformStream.flush(), so even the existing flush-side finalize never
ran on a real client disconnect — losing every diagnostics bundle for
the exact incident class the feature was designed to capture.

Fix: a module-level WeakMap<TransformStream, CaptureBundle> registry.
createSSEStream registers its bundle against the transformStream object
it returns. createDisconnectAwareStream looks up the bundle on cancel
and calls finalizeBundle(bundle, 'client_abort', reason). The
bundle._finalized idempotency guard from step 7 ensures a subsequent
flush-side call (which won't actually run, but is defended against
anyway) is a no-op. The WeakMap entry is garbage-collected when the
TransformStream is — no manual cleanup.

Issue 2 — flush() remaining buffer was missing the 3 capture hooks

The transform() loop captures raw lines / parsed events / translated
chunks. The flush() handler reprocesses any remaining buffered text
(provider responses without trailing \n) through parseSSELine and
translateResponse, but it bypassed the 3 capture hooks. A provider
that emits its final usage-bearing event (message_delta, response.completed,
etc.) without a trailing newline would land that event in
provider_response summary but not in the diagnostics bundle — losing
exactly the last event whose evidence the operator most needs.

Fix: added the matching appendRawLine / appendParsedEvent /
appendTranslatedChunk calls inside the flush() translate-mode
remaining-buffer block, mirroring the transform() loop hooks.

Tests
- client_abort path integration test: registers a bundle on a real
  TransformStream, routes through createDisconnectAwareStream, cancels
  the readable, and asserts a bundle file lands with
  metadata.termination === 'client_abort' and metadata.terminationDetail
  carrying the cancel reason.
- flush() remaining buffer static check: greps the translate-mode
  remaining-buffer block in stream.ts and asserts all three append*
  hooks are present.

Verification
- npm run test:unit                → 3011 / 3011 pass (+2 cases)
- npm run typecheck:core           → clean
- npm run lint                     → 0 errors (75 baseline warnings)
- prettier --write on changed files → no diff after format
- LSP diagnostics on changed files  → clean

* fix(diagnostics): bridge bundle registration across pipeWithDisconnect wrapper

createSSEStream() registers the bundle on the original TransformStream, but
pipeWithDisconnect() pipes through that stream and hands a wrapper object to
createDisconnectAwareStream(), which calls lookupBundle(wrapper). The keys
do not match and client_abort finalize silently no-ops in production.

Fix: lookup the bundle on the original transformStream inside
pipeWithDisconnect() and re-register it against the wrapper before passing
it down.

Test: new regression test exercises the full pipeWithDisconnect() path
(not just createDisconnectAwareStream() in isolation) and uses bounded
polling for finalize landing on disk to avoid flake.

Per reviewer change request on PR #30.

* fix(ci+review): noImplicitAny + hot-path settings cache + diagnostics validation + sub-1% display

Resolves CI Lint failure and four Copilot review comments on PR #30.

1. src/shared/constants/providers.ts (CI blocker)
   - Add ProviderDefinition type
   - Annotate helper signatures: supportsApiKeyOnFreeProvider, isOpenAI/Anthropic/ClaudeCodeCompatibleProvider, getProviderByAlias, resolveProviderId, getProviderAlias
   - reduce<Record<string, string>> on ALIAS_TO_ID / ID_TO_ALIAS
   - Removes 9 implicit-any errors that block typecheck:noimplicit:core

2. open-sse/handlers/chatCore.ts (Copilot review #1 + #2)
   - Streaming translation hot path now reads settings via getCachedSettings()
     (5s TTL, invalidated on updateSettings) instead of a fresh SQLite read
     per request.
   - sseDiagnostics config is no longer a raw type-cast: merged with
     SSE_DIAGNOSTICS_DEFAULT and validated by sseDiagnosticsSettingsSchema
     before being passed to the stream factory. Falls back to defaults on
     malformed DB rows.

3. open-sse/utils/sseDiagnosticsBundle.ts (Copilot review #3)
   - tryCreateBundle() now defensively normalizes incoming numeric config
     (finite integer, in-range) before applying caps. NaN size cap,
     negative keepLastN, non-integer maxActive, and non-boolean truthy
     capture flags all produce null instead of silently disabling caps.

4. src/app/(dashboard)/dashboard/usage/components/ProviderLimits/RoutingBadge.tsx (Copilot review #4)
   - New formatPct helper renders 0 < value < 1 as '<1' instead of rounding
     to '0', preventing visual collision with the hard-excluded <=0% case
     in both the badge percent display and the near-depletion tooltip.

Tests:
- 2 new bundle-validation cases (malformed numeric + non-boolean coerce)
- 1 new formatPct test (sub-1%, NaN, integers)
- existing maxDebugBundleSizeMB-enforcement test rewritten to use 1MB cap
  + oversized line (the old 0MB cap is now correctly rejected by
  normalizeBundleConfig)

Verification:
- typecheck:noimplicit:core: clean (was failing with 9 errors)
- typecheck:core: clean
- npm test: 3015/3015 pass (PR baseline 3011 + 4 new)
- npm run lint: 0 errors, 75 baseline warnings (unchanged)

Out of scope: the two Unit Tests (22) failures (chat-context-relay /
sse-auth) on previous CI runs were ENOTEMPTY directory-cleanup races in
resetStorage(), unrelated to this PR. Local 'npm test' passes cleanly.

* fix(security+docs): forwardingLane trust boundary + pressure() doc

Addresses both Copilot review comments on the second pass of PR #30.

1. open-sse/handlers/chatCore.ts — forwardingLane trust boundary
   The previous code read forwardingLane directly from translatedBody._forwardingLane
   whenever it was a string. Because translatedBody is derived from the client
   request (often via { ...body }), a client could inject
   '_forwardingLane: "claude-oauth-prefixed"' to spoof lane-scoped routing
   (e.g. toolArgumentMode byLane overrides) without the Claude OAuth lexical
   rewrite ever running.

   Three defenses, defense-in-depth:
   a) stripInternalMarkers(body) at function entry — drops _forwardingLane,
      _toolNameMap, _disableToolPrefix, _nativeCodexPassthrough from the
      incoming client body before any translation branch runs. Every branch
      now sees a sanitized body, regardless of which path is taken.
   b) Lane is only honored when translatedToolNameMap instanceof Map &&
      size > 0. Maps survive structuredClone but cannot be JSON-injected by
      a client, so this proves the value came from
      applyClaudeOAuthLexicalRewrite, not the client.
   c) isValidForwardingLane() whitelist (new export on toolArgumentMode.ts)
      blocks unknown lane names — future lane additions must explicitly
      register, no silent acceptance.

2. src/sse/services/strategies/earliestResetFirst.ts — pressure() doc
   The doc comment unconditionally said 'Q < MIN_USABLE_REMAINING_PCT returns
   -Infinity'. The new PressureOptions.bypassMinUsable flag explicitly allows
   finite scores for Q<5 (LowQuota bypass). Comment now distinguishes the two
   cases so callers don't misinterpret the contract.

Tests (4 new):
- tests/unit/chat-core-forwarding-lane-trust-boundary.test.mjs (3 cases):
  * stripInternalMarkers helper deletes all 4 client-controllable fields
  * strip happens before the translation if-chain (every branch sees sanitized body)
  * forwardingLane requires both instanceof Map gate AND lane whitelist
- tests/unit/translator-resp-claude-to-openai-buffered-final.test.mjs:
  * isValidForwardingLane whitelist (8 cases: valid, casing, empty, null,
    undefined, number, object with toString)

Verification:
- typecheck:noimplicit:core: clean
- typecheck:core: clean
- npm test: 3019/3019 pass (PR baseline 3015 + 4 new)
- npm run lint: 0 errors, 75 baseline warnings (unchanged)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants