Skip to content

fix(test): make the proxy-fallback and subprocess cases actually boundable - #1490

Merged
murdore merged 1 commit into
releasefrom
fix/proxy-fallback-timeout-injection
Aug 23, 2026
Merged

murdore merged 1 commit into
releasefrom
fix/proxy-fallback-timeout-injection

Conversation

@murdore

@murdore murdore commented Aug 23, 2026

Copy link
Copy Markdown
Contributor

Why

Two cases in continuous-test-suite-bugfixes fail on CI runners at a 240s bound while passing everywhere locally, blocking #1487. Different bugs, same consequence: a case that cannot be bounded.

1. The proxy-fallback case patched a global to force its timeout

It set globalThis.setTimeout to fire every timer at 0ms, restoring it in a finally. Two problems:

  • The patch is live across an await inside a 280-case suite sharing one process, so for that window it rewrites the delay of every timer created anywhere — including ones belonging to other cases' pending work.
  • If the awaited call never settles, the finally never runs and the patch leaks into the rest of the suite.

Not theoretical. A peer session trying to measure this case wrote a setTimeout-based watchdog around it — and the watchdog was rewritten to 0ms by the patch it was measuring, reporting an instant false hang. Twice, at two different bounds, before the instrument was suspected instead of the code.

executeClaudeFallbackWithRetry now takes an optional idleTimeoutMs (defaulting to FALLBACK_STREAM_IDLE_TIMEOUT_MS); the case passes 1. No global is touched, so the class is gone rather than the instance.

2. spawnSync's timeout is not a guarantee — and it defeats #1487 by construction

The second case drives the CLI through spawnSync with timeout: 10_000. That timeout sends killSignal — SIGTERM by default — and then keeps waiting. A child that ignores SIGTERM is never killed and spawnSync never returns.

Verified directly:

child ignores SIGTERM, killSignal default   → spawnSync NEVER returns
                                               (outer 30s kill, exit 124)
same child, killSignal: "SIGKILL"           → returns at 2003ms,
                                               signal SIGKILL, error ETIMEDOUT

This is the one failure mode that defeats #1487's whole approach. spawnSync blocks the event loop, so a Promise.race per-case bound physically cannot fire while it's stuck — the bound reports only after spawnSync returns, and if it never returns, never. All seven timed subprocess calls in this suite now pass killSignal: "SIGKILL".

Worth noting the suite already asserts production code does this: the audioPlayer case checks source.includes("killSignal") so a hung decoder can't block the CLI. The rule existed; the tests weren't following it.

Proof

The fallback case still catches a real defect — confirmed by removing the abort from the idle-timeout path:

abort removed   ✗ proxy fallback: an idle stream is aborted without ambiguous replay
                  Passed 279  Failed 1
restored        ✓ 280 passed, 0 failed  (3 consecutive runs, 58s each)

What this does not claim

The CI hang itself is still unexplained. Two theories were refuted before this PR: the unpatched path would have waited the real 120s and passed slowly rather than exceeding 240s; and isRetryableNetworkError matches on code, which a TimeoutError doesn't carry, so there's no 2×120s retry path either.

What this removes is the reason those two cases couldn't be bounded or measured honestly when it happens again. If the fallback case hangs after this, the suite bound will actually report it, and any watchdog pointed at it will actually work.

Root cause of the CI failure found and reported by a peer session; the spawnSync mechanism is mine.

…dable

Two cases in continuous-test-suite-bugfixes have been failing on CI runners at
a 240s bound while passing everywhere locally, blocking #1487. They are
unrelated bugs with the same consequence: a case that cannot be bounded.

The proxy-fallback case forced the idle-timeout path by patching
`globalThis.setTimeout` to fire every timer at 0ms, then restoring it in a
`finally`. Two problems. The patch is installed across an await inside a
280-case suite sharing one process, so for that window it rewrites the delay of
EVERY timer created anywhere — including ones belonging to other cases' pending
work. And if the awaited call never settles, the `finally` never runs and the
patch leaks into the rest of the suite.

That is not theoretical. A peer session trying to measure this very case wrote
a `setTimeout`-based watchdog around it, and the watchdog was itself rewritten
to 0ms by the patch it was measuring — reporting an instant false hang, twice,
before the instrument was suspected rather than the code. Same failure at a 90s
bound, same cause.

`executeClaudeFallbackWithRetry` now takes an optional `idleTimeoutMs`,
defaulting to FALLBACK_STREAM_IDLE_TIMEOUT_MS, and the case passes 1. No global
is touched, so the class of problem is gone rather than this instance of it.

The second case is a different mechanism with a sharper edge. It drives the CLI
through `spawnSync` with `timeout: 10_000`, and spawnSync's timeout is not a
guarantee: it sends `killSignal` — SIGTERM by default — and then keeps waiting.
A child that ignores SIGTERM is never killed and spawnSync never returns.
Verified directly: a child running `process.on("SIGTERM",()=>{})` with a live
interval hangs spawnSync indefinitely, while the same call with
`killSignal: "SIGKILL"` returns at the timeout with ETIMEDOUT.

This is the one failure mode that defeats #1487's whole approach. spawnSync
blocks the event loop, so a Promise.race per-case bound physically cannot fire
while it is stuck — the bound reports only after spawnSync returns, and if it
never returns, never. All seven timed subprocess calls in this suite now pass
killSignal SIGKILL.

Worth noting the suite already asserts that production code does this: the
audioPlayer case checks `source.includes("killSignal")` so a hung decoder
cannot block the CLI. The rule existed; the tests just were not following it.

Confirmed the fallback case still catches a real defect by removing the abort
from the idle-timeout path — it fails rather than passing quietly. Three
consecutive full-suite runs: 280 passed, 0 failed, 58s each.

The CI hang itself is not yet explained, and this does not claim to explain it.
What it removes is the reason the two cases could not be bounded or measured
honestly when it happens again.
Copilot AI lite review requested due to automatic review settings August 23, 2026 00:17
@github-actions

Copy link
Copy Markdown
Contributor

✅ Single Commit Policy - COMPLIANT

Status: Policy requirements met • 1 commit • Valid format • Ready for merge

📊 View validation details

📝 Commit Details

  • Hash: 6f8f475b64d9d4b0ca48815ff760fa864c7e5f16
  • Message: fix(test): make the proxy-fallback and subprocess cases actually boundable
  • Author: Sachin Sharma

✅ Validation Results

  • Single commit requirement met
  • No merge commits in branch
  • Semantic commit message format verified
  • Ready for squash merge to release branch

🤖 Automated validation by NeuroLink Single Commit Enforcement

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@github-actions

Copy link
Copy Markdown
Contributor

🤖 AI Review & Build Compliance ✅

Status: AI analysis complete • Build rules validated • Ready for review

📊 View detailed analysis results

🛡️ Analysis Complete

  • ✅ Security scan (vulnerabilities, API keys)
  • ✅ TypeScript safety & code quality
  • ✅ Error handling & best practices
  • ✅ Build rule enforcement validated
  • ✅ Commit format & compliance checks

📋 Ready for Merge When

  • All CI checks passing
  • Manual review approved
  • Any AI-flagged issues resolved

🤖 AI analysis complete - check individual code comments for specific feedback

@coderabbitai

coderabbitai Bot commented Aug 23, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@murdore, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 32 minutes

Limit details: You’ve used all 2 included reviews currently available.

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

How can I continue?

Wait for the limit to reset, then comment @coderabbitai review or push new commits to the PR.

An organization admin can change what happens after included review limits in Billing.

How do review limits work?

CodeRabbit enforces per-developer PR review limits within each organization.

For paid Pro and Pro+ reviews, CodeRabbit uses a developer's included PR review attempts over the past 7 days to set the current hourly allowance. At typical activity levels, the full plan allowance applies. Higher sustained activity can lower the allowance until earlier attempts leave the 7-day window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: 565b6c7e-5841-4aee-a364-0cc16b2b7d8a

📥 Commits

Reviewing files that changed from the base of the PR and between 5c0db3d and 6f8f475.

📒 Files selected for processing (2)
  • src/lib/server/routes/claudeProxyRoutes.ts
  • test/continuous-test-suite-bugfixes.ts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

murdore added a commit that referenced this pull request Aug 23, 2026
`defineSuite` applies a 240s per-case timeout, but it lives in the
`test(name, fn)` helper it hands back. Eighteen suites never use that helper:
they keep their own `tests[]` array and loop `await test.fn()` inside a
`runAllTests` they pass to `runSuite`. `runSuite(body)` only awaits the body —
it adds no timeout of its own — so every case in those suites was UNBOUNDED.

One hung case then hangs the whole run until CI kills the job, and the report
says nothing about which case it was. That is not hypothetical: two of my own
runs of continuous-test-suite-context died on a wall-clock limit this session
with no per-case attribution, and this is why.

`withCaseTimeout` moves to the shared harness rather than being copied per
suite. #1484 added an identical local helper to the proxy suite while fixing
the same bug there; the semantics here match it deliberately so the two cannot
drift, and that suite is left alone to avoid conflicting with an open PR.

19 call sites across 18 suites. The one non-uniform case is autoresearch, which
has two runners — the assertion in the conversion script caught that before
anything was written, rather than silently converting one and leaving the other.

Verified the bound actually fires, since a timeout helper that never triggers
is worse than none:
    normal case returns its value untouched
    a hung case rejects at 301ms against a 300ms bound, naming the case
    the timer does not keep the process alive

AND IT IMMEDIATELY CAUGHT A REAL HANG. `continuous-test-suite-bugfixes` now
reports:
    ✗ proxy fallback: an idle stream is aborted without ambiguous replay
      exceeded 240000ms and was aborted — treat as a hang, not slowness
That case previously hung silently inside a suite that reported 280 passed.
The failure is pre-existing and intermittent; what changed is that it is now
attributable instead of costing four minutes of wall-clock and no information.

Folded in continuous-test-suite-proxy.ts, at its author's request now that
#1484 has merged. Its local `withCaseTimeout` is deleted in favour of the
shared one — two implementations that match today are two that can drift
tomorrow, and there is now exactly one definition in the repo.

The proxy suite keeps its own 180s bound rather than the shared 240s default,
for the reason its author gave: every case there talks to a proxy this repo
builds and spawns, not a live provider, so a hang is a defect and does not
deserve the slack a remote endpoint gets.

Checked the swallow hazard they flagged rather than assuming the shared wording
is safe. `defineSuite` classifies a thrown error as SKIP when the message
matches `isExpectedProviderError()`, and this message contains "aborted", which
is the shape that can be swallowed. Ran the messages through it directly:

    "... exceeded 240000ms and was aborted — treat as a hang, not slowness"
    "... exceeded 180000ms and was aborted — treat as a hang, not slowness"
    "... exceeded 240000ms and was aborted"

all three report as FAILURE, none are swallowed. A timeout that reported as a
skip would be the exact false green this change exists to remove.

Verified after folding: typecheck 4821 files 0 errors, eslint 0 errors,
prettier clean, proxy suite 74 passed / 0 failed / 6 skipped, servers 41
passed / 0 failed, and exactly one `withCaseTimeout` definition repo-wide.

Narrowed the globalThis.setTimeout patch in the proxy-fallback case, because
bounding the cases exposed a cascade that the bound itself creates.

`Promise.race` does not cancel the losing promise. When the harness bound fires
on a hung case, that case KEEPS RUNNING — still inside its `try`, with its
`finally` not yet reached. A case that has patched a global therefore leaks
that patch to every case after it. In CI this showed as two failures rather
than one: the proxy-fallback case timed out, and then

    ✗ CLI setup: --provider <p> --check takes the check-only path
      (no interactive prompt, no hang)

which sits later in the same file, failed in its wake. One hang became two
failures, and the second looked unrelated.

The patch now matches on the delay and collapses only the fallback idle
timeout, passing every other timer through untouched. That keeps the case
testing exactly what it tested, makes the patch inert for other pending work in
a 280-case shared process, and makes it harmless even in the window after an
abandoned run. Proven both directions: an unrelated 400ms timer still waits
403ms, the targeted 120s timer collapses to 1ms.

This does not explain why the case hangs in CI and passes everywhere else —
that is still open, and it is not being papered over. What it removes is the
mechanism by which one unexplained hang corrupts unrelated cases.

The constant is mirrored locally with its provenance rather than exported from
the product: if the product value ever changes, the collapse stops matching and
the case fails loudly at the harness bound instead of passing for the wrong
reason.

bugfixes 280 passed / 0 failed after the change.

Documented what this helper CANNOT bound, because a bound that is believed and
does not work is worse than no bound.

It is `Promise.race`, so the timer only fires when the event loop is free.
Anything that blocks the loop is unprotected: `spawnSync`, `execSync`, a slow
synchronous read. Measured rather than reasoned about — a 300ms bound around a
3s `spawnSync` returned at 3025ms without firing.

That matters here specifically. continuous-test-suite-bugfixes makes 20
`spawnSync` calls, 8 of them with a `timeout`, and `spawnSync`'s timeout is not
a guarantee either: it sends `killSignal`, SIGTERM by default, then keeps
waiting, so a child that ignores SIGTERM hangs it forever. Those cases were
unboundable by construction and this change would have left them looking
protected. Credit to cli-support-83, who found it from the second CI failure I
sent them and is fixing the call sites in #1490.

Also recorded on the helper: `Promise.race` does not cancel the losing promise,
so an abandoned case keeps running with its `finally` unreached — which is the
cascade that made one hang produce two CI failures, and the reason the
proxy-fallback case's global patch is now scoped by delay.
@Tara-ag

Tara-ag commented Aug 23, 2026

Copy link
Copy Markdown
Contributor

Review Summary

Decision: APPROVED ✅

This PR addresses two critical issues blocking issue #1487 from progressing in CI:

Changes Overview

1. src/lib/server/routes/claudeProxyRoutes.ts - Proxy Fallback Timeout Fix

  • Added optional idleTimeoutMs parameter to executeClaudeFallbackTranslation()
  • Allows tests to inject specific timeout values instead of patching global timers
  • Critical improvement: Eliminates dangerous side effect where patching globalThis.setTimeout would affect ALL timers in a 280-case test suite sharing one process
  • No breaking changes - parameter is optional with sensible default

2. test/continuous-test-suite-bugfixes.ts - Multiple Fixes

  • Removed global timer patching from proxy fallback test (was rewriting all timers)
  • Added killSignal: "SIGKILL" to subprocess spawns (prevents indefinite hangs when children ignore SIGTERM)
  • Comprehensive inline documentation explaining why these fixes matter

Impact on Existing Code

  • ✅ Zero breaking changes - only adding optional parameters
  • ✅ Test improvements - makes the test suite more reliable and isolated
  • ✅ Production safety - prevents processes from blocking the event loop indefinitely
  • ✅ Well documented - every change explains the "why" with context

Quality Checks Passed

  • TypeScript types are correct (no any, proper optionality)
  • No interface usage (follows project standard of type)
  • No double type assertions
  • Comments explain reasoning, not just mechanics
  • No hardcoded secrets or credentials
  • No breaking changes to public API
  • Changes are minimal and targeted

Recommendation

APPROVE this PR because:

  1. It fixes real, documented bugs causing CI failures
  2. All changes are well-motivated and explained
  3. No risks introduced to production code or existing functionality
  4. Improves test suite reliability long-term
  5. Follows all project standards and conventions

@murdore
murdore merged commit c91edf0 into release Aug 23, 2026
20 of 21 checks passed
@murdore
murdore deleted the fix/proxy-fallback-timeout-injection branch August 23, 2026 00:36
@Tara-ag

Tara-ag commented Aug 23, 2026

Copy link
Copy Markdown
Contributor

Review Summary

Decision: APPROVED ✅

This PR fixes test infrastructure issues in continuous-test-suite-bugfixes.ts that were causing CI failures. The changes are focused on test-only code and do not modify production behavior.

Changes Overview

File 1: src/lib/server/routes/claudeProxyRoutes.ts

  • Added optional idleTimeoutMs?: number parameter to executeClaudeFallbackTranslation function
  • This allows tests to inject a custom timeout value instead of patching global timers
  • Default value maintains existing behavior (uses FALLBACK_STREAM_IDLE_TIMEOUT_MS)
  • Backward compatible - all existing callers continue to work without modification

File 2: test/continuous-test-suite-bugfixes.ts

  • Removed dangerous global setTimeout patching: Previously the test patched globalThis.setTimeout to fire all timers at 0ms. This caused race conditions because it affected ALL timers in the process, including those belonging to other concurrent test cases.
  • Added killSignal: "SIGKILL" to subprocess calls: Seven instances of spawnSync now use killSignal: "SIGKILL" instead of relying on default SIGTERM. This prevents hung child processes from blocking the event loop, as SIGKILL immediately terminates them while SIGTERM can be ignored by certain processes.
  • Added comprehensive comments explaining why these fixes are necessary

Impact Analysis

✅ Safe backward compatibility: The added parameter is optional with a sensible default. All production callers pass through executeClaudeFallbackWithRetry → executeClaudeFallbackTranslation, so they automatically get the default value.

✅ No breaking changes: Only test infrastructure is modified; production API surface unchanged.

✅ Test suite improvements: These fixes address real bugs that prevented the test suite from running reliably on CI runners with time bounds.

Findings

No issues found. The PR is clean and ready for merge.


Review scope: Test infrastructure fixes for timing and process handling in bugfix test suite

@github-actions

Copy link
Copy Markdown
Contributor

🎉 This PR is included in version 11.20.1 🎉

The release is available on:

Your semantic-release bot 📦🚀

murdore added a commit that referenced this pull request Aug 23, 2026
`defineSuite` applies a 240s per-case timeout, but it lives in the
`test(name, fn)` helper it hands back. Nineteen suites never use that helper:
they keep their own `tests[]` array and loop `await test.fn()` inside a runner
passed to `runSuite`. `runSuite(body)` only awaits the body — it adds no
timeout of its own — so every case in those suites was UNBOUNDED.

One hung case then hangs the whole run until CI kills the job, and the report
says nothing about which case it was. Two of my own runs of
continuous-test-suite-context died on a wall-clock limit with no per-case
attribution; this is why.

`withCaseTimeout` lives in the shared harness rather than being copied per
suite. 20 call sites across 19 suites. The proxy suite keeps its own 180s bound
because every case there talks to a proxy this repo spawns, not a live
provider, so a hang is a defect rather than an environment problem.

WHAT IT CANNOT BOUND is documented on the helper, because a bound that is
believed and does not work is worse than no bound. It is `Promise.race`, so the
timer only fires when the event loop is free: `spawnSync`, `execSync` and other
synchronous blocking are unprotected. Measured — a 300ms bound around a 3s
`spawnSync` returned at 3025ms without firing. The corollary is recorded too:
`spawnSync`'s own `timeout` sends SIGTERM and keeps waiting, so a child that
ignores it hangs forever; every call site needs `killSignal: "SIGKILL"`.
cli-support-83 found that from a CI failure this change surfaced, and fixed all
ten call sites in #1490.

AND IT FAILS CLOSED. `Promise.race` does not cancel the loser, so an abandoned
case keeps running and still holds whatever it mutated — a patched global, an
open server, a temp dir. Every later case then runs in a process it does not
own, and its PASS or FAIL means nothing. That was not theoretical: one hang
produced a second, unrelated CI failure 176 lines further down the same file.
So once a case has been abandoned the helper refuses to START another, and says
why. Reporting the rest as skips would be exactly the false green this change
exists to remove.

`CaseTimeoutError` is exported so a runner can tell a bound from a real failure.

Verified rather than asserted:
  a normal case passes through untouched, three in sequence
  a hung case rejects at 301ms against a 300ms bound, naming the case
  the rejection is a CaseTimeoutError, distinguishable from a test failure
  the next case is refused WITHOUT being started
  the timer does not keep the process alive
  a 400ms timer is unaffected while a targeted one collapses — no global perturbation

typecheck 4821 files 0 errors, eslint 0 errors, prettier clean.
bugfixes 280 passed, servers 41 passed, proxy 74 passed / 6 skipped.
tts reports 17 passed / 1 failed (Fish Audio) here AND on release — pre-existing.
murdore added a commit that referenced this pull request Aug 23, 2026
`defineSuite` applies a 240s per-case timeout, but it lives in the
`test(name, fn)` helper it hands back. Nineteen suites never use that helper:
they keep their own `tests[]` array and loop `await test.fn()` inside a runner
passed to `runSuite`. `runSuite(body)` only awaits the body — it adds no
timeout of its own — so every case in those suites was UNBOUNDED.

One hung case then hangs the whole run until CI kills the job, and the report
says nothing about which case it was. Two of my own runs of
continuous-test-suite-context died on a wall-clock limit with no per-case
attribution; this is why.

`withCaseTimeout` lives in the shared harness rather than being copied per
suite. 20 call sites across 19 suites. The proxy suite keeps its own 180s bound
because every case there talks to a proxy this repo spawns, not a live
provider, so a hang is a defect rather than an environment problem.

WHAT IT CANNOT BOUND is documented on the helper, because a bound that is
believed and does not work is worse than no bound. It is `Promise.race`, so the
timer only fires when the event loop is free: `spawnSync`, `execSync` and other
synchronous blocking are unprotected. Measured — a 300ms bound around a 3s
`spawnSync` returned at 3025ms without firing. The corollary is recorded too:
`spawnSync`'s own `timeout` sends SIGTERM and keeps waiting, so a child that
ignores it hangs forever; every call site needs `killSignal: "SIGKILL"`.
cli-support-83 found that from a CI failure this change surfaced, and fixed all
ten call sites in #1490.

AND IT FAILS CLOSED. `Promise.race` does not cancel the loser, so an abandoned
case keeps running and still holds whatever it mutated — a patched global, an
open server, a temp dir. Every later case then runs in a process it does not
own, and its PASS or FAIL means nothing. That was not theoretical: one hang
produced a second, unrelated CI failure 176 lines further down the same file.
So once a case has been abandoned the helper refuses to START another, and says
why. Reporting the rest as skips would be exactly the false green this change
exists to remove.

`CaseTimeoutError` is exported so a runner can tell a bound from a real failure.

Verified rather than asserted:
  a normal case passes through untouched, three in sequence
  a hung case rejects at 301ms against a 300ms bound, naming the case
  the rejection is a CaseTimeoutError, distinguishable from a test failure
  the next case is refused WITHOUT being started
  the timer does not keep the process alive
  a 400ms timer is unaffected while a targeted one collapses — no global perturbation

typecheck 4821 files 0 errors, eslint 0 errors, prettier clean.
bugfixes 280 passed, servers 41 passed, proxy 74 passed / 6 skipped.
tts reports 17 passed / 1 failed (Fish Audio) here AND on release — pre-existing.

Corrected the proxy suite's timeout rationale. The comment claimed every case
there drives a locally spawned proxy and never a live provider. That is false:
nine cases gate on `hasValidCredentials()` and, when it is true, reach Anthropic
through the spawned proxy. Verified by counting the call sites rather than
taking the review comment's word for it.

180s stays — it is comfortably above a live round-trip and the proxy under test
is still one this repo builds — but the justification now says what is
actually true, including that a breach is usually a defect and occasionally a
slow upstream. A case that legitimately needs longer should carry its own bound
rather than have this one raised for everyone.
murdore added a commit that referenced this pull request Aug 23, 2026
`defineSuite` applies a 240s per-case timeout, but it lives in the
`test(name, fn)` helper it hands back. Nineteen suites never use that helper:
they keep their own `tests[]` array and loop `await test.fn()` inside a runner
passed to `runSuite`. `runSuite(body)` only awaits the body — it adds no
timeout of its own — so every case in those suites was UNBOUNDED.

One hung case then hangs the whole run until CI kills the job, and the report
says nothing about which case it was. Two of my own runs of
continuous-test-suite-context died on a wall-clock limit with no per-case
attribution; this is why.

`withCaseTimeout` lives in the shared harness rather than being copied per
suite. 20 call sites across 19 suites. The proxy suite keeps its own 180s bound
because every case there talks to a proxy this repo spawns, not a live
provider, so a hang is a defect rather than an environment problem.

WHAT IT CANNOT BOUND is documented on the helper, because a bound that is
believed and does not work is worse than no bound. It is `Promise.race`, so the
timer only fires when the event loop is free: `spawnSync`, `execSync` and other
synchronous blocking are unprotected. Measured — a 300ms bound around a 3s
`spawnSync` returned at 3025ms without firing. The corollary is recorded too:
`spawnSync`'s own `timeout` sends SIGTERM and keeps waiting, so a child that
ignores it hangs forever; every call site needs `killSignal: "SIGKILL"`.
cli-support-83 found that from a CI failure this change surfaced, and fixed all
ten call sites in #1490.

AND IT FAILS CLOSED. `Promise.race` does not cancel the loser, so an abandoned
case keeps running and still holds whatever it mutated — a patched global, an
open server, a temp dir. Every later case then runs in a process it does not
own, and its PASS or FAIL means nothing. That was not theoretical: one hang
produced a second, unrelated CI failure 176 lines further down the same file.
So once a case has been abandoned the helper refuses to START another, and says
why. Reporting the rest as skips would be exactly the false green this change
exists to remove.

`CaseTimeoutError` is exported so a runner can tell a bound from a real failure.

Verified rather than asserted:
  a normal case passes through untouched, three in sequence
  a hung case rejects at 301ms against a 300ms bound, naming the case
  the rejection is a CaseTimeoutError, distinguishable from a test failure
  the next case is refused WITHOUT being started
  the timer does not keep the process alive
  a 400ms timer is unaffected while a targeted one collapses — no global perturbation

typecheck 4821 files 0 errors, eslint 0 errors, prettier clean.
bugfixes 280 passed, servers 41 passed, proxy 74 passed / 6 skipped.
tts reports 17 passed / 1 failed (Fish Audio) here AND on release — pre-existing.

Corrected the proxy suite's timeout rationale. The comment claimed every case
there drives a locally spawned proxy and never a live provider. That is false:
nine cases gate on `hasValidCredentials()` and, when it is true, reach Anthropic
through the spawned proxy. Verified by counting the call sites rather than
taking the review comment's word for it.

180s stays — it is comfortably above a live round-trip and the proxy under test
is still one this repo builds — but the justification now says what is
actually true, including that a breach is usually a defect and occasionally a
slow upstream. A case that legitimately needs longer should carry its own bound
rather than have this one raised for everyone.
murdore added a commit that referenced this pull request Aug 23, 2026
`defineSuite` applies a 240s per-case timeout, but it lives in the
`test(name, fn)` helper it hands back. Nineteen suites never use that helper:
they keep their own `tests[]` array and loop `await test.fn()` inside a runner
passed to `runSuite`. `runSuite(body)` only awaits the body — it adds no
timeout of its own — so every case in those suites was UNBOUNDED.

One hung case then hangs the whole run until CI kills the job, and the report
says nothing about which case it was. Two of my own runs of
continuous-test-suite-context died on a wall-clock limit with no per-case
attribution; this is why.

`withCaseTimeout` lives in the shared harness rather than being copied per
suite. 20 call sites across 19 suites. The proxy suite keeps its own 180s bound
because every case there talks to a proxy this repo spawns, not a live
provider, so a hang is a defect rather than an environment problem.

WHAT IT CANNOT BOUND is documented on the helper, because a bound that is
believed and does not work is worse than no bound. It is `Promise.race`, so the
timer only fires when the event loop is free: `spawnSync`, `execSync` and other
synchronous blocking are unprotected. Measured — a 300ms bound around a 3s
`spawnSync` returned at 3025ms without firing. The corollary is recorded too:
`spawnSync`'s own `timeout` sends SIGTERM and keeps waiting, so a child that
ignores it hangs forever; every call site needs `killSignal: "SIGKILL"`.
cli-support-83 found that from a CI failure this change surfaced, and fixed all
ten call sites in #1490.

AND IT FAILS CLOSED. `Promise.race` does not cancel the loser, so an abandoned
case keeps running and still holds whatever it mutated — a patched global, an
open server, a temp dir. Every later case then runs in a process it does not
own, and its PASS or FAIL means nothing. That was not theoretical: one hang
produced a second, unrelated CI failure 176 lines further down the same file.
So once a case has been abandoned the helper refuses to START another, and says
why. Reporting the rest as skips would be exactly the false green this change
exists to remove.

`CaseTimeoutError` is exported so a runner can tell a bound from a real failure.

Verified rather than asserted:
  a normal case passes through untouched, three in sequence
  a hung case rejects at 301ms against a 300ms bound, naming the case
  the rejection is a CaseTimeoutError, distinguishable from a test failure
  the next case is refused WITHOUT being started
  the timer does not keep the process alive
  a 400ms timer is unaffected while a targeted one collapses — no global perturbation

typecheck 4821 files 0 errors, eslint 0 errors, prettier clean.
bugfixes 280 passed, servers 41 passed, proxy 74 passed / 6 skipped.
tts reports 17 passed / 1 failed (Fish Audio) here AND on release — pre-existing.

Corrected the proxy suite's timeout rationale. The comment claimed every case
there drives a locally spawned proxy and never a live provider. That is false:
nine cases gate on `hasValidCredentials()` and, when it is true, reach Anthropic
through the spawned proxy. Verified by counting the call sites rather than
taking the review comment's word for it.

180s stays — it is comfortably above a live round-trip and the proxy under test
is still one this repo builds — but the justification now says what is
actually true, including that a breach is usually a defect and occasionally a
slow upstream. A case that legitimately needs longer should carry its own bound
rather than have this one raised for everyone.
murdore added a commit that referenced this pull request Aug 23, 2026
`defineSuite` applies a 240s per-case timeout, but it lives in the
`test(name, fn)` helper it hands back. Nineteen suites never use that helper:
they keep their own `tests[]` array and loop `await test.fn()` inside a runner
passed to `runSuite`. `runSuite(body)` only awaits the body — it adds no
timeout of its own — so every case in those suites was UNBOUNDED.

One hung case then hangs the whole run until CI kills the job, and the report
says nothing about which case it was. Two of my own runs of
continuous-test-suite-context died on a wall-clock limit with no per-case
attribution; this is why.

`withCaseTimeout` lives in the shared harness rather than being copied per
suite. 20 call sites across 19 suites. The proxy suite keeps its own 180s bound
because every case there talks to a proxy this repo spawns, not a live
provider, so a hang is a defect rather than an environment problem.

WHAT IT CANNOT BOUND is documented on the helper, because a bound that is
believed and does not work is worse than no bound. It is `Promise.race`, so the
timer only fires when the event loop is free: `spawnSync`, `execSync` and other
synchronous blocking are unprotected. Measured — a 300ms bound around a 3s
`spawnSync` returned at 3025ms without firing. The corollary is recorded too:
`spawnSync`'s own `timeout` sends SIGTERM and keeps waiting, so a child that
ignores it hangs forever; every call site needs `killSignal: "SIGKILL"`.
cli-support-83 found that from a CI failure this change surfaced, and fixed all
ten call sites in #1490.

AND IT FAILS CLOSED. `Promise.race` does not cancel the loser, so an abandoned
case keeps running and still holds whatever it mutated — a patched global, an
open server, a temp dir. Every later case then runs in a process it does not
own, and its PASS or FAIL means nothing. That was not theoretical: one hang
produced a second, unrelated CI failure 176 lines further down the same file.
So once a case has been abandoned the helper refuses to START another, and says
why. Reporting the rest as skips would be exactly the false green this change
exists to remove.

`CaseTimeoutError` is exported so a runner can tell a bound from a real failure.

Verified rather than asserted:
  a normal case passes through untouched, three in sequence
  a hung case rejects at 301ms against a 300ms bound, naming the case
  the rejection is a CaseTimeoutError, distinguishable from a test failure
  the next case is refused WITHOUT being started
  the timer does not keep the process alive
  a 400ms timer is unaffected while a targeted one collapses — no global perturbation

typecheck 4821 files 0 errors, eslint 0 errors, prettier clean.
bugfixes 280 passed, servers 41 passed, proxy 74 passed / 6 skipped.
tts reports 17 passed / 1 failed (Fish Audio) here AND on release — pre-existing.

Corrected the proxy suite's timeout rationale. The comment claimed every case
there drives a locally spawned proxy and never a live provider. That is false:
nine cases gate on `hasValidCredentials()` and, when it is true, reach Anthropic
through the spawned proxy. Verified by counting the call sites rather than
taking the review comment's word for it.

180s stays — it is comfortably above a live round-trip and the proxy under test
is still one this repo builds — but the justification now says what is
actually true, including that a breach is usually a defect and occasionally a
slow upstream. A case that legitimately needs longer should carry its own bound
rather than have this one raised for everyone.
murdore added a commit that referenced this pull request Aug 23, 2026
`defineSuite` applies a 240s per-case timeout, but it lives in the
`test(name, fn)` helper it hands back. Nineteen suites never use that helper:
they keep their own `tests[]` array and loop `await test.fn()` inside a runner
passed to `runSuite`. `runSuite(body)` only awaits the body — it adds no
timeout of its own — so every case in those suites was UNBOUNDED.

One hung case then hangs the whole run until CI kills the job, and the report
says nothing about which case it was. Two of my own runs of
continuous-test-suite-context died on a wall-clock limit with no per-case
attribution; this is why.

`withCaseTimeout` lives in the shared harness rather than being copied per
suite. 20 call sites across 19 suites. The proxy suite keeps its own 180s bound
because every case there talks to a proxy this repo spawns, not a live
provider, so a hang is a defect rather than an environment problem.

WHAT IT CANNOT BOUND is documented on the helper, because a bound that is
believed and does not work is worse than no bound. It is `Promise.race`, so the
timer only fires when the event loop is free: `spawnSync`, `execSync` and other
synchronous blocking are unprotected. Measured — a 300ms bound around a 3s
`spawnSync` returned at 3025ms without firing. The corollary is recorded too:
`spawnSync`'s own `timeout` sends SIGTERM and keeps waiting, so a child that
ignores it hangs forever; every call site needs `killSignal: "SIGKILL"`.
cli-support-83 found that from a CI failure this change surfaced, and fixed all
ten call sites in #1490.

AND IT FAILS CLOSED. `Promise.race` does not cancel the loser, so an abandoned
case keeps running and still holds whatever it mutated — a patched global, an
open server, a temp dir. Every later case then runs in a process it does not
own, and its PASS or FAIL means nothing. That was not theoretical: one hang
produced a second, unrelated CI failure 176 lines further down the same file.
So once a case has been abandoned the helper refuses to START another, and says
why. Reporting the rest as skips would be exactly the false green this change
exists to remove.

`CaseTimeoutError` is exported so a runner can tell a bound from a real failure.

Verified rather than asserted:
  a normal case passes through untouched, three in sequence
  a hung case rejects at 301ms against a 300ms bound, naming the case
  the rejection is a CaseTimeoutError, distinguishable from a test failure
  the next case is refused WITHOUT being started
  the timer does not keep the process alive
  a 400ms timer is unaffected while a targeted one collapses — no global perturbation

typecheck 4821 files 0 errors, eslint 0 errors, prettier clean.
bugfixes 280 passed, servers 41 passed, proxy 74 passed / 6 skipped.
tts reports 17 passed / 1 failed (Fish Audio) here AND on release — pre-existing.

Corrected the proxy suite's timeout rationale. The comment claimed every case
there drives a locally spawned proxy and never a live provider. That is false:
nine cases gate on `hasValidCredentials()` and, when it is true, reach Anthropic
through the spawned proxy. Verified by counting the call sites rather than
taking the review comment's word for it.

180s stays — it is comfortably above a live round-trip and the proxy under test
is still one this repo builds — but the justification now says what is
actually true, including that a breach is usually a defect and occasionally a
slow upstream. A case that legitimately needs longer should carry its own bound
rather than have this one raised for everyone.
murdore added a commit that referenced this pull request Aug 23, 2026
`defineSuite` applies a 240s per-case timeout, but it lives in the
`test(name, fn)` helper it hands back. Nineteen suites never use that helper:
they keep their own `tests[]` array and loop `await test.fn()` inside a runner
passed to `runSuite`. `runSuite(body)` only awaits the body — it adds no
timeout of its own — so every case in those suites was UNBOUNDED.

One hung case then hangs the whole run until CI kills the job, and the report
says nothing about which case it was. Two of my own runs of
continuous-test-suite-context died on a wall-clock limit with no per-case
attribution; this is why.

`withCaseTimeout` lives in the shared harness rather than being copied per
suite. 20 call sites across 19 suites. The proxy suite keeps its own 180s bound
because every case there talks to a proxy this repo spawns, not a live
provider, so a hang is a defect rather than an environment problem.

WHAT IT CANNOT BOUND is documented on the helper, because a bound that is
believed and does not work is worse than no bound. It is `Promise.race`, so the
timer only fires when the event loop is free: `spawnSync`, `execSync` and other
synchronous blocking are unprotected. Measured — a 300ms bound around a 3s
`spawnSync` returned at 3025ms without firing. The corollary is recorded too:
`spawnSync`'s own `timeout` sends SIGTERM and keeps waiting, so a child that
ignores it hangs forever; every call site needs `killSignal: "SIGKILL"`.
cli-support-83 found that from a CI failure this change surfaced, and fixed all
ten call sites in #1490.

AND IT FAILS CLOSED. `Promise.race` does not cancel the loser, so an abandoned
case keeps running and still holds whatever it mutated — a patched global, an
open server, a temp dir. Every later case then runs in a process it does not
own, and its PASS or FAIL means nothing. That was not theoretical: one hang
produced a second, unrelated CI failure 176 lines further down the same file.
So once a case has been abandoned the helper refuses to START another, and says
why. Reporting the rest as skips would be exactly the false green this change
exists to remove.

`CaseTimeoutError` is exported so a runner can tell a bound from a real failure.

Verified rather than asserted:
  a normal case passes through untouched, three in sequence
  a hung case rejects at 301ms against a 300ms bound, naming the case
  the rejection is a CaseTimeoutError, distinguishable from a test failure
  the next case is refused WITHOUT being started
  the timer does not keep the process alive
  a 400ms timer is unaffected while a targeted one collapses — no global perturbation

typecheck 4821 files 0 errors, eslint 0 errors, prettier clean.
bugfixes 280 passed, servers 41 passed, proxy 74 passed / 6 skipped.
tts reports 17 passed / 1 failed (Fish Audio) here AND on release — pre-existing.

Corrected the proxy suite's timeout rationale. The comment claimed every case
there drives a locally spawned proxy and never a live provider. That is false:
nine cases gate on `hasValidCredentials()` and, when it is true, reach Anthropic
through the spawned proxy. Verified by counting the call sites rather than
taking the review comment's word for it.

180s stays — it is comfortably above a live round-trip and the proxy under test
is still one this repo builds — but the justification now says what is
actually true, including that a breach is usually a defect and occasionally a
slow upstream. A case that legitimately needs longer should carry its own bound
rather than have this one raised for everyone.
murdore added a commit that referenced this pull request Aug 23, 2026
`defineSuite` applies a 240s per-case timeout, but it lives in the
`test(name, fn)` helper it hands back. Nineteen suites never use that helper:
they keep their own `tests[]` array and loop `await test.fn()` inside a runner
passed to `runSuite`. `runSuite(body)` only awaits the body — it adds no
timeout of its own — so every case in those suites was UNBOUNDED.

One hung case then hangs the whole run until CI kills the job, and the report
says nothing about which case it was. Two of my own runs of
continuous-test-suite-context died on a wall-clock limit with no per-case
attribution; this is why.

`withCaseTimeout` lives in the shared harness rather than being copied per
suite. 20 call sites across 19 suites. The proxy suite keeps its own 180s bound
because every case there talks to a proxy this repo spawns, not a live
provider, so a hang is a defect rather than an environment problem.

WHAT IT CANNOT BOUND is documented on the helper, because a bound that is
believed and does not work is worse than no bound. It is `Promise.race`, so the
timer only fires when the event loop is free: `spawnSync`, `execSync` and other
synchronous blocking are unprotected. Measured — a 300ms bound around a 3s
`spawnSync` returned at 3025ms without firing. The corollary is recorded too:
`spawnSync`'s own `timeout` sends SIGTERM and keeps waiting, so a child that
ignores it hangs forever; every call site needs `killSignal: "SIGKILL"`.
cli-support-83 found that from a CI failure this change surfaced, and fixed all
ten call sites in #1490.

AND IT FAILS CLOSED. `Promise.race` does not cancel the loser, so an abandoned
case keeps running and still holds whatever it mutated — a patched global, an
open server, a temp dir. Every later case then runs in a process it does not
own, and its PASS or FAIL means nothing. That was not theoretical: one hang
produced a second, unrelated CI failure 176 lines further down the same file.
So once a case has been abandoned the helper refuses to START another, and says
why. Reporting the rest as skips would be exactly the false green this change
exists to remove.

`CaseTimeoutError` is exported so a runner can tell a bound from a real failure.

Verified rather than asserted:
  a normal case passes through untouched, three in sequence
  a hung case rejects at 301ms against a 300ms bound, naming the case
  the rejection is a CaseTimeoutError, distinguishable from a test failure
  the next case is refused WITHOUT being started
  the timer does not keep the process alive
  a 400ms timer is unaffected while a targeted one collapses — no global perturbation

typecheck 4821 files 0 errors, eslint 0 errors, prettier clean.
bugfixes 280 passed, servers 41 passed, proxy 74 passed / 6 skipped.
tts reports 17 passed / 1 failed (Fish Audio) here AND on release — pre-existing.

Corrected the proxy suite's timeout rationale. The comment claimed every case
there drives a locally spawned proxy and never a live provider. That is false:
nine cases gate on `hasValidCredentials()` and, when it is true, reach Anthropic
through the spawned proxy. Verified by counting the call sites rather than
taking the review comment's word for it.

180s stays — it is comfortably above a live round-trip and the proxy under test
is still one this repo builds — but the justification now says what is
actually true, including that a breach is usually a defect and occasionally a
slow upstream. A case that legitimately needs longer should carry its own bound
rather than have this one raised for everyone.
murdore added a commit that referenced this pull request Aug 23, 2026
`defineSuite` applies a 240s per-case timeout, but it lives in the
`test(name, fn)` helper it hands back. Nineteen suites never use that helper:
they keep their own `tests[]` array and loop `await test.fn()` inside a runner
passed to `runSuite`. `runSuite(body)` only awaits the body — it adds no
timeout of its own — so every case in those suites was UNBOUNDED.

One hung case then hangs the whole run until CI kills the job, and the report
says nothing about which case it was. Two of my own runs of
continuous-test-suite-context died on a wall-clock limit with no per-case
attribution; this is why.

`withCaseTimeout` lives in the shared harness rather than being copied per
suite. 20 call sites across 19 suites. The proxy suite keeps its own 180s bound
because every case there talks to a proxy this repo spawns, not a live
provider, so a hang is a defect rather than an environment problem.

WHAT IT CANNOT BOUND is documented on the helper, because a bound that is
believed and does not work is worse than no bound. It is `Promise.race`, so the
timer only fires when the event loop is free: `spawnSync`, `execSync` and other
synchronous blocking are unprotected. Measured — a 300ms bound around a 3s
`spawnSync` returned at 3025ms without firing. The corollary is recorded too:
`spawnSync`'s own `timeout` sends SIGTERM and keeps waiting, so a child that
ignores it hangs forever; every call site needs `killSignal: "SIGKILL"`.
cli-support-83 found that from a CI failure this change surfaced, and fixed all
ten call sites in #1490.

AND IT FAILS CLOSED. `Promise.race` does not cancel the loser, so an abandoned
case keeps running and still holds whatever it mutated — a patched global, an
open server, a temp dir. Every later case then runs in a process it does not
own, and its PASS or FAIL means nothing. That was not theoretical: one hang
produced a second, unrelated CI failure 176 lines further down the same file.
So once a case has been abandoned the helper refuses to START another, and says
why. Reporting the rest as skips would be exactly the false green this change
exists to remove.

`CaseTimeoutError` is exported so a runner can tell a bound from a real failure.

Verified rather than asserted:
  a normal case passes through untouched, three in sequence
  a hung case rejects at 301ms against a 300ms bound, naming the case
  the rejection is a CaseTimeoutError, distinguishable from a test failure
  the next case is refused WITHOUT being started
  the timer does not keep the process alive
  a 400ms timer is unaffected while a targeted one collapses — no global perturbation

typecheck 4821 files 0 errors, eslint 0 errors, prettier clean.
bugfixes 280 passed, servers 41 passed, proxy 74 passed / 6 skipped.
tts reports 17 passed / 1 failed (Fish Audio) here AND on release — pre-existing.

Corrected the proxy suite's timeout rationale. The comment claimed every case
there drives a locally spawned proxy and never a live provider. That is false:
nine cases gate on `hasValidCredentials()` and, when it is true, reach Anthropic
through the spawned proxy. Verified by counting the call sites rather than
taking the review comment's word for it.

180s stays — it is comfortably above a live round-trip and the proxy under test
is still one this repo builds — but the justification now says what is
actually true, including that a breach is usually a defect and occasionally a
slow upstream. A case that legitimately needs longer should carry its own bound
rather than have this one raised for everyone.
murdore added a commit that referenced this pull request Aug 23, 2026
`defineSuite` applies a 240s per-case timeout, but it lives in the
`test(name, fn)` helper it hands back. Nineteen suites never use that helper:
they keep their own `tests[]` array and loop `await test.fn()` inside a runner
passed to `runSuite`. `runSuite(body)` only awaits the body — it adds no
timeout of its own — so every case in those suites was UNBOUNDED.

One hung case then hangs the whole run until CI kills the job, and the report
says nothing about which case it was. Two of my own runs of
continuous-test-suite-context died on a wall-clock limit with no per-case
attribution; this is why.

`withCaseTimeout` lives in the shared harness rather than being copied per
suite. 20 call sites across 19 suites. The proxy suite keeps its own 180s bound
because every case there talks to a proxy this repo spawns, not a live
provider, so a hang is a defect rather than an environment problem.

WHAT IT CANNOT BOUND is documented on the helper, because a bound that is
believed and does not work is worse than no bound. It is `Promise.race`, so the
timer only fires when the event loop is free: `spawnSync`, `execSync` and other
synchronous blocking are unprotected. Measured — a 300ms bound around a 3s
`spawnSync` returned at 3025ms without firing. The corollary is recorded too:
`spawnSync`'s own `timeout` sends SIGTERM and keeps waiting, so a child that
ignores it hangs forever; every call site needs `killSignal: "SIGKILL"`.
cli-support-83 found that from a CI failure this change surfaced, and fixed all
ten call sites in #1490.

AND IT FAILS CLOSED. `Promise.race` does not cancel the loser, so an abandoned
case keeps running and still holds whatever it mutated — a patched global, an
open server, a temp dir. Every later case then runs in a process it does not
own, and its PASS or FAIL means nothing. That was not theoretical: one hang
produced a second, unrelated CI failure 176 lines further down the same file.
So once a case has been abandoned the helper refuses to START another, and says
why. Reporting the rest as skips would be exactly the false green this change
exists to remove.

`CaseTimeoutError` is exported so a runner can tell a bound from a real failure.

Verified rather than asserted:
  a normal case passes through untouched, three in sequence
  a hung case rejects at 301ms against a 300ms bound, naming the case
  the rejection is a CaseTimeoutError, distinguishable from a test failure
  the next case is refused WITHOUT being started
  the timer does not keep the process alive
  a 400ms timer is unaffected while a targeted one collapses — no global perturbation

typecheck 4821 files 0 errors, eslint 0 errors, prettier clean.
bugfixes 280 passed, servers 41 passed, proxy 74 passed / 6 skipped.
tts reports 17 passed / 1 failed (Fish Audio) here AND on release — pre-existing.

Corrected the proxy suite's timeout rationale. The comment claimed every case
there drives a locally spawned proxy and never a live provider. That is false:
nine cases gate on `hasValidCredentials()` and, when it is true, reach Anthropic
through the spawned proxy. Verified by counting the call sites rather than
taking the review comment's word for it.

180s stays — it is comfortably above a live round-trip and the proxy under test
is still one this repo builds — but the justification now says what is
actually true, including that a breach is usually a defect and occasionally a
slow upstream. A case that legitimately needs longer should carry its own bound
rather than have this one raised for everyone.
murdore added a commit that referenced this pull request Aug 23, 2026
`defineSuite` applies a 240s per-case timeout, but it lives in the
`test(name, fn)` helper it hands back. Nineteen suites never use that helper:
they keep their own `tests[]` array and loop `await test.fn()` inside a runner
passed to `runSuite`. `runSuite(body)` only awaits the body — it adds no
timeout of its own — so every case in those suites was UNBOUNDED.

One hung case then hangs the whole run until CI kills the job, and the report
says nothing about which case it was. Two of my own runs of
continuous-test-suite-context died on a wall-clock limit with no per-case
attribution; this is why.

`withCaseTimeout` lives in the shared harness rather than being copied per
suite. 20 call sites across 19 suites. The proxy suite keeps its own 180s bound
because every case there talks to a proxy this repo spawns, not a live
provider, so a hang is a defect rather than an environment problem.

WHAT IT CANNOT BOUND is documented on the helper, because a bound that is
believed and does not work is worse than no bound. It is `Promise.race`, so the
timer only fires when the event loop is free: `spawnSync`, `execSync` and other
synchronous blocking are unprotected. Measured — a 300ms bound around a 3s
`spawnSync` returned at 3025ms without firing. The corollary is recorded too:
`spawnSync`'s own `timeout` sends SIGTERM and keeps waiting, so a child that
ignores it hangs forever; every call site needs `killSignal: "SIGKILL"`.
cli-support-83 found that from a CI failure this change surfaced, and fixed all
ten call sites in #1490.

AND IT FAILS CLOSED. `Promise.race` does not cancel the loser, so an abandoned
case keeps running and still holds whatever it mutated — a patched global, an
open server, a temp dir. Every later case then runs in a process it does not
own, and its PASS or FAIL means nothing. That was not theoretical: one hang
produced a second, unrelated CI failure 176 lines further down the same file.
So once a case has been abandoned the helper refuses to START another, and says
why. Reporting the rest as skips would be exactly the false green this change
exists to remove.

`CaseTimeoutError` is exported so a runner can tell a bound from a real failure.

Verified rather than asserted:
  a normal case passes through untouched, three in sequence
  a hung case rejects at 301ms against a 300ms bound, naming the case
  the rejection is a CaseTimeoutError, distinguishable from a test failure
  the next case is refused WITHOUT being started
  the timer does not keep the process alive
  a 400ms timer is unaffected while a targeted one collapses — no global perturbation

typecheck 4821 files 0 errors, eslint 0 errors, prettier clean.
bugfixes 280 passed, servers 41 passed, proxy 74 passed / 6 skipped.
tts reports 17 passed / 1 failed (Fish Audio) here AND on release — pre-existing.

Corrected the proxy suite's timeout rationale. The comment claimed every case
there drives a locally spawned proxy and never a live provider. That is false:
nine cases gate on `hasValidCredentials()` and, when it is true, reach Anthropic
through the spawned proxy. Verified by counting the call sites rather than
taking the review comment's word for it.

180s stays — it is comfortably above a live round-trip and the proxy under test
is still one this repo builds — but the justification now says what is
actually true, including that a breach is usually a defect and occasionally a
slow upstream. A case that legitimately needs longer should carry its own bound
rather than have this one raised for everyone.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants