fix(antigravity): wrap Pro fallback chain in try/catch for timeout resilience - #7290
diegosouzapw merged 10 commits into
Conversation
When a Pro-tier candidate times out or throws a network error, the exception now continues to the next candidate instead of aborting the entire chain. Includes diagnostic logging and unit tests. Signed-off-by: Minxi Hou <houminxi@gmail.com>
Re-throw AbortError and signal.aborted immediately instead of retrying the next candidate. Prevents wasted upstream requests after client disconnect. Signed-off-by: Minxi Hou <houminxi@gmail.com>
There was a problem hiding this comment.
Code Review
This pull request implements exception handling and retry logic within the Pro fallback chain of the AntigravityExecutor, allowing the execution to fall back to subsequent candidates when an exception occurs, while immediately propagating abort signals. It also adds comprehensive unit tests to verify these behaviors. The review feedback suggests improving the abort signal detection to support standard Error objects with an "AbortError" name, and refining the final error handling when the chain is exhausted to return a prior 400 response if one exists, rather than throwing a generic error.
Important
The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.
| // Abort signal (user disconnect) — propagate immediately, do not retry. | ||
| if (input.signal?.aborted || (error instanceof DOMException && error.name === "AbortError")) { | ||
| throw error; | ||
| } |
There was a problem hiding this comment.
To ensure maximum robustness across different environments, shims, or test mocks (which might throw a standard Error with name: "AbortError" instead of a true DOMException), we should also check if error is an instance of Error with name === "AbortError".
const isAbort =
input.signal?.aborted ||
(error instanceof DOMException && error.name === "AbortError") ||
(error instanceof Error && error.name === "AbortError");
if (isAbort) {
throw error;
}| // Last candidate also threw -- wrap all-failed context. | ||
| throw new Error( | ||
| `Pro fallback chain exhausted (all ${chain.length} candidates threw). Last error: ${error instanceof Error ? error.message : String(error)}` | ||
| ); |
There was a problem hiding this comment.
There are two issues here:
- The error message states that
all ${chain.length} candidates threw, which is misleading if some candidates returned a400response instead of throwing an exception (e.g., in a mixed path scenario). - If the first candidate returned a
400response (stored infirstResult), returningfirstResultis much more consistent and robust than throwing a hard 500 error. This matches the behavior of the all-400 case below where we returnfirstResultto surface the original sanitized error to the client.
// Last candidate also threw -- wrap all-failed context or return original 400 if available.
if (firstResult) {
input.log?.warn?.(
'AG_PRO_FALLBACK',
'Pro fallback chain exhausted (last candidate threw, but first candidate returned 400) for "' + resolvedUpstreamId + '". Returning original 400.'
);
return firstResult;
}
throw new Error(
'Pro fallback chain exhausted (all ' + chain.length + ' candidates failed). Last error: ' + (error instanceof Error ? error.message : String(error))
);In containers (USER node, no sudo, not root) provisionDnsEntries() now detects the condition up-front and logs a clear message instead of attempting sudo and silently swallowing the error. Adds canElevate() to the injectable deps interface for testability, and supports SKIP_ANTIGRAVITY_DNS=true for explicit opt-out.
Check Error.name === 'AbortError' for non-DOMException environments (polyfills, test harnesses). Capture first 400 from any candidate (not just i===0) so mixed paths surface the 400 instead of a generic error. Return firstResult when last candidate throws, consistent with the all-400 case.
|
Thanks for this — the try/catch resilience fix for the Pro fallback chain is well-reasoned and I traced the call graph to confirm it only retries pre-response (before any bytes reach the client), so it doesn't violate our no-swallow-in-SSE-streams rule, and the "chain exhausted" error message flows through our existing sanitization path correctly. All 13 tests in agy-pro-fallback-chain-3786.test.ts pass locally (8 existing + 5 new across your commits). One gap before merge: the |
The SKIP_ANTIGRAVITY_DNS=true and canElevate()=false tests used empty agentStates/customHosts, so they could not distinguish 'all steps skipped' from 'only the default step skipped'. Provide non-empty mocks and assert addHostsDns was NOT called. Also add a SKIP_ANTIGRAVITY_DNS=false boundary test confirming the strict === "true" comparison does not block normal provisioning, and verify sudoPassword passthrough in the canElevate=true happy-path test.
…size cap) diegosouzapw#7408 added the non-streaming SSE pass-through (createCreditsExtractionTransform plus its two call sites: the credits-retry path and the main non-streaming path) inline in antigravity.ts, growing it to 1806 lines. Combined with two other authorized PRs touching the same file (diegosouzapw#6979 +11, diegosouzapw#7290 +30), the projected total exceeds the frozen file-size gate (1813). Extract the new streaming-passthrough logic verbatim into open-sse/executors/antigravity/streamingPassthrough.ts (createCreditsExtractionTransform + a new buildSsePassthroughResult that deduplicates the two near-identical call sites), following the existing sseCollect.ts submodule pattern -- pure, no host state, no fetch/auth. antigravity.ts keeps a thin wrapper for createCreditsExtractionTransform (same public signature the existing unit tests import) that injects updateAntigravityRemainingCredits so the two modules don't import each other. No behavior change: same abort handling, same 499-on-early-disconnect, same 16KB credits sliding-window cap. antigravity.ts: 1806 -> 1693 lines (under the 1755 pre-PR baseline, with margin). New module: 176 lines (cap 800). Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
provisionDnsEntries() (complexity ~18, this PR's try/catch/log additions pushed it over check-complexity.mjs's threshold of 15) and execute()'s Pro-fallback loop (complexity 27, from wrapping executeOnce() in try/catch for timeout resilience) were both over the gate. Decomposed each into small named helpers, no behavior change: - provision.ts: split into provisionDefaultDns/provisionAgentDns/ provisionCustomHostsDns, each wrapping one best-effort DNS step. - antigravity.ts: extracted the fallback-chain catch/400-handling decisions (handleAntigravityFallbackChainError, isAntigravityAbortError, handleAntigravityFallback400) into a new antigravity/proFallbackChain.ts submodule (pure, no executor instance state), mirroring the existing antigravity/sseCollect.ts submodule pattern. Also fixes the antigravity.ts file-size cap (was pushed to 1854 lines > 1813 frozen ceiling by this PR's own try/catch addition; now 1771). execute/provisionDnsEntries no longer appear with ruleId complexity or max-lines-per-function in the check-complexity.mjs report. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
…xity gate) Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
…gnitive gate) Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
|
Thanks @HouMinXi! Merged into release/v3.8.49 after merge-train validation on the combined tree. The fallback-chain try/catch was split into the proFallbackChain module to clear the complexity gates — authorship preserved. |
…ough Resolves conflict in open-sse/executors/antigravity.ts between this branch's streaming-passthrough decomposition and diegosouzapw#7290's fallback-chain decomposition (already merged into release/v3.8.49) — both sides added imports from the same new antigravity/ submodule files, kept both. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
) * chore(ci): add .mergify.yml to main — Mergify only reads config from the default branch (#7168) * fix(ci): add the auto-enqueue pull_request_rule to the Mergify config (queue_conditions alone are eligibility-only) (#7179) * fix(ci): migrate Mergify auto-enqueue to merge_protections_settings.auto_merge_conditions (rules-based path is EOL 2026-07-16) (#7216) * fix(ci): drop Mergify batch settings (batching is a paid-tier feature; free plan queue is serial) (#7220) * fix(ci): merge queue tolerates the advisory dast-smoke failure (its GH-hosted build hang dequeued every attempt) (#7225) * test(ci): make the #6634 selfref guard hermetic — main's copy hard-fails every PR (#7341) main's copy of this test still does git I/O inside a unit test: const baseSrc = git(['show', 'origin/main:' + FILE]); Runners check out a shallow single ref, so origin/main does not resolve and the test dies with 'fatal: invalid object name origin/main'. Every PR into main fails Unit Tests (7/8) on it — today that is #7313, #7315, #7316, #7334, #7336 and #7337, six PRs red on a defect none of them introduced. #7313 has no other red at all. release/v3.8.49 already carries a fix (2e42b8e, #7174: try/catch, fetch origin/main on demand, t.skip() when unreachable), but it only reaches main at release time — so main stays broken for the whole cycle. Cherry-picking it would also import a new problem: PR Test Policy classifies t.skip() as a silenced assertion, which we watched it correctly catch on #7300 today. This is the hermetic version instead (ported from #7327, which does the same for the release branch): read the file straight off disk, compare against an empty base so baseTaut/baseExtTaut are 0 — the strictest possible comparison point — and call evaluateMasking() directly. No git ref, no fetch, no skip, nothing the runner's checkout depth can break. The #6634 regression stays covered: the guard's logic lives in SELF_TEST_FIXTURE_RE (check-test-masking.mjs:337), not in the test. Proven both ways on main before committing — neutralise SELF_TEST_FIXTURE_RE to /$^/ and the test FAILS; restore it and it passes 2/2, with check-test-masking.mjs left byte-identical. Co-authored-by: growab <nekron@icloud.com> * chore(quality): tighten main's coverage baseline to the CI's real numbers (#7347) main's ratchet had been failing --require-tighten on every PR: 11 metrics improved but the baseline was never tightened. Same class as the #6634 selfref guard — an infra fix that lands only on the release branch leaves main red for the whole cycle, and every PR into main pays for it. Values are the merged-coverage numbers from a run on main itself (a local run measures ~68% vs CI's ~80%; the baseline's own note warns about that gap). Only the 11 coverage values change — gitleaks and semgrepFindings keep main's own state. No changelog fragment: #7326 carries it on release/v3.8.49, and a second one here would double the entry at release time. * fix(antigravity): remove hardcoded 120s SSE collect timeout The SSE collection in collectStreamToResponse had a hardcoded 120 s timeout. Reasoning-heavy models like gemini-3.1-pro-high on large prompts (>30 KB) regularly exceed 120 s of generation time, causing the executor to return a synthetic 504 before the model finishes. Replace the hardcoded value with FETCH_TIMEOUT_MS (default 600 s, overridable via FETCH_TIMEOUT_MS env var), which is the standard upstream-request budget across all OmniRoute providers. Signed-off-by: Minxi Hou <houminxi@gmail.com> * fix(antigravity): streaming passthrough for non-streaming clients When a client sends stream: false to the Antigravity executor (Gemini models), OmniRoute buffered the entire SSE stream before responding. Long-thinking models exceeded the 120s timeout. Remove hardcoded SSE_COLLECT_TIMEOUT_MS. Extract shared createCreditsExtractionTransform with 16KB buffer cap and abort handling for client disconnect. Add parseSSEToGeminiResponse for the non-streaming drain path. Fix hasGeminiTerminalFinishReason to check top-level candidates (no response wrapper). Add signal null guards for credits retry path. Return 499 on early abort instead of piping cancelled body. Also remove duplicate SKILLS_SANDBOX_RUNTIME from .env.example and clarify .artifacts/ vs _artifacts/ in .gitignore. Signed-off-by: Minxi Hou <houminxi@gmail.com> * refactor(antigravity): extract streaming passthrough to module (file-size cap) #7408 added the non-streaming SSE pass-through (createCreditsExtractionTransform plus its two call sites: the credits-retry path and the main non-streaming path) inline in antigravity.ts, growing it to 1806 lines. Combined with two other authorized PRs touching the same file (#6979 +11, #7290 +30), the projected total exceeds the frozen file-size gate (1813). Extract the new streaming-passthrough logic verbatim into open-sse/executors/antigravity/streamingPassthrough.ts (createCreditsExtractionTransform + a new buildSsePassthroughResult that deduplicates the two near-identical call sites), following the existing sseCollect.ts submodule pattern -- pure, no host state, no fetch/auth. antigravity.ts keeps a thin wrapper for createCreditsExtractionTransform (same public signature the existing unit tests import) that injects updateAntigravityRemainingCredits so the two modules don't import each other. No behavior change: same abort handling, same 499-on-early-disconnect, same 16KB credits sliding-window cap. antigravity.ts: 1806 -> 1693 lines (under the 1755 pre-PR baseline, with margin). New module: 176 lines (cap 800). Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * refactor(antigravity): split incremental parser + move new tests to own file (file-size caps) Two remaining frozen file-size violations from #7408, resolved by extraction/move with zero behavior or assert changes: - open-sse/handlers/sseParser.ts (979 > frozen 830): the PR appended parseSSEToGeminiResponse (+153, the Gemini buffered-SSE -> chat.completion parser). Moved verbatim to open-sse/handlers/sseParser/geminiResponse.ts, following the handlers submodule pattern (chatCore/, responseSanitizer/). sseParser.ts is now byte-identical to its pre-PR content (825 lines; PR delta 0). Importers (chatCore/nonStreamingSse.ts, tests) point at the new module. - tests/unit/executor-antigravity.test.ts (1058 > testFrozen 942): the PR's new streaming-passthrough tests moved verbatim (same tests, same asserts) to tests/unit/antigravity-streaming-passthrough.test.ts: the 3 createCreditsExtractionTransform tests plus the non-streaming passthrough drain test ("auto-retries short 429 ... collects SSE for non-stream clients"), which the PR rewired onto the new raw-SSE path. The frozen file drops to 888 lines (below its pre-PR 941). New files: geminiResponse.ts 156 lines, passthrough test 202 lines (caps 800). Also fixes the stale sseParser.ts path in collectStreamToResponse's deprecation note. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * refactor(antigravity): decompose execute + gemini parser below complexity gate executeOnce() (complexity 127, 436 lines) and parseSSEToGeminiResponse() (complexity 39, 117 lines) were both over the check-complexity.mjs gate (complexity>15, max-lines-per-function>80). Decomposed each into small named helpers, no behavior change: - geminiResponse.ts: split into pure per-concern functions (markdown shortcut, candidate-parts walk, finishReason, usageMetadata, final response assembly). - antigravity.ts: extracted the per-url-index attempt pipeline (runAntigravityAttempt, handleAntigravityRateLimit, tryResolveRetryFromErrorBody, shouldAutoRetryTransient) and moved the request/result-building helpers (send, credits-retry, embed-retry, non-streaming/streaming result builders) into a new antigravity/executeAttempt.ts submodule, mirroring the existing streamingPassthrough.ts/sseCollect.ts pattern. Also fixes the antigravity.ts file-size cap (was pushed to 2084 lines > 1813 frozen ceiling by the decomposition itself; now 1428). check-complexity.mjs: 2054 violations (baseline 2058) — net improvement. execute/executeOnce/parseSSEToGeminiResponse no longer appear with ruleId complexity or max-lines-per-function. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * Merge branch 'release/v3.8.49' into fix/antigravity-streaming-passthrough Resolves conflict in open-sse/executors/antigravity.ts between this branch's streaming-passthrough decomposition and #7290's fallback-chain decomposition (already merged into release/v3.8.49) — both sides added imports from the same new antigravity/ submodule files, kept both. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * fix(antigravity): keep buffered JSON contract for non-streaming callers #3786's Pro-family fallback-chain retry loop (execute()) calls executeOnce() per candidate and inspects result.response directly, expecting a synthesized chat.completion JSON body on success. The streaming-passthrough migration made ALL non-streaming (stream: false) responses a raw SSE pass-through instead, so a successful retry candidate's response.json() threw ("data: {...}" is not valid JSON) — breaking the fallback chain (tests/unit/agy-pro-fallback-chain-3786.test.ts, 3 of 13 red). Route non-streaming (stream: false) responses back through collectStreamToResponse (buffered collect-to-JSON), which already uses FETCH_TIMEOUT_MS with no hardcoded 120s ceiling, so long-thinking models are not penalized. Passthrough is reserved for actual streaming clients (stream: true), which was the PR's real target scenario. Extracted the branch into buildAntigravityAttemptResult() to keep runAntigravityAttempt under the 80-line ratchet cap. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Signed-off-by: Minxi Hou <houminxi@gmail.com> Co-authored-by: Diego Rodrigues de Sa e Souza <8016841+diegosouzapw@users.noreply.github.com> Co-authored-by: growab <nekron@icloud.com> Co-authored-by: HouMinXi <1000+HouMinXi@users.noreply.github.com> Co-authored-by: HouMinXi <19586012+HouMinXi@users.noreply.github.com>
…silience (diegosouzapw#7290) * fix(antigravity): wrap executeOnce in try/catch for Pro fallback chain When a Pro-tier candidate times out or throws a network error, the exception now continues to the next candidate instead of aborting the entire chain. Includes diagnostic logging and unit tests. Signed-off-by: Minxi Hou <houminxi@gmail.com> * fix(antigravity): propagate abort signal in Pro fallback catch block Re-throw AbortError and signal.aborted immediately instead of retrying the next candidate. Prevents wasted upstream requests after client disconnect. Signed-off-by: Minxi Hou <houminxi@gmail.com> * fix(mitm): skip DNS modification when sudo unavailable (container) In containers (USER node, no sudo, not root) provisionDnsEntries() now detects the condition up-front and logs a clear message instead of attempting sudo and silently swallowing the error. Adds canElevate() to the injectable deps interface for testability, and supports SKIP_ANTIGRAVITY_DNS=true for explicit opt-out. * fix(antigravity): improve abort detection and fallback error handling Check Error.name === 'AbortError' for non-DOMException environments (polyfills, test harnesses). Capture first 400 from any candidate (not just i===0) so mixed paths surface the 400 instead of a generic error. Return firstResult when last candidate throws, consistent with the all-400 case. * test(mitm): add coverage for container-skip DNS provisioning * test(mitm): harden container-skip DNS test assertions The SKIP_ANTIGRAVITY_DNS=true and canElevate()=false tests used empty agentStates/customHosts, so they could not distinguish 'all steps skipped' from 'only the default step skipped'. Provide non-empty mocks and assert addHostsDns was NOT called. Also add a SKIP_ANTIGRAVITY_DNS=false boundary test confirming the strict === "true" comparison does not block normal provisioning, and verify sudoPassword passthrough in the canElevate=true happy-path test. * refactor(mitm): split provisionDnsEntries below complexity gate provisionDnsEntries() (complexity ~18, this PR's try/catch/log additions pushed it over check-complexity.mjs's threshold of 15) and execute()'s Pro-fallback loop (complexity 27, from wrapping executeOnce() in try/catch for timeout resilience) were both over the gate. Decomposed each into small named helpers, no behavior change: - provision.ts: split into provisionDefaultDns/provisionAgentDns/ provisionCustomHostsDns, each wrapping one best-effort DNS step. - antigravity.ts: extracted the fallback-chain catch/400-handling decisions (handleAntigravityFallbackChainError, isAntigravityAbortError, handleAntigravityFallback400) into a new antigravity/proFallbackChain.ts submodule (pure, no executor instance state), mirroring the existing antigravity/sseCollect.ts submodule pattern. Also fixes the antigravity.ts file-size cap (was pushed to 1854 lines > 1813 frozen ceiling by this PR's own try/catch addition; now 1771). execute/provisionDnsEntries no longer appear with ruleId complexity or max-lines-per-function in the check-complexity.mjs report. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * refactor(antigravity): drop redundant loop continue (cognitive-complexity gate) Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * refactor(antigravity): fold fallback outcome dispatch into switch (cognitive gate) Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Signed-off-by: Minxi Hou <houminxi@gmail.com> Co-authored-by: Diego Rodrigues de Sa e Souza <diegosouza.pw@gmail.com> Co-authored-by: HouMinXi <19586012+HouMinXi@users.noreply.github.com> Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
…egosouzapw#7408) * chore(ci): add .mergify.yml to main — Mergify only reads config from the default branch (diegosouzapw#7168) * fix(ci): add the auto-enqueue pull_request_rule to the Mergify config (queue_conditions alone are eligibility-only) (diegosouzapw#7179) * fix(ci): migrate Mergify auto-enqueue to merge_protections_settings.auto_merge_conditions (rules-based path is EOL 2026-07-16) (diegosouzapw#7216) * fix(ci): drop Mergify batch settings (batching is a paid-tier feature; free plan queue is serial) (diegosouzapw#7220) * fix(ci): merge queue tolerates the advisory dast-smoke failure (its GH-hosted build hang dequeued every attempt) (diegosouzapw#7225) * test(ci): make the diegosouzapw#6634 selfref guard hermetic — main's copy hard-fails every PR (diegosouzapw#7341) main's copy of this test still does git I/O inside a unit test: const baseSrc = git(['show', 'origin/main:' + FILE]); Runners check out a shallow single ref, so origin/main does not resolve and the test dies with 'fatal: invalid object name origin/main'. Every PR into main fails Unit Tests (7/8) on it — today that is diegosouzapw#7313, diegosouzapw#7315, diegosouzapw#7316, diegosouzapw#7334, diegosouzapw#7336 and diegosouzapw#7337, six PRs red on a defect none of them introduced. diegosouzapw#7313 has no other red at all. release/v3.8.49 already carries a fix (8bbd411, diegosouzapw#7174: try/catch, fetch origin/main on demand, t.skip() when unreachable), but it only reaches main at release time — so main stays broken for the whole cycle. Cherry-picking it would also import a new problem: PR Test Policy classifies t.skip() as a silenced assertion, which we watched it correctly catch on diegosouzapw#7300 today. This is the hermetic version instead (ported from diegosouzapw#7327, which does the same for the release branch): read the file straight off disk, compare against an empty base so baseTaut/baseExtTaut are 0 — the strictest possible comparison point — and call evaluateMasking() directly. No git ref, no fetch, no skip, nothing the runner's checkout depth can break. The diegosouzapw#6634 regression stays covered: the guard's logic lives in SELF_TEST_FIXTURE_RE (check-test-masking.mjs:337), not in the test. Proven both ways on main before committing — neutralise SELF_TEST_FIXTURE_RE to /$^/ and the test FAILS; restore it and it passes 2/2, with check-test-masking.mjs left byte-identical. Co-authored-by: growab <nekron@icloud.com> * chore(quality): tighten main's coverage baseline to the CI's real numbers (diegosouzapw#7347) main's ratchet had been failing --require-tighten on every PR: 11 metrics improved but the baseline was never tightened. Same class as the diegosouzapw#6634 selfref guard — an infra fix that lands only on the release branch leaves main red for the whole cycle, and every PR into main pays for it. Values are the merged-coverage numbers from a run on main itself (a local run measures ~68% vs CI's ~80%; the baseline's own note warns about that gap). Only the 11 coverage values change — gitleaks and semgrepFindings keep main's own state. No changelog fragment: diegosouzapw#7326 carries it on release/v3.8.49, and a second one here would double the entry at release time. * fix(antigravity): remove hardcoded 120s SSE collect timeout The SSE collection in collectStreamToResponse had a hardcoded 120 s timeout. Reasoning-heavy models like gemini-3.1-pro-high on large prompts (>30 KB) regularly exceed 120 s of generation time, causing the executor to return a synthetic 504 before the model finishes. Replace the hardcoded value with FETCH_TIMEOUT_MS (default 600 s, overridable via FETCH_TIMEOUT_MS env var), which is the standard upstream-request budget across all OmniRoute providers. Signed-off-by: Minxi Hou <houminxi@gmail.com> * fix(antigravity): streaming passthrough for non-streaming clients When a client sends stream: false to the Antigravity executor (Gemini models), OmniRoute buffered the entire SSE stream before responding. Long-thinking models exceeded the 120s timeout. Remove hardcoded SSE_COLLECT_TIMEOUT_MS. Extract shared createCreditsExtractionTransform with 16KB buffer cap and abort handling for client disconnect. Add parseSSEToGeminiResponse for the non-streaming drain path. Fix hasGeminiTerminalFinishReason to check top-level candidates (no response wrapper). Add signal null guards for credits retry path. Return 499 on early abort instead of piping cancelled body. Also remove duplicate SKILLS_SANDBOX_RUNTIME from .env.example and clarify .artifacts/ vs _artifacts/ in .gitignore. Signed-off-by: Minxi Hou <houminxi@gmail.com> * refactor(antigravity): extract streaming passthrough to module (file-size cap) diegosouzapw#7408 added the non-streaming SSE pass-through (createCreditsExtractionTransform plus its two call sites: the credits-retry path and the main non-streaming path) inline in antigravity.ts, growing it to 1806 lines. Combined with two other authorized PRs touching the same file (diegosouzapw#6979 +11, diegosouzapw#7290 +30), the projected total exceeds the frozen file-size gate (1813). Extract the new streaming-passthrough logic verbatim into open-sse/executors/antigravity/streamingPassthrough.ts (createCreditsExtractionTransform + a new buildSsePassthroughResult that deduplicates the two near-identical call sites), following the existing sseCollect.ts submodule pattern -- pure, no host state, no fetch/auth. antigravity.ts keeps a thin wrapper for createCreditsExtractionTransform (same public signature the existing unit tests import) that injects updateAntigravityRemainingCredits so the two modules don't import each other. No behavior change: same abort handling, same 499-on-early-disconnect, same 16KB credits sliding-window cap. antigravity.ts: 1806 -> 1693 lines (under the 1755 pre-PR baseline, with margin). New module: 176 lines (cap 800). Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * refactor(antigravity): split incremental parser + move new tests to own file (file-size caps) Two remaining frozen file-size violations from diegosouzapw#7408, resolved by extraction/move with zero behavior or assert changes: - open-sse/handlers/sseParser.ts (979 > frozen 830): the PR appended parseSSEToGeminiResponse (+153, the Gemini buffered-SSE -> chat.completion parser). Moved verbatim to open-sse/handlers/sseParser/geminiResponse.ts, following the handlers submodule pattern (chatCore/, responseSanitizer/). sseParser.ts is now byte-identical to its pre-PR content (825 lines; PR delta 0). Importers (chatCore/nonStreamingSse.ts, tests) point at the new module. - tests/unit/executor-antigravity.test.ts (1058 > testFrozen 942): the PR's new streaming-passthrough tests moved verbatim (same tests, same asserts) to tests/unit/antigravity-streaming-passthrough.test.ts: the 3 createCreditsExtractionTransform tests plus the non-streaming passthrough drain test ("auto-retries short 429 ... collects SSE for non-stream clients"), which the PR rewired onto the new raw-SSE path. The frozen file drops to 888 lines (below its pre-PR 941). New files: geminiResponse.ts 156 lines, passthrough test 202 lines (caps 800). Also fixes the stale sseParser.ts path in collectStreamToResponse's deprecation note. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * refactor(antigravity): decompose execute + gemini parser below complexity gate executeOnce() (complexity 127, 436 lines) and parseSSEToGeminiResponse() (complexity 39, 117 lines) were both over the check-complexity.mjs gate (complexity>15, max-lines-per-function>80). Decomposed each into small named helpers, no behavior change: - geminiResponse.ts: split into pure per-concern functions (markdown shortcut, candidate-parts walk, finishReason, usageMetadata, final response assembly). - antigravity.ts: extracted the per-url-index attempt pipeline (runAntigravityAttempt, handleAntigravityRateLimit, tryResolveRetryFromErrorBody, shouldAutoRetryTransient) and moved the request/result-building helpers (send, credits-retry, embed-retry, non-streaming/streaming result builders) into a new antigravity/executeAttempt.ts submodule, mirroring the existing streamingPassthrough.ts/sseCollect.ts pattern. Also fixes the antigravity.ts file-size cap (was pushed to 2084 lines > 1813 frozen ceiling by the decomposition itself; now 1428). check-complexity.mjs: 2054 violations (baseline 2058) — net improvement. execute/executeOnce/parseSSEToGeminiResponse no longer appear with ruleId complexity or max-lines-per-function. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * Merge branch 'release/v3.8.49' into fix/antigravity-streaming-passthrough Resolves conflict in open-sse/executors/antigravity.ts between this branch's streaming-passthrough decomposition and diegosouzapw#7290's fallback-chain decomposition (already merged into release/v3.8.49) — both sides added imports from the same new antigravity/ submodule files, kept both. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * fix(antigravity): keep buffered JSON contract for non-streaming callers diegosouzapw#3786's Pro-family fallback-chain retry loop (execute()) calls executeOnce() per candidate and inspects result.response directly, expecting a synthesized chat.completion JSON body on success. The streaming-passthrough migration made ALL non-streaming (stream: false) responses a raw SSE pass-through instead, so a successful retry candidate's response.json() threw ("data: {...}" is not valid JSON) — breaking the fallback chain (tests/unit/agy-pro-fallback-chain-3786.test.ts, 3 of 13 red). Route non-streaming (stream: false) responses back through collectStreamToResponse (buffered collect-to-JSON), which already uses FETCH_TIMEOUT_MS with no hardcoded 120s ceiling, so long-thinking models are not penalized. Passthrough is reserved for actual streaming clients (stream: true), which was the PR's real target scenario. Extracted the branch into buildAntigravityAttemptResult() to keep runAntigravityAttempt under the 80-line ratchet cap. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Signed-off-by: Minxi Hou <houminxi@gmail.com> Co-authored-by: Diego Rodrigues de Sa e Souza <8016841+diegosouzapw@users.noreply.github.com> Co-authored-by: growab <nekron@icloud.com> Co-authored-by: HouMinXi <1000+HouMinXi@users.noreply.github.com> Co-authored-by: HouMinXi <19586012+HouMinXi@users.noreply.github.com>
…silience (diegosouzapw#7290) * fix(antigravity): wrap executeOnce in try/catch for Pro fallback chain When a Pro-tier candidate times out or throws a network error, the exception now continues to the next candidate instead of aborting the entire chain. Includes diagnostic logging and unit tests. Signed-off-by: Minxi Hou <houminxi@gmail.com> * fix(antigravity): propagate abort signal in Pro fallback catch block Re-throw AbortError and signal.aborted immediately instead of retrying the next candidate. Prevents wasted upstream requests after client disconnect. Signed-off-by: Minxi Hou <houminxi@gmail.com> * fix(mitm): skip DNS modification when sudo unavailable (container) In containers (USER node, no sudo, not root) provisionDnsEntries() now detects the condition up-front and logs a clear message instead of attempting sudo and silently swallowing the error. Adds canElevate() to the injectable deps interface for testability, and supports SKIP_ANTIGRAVITY_DNS=true for explicit opt-out. * fix(antigravity): improve abort detection and fallback error handling Check Error.name === 'AbortError' for non-DOMException environments (polyfills, test harnesses). Capture first 400 from any candidate (not just i===0) so mixed paths surface the 400 instead of a generic error. Return firstResult when last candidate throws, consistent with the all-400 case. * test(mitm): add coverage for container-skip DNS provisioning * test(mitm): harden container-skip DNS test assertions The SKIP_ANTIGRAVITY_DNS=true and canElevate()=false tests used empty agentStates/customHosts, so they could not distinguish 'all steps skipped' from 'only the default step skipped'. Provide non-empty mocks and assert addHostsDns was NOT called. Also add a SKIP_ANTIGRAVITY_DNS=false boundary test confirming the strict === "true" comparison does not block normal provisioning, and verify sudoPassword passthrough in the canElevate=true happy-path test. * refactor(mitm): split provisionDnsEntries below complexity gate provisionDnsEntries() (complexity ~18, this PR's try/catch/log additions pushed it over check-complexity.mjs's threshold of 15) and execute()'s Pro-fallback loop (complexity 27, from wrapping executeOnce() in try/catch for timeout resilience) were both over the gate. Decomposed each into small named helpers, no behavior change: - provision.ts: split into provisionDefaultDns/provisionAgentDns/ provisionCustomHostsDns, each wrapping one best-effort DNS step. - antigravity.ts: extracted the fallback-chain catch/400-handling decisions (handleAntigravityFallbackChainError, isAntigravityAbortError, handleAntigravityFallback400) into a new antigravity/proFallbackChain.ts submodule (pure, no executor instance state), mirroring the existing antigravity/sseCollect.ts submodule pattern. Also fixes the antigravity.ts file-size cap (was pushed to 1854 lines > 1813 frozen ceiling by this PR's own try/catch addition; now 1771). execute/provisionDnsEntries no longer appear with ruleId complexity or max-lines-per-function in the check-complexity.mjs report. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * refactor(antigravity): drop redundant loop continue (cognitive-complexity gate) Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * refactor(antigravity): fold fallback outcome dispatch into switch (cognitive gate) Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Signed-off-by: Minxi Hou <houminxi@gmail.com> Co-authored-by: Diego Rodrigues de Sa e Souza <diegosouza.pw@gmail.com> Co-authored-by: HouMinXi <19586012+HouMinXi@users.noreply.github.com> Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
…egosouzapw#7408) * chore(ci): add .mergify.yml to main — Mergify only reads config from the default branch (diegosouzapw#7168) * fix(ci): add the auto-enqueue pull_request_rule to the Mergify config (queue_conditions alone are eligibility-only) (diegosouzapw#7179) * fix(ci): migrate Mergify auto-enqueue to merge_protections_settings.auto_merge_conditions (rules-based path is EOL 2026-07-16) (diegosouzapw#7216) * fix(ci): drop Mergify batch settings (batching is a paid-tier feature; free plan queue is serial) (diegosouzapw#7220) * fix(ci): merge queue tolerates the advisory dast-smoke failure (its GH-hosted build hang dequeued every attempt) (diegosouzapw#7225) * test(ci): make the diegosouzapw#6634 selfref guard hermetic — main's copy hard-fails every PR (diegosouzapw#7341) main's copy of this test still does git I/O inside a unit test: const baseSrc = git(['show', 'origin/main:' + FILE]); Runners check out a shallow single ref, so origin/main does not resolve and the test dies with 'fatal: invalid object name origin/main'. Every PR into main fails Unit Tests (7/8) on it — today that is diegosouzapw#7313, diegosouzapw#7315, diegosouzapw#7316, diegosouzapw#7334, diegosouzapw#7336 and diegosouzapw#7337, six PRs red on a defect none of them introduced. diegosouzapw#7313 has no other red at all. release/v3.8.49 already carries a fix (83a7551, diegosouzapw#7174: try/catch, fetch origin/main on demand, t.skip() when unreachable), but it only reaches main at release time — so main stays broken for the whole cycle. Cherry-picking it would also import a new problem: PR Test Policy classifies t.skip() as a silenced assertion, which we watched it correctly catch on diegosouzapw#7300 today. This is the hermetic version instead (ported from diegosouzapw#7327, which does the same for the release branch): read the file straight off disk, compare against an empty base so baseTaut/baseExtTaut are 0 — the strictest possible comparison point — and call evaluateMasking() directly. No git ref, no fetch, no skip, nothing the runner's checkout depth can break. The diegosouzapw#6634 regression stays covered: the guard's logic lives in SELF_TEST_FIXTURE_RE (check-test-masking.mjs:337), not in the test. Proven both ways on main before committing — neutralise SELF_TEST_FIXTURE_RE to /$^/ and the test FAILS; restore it and it passes 2/2, with check-test-masking.mjs left byte-identical. Co-authored-by: growab <nekron@icloud.com> * chore(quality): tighten main's coverage baseline to the CI's real numbers (diegosouzapw#7347) main's ratchet had been failing --require-tighten on every PR: 11 metrics improved but the baseline was never tightened. Same class as the diegosouzapw#6634 selfref guard — an infra fix that lands only on the release branch leaves main red for the whole cycle, and every PR into main pays for it. Values are the merged-coverage numbers from a run on main itself (a local run measures ~68% vs CI's ~80%; the baseline's own note warns about that gap). Only the 11 coverage values change — gitleaks and semgrepFindings keep main's own state. No changelog fragment: diegosouzapw#7326 carries it on release/v3.8.49, and a second one here would double the entry at release time. * fix(antigravity): remove hardcoded 120s SSE collect timeout The SSE collection in collectStreamToResponse had a hardcoded 120 s timeout. Reasoning-heavy models like gemini-3.1-pro-high on large prompts (>30 KB) regularly exceed 120 s of generation time, causing the executor to return a synthetic 504 before the model finishes. Replace the hardcoded value with FETCH_TIMEOUT_MS (default 600 s, overridable via FETCH_TIMEOUT_MS env var), which is the standard upstream-request budget across all OmniRoute providers. Signed-off-by: Minxi Hou <houminxi@gmail.com> * fix(antigravity): streaming passthrough for non-streaming clients When a client sends stream: false to the Antigravity executor (Gemini models), OmniRoute buffered the entire SSE stream before responding. Long-thinking models exceeded the 120s timeout. Remove hardcoded SSE_COLLECT_TIMEOUT_MS. Extract shared createCreditsExtractionTransform with 16KB buffer cap and abort handling for client disconnect. Add parseSSEToGeminiResponse for the non-streaming drain path. Fix hasGeminiTerminalFinishReason to check top-level candidates (no response wrapper). Add signal null guards for credits retry path. Return 499 on early abort instead of piping cancelled body. Also remove duplicate SKILLS_SANDBOX_RUNTIME from .env.example and clarify .artifacts/ vs _artifacts/ in .gitignore. Signed-off-by: Minxi Hou <houminxi@gmail.com> * refactor(antigravity): extract streaming passthrough to module (file-size cap) diegosouzapw#7408 added the non-streaming SSE pass-through (createCreditsExtractionTransform plus its two call sites: the credits-retry path and the main non-streaming path) inline in antigravity.ts, growing it to 1806 lines. Combined with two other authorized PRs touching the same file (diegosouzapw#6979 +11, diegosouzapw#7290 +30), the projected total exceeds the frozen file-size gate (1813). Extract the new streaming-passthrough logic verbatim into open-sse/executors/antigravity/streamingPassthrough.ts (createCreditsExtractionTransform + a new buildSsePassthroughResult that deduplicates the two near-identical call sites), following the existing sseCollect.ts submodule pattern -- pure, no host state, no fetch/auth. antigravity.ts keeps a thin wrapper for createCreditsExtractionTransform (same public signature the existing unit tests import) that injects updateAntigravityRemainingCredits so the two modules don't import each other. No behavior change: same abort handling, same 499-on-early-disconnect, same 16KB credits sliding-window cap. antigravity.ts: 1806 -> 1693 lines (under the 1755 pre-PR baseline, with margin). New module: 176 lines (cap 800). Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * refactor(antigravity): split incremental parser + move new tests to own file (file-size caps) Two remaining frozen file-size violations from diegosouzapw#7408, resolved by extraction/move with zero behavior or assert changes: - open-sse/handlers/sseParser.ts (979 > frozen 830): the PR appended parseSSEToGeminiResponse (+153, the Gemini buffered-SSE -> chat.completion parser). Moved verbatim to open-sse/handlers/sseParser/geminiResponse.ts, following the handlers submodule pattern (chatCore/, responseSanitizer/). sseParser.ts is now byte-identical to its pre-PR content (825 lines; PR delta 0). Importers (chatCore/nonStreamingSse.ts, tests) point at the new module. - tests/unit/executor-antigravity.test.ts (1058 > testFrozen 942): the PR's new streaming-passthrough tests moved verbatim (same tests, same asserts) to tests/unit/antigravity-streaming-passthrough.test.ts: the 3 createCreditsExtractionTransform tests plus the non-streaming passthrough drain test ("auto-retries short 429 ... collects SSE for non-stream clients"), which the PR rewired onto the new raw-SSE path. The frozen file drops to 888 lines (below its pre-PR 941). New files: geminiResponse.ts 156 lines, passthrough test 202 lines (caps 800). Also fixes the stale sseParser.ts path in collectStreamToResponse's deprecation note. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * refactor(antigravity): decompose execute + gemini parser below complexity gate executeOnce() (complexity 127, 436 lines) and parseSSEToGeminiResponse() (complexity 39, 117 lines) were both over the check-complexity.mjs gate (complexity>15, max-lines-per-function>80). Decomposed each into small named helpers, no behavior change: - geminiResponse.ts: split into pure per-concern functions (markdown shortcut, candidate-parts walk, finishReason, usageMetadata, final response assembly). - antigravity.ts: extracted the per-url-index attempt pipeline (runAntigravityAttempt, handleAntigravityRateLimit, tryResolveRetryFromErrorBody, shouldAutoRetryTransient) and moved the request/result-building helpers (send, credits-retry, embed-retry, non-streaming/streaming result builders) into a new antigravity/executeAttempt.ts submodule, mirroring the existing streamingPassthrough.ts/sseCollect.ts pattern. Also fixes the antigravity.ts file-size cap (was pushed to 2084 lines > 1813 frozen ceiling by the decomposition itself; now 1428). check-complexity.mjs: 2054 violations (baseline 2058) — net improvement. execute/executeOnce/parseSSEToGeminiResponse no longer appear with ruleId complexity or max-lines-per-function. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * Merge branch 'release/v3.8.49' into fix/antigravity-streaming-passthrough Resolves conflict in open-sse/executors/antigravity.ts between this branch's streaming-passthrough decomposition and diegosouzapw#7290's fallback-chain decomposition (already merged into release/v3.8.49) — both sides added imports from the same new antigravity/ submodule files, kept both. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * fix(antigravity): keep buffered JSON contract for non-streaming callers diegosouzapw#3786's Pro-family fallback-chain retry loop (execute()) calls executeOnce() per candidate and inspects result.response directly, expecting a synthesized chat.completion JSON body on success. The streaming-passthrough migration made ALL non-streaming (stream: false) responses a raw SSE pass-through instead, so a successful retry candidate's response.json() threw ("data: {...}" is not valid JSON) — breaking the fallback chain (tests/unit/agy-pro-fallback-chain-3786.test.ts, 3 of 13 red). Route non-streaming (stream: false) responses back through collectStreamToResponse (buffered collect-to-JSON), which already uses FETCH_TIMEOUT_MS with no hardcoded 120s ceiling, so long-thinking models are not penalized. Passthrough is reserved for actual streaming clients (stream: true), which was the PR's real target scenario. Extracted the branch into buildAntigravityAttemptResult() to keep runAntigravityAttempt under the 80-line ratchet cap. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Signed-off-by: Minxi Hou <houminxi@gmail.com> Co-authored-by: Diego Rodrigues de Sa e Souza <8016841+diegosouzapw@users.noreply.github.com> Co-authored-by: growab <nekron@icloud.com> Co-authored-by: HouMinXi <1000+HouMinXi@users.noreply.github.com> Co-authored-by: HouMinXi <19586012+HouMinXi@users.noreply.github.com>
Problem
When
gemini-3.1-pro-highreturns HTTP 400, the Pro fallback chain should retry with alternative model IDs (gemini-pro-agent,gemini-3-pro-high). However, if a subsequent candidate hangs and times out, the exception escapes the fallback loop and the remaining candidates are never tried.Reproduction: Request
gemini-3.1-pro-high→ 400 → retrygemini-pro-agent→ timeout → exception escapes →gemini-3-pro-highnever attempted. User sees 120s timeout instead of a response.Fix
Wrap
executeOnce()in try/catch within the Pro fallback loop (execute()method):Tests
3 new test cases in
agy-pro-fallback-chain-3786.test.ts:All 11 tests pass (8 existing + 3 new).
Verification
gemini-3.1-pro-highnow completes in ~3s instead of hanging 120s