Skip to content

fix(executors): repair DuckDuckGo AI Chat challenge solver (418 ERR_CHALLENGE) - #9733

Merged
diegosouzapw merged 1 commit into
diegosouzapw:release/v3.8.50from
Mynacol:fix/duckduckgo-challenge-solver
Aug 11, 2026
Merged

diegosouzapw merged 1 commit into
diegosouzapw:release/v3.8.50from
Mynacol:fix/duckduckgo-challenge-solver

Conversation

@Mynacol

@Mynacol Mynacol commented Aug 7, 2026

Copy link
Copy Markdown
Contributor

Summary

  • Fixes the duckduckgo-web provider, which returned HTTP 418 ERR_CHALLENGE on every chat request while duck.ai worked normally in a browser from the same IP. The provider was fully unusable.
  • Ground truth was established by driving a real headful Chromium at duck.ai from the same IP, which returned 200. That ruled out IP blocking / rate limiting and proved the fault was in our anti-abuse challenge solver.
  • Six independent defects were found and fixed. The first alone disabled the solver outright; the rest were exposed only once it started running again.
  • Also removes a redundant "seed" chat POST that doubled request volume against an IP-rate-limited endpoint and surfaced as spurious 429 ERR_RATE_LIMIT.

Root causes

# Defect Effect
1 export keywords inside CHALLENGE_STUBS. That string is run by vm.runInContext, which compiles in script mode, so module syntax is a hard SyntaxError. A refactor mass-added export to its 5 function declarations. Solver threw on every call → raw unsolved challenge sent → 418
2 \\s inside a String.raw template reached the sandbox as a literal backslash getComputedStyle display probe silently read empty
3 buildHtmlLookup subtracted 1 from the descendant count Chromium reports 3 for <li><div></li><li></div; we sent 2
4 Sandbox failed 9 of 13 browser-fidelity probes (prototype chains, NodeList identity, live body.children, native-code toString, sloppy-mode this) Challenge classified us as a bot
5 Solved payload omitted meta.origin / meta.stack / meta.duration, which the real duck.ai bundle always sends 418 even when every client_hash was correct
6 reasoningEffort is now mandatory on duckchat/v1/chat 400 ERR_BAD_REQUEST

Note on #4: Math must not be sealed — real Chromium reports Object.isSealed(Math) === false, and sealing it made our vector differ by one.

Related Issues

The regression was introduced by the challenge-solver extraction in #5925 (Release v3.8.44), which added export to the functions inside CHALLENGE_STUBS while moving them.

Validation

  • Change type: provider
  • Focused tests and category gates from the golden path
  • npm run lint
  • Reconciled with the current active release base; focused checks rerun afterward
  • Production-code changes include a new or updated automated test in this PR
  • SonarQube PR analysis is green or any remaining issues are explicitly documented below — not runnable locally; deferred to CI

Commands run (Node 22.23.2 provisioned via nix, since the sandbox had no system Node):

# Provider golden-path gates
npm run check:provider-consistency   # OK — 225 REGISTRY entries, 304 canonical providers
npm run check:provider-assets        # Provider asset budget passed

# Focused DuckDuckGo + adjacent suites — 177/177 pass
node --import tsx/esm --test \
  tests/unit/duckduckgo-*.test.ts \
  tests/unit/ddg-circuit-breaker-null-content-6999-7000.test.ts \
  tests/unit/executor-web-cookie-sweep.test.ts \
  tests/unit/search-handler-duckduckgo.test.ts \
  tests/unit/free-web-search.test.ts \
  tests/unit/session-pool-modular.test.ts \
  tests/unit/noauth-provider-validation.test.ts

# Pre-commit gates (hooks were skipped locally: npx not on PATH)
npm run lint
npm run check:any-budget:t11         # PASS
npm run check:tracked-artifacts      # OK
npm run check:docs-sync              # PASS
npx prettier --check <changed files> # All matched files use Prettier code style

Live end-to-end validation — the executor now returns 200 against real DuckDuckGo, 4/4 cases:

Case Result
non-streaming, gpt-4o-mini (alias path) 200 → "OK"
streaming, gpt-5.4-mini 200 → "OK"
claude-haiku-4-5 (reasoningEffort: low) 200 → "OK"
math prompt, gpt-5.4-nano 200 → "42"

This satisfies the AGENTS.md hard-rule #18 bug-fix gate via both paths: automated regression tests (TDD) and a documented live real-environment test.

Tests Added Or Updated

  • tests/unit/duckduckgo-challenge-solver-regression.test.ts — new, 32 tests
  • tests/unit/duckduckgo-reasoning-effort-required.test.ts — new, 5 tests
  • tests/unit/duckduckgo-challenge-split.test.ts — updated, +4 tests guarding the script-mode/export invariant
  • tests/fixtures/duckduckgo/challenge-variants.json — new fixture: 8 real challenge programs captured from duckduckgo.com, each paired with the probe vectors a real headful Chromium produced for that exact program

The suite asserts against recorded real-browser behaviour, not against our own output — matching Chromium bit-for-bit is the actual correctness criterion here. Every fix was individually reverted to confirm its test fails:

Reverted fix Failing tests
re-add export to stubs 30
re-break \\s escaping 3
restore count - 1 4
drop meta augmentation 2
seal Math 4
reasoningEffort: null 2

Coverage Notes

This PR changes open-sse/executors/duckduckgo-web.ts and open-sse/executors/duckduckgo-web/challenge.ts.

  • solveDuckDuckGoChallenge, CHALLENGE_STUBS, and buildHtmlLookup are covered by the 8 real-variant fixtures plus 16 individually-named browser-fidelity probe tests, so a future stub regression names itself instead of failing as one opaque hash mismatch.
  • The reasoningEffort change and the seed-POST removal are covered by duckduckgo-reasoning-effort-required.test.ts, which stubs fetch and asserts on the actual outgoing wire payload (including a test that exactly one POST /duckchat/v1/chat is issued per user request).
  • Coverage moves up on both touched files: previously the solver had no test that executed the sandbox at all — which is precisely why a SyntaxError inside CHALLENGE_STUBS shipped unnoticed.

Reviewer Notes

Base-red inherited — not caused by this PR. Three failures pre-exist on origin/release/v3.8.50 and reproduce identically on a pristine checkout of that tip:

  1. tests/unit/models-catalog-route.test.ts — ERR_MODULE_NOT_FOUND: open-sse/services/antigravityProjectPersistence.ts. The file is imported by open-sse/services/combo/quotaStrategies.ts but does not exist in git at that ref.
  2. tests/unit/provider-translate-path-golden.test.ts — golden drift from devin-cli-agentic and raycast missing from the snapshot. Verified the duckduckgo-web entry is byte-identical and untouched by this PR.
  3. npm run lint — 3 no-explicit-any errors in tests/unit/vertex-functioncall-id-3440.test.ts, a file this PR does not touch.

Per CONTRIBUTING.md these are not fixed here; a base-red fix belongs in its own fix/release-v3.8.50-basereds-* PR.

Risk areas

  • vm sandbox is a supply-chain surface. solveDuckDuckGoChallenge executes upstream-supplied JS. The security posture is unchanged — still vm.runInContext with the 5s timeout — and the SECURITY NOTE remains. This PR only makes the emulated browser environment more faithful; it does not widen what the sandbox can reach.
  • The stubs are anti-bot-sensitive. DuckDuckGo rotates challenge variants (8 distinct programs observed in one session). If upstream adds a new probe this can regress again. The fixtures make that failure mode legible: a new variant will show as a specific probe mismatch against recorded browser behaviour.
  • Two comment-formatting traps inside CHALLENGE_STUBS, both hit during development and now documented inline: (a) a backtick in a comment terminates the String.raw template; (b) \\x in a comment is still a double escape. A WARNING block at the top of the literal spells out the script-mode constraint.
  • Behaviour change: every request now sends reasoningEffort ("none" by default, "low" for claude-haiku-4-5 / tinfoil/gpt-oss-120b). Verified live as strictly required — the previous "let the server default it" path returns 400.

No migrations, no feature flags, no DB changes.

Manual validation reviewers may want to repeat: the live 4/4 table above. Note DuckDuckGo aggressively IP-rate-limits (429 ERR_RATE_LIMIT) under repeated testing; that is environmental and distinct from the 418 this PR fixes.

…HALLENGE)

Every duckduckgo-web chat request failed with HTTP 418 ERR_CHALLENGE while
duck.ai worked normally in a browser from the same IP. Ground truth was
established by driving a real headful Chromium at duck.ai from that IP (it
returned 200), so the environment was never the problem — the anti-abuse
challenge solver was. Six independent defects were found; the first alone
disabled the solver completely.

1. Module syntax inside the vm sandbox source.
   CHALLENGE_STUBS is executed with vm.runInContext, which compiles in SCRIPT
   mode. A refactor mass-added `export` to the five `function` declarations
   inside that template literal (they read as ordinary top-level TS functions),
   so every solve threw SyntaxError. The executor swallows solve failures and
   posts the raw unsolved challenge, which upstream answers with 418.

2. Double-escaped regex in a String.raw template.
   `\\s` in __parseCssDisplay reached the sandbox as a literal backslash, so the
   display regex never matched and a getComputedStyle probe silently read empty.

3. buildHtmlLookup undercounted descendants by one.
   `count` backs el.querySelectorAll('*').length; that returns DESCENDANTS and
   countHtmlElements already skips the #document-fragment root, so the `- 1` was
   wrong. Chromium reports 3 for '<li><div></li><li></div'; we reported 2, and a
   variant multiplies innerHTML.length by that count.

4. Browser-fidelity probes.
   Newer challenge variants assert JS/DOM invariants a flat stub cannot satisfy:
   real prototype chains (HTMLDivElement -> HTMLElement -> Element), NodeList
   identity, a live body.children HTMLCollection, native-code toString, and
   sloppy-mode `this === window`. Nine of thirteen failed. Notably Math must NOT
   be sealed — Chromium reports Object.isSealed(Math) === false, and sealing it
   made our vector differ by one.

5. The solved payload dropped meta.origin / meta.stack / meta.duration.
   The duck.ai bundle always sends all three; captured browser requests confirm
   it. Without them upstream returns 418 even when every client_hash is correct.

6. reasoningEffort is now mandatory on duckchat/v1/chat.
   An otherwise byte-identical payload returns 200 with the field and 400
   ERR_BAD_REQUEST without it (A/B verified live, repeated).

Also removes the throwaway "seed" chat POST that ran before every real request.
It existed to coax a usable challenge out of the upstream while the solver was
broken; it only doubled chat calls against an IP-rate-limited endpoint, showing
up as spurious 429 ERR_RATE_LIMIT.

Verification: the solver now reproduces real Chromium's probe vectors exactly
for all 8 captured challenge variants, and the executor returns 200 end-to-end
live (non-streaming, streaming, claude-haiku-4-5, and a math prompt returning
"42").

Tests: tests/unit/duckduckgo-challenge-solver-regression.test.ts (32 tests) and
tests/unit/duckduckgo-reasoning-effort-required.test.ts (5 tests), backed by
tests/fixtures/duckduckgo/challenge-variants.json — real captured challenge
programs plus the probe vectors a real browser produced for them, so the suite
asserts against recorded browser behaviour rather than our own output. Each fix
was confirmed to fail its test when individually reverted.
@Mynacol

Mynacol commented Aug 7, 2026

Copy link
Copy Markdown
Contributor Author

I manually tested the DuckDuckGo provider on a fully running omniroute server.

@Mynacol
Mynacol marked this pull request as ready for review August 7, 2026 21:27
@Mynacol
Mynacol requested a review from diegosouzapw as a code owner August 7, 2026 21:28
@diegosouzapw

Copy link
Copy Markdown
Owner

Thanks for the PR. Please address the mandatory items (tests and/or merge blockers) in this branch, then rerun checks before /merge-prs.

2 similar comments
@diegosouzapw

Copy link
Copy Markdown
Owner

Thanks for the PR. Please address the mandatory items (tests and/or merge blockers) in this branch, then rerun checks before /merge-prs.

@diegosouzapw

Copy link
Copy Markdown
Owner

Thanks for the PR. Please address the mandatory items (tests and/or merge blockers) in this branch, then rerun checks before /merge-prs.

@diegosouzapw

Copy link
Copy Markdown
Owner

Obrigado pelo PR. Mantive a revisão de fix-in-place e não foi possível concluir o ajuste completo aqui:

  • Para os PRs em fork: não consigo aplicar push de correção diretamente na sua branch.
    Por favor, faça um rebase/sync com release/v3.8.50, resolva conflitos se houver, e rode os checks dessa branch.
    Se preferir, posso aplicar a correção na próxima rodada assim que você mandar o branch atualizado ou confirmar que o PR está limpo pra esse merge.

@diegosouzapw
diegosouzapw merged commit acae259 into diegosouzapw:release/v3.8.50 Aug 11, 2026
3 checks passed
@Mynacol
Mynacol deleted the fix/duckduckgo-challenge-solver branch August 15, 2026 16:10
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…HALLENGE) (diegosouzapw#9733)

Every duckduckgo-web chat request failed with HTTP 418 ERR_CHALLENGE while
duck.ai worked normally in a browser from the same IP. Ground truth was
established by driving a real headful Chromium at duck.ai from that IP (it
returned 200), so the environment was never the problem — the anti-abuse
challenge solver was. Six independent defects were found; the first alone
disabled the solver completely.

1. Module syntax inside the vm sandbox source.
   CHALLENGE_STUBS is executed with vm.runInContext, which compiles in SCRIPT
   mode. A refactor mass-added `export` to the five `function` declarations
   inside that template literal (they read as ordinary top-level TS functions),
   so every solve threw SyntaxError. The executor swallows solve failures and
   posts the raw unsolved challenge, which upstream answers with 418.

2. Double-escaped regex in a String.raw template.
   `\\s` in __parseCssDisplay reached the sandbox as a literal backslash, so the
   display regex never matched and a getComputedStyle probe silently read empty.

3. buildHtmlLookup undercounted descendants by one.
   `count` backs el.querySelectorAll('*').length; that returns DESCENDANTS and
   countHtmlElements already skips the #document-fragment root, so the `- 1` was
   wrong. Chromium reports 3 for '<li><div></li><li></div'; we reported 2, and a
   variant multiplies innerHTML.length by that count.

4. Browser-fidelity probes.
   Newer challenge variants assert JS/DOM invariants a flat stub cannot satisfy:
   real prototype chains (HTMLDivElement -> HTMLElement -> Element), NodeList
   identity, a live body.children HTMLCollection, native-code toString, and
   sloppy-mode `this === window`. Nine of thirteen failed. Notably Math must NOT
   be sealed — Chromium reports Object.isSealed(Math) === false, and sealing it
   made our vector differ by one.

5. The solved payload dropped meta.origin / meta.stack / meta.duration.
   The duck.ai bundle always sends all three; captured browser requests confirm
   it. Without them upstream returns 418 even when every client_hash is correct.

6. reasoningEffort is now mandatory on duckchat/v1/chat.
   An otherwise byte-identical payload returns 200 with the field and 400
   ERR_BAD_REQUEST without it (A/B verified live, repeated).

Also removes the throwaway "seed" chat POST that ran before every real request.
It existed to coax a usable challenge out of the upstream while the solver was
broken; it only doubled chat calls against an IP-rate-limited endpoint, showing
up as spurious 429 ERR_RATE_LIMIT.

Verification: the solver now reproduces real Chromium's probe vectors exactly
for all 8 captured challenge variants, and the executor returns 200 end-to-end
live (non-streaming, streaming, claude-haiku-4-5, and a math prompt returning
"42").

Tests: tests/unit/duckduckgo-challenge-solver-regression.test.ts (32 tests) and
tests/unit/duckduckgo-reasoning-effort-required.test.ts (5 tests), backed by
tests/fixtures/duckduckgo/challenge-variants.json — real captured challenge
programs plus the probe vectors a real browser produced for them, so the suite
asserts against recorded browser behaviour rather than our own output. Each fix
was confirmed to fail its test when individually reverted.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants