fix(executors): repair DuckDuckGo AI Chat challenge solver (418 ERR_CHALLENGE) - #9733
Merged
diegosouzapw merged 1 commit intoAug 11, 2026
Conversation
…HALLENGE)
Every duckduckgo-web chat request failed with HTTP 418 ERR_CHALLENGE while
duck.ai worked normally in a browser from the same IP. Ground truth was
established by driving a real headful Chromium at duck.ai from that IP (it
returned 200), so the environment was never the problem — the anti-abuse
challenge solver was. Six independent defects were found; the first alone
disabled the solver completely.
1. Module syntax inside the vm sandbox source.
CHALLENGE_STUBS is executed with vm.runInContext, which compiles in SCRIPT
mode. A refactor mass-added `export` to the five `function` declarations
inside that template literal (they read as ordinary top-level TS functions),
so every solve threw SyntaxError. The executor swallows solve failures and
posts the raw unsolved challenge, which upstream answers with 418.
2. Double-escaped regex in a String.raw template.
`\\s` in __parseCssDisplay reached the sandbox as a literal backslash, so the
display regex never matched and a getComputedStyle probe silently read empty.
3. buildHtmlLookup undercounted descendants by one.
`count` backs el.querySelectorAll('*').length; that returns DESCENDANTS and
countHtmlElements already skips the #document-fragment root, so the `- 1` was
wrong. Chromium reports 3 for '<li><div></li><li></div'; we reported 2, and a
variant multiplies innerHTML.length by that count.
4. Browser-fidelity probes.
Newer challenge variants assert JS/DOM invariants a flat stub cannot satisfy:
real prototype chains (HTMLDivElement -> HTMLElement -> Element), NodeList
identity, a live body.children HTMLCollection, native-code toString, and
sloppy-mode `this === window`. Nine of thirteen failed. Notably Math must NOT
be sealed — Chromium reports Object.isSealed(Math) === false, and sealing it
made our vector differ by one.
5. The solved payload dropped meta.origin / meta.stack / meta.duration.
The duck.ai bundle always sends all three; captured browser requests confirm
it. Without them upstream returns 418 even when every client_hash is correct.
6. reasoningEffort is now mandatory on duckchat/v1/chat.
An otherwise byte-identical payload returns 200 with the field and 400
ERR_BAD_REQUEST without it (A/B verified live, repeated).
Also removes the throwaway "seed" chat POST that ran before every real request.
It existed to coax a usable challenge out of the upstream while the solver was
broken; it only doubled chat calls against an IP-rate-limited endpoint, showing
up as spurious 429 ERR_RATE_LIMIT.
Verification: the solver now reproduces real Chromium's probe vectors exactly
for all 8 captured challenge variants, and the executor returns 200 end-to-end
live (non-streaming, streaming, claude-haiku-4-5, and a math prompt returning
"42").
Tests: tests/unit/duckduckgo-challenge-solver-regression.test.ts (32 tests) and
tests/unit/duckduckgo-reasoning-effort-required.test.ts (5 tests), backed by
tests/fixtures/duckduckgo/challenge-variants.json — real captured challenge
programs plus the probe vectors a real browser produced for them, so the suite
asserts against recorded browser behaviour rather than our own output. Each fix
was confirmed to fail its test when individually reverted.
Contributor
Author
|
I manually tested the DuckDuckGo provider on a fully running omniroute server. |
Mynacol
marked this pull request as ready for review
August 7, 2026 21:27
Owner
|
Thanks for the PR. Please address the mandatory items (tests and/or merge blockers) in this branch, then rerun checks before /merge-prs. |
2 similar comments
Owner
|
Thanks for the PR. Please address the mandatory items (tests and/or merge blockers) in this branch, then rerun checks before /merge-prs. |
Owner
|
Thanks for the PR. Please address the mandatory items (tests and/or merge blockers) in this branch, then rerun checks before /merge-prs. |
Owner
|
Obrigado pelo PR. Mantive a revisão de
|
muhamadgalihsaputra
pushed a commit
to niyatna/NiyatnaRoute
that referenced
this pull request
Sep 27, 2026
…HALLENGE) (diegosouzapw#9733) Every duckduckgo-web chat request failed with HTTP 418 ERR_CHALLENGE while duck.ai worked normally in a browser from the same IP. Ground truth was established by driving a real headful Chromium at duck.ai from that IP (it returned 200), so the environment was never the problem — the anti-abuse challenge solver was. Six independent defects were found; the first alone disabled the solver completely. 1. Module syntax inside the vm sandbox source. CHALLENGE_STUBS is executed with vm.runInContext, which compiles in SCRIPT mode. A refactor mass-added `export` to the five `function` declarations inside that template literal (they read as ordinary top-level TS functions), so every solve threw SyntaxError. The executor swallows solve failures and posts the raw unsolved challenge, which upstream answers with 418. 2. Double-escaped regex in a String.raw template. `\\s` in __parseCssDisplay reached the sandbox as a literal backslash, so the display regex never matched and a getComputedStyle probe silently read empty. 3. buildHtmlLookup undercounted descendants by one. `count` backs el.querySelectorAll('*').length; that returns DESCENDANTS and countHtmlElements already skips the #document-fragment root, so the `- 1` was wrong. Chromium reports 3 for '<li><div></li><li></div'; we reported 2, and a variant multiplies innerHTML.length by that count. 4. Browser-fidelity probes. Newer challenge variants assert JS/DOM invariants a flat stub cannot satisfy: real prototype chains (HTMLDivElement -> HTMLElement -> Element), NodeList identity, a live body.children HTMLCollection, native-code toString, and sloppy-mode `this === window`. Nine of thirteen failed. Notably Math must NOT be sealed — Chromium reports Object.isSealed(Math) === false, and sealing it made our vector differ by one. 5. The solved payload dropped meta.origin / meta.stack / meta.duration. The duck.ai bundle always sends all three; captured browser requests confirm it. Without them upstream returns 418 even when every client_hash is correct. 6. reasoningEffort is now mandatory on duckchat/v1/chat. An otherwise byte-identical payload returns 200 with the field and 400 ERR_BAD_REQUEST without it (A/B verified live, repeated). Also removes the throwaway "seed" chat POST that ran before every real request. It existed to coax a usable challenge out of the upstream while the solver was broken; it only doubled chat calls against an IP-rate-limited endpoint, showing up as spurious 429 ERR_RATE_LIMIT. Verification: the solver now reproduces real Chromium's probe vectors exactly for all 8 captured challenge variants, and the executor returns 200 end-to-end live (non-streaming, streaming, claude-haiku-4-5, and a math prompt returning "42"). Tests: tests/unit/duckduckgo-challenge-solver-regression.test.ts (32 tests) and tests/unit/duckduckgo-reasoning-effort-required.test.ts (5 tests), backed by tests/fixtures/duckduckgo/challenge-variants.json — real captured challenge programs plus the probe vectors a real browser produced for them, so the suite asserts against recorded browser behaviour rather than our own output. Each fix was confirmed to fail its test when individually reverted.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
duckduckgo-webprovider, which returned HTTP 418ERR_CHALLENGEon every chat request while duck.ai worked normally in a browser from the same IP. The provider was fully unusable.200. That ruled out IP blocking / rate limiting and proved the fault was in our anti-abuse challenge solver.429 ERR_RATE_LIMIT.Root causes
exportkeywords insideCHALLENGE_STUBS. That string is run byvm.runInContext, which compiles in script mode, so module syntax is a hardSyntaxError. A refactor mass-addedexportto its 5functiondeclarations.\\sinside aString.rawtemplate reached the sandbox as a literal backslashgetComputedStyledisplay probe silently read emptybuildHtmlLookupsubtracted 1 from the descendant count<li><div></li><li></div; we sent 2NodeListidentity, livebody.children, native-codetoString, sloppy-modethis)meta.origin/meta.stack/meta.duration, which the real duck.ai bundle always sendsclient_hashwas correctreasoningEffortis now mandatory onduckchat/v1/chat400 ERR_BAD_REQUESTNote on #4:
Mathmust not be sealed — real Chromium reportsObject.isSealed(Math) === false, and sealing it made our vector differ by one.Related Issues
Validation
npm run lintCommands run (Node 22.23.2 provisioned via
nix, since the sandbox had no system Node):Live end-to-end validation — the executor now returns
200against real DuckDuckGo, 4/4 cases:gpt-4o-mini(alias path)200→"OK"gpt-5.4-mini200→"OK"claude-haiku-4-5(reasoningEffort: low)200→"OK"gpt-5.4-nano200→"42"This satisfies the AGENTS.md hard-rule #18 bug-fix gate via both paths: automated regression tests (TDD) and a documented live real-environment test.
Tests Added Or Updated
tests/unit/duckduckgo-challenge-solver-regression.test.ts— new, 32 teststests/unit/duckduckgo-reasoning-effort-required.test.ts— new, 5 teststests/unit/duckduckgo-challenge-split.test.ts— updated, +4 tests guarding the script-mode/exportinvarianttests/fixtures/duckduckgo/challenge-variants.json— new fixture: 8 real challenge programs captured from duckduckgo.com, each paired with the probe vectors a real headful Chromium produced for that exact programThe suite asserts against recorded real-browser behaviour, not against our own output — matching Chromium bit-for-bit is the actual correctness criterion here. Every fix was individually reverted to confirm its test fails:
exportto stubs\\sescapingcount - 1metaaugmentationMathreasoningEffort: nullCoverage Notes
This PR changes
open-sse/executors/duckduckgo-web.tsandopen-sse/executors/duckduckgo-web/challenge.ts.solveDuckDuckGoChallenge,CHALLENGE_STUBS, andbuildHtmlLookupare covered by the 8 real-variant fixtures plus 16 individually-named browser-fidelity probe tests, so a future stub regression names itself instead of failing as one opaque hash mismatch.reasoningEffortchange and the seed-POST removal are covered byduckduckgo-reasoning-effort-required.test.ts, which stubsfetchand asserts on the actual outgoing wire payload (including a test that exactly onePOST /duckchat/v1/chatis issued per user request).SyntaxErrorinsideCHALLENGE_STUBSshipped unnoticed.Reviewer Notes
Base-red inherited — not caused by this PR. Three failures pre-exist on
origin/release/v3.8.50and reproduce identically on a pristine checkout of that tip:tests/unit/models-catalog-route.test.ts—ERR_MODULE_NOT_FOUND: open-sse/services/antigravityProjectPersistence.ts. The file is imported byopen-sse/services/combo/quotaStrategies.tsbut does not exist in git at that ref.tests/unit/provider-translate-path-golden.test.ts— golden drift fromdevin-cli-agenticandraycastmissing from the snapshot. Verified theduckduckgo-webentry is byte-identical and untouched by this PR.npm run lint— 3no-explicit-anyerrors intests/unit/vertex-functioncall-id-3440.test.ts, a file this PR does not touch.Per
CONTRIBUTING.mdthese are not fixed here; a base-red fix belongs in its ownfix/release-v3.8.50-basereds-*PR.Risk areas
vmsandbox is a supply-chain surface.solveDuckDuckGoChallengeexecutes upstream-supplied JS. The security posture is unchanged — stillvm.runInContextwith the 5s timeout — and theSECURITY NOTEremains. This PR only makes the emulated browser environment more faithful; it does not widen what the sandbox can reach.CHALLENGE_STUBS, both hit during development and now documented inline: (a) a backtick in a comment terminates theString.rawtemplate; (b)\\xin a comment is still a double escape. AWARNINGblock at the top of the literal spells out the script-mode constraint.reasoningEffort("none"by default,"low"forclaude-haiku-4-5/tinfoil/gpt-oss-120b). Verified live as strictly required — the previous "let the server default it" path returns400.No migrations, no feature flags, no DB changes.
Manual validation reviewers may want to repeat: the live 4/4 table above. Note DuckDuckGo aggressively IP-rate-limits (
429 ERR_RATE_LIMIT) under repeated testing; that is environmental and distinct from the418this PR fixes.