Skip to content

feat(scraping): browser automation as a job factory, with zero new dependencies - #140

Merged
sebyx07 merged 1 commit into
mainfrom
feat/scraping-package
Aug 18, 2026
Merged

sebyx07 merged 1 commit into
mainfrom
feat/scraping-package

Conversation

@sebyx07

@sebyx07 sebyx07 commented Aug 18, 2026 •

Copy link
Copy Markdown
Contributor

Adds @ultimat3/scraping at tier 5. Designed from an audit of ~12,000 lines of real production scraping code across five repos, not from first principles.

scrape() is a job factory — the fourth, after llm(), agent() and backfill()

A scrape has an input schema, needs a required idempotency key (re-logging into a bank after a worker kill is the exact bug), a tenant, a retry policy, a timeout and concurrency — and decisively, step.run checkpoints, because recovery must resume at the broken page rather than restart the run. That is a job. No ninth primitive.

Tier 5 is the lowest its real imports allow: core/schema (0), storage (1, artifacts), jobs (3, it returns a JobHandle), ai (4, the recovery seam). Same argument CLAUDE.md makes for db at tier 1, run upward.

Zero new dependencies

puppeteer-core is not a dependency. The launcher is injected — localBrowser({ launcher }) — and the CDP port is declared structurally, the same shape s3Driver({ client }) already uses. The app owns the browser binary it had to own anyway.

Puppeteer over Playwright is not merely preference: Playwright's connectOverCDP cannot upgrade the WebSocket under Bun (oven-sh/bun#9911), which forced a real repo in this workspace into a two-runtime monorepo — disqualifying for a Bun-only framework. Verified live before committing to the design: Bun 1.3.14, puppeteer-core 25.8.0, headless Chrome 150, both launch() and connect({ browserWSEndpoint }) including the WebSocket upgrade.

What the audit changed about the design

Requirement Why, from production evidence
Remote CDP attach is the primary path production creates a stealth browser elsewhere and attaches; close() must stop both halves or the remote session bills forever
Hybrid browser → HTTP drive the browser through login/2FA, then hit the site's own JSON endpoints for bulk. http is session-bound: the browser's cookies, the same proxy (a different exit IP mid-session is itself an anti-bot trip), the same allowHosts, timeout and signal
One fixture format for both legs a hybrid path with two fixture stories is the untestable path, and that is where the real code lives
Session burn anti-bot cookies get poisoned; a retry that reloads a flagged profile re-trips the same block every time
Auth failure is terminal, never retried a surveyed site locks the account after three wrong attempts — a retrying framework becomes the thing that destroys the user's account
expect: { minRows, maxDrop } the worst failure a scraper has is the run that succeeds and returns nothing, and stays green for weeks. The collapsed run is deliberately not recorded, so the baseline cannot follow a collapse downward
Wedge/zombie discipline a graceful-quit ceiling, then SIGKILL, plus an inactivity watchdog — two named production incidents, one of which ran 3h11m and persisted nothing

Not shipped, deliberately

Stealth payloads and captcha. Two teams reached this independently — one moved stealth into a Chromium fork on purpose, another removed an injected script because the injection was itself detectable. A framework-shipped stealth payload is a shared fingerprint handed to every user.

No plugin API, per axiom 8. Extensibility is the ScrapeDriver seam (implementable from public exports alone) and wrapping scrape() — primitives are functions returning values.

Testing

  • fakePage is the default under bun test: Chrome is not required to run the suite.
  • An unrecorded fixture request throws rather than reaching the network — a silently-live "offline" test is the failure this prevents.
  • driver-parity.test.ts runs one suite across fake, fixture and the real driver's code path (over an injected fake CDP browser, never mock.module), so the fake cannot drift from the real one.
  • mock.module is banned in the package's own CLAUDE.md, with the observed cross-file leak that motivated it.
  • 114 tests. Writing the README examples so they compile found a real API bug: AuthContext had no secrets, so a login body had no way to reach the credential it is supposed to type. Fixed in source, not by weakening the example.

Note on local verification

bun run verify is green on this branch except for one test — scripts/verify.test.ts's four-full-repo-scan test — which times out under an 8-worker shard on a machine currently at load average 15 from unrelated processes. It costs 10.9s in isolation and has passed CI's 30s budget on the three PRs merged ahead of this one. Every other step and all 1239 other tests pass. CI is the uncontended gate here.

🤖 Generated with Claude Code

https://claude.ai/code/session_01J7WVaWYBVFBCdtHrVJHD5n


View with [code]smith Autofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.

…pendencies

`scrape()` is the fourth factory over an existing primitive, after `llm()`, `agent()` and
`backfill()`. A scrape has an input schema, needs a required idempotency key, a tenant, a
retry policy, a timeout and concurrency — and, decisively, `step.run` checkpoints, because
recovery must resume at the broken page rather than restart the run. That is a `job`, so
this returns one. No ninth primitive.

**puppeteer-core is not a dependency.** The launcher is injected — `localBrowser({ launcher })`
— and the CDP port is declared structurally, the same shape `s3Driver({ client })` already
uses. The app owns the browser binary it had to own anyway.

Puppeteer over Playwright is not only preference: Playwright's `connectOverCDP` cannot
upgrade the WebSocket under Bun (oven-sh/bun#9911), which forced a real repo into a
two-runtime monorepo. Verified live before committing to it — Bun 1.3.14, puppeteer-core
25.8.0, headless Chrome 150, both `launch()` and `connect({ browserWSEndpoint })`.

**Remote CDP attach is the primary path**, not an afterthought: production creates a
stealth browser elsewhere and attaches. `close()` stops both halves, or the remote bills
forever.

**Hybrid browser + HTTP.** Drive the browser through login and 2FA, then hit the site's own
JSON endpoints for the bulk. `http` is session-bound — the browser's cookies, the same
proxy, the same `allowHosts`, the same timeout and signal — and both legs replay from one
fixture format, because a hybrid path with two fixture stories is the untestable path.

**Sessions: acquire, persist, reuse, validate, burn.** The burn is the non-obvious half:
anti-bot cookies get poisoned, so a retry that reloads a flagged profile re-trips the same
block every time. An authentication failure is terminal and never retried — a site that
locks an account after three attempts turns a retrying framework into the thing that
destroys the user's account.

**`expect: { minRows, maxDrop }`** is the alarm for the worst failure a scraper has: the run
that succeeds and returns nothing, and stays green for weeks. The collapsed run is not
recorded, so the baseline cannot follow a collapse downward.

Not shipped, deliberately: stealth payloads and captcha. Two teams reached that
independently, one having removed an injected script because the injection was itself
detectable — a framework-shipped payload is a shared fingerprint handed to every user.

Testing: `fakePage` is the default under `bun test`, so Chrome is not required; an
unrecorded fixture request throws rather than reaching the network; and `driver-parity.test.ts`
runs one suite across fake, fixture and the real driver's code path, so the fake cannot
drift. `mock.module` is banned in the package's own CLAUDE.md, with the observed
cross-file leak that motivated it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01J7WVaWYBVFBCdtHrVJHD5n
@coderabbitai

coderabbitai Bot commented Aug 18, 2026

Copy link
Copy Markdown

Warning

Review limit reached

You’ve reached a temporary PR review limit under our Fair Usage Limits Policy.

Your current included review allowance is based on your included PR review attempts over the past 7 days.

Next review available in: 25 minutes

Limit details: You’ve used the included review currently available. Your 70 included PR review attempts over the past 7 days set your current allowance at 1 review per hour.

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits within each organization.

For paid Pro and Pro+ reviews, CodeRabbit uses a developer's included PR review attempts over the past 7 days to set the current hourly allowance. At typical activity levels, the full plan allowance applies. Higher sustained activity can lower the allowance until earlier attempts leave the 7-day window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 4cbc7297-309a-46f3-8067-520ae29d79e3

📥 Commits

Reviewing files that changed from the base of the PR and between 47c1a5d and 8029eb0.

⛔ Files ignored due to path filters (1)
  • bun.lock is excluded by !**/*.lock, !**/bun.lock
📒 Files selected for processing (67)
  • .env.example
  • CLAUDE.md
  • docs/architecture/01-package-map.md
  • framework.manifest.json
  • packages/cli/src/error-catalog.ts
  • packages/scraping/CLAUDE.md
  • packages/scraping/LICENSE
  • packages/scraping/README.md
  • packages/scraping/package.json
  • packages/scraping/src/actionability.test.ts
  • packages/scraping/src/actionability.ts
  • packages/scraping/src/artifacts.ts
  • packages/scraping/src/auth.test.ts
  • packages/scraping/src/auth.ts
  • packages/scraping/src/cdp-fake.ts
  • packages/scraping/src/cdp-port.ts
  • packages/scraping/src/cdp-snapshot.ts
  • packages/scraping/src/cdp-target.ts
  • packages/scraping/src/clock-discipline.test.ts
  • packages/scraping/src/clock.ts
  • packages/scraping/src/driver-cdp.ts
  • packages/scraping/src/driver-fake.ts
  • packages/scraping/src/driver-fixture.ts
  • packages/scraping/src/driver-parity.test.ts
  • packages/scraping/src/driver.ts
  • packages/scraping/src/error-throws.ts
  • packages/scraping/src/errors.test.ts
  • packages/scraping/src/errors.ts
  • packages/scraping/src/events.ts
  • packages/scraping/src/expect.test.ts
  • packages/scraping/src/expect.ts
  • packages/scraping/src/failures.ts
  • packages/scraping/src/hosts.test.ts
  • packages/scraping/src/hosts.ts
  • packages/scraping/src/html-query.test.ts
  • packages/scraping/src/html-query.ts
  • packages/scraping/src/html-requests.ts
  • packages/scraping/src/html-target.ts
  • packages/scraping/src/http-recorded.ts
  • packages/scraping/src/http.test.ts
  • packages/scraping/src/http.ts
  • packages/scraping/src/index.ts
  • packages/scraping/src/intercept.ts
  • packages/scraping/src/offline-session.ts
  • packages/scraping/src/page-over-target.test.ts
  • packages/scraping/src/page-over-target.ts
  • packages/scraping/src/page.ts
  • packages/scraping/src/rate.ts
  • packages/scraping/src/recording.ts
  • packages/scraping/src/recover.test.ts
  • packages/scraping/src/recover.ts
  • packages/scraping/src/rings.ts
  • packages/scraping/src/robots.test.ts
  • packages/scraping/src/robots.ts
  • packages/scraping/src/scrape-run.ts
  • packages/scraping/src/scrape.test.ts
  • packages/scraping/src/scrape.ts
  • packages/scraping/src/secrets.test.ts
  • packages/scraping/src/secrets.ts
  • packages/scraping/src/session-state.ts
  • packages/scraping/src/target.ts
  • packages/scraping/src/watchdog.test.ts
  • packages/scraping/src/watchdog.ts
  • packages/scraping/tsconfig.json
  • scripts/lib/tiers.ts
  • tsconfig.json
  • wiki/Error-Codes.md

Comment @coderabbitai help to get the list of available commands.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant