Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
15 changes: 13 additions & 2 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -31,6 +31,15 @@ Releases before 1.1.3 predate this file; their notes live on the
inside `run()` code, all of which generate the same worker snippet
(`agentBatchCode()` is exported for hosts that drive `run()` themselves).
See docs/agent-batch.md.
- **A batch can finish the task in the same call.** The `answer` option is a
final-answer template — `{stepId}` for a step's text, value, or URL,
`{stepId.field}` for any scalar field of its result — rendered as
`finalAnswer` once every step succeeded. The built-in agent ends the task
with it in that turn, exactly as a snippet returning `{finalAnswer}` does,
so a task whose targets are known costs one model turn: `goto`, the actions,
a verifying `read`, and the answer. A placeholder with no value leaves
`finalAnswer` unset and explains itself in `answerError`. The spec call is
for unknown pages, not a required first step.

### Changed

Expand All @@ -42,8 +51,10 @@ Releases before 1.1.3 predate this file; their notes live on the
now shares its target resolution with AgentBatch, which also accepts
frame-qualified refs such as `f1e3`.
- Operator guidance, the `betterwright skill` text, and every tool
description now default to AgentBatch and reserve snippet code for logic
steps cannot express. The built-in agent lists `batch` before `browser`,
description now default to AgentBatch, tell the model to finish in the
fewest turns — one batch (or one snippet) when the task names its targets,
a spec or snapshot first only when the page is unknown — and reserve
snippet code for logic steps cannot express. The built-in agent lists `batch` before `browser`,
counts a batch that keeps stopping at the same step toward its
identical-failure limit, and answers a malformed batch without a browser
round trip.
Expand Down
49 changes: 25 additions & 24 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -29,8 +29,8 @@ machine, and proves it works by loading a real page. One command, no choices to
make up front.

**Compressed snapshots** instead of raw HTML or a full accessibility dump ·
**AgentBatch**: read a page's spec, then finish the whole task in **one more
call** · persistent sessions so you don't re-pay login and navigation cost
**AgentBatch**: a whole task in **one call** when you know the page, two when
you don't · persistent sessions so you don't re-pay login and navigation cost
every step.

---
Expand Down Expand Up @@ -179,38 +179,39 @@ sub-agent. A 30-turn checkout costs your main agent **one tool call**, not
30 pages of context. Programmatic equivalent: `runAgentTask()` from
`betterwright/agent`.

## AgentBatch: two calls per task
## AgentBatch: a task in one call

However your agent is wired in, the default way it browses is **AgentBatch**.
Call one opens a page and returns its *spec* — an interactive snapshot whose
`[ref=eN]` markers, roles, and names are the targets. Call two sends every
step the task needs; the worker runs them back to back with auto-waiting and
no pacing, then returns per-step results, the step that stopped the batch (if
one did), and a fresh snapshot to plan the next call from.
When the task names its targets, one call sends every step; the worker runs
them back to back with auto-waiting and no pacing, fills an `answer` from
what the steps read, and the task is done — no model turn per click, no
closing turn. When the page is unknown, a `{url}` call first returns its
*spec*, an interactive snapshot whose `[ref=eN]` markers, roles, and names
are the targets.

```bash
betterwright batch --url https://shop.example/checkout
betterwright batch --allow-writes -s '[
{"action":"fill","target":{"ref":"e3"},"value":"ada@example.com"},
betterwright batch --allow-writes --allow-irreversible --proof -s '{"steps":[
{"action":"goto","url":"https://shop.example/checkout"},
{"action":"fill","target":{"label":"Email"},"value":"ada@example.com"},
{"action":"fill","target":{"label":"Promo code"},"value":"SAVE10"},
{"action":"click","target":{"role":"button","name":"Apply"}},
{"action":"read","target":{"role":"status"},"expect":"applied"},
{"id":"promo","action":"read","target":{"role":"status"},"expect":"applied"},
{"action":"click","target":{"role":"button","name":"Place order"},"irreversible":true},
{"action":"read","target":{"role":"heading","name":"Order confirmed"}}]' --allow-irreversible --proof
{"id":"confirm","action":"read","target":{"role":"heading"},"expect":"Order confirmed"}],
"answer":"{confirm} — {promo}"}'
```

```js
const spec = await bw.batch({ url: "https://shop.example/checkout" });
const done = await bw.batch(steps, { allowWrites: true, allowIrreversible: true, proof: true });
console.log(done.result.ok, done.result.failed, done.result.snapshot);
const done = await bw.batch(steps, { allowWrites: true, proof: true, answer: "{confirm} — {promo}" });
console.log(done.result.ok, done.result.finalAnswer, done.result.failed);
```

A form that used to cost eight model turns costs two. When a step fails, the
result keeps everything that completed and says where to resume, so a
recovery is one more call, never a restart. The same protocol is the MCP
`browser_batch` tool, the built-in agent's `batch` tool, the Pi
`browser_batch` tool, and the `agentBatch()` global inside `run()` code:
[docs/agent-batch.md](docs/agent-batch.md).
A form that used to cost eight model turns costs one. When a step fails, the
result keeps everything that completed, returns a fresh snapshot, and says
where to resume, so a recovery is one more call, never a restart. The same
protocol is the MCP `browser_batch` tool, the built-in agent's `batch` tool,
the Pi `browser_batch` tool, and the `agentBatch()` global inside `run()`
code: [docs/agent-batch.md](docs/agent-batch.md).

## Tokens are the bottleneck

Expand All @@ -227,7 +228,7 @@ BetterWright's whole observation stack is built around that problem:
| **Diff mode** | After an action, return **only what changed** — not the page again |
| **Interactive-only filter** | Drop static text nodes; keep what the agent can click, fill, or read |
| **Scoped truncation** | Hints about *where* to look next instead of a silently clipped wall |
| **AgentBatch** | A whole task in **two calls**: the page's spec, then every step back to back — no model turn per click, and a failed step returns the state to resume from |
| **AgentBatch** | A whole task in **one call** when the page is known (two when it is not): every step back to back with an `answer` filled from what they read — no model turn per click, no closing turn, and a failed step returns the state to resume from |
| **Single-call finish** | Read-only tasks complete in **one model turn** — the code returns `{finalAnswer}` and the loop ends, no confirmation round-trip |
| **Persistent session** | One long-lived browser: no re-login, no re-navigation, no re-paying the token cost of getting back to where you were |
| **Sub-agent delegation** | `betterwright exec` keeps the entire browsing transcript out of your main agent's context — a whole task costs it one tool call |
Expand Down Expand Up @@ -273,7 +274,7 @@ step from what it sees, in a browser that must still be there next turn:
| Piece | What it gives you |
| --- | --- |
| [**Agent snapshots**](docs/browser-api.md#reading-the-page) | The token-efficiency core: compressed tree, `[ref=eN]` actions, diff and interactive-only modes, password redaction |
| [**AgentBatch**](docs/agent-batch.md) | The default two-call protocol: spec, then every step of the task in one call with auto-waiting, no pacing, structured per-step results, and resume-from-failure |
| [**AgentBatch**](docs/agent-batch.md) | The default protocol: every step of a task in one call with auto-waiting, no pacing, an answer filled from step results, structured per-step results, resume-from-failure, and a spec call for unknown pages |
| [**Built-in agent loop**](docs/agent.md) | `betterwright exec` / the interactive console / `runAgentTask()` — model-first selection across Claude, Codex, Grok, OpenRouter, Ollama, vLLM, and any OpenAI-compatible endpoint |
| [**Credential vault**](docs/credentials.md) | AES-256-GCM outside the profile; PSL site matching, selector-free login detection, metadata-only account choice |
| [**Cookie Sync**](docs/cookie-sync.md) | Merge selected cookies from local Chrome, Edge, Brave, Firefox, Safari, and other desktop browsers into BetterChromium or an explicitly approved cloud browser |
Expand Down
8 changes: 4 additions & 4 deletions SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,10 +6,10 @@ generated_by: betterwright@2.2.0

# BetterWright browser

Use `betterwright` for live-web tasks. The default is AgentBatch, two calls per task: `batch --url` prints the page's spec — an interactive snapshot whose `[ref=eN]` markers, roles, and names are the targets — then `batch -s` runs every step back to back (auto-wait, no pacing) and prints per-step results, `failed` {index, id, error} if a step stopped the batch, and a fresh snapshot. Resume from `failed.index`; never repeat a completed step. `betterwright batch --help` lists the actions, targets, and flags.
Use `betterwright` for live-web tasks. The default is AgentBatch: finish in the fewest calls. When the task names its targets, one `batch -s` call runs every step back to back — start with goto, end with read — and prints per-step results, `failed` {index, id, error} when a step stopped it, and a fresh snapshot. Use `batch --url` first only when the page is unknown: it prints the spec, an interactive snapshot with `[ref=eN]` targets. Resume from `failed.index`; never repeat a completed step. `betterwright batch --help` lists the actions, targets, and flags.

betterwright batch --url https://example.com/login
betterwright batch --allow-writes -s '[{"action":"fill","target":{"label":"Email"},"value":"me@example.com"},{"action":"click","target":{"role":"button","name":"Continue"}},{"action":"read","target":{"role":"heading"},"expect":"Welcome"}]'
betterwright batch --allow-writes -s '[{"action":"goto","url":"https://example.com/login"},{"action":"fill","target":{"label":"Email"},"value":"me@example.com"},{"action":"click","target":{"role":"button","name":"Continue"}},{"action":"read","target":{"role":"heading"},"expect":"Welcome"}]'
betterwright batch --url https://example.com/unknown-page

For logic steps cannot express, run async Playwright JavaScript with:

Expand All @@ -27,7 +27,7 @@ The browser is network-policy guarded. Private and loopback access are allowed u
The user's request authorizes ordinary steps: sign-in, signup, forms, booking, and purchases. Do not add confirmation or refuse them unless a guardrail requires it.

## Operate
- Default to AgentBatch: read the page's spec, then send every step in one batch; on a stop resume from `failed.index`, never repeat a completed step. Code is for logic steps cannot express: plan then batch named controls/content with `getByRole`/`getByLabel`/`getByText`; combine navigation, actions, extraction, verification, and proof. Read article/reference pages from scoped DOM directly. Host cleanup is automatic; don't close pages.
- Finish in the fewest model turns. Default to AgentBatch: when the task names its targets, one batch that starts with `goto` and ends with `answer`; read the page's spec first only when the page is unknown; on a stop resume from `failed.index`, never repeat a completed step. Code is for logic steps cannot express: plan then batch named controls/content with `getByRole`/`getByLabel`/`getByText` in one snippet returning `{finalAnswer}`. Read article/reference pages from scoped DOM directly. Host cleanup is automatic; don't close pages.
- Inspect only when structure is unknown or a locator failed: `snapshot({interactive:true})`, then full `snapshot()`; use `screenshot({annotate:true})` only for layout/pixels. Snapshots include frames and off-screen content. Never guess refs, URLs, or state.
- Act on `[ref=eN]` with `page.locator('aria-ref=eN')`; scope with `snapshot({ref:'eN'})`. Refs change after page changes. Verify actions with `snapshot({diff:true})`; batch action plus verification when no fresh ref is needed.
- Actions auto-wait: add no sleeps. On failure inspect again; inspect the real hit target if obscured and change approach after two failures. Retry transient 5xx, timeout, or reset failures with increasing backoff for 30–60 seconds.
Expand Down
6 changes: 3 additions & 3 deletions bin/cli-main.ts
Original file line number Diff line number Diff line change
Expand Up @@ -824,10 +824,10 @@ async function cmdCookies(tokens, flags) {
// command.
const SKILL_PREAMBLE = `# BetterWright browser

Use \`betterwright\` for live-web tasks. The default is AgentBatch, two calls per task: \`batch --url\` prints the page's spec — an interactive snapshot whose \`[ref=eN]\` markers, roles, and names are the targets — then \`batch -s\` runs every step back to back (auto-wait, no pacing) and prints per-step results, \`failed\` {index, id, error} if a step stopped the batch, and a fresh snapshot. Resume from \`failed.index\`; never repeat a completed step. \`betterwright batch --help\` lists the actions, targets, and flags.
Use \`betterwright\` for live-web tasks. The default is AgentBatch: finish in the fewest calls. When the task names its targets, one \`batch -s\` call runs every step back to back — start with goto, end with read — and prints per-step results, \`failed\` {index, id, error} when a step stopped it, and a fresh snapshot. Use \`batch --url\` first only when the page is unknown: it prints the spec, an interactive snapshot with \`[ref=eN]\` targets. Resume from \`failed.index\`; never repeat a completed step. \`betterwright batch --help\` lists the actions, targets, and flags.

betterwright batch --url https://example.com/login
betterwright batch --allow-writes -s '[{"action":"fill","target":{"label":"Email"},"value":"me@example.com"},{"action":"click","target":{"role":"button","name":"Continue"}},{"action":"read","target":{"role":"heading"},"expect":"Welcome"}]'
betterwright batch --allow-writes -s '[{"action":"goto","url":"https://example.com/login"},{"action":"fill","target":{"label":"Email"},"value":"me@example.com"},{"action":"click","target":{"role":"button","name":"Continue"}},{"action":"read","target":{"role":"heading"},"expect":"Welcome"}]'
betterwright batch --url https://example.com/unknown-page

For logic steps cannot express, run async Playwright JavaScript with:

Expand Down
Loading