Skip to content

feat: let an AgentBatch finish the task in one call - #161

Merged
SSHdotCodes merged 3 commits into
mainfrom
agent/agent-batch-one-call
Sep 3, 2026
Merged

SSHdotCodes merged 3 commits into
mainfrom
agent/agent-batch-one-call

Conversation

@SSHdotCodes

Copy link
Copy Markdown
Collaborator

What and why

Follow-up to #160. Benchmarking the AgentBatch default against the previous build with gpt-5.6-sol (medium, 6 tasks × 2 rounds, arms interleaved) came out 0.90× on wall clock: a strong model already wrote whole tasks as one browser snippet returning {finalAnswer} (one model turn), while the batch protocol cost a spec turn, a steps turn, and a done turn. Model turns decide task time; the delay between browser actions does not.

This PR makes a batch as cheap in turns as the snippet path:

  • answer: a final-answer template rendered from step results once every step succeeded ({stepId} = the step's text, value, or URL; {stepId.field} = any scalar field). The built-in agent finishes the task with the rendered finalAnswer in the same turn. A placeholder with no value leaves finalAnswer unset and explains itself in answerError, so a stopped or empty read never becomes a claimed answer.
  • Spec call optional: guidance, tool descriptions, and the skill text are now turn-centric — when the task names its targets, one batch (goto first, read last, answer) or one snippet returning {finalAnswer}; {url} is for unknown pages.
  • Docs, changelog, regenerated SKILL.md, and the public types updated; budgets for the agent and MCP tool lists raised (6,000 → 6,500 and 9,000 → 9,250) for the new option and its guidance.

Benchmark, honestly

Same rig as before (gpt-5.6-sol, effort medium, Codex OAuth; 6 tasks × 2 rounds, arms interleaved; "old" = pre-#160 build, "new" = this branch). All 24 runs correct on both arms.

old (snippets) #160 (batch default) this branch
model turns 27 33 28
total wall clock 167 s / 156 s 185 s 207 s
input tokens 49.8k / 48.8k 72.3k 65.8k
output tokens 5.7k / 5.5k 5.9k 6.9k

(old was measured in both sessions; both values shown.) The turn penalty of #160 is gone: turn counts are at parity, and login dropped from 4 turns to 2. Wall clock is not better in this sample: the new arm emits about 25% more output tokens per turn, and with only two rounds per cell the per-turn latency noise is large (the same 3-call form task took 25.6 s on one arm and 10.8 s on the other). gpt-5.6-sol did not use answer in any run (it chose browser snippets in 9 of 12 runs, batch in 3), so the one-call finish is exercised by the tests, not yet by this model in the wild. So: this PR removes the structural turn cost; it does not demonstrate a speedup with gpt-5.6-sol. Raw results: ~/bw-bench/results-2 on the author's machine.

Checklist

  • bun run release:check passes locally (versions, lint, typecheck, build,
    unit tests, published declarations, tarball).
  • Every browser connection still goes through the guard proxy. No new
    launch path, transport, or fetch.
  • No secret value reaches the model sandbox. The answer template is
    rendered from step results that are already redacted (password readings
    are [redacted]), and the rendered result goes through the envelope's
    redaction like any other field.
  • Dependency pins are still mirrored everywhere. No pin changed.
  • Public API changes update types/*.d.ts in this same commit.
    types/common.d.ts gains AgentBatchOptions.answer and
    AgentBatchResult.finalAnswer / answerError.
  • The two branch-protected CI job names are unchanged; no workflow changes.
  • No unit test imports src/worker.ts or dist/src/worker.js directly.
  • User-visible behaviour is reflected in CHANGELOG.md, docs/agent-batch.md,
    README, docs/{agent,javascript,getting-started}.md, and the regenerated SKILL.md.
  • Nothing private is being committed.

How it was verified

  • bun run release:check and the managed-browser suite (BETTERWRIGHT_REQUIRE_BROWSER=1 bun run test) on Bun 1.4.0 / macOS arm64, including a new end-to-end check that a batch with answer returns the rendered finalAnswer.
  • Unit tests for template rendering (shorthand fall-through, explicit fields, unknown step/field refusal, no answer on a stopped batch) and for the agent loop finishing in one turn from a batch finalAnswer.
  • The A/B benchmark above, re-run on this branch.

🤖 Generated with Claude Code

https://claude.ai/code/session_01Td3ZP4tb9UKEQHLaoAo6NN

An A/B of the AgentBatch default against the previous build with
gpt-5.6-sol (medium) came out 0.90x on wall clock: a strong model already
wrote whole tasks as one snippet returning {finalAnswer}, while the batch
protocol cost a spec turn, a steps turn, and a `done` turn. Model turns,
not the delay between browser actions, decide how fast a task finishes.

The batch `answer` option is a final-answer template rendered from step
results once every step succeeded — {stepId} for a step's text, value, or
URL, {stepId.field} for any scalar field — and the built-in agent ends the
task with it in that turn, exactly as the snippet path does. A placeholder
with no value leaves finalAnswer unset and explains itself in answerError,
so an answer never claims what the page did not show.

Guidance across the prompt, skill text, and tool descriptions is now
turn-centric: when the task names its targets, one batch (goto first,
read last, answer) or one snippet returning {finalAnswer}; the {url} spec
call is for unknown pages, not a required first step.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Td3ZP4tb9UKEQHLaoAo6NN
@greptile-apps

greptile-apps Bot commented Sep 3, 2026 •

Copy link
Copy Markdown

Greptile Summary

Summary

Answer templates now reserve braces exclusively for valid step-result placeholders and return a recoverable answer error for malformed input.

Two previously reported issues were disproved by an executed focused runtime check:

  • Unsupported placeholders bypass validation: {1price} and {step.1} now throw during rendering, and batch execution returns answerError without finalAnswer.
  • Nested braces bypass validation: {{flash}} now throws during rendering, and batch execution returns answerError without finalAnswer.

The same check confirmed that stray closing braces and unclosed opening braces are rejected through both rendering and batch execution.

Confidence Score: 5/5

No blocking failure remains.

The executed checks disproved both previously reported answer-template failure paths, and no accepted blocking finding remains.

T-Rex T-Rex Logs

What T-Rex did

  • Ran the brace-template-validation.ts script from the repository root to exercise the focused runtime harness against five invalid brace templates; renderAgentBatchAnswer and executeAgentBatch were invoked with {1price}, {step.1}, {{flash}}, {flash}}, and {flash}, and all five templates threw during direct rendering, returning answerError with no finalAnswer in batch execution.
  • Encountered a prerequisite blocker during the initial setup: Bun 1.3.14 reported Unknown lockfile version for Bun v2, halting progress.
  • Validated the post-blocker behavior: the after-log shows each of the five invalid templates has finalAnswer: null and an expected answerError.

View all artifacts

T-Rex Ran code and verified through T-Rex

Reviews (3): Last reviewed commit: "fix: reserve braces in answer templates ..." | Re-trigger Greptile

Comment thread src/agent-batch.ts Outdated
…text

A brace group that did not fit the placeholder grammar, such as {1price}
or {step.1}, never reached the renderer's callback and stayed in the
rendered finalAnswer, ending the task with an unfilled token. Every brace
group in a template is now checked against the grammar, and a malformed
one leaves finalAnswer unset with the reason in answerError.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Td3ZP4tb9UKEQHLaoAo6NN
@SSHdotCodes

Copy link
Copy Markdown
Collaborator Author

@greptileai review — the latest commit rejects malformed answer placeholders (every brace group must match {stepId} or {stepId.field}; otherwise answerError, no finalAnswer), with tests through both the renderer and executeAgentBatch.

Comment thread src/agent-batch.ts Outdated
A nested placeholder such as {{flash}} rendered its inner token and kept
the outer braces in finalAnswer. The renderer now scans the template
linearly: every "{" must open a placeholder that ends at the next "}",
and a stray, nested, or malformed brace is an answerError, never text.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Td3ZP4tb9UKEQHLaoAo6NN
@SSHdotCodes

Copy link
Copy Markdown
Collaborator Author

@greptileai review — the latest commit reserves braces in answer templates: the renderer scans linearly, every { must open a placeholder ending at the next }, and nested ({{flash}}), stray, or malformed braces produce answerError with no finalAnswer; tests cover {{flash}}, {{flash}, {flash}}, and stray braces through the renderer and executeAgentBatch.

@SSHdotCodes
SSHdotCodes merged commit 3f36b9f into main Sep 3, 2026
26 of 28 checks passed
SSHdotCodes added a commit that referenced this pull request Sep 3, 2026
* Revert "feat: let an AgentBatch finish the task in one call (#161)"

This reverts commit 3f36b9f.

* Revert "feat: AgentBatch — the default two-call way agents browse (#160)"

This reverts commit 45c7642.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant