Repository navigation
feat: let an AgentBatch finish the task in one call - #161
Conversation
An A/B of the AgentBatch default against the previous build with
gpt-5.6-sol (medium) came out 0.90x on wall clock: a strong model already
wrote whole tasks as one snippet returning {finalAnswer}, while the batch
protocol cost a spec turn, a steps turn, and a `done` turn. Model turns,
not the delay between browser actions, decide how fast a task finishes.
The batch `answer` option is a final-answer template rendered from step
results once every step succeeded — {stepId} for a step's text, value, or
URL, {stepId.field} for any scalar field — and the built-in agent ends the
task with it in that turn, exactly as the snippet path does. A placeholder
with no value leaves finalAnswer unset and explains itself in answerError,
so an answer never claims what the page did not show.
Guidance across the prompt, skill text, and tool descriptions is now
turn-centric: when the task names its targets, one batch (goto first,
read last, answer) or one snippet returning {finalAnswer}; the {url} spec
call is for unknown pages, not a required first step.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Td3ZP4tb9UKEQHLaoAo6NN
Greptile SummarySummaryAnswer templates now reserve braces exclusively for valid step-result placeholders and return a recoverable answer error for malformed input. Two previously reported issues were disproved by an executed focused runtime check:
The same check confirmed that stray closing braces and unclosed opening braces are rejected through both rendering and batch execution. Confidence Score: 5/5No blocking failure remains. The executed checks disproved both previously reported answer-template failure paths, and no accepted blocking finding remains.
What T-Rex did
Reviews (3): Last reviewed commit: "fix: reserve braces in answer templates ..." | Re-trigger Greptile |
…text
A brace group that did not fit the placeholder grammar, such as {1price}
or {step.1}, never reached the renderer's callback and stayed in the
rendered finalAnswer, ending the task with an unfilled token. Every brace
group in a template is now checked against the grammar, and a malformed
one leaves finalAnswer unset with the reason in answerError.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Td3ZP4tb9UKEQHLaoAo6NN
|
@greptileai review — the latest commit rejects malformed answer placeholders (every brace group must match {stepId} or {stepId.field}; otherwise answerError, no finalAnswer), with tests through both the renderer and executeAgentBatch. |
A nested placeholder such as {{flash}} rendered its inner token and kept
the outer braces in finalAnswer. The renderer now scans the template
linearly: every "{" must open a placeholder that ends at the next "}",
and a stray, nested, or malformed brace is an answerError, never text.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Td3ZP4tb9UKEQHLaoAo6NN
|
@greptileai review — the latest commit reserves braces in answer templates: the renderer scans linearly, every { must open a placeholder ending at the next }, and nested ({{flash}}), stray, or malformed braces produce answerError with no finalAnswer; tests cover {{flash}}, {{flash}, {flash}}, and stray braces through the renderer and executeAgentBatch. |
What and why
Follow-up to #160. Benchmarking the AgentBatch default against the previous build with gpt-5.6-sol (medium, 6 tasks × 2 rounds, arms interleaved) came out 0.90× on wall clock: a strong model already wrote whole tasks as one
browsersnippet returning{finalAnswer}(one model turn), while the batch protocol cost a spec turn, a steps turn, and adoneturn. Model turns decide task time; the delay between browser actions does not.This PR makes a batch as cheap in turns as the snippet path:
answer: a final-answer template rendered from step results once every step succeeded ({stepId}= the step's text, value, or URL;{stepId.field}= any scalar field). The built-in agent finishes the task with the renderedfinalAnswerin the same turn. A placeholder with no value leavesfinalAnswerunset and explains itself inanswerError, so a stopped or empty read never becomes a claimed answer.gotofirst,readlast,answer) or one snippet returning{finalAnswer};{url}is for unknown pages.SKILL.md, and the public types updated; budgets for the agent and MCP tool lists raised (6,000 → 6,500 and 9,000 → 9,250) for the new option and its guidance.Benchmark, honestly
Same rig as before (gpt-5.6-sol, effort medium, Codex OAuth; 6 tasks × 2 rounds, arms interleaved; "old" = pre-#160 build, "new" = this branch). All 24 runs correct on both arms.
(old was measured in both sessions; both values shown.) The turn penalty of #160 is gone: turn counts are at parity, and login dropped from 4 turns to 2. Wall clock is not better in this sample: the new arm emits about 25% more output tokens per turn, and with only two rounds per cell the per-turn latency noise is large (the same 3-call form task took 25.6 s on one arm and 10.8 s on the other). gpt-5.6-sol did not use
answerin any run (it chosebrowsersnippets in 9 of 12 runs,batchin 3), so the one-call finish is exercised by the tests, not yet by this model in the wild. So: this PR removes the structural turn cost; it does not demonstrate a speedup with gpt-5.6-sol. Raw results:~/bw-bench/results-2on the author's machine.Checklist
bun run release:checkpasses locally (versions, lint, typecheck, build,unit tests, published declarations, tarball).
launch path, transport, or fetch.
rendered from step results that are already redacted (password readings
are
[redacted]), and the rendered result goes through the envelope'sredaction like any other field.
types/*.d.tsin this same commit.types/common.d.tsgainsAgentBatchOptions.answerandAgentBatchResult.finalAnswer/answerError.src/worker.tsordist/src/worker.jsdirectly.CHANGELOG.md,docs/agent-batch.md,README, docs/{agent,javascript,getting-started}.md, and the regenerated
SKILL.md.How it was verified
bun run release:checkand the managed-browser suite (BETTERWRIGHT_REQUIRE_BROWSER=1 bun run test) on Bun 1.4.0 / macOS arm64, including a new end-to-end check that a batch withanswerreturns the renderedfinalAnswer.finalAnswer.🤖 Generated with Claude Code
https://claude.ai/code/session_01Td3ZP4tb9UKEQHLaoAo6NN