Skip to content

Add testing-agents problem document - #14

Merged
ralphbean merged 3 commits into
mainfrom
testing-agents
Mar 17, 2026
Merged

ralphbean merged 3 commits into
mainfrom
testing-agents

Conversation

@twaugh

@twaugh twaugh commented Mar 12, 2026

Copy link
Copy Markdown
Contributor

Explores how to verify agent behavior hasn't regressed when instructions change, covering golden-set evaluation, behavioral contracts, canary deployments, and mutation testing for natural-language instructions.

@twaugh
twaugh requested a review from a team as a code owner March 12, 2026 10:57
@ralphbean

Copy link
Copy Markdown
Member

Tim, this is great - thank you.

It may be an implementation detail, but do you think it would be worthwhile to mention tools like promptfoo, deepeval, or lightspeed-evaluation? If I understand them correctly, those are tools that can take a set of (input, expected output) pairs and then mutate them into dozens or hundreds of copies so that you can score how an agent behaves against a spectrum of inputs. CI could end up looking like establishing a minimum score that the agent instruction has to pass before we would let that merge.

twaugh and others added 2 commits March 17, 2026 14:40
Explores how to verify agent behavior hasn't regressed when instructions
change, covering golden-set evaluation, behavioral contracts, canary
deployments, and mutation testing for natural-language instructions.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Cover promptfoo, deepeval, and lightspeed-evaluation as existing
tooling that implements parts of the golden-set, contract testing,
and adversarial evaluation approaches. Note gaps: none handles
cross-agent composition, mutation testing for natural language,
or absence detection natively.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Describe how eval tools can expand a small seed set of (input,
expected output) pairs into a large eval corpus through mutation,
lowering the barrier to bootstrapping golden-set coverage. Add the
minimum-score CI gate pattern as the natural way to consume
expanded eval results, absorbing non-determinism while still
catching systematic regressions.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
@twaugh

twaugh commented Mar 17, 2026

Copy link
Copy Markdown
Contributor Author

If I understand them correctly, those are tools that can take a set of (input, expected output) pairs and then mutate them into dozens or hundreds of copies so that you can score how an agent behaves against a spectrum of inputs.

Great idea. Added!

@ralphbean ralphbean left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm super excited about this one.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants