Add testing-agents problem document - #14
Conversation
|
Tim, this is great - thank you. It may be an implementation detail, but do you think it would be worthwhile to mention tools like |
Explores how to verify agent behavior hasn't regressed when instructions change, covering golden-set evaluation, behavioral contracts, canary deployments, and mutation testing for natural-language instructions. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Cover promptfoo, deepeval, and lightspeed-evaluation as existing tooling that implements parts of the golden-set, contract testing, and adversarial evaluation approaches. Note gaps: none handles cross-agent composition, mutation testing for natural language, or absence detection natively. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Describe how eval tools can expand a small seed set of (input, expected output) pairs into a large eval corpus through mutation, lowering the barrier to bootstrapping golden-set coverage. Add the minimum-score CI gate pattern as the natural way to consume expanded eval results, absorbing non-determinism while still catching systematic regressions. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Great idea. Added! |
ralphbean
left a comment
There was a problem hiding this comment.
I'm super excited about this one.
Explores how to verify agent behavior hasn't regressed when instructions change, covering golden-set evaluation, behavioral contracts, canary deployments, and mutation testing for natural-language instructions.