-
Notifications
You must be signed in to change notification settings - Fork 52
Pull requests: strands-agents/evals
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
ci: update opentelemetry-instrumentation-langchain requirement from <0.62.0,>=0.40.0 to >=0.40.0,<0.63.0
area-community
Repo health, governance, contributor process, release process, and CI dependency bumps
chore
Maintenance tasks, dependency updates, CI changes, refactoring with no user-facing impact
#369
opened Aug 13, 2026 by
Unshure
Member
Loading…
ci: update langfuse requirement from <4,>=2.0.0 to >=2.0.0,<5
area-community
Repo health, governance, contributor process, release process, and CI dependency bumps
chore
Maintenance tasks, dependency updates, CI changes, refactoring with no user-facing impact
dependencies
Pull requests that update a dependency file
python
Pull requests that update python code
#367
opened Aug 13, 2026 by
dependabot
Bot
Loading…
feat: add OpenAI Agents SDK support to OpenInference mapper
area-tracing
Trace/session ingestion: providers, session mappers, extractors, telemetry/OTEL
enhancement
New feature or request
feat(mappers): add OpenAI Agents OTel session mapper
area-tracing
Trace/session ingestion: providers, session mappers, extractors, telemetry/OTEL
enhancement
New feature or request
#365
opened Aug 11, 2026 by
liramon2
Contributor
Loading…
9 tasks done
feat: add ToolEfficiencyEvaluator for session-level tool usage analysis
area-evaluators
Evaluators: output, trajectory, tool use, interactions, and LLM-as-judge quality metrics
enhancement
New feature or request
#362
opened Aug 10, 2026 by
max-rattray-aws
Contributor
Loading…
feat: add evaluator metadata method and types
area-devx
Developer experience: papercuts, confusing public APIs, error messages, ergonomics, usability
area-evaluators
Evaluators: output, trajectory, tool use, interactions, and LLM-as-judge quality metrics
enhancement
New feature or request
#361
opened Aug 10, 2026 by
max-rattray-aws
Contributor
Loading…
feat: add failure cohort analysis for evaluation reports
area-core
Core eval framework: Case, Experiment, task handler, evaluation data stores
area-evaluators
Evaluators: output, trajectory, tool use, interactions, and LLM-as-judge quality metrics
enhancement
New feature or request
#360
opened Aug 10, 2026 by
max-rattray-aws
Contributor
Loading…
feat: add status field to EvaluationOutput
area-core
Core eval framework: Case, Experiment, task handler, evaluation data stores
area-evaluators
Evaluators: output, trajectory, tool use, interactions, and LLM-as-judge quality metrics
enhancement
New feature or request
#359
opened Aug 10, 2026 by
max-rattray-aws
Contributor
Loading…
feat: add inclusive language evaluator
area-evaluators
Evaluators: output, trajectory, tool use, interactions, and LLM-as-judge quality metrics
enhancement
New feature or request
#358
opened Aug 10, 2026 by
max-rattray-aws
Contributor
Loading…
feat: add could-not-evaluate status so non-gradable cases are excluded from scores
area-core
Core eval framework: Case, Experiment, task handler, evaluation data stores
area-evaluators
Evaluators: output, trajectory, tool use, interactions, and LLM-as-judge quality metrics
enhancement
New feature or request
#357
opened Aug 10, 2026 by
pdebjyot
Contributor
Loading…
fix: CorrectnessEvaluator honors expected_output for reference-mode grading
area-evaluators
Evaluators: output, trajectory, tool use, interactions, and LLM-as-judge quality metrics
bug
Something isn't working
#356
opened Aug 8, 2026 by
AmirF194
Loading…
8 of 9 tasks
feat(tools): TraceIndex — progressive trace disclosure for judge-based evaluators
area-evaluators
Evaluators: output, trajectory, tool use, interactions, and LLM-as-judge quality metrics
area-tracing
Trace/session ingestion: providers, session mappers, extractors, telemetry/OTEL
enhancement
New feature or request
#343
opened Aug 3, 2026 by
pdebjyot
Contributor
Loading…
5 tasks done
feat: add skill-level evaluators for skill-equipped agents
area-evaluators
Evaluators: output, trajectory, tool use, interactions, and LLM-as-judge quality metrics
enhancement
New feature or request
#330
opened Jul 28, 2026 by
sangminwoo
Collaborator
Loading…
9 tasks done
feat(experiment): add opt-in model routing for evaluators
area-core
Core eval framework: Case, Experiment, task handler, evaluation data stores
area-evaluators
Evaluators: output, trajectory, tool use, interactions, and LLM-as-judge quality metrics
enhancement
New feature or request
#325
opened Jul 23, 2026 by
pdebjyot
Contributor
Loading…
4 tasks done
fix: convert Langfuse AGENT observations regardless of parent
area-tracing
Trace/session ingestion: providers, session mappers, extractors, telemetry/OTEL
bug
Something isn't working
#314
opened Jul 15, 2026 by
poshinchen
Contributor
Loading…
7 of 9 tasks
test(redteam): cover RedTeamCase metadata sync and _build_system_prompt assembly
area-redteam
Red teaming: adversarial generation, attack strategies, attack success evaluation
chore
Maintenance tasks, dependency updates, CI changes, refactoring with no user-facing impact
#301
opened Jul 7, 2026 by
zeroshotmind
Loading…
9 of 11 tasks
feat(redteam): reframe attacker prompts to survive aligned attacker models
area-redteam
Red teaming: adversarial generation, attack strategies, attack success evaluation
enhancement
New feature or request
#298
opened Jul 2, 2026 by
kevmyung
Contributor
Loading…
6 of 9 tasks
fix(redteam): separate errored attacks from breaches in ASR
area-redteam
Red teaming: adversarial generation, attack strategies, attack success evaluation
bug
Something isn't working
#296
opened Jul 2, 2026 by
kevmyung
Contributor
Loading…
7 of 9 tasks
feat(redteam): redesign risk categories for agent-centric evaluation
area-redteam
Red teaming: adversarial generation, attack strategies, attack success evaluation
enhancement
New feature or request
#290
opened Jul 1, 2026 by
kevmyung
Contributor
Loading…
9 tasks done
ci: update opentelemetry-instrumentation-langchain requirement from <0.62.0,>=0.40.0 to >=0.40.0,<0.63.0
area-community
Repo health, governance, contributor process, release process, and CI dependency bumps
chore
Maintenance tasks, dependency updates, CI changes, refactoring with no user-facing impact
dependencies
Pull requests that update a dependency file
python
Pull requests that update python code
#288
opened Jun 29, 2026 by
dependabot
Bot
Loading…
feat: add model-output corruption effects
area-chaos
Chaos/fault injection: experiments, recovery strategy, partial completion, failure communication
enhancement
New feature or request
#284
opened Jun 24, 2026 by
venkatkrish543re
Loading…
1 of 9 tasks
chore(detectors): added async detectors execution (WIP)
area-detectors
Failure detection and root cause analysis of agent sessions
chore
Maintenance tasks, dependency updates, CI changes, refactoring with no user-facing impact
#277
opened Jun 17, 2026 by
poshinchen
Contributor
Loading…
7 of 9 tasks
feat(experiment): add verbosity-aware stack traces to evaluation error reasons
area-core
Core eval framework: Case, Experiment, task handler, evaluation data stores
area-devx
Developer experience: papercuts, confusing public APIs, error messages, ergonomics, usability
enhancement
New feature or request
#268
opened Jun 15, 2026 by
AndyMc629
Loading…
9 tasks done
Feature/skills aggregator: add SkillEvalAggregator for batch evaluation comparisons
area-evaluators
Evaluators: output, trajectory, tool use, interactions, and LLM-as-judge quality metrics
enhancement
New feature or request
#229
opened May 14, 2026 by
venkatkrish543re
•
Draft
feat: Add EvaluationPlugin for agent invocation evaluation and retry
area-core
Core eval framework: Case, Experiment, task handler, evaluation data stores
area-evaluators
Evaluators: output, trajectory, tool use, interactions, and LLM-as-judge quality metrics
enhancement
New feature or request
#166
opened Mar 18, 2026 by
afarntrog
Contributor
Loading…
5 of 7 tasks
Previous Next
ProTip!
What’s not been updated in a month: updated:<2026-07-13.