Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .claude/skills/create-eval/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -178,4 +178,4 @@ For detailed documentation, consult:
- **`references/test-case-format.md`** — Complete test_case.yaml field reference with all options
- **`references/anti-hallucination.md`** — Anti-cheat testing patterns and prompt design
- **`references/infrastructure-patterns.md`** — Setup scripts, retry loops, port forwards, shared infra
- **`references/running-evals.md`** — CLI flags, environment variables, model comparison, debugging
- **`references/running-evals.md`** — CLI flags, environment variables, model comparison, debugging
Original file line number Diff line number Diff line change
Expand Up @@ -143,4 +143,4 @@ Before finalizing any eval, verify:
- [ ] User prompt uses business language, not implementation details
- [ ] Expected output checks specific discoverable values
- [ ] Test data looks realistic, not synthetic
- [ ] `expected_output` is invisible to the LLM (only the evaluator sees it)
- [ ] `expected_output` is invisible to the LLM (only the evaluator sees it)
2 changes: 1 addition & 1 deletion .claude/skills/create-eval/references/test-case-format.md
Original file line number Diff line number Diff line change
Expand Up @@ -189,4 +189,4 @@ before_test: |

after_test: |
kubectl delete namespace app-212 --ignore-not-found
```
```
12 changes: 9 additions & 3 deletions conftest.py
Original file line number Diff line number Diff line change
Expand Up @@ -121,14 +121,20 @@ def _patched_openai_init(self, *args, **kwargs):
if confluence_base and not os.environ.get("CONFLUENCE_SA_BASE_URL"):
parsed = urllib.parse.urlparse(confluence_base)
if parsed.scheme not in ("http", "https"):
logging.warning(f"CONFLUENCE_BASE_URL has unsupported scheme '{parsed.scheme}', skipping SA URL derivation")
logging.warning(
f"CONFLUENCE_BASE_URL has unsupported scheme '{parsed.scheme}', skipping SA URL derivation"
)
else:
try:
tenant_url = f"{confluence_base.rstrip('/')}/_edge/tenant_info"
with urllib.request.urlopen(tenant_url, timeout=10) as resp:
cloud_id = json.loads(resp.read())["cloudId"]
os.environ["CONFLUENCE_SA_BASE_URL"] = f"https://api.atlassian.com/ex/confluence/{cloud_id}"
logging.info(f"Auto-derived CONFLUENCE_SA_BASE_URL from cloud ID {cloud_id}")
os.environ["CONFLUENCE_SA_BASE_URL"] = (
f"https://api.atlassian.com/ex/confluence/{cloud_id}"
)
logging.info(
f"Auto-derived CONFLUENCE_SA_BASE_URL from cloud ID {cloud_id}"
)
except Exception as e:
logging.warning(f"Could not auto-derive CONFLUENCE_SA_BASE_URL: {e}")

Expand Down
1 change: 0 additions & 1 deletion docs/data-sources/builtin-toolsets/splunk-mcp.md
Original file line number Diff line number Diff line change
Expand Up @@ -161,4 +161,3 @@ holmes ask "Search Splunk for the most recent 10 error events"

- [Splunk MCP Server on Splunkbase](https://splunkbase.splunk.com/app/7931)
- [Splunk MCP Server Tools Reference](https://help.splunk.com/en/splunk-cloud-platform/mcp-server-for-splunk-platform/mcp-server-tools)

Original file line number Diff line number Diff line change
Expand Up @@ -133,4 +133,4 @@ Status of all evaluations across models. Color coding:
| [96_no_matching_runbook](https://github.com/HolmesGPT/holmesgpt/blob/master/tests/llm/fixtures/test_ask_holmes/96_no_matching_runbook/test_case.yaml) [🔗](https://www.braintrust.dev/app/robustadev/p/HolmesGPT/experiments/ci-benchmark-23093326433?c=&search=%7B%22filter%22%3A%20%5B%7B%22text%22%3A%20%22metadata.eval_id%2520%253D%2520%252296_no_matching_runbook%2522%22%2C%20%22label%22%3A%20%22metadata.eval_id%2520equals%252096_no_matching_runbook%22%2C%20%22originType%22%3A%20%22form%22%7D%5D%7D) | [🟡 80% (4/5)](https://www.braintrust.dev/app/robustadev/p/HolmesGPT/experiments/ci-benchmark-23093326433?c=&search=%7B%22filter%22%3A%20%5B%7B%22text%22%3A%20%22span_attributes.name%2520%253D%2520%252296_no_matching_runbook%255Bdeepseek-v3.2-chat%255D%2522%22%2C%20%22label%22%3A%20%22Name%2520equals%252096_no_matching_runbook%255Bdeepseek-v3.2-chat%255D%22%2C%20%22originType%22%3A%20%22form%22%7D%5D%7D) / ⏱️ 196.5s / 💰 $0.02 | [🟡 60% (3/5)](https://www.braintrust.dev/app/robustadev/p/HolmesGPT/experiments/ci-benchmark-23093326433?c=&search=%7B%22filter%22%3A%20%5B%7B%22text%22%3A%20%22span_attributes.name%2520%253D%2520%252296_no_matching_runbook%255Bgemini-3.1-pro-preview%255D%2522%22%2C%20%22label%22%3A%20%22Name%2520equals%252096_no_matching_runbook%255Bgemini-3.1-pro-preview%255D%22%2C%20%22originType%22%3A%20%22form%22%7D%5D%7D) / ⏱️ 37.4s / 💰 $0.16 | [🔴 0% (0/5)](https://www.braintrust.dev/app/robustadev/p/HolmesGPT/experiments/ci-benchmark-23093326433?c=&search=%7B%22filter%22%3A%20%5B%7B%22text%22%3A%20%22span_attributes.name%2520%253D%2520%252296_no_matching_runbook%255Bgpt-5.4%255D%2522%22%2C%20%22label%22%3A%20%22Name%2520equals%252096_no_matching_runbook%255Bgpt-5.4%255D%22%2C%20%22originType%22%3A%20%22form%22%7D%5D%7D) / ⏱️ 51.1s / 💰 $0.19 | [🟢 100% (5/5)](https://www.braintrust.dev/app/robustadev/p/HolmesGPT/experiments/ci-benchmark-23093326433?c=&search=%7B%22filter%22%3A%20%5B%7B%22text%22%3A%20%22span_attributes.name%2520%253D%2520%252296_no_matching_runbook%255Bopus-4.6%255D%2522%22%2C%20%22label%22%3A%20%22Name%2520equals%252096_no_matching_runbook%255Bopus-4.6%255D%22%2C%20%22originType%22%3A%20%22form%22%7D%5D%7D) / ⏱️ 63.6s / 💰 $0.45 | [🟢 100% (5/5)](https://www.braintrust.dev/app/robustadev/p/HolmesGPT/experiments/ci-benchmark-23093326433?c=&search=%7B%22filter%22%3A%20%5B%7B%22text%22%3A%20%22span_attributes.name%2520%253D%2520%252296_no_matching_runbook%255Bsonnet-4.6%255D%2522%22%2C%20%22label%22%3A%20%22Name%2520equals%252096_no_matching_runbook%255Bsonnet-4.6%255D%22%2C%20%22originType%22%3A%20%22form%22%7D%5D%7D) / ⏱️ 60.7s / 💰 $0.33 |

---
*Results are automatically generated and updated weekly. View full traces and detailed analysis in [Braintrust experiment: ci-benchmark-23093326433](https://www.braintrust.dev/app/robustadev/p/HolmesGPT/experiments/ci-benchmark-23093326433).*
*Results are automatically generated and updated weekly. View full traces and detailed analysis in [Braintrust experiment: ci-benchmark-23093326433](https://www.braintrust.dev/app/robustadev/p/HolmesGPT/experiments/ci-benchmark-23093326433).*
Original file line number Diff line number Diff line change
Expand Up @@ -123,4 +123,4 @@ Status of all evaluations across models. Color coding:
| [96_no_matching_runbook](https://github.com/HolmesGPT/holmesgpt/blob/master/tests/llm/fixtures/test_ask_holmes/96_no_matching_runbook/test_case.yaml) [🔗](https://www.braintrust.dev/app/robustadev/p/HolmesGPT/experiments/ci-benchmark-21238150400?c=&search=%7B%22filter%22%3A%20%5B%7B%22text%22%3A%20%22metadata.eval_id%2520%253D%2520%252296_no_matching_runbook%2522%22%2C%20%22label%22%3A%20%22metadata.eval_id%2520equals%252096_no_matching_runbook%22%2C%20%22originType%22%3A%20%22form%22%7D%5D%7D) | [🟢 100% (5/5)](https://www.braintrust.dev/app/robustadev/p/HolmesGPT/experiments/ci-benchmark-21238150400?c=&search=%7B%22filter%22%3A%20%5B%7B%22text%22%3A%20%22span_attributes.name%2520%253D%2520%252296_no_matching_runbook%255Bbedrock%252Feu.anthropic.claude-haiku-4-5-20251001-v1%253A0%255D%2522%22%2C%20%22label%22%3A%20%22Name%2520equals%252096_no_matching_runbook%255Bbedrock%252Feu.anthropic.claude-haiku-4-5-20251001-v1%253A0%255D%22%2C%20%22originType%22%3A%20%22form%22%7D%5D%7D) / ⏱️ 47.8s / 💰 $0.10 | [🟡 40% (2/5)](https://www.braintrust.dev/app/robustadev/p/HolmesGPT/experiments/ci-benchmark-21238150400?c=&search=%7B%22filter%22%3A%20%5B%7B%22text%22%3A%20%22span_attributes.name%2520%253D%2520%252296_no_matching_runbook%255Bbedrock%252Feu.anthropic.claude-opus-4-5-20251101-v1%253A0%255D%2522%22%2C%20%22label%22%3A%20%22Name%2520equals%252096_no_matching_runbook%255Bbedrock%252Feu.anthropic.claude-opus-4-5-20251101-v1%253A0%255D%22%2C%20%22originType%22%3A%20%22form%22%7D%5D%7D) / ⏱️ 61.0s / 💰 $0.35 | [🟢 100% (5/5)](https://www.braintrust.dev/app/robustadev/p/HolmesGPT/experiments/ci-benchmark-21238150400?c=&search=%7B%22filter%22%3A%20%5B%7B%22text%22%3A%20%22span_attributes.name%2520%253D%2520%252296_no_matching_runbook%255Bbedrock%252Feu.anthropic.claude-sonnet-4-5-20250929-v1%253A0%255D%2522%22%2C%20%22label%22%3A%20%22Name%2520equals%252096_no_matching_runbook%255Bbedrock%252Feu.anthropic.claude-sonnet-4-5-20250929-v1%253A0%255D%22%2C%20%22originType%22%3A%20%22form%22%7D%5D%7D) / ⏱️ 54.1s / 💰 $0.26 | [🟡 20% (1/5)](https://www.braintrust.dev/app/robustadev/p/HolmesGPT/experiments/ci-benchmark-21238150400?c=&search=%7B%22filter%22%3A%20%5B%7B%22text%22%3A%20%22span_attributes.name%2520%253D%2520%252296_no_matching_runbook%255Bgemini%252Fgemini-3-flash-preview%255D%2522%22%2C%20%22label%22%3A%20%22Name%2520equals%252096_no_matching_runbook%255Bgemini%252Fgemini-3-flash-preview%255D%22%2C%20%22originType%22%3A%20%22form%22%7D%5D%7D) / ⏱️ 198.1s / 💰 $0.10 | [🟡 40% (2/5)](https://www.braintrust.dev/app/robustadev/p/HolmesGPT/experiments/ci-benchmark-21238150400?c=&search=%7B%22filter%22%3A%20%5B%7B%22text%22%3A%20%22span_attributes.name%2520%253D%2520%252296_no_matching_runbook%255Bgemini%252Fgemini-3-pro-preview%255D%2522%22%2C%20%22label%22%3A%20%22Name%2520equals%252096_no_matching_runbook%255Bgemini%252Fgemini-3-pro-preview%255D%22%2C%20%22originType%22%3A%20%22form%22%7D%5D%7D) / ⏱️ 112.8s / 💰 $0.28 | [🟡 40% (2/5)](https://www.braintrust.dev/app/robustadev/p/HolmesGPT/experiments/ci-benchmark-21238150400?c=&search=%7B%22filter%22%3A%20%5B%7B%22text%22%3A%20%22span_attributes.name%2520%253D%2520%252296_no_matching_runbook%255Bopenai%252Fgpt-5.2%255D%2522%22%2C%20%22label%22%3A%20%22Name%2520equals%252096_no_matching_runbook%255Bopenai%252Fgpt-5.2%255D%22%2C%20%22originType%22%3A%20%22form%22%7D%5D%7D) / ⏱️ 41.9s / 💰 $0.12 |

---
*Results are automatically generated and updated weekly. View full traces and detailed analysis in [Braintrust experiment: ci-benchmark-21238150400](https://www.braintrust.dev/app/robustadev/p/HolmesGPT/experiments/ci-benchmark-21238150400).*
*Results are automatically generated and updated weekly. View full traces and detailed analysis in [Braintrust experiment: ci-benchmark-21238150400](https://www.braintrust.dev/app/robustadev/p/HolmesGPT/experiments/ci-benchmark-21238150400).*
Original file line number Diff line number Diff line change
Expand Up @@ -129,4 +129,4 @@ Status of all evaluations across models. Color coding:
| [96_no_matching_runbook](https://github.com/HolmesGPT/holmesgpt/blob/master/tests/llm/fixtures/test_ask_holmes/96_no_matching_runbook/test_case.yaml) [🔗](https://www.braintrust.dev/app/robustadev/p/HolmesGPT/experiments/ci-benchmark-21401358384?c=&search=%7B%22filter%22%3A%20%5B%7B%22text%22%3A%20%22metadata.eval_id%2520%253D%2520%252296_no_matching_runbook%2522%22%2C%20%22label%22%3A%20%22metadata.eval_id%2520equals%252096_no_matching_runbook%22%2C%20%22originType%22%3A%20%22form%22%7D%5D%7D) | [🟢 100% (1/1)](https://www.braintrust.dev/app/robustadev/p/HolmesGPT/experiments/ci-benchmark-21401358384?c=&search=%7B%22filter%22%3A%20%5B%7B%22text%22%3A%20%22span_attributes.name%2520%253D%2520%252296_no_matching_runbook%255Bdeepseek%252Fdeepseek-chat%255D%2522%22%2C%20%22label%22%3A%20%22Name%2520equals%252096_no_matching_runbook%255Bdeepseek%252Fdeepseek-chat%255D%22%2C%20%22originType%22%3A%20%22form%22%7D%5D%7D) / ⏱️ 160.5s / 💰 $0.03 | [🔴 0% (0/1)](https://www.braintrust.dev/app/robustadev/p/HolmesGPT/experiments/ci-benchmark-21401358384?c=&search=%7B%22filter%22%3A%20%5B%7B%22text%22%3A%20%22span_attributes.name%2520%253D%2520%252296_no_matching_runbook%255Bdeepseek%252Fdeepseek-reasoner%255D%2522%22%2C%20%22label%22%3A%20%22Name%2520equals%252096_no_matching_runbook%255Bdeepseek%252Fdeepseek-reasoner%255D%22%2C%20%22originType%22%3A%20%22form%22%7D%5D%7D) / ⏱️ 366.0s / 💰 $0.03 | [🟢 100% (1/1)](https://www.braintrust.dev/app/robustadev/p/HolmesGPT/experiments/ci-benchmark-21401358384?c=&search=%7B%22filter%22%3A%20%5B%7B%22text%22%3A%20%22span_attributes.name%2520%253D%2520%252296_no_matching_runbook%255Bbedrock%252Feu.anthropic.claude-haiku-4-5-20251001-v1%253A0%255D%2522%22%2C%20%22label%22%3A%20%22Name%2520equals%252096_no_matching_runbook%255Bbedrock%252Feu.anthropic.claude-haiku-4-5-20251001-v1%253A0%255D%22%2C%20%22originType%22%3A%20%22form%22%7D%5D%7D) / ⏱️ 52.2s / 💰 $0.11 | [🔴 0% (0/1)](https://www.braintrust.dev/app/robustadev/p/HolmesGPT/experiments/ci-benchmark-21401358384?c=&search=%7B%22filter%22%3A%20%5B%7B%22text%22%3A%20%22span_attributes.name%2520%253D%2520%252296_no_matching_runbook%255Bbedrock%252Feu.anthropic.claude-opus-4-5-20251101-v1%253A0%255D%2522%22%2C%20%22label%22%3A%20%22Name%2520equals%252096_no_matching_runbook%255Bbedrock%252Feu.anthropic.claude-opus-4-5-20251101-v1%253A0%255D%22%2C%20%22originType%22%3A%20%22form%22%7D%5D%7D) / ⏱️ 83.8s / 💰 $0.48 | [🟢 100% (1/1)](https://www.braintrust.dev/app/robustadev/p/HolmesGPT/experiments/ci-benchmark-21401358384?c=&search=%7B%22filter%22%3A%20%5B%7B%22text%22%3A%20%22span_attributes.name%2520%253D%2520%252296_no_matching_runbook%255Bbedrock%252Feu.anthropic.claude-sonnet-4-5-20250929-v1%253A0%255D%2522%22%2C%20%22label%22%3A%20%22Name%2520equals%252096_no_matching_runbook%255Bbedrock%252Feu.anthropic.claude-sonnet-4-5-20250929-v1%253A0%255D%22%2C%20%22originType%22%3A%20%22form%22%7D%5D%7D) / ⏱️ 57.7s / 💰 $0.29 | [🔴 0% (0/1)](https://www.braintrust.dev/app/robustadev/p/HolmesGPT/experiments/ci-benchmark-21401358384?c=&search=%7B%22filter%22%3A%20%5B%7B%22text%22%3A%20%22span_attributes.name%2520%253D%2520%252296_no_matching_runbook%255Bgemini%252Fgemini-3-flash-preview%255D%2522%22%2C%20%22label%22%3A%20%22Name%2520equals%252096_no_matching_runbook%255Bgemini%252Fgemini-3-flash-preview%255D%22%2C%20%22originType%22%3A%20%22form%22%7D%5D%7D) / ⏱️ 105.7s / 💰 $0.12 | [🟢 100% (1/1)](https://www.braintrust.dev/app/robustadev/p/HolmesGPT/experiments/ci-benchmark-21401358384?c=&search=%7B%22filter%22%3A%20%5B%7B%22text%22%3A%20%22span_attributes.name%2520%253D%2520%252296_no_matching_runbook%255Bgemini%252Fgemini-3-pro-preview%255D%2522%22%2C%20%22label%22%3A%20%22Name%2520equals%252096_no_matching_runbook%255Bgemini%252Fgemini-3-pro-preview%255D%22%2C%20%22originType%22%3A%20%22form%22%7D%5D%7D) / ⏱️ 883.4s / 💰 $0.29 | [🔴 0% (0/1)](https://www.braintrust.dev/app/robustadev/p/HolmesGPT/experiments/ci-benchmark-21401358384?c=&search=%7B%22filter%22%3A%20%5B%7B%22text%22%3A%20%22span_attributes.name%2520%253D%2520%252296_no_matching_runbook%255Bopenai%252Fgpt-5.2%255D%2522%22%2C%20%22label%22%3A%20%22Name%2520equals%252096_no_matching_runbook%255Bopenai%252Fgpt-5.2%255D%22%2C%20%22originType%22%3A%20%22form%22%7D%5D%7D) / ⏱️ 73.8s / 💰 $0.24 |

---
*Results are automatically generated and updated weekly. View full traces and detailed analysis in [Braintrust experiment: ci-benchmark-21401358384](https://www.braintrust.dev/app/robustadev/p/HolmesGPT/experiments/ci-benchmark-21401358384).*
*Results are automatically generated and updated weekly. View full traces and detailed analysis in [Braintrust experiment: ci-benchmark-21401358384](https://www.braintrust.dev/app/robustadev/p/HolmesGPT/experiments/ci-benchmark-21401358384).*
Original file line number Diff line number Diff line change
Expand Up @@ -131,4 +131,4 @@ Status of all evaluations across models. Color coding:
| [96_no_matching_runbook](https://github.com/HolmesGPT/holmesgpt/blob/master/tests/llm/fixtures/test_ask_holmes/96_no_matching_runbook/test_case.yaml) [🔗](https://www.braintrust.dev/app/robustadev/p/HolmesGPT/experiments/ci-benchmark-21471579810?c=&search=%7B%22filter%22%3A%20%5B%7B%22text%22%3A%20%22metadata.eval_id%2520%253D%2520%252296_no_matching_runbook%2522%22%2C%20%22label%22%3A%20%22metadata.eval_id%2520equals%252096_no_matching_runbook%22%2C%20%22originType%22%3A%20%22form%22%7D%5D%7D) | [🔴 0% (0/1)](https://www.braintrust.dev/app/robustadev/p/HolmesGPT/experiments/ci-benchmark-21471579810?c=&search=%7B%22filter%22%3A%20%5B%7B%22text%22%3A%20%22span_attributes.name%2520%253D%2520%252296_no_matching_runbook%255Bdeepseek-chat%255D%2522%22%2C%20%22label%22%3A%20%22Name%2520equals%252096_no_matching_runbook%255Bdeepseek-chat%255D%22%2C%20%22originType%22%3A%20%22form%22%7D%5D%7D) / ⏱️ 205.2s / 💰 $0.03 | [🟢 100% (1/1)](https://www.braintrust.dev/app/robustadev/p/HolmesGPT/experiments/ci-benchmark-21471579810?c=&search=%7B%22filter%22%3A%20%5B%7B%22text%22%3A%20%22span_attributes.name%2520%253D%2520%252296_no_matching_runbook%255Bdeepseek-reasoner%255D%2522%22%2C%20%22label%22%3A%20%22Name%2520equals%252096_no_matching_runbook%255Bdeepseek-reasoner%255D%22%2C%20%22originType%22%3A%20%22form%22%7D%5D%7D) / ⏱️ 422.5s / 💰 $0.03 | [🔴 0% (0/1)](https://www.braintrust.dev/app/robustadev/p/HolmesGPT/experiments/ci-benchmark-21471579810?c=&search=%7B%22filter%22%3A%20%5B%7B%22text%22%3A%20%22span_attributes.name%2520%253D%2520%252296_no_matching_runbook%255Bgemini-3-flash-preview%255D%2522%22%2C%20%22label%22%3A%20%22Name%2520equals%252096_no_matching_runbook%255Bgemini-3-flash-preview%255D%22%2C%20%22originType%22%3A%20%22form%22%7D%5D%7D) / ⏱️ 34.6s / 💰 $0.10 | [🟢 100% (1/1)](https://www.braintrust.dev/app/robustadev/p/HolmesGPT/experiments/ci-benchmark-21471579810?c=&search=%7B%22filter%22%3A%20%5B%7B%22text%22%3A%20%22span_attributes.name%2520%253D%2520%252296_no_matching_runbook%255Bgemini-3-pro-preview%255D%2522%22%2C%20%22label%22%3A%20%22Name%2520equals%252096_no_matching_runbook%255Bgemini-3-pro-preview%255D%22%2C%20%22originType%22%3A%20%22form%22%7D%5D%7D) / ⏱️ 117.1s / 💰 $0.25 | [🔴 0% (0/1)](https://www.braintrust.dev/app/robustadev/p/HolmesGPT/experiments/ci-benchmark-21471579810?c=&search=%7B%22filter%22%3A%20%5B%7B%22text%22%3A%20%22span_attributes.name%2520%253D%2520%252296_no_matching_runbook%255Bgpt-5.2-high-reasoning%255D%2522%22%2C%20%22label%22%3A%20%22Name%2520equals%252096_no_matching_runbook%255Bgpt-5.2-high-reasoning%255D%22%2C%20%22originType%22%3A%20%22form%22%7D%5D%7D) / ⏱️ 836.0s / 💰 $0.75 | [🟢 100% (1/1)](https://www.braintrust.dev/app/robustadev/p/HolmesGPT/experiments/ci-benchmark-21471579810?c=&search=%7B%22filter%22%3A%20%5B%7B%22text%22%3A%20%22span_attributes.name%2520%253D%2520%252296_no_matching_runbook%255Bhaiku-4.5%255D%2522%22%2C%20%22label%22%3A%20%22Name%2520equals%252096_no_matching_runbook%255Bhaiku-4.5%255D%22%2C%20%22originType%22%3A%20%22form%22%7D%5D%7D) / ⏱️ 56.3s / 💰 $0.10 | [🟢 100% (1/1)](https://www.braintrust.dev/app/robustadev/p/HolmesGPT/experiments/ci-benchmark-21471579810?c=&search=%7B%22filter%22%3A%20%5B%7B%22text%22%3A%20%22span_attributes.name%2520%253D%2520%252296_no_matching_runbook%255Bkimi-2.5-openrouter%255D%2522%22%2C%20%22label%22%3A%20%22Name%2520equals%252096_no_matching_runbook%255Bkimi-2.5-openrouter%255D%22%2C%20%22originType%22%3A%20%22form%22%7D%5D%7D) / ⏱️ 218.4s | [🔴 0% (0/1)](https://www.braintrust.dev/app/robustadev/p/HolmesGPT/experiments/ci-benchmark-21471579810?c=&search=%7B%22filter%22%3A%20%5B%7B%22text%22%3A%20%22span_attributes.name%2520%253D%2520%252296_no_matching_runbook%255Bopus-4.5%255D%2522%22%2C%20%22label%22%3A%20%22Name%2520equals%252096_no_matching_runbook%255Bopus-4.5%255D%22%2C%20%22originType%22%3A%20%22form%22%7D%5D%7D) / ⏱️ 62.5s / 💰 $0.37 | [🟢 100% (1/1)](https://www.braintrust.dev/app/robustadev/p/HolmesGPT/experiments/ci-benchmark-21471579810?c=&search=%7B%22filter%22%3A%20%5B%7B%22text%22%3A%20%22span_attributes.name%2520%253D%2520%252296_no_matching_runbook%255Bsonnet-4.5%255D%2522%22%2C%20%22label%22%3A%20%22Name%2520equals%252096_no_matching_runbook%255Bsonnet-4.5%255D%22%2C%20%22originType%22%3A%20%22form%22%7D%5D%7D) / ⏱️ 66.8s / 💰 $0.31 |

---
*Results are automatically generated and updated weekly. View full traces and detailed analysis in [Braintrust experiment: ci-benchmark-21471579810](https://www.braintrust.dev/app/robustadev/p/HolmesGPT/experiments/ci-benchmark-21471579810).*
*Results are automatically generated and updated weekly. View full traces and detailed analysis in [Braintrust experiment: ci-benchmark-21471579810](https://www.braintrust.dev/app/robustadev/p/HolmesGPT/experiments/ci-benchmark-21471579810).*
Loading
Loading