Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
25 commits
Select commit Hold shift + click to select a range
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 0 additions & 1 deletion .github/workflows/llm-evaluation.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -43,7 +43,6 @@ jobs:
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
BRAINTRUST_API_KEY: ${{ secrets.BRAINTRUST_API_KEY }}
UPLOAD_DATASET: "true"
PUSH_EVALS_TO_BRAINTRUST: "true"
EXPERIMENT_ID: github-${{ github.run_id }}.${{ github.run_number }}.${{ github.run_attempt }}
run: |
poetry run pytest --no-cov tests/llm/test_ask_holmes.py tests/llm/test_investigate.py -n 6
Expand Down
3 changes: 2 additions & 1 deletion CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -137,7 +137,8 @@ poetry run pytest tests/llm/test_ask_holmes.py
- `OPENAI_API_KEY`, `ANTHROPIC_API_KEY`: LLM API keys
- `MODEL`: Override default model
- `RUN_LIVE`: Use live tools in tests instead of mocks
- `BRAINTRUST_API_KEY`: For test result tracking
- `BRAINTRUST_API_KEY`: For test result tracking and CI/CD report generation
- `BRAINTRUST_ORG`: Braintrust organization name (default: "robustadev")

## Development Guidelines

Expand Down
2 changes: 2 additions & 0 deletions docs/development/evals/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -100,6 +100,8 @@ Configure evaluations using these environment variables:
| `RUN_LIVE` | `RUN_LIVE=true` | Execute `before-test` and `after-test` commands, ignore mock files |
| `UPLOAD_DATASET` | `UPLOAD_DATASET=true` | Sync dataset to external evaluation platform |
| `EXPERIMENT_ID` | `EXPERIMENT_ID=my_baseline` | Custom experiment name for result tracking |
| `BRAINTRUST_API_KEY` | `BRAINTRUST_API_KEY=sk-...` | Enable Braintrust integration for result tracking and CI/CD report generation |
| `BRAINTRUST_ORG` | `BRAINTRUST_ORG=my-org` | Braintrust organization name (defaults to "robustadev") |

### Simple Example

Expand Down
2 changes: 0 additions & 2 deletions docs/development/evals/reporting.md
Original file line number Diff line number Diff line change
Expand Up @@ -40,7 +40,6 @@ export BRAINTRUST_API_KEY=sk-your-api-key-here
```bash
export BRAINTRUST_API_KEY=sk-your-key
export UPLOAD_DATASET=true
export PUSH_EVALS_TO_BRAINTRUST=true

pytest ./tests/llm/test_ask_holmes.py
```
Expand All @@ -58,7 +57,6 @@ pytest -n 10 ./tests/llm/test_*.py
| Variable | Purpose |
|----------|---------|
| `UPLOAD_DATASET` | Sync test cases to Braintrust |
| `PUSH_EVALS_TO_BRAINTRUST` | Upload evaluation results |
| `EXPERIMENT_ID` | Name your experiment run. This makes it easier to find and track in Braintrust's UI |
| `MODEL` | The LLM model for Holmes to use |
| `CLASSIFIER_MODEL` | The LLM model to use for scoring the answer (LLM as judge) |
Expand Down
Loading