Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
71 changes: 67 additions & 4 deletions docs/development/evals/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -64,6 +64,34 @@ RUN_LIVE=true MODEL=azure/your-deployment-name CLASSIFIER_MODEL=azure/your-deplo
- For any model provider, ensure you have the necessary API keys and environment variables set (e.g., `OPENAI_API_KEY`, `ANTHROPIC_API_KEY`, `AZURE_API_KEY`)
- The model specified here is passed directly to LiteLLM, so any model supported by LiteLLM can be used

### Multi-Model Benchmarking

HolmesGPT supports running evaluations across multiple models simultaneously to compare their performance:

```bash
# Test multiple models in a single run
# Models are specified as comma-separated list
RUN_LIVE=true MODEL=gpt-4o,anthropic/claude-3-5-sonnet-20241022,gpt-4o-mini \
CLASSIFIER_MODEL=gpt-4o \
poetry run pytest -m 'llm and easy' --no-cov

# Run with multiple iterations for statistically significant results
RUN_LIVE=true ITERATIONS=10 \
MODEL=gpt-4o,anthropic/claude-3-5-sonnet-20241022 \
CLASSIFIER_MODEL=gpt-4o \
poetry run pytest -m 'llm and easy' -n 10

# Test specific scenario across models
RUN_LIVE=true MODEL=gpt-4o,gpt-4o-mini \
poetry run pytest tests/llm/test_ask_holmes.py -k "01_how_many_pods"
```

When running multi-model benchmarks:
- Results will show a **Model Comparison Table** with side-by-side performance metrics
- Each model's pass rate, execution times, and P90 percentiles are displayed
- Tests are parameterized by model, so you'll see separate results for each model/test combination
- Use `CLASSIFIER_MODEL` to ensure consistent scoring across all models

### Running Evals with Multiple Iterations

LLMs are non-deterministic - they produce different outputs for the same input. **10 iterations is a good rule of thumb** for reliable results.
Expand Down Expand Up @@ -147,18 +175,53 @@ RUN_LIVE=true pytest -k "test" --skip-setup

## Model Comparison Workflow

Track performance across different models:
### Recommended: Multi-Model Testing (Single Run)

**Use the `MODEL` environment variable to test multiple models in a single run:**

```bash
# Compare multiple models simultaneously - RECOMMENDED approach
RUN_LIVE=true ITERATIONS=10 \
MODEL=gpt-4o,anthropic/claude-3-5-sonnet-20241022,gpt-4o-mini \
CLASSIFIER_MODEL=gpt-4o \
poetry run pytest -m 'llm and easy' -n 10

# This will generate a comparison table showing:
# - Side-by-side pass rates for each model
# - Execution time comparisons
# - Cost comparisons
# - Best performing models summary
```

### Alternative: Single-Model Testing (Separate Runs)

For cases where you need separate experiments or different configurations per model:

```bash
# Run separate experiments for each model
# Useful when you need different settings or want to track experiments separately

# 1. Baseline with GPT-4
RUN_LIVE=true ITERATIONS=10 EXPERIMENT_ID=baseline_gpt4o MODEL=gpt-4o pytest -n 10 tests/llm/

# 2. Compare with Claude
RUN_LIVE=true ITERATIONS=10 EXPERIMENT_ID=claude35 MODEL=anthropic/claude-3-5-sonnet CLASSIFIER_MODEL=gpt-4o pytest -n 10 tests/llm/
# 2. Compare with Claude (using GPT-4 as classifier since Anthropic models can't classify)
RUN_LIVE=true ITERATIONS=10 EXPERIMENT_ID=claude35 MODEL=anthropic/claude-3-5-sonnet-20241022 CLASSIFIER_MODEL=gpt-4o pytest -n 10 tests/llm/

# 3. Test a smaller model
RUN_LIVE=true ITERATIONS=10 EXPERIMENT_ID=gpt4o_mini MODEL=gpt-4o-mini pytest -n 10 tests/llm/
```

# 3. Results will be tracked if both BRAINTRUST_API_KEY and BRAINTRUST_ORG are set
### Braintrust Integration

Results are automatically tracked if Braintrust is configured:

```bash
# Set these once in your environment
export BRAINTRUST_API_KEY=your-key
export BRAINTRUST_ORG=your-org

# Then run any evaluation command - results will be tracked automatically
RUN_LIVE=true MODEL=gpt-4o,anthropic/claude-3-5-sonnet-20241022 pytest -m 'llm and easy'
```

## Test Markers
Expand Down
114 changes: 104 additions & 10 deletions holmes/core/tool_calling_llm.py
Original file line number Diff line number Diff line change
Expand Up @@ -10,7 +10,7 @@
from openai.types.chat.chat_completion_message_tool_call import (
ChatCompletionMessageToolCall,
)
from pydantic import BaseModel
from pydantic import BaseModel, Field
from rich.console import Console

from holmes.common.env_vars import TEMPERATURE, MAX_OUTPUT_TOKEN_RESERVATION
Expand Down Expand Up @@ -43,6 +43,78 @@
get_todo_manager,
)

# Create a named logger for cost tracking
cost_logger = logging.getLogger("holmes.costs")


class LLMCosts(BaseModel):
"""Tracks cost and token usage for LLM calls."""

total_cost: float = 0.0
total_tokens: int = 0
prompt_tokens: int = 0
completion_tokens: int = 0


def _extract_cost_from_response(full_response) -> float:
"""Extract cost value from LLM response.

Args:
full_response: The raw LLM response object

Returns:
The cost as a float, or 0.0 if not available
"""
try:
cost_value = (
full_response._hidden_params.get("response_cost", 0)
if hasattr(full_response, "_hidden_params")
else 0
)
# Ensure cost is a float
return float(cost_value) if cost_value is not None else 0.0
except Exception:
return 0.0


def _process_cost_info(
full_response, costs: Optional[LLMCosts] = None, log_prefix: str = "LLM call"
) -> None:
"""Process cost and token information from LLM response.

Logs the cost information and optionally accumulates it into a costs object.

Args:
full_response: The raw LLM response object
costs: Optional LLMCosts object to accumulate costs into
log_prefix: Prefix for logging messages (e.g., "LLM call", "Post-processing")
"""
try:
cost = _extract_cost_from_response(full_response)
usage = getattr(full_response, "usage", {})

if usage:
prompt_toks = usage.get("prompt_tokens", 0)
completion_toks = usage.get("completion_tokens", 0)
total_toks = usage.get("total_tokens", 0)
cost_logger.debug(
f"{log_prefix} cost: ${cost:.6f} | Tokens: {prompt_toks} prompt + {completion_toks} completion = {total_toks} total"
)
# Accumulate costs and tokens if costs object provided
if costs:
costs.total_cost += cost
costs.prompt_tokens += prompt_toks
costs.completion_tokens += completion_toks
costs.total_tokens += total_toks
elif cost > 0:
cost_logger.debug(
f"{log_prefix} cost: ${cost:.6f} | Token usage not available"
)
if costs:
costs.total_cost += cost
except Exception as e:
logging.debug(f"Could not extract cost information: {e}")


def format_tool_result_data(tool_result: StructuredToolResult) -> str:
tool_response = tool_result.data
Expand Down Expand Up @@ -186,11 +258,11 @@ def as_streaming_tool_result_response(self):
}


class LLMResult(BaseModel):
class LLMResult(LLMCosts):
tool_calls: Optional[List[ToolCallResult]] = None
result: Optional[str] = None
unprocessed_result: Optional[str] = None
instructions: List[str] = []
instructions: List[str] = Field(default_factory=list)
# TODO: clean up these two
prompt: Optional[str] = None
messages: Optional[List[dict]] = None
Expand Down Expand Up @@ -259,6 +331,8 @@ def call( # type: ignore
) -> LLMResult:
perf_timing = PerformanceTiming("tool_calling_llm.call")
tool_calls = [] # type: ignore
costs = LLMCosts()

tools = self.tool_executor.get_all_tools_openai_format(
target_model=self.llm.model
)
Expand Down Expand Up @@ -299,6 +373,9 @@ def call( # type: ignore
)
logging.debug(f"got response {full_response.to_json()}") # type: ignore

# Extract and accumulate cost information
_process_cost_info(full_response, costs, "LLM call")

perf_timing.measure("llm.completion")
# catch a known error that occurs with Azure and replace the error message with something more obvious to the user
except BadRequestError as e:
Expand Down Expand Up @@ -352,11 +429,14 @@ def call( # type: ignore
if post_process_prompt and user_prompt:
logging.info("Running post processing on investigation.")
raw_response = text_response
post_processed_response = self._post_processing_call(
prompt=user_prompt,
investigation=raw_response,
user_prompt=post_process_prompt,
post_processed_response, post_processing_cost = (
self._post_processing_call(
prompt=user_prompt,
investigation=raw_response,
user_prompt=post_process_prompt,
)
)
costs.total_cost += post_processing_cost

Comment thread
aantn marked this conversation as resolved.
perf_timing.end(f"- completed in {i} iterations -")
return LLMResult(
Expand All @@ -365,6 +445,7 @@ def call( # type: ignore
tool_calls=tool_calls,
prompt=json.dumps(messages, indent=2),
messages=messages,
**costs.model_dump(), # Include all cost fields
)

perf_timing.end(f"- completed in {i} iterations -")
Expand All @@ -373,6 +454,7 @@ def call( # type: ignore
tool_calls=tool_calls,
prompt=json.dumps(messages, indent=2),
messages=messages,
**costs.model_dump(), # Include all cost fields
)

if text_response and text_response.strip():
Expand Down Expand Up @@ -543,7 +625,7 @@ def _post_processing_call(
investigation,
user_prompt: Optional[str] = None,
system_prompt: str = "You are an AI assistant summarizing Kubernetes issues.",
) -> Optional[str]:
) -> tuple[Optional[str], float]:
try:
user_prompt = ToolCallingLLM.__load_post_processing_user_prompt(
prompt, investigation, user_prompt
Expand All @@ -562,10 +644,18 @@ def _post_processing_call(
]
full_response = self.llm.completion(messages=messages, temperature=0)
logging.debug(f"Post processing response {full_response}")
return full_response.choices[0].message.content # type: ignore

# Extract and log cost information for post-processing
post_processing_cost = _extract_cost_from_response(full_response)
if post_processing_cost > 0:
cost_logger.debug(
f"Post-processing LLM cost: ${post_processing_cost:.6f}"
)

return full_response.choices[0].message.content, post_processing_cost # type: ignore
except Exception:
logging.exception("Failed to run post processing", exc_info=True)
return investigation
return investigation, 0.0

@sentry_sdk.trace
def truncate_messages_to_fit_context(
Expand Down Expand Up @@ -638,6 +728,10 @@ def call_stream(
stream=False,
drop_params=True,
)

# Log cost information for this iteration (no accumulation in streaming)
_process_cost_info(full_response, log_prefix="LLM iteration")

perf_timing.measure("llm.completion")
# catch a known error that occurs with Azure and replace the error message with something more obvious to the user
except BadRequestError as e:
Expand Down
2 changes: 1 addition & 1 deletion holmes/core/tracing.py
Original file line number Diff line number Diff line change
Expand Up @@ -120,7 +120,7 @@ def __exit__(self, exc_type, exc_val, exc_tb):
class DummyTracer:
"""A no-op tracer implementation for when tracing is disabled."""

def start_experiment(self, experiment_name=None, metadata=None):
def start_experiment(self, experiment_name=None, additional_metadata=None):
"""No-op experiment creation."""
return None

Comment thread
aantn marked this conversation as resolved.
Expand Down
8 changes: 7 additions & 1 deletion holmes/main.py
Original file line number Diff line number Diff line change
Expand Up @@ -104,6 +104,11 @@
"-v",
help="Verbose output. You can pass multiple times to increase the verbosity. e.g. -v or -vv or -vvv",
)
opt_log_costs: bool = typer.Option(
False,
"--log-costs",
help="Show LLM cost information in the output",
)
opt_echo_request: bool = typer.Option(
True,
"--echo/--no-echo",
Expand Down Expand Up @@ -176,6 +181,7 @@ def ask(
custom_toolsets: Optional[List[Path]] = opt_custom_toolsets,
max_steps: Optional[int] = opt_max_steps,
verbose: Optional[List[bool]] = opt_verbose,
log_costs: bool = opt_log_costs,
# semi-common options
destination: Optional[DestinationType] = opt_destination,
slack_token: Optional[str] = opt_slack_token,
Expand Down Expand Up @@ -219,7 +225,7 @@ def ask(
"""
Ask any question and answer using available tools
"""
console = init_logging(verbose) # type: ignore
console = init_logging(verbose, log_costs) # type: ignore
# Detect and read piped input
piped_data = None

Expand Down
7 changes: 6 additions & 1 deletion holmes/utils/console/logging.py
Original file line number Diff line number Diff line change
Expand Up @@ -41,9 +41,14 @@ def suppress_noisy_logs():
warnings.filterwarnings("ignore", category=UserWarning, module="slack_sdk.*")


def init_logging(verbose_flags: Optional[List[bool]] = None):
def init_logging(verbose_flags: Optional[List[bool]] = None, log_costs: bool = False):
verbosity = cli_flags_to_verbosity(verbose_flags) # type: ignore

# Setup cost logger if requested
if log_costs:
cost_logger = logging.getLogger("holmes.costs")
cost_logger.setLevel(logging.DEBUG)

if verbosity == Verbosity.VERY_VERBOSE:
logging.basicConfig(
force=True,
Expand Down
Loading
Loading