Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
99 changes: 1 addition & 98 deletions docs/reference/http-api.md
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
# HolmesGPT API Reference

## Overview
The HolmesGPT API provides endpoints for automated investigations, workload health checks, and conversational troubleshooting. This document describes each endpoint, its purpose, request fields, and example usage.
The HolmesGPT API provides endpoints for automated investigations and conversational troubleshooting. This document describes each endpoint, its purpose, request fields, and example usage.

## Model Parameter Behavior

Expand Down Expand Up @@ -391,103 +391,6 @@ curl -X POST http://<HOLMES-URL>/api/issue_chat \

---

### `/api/workload_health_check` (POST)
**Description:** Performs a health check on a specified workload (e.g., a Kubernetes deployment).

#### Request Fields

| Field | Required | Default | Type | Description |
|-------------------------|----------|--------------------------------------------|-----------|--------------------------------------------------|
| ask | Yes | | string | User's question |
| resource | Yes | | object | Resource details (e.g., name, kind) |
| alert_history_since_hours| No | 24 | float | How many hours back to check alerts |
| alert_history | No | true | boolean | Whether to include alert history |
| stored_instructions | No | true | boolean | Use stored instructions |
| instructions | No | [] | list | Additional instructions |
| include_tool_calls | No | false | boolean | Include tool calls in response |
| include_tool_call_results| No | false | boolean | Include tool call results in response |
| prompt_template | No | "builtin://kubernetes_workload_ask.jinja2" | string | Prompt template to use |
| model | No | | string | Model name from your `modelList` configuration |

**Example**
```bash
curl -X POST http://<HOLMES-URL>/api/workload_health_check \
-H "Content-Type: application/json" \
-d '{
"ask": "Why is my deployment unhealthy?",
"resource": {"name": "my-deployment", "kind": "Deployment"},
"alert_history_since_hours": 12
}'
```

**Example** Response
```json
{
"analysis": "Deployment 'my-deployment' is unhealthy due to repeated CrashLoopBackOff events.",
"sections": null,
"tool_calls": [
{
"tool_call_id": "2",
"tool_name": "kubectl_get_events",
"description": "Fetch recent events",
"result": {"events": "..."}
}
],
"instructions": [...]
}
```

---

### `/api/workload_health_chat` (POST)
**Description:** Conversational interface for discussing the health of a workload.

#### Request Fields

| Field | Required | Default | Type | Description |
|-------------------------|----------|---------|-----------|--------------------------------------------------|
| ask | Yes | | string | User's question |
| workload_health_result | Yes | | object | Previous health check result (see below) |
| resource | Yes | | object | Resource details |
| conversation_history | No | | list | Conversation history (first message must be system)|
| model | No | | string | Model name from your `modelList` configuration |

**workload_health_result** object:
- `analysis` (string, optional): Previous analysis
- `tools` (list, optional): Tools used/results

**Example**
```bash
curl -X POST http://<HOLMES-URL>/api/workload_health_chat \
-H "Content-Type: application/json" \
-d '{
"ask": "Check the workload health.",
"workload_health_result": {
"analysis": "Previous health check: all good.",
"tools": []
},
"resource": {"name": "my-deployment", "kind": "Deployment"},
"conversation_history": [
{"role": "system", "content": "You are a helpful assistant."}
]
}'
```

**Example** Response
```json
{
"analysis": "The deployment 'my-deployment' is healthy. No recent issues detected.",
"conversation_history": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Check the workload health."},
{"role": "assistant", "content": "The deployment 'my-deployment' is healthy. No recent issues detected."}
],
"tool_calls": [...]
}
```

---

### `/api/model` (GET)
**Description:** Returns a list of available AI models that can be used for investigations and chat.

Expand Down
222 changes: 0 additions & 222 deletions holmes/core/conversations.py
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,6 @@
from holmes.core.models import (
IssueChatRequest,
ToolCallConversationResult,
WorkloadHealthChatRequest,
)
from holmes.core.prompt import generate_user_prompt
from holmes.core.tool_calling_llm import ToolCallingLLM
Expand Down Expand Up @@ -480,224 +479,3 @@ def build_chat_messages(
)
truncate_tool_messages(conversation_history, tool_size) # type: ignore
return conversation_history # type: ignore


def build_workload_health_chat_messages(
workload_health_chat_request: WorkloadHealthChatRequest,
ai: ToolCallingLLM,
config: Config,
global_instructions: Optional[Instructions] = None,
runbooks: Optional[RunbookCatalog] = None,
):
"""
This function generates a list of messages for workload health conversation and ensures that the message sequence adheres to the model's context window limitations
by truncating tool outputs as necessary before sending to llm.

We always expect conversation_history to be passed in the openAI format which is supported by litellm and passed back by us.
That's why we assume that first message in the conversation is system message and truncate tools for it.

System prompt handling:
1. For new conversations (empty conversation_history):
- Creates a new system prompt using kubernetes_workload_chat.jinja2 template
- Includes workload analysis, tools (if any), and resource information
- If there are tools, calculates appropriate tool size and truncates tool outputs

2. For existing conversations:
- Preserves the conversation history
- Updates the first message (system prompt) with recalculated content
- Truncates tool outputs if necessary to fit context window
- Maintains the original conversation flow while ensuring context limits

Example structure of conversation history:
conversation_history = [
# System prompt with workload analysis
{"role": "system", "content": "...."},
# User message asking about workload health
{"role": "user", "content": "What's the current health status of my deployment?"},
# Assistant initiates a tool call
{
"role": "assistant",
"content": None,
"tool_call": {
"name": "check_workload_metrics",
"arguments": "{\"namespace\": \"default\", \"workload\": \"my-deployment\"}"
}
},
# Tool/Function response
{
"role": "tool",
"name": "check_workload_metrics",
"content": "{\"cpu_usage\": \"45%\", \"memory_usage\": \"60%\", \"status\": \"Running\"}"
},
# Assistant's final response to the user
{
"role": "assistant",
"content": "Your deployment is running normally with CPU usage at 45% and memory usage at 60%."
},
]
"""

template_path = "builtin://kubernetes_workload_chat.jinja2"

conversation_history = workload_health_chat_request.conversation_history
user_prompt = workload_health_chat_request.ask
workload_analysis = workload_health_chat_request.workload_health_result.analysis
tools_for_workload = workload_health_chat_request.workload_health_result.tools
resource = workload_health_chat_request.resource

if not conversation_history or len(conversation_history) == 0:
runbooks_ctx = generate_runbooks_args(
runbook_catalog=runbooks,
global_instructions=global_instructions,
)
user_prompt = generate_user_prompt(
user_prompt,
runbooks_ctx,
)

number_of_tools_for_workload = len(tools_for_workload) # type: ignore
if number_of_tools_for_workload == 0:
system_prompt = load_and_render_prompt(
template_path,
{
"workload_analysis": workload_analysis,
"tools_called_for_workload": tools_for_workload,
"resource": resource,
"toolsets": ai.tool_executor.toolsets,
"cluster_name": config.cluster_name,
"runbooks_enabled": True if runbooks else False,
},
)
messages = [
{
"role": "system",
"content": system_prompt,
},
{
"role": "user",
"content": user_prompt,
},
]
return messages

template_context_without_tools = {
"workload_analysis": workload_analysis,
"tools_called_for_workload": None,
"resource": resource,
"toolsets": ai.tool_executor.toolsets,
"cluster_name": config.cluster_name,
"runbooks_enabled": True if runbooks else False,
}
system_prompt_without_tools = load_and_render_prompt(
template_path, template_context_without_tools
)
messages_without_tools = [
{
"role": "system",
"content": system_prompt_without_tools,
},
{
"role": "user",
"content": user_prompt,
},
]
tool_size = calculate_tool_size(
ai, messages_without_tools, number_of_tools_for_workload
)

truncated_workload_result_tool_calls = [
ToolCallConversationResult(
name=tool.name,
description=tool.description,
output=tool.output[:tool_size],
)
for tool in tools_for_workload # type: ignore
]

truncated_template_context = {
"workload_analysis": workload_analysis,
"tools_called_for_workload": truncated_workload_result_tool_calls,
"resource": resource,
"toolsets": ai.tool_executor.toolsets,
"cluster_name": config.cluster_name,
"runbooks_enabled": True if runbooks else False,
}
system_prompt_with_truncated_tools = load_and_render_prompt(
template_path, truncated_template_context
)
return [
{
"role": "system",
"content": system_prompt_with_truncated_tools,
},
{
"role": "user",
"content": user_prompt,
},
]

runbooks_ctx = generate_runbooks_args(
runbook_catalog=runbooks,
global_instructions=global_instructions,
)
user_prompt = generate_user_prompt(
user_prompt,
runbooks_ctx,
)

conversation_history.append(
{
"role": "user",
"content": user_prompt,
}
)
number_of_tools = len(tools_for_workload) + len( # type: ignore
[message for message in conversation_history if message.get("role") == "tool"]
)

if number_of_tools == 0:
return conversation_history

conversation_history_without_tools = [
message for message in conversation_history if message.get("role") != "tool"
]
template_context_without_tools = {
"workload_analysis": workload_analysis,
"tools_called_for_workload": None,
"resource": resource,
"toolsets": ai.tool_executor.toolsets,
"cluster_name": config.cluster_name,
"runbooks_enabled": True if runbooks else False,
}
system_prompt_without_tools = load_and_render_prompt(
template_path, template_context_without_tools
)
conversation_history_without_tools[0]["content"] = system_prompt_without_tools

tool_size = calculate_tool_size(
ai, conversation_history_without_tools, number_of_tools
)

truncated_workload_result_tool_calls = [
ToolCallConversationResult(
name=tool.name, description=tool.description, output=tool.output[:tool_size]
)
for tool in tools_for_workload # type: ignore
]

template_context = {
"workload_analysis": workload_analysis,
"tools_called_for_workload": truncated_workload_result_tool_calls,
"resource": resource,
"toolsets": ai.tool_executor.toolsets,
"cluster_name": config.cluster_name,
"runbooks_enabled": True if runbooks else False,
}
system_prompt_with_truncated_tools = load_and_render_prompt(
template_path, template_context
)
conversation_history[0]["content"] = system_prompt_with_truncated_tools

truncate_tool_messages(conversation_history, tool_size)

return conversation_history
Loading
Loading