Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
71 changes: 62 additions & 9 deletions docs/reference/http-api.md
Original file line number Diff line number Diff line change
Expand Up @@ -573,29 +573,80 @@ Emitted periodically to provide token usage updates during the chat. This event

---

#### `conversation_history_compaction_start`

Emitted when the conversation history is about to be compacted. This event fires before the compaction LLM call, allowing clients to show a loading state.

**Payload:**
```json
{
"content": "Compacting conversation history (150000 tokens, 42 messages)...",
"metadata": {
"initial_tokens": 150000,
"num_messages": 42,
"max_context_size": 128000,
"threshold_pct": 95
}
}
```

**Fields:**

- `content` (string): Human-readable status message
- `metadata` (object): Context window state before compaction
- `initial_tokens` (integer): Current token count triggering compaction
- `num_messages` (integer): Number of messages in the conversation
- `max_context_size` (integer): Model's maximum context window size
- `threshold_pct` (integer): Context window usage percentage that triggered compaction

---

#### `conversation_history_compacted`

Emitted when the conversation history has been compacted to fit within the context window. This happens automatically when the conversation grows too large.
Emitted when the conversation history has been compacted to fit within the context window. This happens automatically when the conversation grows too large. Contains detailed statistics about the compaction result.

**Payload:**
```json
{
"content": "Conversation history was compacted to fit within context limits.",
"content": "The conversation history has been compacted from 150000 to 80000 tokens",
"compaction_summary": "<analysis>\n1. Primary Request: User asked to investigate pod crashes...\n2. Key Technical Concepts: OOMKilled, memory limits...\n...\n</analysis>",
"messages": [...],
"metadata": {
"initial_tokens": 150000,
"compacted_tokens": 80000
"compacted_tokens": 80000,
"compression_ratio_pct": 46.7,
"num_messages_before": 42,
"num_messages_after": 4,
"max_context_size": 128000,
"threshold_pct": 95,
"compaction_cost": {
"total_cost": 0.003542,
"prompt_tokens": 12000,
"completion_tokens": 800,
"total_tokens": 12800
}
}
}
```

**Fields:**

- `content` (string): Human-readable description of the compaction
- `compaction_summary` (string|null): The LLM-generated summary of the previous conversation history. This is the full text the model produced to condense the conversation, wrapped in `<analysis>` tags. Useful for debugging to verify that important context was preserved during compaction.
- `messages` (array): The compacted conversation history
- `metadata` (object): Token information about the compaction
- `metadata` (object): Detailed compaction statistics
- `initial_tokens` (integer): Token count before compaction
- `compacted_tokens` (integer): Token count after compaction
- `compression_ratio_pct` (number): Percentage of tokens saved (e.g., 46.7 means 46.7% reduction)
- `num_messages_before` (integer): Number of messages before compaction
- `num_messages_after` (integer): Number of messages after compaction (typically 3-4)
- `max_context_size` (integer): Model's maximum context window size
- `threshold_pct` (integer): Context window usage percentage that triggered compaction
- `compaction_cost` (object, optional): Cost of the compaction LLM call
- `total_cost` (number): Dollar cost of the compaction call
- `prompt_tokens` (integer): Prompt tokens used for compaction
- `completion_tokens` (integer): Completion tokens generated during compaction
- `total_tokens` (integer): Total tokens used for compaction

---

Expand Down Expand Up @@ -646,9 +697,11 @@ Emitted when an error occurs during processing.
### Chat with History Compaction

```
1. conversation_history_compacted
2. start_tool_calling (tool 1)
3. tool_calling_result (tool 1)
4. token_count
5. ai_answer_end
1. conversation_history_compaction_start
2. conversation_history_compacted
3. ai_message (compaction notice)
4. start_tool_calling (tool 1)
5. tool_calling_result (tool 1)
6. token_count
7. ai_answer_end
```
50 changes: 46 additions & 4 deletions holmes/core/truncation/input_context_window_limiter.py
Original file line number Diff line number Diff line change
Expand Up @@ -163,6 +163,22 @@ def limit_input_context_window(
if ENABLE_CONVERSATION_HISTORY_COMPACTION and (
initial_tokens.total_tokens + maximum_output_token
) > (max_context_size * get_context_window_compaction_threshold_pct() / 100):
num_messages_before = len(messages)
events.append(
StreamMessage(
event=StreamEvents.CONVERSATION_HISTORY_COMPACTION_START,
data={
"content": f"Compacting conversation history ({initial_tokens.total_tokens} tokens, {num_messages_before} messages)...",
"metadata": {
"initial_tokens": initial_tokens.total_tokens,
"num_messages": num_messages_before,
"max_context_size": max_context_size,
"threshold_pct": get_context_window_compaction_threshold_pct(),
},
},
)
)

compaction_result = compact_conversation_history(
original_conversation_history=messages, llm=llm
)
Expand All @@ -172,19 +188,45 @@ def limit_input_context_window(

if compacted_total_tokens < initial_tokens.total_tokens:
messages = compaction_result.messages_after_compaction
num_messages_after = len(messages)
compression_ratio = round((1 - compacted_total_tokens / initial_tokens.total_tokens) * 100, 1)
compaction_message = f"The conversation history has been compacted from {initial_tokens.total_tokens} to {compacted_total_tokens} tokens"
logging.info(compaction_message)
conversation_history_compacted = True

# Extract the LLM-generated summary from the compacted messages
# Structure is: [system_prompt?, last_user_prompt?, assistant_summary, continuation_marker]
compaction_summary = None
for msg in compaction_result.messages_after_compaction:
if msg.get("role") == "assistant":
compaction_summary = msg.get("content")
break

compaction_stats: dict = {
"initial_tokens": initial_tokens.total_tokens,
"compacted_tokens": compacted_total_tokens,
"compression_ratio_pct": compression_ratio,
"num_messages_before": num_messages_before,
"num_messages_after": num_messages_after,
"max_context_size": max_context_size,
"threshold_pct": get_context_window_compaction_threshold_pct(),
}
if compaction_usage:
compaction_stats["compaction_cost"] = {
"total_cost": compaction_usage.total_cost,
"prompt_tokens": compaction_usage.prompt_tokens,
"completion_tokens": compaction_usage.completion_tokens,
"total_tokens": compaction_usage.total_tokens,
}

events.append(
StreamMessage(
event=StreamEvents.CONVERSATION_HISTORY_COMPACTED,
data={
"content": compaction_message,
"compaction_summary": compaction_summary,
"messages": compaction_result.messages_after_compaction,
"metadata": {
"initial_tokens": initial_tokens.total_tokens,
"compacted_tokens": compacted_total_tokens,
},
"metadata": compaction_stats,
},
)
)
Expand Down
1 change: 1 addition & 0 deletions holmes/utils/stream.py
Original file line number Diff line number Diff line change
Expand Up @@ -20,6 +20,7 @@ class StreamEvents(str, Enum):
AI_MESSAGE = "ai_message"
APPROVAL_REQUIRED = "approval_required"
TOKEN_COUNT = "token_count"
CONVERSATION_HISTORY_COMPACTION_START = "conversation_history_compaction_start"
CONVERSATION_HISTORY_COMPACTED = "conversation_history_compacted"


Expand Down
Loading