Add token count caching and performance instrumentation - #1692
Conversation
Reuse cached message["token_count"] in count_tokens() instead of re-tokenizing every message on every call. Also use cached counts in truncation sort and per-tool allocation to avoid redundant litellm.token_counter calls. Added debug-level timing logs to all major token counting call sites for performance visibility. https://claude.ai/code/session_01JHHMTC5p9gA1NunoyQb2Y8 Signed-off-by: Claude <noreply@anthropic.com>
…ream - Remove time.monotonic() timing instrumentation from count_tokens calls - Remove unused `import time` - Add token_count SSE event after tool results in call_stream, so streaming clients get updated token counts reflecting tool output (previously only emitted after LLM responses, not after tool execution) https://claude.ai/code/session_01JHHMTC5p9gA1NunoyQb2Y8 Signed-off-by: Claude <noreply@anthropic.com>
📂 Previous Runs📜 Run @ f733759 (#22804511524)✅ Results of HolmesGPT evalsAutomatically triggered by commit f733759 on branch Results of HolmesGPT evals
Benchmark Comparison DetailsBaseline: latest ci-benchmark experiment on master Status: Success - 61 test/model combinations loaded Benchmark experiment:
Time comparison (seconds):
Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run. Comparison indicators:
📜 Run @ 196404f (#22804165959)✅ Results of HolmesGPT evalsAutomatically triggered by commit 196404f on branch Results of HolmesGPT evals
✅ Results of HolmesGPT evalsAutomatically triggered by commit c8d48f7 on branch Results of HolmesGPT evals
Benchmark Comparison DetailsBaseline: latest ci-benchmark experiment on master Status: Success - 61 test/model combinations loaded Benchmark experiment:
Time comparison (seconds):
Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run. Comparison indicators:
📖 Legend
🔄 Re-run evals manually
Option 1: Comment on this PR with Or with more options (one per line): Run evals on a different branch (e.g., master) for comparison:
Quick re-run: Use Option 2: Trigger via GitHub Actions UI → "Run workflow" Option 3: Add PR labels to include extra evals in automatic regression runs:
Examples: 🏷️ Valid tags
Commands: CLI: |
|
✅ Docker images ready for
Use these tags to pull the images for testing. 📋 Copy commandsgcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:b4795d0a
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:b4795d0a me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:b4795d0a
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:b4795d0a
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:b4795d0a
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:b4795d0a me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:b4795d0a
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:b4795d0aPatch Helm values in one line (choose the chart you use): HolmesGPT chart: helm upgrade --install holmesgpt ./helm/holmes \
--set registry=me-west1-docker.pkg.dev/robusta-development/development \
--set image=holmes-dev:b4795d0a \
--set operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set operator.image=holmes-operator-dev:b4795d0aRobusta wrapper chart: helm upgrade --install robusta robusta/robusta \
--reuse-values \
--set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set holmes.image=holmes-dev:b4795d0a \
--set holmes.operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
--set holmes.operator.image=holmes-operator-dev:b4795d0a |
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Pro Run ID: 📒 Files selected for processing (1)
🚧 Files skipped from review as they are similar to previous changes (1)
WalkthroughAdds per-message token-count caching and time-based instrumentation to token counting, sanitizes messages before provider calls, and emits token-count/cost snapshots during tool-call streaming. Several modules receive monotonic-timing logs around counting and truncation operations. Changes
Sequence Diagram(s)sequenceDiagram
autonumber
participant Client
participant Streamer as tool_calling_llm.call_stream
participant Tools
participant LLM as holmes/core/llm.count_tokens
participant Provider
Client->>Streamer: open stream / send messages
Streamer->>Tools: invoke tool(s)
Tools-->>Streamer: tool results
Streamer->>LLM: count_tokens(current messages + tools)
LLM-->>Streamer: token totals (may use cached token_count)
Streamer->>Provider: send sanitized messages (no token_count)
Provider-->>Streamer: provider responses (streamed)
Streamer-->>Client: yield tool result + token-count snapshot (usage/cost metadata)
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~20 minutes Possibly related PRs
Suggested reviewers
🚥 Pre-merge checks | ✅ 2 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (2 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
✅ Deploy Preview for holmes-docs ready!
To edit notification comments on pull requests, go to your Netlify project configuration. |
🔬 CLI Performance Benchmark🟡 Startup Time (no LLM)Measures
🟡 Full CLI with LLMMeasures
PR: |
There was a problem hiding this comment.
Actionable comments posted: 2
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (1)
holmes/core/truncation/input_context_window_limiter.py (1)
104-122:⚠️ Potential issue | 🟠 MajorCount uncached tool messages from the full payload.
When the cache is missing, this fallback only tokenizes
{"role": "tool", "content": ...}. Actual tool messages here also includenameandtool_call_id, soneeded_spacecan be underestimated and truncation can still leave the prompt over budget. Usingorhere also treats a cached0as a miss.Suggested fix
+ def get_tool_message_token_count(message: dict[str, Any]) -> int: + cached_token_count = message.get("token_count") + if cached_token_count is not None: + return cached_token_count + + token_count = count_tokens_fn([message]).total_tokens + message["token_count"] = token_count + return token_count + t_sort = time.monotonic() - tool_call_messages.sort( - key=lambda x: x.get("token_count") or count_tokens_fn( - [{"role": "tool", "content": x["content"]}] - ).total_tokens - ) + tool_call_messages.sort(key=get_tool_message_token_count) @@ - needed_space = msg.get("token_count") or count_tokens_fn( - [{"role": "tool", "content": msg["content"]}] - ).total_tokens + needed_space = get_tool_message_token_count(msg)🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@holmes/core/truncation/input_context_window_limiter.py` around lines 104 - 122, The count fallback currently underestimates tool tokens and treats cached zero as missing; update both the sort key lambda and the needed_space calculation to use the cached token_count only when it is not None (e.g., "if msg.get('token_count') is not None") and otherwise call count_tokens_fn with the full tool message fields (include "name" and "tool_call_id" along with "role" and "content") so the tokenizer sees the same shape as the real payload; adjust references around tool_call_messages, count_tokens_fn, and needed_space accordingly.
🧹 Nitpick comments (1)
holmes/core/tools_utils/tool_context_window_limiter.py (1)
37-41: Renamet0to a descriptive timer variable.The instrumentation is fine, but
t0obscures what is being timed. A name likecount_tokens_started_atkeeps the log path self-documenting.As per coding guidelines,
**/*.{js,ts,tsx,jsx,py,java,cs,go,rb,php}: Use semantic, descriptive names for variables, functions, and components.🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@holmes/core/tools_utils/tool_context_window_limiter.py` around lines 37 - 41, Rename the ambiguous timer variable t0 to a descriptive name like count_tokens_started_at in the block that measures token counting (surrounding message = tool_call_result.as_tool_call_message(), messages_token = llm.count_tokens(...), and the logging.debug call inside prevent_overly_big_tool_response), and update the logging.debug timestamp expression to use the new variable (count_tokens_started_at) so the timing measurement remains correct and self-documenting.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Inline comments:
In `@holmes/core/llm.py`:
- Around line 314-323: count_tokens() currently mutates message dicts by setting
message["token_count"] (seen in count_tokens / the shown block) which then can
be forwarded to providers; before calling completion() (or inside
parse_messages_tags()), strip any non-schema fields by removing "token_count"
from each message (e.g., msg.pop("token_count", None)) so provider chat
completion APIs only receive valid schema fields—implement this sanitization
either at the end of parse_messages_tags() or immediately prior to the
completion() call in the completion pipeline (keeping the existing truncation
behavior in input_context_window_limiter.py for reference).
In `@holmes/core/tool_calling_llm.py`:
- Around line 1204-1214: The token-count snapshot emitted by
build_stream_event_token_count is taken too early and can be stale because
messages and tools are still mutated later (e.g. appending pending_approval to
messages and replacing tools on runbook activation); change the sequence so that
self.llm.count_tokens(...) and add_token_count_to_metadata(...) are called only
after all remaining state changes (including the pending_approval append and any
tools replacement after runbook activation) have been applied to messages and
tools, then update metadata["costs"] and yield
build_stream_event_token_count(metadata=metadata) using the final messages/tools
used in the actual response; keep references to the existing helpers
(self.llm.count_tokens, add_token_count_to_metadata,
build_stream_event_token_count, metadata, messages, tools, pending_approval,
limit_result, full_response) when moving the logic.
---
Outside diff comments:
In `@holmes/core/truncation/input_context_window_limiter.py`:
- Around line 104-122: The count fallback currently underestimates tool tokens
and treats cached zero as missing; update both the sort key lambda and the
needed_space calculation to use the cached token_count only when it is not None
(e.g., "if msg.get('token_count') is not None") and otherwise call
count_tokens_fn with the full tool message fields (include "name" and
"tool_call_id" along with "role" and "content") so the tokenizer sees the same
shape as the real payload; adjust references around tool_call_messages,
count_tokens_fn, and needed_space accordingly.
---
Nitpick comments:
In `@holmes/core/tools_utils/tool_context_window_limiter.py`:
- Around line 37-41: Rename the ambiguous timer variable t0 to a descriptive
name like count_tokens_started_at in the block that measures token counting
(surrounding message = tool_call_result.as_tool_call_message(), messages_token =
llm.count_tokens(...), and the logging.debug call inside
prevent_overly_big_tool_response), and update the logging.debug timestamp
expression to use the new variable (count_tokens_started_at) so the timing
measurement remains correct and self-documenting.
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: CHILL
Plan: Pro
Run ID: 47d16961-bb4e-4563-bc9c-b62676683ce0
📒 Files selected for processing (4)
holmes/core/llm.pyholmes/core/tool_calling_llm.pyholmes/core/tools_utils/tool_context_window_limiter.pyholmes/core/truncation/input_context_window_limiter.py
…iders count_tokens() caches per-message token counts by setting message["token_count"] on the dict objects. These same dicts are passed to litellm.completion() and ultimately to provider APIs (OpenAI, Anthropic, etc.) which may reject unknown fields. Sanitize in DefaultLLM.completion() — the single gateway to litellm — so all callers are covered (including compaction which bypasses parse_messages_tags). Shallow-copies only dicts that have the field to avoid invalidating the cache. Finding 2 (token count snapshot ordering in streaming) was verified as not a real issue: subsequent mutations are either negligible (pending_approval flag) or captured in the next iteration (runbook tool replacement). https://claude.ai/code/session_01JHHMTC5p9gA1NunoyQb2Y8 Signed-off-by: Claude <noreply@anthropic.com>
…38512/git/HolmesGPT/holmesgpt into claude/audit-token-counting-fZkU1
Summary
This PR improves performance of token counting operations by implementing a caching mechanism for per-message token counts and adds comprehensive timing instrumentation to identify bottlenecks.
Key Changes
Token Count Caching: Modified
count_tokens()inllm.pyto reuse cachedtoken_countvalues stored on individual messages, avoiding redundant tokenization calls. The cache is invalidated when messages are modified (e.g., during truncation).Performance Instrumentation: Added timing measurements and debug logging to:
count_tokens(): Tracks elapsed time and reports cache hit/miss statisticslimit_input_context_window(): Measures total execution time and final token countprevent_overly_big_tool_response(): Logs token counting duration per tooltruncate_messages_to_fit_context(): Measures sorting time for tool call messagesOptimized Token Lookups: Updated token counting calls in
input_context_window_limiter.pyto check for cached values before calling the expensivecount_tokens_fn().Token Count Updates in Streaming: Added token count metadata emission after tool results are processed in
tool_calling_llm.pyto provide updated context window information during streaming responses.Implementation Details
message["token_count"]and checked viamessage.get("token_count")before recountingtime.monotonic()for accurate elapsed time measurementhttps://claude.ai/code/session_01JHHMTC5p9gA1NunoyQb2Y8
Summary by CodeRabbit
Performance & Improvements
Observability & Privacy