Skip to content

Add cached tokens tracking to LLM usage reporting - #1680

Merged
aantn merged 11 commits into
masterfrom
claude/add-cached-tokens-column-bdEhY
Mar 7, 2026
Merged

aantn merged 11 commits into
masterfrom
claude/add-cached-tokens-column-bdEhY

Conversation

@aantn

@aantn aantn commented Mar 6, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

This PR adds support for tracking and reporting cached tokens from LLM API responses. Cached tokens represent prompt tokens that were retrieved from the model's cache, reducing both latency and cost. The feature is now integrated throughout the LLM usage tracking pipeline and included in markdown reports.

Key Changes

  • LLMResponseUsage: Added cached_tokens field to the NamedTuple to capture cached token counts from API responses
  • extract_usage_from_response(): Implemented extraction of cached tokens from the prompt_tokens_details field in LLM responses, supporting both dict and object attribute access patterns
  • LLMCosts model: Added cached_tokens field and accumulation logic in _process_cost_info() to aggregate cached tokens across multiple LLM calls
  • Test result tracking: Updated property_manager.py to capture and store cached tokens in test node properties
  • Test result collection: Modified conftest.py to include cached tokens when collecting test results from pytest stats
  • Markdown reporting: Enhanced the detailed results table in generate_markdown_report() to display cached tokens as a new column, with totals aggregation in the summary row

Implementation Details

  • Cached tokens are extracted from the prompt_tokens_details.cached_tokens field when available in the LLM response
  • The implementation handles both dictionary and object attribute access patterns for flexibility across different LLM provider SDKs
  • Cached tokens are formatted with thousand separators in reports (e.g., "1,234") and display "—" when not available
  • The summary row includes total cached tokens aggregated across all test cases

https://claude.ai/code/session_01AcrrH8vM9rV5cA2yV1RmUj

Summary by CodeRabbit

  • New Features

    • Tracks cached_tokens and reasoning_tokens alongside existing token and cost metrics.
    • Exposes max_completion_tokens_per_call to reflect per-call output limits.
  • Reporting

    • Expanded reports show per-row and aggregate columns: Total, Input, Output, Cached, Non-cached, Reasoning, and Max Output.
    • Terminal output adds a Tokens column with formatted totals.
  • Tests

    • Test result emissions include cached_tokens, reasoning_tokens, and max_completion_tokens_per_call.

claude and others added 2 commits March 5, 2026 22:17
Track cached_tokens (from prompt_tokens_details.cached_tokens) through
the full pipeline: LLM response extraction → LLMCosts accumulation →
pytest user_properties → GitHub report. The new "Cached" column appears
next to the existing "Tokens" column.

https://claude.ai/code/session_01AcrrH8vM9rV5cA2yV1RmUj
Signed-off-by: Claude <noreply@anthropic.com>
@github-actions

github-actions Bot commented Mar 6, 2026 •

Copy link
Copy Markdown
Contributor

📂 Previous Runs

📜 Run @ bd76554 (#22758158754)

✅ Results of HolmesGPT evals

Automatically triggered by commit bd76554 on branch claude/add-cached-tokens-column-bdEhY

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Output Cached Non-cached Reasoning Max output Compactions
✅ 09_crashpod 33.0s 5 12 $0.2459 108,211 106,267 1,944 80,970 25,297 — 538 —
✅ 101_loki_historical_logs_pod_deleted 46.8s 7 10 $0.2724 150,259 147,985 2,274 123,307 24,678 — 497 —
✅ 111_pod_names_contain_service 34.6s 5 12 $0.2390 105,949 104,003 1,946 79,732 24,271 — 575 —
✅ 112_find_pvcs_by_uuid 39.4s 7 9 $0.2760 152,910 150,804 2,106 124,904 25,900 — 431 —
✅ 12_job_crashing 34.9s 5 12 $0.2457 108,977 107,068 1,909 81,741 25,327 — 459 —
✅ 176_network_policy_blocking_traffic_no_runbooks 48.0s 8 17 $0.3240 181,646 179,041 2,605 148,857 30,184 — 568 —
✅ 24_misconfigured_pvc 36.0s 5 15 $0.2491 107,689 105,527 2,162 80,611 24,916 — 590 —
✅ 43_current_datetime_from_prompt 6.3s 1 — $0.1127 17,549 17,388 161 0 17,388 — 161 —
✅ 61_exact_match_counting 17.0s 4 4 $0.1670 77,158 76,635 523 56,530 20,105 — 225 —
Total 32.9s avg 5.2 avg 11.4 avg $2.1318 1,010,348 994,718 15,630 776,652 218,066 — 590 —
📜 Run @ 5d34d44 (#22757733306)

✅ Results of HolmesGPT evals

Automatically triggered by commit 5d34d44 on branch claude/add-cached-tokens-column-bdEhY

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Output Cached Non-cached Reasoning Max output Compactions
✅ 09_crashpod 39.3s 7 12 $0.2664 149,812 147,830 1,982 122,925 24,905 — 486 —
✅ 101_loki_historical_logs_pod_deleted 41.4s 5 10 $0.2525 107,981 105,676 2,305 80,782 24,894 — 831 —
✅ 111_pod_names_contain_service 34.4s 5 12 $0.2410 105,885 103,879 2,006 79,498 24,381 — 495 —
✅ 112_find_pvcs_by_uuid 42.4s 7 11 $0.3143 165,701 163,404 2,297 132,516 30,888 — 421 —
✅ 12_job_crashing 33.4s 5 11 $0.2415 108,331 106,558 1,773 81,244 25,314 — 467 —
✅ 176_network_policy_blocking_traffic_no_runbooks 43.2s 7 15 $0.2852 155,751 153,481 2,270 127,008 26,473 — 417 —
✅ 24_misconfigured_pvc 35.5s 5 13 $0.2377 105,925 103,969 1,956 80,006 23,963 — 620 —
✅ 43_current_datetime_from_prompt 4.8s 1 — $0.1120 17,522 17,388 134 0 17,388 — 134 —
✅ 61_exact_match_counting 12.8s 3 2 $0.1490 56,189 55,820 369 36,366 19,454 — 149 —
Total 31.9s avg 5.0 avg 10.8 avg $2.0997 973,097 958,005 15,092 740,345 217,660 — 831 —
📜 Run @ a4aa6df (#22757233529)

✅ Results of HolmesGPT evals

Automatically triggered by commit a4aa6df on branch claude/add-cached-tokens-column-bdEhY

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Output Cached Non-cached Reasoning Max output Compactions
✅ 09_crashpod 33.3s 5 11 $0.1461 107,837 105,996 1,841 97,548 8,448 — 605 —
✅ 101_loki_historical_logs_pod_deleted 37.1s 4 9 $0.1377 85,303 83,076 2,227 75,933 7,143 — 800 —
✅ 111_pod_names_contain_service 35.5s 5 13 $0.1538 109,394 107,425 1,969 98,421 9,004 — 582 —
✅ 112_find_pvcs_by_uuid 36.8s 6 10 $0.1714 133,695 131,735 1,960 121,552 10,183 — 520 —
✅ 12_job_crashing 31.6s 5 11 $0.1399 107,944 106,265 1,679 98,284 7,981 — 460 —
✅ 176_network_policy_blocking_traffic_no_runbooks 46.7s 7 16 $0.1973 155,686 153,406 2,280 141,892 11,514 — 505 —
✅ 24_misconfigured_pvc 34.2s 5 15 $0.1465 106,705 104,652 2,053 97,042 7,610 — 686 —
✅ 43_current_datetime_from_prompt 5.8s 1 — $0.0121 17,522 17,388 134 17,378 10 — 134 —
✅ 61_exact_match_counting 16.7s 4 4 $0.0706 77,184 76,657 527 73,333 3,324 — 229 —
Total 30.9s avg 4.7 avg 11.1 avg $1.1753 901,270 886,600 14,670 821,383 65,217 — 800 —
📜 Run @ 8f0f0b0 (#22756357620)

✅ Results of HolmesGPT evals

Automatically triggered by commit 8f0f0b0 on branch claude/add-cached-tokens-column-bdEhY

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Output Cached Non-cached Reasoning Max output Compactions
✅ 09_crashpod 33.6s 5 11 $0.2459 106,956 104,996 1,960 79,527 25,469 — 501 —
✅ 101_loki_historical_logs_pod_deleted 45.7s 6 11 $0.2814 135,786 133,292 2,494 106,425 26,867 — 805 —
✅ 111_pod_names_contain_service 32.7s 5 10 $0.2251 103,345 101,698 1,647 78,341 23,357 — 475 —
✅ 112_find_pvcs_by_uuid 32.3s 6 7 $0.2333 123,859 122,268 1,591 99,153 23,115 — 427 —
✅ 12_job_crashing 31.8s 5 11 $0.2386 108,076 106,317 1,759 81,524 24,793 — 458 —
✅ 176_network_policy_blocking_traffic_no_runbooks 46.8s 6 14 $0.2835 136,787 134,461 2,326 106,497 27,964 — 687 —
✅ 24_misconfigured_pvc 40.4s 7 16 $0.2716 146,033 143,859 2,174 118,229 25,630 — 471 —
✅ 43_current_datetime_from_prompt 5.1s 1 — $0.1115 17,503 17,388 115 — 17,388 — 115 —
✅ 61_exact_match_counting 17.0s 4 4 $0.1670 77,166 76,643 523 56,535 20,108 — 225 —
Total 31.7s avg 5.0 avg 10.5 avg $2.0580 955,511 940,922 14,589 726,231 214,691 — 805 —
📜 Run @ b68f1c6 (#22755754413)

✅ Results of HolmesGPT evals

Automatically triggered by commit b68f1c6 on branch claude/add-cached-tokens-column-bdEhY

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Input Output Cached Non-cached Reasoning Max output Compactions
✅ 09_crashpod 36.0s 6 12 $0.2579 124,581 2,048 99,252 25,329 — 561 —
✅ 101_loki_historical_logs_pod_deleted 33.3s 4 8 $0.2197 81,345 1,956 58,661 22,684 — 828 —
✅ 111_pod_names_contain_service 30.3s 4 11 $0.2183 81,113 1,766 57,735 23,378 — 695 —
✅ 112_find_pvcs_by_uuid 38.9s 7 9 $0.2780 151,079 2,125 124,941 26,138 — 432 —
✅ 12_job_crashing 32.0s 5 10 $0.2373 105,581 1,727 80,722 24,859 — 454 —
✅ 176_network_policy_blocking_traffic_no_runbooks 44.2s 7 15 $0.2889 152,491 2,578 126,718 25,773 — 707 —
✅ 24_misconfigured_pvc 33.4s 5 14 $0.2420 104,117 1,979 79,397 24,720 — 672 —
✅ 43_current_datetime_from_prompt 5.2s 1 — $0.1118 17,388 126 — 17,388 — 126 —
✅ 61_exact_match_counting 12.8s 3 2 $0.1506 55,977 411 36,446 19,531 — 192 —
Total 29.6s avg 4.7 avg 10.1 avg $2.0044 873,672 14,716 663,872 209,800 — 828 —

✅ Results of HolmesGPT evals

Automatically triggered by commit f911932 on branch claude/add-cached-tokens-column-bdEhY

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Output Cached Non-cached Reasoning Max output Compactions
✅ 09_crashpod 33.0s 5 12 $0.2476 108,632 106,673 1,959 81,154 25,519 — 558 —
✅ 101_loki_historical_logs_pod_deleted 46.3s 7 10 $0.2674 147,430 145,144 2,286 121,098 24,046 — 429 —
✅ 111_pod_names_contain_service 32.1s 5 12 $0.2316 104,832 103,069 1,763 79,262 23,807 — 457 —
✅ 112_find_pvcs_by_uuid 33.5s 6 7 $0.2455 127,765 126,050 1,715 101,761 24,289 — 434 —
✅ 12_job_crashing 39.2s 6 15 $0.2790 139,334 137,147 2,187 109,979 27,168 — 460 —
✅ 176_network_policy_blocking_traffic_no_runbooks 51.1s 7 17 $0.3191 169,610 166,658 2,952 138,398 28,260 — 735 —
✅ 24_misconfigured_pvc 37.2s 6 16 $0.2632 130,430 128,176 2,254 103,158 25,018 — 519 —
✅ 43_current_datetime_from_prompt 5.2s 1 — $0.1117 17,511 17,388 123 0 17,388 — 123 —
✅ 61_exact_match_counting 16.5s 4 4 $0.1671 77,107 76,573 534 56,485 20,088 — 236 —
Total 32.7s avg 5.2 avg 11.6 avg $2.1323 1,022,651 1,006,878 15,773 791,295 215,583 — 735 —
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/add-cached-tokens-column-bdEhY -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
tags: regression

Or with more options (one per line):

/eval
model: gpt-4o
tags: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
tags: regression
Option Description
model Model(s) to test (default: same as automatic runs)
tags Pytest tags / markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

Option 3: Add PR labels to include extra evals in automatic regression runs:

Label Effect
evals-tag-<name> Run tests with tag <name> alongside regression
evals-id-<name> Run a specific eval by test ID

Examples: evals-tag-easy, evals-id-09_crashpod

🏷️ Valid tags

benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, db-connectors, easy, elasticsearch, embeds, fast, frontend, grafana-dashboard, hard, integration, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/add-cached-tokens-column-bdEhY -f markers=regression -f filter=

@github-actions

github-actions Bot commented Mar 6, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker images ready for 80c0eb61 (built in 5m 34s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use these tags to pull the images for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:80c0eb61
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:80c0eb61 me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:80c0eb61
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:80c0eb61
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:80c0eb61
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:80c0eb61 me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:80c0eb61
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:80c0eb61

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:80c0eb61 \
  --set operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set operator.image=holmes-operator-dev:80c0eb61

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:80c0eb61 \
  --set holmes.operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.operator.image=holmes-operator-dev:80c0eb61

@coderabbitai

coderabbitai Bot commented Mar 6, 2026 •

Copy link
Copy Markdown
Contributor

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review

Walkthrough

Adds cached_tokens and reasoning_tokens to LLM usage and cost models, extracts them from LLM responses, accumulates max per-call completion tokens, and surfaces these fields through test result collection and GitHub/terminal reporting.

Changes

Cohort / File(s) Summary
LLM usage & extraction
holmes/core/llm_usage.py, holmes/core/llm.py
Add cached_tokens: Optional[int] and reasoning_tokens: int to LLMResponseUsage; implement helper to extract detail fields; use centralized extract_usage_from_response in get_llm_usage to populate these fields.
Cost accumulation / tool-calling
holmes/core/tool_calling_llm.py
Add cached_tokens, reasoning_tokens, and max_completion_tokens_per_call to LLMCosts; accumulate cached/reasoning tokens and update max completion tokens per call in _process_cost_info.
Test harness — collection & properties
tests/llm/conftest.py, tests/llm/utils/property_manager.py
Emit and record cached_tokens, reasoning_tokens, and max_completion_tokens_per_call in test result dicts and pytest user_properties.
Test reporting — GitHub & terminal
tests/llm/utils/reporting/github_reporter.py, tests/llm/utils/reporting/terminal_reporter.py
Extend per-row and aggregate reporting to include total tokens, cached, non-cached, reasoning, and max-per-call columns and update formatting/aggregates accordingly.

Sequence Diagram(s)

sequenceDiagram
    participant Runner as Test Runner
    participant LLM as LLM Service
    participant Extractor as Usage Extractor
    participant Cost as Cost Accumulator
    participant Collector as Test Result Collector
    participant Reporter as Reporter

    Runner->>LLM: invoke model
    LLM-->>Runner: response (usage, details)
    Runner->>Extractor: extract_usage_from_response(response)
    Extractor->>Extractor: parse prompt/completion/cached/reasoning
    Extractor-->>Cost: LLMResponseUsage
    Cost->>Cost: accumulate cached/reasoning tokens, update max_completion_tokens_per_call
    Cost-->>Collector: attach usage/cost metadata to result
    Collector-->>Reporter: emit rows with input/output/cached/non-cached/reasoning/max
    Reporter->>Reporter: compute aggregates and render report
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related PRs

Suggested reviewers

  • moshemorad
🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The pull request title accurately summarizes the main change: adding support for tracking cached tokens in LLM usage reporting, which is reflected across all modified files.
Docstring Coverage ✅ Passed Docstring coverage is 87.50% which is sufficient. The required threshold is 80.00%.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@netlify

netlify Bot commented Mar 6, 2026

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit ccd4702
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/69aa79f17bb1fd000815db5e
😎 Deploy Preview https://deploy-preview-1680--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@netlify

netlify Bot commented Mar 6, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit f911932
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/69abdb63d8689000084bd355
😎 Deploy Preview https://deploy-preview-1680--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@github-actions

github-actions Bot commented Mar 6, 2026 •

Copy link
Copy Markdown
Contributor

🔬 CLI Performance Benchmark

🟡 Startup Time (no LLM)

Measures holmes version execution time (imports + initialization)

Metric PR Master Change
Cold Start 11.25s 11.10s +1.4%
Warm Mean 5.17s 5.10s +1.2%
Warm Min 5.07s 5.04s
Warm Max 5.20s 5.14s

🟡 Full CLI with LLM

Measures holmes ask execution time (OpenRouter + Haiku 4.5)

Metric PR Master Change
Cold Start 27.71s 26.72s +3.7%
Warm Mean 7.58s 7.65s -0.9%
Warm Min 7.46s 7.40s
Warm Max 7.80s 7.99s

PR: 80c0eb61 | Master: 06ac630e | Iterations: 5

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@holmes/core/llm_usage.py`:
- Line 16: The code currently coerces missing
prompt_tokens_details.cached_tokens into 0, losing the distinction between
"unavailable" and a real zero; update holmes/core/llm_usage.py to preserve
"unavailable" by making cached_tokens Optional[int] (allow None) and stop
defaulting to 0 when reading prompt_tokens_details.cached_tokens (leave as None
if absent). Find the cached_tokens annotation and any parsing/assignment that
reads prompt_tokens_details.cached_tokens (including the similar logic around
lines 35–58) and change those assignments to set None when the provider did not
supply the metric; also update any downstream checks to explicitly test for None
vs 0 where needed.

In `@holmes/core/tool_calling_llm.py`:
- Line 159: The compaction cached token counts are not being added into the
overall LLMCosts totals, so update the compaction accumulation in
ToolCallingLLM.call and ToolCallingLLM.call_stream to include
compaction.cached_tokens into the running LLMCosts (same pattern used for
prompt/completion/total tokens); specifically, when accumulating compaction
results into the costs variable (referencing compaction and costs/LLMCosts), add
costs.cached_tokens += compaction.cached_tokens or 0 so that _process_cost_info
and the final markdown/report reflect cached tokens consumed during compaction.

In `@tests/llm/utils/reporting/github_reporter.py`:
- Around line 301-307: The current logic treats falsy cached_tokens the same as
missing data; change the condition to distinguish None from 0 by checking if
result.get("cached_tokens") is None (i.e., if cached_tokens is None then set
cached_tokens_str = "—"), otherwise format the integer (even if 0) into
cached_tokens_str = f"{cached_tokens:,}" and add to total_cached_tokens_sum;
update the same pattern used at the other occurrence (the block around
cached_tokens handling at lines ~325-327) to use the None check rather than a
truthy check.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 2c97c778-a0e2-4c4d-be04-c5af7a8a8633

📥 Commits

Reviewing files that changed from the base of the PR and between a8b22c7 and 0f8c479.

📒 Files selected for processing (5)
  • holmes/core/llm_usage.py
  • holmes/core/tool_calling_llm.py
  • tests/llm/conftest.py
  • tests/llm/utils/property_manager.py
  • tests/llm/utils/reporting/github_reporter.py

Comment thread holmes/core/llm_usage.py Outdated
Comment thread holmes/core/tool_calling_llm.py Outdated
Comment thread tests/llm/utils/reporting/github_reporter.py Outdated
claude added 3 commits March 6, 2026 07:05
Shows prompt_tokens - cached_tokens as "Non-cached" column, giving
visibility into actual billable input tokens per test.

https://claude.ai/code/session_01AcrrH8vM9rV5cA2yV1RmUj
Signed-off-by: Claude <noreply@anthropic.com>
Replace the single "Tokens" column with detailed breakdown:
- Input: prompt_tokens summed across all LLM calls
- Output: completion_tokens summed across all LLM calls
- Cached / Non-cached: cached vs fresh input tokens
- Reasoning: reasoning_tokens from completion_tokens_details
- Max output: largest single-call completion_tokens (useful for
  sizing output token reservations)

Also adds cached_tokens and reasoning_tokens to SSE token_count
events via get_llm_usage(), and tracks max_completion_tokens_per_call
through the full pipeline (LLMCosts -> property_manager -> conftest
-> github_reporter).

https://claude.ai/code/session_01AcrrH8vM9rV5cA2yV1RmUj
Signed-off-by: Claude <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🧹 Nitpick comments (1)
holmes/core/llm.py (1)

686-696: Consider consolidating with extract_usage_from_response to avoid duplication.

Both get_llm_usage() and extract_usage_from_response() in llm_usage.py extract the same fields from the same response structure using similar logic. This duplication increases maintenance burden and risk of inconsistency.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@holmes/core/llm.py` around lines 686 - 696, get_llm_usage duplicates logic
already implemented in extract_usage_from_response; remove the duplicated
field-extraction in get_llm_usage (functions: get_llm_usage) and delegate to or
call the shared helper extract_usage_from_response in llm_usage.py to populate
prompt_tokens_details.cached_tokens and
completion_tokens_details.reasoning_tokens (symbols: prompt_tokens_details,
cached_tokens, completion_tokens_details, reasoning_tokens) so there is a single
source of truth for usage extraction; if extract_usage_from_response is not
accessible, move the shared extraction logic into a small private helper (e.g.,
_extract_usage_fields) and have both get_llm_usage and
extract_usage_from_response call it.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@tests/llm/utils/reporting/github_reporter.py`:
- Around line 321-329: The current logic sets non_cached_tokens_str to "—"
whenever non_cached_tokens <= 0, which hides the valid case where prompt_tokens
> 0 and cached_tokens == prompt_tokens (a 0 non-cached token result); update the
condition in the block around prompt_tokens, cached_tokens, non_cached_tokens,
non_cached_tokens_str and total_non_cached_tokens_sum so that "—" is only used
when prompt_tokens is falsy/None (missing data), but when prompt_tokens is
present and non_cached_tokens == 0 set non_cached_tokens_str to "0" (formatted
with commas) and still add 0 to total_non_cached_tokens_sum as appropriate.

---

Nitpick comments:
In `@holmes/core/llm.py`:
- Around line 686-696: get_llm_usage duplicates logic already implemented in
extract_usage_from_response; remove the duplicated field-extraction in
get_llm_usage (functions: get_llm_usage) and delegate to or call the shared
helper extract_usage_from_response in llm_usage.py to populate
prompt_tokens_details.cached_tokens and
completion_tokens_details.reasoning_tokens (symbols: prompt_tokens_details,
cached_tokens, completion_tokens_details, reasoning_tokens) so there is a single
source of truth for usage extraction; if extract_usage_from_response is not
accessible, move the shared extraction logic into a small private helper (e.g.,
_extract_usage_fields) and have both get_llm_usage and
extract_usage_from_response call it.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 50b22bae-f474-4c59-a300-4c4f6954cc71

📥 Commits

Reviewing files that changed from the base of the PR and between 0f8c479 and df1b09d.

📒 Files selected for processing (6)
  • holmes/core/llm.py
  • holmes/core/llm_usage.py
  • holmes/core/tool_calling_llm.py
  • tests/llm/conftest.py
  • tests/llm/utils/property_manager.py
  • tests/llm/utils/reporting/github_reporter.py
🚧 Files skipped from review as they are similar to previous changes (1)
  • tests/llm/conftest.py

Comment thread tests/llm/utils/reporting/github_reporter.py Outdated
Adds a "Tokens" column (with comma-formatted count) after the Cost
column in the terminal reporter summary table. The data was already
being tracked in test properties and Braintrust metadata but wasn't
visible in the terminal output.

https://claude.ai/code/session_01AcrrH8vM9rV5cA2yV1RmUj
Signed-off-by: Claude <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@tests/llm/utils/reporting/terminal_reporter.py`:
- Around line 266-271: When formatting total tokens in the terminal report, if
result.get("total_tokens") is missing or zero but prompt_tokens or
completion_tokens exist, compute tokens_str by summing
result.get("prompt_tokens", 0) + result.get("completion_tokens", 0) and format
that sum with thousands separators; otherwise fall back to "—". Update the logic
around the total_tokens/prompt_tokens/completion_tokens handling (referencing
result, total_tokens, prompt_tokens, completion_tokens) so the report shows the
component sum when total_tokens is not provided.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: d870f7cd-1499-4444-8b16-a14aacba20d9

📥 Commits

Reviewing files that changed from the base of the PR and between df1b09d and b68f1c6.

📒 Files selected for processing (1)
  • tests/llm/utils/reporting/terminal_reporter.py

Comment thread tests/llm/utils/reporting/terminal_reporter.py Outdated
Adds a "Total tokens" column after Cost in the GitHub markdown reporter
and a "Tokens" column in the terminal Rich table. This shows the
total_tokens value (prompt + completion) that was already tracked but
not displayed in either reporter.

https://claude.ai/code/session_01AcrrH8vM9rV5cA2yV1RmUj
Signed-off-by: Claude <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

♻️ Duplicate comments (1)
tests/llm/utils/reporting/github_reporter.py (1)

322-338: ⚠️ Potential issue | 🟡 Minor

Render zero cache values as 0, not —.

cached_tokens == 0 and non_cached_tokens == 0 are valid outcomes here, but the truthy checks collapse them into “unavailable”. Because holmes/core/llm_usage.py:32-77 and tests/llm/conftest.py:870-882 normalize these fields to 0, the new columns will otherwise hide cache misses, 100% cache hits, and all-zero totals.

💡 Suggested fix
-        cached_tokens = result.get("cached_tokens", 0)
-        if cached_tokens and cached_tokens > 0:
+        cached_tokens = result.get("cached_tokens")
+        if cached_tokens is None:
+            cached_tokens_str = "—"
+        else:
             cached_tokens_str = f"{cached_tokens:,}"
             total_cached_tokens_sum += cached_tokens
-        else:
-            cached_tokens_str = "—"
...
-        if non_cached_tokens > 0:
+        if prompt_for_calc > 0:
             non_cached_tokens_str = f"{non_cached_tokens:,}"
             total_non_cached_tokens_sum += non_cached_tokens
         else:
             non_cached_tokens_str = "—"
...
-    total_cached_tokens_str = f"{total_cached_tokens_sum:,}" if total_cached_tokens_sum > 0 else "—"
-    total_non_cached_tokens_str = f"{total_non_cached_tokens_sum:,}" if total_non_cached_tokens_sum > 0 else "—"
+    total_cached_tokens_str = f"{total_cached_tokens_sum:,}"
+    total_non_cached_tokens_str = f"{total_non_cached_tokens_sum:,}"

Also applies to: 373-374

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@tests/llm/utils/reporting/github_reporter.py` around lines 322 - 338, The
code currently treats zero as "unavailable" by using truthy checks; change those
checks to explicit None checks so 0 is rendered as "0". Specifically, for
cached_tokens (from result.get(...)) replace the if cached_tokens and
cached_tokens > 0 condition with if cached_tokens is not None to set
cached_tokens_str = f"{cached_tokens:,}" and add cached_tokens to
total_cached_tokens_sum; similarly, compute non_cached_tokens as currently done
and replace if non_cached_tokens > 0 with if non_cached_tokens is not None to
set non_cached_tokens_str = f"{non_cached_tokens:,}" and add to
total_non_cached_tokens_sum. Reference symbols: cached_tokens,
cached_tokens_str, total_cached_tokens_sum, non_cached_tokens,
non_cached_tokens_str, total_non_cached_tokens_sum, and result.get.
🧹 Nitpick comments (1)
tests/llm/utils/reporting/github_reporter.py (1)

230-231: Use a per-call label for the max completion column.

Max output is ambiguous next to Output; it reads like another aggregate instead of a peak-per-call metric. A label like Max output/call or Max completion/call would make the column self-explanatory.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@tests/llm/utils/reporting/github_reporter.py` around lines 230 - 231, The
table header string built into the markdown variable uses an ambiguous column
label "Max output"; update that header to a per-call label like "Max
output/call" or "Max completion/call" in the two places where the header rows
are concatenated (the lines that append to markdown: the first header row
containing column names and the second separator row), so the table column is
clearly described as a peak-per-call metric instead of an aggregate.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Duplicate comments:
In `@tests/llm/utils/reporting/github_reporter.py`:
- Around line 322-338: The code currently treats zero as "unavailable" by using
truthy checks; change those checks to explicit None checks so 0 is rendered as
"0". Specifically, for cached_tokens (from result.get(...)) replace the if
cached_tokens and cached_tokens > 0 condition with if cached_tokens is not None
to set cached_tokens_str = f"{cached_tokens:,}" and add cached_tokens to
total_cached_tokens_sum; similarly, compute non_cached_tokens as currently done
and replace if non_cached_tokens > 0 with if non_cached_tokens is not None to
set non_cached_tokens_str = f"{non_cached_tokens:,}" and add to
total_non_cached_tokens_sum. Reference symbols: cached_tokens,
cached_tokens_str, total_cached_tokens_sum, non_cached_tokens,
non_cached_tokens_str, total_non_cached_tokens_sum, and result.get.

---

Nitpick comments:
In `@tests/llm/utils/reporting/github_reporter.py`:
- Around line 230-231: The table header string built into the markdown variable
uses an ambiguous column label "Max output"; update that header to a per-call
label like "Max output/call" or "Max completion/call" in the two places where
the header rows are concatenated (the lines that append to markdown: the first
header row containing column names and the second separator row), so the table
column is clearly described as a peak-per-call metric instead of an aggregate.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 636d9a31-f6f8-4c11-86e4-35e8be1e9f20

📥 Commits

Reviewing files that changed from the base of the PR and between b68f1c6 and 8f0f0b0.

📒 Files selected for processing (1)
  • tests/llm/utils/reporting/github_reporter.py

- cached_tokens is now Optional[int] (None = provider didn't report it,
  0 = provider reported zero cached). This preserves the distinction
  between "unavailable" and "real zero" throughout the pipeline.

- Deduplicate: get_llm_usage() in llm.py now delegates to
  extract_usage_from_response() instead of reimplementing the same
  prompt_tokens_details/completion_tokens_details parsing.

- GitHub reporter: non_cached_tokens shows "0" when prompt_tokens > 0
  but everything was cached, instead of hiding it as "—".

- Terminal reporter: total_tokens falls back to prompt + completion sum
  when total_tokens is not provided by the API.

https://claude.ai/code/session_01AcrrH8vM9rV5cA2yV1RmUj
Signed-off-by: Claude <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
holmes/core/tool_calling_llm.py (1)

197-217: ⚠️ Potential issue | 🟠 Major

Don't gate usage accumulation on total_tokens alone.

Providers can omit usage.total_tokens while still returning prompt_tokens, completion_tokens, cached_tokens, or reasoning_tokens. The current guard at line 197 (if raw.total_tokens > 0) skips the entire accumulation block when total_tokens is 0 or missing, causing valid token metrics to be dropped from cost tracking and reporting even when individual token counts are available.

The fix should check whether any token metric is present before accumulating:

💡 Suggested fix
-        if raw.total_tokens > 0:
+        has_token_usage = (
+            raw.total_tokens > 0
+            or raw.prompt_tokens > 0
+            or raw.completion_tokens > 0
+            or raw.cached_tokens is not None
+            or raw.reasoning_tokens > 0
+        )
+        if has_token_usage:
             cost_logger.debug(
                 f"{log_prefix} cost: ${raw.cost:.6f} | Tokens: {raw.prompt_tokens} prompt + {raw.completion_tokens} completion = {raw.total_tokens} total"
             )
             if costs:
                 costs.total_cost += raw.cost
                 costs.prompt_tokens += raw.prompt_tokens
                 costs.completion_tokens += raw.completion_tokens
-                costs.total_tokens += raw.total_tokens
+                costs.total_tokens += raw.total_tokens or (
+                    raw.prompt_tokens + raw.completion_tokens
+                )
                 if raw.cached_tokens is not None:
                     costs.cached_tokens = (costs.cached_tokens or 0) + raw.cached_tokens
                 costs.reasoning_tokens += raw.reasoning_tokens
                 costs.max_completion_tokens_per_call = max(
                     costs.max_completion_tokens_per_call, raw.completion_tokens
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@holmes/core/tool_calling_llm.py` around lines 197 - 217, The accumulation
block currently gated by "if raw.total_tokens > 0" should instead run whenever
any token metric is present; update the condition around the accumulation that
references raw and costs (the block using cost_logger, costs.total_cost,
costs.prompt_tokens, costs.completion_tokens, costs.total_tokens,
costs.cached_tokens, costs.reasoning_tokens,
costs.max_completion_tokens_per_call) to check for presence of any of
raw.prompt_tokens, raw.completion_tokens, raw.total_tokens, raw.cached_tokens,
or raw.reasoning_tokens (e.g. any is not None or >0) before aggregating; keep
the separate elif raw.cost > 0 branch for when no token metrics exist, and
ensure cached_tokens handling still guards against None when adding to
costs.cached_tokens.
♻️ Duplicate comments (2)
holmes/core/tool_calling_llm.py (1)

206-211: ⚠️ Potential issue | 🟠 Major

Compaction runs still bypass the new token fields.

The normal response path now accumulates cached/reasoning/max-completion metrics, but the separate compaction accumulation in ToolCallingLLM.call() and ToolCallingLLM.call_stream() still copies only cost/prompt/completion/total. Any compaction cache hits or reasoning tokens will be missing from final totals.

💡 Suggested fix
 # In both compaction accumulation blocks
 if compaction.total_tokens > 0:
     costs.num_compactions += 1
     costs.total_tokens += compaction.total_tokens
     costs.prompt_tokens += compaction.prompt_tokens
     costs.completion_tokens += compaction.completion_tokens
     costs.total_cost += compaction.cost
+    if compaction.cached_tokens is not None:
+        costs.cached_tokens = (costs.cached_tokens or 0) + compaction.cached_tokens
+    costs.reasoning_tokens += compaction.reasoning_tokens
+    costs.max_completion_tokens_per_call = max(
+        costs.max_completion_tokens_per_call,
+        compaction.completion_tokens,
+    )
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@holmes/core/tool_calling_llm.py` around lines 206 - 211, The compaction path
in ToolCallingLLM.call and ToolCallingLLM.call_stream currently only accumulates
cost/prompt/completion/total and omits the new token fields; update the
compaction accumulation logic to also add raw.cached_tokens into
costs.cached_tokens (safely handling None as in the normal path), increment
costs.reasoning_tokens by raw.reasoning_tokens, and update
costs.max_completion_tokens_per_call to max(current, raw.completion_tokens) so
the final totals include cached, reasoning, and max-completion tokens (use the
same null/coalesce behavior used in the non-compaction accumulation).
tests/llm/utils/reporting/github_reporter.py (1)

239-240: ⚠️ Potential issue | 🟡 Minor

Keep explicit zero cache totals visible in the summary row.

Per-row formatting now distinguishes None from 0, but the summary row reverts to — unless the aggregate is positive. If every run is a cache miss (cached_tokens == 0) or fully cached (non_cached_tokens == 0), the total currently hides a valid zero.

💡 Suggested fix
     total_cached_tokens_sum = 0
     total_non_cached_tokens_sum = 0
+    saw_cached_tokens = False
+    saw_non_cached_tokens = False
     total_reasoning_tokens_sum = 0
     max_completion_per_call_max = 0
@@
         if cached_tokens is not None:
             cached_tokens_str = f"{cached_tokens:,}"
             total_cached_tokens_sum += cached_tokens
+            saw_cached_tokens = True
         else:
             cached_tokens_str = "—"
@@
         if prompt_for_calc > 0:
             non_cached_tokens_str = f"{non_cached_tokens:,}"
             total_non_cached_tokens_sum += non_cached_tokens
+            saw_non_cached_tokens = True
         else:
             non_cached_tokens_str = "—"
@@
-    total_cached_tokens_str = f"{total_cached_tokens_sum:,}" if total_cached_tokens_sum > 0 else "—"
-    total_non_cached_tokens_str = f"{total_non_cached_tokens_sum:,}" if total_non_cached_tokens_sum > 0 else "—"
+    total_cached_tokens_str = (
+        f"{total_cached_tokens_sum:,}" if saw_cached_tokens else "—"
+    )
+    total_non_cached_tokens_str = (
+        f"{total_non_cached_tokens_sum:,}" if saw_non_cached_tokens else "—"
+    )

Also applies to: 322-339, 372-375

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@tests/llm/utils/reporting/github_reporter.py` around lines 239 - 240, The
summary row hides valid zero totals because the code uses truthiness to decide
between displaying a numeric total and an em-dash; change those checks to
explicitly test for None so zero (0) is rendered as "0". Update the logic where
total_cached_tokens_sum and total_non_cached_tokens_sum are formatted for the
summary (and the similar blocks around the other occurrences) to use "if
total_cached_tokens_sum is None" / "if total_non_cached_tokens_sum is None"
rather than "if not total_cached_tokens_sum" before outputting the value so
explicit zeros remain visible.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@tests/llm/utils/reporting/github_reporter.py`:
- Around line 298-304: The code sets total_tokens_str to "—" when result lacks
total_tokens, which undercounts rows; instead, if result.get("total_tokens") is
falsy but result contains prompt_tokens and/or completion_tokens, compute
total_tokens = (result.get("prompt_tokens", 0) + result.get("completion_tokens",
0)), set total_tokens_str = f"{total_tokens:,}" and add that value into
total_tokens_sum; update the logic around the variables total_tokens,
prompt_tokens, completion_tokens, total_tokens_str, and total_tokens_sum to use
this computed fallback.

---

Outside diff comments:
In `@holmes/core/tool_calling_llm.py`:
- Around line 197-217: The accumulation block currently gated by "if
raw.total_tokens > 0" should instead run whenever any token metric is present;
update the condition around the accumulation that references raw and costs (the
block using cost_logger, costs.total_cost, costs.prompt_tokens,
costs.completion_tokens, costs.total_tokens, costs.cached_tokens,
costs.reasoning_tokens, costs.max_completion_tokens_per_call) to check for
presence of any of raw.prompt_tokens, raw.completion_tokens, raw.total_tokens,
raw.cached_tokens, or raw.reasoning_tokens (e.g. any is not None or >0) before
aggregating; keep the separate elif raw.cost > 0 branch for when no token
metrics exist, and ensure cached_tokens handling still guards against None when
adding to costs.cached_tokens.

---

Duplicate comments:
In `@holmes/core/tool_calling_llm.py`:
- Around line 206-211: The compaction path in ToolCallingLLM.call and
ToolCallingLLM.call_stream currently only accumulates
cost/prompt/completion/total and omits the new token fields; update the
compaction accumulation logic to also add raw.cached_tokens into
costs.cached_tokens (safely handling None as in the normal path), increment
costs.reasoning_tokens by raw.reasoning_tokens, and update
costs.max_completion_tokens_per_call to max(current, raw.completion_tokens) so
the final totals include cached, reasoning, and max-completion tokens (use the
same null/coalesce behavior used in the non-compaction accumulation).

In `@tests/llm/utils/reporting/github_reporter.py`:
- Around line 239-240: The summary row hides valid zero totals because the code
uses truthiness to decide between displaying a numeric total and an em-dash;
change those checks to explicitly test for None so zero (0) is rendered as "0".
Update the logic where total_cached_tokens_sum and total_non_cached_tokens_sum
are formatted for the summary (and the similar blocks around the other
occurrences) to use "if total_cached_tokens_sum is None" / "if
total_non_cached_tokens_sum is None" rather than "if not
total_cached_tokens_sum" before outputting the value so explicit zeros remain
visible.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: d1a444e4-685e-4bfc-9917-f7a1c516c89e

📥 Commits

Reviewing files that changed from the base of the PR and between 8f0f0b0 and a4aa6df.

📒 Files selected for processing (6)
  • holmes/core/llm.py
  • holmes/core/llm_usage.py
  • holmes/core/tool_calling_llm.py
  • tests/llm/conftest.py
  • tests/llm/utils/reporting/github_reporter.py
  • tests/llm/utils/reporting/terminal_reporter.py
🚧 Files skipped from review as they are similar to previous changes (1)
  • tests/llm/conftest.py

Comment thread tests/llm/utils/reporting/github_reporter.py Outdated
claude and others added 2 commits March 6, 2026 09:37
- Remove unused _extract_cost_from_response function
- Simplify test_cache.py to reuse extract_usage_from_response
- Extract _fmt_tokens helper to reduce formatting repetition in github_reporter
- Fix non_cached_tokens to show "—" when cached is unknown (not misleading number)
- Add total_tokens fallback in github_reporter for consistency with terminal_reporter

https://claude.ai/code/session_01AcrrH8vM9rV5cA2yV1RmUj
Signed-off-by: Claude <noreply@anthropic.com>
@aantn
aantn enabled auto-merge (squash) March 6, 2026 09:50
@aantn
aantn merged commit 86c740c into master Mar 7, 2026
19 of 22 checks passed
@aantn
aantn deleted the claude/add-cached-tokens-column-bdEhY branch March 7, 2026 08:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants