Skip to content

Flush events eagerly on TOKEN_COUNT to prevent memory buildup - #1974

Merged
naomi-robusta merged 6 commits into
masterfrom
claude/flush-events-before-llm-egnHo
May 3, 2026
Merged

naomi-robusta merged 6 commits into
masterfrom
claude/flush-events-before-llm-egnHo

Conversation

@naomi-robusta

@naomi-robusta naomi-robusta commented Apr 30, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Modified the event publisher to flush pending events immediately when a TOOL_RESULT event is encountered, preventing events from sitting in memory during subsequent LLM calls.

Key Changes

  • Added StreamEvents.TOOL_RESULT to the _FLUSH_IMMEDIATELY_EVENTS set in event_publisher.py
  • Updated the comment explaining the rationale: TOOL_RESULT events precede another (typically slow) LLM call, so immediate flushing ensures subscribers see events promptly rather than waiting for the LLM call to complete
  • Added comprehensive test test_publisher_flushes_eagerly_on_tool_result() that verifies:
    • Events are split into two flush calls when TOOL_RESULT is present
    • First flush contains AI_MESSAGE + START_TOOL + TOOL_RESULT (triggered by TOOL_RESULT)
    • Second flush contains TOKEN_COUNT + ANSWER_END (triggered by ANSWER_END)
    • Compact flags are set correctly for each flush

Implementation Details

The change treats TOOL_RESULT similarly to terminal events (ANSWER_END, etc.) in terms of triggering immediate flushes, but with different semantics: rather than ending a turn, TOOL_RESULT marks a transition point where the next operation (another LLM call) may take significant time. This prevents a buildup of pending events in memory during that latency window.

https://claude.ai/code/session_01LjFK8k38XDifb4Q3LbuXcu

Summary by CodeRabbit

  • Bug Fixes

    • Token-count markers now trigger immediate flushes for more timely event delivery; tool-result events no longer force intermediate flushes.
  • Tests

    • Updated tests to verify eager flushing on token-count markers and that intermediate flushes are non-compacting while terminal flushes compact results.

The publisher's interval-based flush only fires when the next event
arrives. Between TOOL_RESULT and the next AI_MESSAGE the LLM is being
re-invoked, which typically takes >1s, so pending tool result events
sit in memory and subscribers see nothing during the wait. Promote
TOOL_RESULT to flush-immediately so subscribers get tool results as
soon as they're produced.

Signed-off-by: Claude <noreply@anthropic.com>

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This repository is configured for manual code reviews. Comment @claude review to trigger a review and subscribe this PR to future pushes, or @claude review once for a one-time review.

Tip: disable this comment in your organization's Code Review settings.

@coderabbitai

coderabbitai Bot commented Apr 30, 2026 •

Copy link
Copy Markdown
Contributor

Walkthrough

The event publisher now treats StreamEvents.TOKEN_COUNT as an eager-flush trigger, causing an immediate _flush() when TOKEN_COUNT messages are consumed. Tests were updated to reflect batching changes and to verify eager, non-compacting intermediate flushes and a compacting terminal flush.

Changes

Cohort / File(s) Summary
Event Publisher Flushing Logic
holmes/core/conversations_worker/event_publisher.py
Added StreamEvents.TOKEN_COUNT to _FLUSH_IMMEDIATELY_EVENTS and updated comment to define when call_stream() emits TOKEN_COUNT (after LLM response and after the last TOOL_RESULT in a batch), changing flush timing to occur at token-count markers.
Event Publisher Tests
tests/core/conversations_worker/test_event_publisher.py
Updated intermediate-event batching test to omit TOKEN_COUNT from the batch and adjusted expectations. Added test_publisher_flushes_eagerly_on_token_count verifying immediate flushes at TOKEN_COUNT markers (intermediate flushes with compact=False) and a final terminal flush at ANSWER_END with compact=True; asserts three DAL post calls across the stream.

Estimated code review effort

🎯 2 (Simple) | ⏱️ ~12 minutes

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The PR title states 'Flush events eagerly on TOKEN_COUNT' but the PR objectives indicate the actual implementation flushes on TOKEN_COUNT (as per the commit message rationale), aligning with the title. The raw summary confirms TOKEN_COUNT is added to _FLUSH_IMMEDIATELY_EVENTS. The title accurately describes the main change.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@github-actions

github-actions Bot commented Apr 30, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker images ready for 8fb6f745 (built in 5m 44s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use these tags to pull the images for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:8fb6f745
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:8fb6f745 me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:8fb6f745
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:8fb6f745
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:8fb6f745
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes-operator:8fb6f745 me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:8fb6f745
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-operator-dev:8fb6f745

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:8fb6f745 \
  --set operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set operator.image=holmes-operator-dev:8fb6f745

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:8fb6f745 \
  --set holmes.operator.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.operator.image=holmes-operator-dev:8fb6f745

@github-actions

github-actions Bot commented Apr 30, 2026 •

Copy link
Copy Markdown
Contributor

📂 Previous Runs

📜 #4 · Run @ __5bed630__ (#25191649474) — Apr 30, 22:12 UTC

✅ Results of HolmesGPT evals

Automatically triggered by commit 5bed630 on branch claude/flush-events-before-llm-egnHo

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/11 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 30.8s 5 8 $0.2228 98,221 96,456 22,236 1,765 849 73,649 22,807 — —
✅ 101_loki_historical_logs_pod_deleted 61.6s 8 16 $0.3503 186,690 183,046 28,011 3,644 741 153,555 29,491 — —
✅ 112_find_pvcs_by_uuid 24.9s 5 4 $0.2037 97,517 96,353 21,981 1,164 290 74,359 21,994 — —
✅ 12_job_crashing 37.2s 5 14 $0.2683 113,868 111,441 25,801 2,427 633 84,849 26,592 — —
✅ 176_network_policy_blocking_traffic_no_skills 47.0s 6 15 $0.2813 136,850 134,284 25,848 2,566 649 108,135 26,149 — —
✅ 227_count_configmaps_per_namespace[0] 27.1s 6 10 $0.2132 113,841 112,316 20,677 1,525 519 91,625 20,691 — —
✅ 243_pod_names_contain_service 36.3s 5 9 $0.2314 100,704 98,745 22,393 1,959 766 75,406 23,339 — —
✅ 24_misconfigured_pvc 35.5s 5 14 $0.2546 105,332 102,979 23,558 2,353 679 77,466 25,513 — —
✅ 43_current_datetime_from_prompt 4.5s 1 — $0.1093 17,110 16,982 16,982 128 128 0 16,982 — —
✅ 51_logs_summarize_errors 22.5s 4 5 $0.1908 78,766 77,675 21,668 1,091 349 55,995 21,680 — —
✅ 61_exact_match_counting 11.2s 3 3 $0.1390 53,006 52,632 17,967 374 227 34,654 17,978 — —
Total 30.8s avg 4.8 avg 9.8 avg $2.4647 1,101,905 1,082,909 28,011 18,996 849 829,693 253,216 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 74 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 #3 · Run @ __0573c65__ (#25177025163) — Apr 30, 16:35 UTC

✅ Results of HolmesGPT evals

Automatically triggered by commit 0573c65 on branch claude/flush-events-before-llm-egnHo

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/11 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 42.5s 7 12 $0.2770 149,102 146,671 24,616 2,431 611 121,759 24,912 — —
✅ 101_loki_historical_logs_pod_deleted 42.6s 5 10 $0.2583 109,445 106,969 24,643 2,476 859 82,044 24,925 — —
✅ 112_find_pvcs_by_uuid 22.7s 4 4 $0.1849 77,190 76,064 20,636 1,126 558 55,416 20,648 — —
✅ 12_job_crashing 34.4s 5 11 $0.2455 109,413 107,555 24,145 1,858 529 81,968 25,587 — —
✅ 176_network_policy_blocking_traffic_no_skills 51.6s 7 16 $0.3178 161,530 158,553 27,317 2,977 975 129,666 28,887 — —
✅ 227_count_configmaps_per_namespace[0] 24.2s 5 9 $0.1973 94,053 92,698 20,346 1,355 703 72,339 20,359 — —
✅ 243_pod_names_contain_service 34.0s 5 9 $0.2230 100,638 98,799 22,181 1,839 554 76,605 22,194 — —
✅ 24_misconfigured_pvc 35.8s 6 11 $0.2397 119,758 117,815 22,230 1,943 603 94,608 23,207 — —
✅ 43_current_datetime_from_prompt 5.7s 1 — $0.1092 17,106 16,982 16,982 124 124 0 16,982 — —
✅ 51_logs_summarize_errors 25.7s 4 5 $0.1852 77,117 76,023 20,835 1,094 333 55,176 20,847 — —
✅ 61_exact_match_counting 12.6s 3 3 $0.1387 52,980 52,616 17,959 364 217 34,646 17,970 — —
Total 30.2s avg 4.7 avg 9.0 avg $2.3767 1,068,332 1,050,745 27,317 17,587 975 804,227 246,518 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 74 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 #2 · Run @ __3950bce__ (#25176554505) — Apr 30, 16:25 UTC

✅ Results of HolmesGPT evals

Automatically triggered by commit 3950bce on branch claude/flush-events-before-llm-egnHo

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/11 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 34.5s 5 10 $0.2457 105,226 103,183 23,662 2,043 598 77,974 25,209 — —
✅ 101_loki_historical_logs_pod_deleted 51.5s 6 13 $0.2964 134,787 131,638 25,963 3,149 864 105,111 26,527 — —
✅ 112_find_pvcs_by_uuid 26.7s 4 4 $0.1916 78,998 77,927 21,886 1,071 383 56,029 21,898 — —
✅ 12_job_crashing 46.2s 6 16 $0.2994 141,269 138,689 26,998 2,580 701 109,400 29,289 — —
✅ 176_network_policy_blocking_traffic_no_skills 55.7s 8 16 $0.3271 185,071 182,010 27,220 3,061 646 154,116 27,894 — —
✅ 227_count_configmaps_per_namespace[0] 29.1s 6 10 $0.2161 114,021 112,467 20,743 1,554 519 91,339 21,128 — —
✅ 243_pod_names_contain_service 33.8s 4 8 $0.2153 79,998 78,082 21,868 1,916 864 55,629 22,453 — —
✅ 24_misconfigured_pvc 35.7s 5 13 $0.2494 106,793 104,525 23,565 2,268 884 79,869 24,656 — —
✅ 43_current_datetime_from_prompt 4.9s 1 — $0.1094 17,115 16,982 16,982 133 133 0 16,982 — —
✅ 51_logs_summarize_errors 24.1s 4 5 $0.1875 77,693 76,585 21,123 1,108 350 55,450 21,135 — —
✅ 61_exact_match_counting 12.1s 3 3 $0.1389 52,997 52,629 17,966 368 221 34,652 17,977 — —
Total 32.2s avg 4.7 avg 9.8 avg $2.4767 1,093,968 1,074,717 27,220 19,251 884 819,569 255,148 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 74 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📜 #1 · Run @ __259f00e__ (#25173791298) — Apr 30, 15:28 UTC

✅ Results of HolmesGPT evals

Automatically triggered by commit 259f00e on branch claude/flush-events-before-llm-egnHo

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/11 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 39.1s 6 10 $0.2494 126,564 124,583 23,470 1,981 857 100,537 24,046 — —
✅ 101_loki_historical_logs_pod_deleted 48.8s 6 11 $0.2875 133,540 130,733 25,840 2,807 854 104,152 26,581 — —
✅ 112_find_pvcs_by_uuid 26.5s 4 5 $0.2038 80,856 79,503 22,353 1,353 513 56,784 22,719 — —
✅ 12_job_crashing 39.4s 6 13 $0.2726 136,443 134,292 25,191 2,151 577 107,575 26,717 — —
✅ 176_network_policy_blocking_traffic_no_skills 43.9s 6 13 $0.2945 130,743 128,236 26,412 2,507 538 98,298 29,938 — —
✅ 227_count_configmaps_per_namespace[0] 25.6s 5 9 $0.1989 94,177 92,798 20,377 1,379 710 72,234 20,564 — —
✅ 243_pod_names_contain_service 31.5s 4 8 $0.2120 79,605 77,848 21,750 1,757 851 55,180 22,668 — —
✅ 24_misconfigured_pvc 36.4s 6 11 $0.2330 118,455 116,588 21,889 1,867 430 94,227 22,361 — —
✅ 43_current_datetime_from_prompt 5.0s 1 — $0.1094 17,111 16,982 16,982 129 129 0 16,982 — —
✅ 51_logs_summarize_errors 23.8s 4 5 $0.1871 77,258 76,098 20,877 1,160 419 55,209 20,889 — —
✅ 61_exact_match_counting 8.4s 2 1 $0.1229 34,578 34,342 17,351 236 168 16,981 17,361 — —
Total 29.9s avg 4.5 avg 8.6 avg $2.3710 1,029,330 1,012,003 26,412 17,327 857 761,177 250,826 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 74 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)

✅ Results of HolmesGPT evals

Automatically triggered by commit 5e033ba on branch claude/flush-events-before-llm-egnHo

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 11/11 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost Total tokens Input Max input Output Max output Cached Non-cached Reasoning Compactions
✅ 09_crashpod 32.4s 5 11 $0.2502 104,960 102,683 23,820 2,277 788 77,751 24,932 — —
✅ 101_loki_historical_logs_pod_deleted 47.6s 6 13 $0.2992 137,974 134,984 26,702 2,990 883 107,514 27,470 — —
✅ 112_find_pvcs_by_uuid 16.3s 3 3 $0.1790 60,680 59,687 21,614 993 551 38,062 21,625 — —
✅ 12_job_crashing 29.7s 5 12 $0.2479 110,750 108,764 24,527 1,986 589 83,607 25,157 — —
✅ 176_network_policy_blocking_traffic_no_skills 40.4s 6 15 $0.3062 132,979 130,064 26,527 2,915 885 100,028 30,036 — —
✅ 227_count_configmaps_per_namespace[0] 20.6s 5 9 $0.2008 94,353 92,942 20,426 1,411 705 72,156 20,786 — —
✅ 243_pod_names_contain_service 27.8s 5 8 $0.2201 99,603 97,869 21,730 1,734 423 75,491 22,378 — —
✅ 24_misconfigured_pvc 32.1s 5 12 $0.2501 105,786 103,427 23,541 2,359 697 78,982 24,445 — —
✅ 43_current_datetime_from_prompt 3.9s 1 — $0.0115 17,100 16,982 16,982 118 118 16,972 10 — —
✅ 51_logs_summarize_errors 20.1s 4 5 $0.1864 77,621 76,548 21,086 1,073 329 55,450 21,098 — —
✅ 61_exact_match_counting 9.5s 3 3 $0.1390 53,005 52,633 17,967 372 225 34,655 17,978 — —
Total 25.5s avg 4.4 avg 9.1 avg $2.2905 994,811 976,583 26,702 18,228 885 740,668 235,915 — —
Benchmark Comparison Details

Baseline: latest ci-benchmark experiment on master

Status: Success - 73 test/model combinations loaded

Benchmark experiment:

No benchmark data available for comparison.

Benchmark has no cost, total tokens, cached tokens data. Will appear after the next weekly benchmark run.

Comparison indicators:

  • ±0% — diff under 10% (within noise threshold)
  • ↑N%/↓N% — diff 10-25%
  • ↑N%/↓N% — diff over 25% (significant)
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/flush-events-before-llm-egnHo -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
tags: regression

Or with more options (one per line):

/eval
model: gpt-4o
tags: regression
id: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
tags: regression
Option Description
model Model(s) to test (default: same as automatic runs)
tags Pytest tags / markers (no default - runs all tests!)
id Eval ID / pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

Option 3: Add PR labels to include extra evals (applies to both automatic runs and /eval comments):

Label Effect
evals-tag-<name> Run tests with tag <name> alongside regression
evals-id-<name> Run a specific eval by test ID
evals-model-<name> Override the model (use model list name, e.g. sonnet-4.5)

Examples: evals-tag-easy, evals-id-09_crashpod, evals-model-sonnet-4.5

🏷️ Valid tags

benchmark, chain-of-causation, compaction, confluence, context_window, conversation_worker, coralogix, counting, database, datadog, datetime, db-connectors, easy, elasticsearch, embeds, fast, frontend, grafana, hard, images, integration, kafka, kubernetes, leaked-information, logs, loki, manual, mcp, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, skills, slackbot, storage, token-limit, toolset-limitation, traces, transparency, victorialogs

🤖 Valid models

deepseek-chat, deepseek-r1-reasoner, deepseek-reasoner, deepseek-v3.2-chat, gemini-3-flash-preview, gemini-3-pro-preview, gemini-3.1-pro-preview, gpt-4.1, gpt-5.2-high-reasoning, gpt-5.3-codex, gpt-5.4, haiku-4.5, kimi-2.5, kimi-2.5-openrouter, opus-4.5, opus-4.6, qwen-next-80B-instruct, qwen-next-80B-thinking, sonnet-4.5, sonnet-4.6


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/flush-events-before-llm-egnHo -f markers=regression -f filter=

@netlify

netlify Bot commented Apr 30, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit 5e033ba
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/69f7492d37b7a000086980ff
😎 Deploy Preview https://deploy-preview-1974--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@naomi-robusta
naomi-robusta enabled auto-merge (squash) April 30, 2026 15:23
@github-actions

github-actions Bot commented Apr 30, 2026 •

Copy link
Copy Markdown
Contributor

🔬 CLI Performance Benchmark

🟡 Startup Time (no LLM)

Measures holmes version execution time (imports + initialization)

Metric PR Master Change
Cold Start 10.24s 10.76s -4.9%
Warm Mean 4.85s 4.45s +9.0%
Warm Min 4.84s 4.40s
Warm Max 4.87s 4.50s

🟡 Full CLI with LLM

Measures holmes ask execution time (OpenRouter + Haiku 4.5)

Metric PR Master Change
Cold Start 13.39s 17.17s -22.0%
Warm Mean 6.28s 5.85s +7.5%
Warm Min 6.08s 5.77s
Warm Max 6.46s 5.88s

PR: 8fb6f745 | Master: e270ea66 | Iterations: 5

moshemorad
moshemorad previously approved these changes Apr 30, 2026
Comment thread holmes/core/conversations_worker/event_publisher.py Outdated
naomi-robusta and others added 3 commits April 30, 2026 19:17
call_stream() emits TOKEN_COUNT at two boundaries per loop iteration: right
after the LLM response (before tool execution) and right after the last
TOOL_RESULT of a parallel batch (before the next LLM call). Both precede a
long-running step (>1s tool work or LLM call), so flushing here keeps
subscribers up to date without the per-tool write amplification of flushing
on every TOOL_RESULT.

Signed-off-by: Claude <noreply@anthropic.com>
…gnHo' into claude/flush-events-before-llm-egnHo

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@tests/core/conversations_worker/test_event_publisher.py`:
- Line 324: Replace the ambiguous multiplication glyph in the comment "Second
flush: START_TOOL + TOOL_RESULT × 2 + TOKEN_COUNT" in test_event_publisher (the
test_event_publisher module) with a plain ASCII "x" to satisfy Ruff (RUF003);
i.e., change "×" to "x" so the comment reads "Second flush: START_TOOL +
TOOL_RESULT x 2 + TOKEN_COUNT".
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: b3806a66-a802-4518-b1b0-e30da9c5c813

📥 Commits

Reviewing files that changed from the base of the PR and between 259f00e and 0573c65.

📒 Files selected for processing (2)
  • holmes/core/conversations_worker/event_publisher.py
  • tests/core/conversations_worker/test_event_publisher.py

Comment thread tests/core/conversations_worker/test_event_publisher.py
@naomi-robusta naomi-robusta changed the title Flush events eagerly on TOOL_RESULT to prevent memory buildup Flush events eagerly on TOKEN_COUNT to prevent memory buildup Apr 30, 2026
@naomi-robusta
naomi-robusta merged commit 1f646ac into master May 3, 2026
21 of 22 checks passed
@naomi-robusta
naomi-robusta deleted the claude/flush-events-before-llm-egnHo branch May 3, 2026 13:13
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants