Skip to content

Fix CLI toolset cache not invalidated when config changes - #1502

Open
aantn wants to merge 1 commit into
masterfrom
claude/fix-cli-cache-invalidation-5v1wI
Open

aantn wants to merge 1 commit into
masterfrom
claude/fix-cli-cache-invalidation-5v1wI

Conversation

@aantn

@aantn aantn commented Feb 6, 2026 •

Copy link
Copy Markdown
Collaborator

The toolset status cache (toolsets_status.json) was never invalidated
when the user modified their config file (e.g., enabling/disabling
toolsets, changing custom_toolsets). The cache check only looked at
whether the file existed or if --refresh-toolsets was passed.

Add a config fingerprint (MD5 hash of toolset-affecting config values)
that is stored in the cache file. On load, the fingerprint is compared
to detect config changes and automatically trigger a refresh. This
also handles custom toolset file content changes via mtime tracking.

The cache format is backward-compatible: legacy format (plain list) is
detected and triggers a refresh to migrate to the new format.

https://claude.ai/code/session_01Pmbke45TMd7nHqm1MNfVd3
Signed-off-by: Claude noreply@anthropic.com

Summary by CodeRabbit

  • Bug Fixes

    • Enhanced cache handling with configuration change detection to ensure toolset updates are properly recognized and applied.
  • Tests

    • Updated test suite to reflect new cache validation structure.

The toolset status cache (toolsets_status.json) was never invalidated
when the user modified their config file (e.g., enabling/disabling
toolsets, changing custom_toolsets). The cache check only looked at
whether the file existed or if --refresh-toolsets was passed.

Add a config fingerprint (MD5 hash of toolset-affecting config values)
that is stored in the cache file. On load, the fingerprint is compared
to detect config changes and automatically trigger a refresh. This
also handles custom toolset file content changes via mtime tracking.

The cache format is backward-compatible: legacy format (plain list) is
detected and triggers a refresh to migrate to the new format.

https://claude.ai/code/session_01Pmbke45TMd7nHqm1MNfVd3
Signed-off-by: Claude <noreply@anthropic.com>
@linux-foundation-easycla

Copy link
Copy Markdown

CLA Not Signed

@netlify

netlify Bot commented Feb 6, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit b74bd7d
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/6985afc3190b9f0008c9d5d1
😎 Deploy Preview https://deploy-preview-1502--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@github-actions

github-actions Bot commented Feb 6, 2026 •

Copy link
Copy Markdown
Contributor

✅ Docker image ready for f60125b (built in 3m 41s)

⚠️ Warning: does not support ARM (ARM images are built on release only - not on every PR)

Use this tag to pull the image for testing.

📋 Copy commands

⚠️ Temporary images are deleted after 30 days. Copy to a permanent registry before using them:

gcloud auth configure-docker us-central1-docker.pkg.dev
docker pull us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:f60125b
docker tag us-central1-docker.pkg.dev/robusta-development/temporary-builds/holmes:f60125b me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:f60125b
docker push me-west1-docker.pkg.dev/robusta-development/development/holmes-dev:f60125b

Patch Helm values in one line (choose the chart you use):

HolmesGPT chart:

helm upgrade --install holmesgpt ./helm/holmes \
  --set registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set image=holmes-dev:f60125b

Robusta wrapper chart:

helm upgrade --install robusta robusta/robusta \
  --reuse-values \
  --set holmes.registry=me-west1-docker.pkg.dev/robusta-development/development \
  --set holmes.image=holmes-dev:f60125b

@github-actions

github-actions Bot commented Feb 6, 2026 •

Copy link
Copy Markdown
Contributor

✅ Results of HolmesGPT evals

Automatically triggered by commit b74bd7d on branch claude/fix-cli-cache-invalidation-5v1wI

View workflow logs

Results of HolmesGPT evals

  • ask_holmes: 9/9 test cases were successful, 0 regressions
Status Test case Time Turns Tools Cost
✅ 09_crashpod 33.9s 5 11 $0.2310
✅ 101_loki_historical_logs_pod_deleted 42.6s 5 11 $0.2672
✅ 111_pod_names_contain_service 34.3s 5 11 $0.2287
✅ 112_find_pvcs_by_uuid 35.0s 6 8 $0.2642
✅ 12_job_crashing 34.5s 5 12 $0.2331
✅ 176_network_policy_blocking_traffic_no_runbooks 45.6s 6 17 $0.2744
✅ 24_misconfigured_pvc 43.7s 7 15 $0.2569
✅ 43_current_datetime_from_prompt 4.8s 1 — $0.1049
✅ 61_exact_match_counting 17.7s 4 4 $0.1585
Total 32.4s avg 4.9 avg 11.1 avg $2.0190
📖 Legend
Icon Meaning
✅ The test was successful
➖ The test was skipped
⚠️ The test failed but is known to be flaky or known to fail
🚧 The test had a setup failure (not a code regression)
🔧 The test failed due to mock data issues (not a code regression)
🚫 The test was throttled by API rate limits/overload
❌ The test failed and should be fixed before merging the PR
🔄 Re-run evals manually

⚠️ Warning: /eval comments always run using the workflow from master, not from this PR branch. If you modified the GitHub Action (e.g., added secrets or env vars), those changes won't take effect.

To test workflow changes, use the GitHub CLI or Actions UI instead:

gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/fix-cli-cache-invalidation-5v1wI -f markers=regression -f filter=

Option 1: Comment on this PR with /eval:

/eval
markers: regression

Or with more options (one per line):

/eval
model: gpt-4o
markers: regression
filter: 09_crashpod
iterations: 5

Run evals on a different branch (e.g., master) for comparison:

/eval
branch: master
markers: regression
Option Description
model Model(s) to test (default: same as automatic runs)
markers Pytest markers (no default - runs all tests!)
filter Pytest -k filter (use /list to see valid eval names)
iterations Number of runs, max 10
branch Run evals on a different branch (for cross-branch comparison)

Quick re-run: Use /rerun to re-run the most recent /eval on this PR with the same parameters.

Option 2: Trigger via GitHub Actions UI → "Run workflow"

🏷️ Valid markers

benchmark, chain-of-causation, compaction, confluence, context_window, coralogix, counting, database, datadog, datetime, easy, elasticsearch, embeds, frontend, grafana-dashboard, hard, integration, kafka, kubernetes, leaked-information, logs, loki, medium, metrics, network, newrelic, no-cicd, numerical, one-test, port-forward, prometheus, question-answer, regression, runbooks, slackbot, storage, toolset-limitation, traces, transparency


Commands: /eval · /rerun · /list

CLI: gh workflow run eval-regression.yaml --repo HolmesGPT/holmesgpt --ref claude/fix-cli-cache-invalidation-5v1wI -f markers=regression -f filter=

@coderabbitai

coderabbitai Bot commented Feb 6, 2026 •

Copy link
Copy Markdown
Contributor

Walkthrough

The changes introduce a configuration fingerprint mechanism to the ToolsetManager to detect when cached toolsets become stale. Three new private methods compute fingerprints from configuration, check cache freshness by comparing stored and current fingerprints, and read cached toolsets. The cache format extends from storing raw toolset lists to storing a dictionary containing both the fingerprint and toolsets.

Changes

Cohort / File(s) Summary
Cache Fingerprinting
holmes/core/toolset_manager.py
Adds _config_fingerprint attribute and three new methods: _compute_config_fingerprint() hashes toolset configuration via MD5, _is_cache_stale() compares cached vs. current fingerprints, and _read_cached_toolsets() reads from new cache dict format. Updates refresh_toolset_status() to write cache as {"config_fingerprint": ..., "toolsets": ...} instead of raw list. Modifies load_toolset_with_status() to check cache staleness before deciding whether to refresh or load from cache. Adds hashlib import.
Test Updates
tests/core/test_toolset_manager.py
Updates test assertions to reflect new cache structure: data["toolsets"][0]["name"] instead of data[0]["name"]. Adds config_fingerprint validation in cache setup. Aligns test expectations with fingerprint and dict-based cache format.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~22 minutes

Possibly related PRs

Suggested reviewers

  • nherment
🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately describes the main change: adding cache invalidation when configuration changes. It directly reflects the core objective of detecting config changes and invalidating the toolset cache via a fingerprint mechanism.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@github-actions

github-actions Bot commented Feb 6, 2026

Copy link
Copy Markdown
Contributor

🔬 CLI Performance Benchmark

🟡 Startup Time (no LLM)

Measures holmes version execution time (imports + initialization)

Metric PR Master Change
Cold Start 10.36s 10.64s -2.6%
Warm Mean 4.86s 4.67s +4.1%
Warm Min 4.81s 4.61s
Warm Max 4.98s 4.72s

🟡 Full CLI with LLM

Measures holmes ask execution time (OpenRouter + Haiku 4.5)

Metric PR Master Change
Cold Start 26.30s 15.60s +68.6%
Warm Mean 8.34s 6.95s +19.9%
Warm Min 8.17s 6.67s
Warm Max 8.59s 7.13s

PR: f60125b1 | Master: aa839918 | Iterations: 5

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Fix all issues with AI agents
In `@holmes/core/toolset_manager.py`:
- Around line 125-131: The _read_cached_toolsets function should defensively
handle missing/corrupt cache files: wrap the open + json.load block in a
try/except that catches OSError and json.JSONDecodeError (and ValueError if
needed), log a short warning via the instance logger (e.g., self.logger.warning)
with the error, and return an empty list so callers like
load_toolset_with_status will treat the cache as missing and trigger a refresh;
keep the existing legacy-format handling after the try block.

In `@tests/core/test_toolset_manager.py`:
- Around line 140-158: Add unit tests for stale-cache and fingerprint logic:
create tests that write a cache file with a non-matching "config_fingerprint"
and assert that _is_cache_stale() returns True and that
load_toolset_with_status() triggers refresh_toolset_status (mock refresh);
create a test that writes a legacy plain JSON list as the cache and assert it is
treated as stale; and add deterministic fingerprint tests that call
_compute_config_fingerprint() twice on the same config (assert equal) and once
after mutating any config field (assert different). Reference the existing test
harness that sets toolset_manager.toolset_status_location and uses
load_toolset_with_status(), mocking refresh_toolset_status where needed.
🧹 Nitpick comments (2)
holmes/core/toolset_manager.py (2)

81-105: Consider using hashlib.sha256 instead of md5 to satisfy the Ruff S324 lint rule.

MD5 isn't being used for security here (only cache fingerprinting), so it's functionally fine. However, Ruff flags it (S324: Probable use of insecure hash functions), and sha256 is equally fast for small payloads and avoids the lint suppression.

♻️ Swap to sha256
-        config_str = json.dumps(fingerprint_data, sort_keys=True, default=str)
-        return hashlib.md5(config_str.encode()).hexdigest()
+        config_str = json.dumps(fingerprint_data, sort_keys=True, default=str)
+        return hashlib.sha256(config_str.encode()).hexdigest()

107-123: _is_cache_stale reads and parses the entire cache file, then _read_cached_toolsets re-reads it — consider unifying.

When the cache is not stale, the JSON is parsed twice: once inside _is_cache_stale (Line 111) and again inside _read_cached_toolsets (Line 128, called from Line 350). For a CLI tool this is unlikely to be a real performance problem, but it's a straightforward refactor to parse once and reuse the result.

Also, Ruff TRY300 suggests the return False on Line 121 would be better inside an else block — this is a minor style nit.

Comment on lines +125 to +131
def _read_cached_toolsets(self) -> List[dict[str, Any]]:
"""Read cached toolset status, handling both legacy and new formats."""
with open(self.toolset_status_location, "r") as f:
raw_data = json.load(f)
if isinstance(raw_data, list):
return raw_data
return raw_data.get("toolsets", [])

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

_read_cached_toolsets has no error handling — verify this is intentional.

If the cache file is missing or corrupt when this method is called (e.g., deleted between the _is_cache_stale check and this read), an unhandled OSError or JSONDecodeError will propagate up to load_toolset_with_status. Since this is a CLI tool and the race window is tiny, this is low-risk, but wrapping in a try/except that falls back to triggering a refresh would make it more robust.

🛡️ Proposed defensive handling
     def _read_cached_toolsets(self) -> List[dict[str, Any]]:
         """Read cached toolset status, handling both legacy and new formats."""
-        with open(self.toolset_status_location, "r") as f:
-            raw_data = json.load(f)
-        if isinstance(raw_data, list):
-            return raw_data
-        return raw_data.get("toolsets", [])
+        try:
+            with open(self.toolset_status_location, "r") as f:
+                raw_data = json.load(f)
+            if isinstance(raw_data, list):
+                return raw_data
+            return raw_data.get("toolsets", [])
+        except (json.JSONDecodeError, OSError):
+            logging.warning("Failed to read cached toolsets, returning empty list")
+            return []
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
def _read_cached_toolsets(self) -> List[dict[str, Any]]:
"""Read cached toolset status, handling both legacy and new formats."""
with open(self.toolset_status_location, "r") as f:
raw_data = json.load(f)
if isinstance(raw_data, list):
return raw_data
return raw_data.get("toolsets", [])
def _read_cached_toolsets(self) -> List[dict[str, Any]]:
"""Read cached toolset status, handling both legacy and new formats."""
try:
with open(self.toolset_status_location, "r") as f:
raw_data = json.load(f)
if isinstance(raw_data, list):
return raw_data
return raw_data.get("toolsets", [])
except (json.JSONDecodeError, OSError):
logging.warning("Failed to read cached toolsets, returning empty list")
return []
🤖 Prompt for AI Agents
In `@holmes/core/toolset_manager.py` around lines 125 - 131, The
_read_cached_toolsets function should defensively handle missing/corrupt cache
files: wrap the open + json.load block in a try/except that catches OSError and
json.JSONDecodeError (and ValueError if needed), log a short warning via the
instance logger (e.g., self.logger.warning) with the error, and return an empty
list so callers like load_toolset_with_status will treat the cache as missing
and trigger a refresh; keep the existing legacy-format handling after the try
block.

Comment on lines +140 to 158
toolset_manager.toolset_status_location = cache_path
cache_data = {
"config_fingerprint": toolset_manager._config_fingerprint,
"toolsets": [
{
"name": "test",
"status": "enabled",
"enabled": True,
"type": "built-in",
"path": None,
"error": None,
}
],
}
with open(cache_path, "w") as f:
json.dump(cache_data, f)
toolset_manager.toolset_status_location = cache_path
result = toolset_manager.load_toolset_with_status()
assert result[0].name == "test"
assert result[0].enabled is True

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

🧩 Analysis chain

🏁 Script executed:

rg -n '_is_cache_stale\|_compute_config_fingerprint\|_read_cached_toolsets\|config_fingerprint' --type=py -g '*test*'

Repository: HolmesGPT/holmesgpt

Length of output: 45


🏁 Script executed:

# Also check the actual test file to see what tests exist
cat -n tests/core/test_toolset_manager.py | head -200

Repository: HolmesGPT/holmesgpt

Length of output: 8721


🏁 Script executed:

# Find the toolset_manager implementation to understand what was changed
find . -type f -name "toolset_manager.py" | grep -v test

Repository: HolmesGPT/holmesgpt

Length of output: 95


🏁 Script executed:

# Get the toolset_manager.py implementation
wc -l ./holmes/core/toolset_manager.py

Repository: HolmesGPT/holmesgpt

Length of output: 99


🏁 Script executed:

# View the toolset_manager implementation focusing on fingerprint and cache logic
cat -n ./holmes/core/toolset_manager.py | grep -A 5 -B 5 'fingerprint\|_is_cache_stale\|_read_cached_toolsets' | head -100

Repository: HolmesGPT/holmesgpt

Length of output: 4071


🏁 Script executed:

# Get a better view of the load_toolset_with_status and refresh_toolset_status methods
ast-grep --pattern $'def load_toolset_with_status($_) {
  $$$
}'

Repository: HolmesGPT/holmesgpt

Length of output: 45


🏁 Script executed:

# Check the rest of the test file for any tests covering stale cache or fingerprint mismatch
tail -n +160 tests/core/test_toolset_manager.py

Repository: HolmesGPT/holmesgpt

Length of output: 17130


🏁 Script executed:

# Look at the full load_toolset_with_status method implementation
sed -n '320,360p' ./holmes/core/toolset_manager.py

Repository: HolmesGPT/holmesgpt

Length of output: 1644


Add tests for stale cache detection and fingerprint logic to cover the core feature of this PR.

The test at lines 140–158 only exercises the cache-hit path (matching fingerprint). The stale-cache detection logic—the main behavioral change—lacks coverage:

  • Stale fingerprint triggers refresh: Cache with a different fingerprint should cause _is_cache_stale() to return True and trigger refresh_toolset_status.
  • Legacy (list) format triggers refresh: A plain JSON list in the cache file should be detected as stale.
  • _compute_config_fingerprint determinism: Same config must produce the same fingerprint; changing any config field must produce a different one.

Per coding guidelines, all new features require unit tests.

🤖 Prompt for AI Agents
In `@tests/core/test_toolset_manager.py` around lines 140 - 158, Add unit tests
for stale-cache and fingerprint logic: create tests that write a cache file with
a non-matching "config_fingerprint" and assert that _is_cache_stale() returns
True and that load_toolset_with_status() triggers refresh_toolset_status (mock
refresh); create a test that writes a legacy plain JSON list as the cache and
assert it is treated as stale; and add deterministic fingerprint tests that call
_compute_config_fingerprint() twice on the same config (assert equal) and once
after mutating any config field (assert different). Reference the existing test
harness that sets toolset_manager.toolset_status_location and uses
load_toolset_with_status(), mocking refresh_toolset_status where needed.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants