Readme update and toolset fix - #3
Merged
Merged
Conversation
arikalon1
added a commit
that referenced
this pull request
Dec 7, 2025
When a Todo task fails, the LLM returns a "Failed" state. But since this state is not present, an error is raised and thrown in the terminal. ``` Running tool #3 TodoWrite: Update investigation tasks error using todowrite tool Traceback (most recent call last): File "/Users/pavan/Documents/repos/holmesgpt/holmes/plugins/toolsets/investigator/core_investigation.py", line 102, in _invoke tasks = parse_tasks(todos_data=todos_data) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/Users/pavan/Documents/repos/holmesgpt/holmes/plugins/toolsets/investigator/core_investigation.py", line 30, in parse_tasks status=TaskStatus(todo_item.get("status", "pending")), ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/Users/pavan/.pyenv/versions/3.12.1/lib/python3.12/enum.py", line 744, in __call__ return cls.__new__(cls, value) ^^^^^^^^^^^^^^^^^^^^^^^ File "/Users/pavan/.pyenv/versions/3.12.1/lib/python3.12/enum.py", line 1158, in __new__ raise ve_exc ValueError: 'failed' is not a valid TaskStatus ``` Handle this error and display it properly see **status** for failed tasks <img width="2048" height="614" alt="image" src="https://github.com/user-attachments/assets/77c373da-a723-4ecf-a57f-902539a63c64" /> Co-authored-by: arik <alon.arik@gmail.com>
aantn
pushed a commit
that referenced
this pull request
Dec 10, 2025
When a Todo task fails, the LLM returns a "Failed" state. But since this state is not present, an error is raised and thrown in the terminal. ``` Running tool #3 TodoWrite: Update investigation tasks error using todowrite tool Traceback (most recent call last): File "/Users/pavan/Documents/repos/holmesgpt/holmes/plugins/toolsets/investigator/core_investigation.py", line 102, in _invoke tasks = parse_tasks(todos_data=todos_data) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/Users/pavan/Documents/repos/holmesgpt/holmes/plugins/toolsets/investigator/core_investigation.py", line 30, in parse_tasks status=TaskStatus(todo_item.get("status", "pending")), ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/Users/pavan/.pyenv/versions/3.12.1/lib/python3.12/enum.py", line 744, in __call__ return cls.__new__(cls, value) ^^^^^^^^^^^^^^^^^^^^^^^ File "/Users/pavan/.pyenv/versions/3.12.1/lib/python3.12/enum.py", line 1158, in __new__ raise ve_exc ValueError: 'failed' is not a valid TaskStatus ``` Handle this error and display it properly see **status** for failed tasks <img width="2048" height="614" alt="image" src="https://github.com/user-attachments/assets/77c373da-a723-4ecf-a57f-902539a63c64" /> Co-authored-by: arik <alon.arik@gmail.com> Signed-off-by: Robusta Runner <aantny@gmail.com>
goyamegh
added a commit
to goyamegh/holmesgpt
that referenced
this pull request
Mar 9, 2026
- Rename OTEL_AWS_SERVICE → HOLMES_AWS_OSIS_SERVICE with backwards-compat fallback (Comment HolmesGPT#10, svrnm) - Align OTEL_DEBUG with OTEL spec OTEL_LOG_LEVEL=debug with backwards-compat fallback (Comment HolmesGPT#9, svrnm) - Add ml-commons AgentTracer.java GitHub permalink (Comment HolmesGPT#3, kylehounslow) - Document needs_aws_auth() as single source of truth with consumer list (Comment HolmesGPT#4, kylehounslow) - Clarify otel_logging.py: logs NOT exported via OTLP, naming avoids shadowing builtin logging (Comment HolmesGPT#5, kylehounslow) - Explain experimental/ placement: API evolving, removable via try/except no-op fallbacks (Comment HolmesGPT#6, kylehounslow) - Clarify server.py middleware: tracing/metrics init in init_otel() above, middleware only handles per-request spans (Comment HolmesGPT#7, kylehounslow) - Split env var docs into Standard OTEL / Holmes-Specific tables, add missing vars, add OSIS hyperlink (Comments HolmesGPT#1, HolmesGPT#9, HolmesGPT#10) - Fix no-op tracer fallback in server-agui.py for when OTEL is unavailable Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> Signed-off-by: Megha Goyal <goyamegh@amazon.com>
aantn
pushed a commit
that referenced
this pull request
Mar 21, 2026
Address 5 CodeRabbit review comments from PR #1811: Issue #3 - Cache key inadequacy: create_tool_executor now stores a cache key (tags, enable_all) alongside the executor and validates it on cache hit. Callers with different parameters get a fresh executor. Issue #4/#5 - Change detection misses added/removed toolsets: refresh_toolsets_and_get_changes now compares name sets between old and new toolset lists, reporting additions and removals alongside status transitions. refresh_tool_executor always updates the cached executor (not just when status changes exist). Issue #6 - Race condition: Added threading.Lock around all reads and writes of _cached_tool_executor. Added Config.cached_tool_executor property for thread-safe external access. Updated server.py to use it. Issue #7 - Cold-cache refresh: refresh_tool_executor now uses FORCE_REFRESH on cold start instead of ENABLED, ensuring live prerequisite checks rather than stale disk cache. https://claude.ai/code/session_01Qr7wtBgYGnh1G6Bq2wPp79 Signed-off-by: Claude <noreply@anthropic.com>
aantn
pushed a commit
that referenced
this pull request
Mar 21, 2026
- Embed timestamp as hidden HTML comment in buildBody() - Extract and display date in "Mar 21, 14:32 UTC" format in run summaries - Add sequential run numbers (#1, #2, ...) to previous runs (newest = highest) - Strip existing run numbers on re-parse to avoid double-numbering - Bold commit SHA with __underscores__ for visual clarity - Handle multiline header matching for timestamp comment prefix Example: 📜 #3 · Run @ __8d93be9__ (#23379683267) — Mar 21, 14:32 UTC https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt Signed-off-by: Claude <noreply@anthropic.com>
This was referenced Mar 21, 2026
This was referenced Mar 31, 2026
This was referenced Apr 30, 2026
Merged
Merged
Merged
This was referenced May 10, 2026
This was referenced May 17, 2026
Merged
This was referenced Jun 7, 2026
aantn
pushed a commit
that referenced
this pull request
Jun 12, 2026
…the proxy hop Run #3's surfaced error showed the claude CLI reporting 'Not logged in' and every turn erroring. Cause: CI sets the ANTHROPIC_API_KEY secret to empty (it uses OpenRouter), and the engine/bootstrap inherited that empty value; the CLI refuses to run without a non-empty key. The local proxy ignores the key (it authenticates to the provider via the model_list), so a placeholder is correct. Engine and bootstrap now coerce an empty/unset ANTHROPIC_API_KEY to a non-empty placeholder for the CLI<->proxy hop. Signed-off-by: Claude <noreply@anthropic.com>
naomi-robusta
pushed a commit
that referenced
this pull request
Jul 20, 2026
Second optimization pass, keeping only changes that held up under repeated eval runs on the production model mix: - conversation_history_compaction.jinja2: 1,200 -> ~390 tokens; same section contract. Verified with the compaction eval suite: 7/7 before, 7/7 after (opus-4.6). - _toolsets_instructions.jinja2: drop the per-entry status label on disabled toolsets (header states it once) and merge the missing-integration guidance into one paragraph. - investigator_instructions.jinja2: further condensed TodoWrite mechanics. - fetch_pod_logs tool schema: compressed start/end_time and filter parameter descriptions (~200 tokens per prompt with kubernetes/logs enabled). - Updated string-pin unit tests (fetch_logs wording, disabled-list format). An even more aggressive generic_ask rewrite (extra ~480 tokens saved) was built and evaluated but NOT kept: over 9 runs of the multicluster suite, haiku-4.5 averaged 6.75/13 with it vs 8.6/13 with the current version, while opus was unaffected. Since haiku is the #3 production model, the current generic_ask stays. Eval evidence for the final state (this commit): - multicluster suite: opus-4.7 13/13, sonnet-4.6 10/13 (only pre-existing 235 + known-flaky 254 + one skills test), haiku-4.5 9/13 and 7/13 (within its 7-11 baseline band), opus-4.6 12/13 in prior rounds - k8s regression: opus-4.6 12/13 (only the known-flaky 254; 24_misconfigured_pvc passes) - ES quirks/skills: opus-4.6 6/6 (one flake passed on rerun), opus-4.7 6/7, sonnet-4.6 6/7 - bash prefix suite: opus-4.6 9/9, opus-4.7 8/9 (205 is flaky ~1/3 runs on all configs including baseline) - compaction: 7/7 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XgDuG2hb6jmSqYpD7iGUkd Signed-off-by: Claude <noreply@anthropic.com>
naomi-robusta
added a commit
that referenced
this pull request
Jul 28, 2026
…60% fewer prompt tokens) (#2301) ## What Condenses the Holmes prompts for Claude 4.x-class models while keeping every behavioral contract. Newer models over-trigger on repetition, ALL-CAPS emphasis and consequence language ("VIOLATION CONSEQUENCES", "FAILURE = INCOMPLETE INVESTIGATION"), so this cuts the scaffolding and keeps the domain facts, output contracts and non-obvious constraints. **Sizes (cl100k tokens, representative server config with core_investigation enabled):** - System prompt: **~15.3K → ~6.1K (−60%)**; full request payload incl. tool schemas: **−45%** - `generic_ask.jinja2` TodoWrite/phases section: 2,536 → 331 tokens - `investigator_instructions.jinja2`: 13.9KB → 1.7KB (was a verbatim coding-assistant prompt with dark-mode/npm examples and a "4 lines max" rule conflicting with the style guide) - `conversation_history_compaction.jinja2`: ~1,200 → ~390 tokens - Disabled-toolsets listing: one line per toolset; docs-URL pattern stated once, only exceptions listed - bash prefix guidance ~40% smaller; prometheus/robusta llm-instructions deduped (rules stated 2-3× → once); `fetch_pod_logs` + kubernetes.yaml tool schemas tightened; per-message TodoWrite reminder ~120 → ~45 tokens ## Verification (live evals via OpenRouter, k3s + Elasticsearch Cloud) Baselined every suite on the old prompts first, then re-ran per change group. Models chosen by production traffic distribution (opus-4.6 #1, opus-4.7 #2, haiku-4.5 #3, sonnet-4.6). | Suite | opus-4.6 | opus-4.7 | haiku-4.5 | sonnet-4.6 | |---|---|---|---|---| | multicluster/transparency ES (13) | 12/13 baseline → 12/13 (same pre-existing 235 failure) | 13/13 | 9/13 baseline → 7–9/13 (within its 7–11 variance band, n=7 runs) | 10/13 | | k8s regression (13) | 12/13 baseline → 12–13/13 (`24_misconfigured_pvc` baseline failure now passes) | — | — | — | | ES quirks/skills (6) | 6/6 → 6/6 | 6/7 | n/a (0/6 at baseline, pre-existing) | 6/7 | | bash prefix rules (9) | 9/9 | 8/9 (205 flaky ~1/3 runs on all configs) | — | — | | compaction (7) | 7/7 baseline → 7/7 | — | — | — | A more aggressive rewrite (extra ~480 tokens saved) was built and **rejected by evals**: haiku-4.5 averaged 6.75/13 with it vs 8.6/13 (n=9 runs), while opus was unaffected — weaker models need the extra structure, so this PR keeps the version that's safe across the whole prod model mix. **Not eval-verified here** (no creds/infra in the dev sandbox): datadog/newrelic/coralogix instructions (untouched), OpenSearch PPL query-assist docs (untouched — at 56KB it's the biggest remaining candidate, needs a PPL eval first), prometheus eval fixtures can't run in the sandbox so the prometheus edits are strictly dedup-only. The `/eval` comment on this PR covers these areas in CI. **Known-flaky tests observed during this work** (fail ~1/3 runs on baseline AND reduced prompts): `254_elasticsearch_dr_test_log_check`, `205_bash_deployment_logs_all_pods`; `235_elasticsearch_cluster_mismatch` fails consistently on every config/model except opus-4.7. Worth hardening separately. ## Test changes - `test_toolsets_instructions.py`, `test_fetch_logs.py`: string pins updated to the new wording (same semantics). 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01XgDuG2hb6jmSqYpD7iGUkd --- _Generated by [Claude Code](https://claude.ai/code/session_01XgDuG2hb6jmSqYpD7iGUkd)_ <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Improvements** * Strengthened TodoWrite multi-step workflows, task state handling, and evidence-backed completion with clearer blocker follow-ups. * Refined user-facing prompt and tool instructions for log retrieval, including consistent summaries, pod/namespace identification, and timeframe behavior. * Updated guidance for Bash command approval/allow-lists, Kubernetes/Prometheus/Robusta investigations, permissions-focused troubleshooting, and conversation-history compaction. * Improved messaging when toolsets are disabled or unavailable, including clearer next actions. * **Tests** * Updated prompt wording assertions and added a new TodoWrite multi-step audit regression fixture. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Claude <noreply@anthropic.com> Co-authored-by: Claude <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.