Skip to content

Readme update and toolset fix - #3

Merged
aantn merged 1 commit into
masterfrom
readme-update
May 30, 2024
Merged

aantn merged 1 commit into
masterfrom
readme-update

Conversation

@pavangudiwada

Copy link
Copy Markdown
Contributor

No description provided.

@pavangudiwada
pavangudiwada requested a review from aantn May 30, 2024 14:26
@CLAassistant

CLAassistant commented May 30, 2024 •

Copy link
Copy Markdown

CLA assistant check
All committers have signed the CLA.

@aantn aantn left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Awesome.

@aantn
aantn merged commit 5cedc8a into master May 30, 2024
@aantn
aantn deleted the readme-update branch May 30, 2024 14:40
@sebv004 sebv004 mentioned this pull request Oct 16, 2025
arikalon1 added a commit that referenced this pull request Dec 7, 2025
When a Todo task fails, the LLM returns a "Failed" state. But since this
state is not present, an error is raised and thrown in the terminal.

```
Running tool #3 TodoWrite: Update investigation tasks                                                                                                    
error using todowrite tool                                                                                                                               
Traceback (most recent call last):                                                                                                                       
  File "/Users/pavan/Documents/repos/holmesgpt/holmes/plugins/toolsets/investigator/core_investigation.py", line 102, in _invoke                         
    tasks = parse_tasks(todos_data=todos_data)                                                                                                           
            ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^                                                                                                           
  File "/Users/pavan/Documents/repos/holmesgpt/holmes/plugins/toolsets/investigator/core_investigation.py", line 30, in parse_tasks                      
    status=TaskStatus(todo_item.get("status", "pending")),                                                                                               
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^                                                                                                
  File "/Users/pavan/.pyenv/versions/3.12.1/lib/python3.12/enum.py", line 744, in __call__                                                               
    return cls.__new__(cls, value)                                                                                                                       
           ^^^^^^^^^^^^^^^^^^^^^^^                                                                                                                       
  File "/Users/pavan/.pyenv/versions/3.12.1/lib/python3.12/enum.py", line 1158, in __new__                                                               
    raise ve_exc                                                                                                                                         
ValueError: 'failed' is not a valid TaskStatus                                                                                                           
```

Handle this error and display it properly see **status** for failed
tasks

<img width="2048" height="614" alt="image"
src="https://github.com/user-attachments/assets/77c373da-a723-4ecf-a57f-902539a63c64"
/>

Co-authored-by: arik <alon.arik@gmail.com>
aantn pushed a commit that referenced this pull request Dec 10, 2025
When a Todo task fails, the LLM returns a "Failed" state. But since this
state is not present, an error is raised and thrown in the terminal.

```
Running tool #3 TodoWrite: Update investigation tasks
error using todowrite tool
Traceback (most recent call last):
  File "/Users/pavan/Documents/repos/holmesgpt/holmes/plugins/toolsets/investigator/core_investigation.py", line 102, in _invoke
    tasks = parse_tasks(todos_data=todos_data)
            ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/Users/pavan/Documents/repos/holmesgpt/holmes/plugins/toolsets/investigator/core_investigation.py", line 30, in parse_tasks
    status=TaskStatus(todo_item.get("status", "pending")),
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/Users/pavan/.pyenv/versions/3.12.1/lib/python3.12/enum.py", line 744, in __call__
    return cls.__new__(cls, value)
           ^^^^^^^^^^^^^^^^^^^^^^^
  File "/Users/pavan/.pyenv/versions/3.12.1/lib/python3.12/enum.py", line 1158, in __new__
    raise ve_exc
ValueError: 'failed' is not a valid TaskStatus
```

Handle this error and display it properly see **status** for failed
tasks

<img width="2048" height="614" alt="image"
src="https://github.com/user-attachments/assets/77c373da-a723-4ecf-a57f-902539a63c64"
/>

Co-authored-by: arik <alon.arik@gmail.com>
Signed-off-by: Robusta Runner <aantny@gmail.com>
goyamegh added a commit to goyamegh/holmesgpt that referenced this pull request Mar 9, 2026
- Rename OTEL_AWS_SERVICE → HOLMES_AWS_OSIS_SERVICE with backwards-compat
  fallback (Comment HolmesGPT#10, svrnm)
- Align OTEL_DEBUG with OTEL spec OTEL_LOG_LEVEL=debug with backwards-compat
  fallback (Comment HolmesGPT#9, svrnm)
- Add ml-commons AgentTracer.java GitHub permalink (Comment HolmesGPT#3, kylehounslow)
- Document needs_aws_auth() as single source of truth with consumer list
  (Comment HolmesGPT#4, kylehounslow)
- Clarify otel_logging.py: logs NOT exported via OTLP, naming avoids
  shadowing builtin logging (Comment HolmesGPT#5, kylehounslow)
- Explain experimental/ placement: API evolving, removable via try/except
  no-op fallbacks (Comment HolmesGPT#6, kylehounslow)
- Clarify server.py middleware: tracing/metrics init in init_otel() above,
  middleware only handles per-request spans (Comment HolmesGPT#7, kylehounslow)
- Split env var docs into Standard OTEL / Holmes-Specific tables, add
  missing vars, add OSIS hyperlink (Comments HolmesGPT#1, HolmesGPT#9, HolmesGPT#10)
- Fix no-op tracer fallback in server-agui.py for when OTEL is unavailable

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Signed-off-by: Megha Goyal <goyamegh@amazon.com>
aantn pushed a commit that referenced this pull request Mar 21, 2026
Address 5 CodeRabbit review comments from PR #1811:

Issue #3 - Cache key inadequacy: create_tool_executor now stores a cache
key (tags, enable_all) alongside the executor and validates it on cache
hit. Callers with different parameters get a fresh executor.

Issue #4/#5 - Change detection misses added/removed toolsets:
refresh_toolsets_and_get_changes now compares name sets between old and
new toolset lists, reporting additions and removals alongside status
transitions. refresh_tool_executor always updates the cached executor
(not just when status changes exist).

Issue #6 - Race condition: Added threading.Lock around all reads and
writes of _cached_tool_executor. Added Config.cached_tool_executor
property for thread-safe external access. Updated server.py to use it.

Issue #7 - Cold-cache refresh: refresh_tool_executor now uses
FORCE_REFRESH on cold start instead of ENABLED, ensuring live
prerequisite checks rather than stale disk cache.

https://claude.ai/code/session_01Qr7wtBgYGnh1G6Bq2wPp79
Signed-off-by: Claude <noreply@anthropic.com>
aantn pushed a commit that referenced this pull request Mar 21, 2026
- Embed timestamp as hidden HTML comment in buildBody()
- Extract and display date in "Mar 21, 14:32 UTC" format in run summaries
- Add sequential run numbers (#1, #2, ...) to previous runs (newest = highest)
- Strip existing run numbers on re-parse to avoid double-numbering
- Bold commit SHA with __underscores__ for visual clarity
- Handle multiline header matching for timestamp comment prefix

Example: 📜 #3 · Run @ __8d93be9__ (#23379683267) — Mar 21, 14:32 UTC

https://claude.ai/code/session_01UoX9TpBtKvD4cjogGmXFrt
Signed-off-by: Claude <noreply@anthropic.com>
aantn pushed a commit that referenced this pull request Jun 12, 2026
…the proxy hop

Run #3's surfaced error showed the claude CLI reporting 'Not logged in' and
every turn erroring. Cause: CI sets the ANTHROPIC_API_KEY secret to empty
(it uses OpenRouter), and the engine/bootstrap inherited that empty value;
the CLI refuses to run without a non-empty key. The local proxy ignores the
key (it authenticates to the provider via the model_list), so a placeholder is
correct. Engine and bootstrap now coerce an empty/unset ANTHROPIC_API_KEY to a
non-empty placeholder for the CLI<->proxy hop.

Signed-off-by: Claude <noreply@anthropic.com>
naomi-robusta pushed a commit that referenced this pull request Jul 20, 2026
Second optimization pass, keeping only changes that held up under repeated
eval runs on the production model mix:

- conversation_history_compaction.jinja2: 1,200 -> ~390 tokens; same section
  contract. Verified with the compaction eval suite: 7/7 before, 7/7 after
  (opus-4.6).
- _toolsets_instructions.jinja2: drop the per-entry status label on disabled
  toolsets (header states it once) and merge the missing-integration guidance
  into one paragraph.
- investigator_instructions.jinja2: further condensed TodoWrite mechanics.
- fetch_pod_logs tool schema: compressed start/end_time and filter parameter
  descriptions (~200 tokens per prompt with kubernetes/logs enabled).
- Updated string-pin unit tests (fetch_logs wording, disabled-list format).

An even more aggressive generic_ask rewrite (extra ~480 tokens saved) was
built and evaluated but NOT kept: over 9 runs of the multicluster suite,
haiku-4.5 averaged 6.75/13 with it vs 8.6/13 with the current version, while
opus was unaffected. Since haiku is the #3 production model, the current
generic_ask stays.

Eval evidence for the final state (this commit):
- multicluster suite: opus-4.7 13/13, sonnet-4.6 10/13 (only pre-existing
  235 + known-flaky 254 + one skills test), haiku-4.5 9/13 and 7/13
  (within its 7-11 baseline band), opus-4.6 12/13 in prior rounds
- k8s regression: opus-4.6 12/13 (only the known-flaky 254;
  24_misconfigured_pvc passes)
- ES quirks/skills: opus-4.6 6/6 (one flake passed on rerun),
  opus-4.7 6/7, sonnet-4.6 6/7
- bash prefix suite: opus-4.6 9/9, opus-4.7 8/9 (205 is flaky ~1/3 runs on
  all configs including baseline)
- compaction: 7/7

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XgDuG2hb6jmSqYpD7iGUkd
Signed-off-by: Claude <noreply@anthropic.com>
naomi-robusta added a commit that referenced this pull request Jul 28, 2026
…60% fewer prompt tokens) (#2301)

## What

Condenses the Holmes prompts for Claude 4.x-class models while keeping
every behavioral contract. Newer models over-trigger on repetition,
ALL-CAPS emphasis and consequence language ("VIOLATION CONSEQUENCES",
"FAILURE = INCOMPLETE INVESTIGATION"), so this cuts the scaffolding and
keeps the domain facts, output contracts and non-obvious constraints.

**Sizes (cl100k tokens, representative server config with
core_investigation enabled):**
- System prompt: **~15.3K → ~6.1K (−60%)**; full request payload incl.
tool schemas: **−45%**
- `generic_ask.jinja2` TodoWrite/phases section: 2,536 → 331 tokens
- `investigator_instructions.jinja2`: 13.9KB → 1.7KB (was a verbatim
coding-assistant prompt with dark-mode/npm examples and a "4 lines max"
rule conflicting with the style guide)
- `conversation_history_compaction.jinja2`: ~1,200 → ~390 tokens
- Disabled-toolsets listing: one line per toolset; docs-URL pattern
stated once, only exceptions listed
- bash prefix guidance ~40% smaller; prometheus/robusta llm-instructions
deduped (rules stated 2-3× → once); `fetch_pod_logs` + kubernetes.yaml
tool schemas tightened; per-message TodoWrite reminder ~120 → ~45 tokens

## Verification (live evals via OpenRouter, k3s + Elasticsearch Cloud)

Baselined every suite on the old prompts first, then re-ran per change
group. Models chosen by production traffic distribution (opus-4.6 #1,
opus-4.7 #2, haiku-4.5 #3, sonnet-4.6).

| Suite | opus-4.6 | opus-4.7 | haiku-4.5 | sonnet-4.6 |
|---|---|---|---|---|
| multicluster/transparency ES (13) | 12/13 baseline → 12/13 (same
pre-existing 235 failure) | 13/13 | 9/13 baseline → 7–9/13 (within its
7–11 variance band, n=7 runs) | 10/13 |
| k8s regression (13) | 12/13 baseline → 12–13/13
(`24_misconfigured_pvc` baseline failure now passes) | — | — | — |
| ES quirks/skills (6) | 6/6 → 6/6 | 6/7 | n/a (0/6 at baseline,
pre-existing) | 6/7 |
| bash prefix rules (9) | 9/9 | 8/9 (205 flaky ~1/3 runs on all configs)
| — | — |
| compaction (7) | 7/7 baseline → 7/7 | — | — | — |

A more aggressive rewrite (extra ~480 tokens saved) was built and
**rejected by evals**: haiku-4.5 averaged 6.75/13 with it vs 8.6/13 (n=9
runs), while opus was unaffected — weaker models need the extra
structure, so this PR keeps the version that's safe across the whole
prod model mix.

**Not eval-verified here** (no creds/infra in the dev sandbox):
datadog/newrelic/coralogix instructions (untouched), OpenSearch PPL
query-assist docs (untouched — at 56KB it's the biggest remaining
candidate, needs a PPL eval first), prometheus eval fixtures can't run
in the sandbox so the prometheus edits are strictly dedup-only. The
`/eval` comment on this PR covers these areas in CI.

**Known-flaky tests observed during this work** (fail ~1/3 runs on
baseline AND reduced prompts): `254_elasticsearch_dr_test_log_check`,
`205_bash_deployment_logs_all_pods`;
`235_elasticsearch_cluster_mismatch` fails consistently on every
config/model except opus-4.7. Worth hardening separately.

## Test changes
- `test_toolsets_instructions.py`, `test_fetch_logs.py`: string pins
updated to the new wording (same semantics).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01XgDuG2hb6jmSqYpD7iGUkd

---
_Generated by [Claude
Code](https://claude.ai/code/session_01XgDuG2hb6jmSqYpD7iGUkd)_

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Improvements**
* Strengthened TodoWrite multi-step workflows, task state handling, and
evidence-backed completion with clearer blocker follow-ups.
* Refined user-facing prompt and tool instructions for log retrieval,
including consistent summaries, pod/namespace identification, and
timeframe behavior.
* Updated guidance for Bash command approval/allow-lists,
Kubernetes/Prometheus/Robusta investigations, permissions-focused
troubleshooting, and conversation-history compaction.
* Improved messaging when toolsets are disabled or unavailable,
including clearer next actions.
* **Tests**
* Updated prompt wording assertions and added a new TodoWrite multi-step
audit regression fixture.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Claude <noreply@anthropic.com>
Co-authored-by: Claude <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants