More doc improvements - #1094
Conversation
WalkthroughThis pull request updates documentation across AI provider guides, installation instructions, and evaluation tooling. Changes include adding model recommendations, updating model references (gpt-4o→gpt-4.1, gpt-4o-mini→gpt-5), expanding Ollama configuration with OpenAI-compatible gateway options, restructuring evaluation documentation, and replacing demo content with Loom video embeds. Changes
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~20 minutes
Possibly related PRs
Suggested reviewers
Pre-merge checks and finishing touches❌ Failed checks (1 warning, 2 inconclusive)
✨ Finishing touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Actionable comments posted: 1
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (2)
docs/installation/kubernetes-installation.md (1)
148-148: Update outdated model name in Usage example – The test curl request still references"model": "gpt-4o-mini", which should be updated to"model": "gpt-5"to align with the modelList configuration changes and PR objectives (gpt-4o-mini → gpt-5).Apply this diff to fix the model reference:
- -d '{"ask": "list pods in namespace default?", "model": "gpt-4o-mini"}' + -d '{"ask": "list pods in namespace default?", "model": "gpt-5"}'docs/development/evaluations/running-evals.md (1)
155-156: Update outdated model name in multi-model example – Line 155 referencesgpt-4o-miniin the comma-separated model list, which should be updated togpt-5to align with the PR's model migration (gpt-4o-mini → gpt-5).Apply this diff to fix the model reference:
- RUN_LIVE=true MODEL=gpt-4o,gpt-4o-mini \ + RUN_LIVE=true MODEL=gpt-4o,gpt-5 \
📜 Review details
Configuration used: CodeRabbit UI
Review profile: CHILL
Plan: Pro
📒 Files selected for processing (8)
docs/ai-providers/aws-bedrock.md(1 hunks)docs/ai-providers/ollama.md(2 hunks)docs/development/evaluations/adding-evals.md(3 hunks)docs/development/evaluations/running-evals.md(2 hunks)docs/installation/.nav.yml(1 hunks)docs/installation/cli-installation.md(6 hunks)docs/installation/kubernetes-installation.md(2 hunks)docs/installation/ui-installation.md(2 hunks)
🧰 Additional context used
📓 Path-based instructions (1)
docs/**/*.md
📄 CodeRabbit inference engine (CLAUDE.md)
In MkDocs content, always add a blank line between a header or bold text and a following list so lists render correctly
Files:
docs/ai-providers/aws-bedrock.mddocs/installation/kubernetes-installation.mddocs/development/evaluations/adding-evals.mddocs/development/evaluations/running-evals.mddocs/installation/cli-installation.mddocs/installation/ui-installation.mddocs/ai-providers/ollama.md
🧠 Learnings (9)
📚 Learning: 2025-08-05T06:14:39.523Z
Learnt from: aantn
Repo: robusta-dev/holmesgpt PR: 783
File: tests/llm/fixtures/test_ask_holmes/100_historical_logs/payment-api.yaml:49-70
Timestamp: 2025-08-05T06:14:39.523Z
Learning: For evaluation test fixtures in the holmesgpt project, security contexts and security hardening are not priorities. The focus should be on functionality and test reliability rather than adding security configurations to Kubernetes manifests used in evals.
Applied to files:
docs/development/evaluations/adding-evals.md
📚 Learning: 2025-08-30T18:12:58.187Z
Learnt from: CR
Repo: robusta-dev/holmesgpt PR: 0
File: holmes/plugins/runbooks/CLAUDE.md:0-0
Timestamp: 2025-08-30T18:12:58.187Z
Learning: Applies to holmes/plugins/runbooks/**/*.md : In Recommended Remediation Steps, include Immediate Actions, Permanent Solutions, Verification Steps, Documentation References, Escalation Criteria, and Post-Remediation Monitoring.
Applied to files:
docs/development/evaluations/adding-evals.md
📚 Learning: 2025-08-13T05:57:40.420Z
Learnt from: mainred
Repo: robusta-dev/holmesgpt PR: 829
File: holmes/plugins/runbooks/runbook-format.prompt.md:11-22
Timestamp: 2025-08-13T05:57:40.420Z
Learning: In holmes/plugins/runbooks/runbook-format.prompt.md, the user (mainred) prefers to keep the runbook step specifications simple without detailed orchestration metadata like step IDs, dependencies, retry policies, timeouts, and storage variables. The current format with Action, Function Description, Parameters, Expected Output, and Success/Failure Criteria is sufficient for their AI agent troubleshooting use case.
Applied to files:
docs/development/evaluations/adding-evals.md
📚 Learning: 2025-07-08T08:45:41.069Z
Learnt from: nherment
Repo: robusta-dev/holmesgpt PR: 610
File: .github/workflows/llm-evaluation.yaml:39-42
Timestamp: 2025-07-08T08:45:41.069Z
Learning: The robusta-dev/holmesgpt codebase has comprehensive existing validation for Azure environment variables (AZURE_API_BASE, AZURE_API_KEY, AZURE_API_VERSION) and MODEL in tests/llm/utils/classifiers.py, tests/llm/conftest.py, and holmes/core/llm.py. Don't suggest adding redundant validation logic.
Applied to files:
docs/development/evaluations/adding-evals.mddocs/development/evaluations/running-evals.mddocs/ai-providers/ollama.md
📚 Learning: 2025-07-02T10:27:17.231Z
Learnt from: Sheeproid
Repo: robusta-dev/holmesgpt PR: 586
File: tests/llm/fixtures/test_ask_holmes/03_what_is_the_command_to_port_forward/test_case.yaml:4-4
Timestamp: 2025-07-02T10:27:17.231Z
Learning: In LLM-as-judge test cases for HolmesGPT, expected outputs should be descriptive rather than prescriptive when testing for flexible responses like port numbers. Using specific values in expected outputs can cause unnecessary test failures when the AI generates different but equally valid responses.
Applied to files:
docs/development/evaluations/adding-evals.mddocs/development/evaluations/running-evals.md
📚 Learning: 2025-10-05T13:01:12.288Z
Learnt from: CR
Repo: robusta-dev/holmesgpt PR: 0
File: CLAUDE.md:0-0
Timestamp: 2025-10-05T13:01:12.288Z
Learning: Applies to tests/llm/**/*.sh : When scripting kubectl operations in evals, never use a bare 'kubectl wait' immediately after creating resources; use a retry loop to avoid race conditions
Applied to files:
docs/development/evaluations/adding-evals.md
📚 Learning: 2025-10-05T13:01:12.288Z
Learnt from: CR
Repo: robusta-dev/holmesgpt PR: 0
File: CLAUDE.md:0-0
Timestamp: 2025-10-05T13:01:12.288Z
Learning: Applies to tests/llm/**/*.yaml : In eval manifests, ALWAYS use Kubernetes Secrets for scripts rather than inline manifests or ConfigMaps to prevent script exposure via kubectl describe
Applied to files:
docs/development/evaluations/adding-evals.md
📚 Learning: 2025-10-05T13:01:12.288Z
Learnt from: CR
Repo: robusta-dev/holmesgpt PR: 0
File: CLAUDE.md:0-0
Timestamp: 2025-10-05T13:01:12.288Z
Learning: Applies to tests/**/*.py : Only use pytest markers that are defined in pyproject.toml; never introduce undefined markers/tags
Applied to files:
docs/development/evaluations/adding-evals.md
📚 Learning: 2025-10-05T13:01:12.288Z
Learnt from: CR
Repo: robusta-dev/holmesgpt PR: 0
File: CLAUDE.md:0-0
Timestamp: 2025-10-05T13:01:12.288Z
Learning: Applies to tests/llm/**/*.yaml : Eval resources must use neutral names, each test must use a dedicated namespace 'app-<testid>', and all pod names must be unique across tests
Applied to files:
docs/development/evaluations/adding-evals.md
🪛 LanguageTool
docs/installation/cli-installation.md
[style] ~80-~80: Using many exclamation marks might seem excessive (in this case: 3 exclamation marks for a text that’s 1434 characters long)
Context: ...providers/index.md) for more options). !!! tip "Which Model to Use" We highly ...
(EN_EXCESSIVE_EXCLAMATION)
docs/ai-providers/ollama.md
[grammar] ~6-~6: Use a hyphen to join words.
Context: ...duce inconsistent results. Only [LiteLLM supported Ollama models](https://docs.li...
(QB_NEW_EN_HYPHEN)
🔇 Additional comments (10)
docs/ai-providers/aws-bedrock.md (1)
5-8: Tip block structure looks good – The blank line between the tip and the "Setup" header follows proper MkDocs formatting.docs/installation/kubernetes-installation.md (1)
42-50: Model name updates are consistent – The changes from gpt-4o to gpt-4.1 and introduction of gpt-5 are applied correctly across both OpenAI examples and the Multiple Providers section.Also applies to: 113-124
docs/ai-providers/ollama.md (1)
21-39: Ollama CLI configuration expansion is well-structured – The addition of MODEL environment variable option and OpenAI-compatible gateway path provides good flexibility for users facing compatibility issues.docs/installation/cli-installation.md (2)
77-103: Provider-specific quick start is well-organized – The new section with model recommendations tip and detailed Anthropic provider example (lines 81-103) provides excellent guidance. The tip block is properly separated from content.
119-126: OpenAI examples properly reference updated models – References to gpt-4.1 and gpt-5 are consistent and include helpful comments about defaults and options.docs/installation/ui-installation.md (1)
139-175: Video embed structure and Get Started section look good – The three Loom video tabs are properly formatted with correct iframe markup, and the simplified Get Started steps maintain clarity while reducing verbosity.docs/installation/.nav.yml (1)
3-3: Navigation label correctly updated – The change from "Install UI/TUI" to "Install UI/Slack/K9s" aligns with the heading update in ui-installation.md and accurately reflects the supported interfaces.docs/development/evaluations/running-evals.md (1)
28-52: Quick Start examples are comprehensive and well-explained – The multi-model comparison with clear notes about model performance differences (Sonnet 4.5 vs weaker models) helps users understand benchmarking effectively.docs/development/evaluations/adding-evals.md (2)
13-37: Quick Start examples demonstrate proper model selection – The examples showing Sonnet 4.5 performance comparison with weaker models, plus multi-model testing, help users understand eval effectiveness with different AI providers.
81-117: test_case.yaml configuration section is thorough – Clear distinction between required and optional fields, with practical examples for advanced configurations (runbooks, toolsets, mock_policy) helps users create robust evals.
No description provided.