Skip to content

Add benchmarks for Claude 4.5 - #1017

Merged
aantn merged 5 commits into
masterfrom
claude-4.5
Sep 30, 2025
Merged

aantn merged 5 commits into
masterfrom
claude-4.5

Conversation

@aantn

@aantn aantn commented Sep 30, 2025

Copy link
Copy Markdown
Collaborator

Also includes improvements to mkdocs so new benchmark results are automatically listed. (This requires changing the way navigation is done in mkdocs and switching to a plugin.)

@aantn
aantn requested a review from Sheeproid September 30, 2025 06:13
@coderabbitai

coderabbitai Bot commented Sep 30, 2025 •

Copy link
Copy Markdown
Contributor

Walkthrough

Introduces timestamped benchmark report generation in CI, adjusts report header logic based on filename, restructures docs navigation into per-directory .nav.yml files using mkdocs-awesome-nav, updates MkDocs config and dev dependencies, and refreshes evaluation docs including history index and a new timestamped results file.

Changes

Cohort / File(s) Summary
CI benchmark report history
.github/workflows/eval-benchmarks.yaml
Adds TIMESTAMP var and writes an additional results file to docs/development/evaluations/history/results_<TIMESTAMP>.md while retaining latest-results generation.
Eval report header logic
scripts/generate_eval_report.py
If output filename matches results_YYYYMMDD_HHMMSS.md, sets title to formatted date; otherwise uses default header.
MkDocs plugin and nav restructure
mkdocs.yml, pyproject.toml
Adds awesome-nav plugin and dev dependency mkdocs-awesome-nav. Removes top-level nav from mkdocs.yml in favor of per-directory .nav.yml files.
Root docs nav
docs/.nav.yml
Introduces site-level nav mapping core sections (Installation, Walkthrough, AI Providers, Data Sources, Development, Reference, Community).
AI providers nav
docs/ai-providers/.nav.yml
Adds provider-specific nav entries (Anthropic, Bedrock, Azure OpenAI, Gemini, Vertex AI, Ollama, OpenAI, OpenAI-Compatible, Robusta AI, Multi-Provider).
Data sources nav
docs/data-sources/.nav.yml, docs/data-sources/builtin-toolsets/.nav.yml
Adds data-sources nav and a comprehensive builtin-toolsets nav list mapping labels to markdown files.
Development/evaluations nav
docs/development/.nav.yml, docs/development/evaluations/.nav.yml, docs/development/evaluations/history/.nav.yml
Adds nav entries for development, evaluations, and history (including wildcard to list all history pages).
Evaluations docs content updates
docs/development/evaluations/index.md, docs/development/evaluations/history/index.md, docs/development/evaluations/history/results_20250928_001434.md
Updates links to explicit history/index.md, simplifies history index content, and adds a timestamped results page with updated metadata.
Installation nav
docs/installation/.nav.yml
Adds nav entries for CLI, UI/TUI, Helm Chart, and Python SDK installation.
Reference nav
docs/reference/.nav.yml
Adds reference section entries (Env Vars, Helm Configuration, HTTP API, Slash Commands, Troubleshooting).
Walkthrough nav
docs/walkthrough/.nav.yml
Adds walkthrough index and specific guides (Interactive Mode, CI/CD Troubleshooting, Prometheus Alerts, AKS MCP).

Sequence Diagram(s)

sequenceDiagram
  autonumber
  participant Dev as GitHub Actions Workflow
  participant Script as generate_eval_report.py
  participant FS as Docs Filesystem

  Dev->>Script: Run report generator (latest)
  Script->>FS: Write latest-results.md
  Note over Script: Default header used (non-matching filename)

  Dev->>Dev: Compute TIMESTAMP
  Dev->>Script: Run report generator (results_<TIMESTAMP>.md)
  Script->>Script: Filename matches results_YYYYMMDD_HHMMSS.md
  Script->>Script: Set header to formatted date
  Script->>FS: Write history/results_<TIMESTAMP>.md

  Note over Dev,FS: MkDocs uses awesome-nav with per-directory .nav.yml to surface new history entry
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Suggested reviewers

  • Sheeproid
  • arikalon1

Pre-merge checks and finishing touches

❌ Failed checks (2 warnings)
Check name Status Explanation Resolution
Title Check ⚠️ Warning The title “Add benchmarks for Claude 4.5” focuses on introducing a new model’s benchmarks, but the changes in this pull request primarily update the MkDocs navigation, introduce a timestamped history report, and adjust report generation logic without explicitly adding Claude 4.5 benchmarks, so it does not accurately reflect the main modifications. Please update the title to highlight the key changes—such as adding timestamped benchmark history files and switching to the MkDocs navigation plugin—rather than referencing Claude 4.5 benchmarks which are not clearly included.
Docstring Coverage ⚠️ Warning Docstring coverage is 0.00% which is insufficient. The required threshold is 80.00%. You can run @coderabbitai generate docstrings to improve docstring coverage.
✅ Passed checks (1 passed)
Check name Status Explanation
Description Check ✅ Passed The description clearly references the MkDocs navigation improvements and automatic listing of benchmark results, which directly correspond to the updates in navigation configuration and the new plugin usage present in this changeset.
✨ Finishing touches
  • 📝 Generate Docstrings
🧪 Generate unit tests
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch claude-4.5

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share
🧪 Early access (Sonnet 4.5): enabled

We are currently testing the Sonnet 4.5 model, which is expected to improve code review quality. However, this model may lead to increased noise levels in the review comments. Please disable the early access features if the noise level causes any inconvenience.

Note:

  • Public repositories are always opted into early access features.
  • You can enable or disable early access features from the CodeRabbit UI or by updating the CodeRabbit configuration file.

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

📜 Review details

Configuration used: CodeRabbit UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between d00778f and 9987131.

⛔ Files ignored due to path filters (1)
  • poetry.lock is excluded by !**/*.lock
📒 Files selected for processing (17)
  • .github/workflows/eval-benchmarks.yaml (1 hunks)
  • docs/.nav.yml (1 hunks)
  • docs/ai-providers/.nav.yml (1 hunks)
  • docs/data-sources/.nav.yml (1 hunks)
  • docs/data-sources/builtin-toolsets/.nav.yml (1 hunks)
  • docs/development/.nav.yml (1 hunks)
  • docs/development/evaluations/.nav.yml (1 hunks)
  • docs/development/evaluations/history/.nav.yml (1 hunks)
  • docs/development/evaluations/history/index.md (1 hunks)
  • docs/development/evaluations/history/results_20250928_001434.md (1 hunks)
  • docs/development/evaluations/index.md (2 hunks)
  • docs/installation/.nav.yml (1 hunks)
  • docs/reference/.nav.yml (1 hunks)
  • docs/walkthrough/.nav.yml (1 hunks)
  • mkdocs.yml (1 hunks)
  • pyproject.toml (1 hunks)
  • scripts/generate_eval_report.py (1 hunks)
🧰 Additional context used
📓 Path-based instructions (2)
docs/**/*.md

📄 CodeRabbit inference engine (CLAUDE.md)

In MkDocs docs, always add a blank line between a header/bold text and a following list to render correctly

Files:

  • docs/development/evaluations/history/index.md
  • docs/development/evaluations/index.md
  • docs/development/evaluations/history/results_20250928_001434.md
pyproject.toml

📄 CodeRabbit inference engine (CLAUDE.md)

Define and maintain the canonical list of pytest markers in pyproject.toml

Files:

  • pyproject.toml
🪛 Ruff (0.13.1)
scripts/generate_eval_report.py

1364-1364: Unpacked variable hour is never used

Prefix it with an underscore or any other dummy variable pattern

(RUF059)


1364-1364: Unpacked variable minute is never used

Prefix it with an underscore or any other dummy variable pattern

(RUF059)


1364-1364: Unpacked variable second is never used

Prefix it with an underscore or any other dummy variable pattern

(RUF059)

⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (2)
  • GitHub Check: Pre-commit checks
  • GitHub Check: llm_evals
🔇 Additional comments (18)
pyproject.toml (2)

79-79: LGTM! Dev dependency addition aligns with PR objectives.

The addition of mkdocs-awesome-nav plugin properly supports the new navigation approach described in the PR objectives. The dependency is correctly placed in the dev group with an appropriate version constraint.


95-123: Pytest markers are well-maintained.

The canonical list of pytest markers in pyproject.toml is comprehensive and well-documented, following the coding guidelines.

docs/reference/.nav.yml (1)

1-6: All referenced markdown files are present. No further changes needed.

docs/development/evaluations/history/results_20250928_001434.md (1)

1-7: Metadata update looks correct.

Heading timestamp, generated time, and judge model attribution all line up with the new benchmarking flow. Nothing else to block here.

docs/ai-providers/.nav.yml (1)

1-13: LGTM! Clean navigation structure.

The AI providers navigation configuration is well-organized and follows a logical structure. All entries are properly formatted for the mkdocs-awesome-nav plugin.

docs/walkthrough/.nav.yml (1)

1-7: LGTM! Walkthrough navigation is well-structured.

The navigation configuration properly organizes the walkthrough documentation with clear, descriptive titles.

docs/data-sources/.nav.yml (1)

1-7: LGTM! Data sources navigation properly configured.

The navigation structure correctly includes both direct markdown files and a directory reference for nested navigation (builtin-toolsets), which aligns with the mkdocs-awesome-nav plugin's directory-based navigation approach.

.github/workflows/eval-benchmarks.yaml (1)

118-129: Verify artifact upload includes timestamped report.

The timestamped benchmark report is generated but not included in the artifact upload (line 137 only uploads latest-results.md). Since the git commit step (lines 140-153) is currently disabled for testing, these timestamped files may be lost.

Consider updating the artifact upload to preserve timestamped reports during the testing phase:

       - name: Upload eval results
         if: always()
         uses: actions/upload-artifact@v4
         with:
           name: eval-results-${{ github.run_id }}
           path: |
             docs/development/evaluations/latest-results.md
+            docs/development/evaluations/history/results_*.md

This ensures timestamped reports are accessible even while the automatic commit is disabled.

docs/data-sources/builtin-toolsets/.nav.yml (1)

1-30: LGTM! Navigation structure is well-organized.

The navigation configuration properly defines the Built-in Toolsets section with clear, descriptive labels for each tool. The YAML structure is valid and entries are logically organized.

docs/development/.nav.yml (1)

1-4: LGTM! Clean navigation structure.

The navigation configuration correctly defines the Development section with an index and reference to the Evaluations subdirectory.

docs/.nav.yml (1)

1-10: LGTM! Top-level navigation structure is well-organized.

The root navigation configuration properly defines the main sections of the documentation site with a logical flow from installation through reference and community resources.

docs/development/evaluations/index.md (2)

69-69: LGTM! Link update is correct.

The link now properly points to the history index file, ensuring consistent navigation to the historical results section.


89-89: LGTM! Link update is correct.

The historical results link now properly references the history index file, maintaining consistency with the navigation restructuring.

docs/development/evaluations/history/.nav.yml (1)

1-4: LGTM! Clever use of wildcard for dynamic content.

The navigation configuration uses the wildcard pattern to automatically include all timestamped benchmark result files in the history directory, which aligns well with the PR's goal of automatically listing new benchmark results.

docs/installation/.nav.yml (1)

1-5: LGTM!

The navigation structure is well-organized and follows the YAML format correctly. The entries provide clear labels for the installation documentation.

docs/development/evaluations/history/index.md (1)

1-5: LGTM!

The simplified documentation correctly includes blank lines between headers and content, ensuring proper MkDocs rendering.

docs/development/evaluations/.nav.yml (1)

1-7: LGTM!

The navigation structure is comprehensive and well-organized, covering all evaluation documentation sections. The reference to history directory on Line 4 will correctly expand to include the timestamped results files.

mkdocs.yml (1)

59-59: mkdocs-awesome-nav dependency confirmed
Found in pyproject.toml on line 79.

Comment thread scripts/generate_eval_report.py
@aantn
aantn enabled auto-merge (squash) September 30, 2025 06:53
@github-actions

Copy link
Copy Markdown
Contributor

Results of HolmesGPT evals

  • ask_holmes: 33/36 test cases were successful, 1 regressions, 1 setup failures
Test suite Test case Status
ask 01_how_many_pods ✅
ask 02_what_is_wrong_with_pod ✅
ask 04_related_k8s_events ✅
ask 05_image_version ✅
ask 09_crashpod ✅
ask 10_image_pull_backoff ✅
ask 110_k8s_events_image_pull ✅
ask 11_init_containers ✅
ask 13a_pending_node_selector_basic ✅
ask 14_pending_resources ✅
ask 15_failed_readiness_probe ✅
ask 17_oom_kill ✅
ask 19_detect_missing_app_details ✅
ask 20_long_log_file_search ✅
ask 24_misconfigured_pvc ✅
ask 24a_misconfigured_pvc_basic ✅
ask 28_permissions_error 🚧
ask 33_cpu_metrics_discovery ❌
ask 39_failed_toolset ✅
ask 41_setup_argo ✅
ask 42_dns_issues_steps_new_tools ⚠️
ask 43_current_datetime_from_prompt ✅
ask 45_fetch_deployment_logs_simple ✅
ask 51_logs_summarize_errors ✅
ask 53_logs_find_term ✅
ask 54_not_truncated_when_getting_pods ✅
ask 59_label_based_counting ✅
ask 60_count_less_than ✅
ask 61_exact_match_counting ✅
ask 63_fetch_error_logs_no_errors ✅
ask 79_configmap_mount_issue ✅
ask 83_secret_not_found ✅
ask 86_configmap_like_but_secret ✅
ask 93_calling_datadog[0] ✅
ask 93_calling_datadog[1] ✅
ask 93_calling_datadog[2] ✅

Legend

  • ✅ the test was successful
  • :minus: the test was skipped
  • ⚠️ the test failed but is known to be flaky or known to fail
  • 🚧 the test had a setup failure (not a code regression)
  • 🔧 the test failed due to mock data issues (not a code regression)
  • 🚫 the test was throttled by API rate limits/overload
  • ❌ the test failed and should be fixed before merging the PR

@aantn
aantn merged commit 92c3ca1 into master Sep 30, 2025
7 checks passed
@aantn
aantn deleted the claude-4.5 branch September 30, 2025 07:06
@coderabbitai coderabbitai Bot mentioned this pull request Dec 27, 2025
@coderabbitai coderabbitai Bot mentioned this pull request Mar 15, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants