Repository navigation
Add benchmarks for Claude 4.5 - #1017
Conversation
WalkthroughIntroduces timestamped benchmark report generation in CI, adjusts report header logic based on filename, restructures docs navigation into per-directory .nav.yml files using mkdocs-awesome-nav, updates MkDocs config and dev dependencies, and refreshes evaluation docs including history index and a new timestamped results file. Changes
Sequence Diagram(s)sequenceDiagram
autonumber
participant Dev as GitHub Actions Workflow
participant Script as generate_eval_report.py
participant FS as Docs Filesystem
Dev->>Script: Run report generator (latest)
Script->>FS: Write latest-results.md
Note over Script: Default header used (non-matching filename)
Dev->>Dev: Compute TIMESTAMP
Dev->>Script: Run report generator (results_<TIMESTAMP>.md)
Script->>Script: Filename matches results_YYYYMMDD_HHMMSS.md
Script->>Script: Set header to formatted date
Script->>FS: Write history/results_<TIMESTAMP>.md
Note over Dev,FS: MkDocs uses awesome-nav with per-directory .nav.yml to surface new history entry
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~25 minutes Suggested reviewers
Pre-merge checks and finishing touches❌ Failed checks (2 warnings)
✅ Passed checks (1 passed)
✨ Finishing touches
🧪 Generate unit tests
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. 🧪 Early access (Sonnet 4.5): enabledWe are currently testing the Sonnet 4.5 model, which is expected to improve code review quality. However, this model may lead to increased noise levels in the review comments. Please disable the early access features if the noise level causes any inconvenience. Note:
Comment |
There was a problem hiding this comment.
Actionable comments posted: 1
📜 Review details
Configuration used: CodeRabbit UI
Review profile: CHILL
Plan: Pro
⛔ Files ignored due to path filters (1)
poetry.lockis excluded by!**/*.lock
📒 Files selected for processing (17)
.github/workflows/eval-benchmarks.yaml(1 hunks)docs/.nav.yml(1 hunks)docs/ai-providers/.nav.yml(1 hunks)docs/data-sources/.nav.yml(1 hunks)docs/data-sources/builtin-toolsets/.nav.yml(1 hunks)docs/development/.nav.yml(1 hunks)docs/development/evaluations/.nav.yml(1 hunks)docs/development/evaluations/history/.nav.yml(1 hunks)docs/development/evaluations/history/index.md(1 hunks)docs/development/evaluations/history/results_20250928_001434.md(1 hunks)docs/development/evaluations/index.md(2 hunks)docs/installation/.nav.yml(1 hunks)docs/reference/.nav.yml(1 hunks)docs/walkthrough/.nav.yml(1 hunks)mkdocs.yml(1 hunks)pyproject.toml(1 hunks)scripts/generate_eval_report.py(1 hunks)
🧰 Additional context used
📓 Path-based instructions (2)
docs/**/*.md
📄 CodeRabbit inference engine (CLAUDE.md)
In MkDocs docs, always add a blank line between a header/bold text and a following list to render correctly
Files:
docs/development/evaluations/history/index.mddocs/development/evaluations/index.mddocs/development/evaluations/history/results_20250928_001434.md
pyproject.toml
📄 CodeRabbit inference engine (CLAUDE.md)
Define and maintain the canonical list of pytest markers in pyproject.toml
Files:
pyproject.toml
🪛 Ruff (0.13.1)
scripts/generate_eval_report.py
1364-1364: Unpacked variable hour is never used
Prefix it with an underscore or any other dummy variable pattern
(RUF059)
1364-1364: Unpacked variable minute is never used
Prefix it with an underscore or any other dummy variable pattern
(RUF059)
1364-1364: Unpacked variable second is never used
Prefix it with an underscore or any other dummy variable pattern
(RUF059)
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (2)
- GitHub Check: Pre-commit checks
- GitHub Check: llm_evals
🔇 Additional comments (18)
pyproject.toml (2)
79-79: LGTM! Dev dependency addition aligns with PR objectives.The addition of
mkdocs-awesome-navplugin properly supports the new navigation approach described in the PR objectives. The dependency is correctly placed in the dev group with an appropriate version constraint.
95-123: Pytest markers are well-maintained.The canonical list of pytest markers in pyproject.toml is comprehensive and well-documented, following the coding guidelines.
docs/reference/.nav.yml (1)
1-6: All referenced markdown files are present. No further changes needed.docs/development/evaluations/history/results_20250928_001434.md (1)
1-7: Metadata update looks correct.Heading timestamp, generated time, and judge model attribution all line up with the new benchmarking flow. Nothing else to block here.
docs/ai-providers/.nav.yml (1)
1-13: LGTM! Clean navigation structure.The AI providers navigation configuration is well-organized and follows a logical structure. All entries are properly formatted for the mkdocs-awesome-nav plugin.
docs/walkthrough/.nav.yml (1)
1-7: LGTM! Walkthrough navigation is well-structured.The navigation configuration properly organizes the walkthrough documentation with clear, descriptive titles.
docs/data-sources/.nav.yml (1)
1-7: LGTM! Data sources navigation properly configured.The navigation structure correctly includes both direct markdown files and a directory reference for nested navigation (builtin-toolsets), which aligns with the mkdocs-awesome-nav plugin's directory-based navigation approach.
.github/workflows/eval-benchmarks.yaml (1)
118-129: Verify artifact upload includes timestamped report.The timestamped benchmark report is generated but not included in the artifact upload (line 137 only uploads
latest-results.md). Since the git commit step (lines 140-153) is currently disabled for testing, these timestamped files may be lost.Consider updating the artifact upload to preserve timestamped reports during the testing phase:
- name: Upload eval results if: always() uses: actions/upload-artifact@v4 with: name: eval-results-${{ github.run_id }} path: | docs/development/evaluations/latest-results.md + docs/development/evaluations/history/results_*.mdThis ensures timestamped reports are accessible even while the automatic commit is disabled.
docs/data-sources/builtin-toolsets/.nav.yml (1)
1-30: LGTM! Navigation structure is well-organized.The navigation configuration properly defines the Built-in Toolsets section with clear, descriptive labels for each tool. The YAML structure is valid and entries are logically organized.
docs/development/.nav.yml (1)
1-4: LGTM! Clean navigation structure.The navigation configuration correctly defines the Development section with an index and reference to the Evaluations subdirectory.
docs/.nav.yml (1)
1-10: LGTM! Top-level navigation structure is well-organized.The root navigation configuration properly defines the main sections of the documentation site with a logical flow from installation through reference and community resources.
docs/development/evaluations/index.md (2)
69-69: LGTM! Link update is correct.The link now properly points to the history index file, ensuring consistent navigation to the historical results section.
89-89: LGTM! Link update is correct.The historical results link now properly references the history index file, maintaining consistency with the navigation restructuring.
docs/development/evaluations/history/.nav.yml (1)
1-4: LGTM! Clever use of wildcard for dynamic content.The navigation configuration uses the wildcard pattern to automatically include all timestamped benchmark result files in the history directory, which aligns well with the PR's goal of automatically listing new benchmark results.
docs/installation/.nav.yml (1)
1-5: LGTM!The navigation structure is well-organized and follows the YAML format correctly. The entries provide clear labels for the installation documentation.
docs/development/evaluations/history/index.md (1)
1-5: LGTM!The simplified documentation correctly includes blank lines between headers and content, ensuring proper MkDocs rendering.
docs/development/evaluations/.nav.yml (1)
1-7: LGTM!The navigation structure is comprehensive and well-organized, covering all evaluation documentation sections. The reference to
historydirectory on Line 4 will correctly expand to include the timestamped results files.mkdocs.yml (1)
59-59: mkdocs-awesome-nav dependency confirmed
Found in pyproject.toml on line 79.
Also includes improvements to mkdocs so new benchmark results are automatically listed. (This requires changing the way navigation is done in mkdocs and switching to a plugin.)