Repository navigation
Conversation
WalkthroughAdds timestamped historical evaluation report generation to CI and updates the report script to set headers based on filename. Overhauls MkDocs navigation to use mkdocs-awesome-nav with distributed .nav.yml files across docs sections. Updates evaluation docs (latest, history index, specific historical result) and fixes internal links. Changes
Sequence Diagram(s)sequenceDiagram
autonumber
actor Dev as Developer
participant GH as GitHub Actions (eval-benchmarks)
participant Script as generate_eval_report.py
participant Docs as docs/development/evaluations/...
Dev->>GH: Push/Trigger workflow
GH->>Script: Generate latest-results.md
Script-->>Docs: Write latest-results.md (static header)
GH->>GH: Set TIMESTAMP (YYYYMMDD_HHMMSS)
GH->>Script: Generate results_<TIMESTAMP>.md
Script-->>Docs: Write history/results_<TIMESTAMP>.md (date-based header)
sequenceDiagram
autonumber
participant WF as Workflow step
participant RG as Report Generator
participant FS as Filesystem
WF->>RG: Invoke with output_path
alt output filename matches results_YYYYMMDD_HHMMSS.md
RG->>RG: Parse date from filename
RG->>FS: Write header "Month Day, Year - HH:MM:SS"
else default case
RG->>FS: Write static header "Latest Benchmark Results"
end
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~25 minutes Possibly related PRs
Suggested reviewers
Pre-merge checks and finishing touches❌ Failed checks (1 warning)
✅ Passed checks (2 passed)
✨ Finishing touches
🧪 Generate unit tests
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Actionable comments posted: 1
📜 Review details
Configuration used: CodeRabbit UI
Review profile: CHILL
Plan: Pro
⛔ Files ignored due to path filters (1)
poetry.lockis excluded by!**/*.lock
📒 Files selected for processing (18)
.github/workflows/eval-benchmarks.yaml(1 hunks)docs/.nav.yml(1 hunks)docs/ai-providers/.nav.yml(1 hunks)docs/data-sources/.nav.yml(1 hunks)docs/data-sources/builtin-toolsets/.nav.yml(1 hunks)docs/development/.nav.yml(1 hunks)docs/development/evaluations/.nav.yml(1 hunks)docs/development/evaluations/history/.nav.yml(1 hunks)docs/development/evaluations/history/index.md(1 hunks)docs/development/evaluations/history/results_20250928_001434.md(1 hunks)docs/development/evaluations/index.md(2 hunks)docs/development/evaluations/latest-results.md(1 hunks)docs/installation/.nav.yml(1 hunks)docs/reference/.nav.yml(1 hunks)docs/walkthrough/.nav.yml(1 hunks)mkdocs.yml(1 hunks)pyproject.toml(1 hunks)scripts/generate_eval_report.py(1 hunks)
🧰 Additional context used
📓 Path-based instructions (2)
docs/**/*.md
📄 CodeRabbit inference engine (CLAUDE.md)
In MkDocs docs, always add a blank line between a header/bold text and a following list to render correctly
Files:
docs/development/evaluations/history/index.mddocs/development/evaluations/latest-results.mddocs/development/evaluations/index.mddocs/development/evaluations/history/results_20250928_001434.md
pyproject.toml
📄 CodeRabbit inference engine (CLAUDE.md)
Define and maintain the canonical list of pytest markers in pyproject.toml
Files:
pyproject.toml
🧠 Learnings (1)
📚 Learning: 2025-07-02T10:27:17.231Z
Learnt from: Sheeproid
PR: robusta-dev/holmesgpt#586
File: tests/llm/fixtures/test_ask_holmes/03_what_is_the_command_to_port_forward/test_case.yaml:4-4
Timestamp: 2025-07-02T10:27:17.231Z
Learning: In LLM-as-judge test cases for HolmesGPT, expected outputs should be descriptive rather than prescriptive when testing for flexible responses like port numbers. Using specific values in expected outputs can cause unnecessary test failures when the AI generates different but equally valid responses.
Applied to files:
docs/development/evaluations/history/results_20250928_001434.md
🪛 Ruff (0.13.1)
scripts/generate_eval_report.py
1364-1364: Unpacked variable hour is never used
Prefix it with an underscore or any other dummy variable pattern
(RUF059)
1364-1364: Unpacked variable minute is never used
Prefix it with an underscore or any other dummy variable pattern
(RUF059)
1364-1364: Unpacked variable second is never used
Prefix it with an underscore or any other dummy variable pattern
(RUF059)
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (4)
- GitHub Check: build (3.12)
- GitHub Check: build (3.11)
- GitHub Check: build (3.10)
- GitHub Check: llm_evals
🔇 Additional comments (3)
docs/development/evaluations/latest-results.md (1)
1-6: Nice refresh on the latest-results header and metadata.The new title, timestamp, and judge model line accurately reflect the most recent GPT‑4.1-based run and keep the page aligned with OpenAI’s current release cadence. (openai.com)
docs/development/evaluations/index.md (1)
69-69: Updated history links look good.Linking directly to
./history/index.mdmatches the new navigation layout and keeps the docs consistent.Also applies to: 89-89
mkdocs.yml (1)
59-59: Good call enablingawesome-nav.The plugin expects the
navblock to be omitted so it can source.nav.ymlfiles, and listing it here is exactly how the upstream docs recommend activating it. (github.com)
| match = re.match( | ||
| r"results_(\d{4})(\d{2})(\d{2})_(\d{2})(\d{2})(\d{2})\.md", output_filename | ||
| ) | ||
| if match: | ||
| # Extract date components and format as title | ||
| year, month, day, hour, minute, second = match.groups() | ||
| date_obj = datetime(int(year), int(month), int(day)) | ||
| # Format as "Month Day, Year" (no time) | ||
| title = date_obj.strftime("%B %d, %Y") | ||
| report_lines.append(f"# {title}") |
There was a problem hiding this comment.
Fix unused timestamp captures to satisfy lint.
Ruff is flagging the unpacked hour, minute, and second values because we never use them. This will fail CI unless we either prefix them with underscores or remove them entirely.
- year, month, day, hour, minute, second = match.groups()
+ year, month, day, _hour, _minute, _second = match.groups()Based on static analysis hints.
📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| match = re.match( | |
| r"results_(\d{4})(\d{2})(\d{2})_(\d{2})(\d{2})(\d{2})\.md", output_filename | |
| ) | |
| if match: | |
| # Extract date components and format as title | |
| year, month, day, hour, minute, second = match.groups() | |
| date_obj = datetime(int(year), int(month), int(day)) | |
| # Format as "Month Day, Year" (no time) | |
| title = date_obj.strftime("%B %d, %Y") | |
| report_lines.append(f"# {title}") | |
| match = re.match( | |
| r"results_(\d{4})(\d{2})(\d{2})_(\d{2})(\d{2})(\d{2})\.md", output_filename | |
| ) | |
| if match: | |
| # Extract date components and format as title | |
| year, month, day, _hour, _minute, _second = match.groups() | |
| date_obj = datetime(int(year), int(month), int(day)) | |
| # Format as "Month Day, Year" (no time) | |
| title = date_obj.strftime("%B %d, %Y") | |
| report_lines.append(f"# {title}") |
🧰 Tools
🪛 Ruff (0.13.1)
1364-1364: Unpacked variable hour is never used
Prefix it with an underscore or any other dummy variable pattern
(RUF059)
1364-1364: Unpacked variable minute is never used
Prefix it with an underscore or any other dummy variable pattern
(RUF059)
1364-1364: Unpacked variable second is never used
Prefix it with an underscore or any other dummy variable pattern
(RUF059)
🤖 Prompt for AI Agents
In scripts/generate_eval_report.py around lines 1359 to 1368, the regex match
unpacks six capture groups but only uses the first three (year, month, day),
causing lint errors for unused variables; change the unpack to ignore the unused
time captures (e.g., year, month, day, *_ = match.groups() or year, month, day,
_, _, _ = match.groups()) or extract only the first three groups via
match.groups()[:3], then proceed to construct the date_obj and title as before.
|
Closing, replaced by #1017 which includes all changes here. |
By using an mkdocs plugin to generate dynamically instead of relying on hardcoded markdown links. This also requires changing navigation across the docs due to the new plugin