Skip to content

chore(e2e): overhaul nightly benchmark summary and trim model list - #392

Merged
slin1237 merged 2 commits into
mainfrom
bench-fix-2
Feb 10, 2026
Merged

slin1237 merged 2 commits into
mainfrom
bench-fix-2

Conversation

@slin1237

@slin1237 slin1237 commented Feb 10, 2026 •

Copy link
Copy Markdown
Member

Summary

Overhauls the nightly benchmark summary script with better presentation, new sections, chart generation, and removes unused models from the benchmark matrix. Follows up on #381.

What changed

Benchmark model list trimmed

  • Removed models: Mistral-7B-Instruct-v0.3, gpt-oss (8B/70B), Llama-3.2-1B-Instruct, Qwen2.5-14B-Instruct, DeepSeek-R1-Distill-Qwen-7B
  • Remaining models: Llama-3.1-8B, Qwen2.5-7B, Qwen3-30B-A3B (single+multi), Llama-4-Maverick (single only)
  • Removed from workflow matrix (single-gpu, multi-gpu, TP overrides) and test file

Summary script rewrite (nightly_summarize.py)

  • Glossary section: explains all metrics (TTFT, TPOT, E2E, p99, mean), traffic scenarios (D/N/E patterns), and comparison column format
  • Key Findings: auto-generated executive summary with per-metric verdict, error summary, and biggest outlier wins
  • Error rate tracking: collapsible table surfacing any runs with non-zero error rates
  • TPOT metrics: added to aggregate, per-concurrency, per-model, and top wins sections
  • p99 for detail views: per-model detail tables now show p99 instead of mean (more relevant for SLAs)
  • Top wins revamped: shows all gRPC wins >30% (was capped top-15 at >10%), removed HTTP wins section
  • Chart generation: matplotlib-based 2x2 comparison grids (TTFT p99, TPOT p99, E2E p99, Output Throughput vs Concurrency) — aggregate + per-model/config charts saved as artifacts

Code quality refactoring

  • Canonical metric registry: each metric defined once, section lists composed by reference
  • Eliminated ~15 repeated getattr(cp.grpc, fld) + _advantage() patterns via _cp_advantage() / _avg_advantage() helpers
  • Unified _fmt_winner() with bold param (was two nearly-identical functions)
  • Single _plot_comparison_grid() for chart generation (was duplicated for aggregate vs per-model)
  • _group_by_concurrency() helper (was computed 3 times)
  • _fmt_metric_value() for type-based formatting (eliminates if/elif dispatch)
  • SummaryResult dataclass instead of bare tuple return
  • ComparisonPoint.config property

Gateway debug logging

  • Added log_level and log_dir parameters to gateway start
  • Plumbed through pytest markers (@pytest.mark.gateway(log_level="debug", log_dir="nightly_gateway_logs"))
  • Workflow uploads chart artifacts alongside benchmark results

Test plan

  • Verify python3 e2e_test/benchmarks/nightly_summarize.py parses without errors
  • Verify pre-commit hooks pass (ruff lint + format)
  • Next nightly run validates end-to-end with trimmed model list
  • Chart artifacts appear in workflow run artifacts

Summary by CodeRabbit

Release Notes

  • New Features

    • Nightly benchmarks now generate and upload comparison charts as artifacts for easier result visualization.
    • Added configurable gateway logging with debug level and log directory options for enhanced diagnostics.
  • Improvements

    • Streamlined nightly benchmark test matrix by reducing model set for faster feedback cycles.
    • Refactored benchmark summary reporting with improved metric organization and formatted output.

Rewrite nightly_summarize.py for better readability and richer output:

- Add glossary explaining all metrics (TTFT, TPOT, E2E, p99, mean),
  traffic scenarios (D/N/E patterns), and comparison columns
- Add auto-generated Key Findings executive summary at the top
- Add TPOT (Time Per Output Token) to all summary sections — was
  parsed from JSON but mostly not displayed
- Switch per-model summary and detail tables from mean to p99 metrics
  which better capture tail latency for production SLAs
- Replace confusing +/- percentage format with clear "gRPC X%" /
  "HTTP X%" / "~" labels using normalized gRPC-advantage direction
- Add matplotlib comparison chart generation (aggregate + per-model
  2x2 grids of TTFT p99, TPOT p99, E2E p99, Output Throughput vs
  Concurrency) uploaded as separate artifact
- Add error rate tracking section for runs with non-zero errors
- Show all gRPC wins >30% instead of capped top-15 at >10%
- Add TPOT p99 to aggregate metrics table

Remove 5 models that don't add value to the benchmark:
- Llama-3.2-1B-Instruct, Qwen2.5-14B-Instruct,
  DeepSeek-R1-Distill-Qwen-7B (this PR)
- Mistral-7B, gpt-oss (removed from matrix + test file)

Enable gateway debug logging during nightly benchmarks:
- Make log_level and log_dir configurable on Gateway class
  (was hardcoded to --log-level warn with no --log-dir)
- Plumb log_level/log_dir through pytest markers → gateway_config →
  gateway.start() in setup_backend.py
- Set log_level="debug" and log_dir="nightly_gateway_logs" in
  the nightly test marker so logs are captured as artifacts

Workflow changes:
- Install matplotlib in summarize job for chart generation
- Pass --charts-dir nightly_charts to generate comparison plots
- Upload charts as separate artifact
- Remove trimmed models from single/multi GPU matrices and
  TP override JSON

Remaining benchmark models: Llama-3.1-8B, Qwen2.5-7B,
Qwen3-30B-A3B (single+multi), Llama-4-Maverick (single only)
…cation

- Define each metric once in a canonical registry (_M_TTFT_P99, etc.)
  and compose section-specific lists by reference instead of copy-pasting
  tuples across 6 separate lists
- Merge PER_MODEL_SUMMARY_METRICS and PER_MODEL_DETAIL_METRICS into
  single PER_MODEL_METRICS used for both summary and detail tables
- Replace _raw_pct() + _advantage() with single _advantage() that
  inlines the percentage calculation
- Add _cp_advantage() and _avg_advantage() helpers to eliminate the
  repeated getattr(cp.grpc, fld) + _advantage() pattern (~15 call sites)
- Unify _fmt_winner() and _fmt_winner_bold() into single function with
  bold parameter
- Add _fmt_metric_value() for type-based value formatting, reuse in
  _section_by_concurrency instead of duplicated if/elif/else dispatch
- Extract _group_by_concurrency() helper (was computed 3 times)
- Extract _plot_comparison_grid() to deduplicate aggregate vs per-model
  chart generation (was ~40 lines of identical plotting code)
- Add ComparisonPoint.config property to replace f-string duplication
- Add SummaryResult dataclass for generate_summary() return type
  instead of bare tuple
- Extract glossary text into _GLOSSARY_LINES constant
- Remove unused experiments parameter from generate_charts()
- Shorten glossary descriptions to satisfy 120-char line limit
@github-actions github-actions Bot added ci CI/CD configuration changes tests Test changes labels Feb 10, 2026
@coderabbitai

coderabbitai Bot commented Feb 10, 2026 •

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

The pull request reduces the nightly benchmark test matrix by removing several model entries across multiple frameworks, introduces chart generation and upload capabilities to the benchmark summarization workflow, refactors summary generation with a centralized metrics registry and new data structures, and adds logging configuration controls to the gateway backend setup.

Changes

Cohort / File(s) Summary
Nightly Benchmark Workflow
.github/workflows/nightly-benchmark.yml
Removed model entries from benchmark matrices (Llama, Qwen, DeepSeek, Mistral, OpenAI variants), simplified E2E model overrides, and added matplotlib installation and chart generation/upload steps to the summarize phase.
Summary Generation Refactoring
e2e_test/benchmarks/nightly_summarize.py
Introduced centralized metric registry and SummaryResult dataclass; replaced string return with structured output; added formatting helpers for advantage calculation, latency/throughput values, and metric rendering; implemented optional chart plotting with _plot_comparison_grid and generate_charts; restructured summary sections for key findings, error rates, and per-model details; added --charts-dir CLI option.
Test Configuration & Logging
e2e_test/benchmarks/test_nightly_perf.py
Removed model entries from _NIGHTLY_MODELS and updated gateway pytest marker to include debug-level logging configuration (log_level="debug", log_dir="nightly_gateway_logs").
Backend Gateway Logging Setup
e2e_test/fixtures/setup_backend.py, e2e_test/infra/gateway.py
Added log_level and log_dir fields to gateway configuration; threaded these parameters through to Gateway.start() calls across HTTP, gRPC, and local backends; updated command building to use dynamic log level and conditionally append log directory.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~35 minutes

Possibly related PRs

Suggested labels

ci, tests

Suggested reviewers

  • key4ng
  • CatherineSue

Poem

🐰 Charts and logs now light the way,
Benchmarks streamlined every day,
Summaries rich with metrics new,
Gateway whispers secrets true,
Testing cleaner, brighter still! ✨

🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title 'chore(e2e): overhaul nightly benchmark summary and trim model list' directly and specifically summarizes the main changes: rewriting the benchmark summary script and reducing the model matrix.
Docstring Coverage ✅ Passed Docstring coverage is 88.00% which is sufficient. The required threshold is 80.00%.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing touches
  • 📝 Generate docstrings
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch bench-fix-2

Tip

Issue Planner is now in beta. Read the docs and try it out! Share your feedback on Discord.


Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello @slin1237, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request introduces a comprehensive overhaul of the nightly benchmark reporting system, significantly enhancing the clarity, detail, and visual presentation of performance metrics. By refining the benchmark model list, implementing advanced summary generation features, and improving underlying code quality, the changes aim to provide more actionable insights and a more efficient benchmarking workflow.

Highlights

  • Benchmark Model List Trimmed: Several models, including Mistral-7B-Instruct-v0.3, gpt-oss, Llama-3.2-1B-Instruct, Qwen2.5-14B-Instruct, and DeepSeek-R1-Distill-Qwen-7B, were removed from the nightly benchmark matrix and test files to streamline the benchmark process. The remaining models are Llama-3.1-8B, Qwen2.5-7B, Qwen3-30B-A3B, and Llama-4-Maverick.
  • Nightly Summary Script Overhaul (nightly_summarize.py): The benchmark summary script was significantly rewritten to include a comprehensive glossary for metrics and traffic scenarios, an auto-generated 'Key Findings' executive summary with per-metric verdicts and error summaries, and detailed error rate tracking. It now incorporates TPOT metrics, uses p99 for per-model detail tables, revamps 'Top Wins' to show all gRPC wins >30%, and generates matplotlib-based 2x2 comparison charts for aggregate and per-model/config data.
  • Code Quality Refactoring: The nightly_summarize.py script underwent extensive refactoring, introducing a canonical metric registry for single-source-of-truth definitions, eliminating repeated code patterns with new helper functions (_cp_advantage, _avg_advantage, _fmt_winner, _group_by_concurrency, _fmt_metric_value), unifying chart generation logic, and adopting a SummaryResult dataclass for cleaner return values.
  • Gateway Debug Logging: New log_level and log_dir parameters were added to the gateway startup process, enabling more granular control over logging. These parameters are now plumbed through pytest markers, and the workflow is configured to upload generated chart artifacts alongside benchmark results.
Changelog
  • e2e_test/benchmarks/nightly_summarize.py
    • Introduced a canonical metric registry and section-specific metric lists for consistent metric definitions.
    • Added a comprehensive glossary section explaining metrics, traffic scenarios, and comparison column formats.
    • Defined a SummaryResult dataclass to encapsulate the output of the generate_summary function.
    • Refactored formatting helpers, including _advantage, _cp_advantage, _avg_advantage, _fmt_winner, _fmt_metric_value, and _group_by_concurrency, to improve code reusability and clarity.
    • Implemented chart generation logic (_plot_comparison_grid, generate_charts) using matplotlib for visual comparisons.
    • Added new sections: _section_key_findings for an executive summary and _section_error_rates for surfacing non-zero error rates.
    • Updated _section_overview to filter models and note excluded ones based on comparison data.
    • Revised _section_aggregate, _section_by_concurrency, _section_scorecard, _section_top_wins, and _section_per_model to leverage new helpers, metric definitions, and improved output formatting.
    • Modified generate_summary to return SummaryResult, include the glossary, key findings, and error rates, and handle chart generation.
    • Updated the main function to parse a --charts-dir argument and trigger chart generation.
  • e2e_test/benchmarks/test_nightly_perf.py
    • Removed several models (Llama-3.2-1B-Instruct, Qwen2.5-14B-Instruct, DeepSeek-R1-Distill-Qwen-7B, Mistral-7B-Instruct-v0.3, gpt-oss-20b) from the _NIGHTLY_MODELS list.
    • Updated the @pytest.mark.gateway decorator to include log_level="debug" and log_dir="nightly_gateway_logs" for enhanced test logging.
  • e2e_test/fixtures/setup_backend.py
    • Added log_level and log_dir with None defaults to the gateway_config dictionary.
    • Passed log_level and log_dir from gateway_config to the gateway.start method in _setup_pd_backend_common, _setup_grpc_backend, and _setup_local_backend functions.
  • e2e_test/infra/gateway.py
    • Added log_level (default 'warn') and log_dir (default None) attributes to the Gateway class constructor.
    • Modified the start method to accept log_level and log_dir parameters and update the corresponding instance attributes if provided.
    • Updated the _build_base_cmd method to include --log-level and optionally --log-dir arguments in the gateway launch command.
Ignored Files
  • Ignored by pattern: .github/workflows/** (1)
    • .github/workflows/nightly-benchmark.yml
Activity
  • No specific human activity (comments, reviews, progress updates) was provided in the context for this pull request.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@slin1237
slin1237 merged commit f80dca4 into main Feb 10, 2026
17 of 18 checks passed
@slin1237
slin1237 deleted the bench-fix-2 branch February 10, 2026 16:30

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Fix all issues with AI agents
In `@e2e_test/benchmarks/nightly_summarize.py`:
- Around line 562-627: The summary in _section_key_findings uses max(grpc_wins,
http_wins) and min(...) when reporting "{winner} wins X/Y comparisons", which
can misattribute counts if the average direction (avg_adv) differs from the
majority; change it to pick counts based on the winner determined by avg_adv
(use winner_count = grpc_wins if avg_adv > 0 else http_wins and loser_count =
the other) and report winner_count/len(advs) and loser_count instead of max/min
to ensure wins/losses align with the declared winner.
🧹 Nitpick comments (2)
e2e_test/infra/gateway.py (1)

198-201: Allow explicit falsy overrides for log settings.
Truthy checks skip empty-string overrides and can retain prior values if a Gateway instance is reused; using is not None makes the override behavior explicit.

Suggested tweak
-        if log_level:
-            self.log_level = log_level
-        if log_dir:
-            self.log_dir = log_dir
+        if log_level is not None:
+            self.log_level = log_level
+        if log_dir is not None:
+            self.log_dir = log_dir
e2e_test/benchmarks/test_nightly_perf.py (1)

126-126: Consider per-model log directories to avoid collisions.
Using a single nightly_gateway_logs folder can overwrite logs across parametrized runs in the same workspace; scoping by model (and optionally backend) keeps artifacts distinct.

Possible refactor
-    _worker_count = worker_count
+    _worker_count = worker_count
+    safe_model_id = model_id.replace("/", "__")
@@
-    `@pytest.mark.gateway`(policy="round_robin", log_level="debug", log_dir="nightly_gateway_logs")
+    `@pytest.mark.gateway`(
+        policy="round_robin",
+        log_level="debug",
+        log_dir=f"nightly_gateway_logs/{safe_model_id}",
+    )

Comment on lines +562 to +627
def _section_key_findings(
comparisons: list[ComparisonPoint],
experiments: list[ExperimentInfo],
) -> list[str]:
"""Auto-generated executive summary of the benchmark results."""
if not comparisons:
return []

lines = ["### Key Findings", ""]

# 1. Overall verdict per key metric
for label, fld, lower_better, _unit in CHART_METRICS:
advs = [a for cp in comparisons if (a := _cp_advantage(cp, fld, lower_better)) is not None]
if not advs:
continue
avg_adv = sum(advs) / len(advs)
grpc_wins = sum(1 for a in advs if a > 2)
http_wins = sum(1 for a in advs if a < -2)
ties = len(advs) - grpc_wins - http_wins
if abs(avg_adv) < 1:
lines.append(f"- **{label}**: No clear winner — essentially tied across all scenarios")
else:
winner = "gRPC" if avg_adv > 0 else "HTTP"
lines.append(
f"- **{label}**: {winner} wins {max(grpc_wins, http_wins)}/{len(advs)} "
f"comparisons (avg {abs(avg_adv):.1f}% better), "
f"{min(grpc_wins, http_wins)} losses, {ties} ties"
)

# 2. Error rates
total_runs = sum(len(e.runs) for e in experiments)
error_runs = [(e, r) for e in experiments for r in e.runs if r.error_rate > 0]
if error_runs:
lines.append(f"- **Errors**: {len(error_runs)}/{total_runs} runs had non-zero error rates")
else:
lines.append(f"- **Errors**: All {total_runs} runs completed with 0% error rate")

# 3. Biggest outliers
biggest_grpc_win = biggest_http_win = None
biggest_grpc_adv = biggest_http_adv = 0.0
for cp in comparisons:
for label, fld, lower_better, _ in CHART_METRICS:
adv = _cp_advantage(cp, fld, lower_better)
if adv is not None and adv > biggest_grpc_adv:
biggest_grpc_adv = adv
biggest_grpc_win = (cp, label)
if adv is not None and adv < biggest_http_adv:
biggest_http_adv = adv
biggest_http_win = (cp, label)

if biggest_grpc_win and biggest_grpc_adv > 10:
cp, metric = biggest_grpc_win
lines.append(
f"- **Largest gRPC win**: {biggest_grpc_adv:.0f}% on {metric} "
f"— {cp.model} `{cp.scenario}` C={cp.concurrency}"
)
if biggest_http_win and abs(biggest_http_adv) > 10:
cp, metric = biggest_http_win
lines.append(
f"- **Largest HTTP win**: {abs(biggest_http_adv):.0f}% on {metric} "
f"— {cp.model} `{cp.scenario}` C={cp.concurrency}"
)

lines.append("")
return lines

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

Fix win/loss counts in key findings summary.
The message uses max/min of win counts, which can misattribute wins when the average winner differs from the majority; select counts based on the winner.

Proposed fix
-        else:
-            winner = "gRPC" if avg_adv > 0 else "HTTP"
-            lines.append(
-                f"- **{label}**: {winner} wins {max(grpc_wins, http_wins)}/{len(advs)} "
-                f"comparisons (avg {abs(avg_adv):.1f}% better), "
-                f"{min(grpc_wins, http_wins)} losses, {ties} ties"
-            )
+        else:
+            winner = "gRPC" if avg_adv > 0 else "HTTP"
+            win_count = grpc_wins if avg_adv > 0 else http_wins
+            loss_count = http_wins if avg_adv > 0 else grpc_wins
+            lines.append(
+                f"- **{label}**: {winner} wins {win_count}/{len(advs)} "
+                f"comparisons (avg {abs(avg_adv):.1f}% better), "
+                f"{loss_count} losses, {ties} ties"
+            )
🤖 Prompt for AI Agents
In `@e2e_test/benchmarks/nightly_summarize.py` around lines 562 - 627, The summary
in _section_key_findings uses max(grpc_wins, http_wins) and min(...) when
reporting "{winner} wins X/Y comparisons", which can misattribute counts if the
average direction (avg_adv) differs from the majority; change it to pick counts
based on the winner determined by avg_adv (use winner_count = grpc_wins if
avg_adv > 0 else http_wins and loser_count = the other) and report
winner_count/len(advs) and loser_count instead of max/min to ensure wins/losses
align with the declared winner.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request significantly overhauls the nightly benchmark summary script, enhancing report readability, informativeness, and maintainability through refactoring, new features like chart generation and a "Key Findings" summary, and enabling gateway debug logging. A security audit identified two high-severity vulnerabilities in the nightly_summarize.py script: a Stored Cross-Site Scripting (XSS) vulnerability due to unsanitized data and an Arbitrary File Write vulnerability from unvalidated GITHUB_STEP_SUMMARY usage. Remediation by sanitizing all external data and validating file paths is strongly recommended. Additionally, there is a suggestion to improve the command-line argument parsing in the summary script for robustness.

Comment on lines +614 to +617
lines.append(
f"- **Largest gRPC win**: {biggest_grpc_adv:.0f}% on {metric} "
f"— {cp.model} `{cp.scenario}` C={cp.concurrency}"
)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

security-high high

The script generates a markdown summary by embedding data (e.g., cp.model, cp.scenario) directly from parsed JSON benchmark files without sanitization. An attacker who can control the content of these JSON files can inject malicious HTML (<script> tags). This will be executed in the browser of anyone viewing the generated report, leading to a Stored XSS vulnerability. It is recommended to sanitize all data read from external files before embedding it into the report by using html.escape().

Comment on lines 1017 to 1019
with open(summary_file, "a") as f:
f.write(summary)
f.write(result.markdown)
f.write("\n")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

security-high high

The script reads the GITHUB_STEP_SUMMARY environment variable and uses its value as a file path to write the summary report. An attacker who can control this environment variable can cause the script to write to arbitrary files on the system, potentially leading to code execution or system compromise. While using GITHUB_STEP_SUMMARY is standard in GitHub Actions, it's crucial to validate the path to ensure it resides within an expected directory, especially if the execution environment can be influenced by external actors.

Comment on lines +989 to +1002
args = sys.argv[1:]
base_dir = Path.cwd()
charts_dir: Path | None = None

i = 0
while i < len(args):
if args[i] == "--charts-dir" and i + 1 < len(args):
charts_dir = Path(args[i + 1])
i += 2
elif not args[i].startswith("-"):
base_dir = Path(args[i])
i += 1
else:
i += 1

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The manual command-line argument parsing using sys.argv is a bit fragile. It might not handle unexpected arguments gracefully and lacks features like auto-generated help text. Using Python's built-in argparse module would make the script more robust and user-friendly.

You'll need to add import argparse at the top of the file.

    parser = argparse.ArgumentParser(
        description="Generate nightly benchmark summary.",
        formatter_class=argparse.ArgumentDefaultsHelpFormatter,
    )
    parser.add_argument(
        "base_dir",
        nargs="?",
        default=Path.cwd(),
        type=Path,
        help="Base directory containing benchmark results.",
    )
    parser.add_argument(
        "--charts-dir",
        type=Path,
        default=None,
        help="Directory to save generated charts.",
    )
    args = parser.parse_args()
    base_dir = args.base_dir
    charts_dir = args.charts_dir

ppraneth pushed a commit that referenced this pull request Feb 18, 2026
)

Signed-off-by: ppraneth <pranethparuchuri@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci CI/CD configuration changes tests Test changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant