Skip to content

refactor(e2e): clean up test infrastructure code quality and efficiency - #686

Merged
slin1237 merged 1 commit into
mainfrom
refactor/e2e-test-infra-cleanup
Mar 9, 2026
Merged

slin1237 merged 1 commit into
mainfrom
refactor/e2e-test-infra-cleanup

Conversation

@slin1237

@slin1237 slin1237 commented Mar 9, 2026 •

Copy link
Copy Markdown
Member

Summary

Six targeted code quality and efficiency improvements across the E2E test infrastructure, reducing net 15 lines while improving consistency and eliminating redundant work.

What changed

  • e2e_test/fixtures/hooks.py: Removed duplicate _get_own_class_marker() (26 lines) that reimplemented the same MRO-walking logic already in markers.resolve_class_marker(). The _get_marker() helper now delegates directly.
  • e2e_test/infra/constants.py: Removed redundant import os inside get_runtime() — os is already imported at module level.
  • e2e_test/infra/gateway.py: Moved import time from inside add_worker() to module level; replaced time.time() with time.perf_counter() for consistency with the rest of the codebase.
  • e2e_test/infra/gpu_monitor.py: Changed _percentile() to accept a pre-sorted list and sort once in _compute_stats() instead of re-sorting 7 times per call. Also use sorted_list[0]/sorted_list[-1] for min/max instead of separate min()/max() traversals.
  • e2e_test/infra/model_specs.py: Cached E2E_MODEL_TP_OVERRIDES JSON parsing at module load via _parse_tp_overrides() instead of re-parsing the env var on every get_model_spec() call.
  • e2e_test/infra/worker.py: Create one gRPC channel before the health check retry loop and reuse it across iterations (with finally cleanup) instead of creating+closing a new channel per attempt.

Why

These are low-risk housekeeping fixes identified via systematic code review:

  • Code reuse: hooks.py duplicated logic already available in markers.py
  • Consistency: mixed time.time() vs time.perf_counter(), inline vs module-level imports
  • Efficiency: redundant sorting, repeated JSON parsing, unnecessary gRPC channel churn

Test plan

  • Verify pytest --collect-only still collects tests correctly (hooks.py marker resolution)
  • Verify GPU monitor stats produce identical results (gpu_monitor.py sort change is behavioral no-op)
  • Verify gRPC health checks still work for workers (worker.py channel reuse)
  • CI passes on existing E2E test suite

Summary by CodeRabbit

  • Refactor
    • Optimized internal test infrastructure including timer resolution, data processing efficiency, and gRPC health-check lifecycle management.
    • Consolidated marker resolution logic and streamlined environment variable parsing with caching for improved startup performance.

- fixtures/hooks.py: Remove duplicate _get_own_class_marker() implementation
  (26 lines) and delegate to markers.resolve_class_marker() which already
  implements the same MRO-walking logic
- infra/constants.py: Remove redundant `import os` inside get_runtime();
  os is already imported at module level
- infra/gateway.py: Move `import time` to module level and replace
  time.time() with time.perf_counter() in add_worker() for consistency
  with the rest of the codebase
- infra/gpu_monitor.py: Sort samples once in _compute_stats() instead of
  re-sorting 7 times per call through _percentile(); use sorted list for
  min/max instead of separate min()/max() calls
- infra/model_specs.py: Cache E2E_MODEL_TP_OVERRIDES JSON parsing at module
  load via _parse_tp_overrides() instead of re-parsing on every
  get_model_spec() call
- infra/worker.py: Create one gRPC channel before the health check retry
  loop and reuse it across iterations instead of creating and closing a new
  channel per attempt; clean up channel in a finally block

Signed-off-by: SimonOlin <simon@simolin.dev>
Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
@github-actions github-actions Bot added the tests Test changes label Mar 9, 2026
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request focuses on enhancing the code quality and efficiency of the E2E test infrastructure. It introduces six targeted improvements that streamline operations, reduce redundant code, and improve performance, resulting in a net reduction of 15 lines of code. The changes address issues related to code reuse, consistency in utility usage, and overall operational efficiency within the testing framework.

Highlights

  • Code Reuse: Removed duplicate logic for resolving class markers in e2e_test/fixtures/hooks.py by delegating to an existing utility function, markers.resolve_class_marker().
  • Consistency: Standardized time measurement in e2e_test/infra/gateway.py by replacing time.time() with time.perf_counter() and moved import time to the module level. Also removed a redundant import os in e2e_test/infra/constants.py.
  • Efficiency: Optimized e2e_test/infra/gpu_monitor.py by sorting samples once in _compute_stats() and passing pre-sorted lists to _percentile(), avoiding multiple sorts. Improved e2e_test/infra/model_specs.py by caching JSON parsing of E2E_MODEL_TP_OVERRIDES at module load. Enhanced e2e_test/infra/worker.py by reusing a single gRPC channel across health check retries instead of creating a new one for each attempt.
Changelog
  • e2e_test/fixtures/hooks.py
    • Removed duplicate _get_own_class_marker() function, which reimplemented MRO-walking logic already present in markers.resolve_class_marker().
    • Updated _get_marker() to delegate directly to markers.resolve_class_marker() for resolving class markers.
  • e2e_test/infra/constants.py
    • Removed a redundant import os statement from inside the get_runtime() function, as os is already imported at the module level.
  • e2e_test/infra/gateway.py
    • Moved the import time statement from inside add_worker() to the module level.
    • Replaced time.time() with time.perf_counter() for consistency with other performance measurements in the codebase.
  • e2e_test/infra/gpu_monitor.py
    • Modified _percentile() to accept a pre-sorted list of samples.
    • Updated _compute_stats() to sort the sample list once and then use array indexing (sorted_list[0], sorted_list[-1]) for min/max values, and pass the sorted list to _percentile().
  • e2e_test/infra/model_specs.py
    • Implemented caching for E2E_MODEL_TP_OVERRIDES JSON parsing at module load via a new _parse_tp_overrides() function.
    • Modified get_model_spec() to use the cached _TP_OVERRIDES instead of re-parsing the environment variable on every call.
  • e2e_test/infra/worker.py
    • Refactored _wait_grpc_healthy() to create a single gRPC channel before the health check retry loop and reuse it across iterations.
    • Ensured the gRPC channel is properly closed using a finally block.
Activity
  • No human activity has been recorded on this pull request yet.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@coderabbitai

coderabbitai Bot commented Mar 9, 2026 •

Copy link
Copy Markdown

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 9e819ab0-a154-4d6a-94c2-a3b528f91f83

📥 Commits

Reviewing files that changed from the base of the PR and between c9bbade and 8879591.

📒 Files selected for processing (6)
  • e2e_test/fixtures/hooks.py
  • e2e_test/infra/constants.py
  • e2e_test/infra/gateway.py
  • e2e_test/infra/gpu_monitor.py
  • e2e_test/infra/model_specs.py
  • e2e_test/infra/worker.py
💤 Files with no reviewable changes (1)
  • e2e_test/infra/constants.py

📝 Walkthrough

Walkthrough

This PR refactors e2e test infrastructure across six files, improving marker resolution logic, removing redundant code, optimizing timing measurements with monotonic clocks, refactoring percentile computation for pre-sorted data, centralizing environment variable parsing with caching, and optimizing gRPC channel lifecycle management.

Changes

Cohort / File(s) Summary
Fixture and Marker Resolution
e2e_test/fixtures/hooks.py
Consolidates marker resolution into a single resolve_class_marker delegation, replacing direct __dict__ access and get_closest_marker fallback logic with a unified child-first MRO-aware strategy.
Code Cleanup
e2e_test/infra/constants.py
Removes redundant local os module import inside get_runtime() function; behavior unchanged as global import remains available.
Timing and Channel Optimization
e2e_test/infra/gateway.py, e2e_test/infra/worker.py
Replaces time.time() with monotonic time.perf_counter() in timing loops; refactors gRPC health-check to create channel once before retry loop and close in finally block instead of recreating on each iteration.
Data Processing
e2e_test/infra/gpu_monitor.py
Refactors _percentile to accept pre-sorted input, moving sort operation to _compute_stats; adjusts min/max extraction to use sorted list indexing instead of built-in functions.
Configuration Caching
e2e_test/infra/model_specs.py
Introduces _parse_tp_overrides() helper and module-level _TP_OVERRIDES cache to centralize E2E_MODEL_TP_OVERRIDES parsing at import time, replacing per-call parsing in get_model_spec().

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related PRs

Suggested labels

tests

Suggested reviewers

  • CatherineSue
  • key4ng
  • XinyueZhang369

Poem

🐰 Markers now resolve with grace so true,
Channels live once, not born anew,
Sorted percentiles, caches in place,
E2E tests hop at a faster pace! ✨

🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately summarizes the main objective: refactoring E2E test infrastructure to improve code quality and efficiency across six files.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
  • 📝 Generate docstrings (stacked PR)
  • 📝 Generate docstrings (commit on current branch)
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch refactor/e2e-test-infra-cleanup

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a series of well-targeted refactorings across the E2E test infrastructure, enhancing code quality, consistency, and efficiency. The changes include removing duplicated marker resolution logic, eliminating redundant imports and JSON parsing, optimizing statistical calculations by avoiding repeated sorting, and improving gRPC connection handling by reusing channels. All changes are implemented correctly and contribute to a cleaner and more performant test suite. I have reviewed the changes and found no issues.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 887959124e

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

def _get_marker(item: pytest.Item, name: str):
"""Get the most specific marker, preferring child class over parent."""
return _get_own_class_marker(item, name) or item.get_closest_marker(name)
return resolve_class_marker(item, name)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Honor method markers over inherited class markers

Switching _get_marker() to resolve_class_marker() changes precedence when a test method has an engine/vendor/gpu marker but its class only inherits a parent class marker: resolve_class_marker() returns the inherited class marker before consulting item.get_closest_marker(), so method-level overrides are ignored. Under E2E_ENGINE/E2E_VENDOR/E2E_GPU_TIER filtering this can incorrectly include/exclude tests in subclass hierarchies where only the parent class carries the broad marker.

Useful? React with 👍 / 👎.

return None


_TP_OVERRIDES = _parse_tp_overrides()

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Keep TP override env lookup dynamic

Caching E2E_MODEL_TP_OVERRIDES at import time freezes the override map for the lifetime of the process, so any later environment updates are silently ignored by get_model_spec(). This is a behavioral regression from the previous per-call lookup and breaks workflows/tests that set or mutate this env var after module import (for example via monkeypatch.setenv) to control GPU parallelism per run.

Useful? React with 👍 / 👎.

@slin1237
slin1237 merged commit 11479d1 into main Mar 9, 2026
34 checks passed
@slin1237
slin1237 deleted the refactor/e2e-test-infra-cleanup branch March 9, 2026 22:33
slin1237 added a commit that referenced this pull request Mar 10, 2026
…ssing

- Extract _make_user_message() helper in test_realtime_ws.py to replace
  7 copy-pasted conversation.item.create dict structures, reducing ~70
  lines of boilerplate
- Remove "HF_HOME" from the env-var pass-through tuple in
  benchmarks/conftest.py — it was already explicitly set on line 61 when
  mounting the HF cache directory, causing the variable to be passed
  twice to Docker
- Simplify redundant hasattr+getattr guard in _cleanup_procs() to plain
  getattr with default, since getattr(p, "proc", p) already handles the
  missing attribute case

Refs: #686
Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
slin1237 added a commit that referenced this pull request Mar 10, 2026
…ssing

- Extract _make_user_message() helper in test_realtime_ws.py to replace
  7 copy-pasted conversation.item.create dict structures, reducing ~70
  lines of boilerplate
- Remove "HF_HOME" from the env-var pass-through tuple in
  benchmarks/conftest.py — it was already explicitly set on line 61 when
  mounting the HF cache directory, causing the variable to be passed
  twice to Docker
- Simplify redundant hasattr+getattr guard in _cleanup_procs() to plain
  getattr with default, since getattr(p, "proc", p) already handles the
  missing attribute case

Refs: #686
Signed-off-by: Simo Lin <linsimo.mark@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

tests Test changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant