Skip to content

perf(e2e): reorder tests by model affinity to minimize GPU swaps - #556

Merged
slin1237 merged 3 commits into
mainfrom
keyang/model-pool
Feb 28, 2026
Merged

slin1237 merged 3 commits into
mainfrom
keyang/model-pool

Conversation

@key4ng

@key4ng key4ng commented Feb 27, 2026 •

Copy link
Copy Markdown
Member

Description

Problem

E2E tests take ~16 minutes, with a significant portion spent on GPU model swaps rather than actual testing. Tests alternate between models (openai/gpt-oss-20b for Harmony tests, Qwen/Qwen2.5-14B-Instruct for Local tests) across different files, causing repeated MRU eviction + relaunch cycles (~35–100s each). Both models require 2 GPUs (tp=2), but only GPUs 0–1 are available for swapping.

Solution

Reorder tests at collection time to group them by GPU model affinity — cloud tests first, then all tests for each model contiguously — reducing model swaps from ~7 to 1 and bringing E2E time from ~16 min down to ~12 min.

How it works

 BEFORE (default pytest collection order)
 ─────────────────────────────────────────────────────────────
  cloud │ModelA │ ModelB │ ModelA │ ModelB │ ModelA │ ModelB │ ...
        ↑swap    ↑swap    ↑swap    ↑swap    ↑swap    ↑swap    ↑swap
                         7 model swaps ≈ 4+ min wasted


 AFTER (reordered by model affinity)
 ─────────────────────────────────────────────────────────────
  cloud tests        │ all ModelA tests     │ all ModelB tests
  (no GPU needed)    │ (contiguous)         │ (contiguous)
                     ↑ load ModelA          ↑ 1 swap to ModelB

                         1 model swap ≈ 35–100s

The reordering happens in pytest_collection_modifyitems():

  1. During the existing marker scan, each test's model affinity is recorded — the GPU model it needs, or None for cloud-only tests
  2. After scanning, _reorder_to_minimize_model_swaps() groups tests:
    • Cloud tests (affinity=None) run first — no GPU model needed
    • Per-model groups run contiguously, in first-appearance order
    • Within each group, original collection order is preserved
  3. Test semantics are unchanged — only execution order changes

Changes

  • e2e_test/fixtures/hooks.py: Added _item_model_affinity dict and _LOCAL_BACKENDS frozenset to track each test's GPU model requirement. Populated affinity during the existing marker scan loop. Added _reorder_to_minimize_model_swaps() function that groups tests by model affinity (cloud first, then per-model). Updated reset_collection_state() to clear the new dict.

Test plan

  • Run full E2E suite — all tests pass with no functional change
  • Verify log shows Reordered tests to minimize model swaps: cloud=N, ModelA=N, ModelB=N
  • Confirm cloud tests execute before any local GPU tests
  • Observe reduced model swap count in model_pool logs

🤖 Generated with Claude Code

Summary by CodeRabbit

  • Tests
    • Tests now track per-test model affinity to distinguish local GPU runs from cloud runs.
    • Collection resets this affinity between runs to avoid stale assignments.
    • Collected tests are reordered to run cloud tests first, then contiguous groups per model (preserving within-group order) to reduce model switching and improve test-suite performance.

@coderabbitai

coderabbitai Bot commented Feb 27, 2026 •

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

Per-test model affinity mapping added and persisted during collection; local backend set introduced. Collection reset clears affinity. Tests are reordered so cloud tests run first, then contiguous groups per model (preserving within-group order) to minimize model swaps.

Changes

Cohort / File(s) Summary
Test Affinity Tracking & Reordering
e2e_test/fixtures/hooks.py
Added `_item_model_affinity: dict[str, str

Sequence Diagram(s)

sequenceDiagram
    actor Pytest
    participant Collector as Collection Logic
    participant Affinity as Affinity Store
    participant Reorder as Reordering Logic
    participant Items as Test Items

    Pytest->>Collector: pytest_collection_modifyitems(items)
    Collector->>Items: iterate items to compute backend/model info
    Collector->>Affinity: record _item_model_affinity[item.nodeid] = model_id or None
    loop per item
        Affinity->>Affinity: store mapping
    end
    Collector->>Reorder: call _reorder_to_minimize_model_swaps(items)
    Reorder->>Items: move cloud (None) tests first
    Reorder->>Items: then group items by model_id in first-appearance order, preserving intra-group order
    Reorder->>Pytest: return reordered items
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Poem

🐰 I hopped through collection's night,

Marked each test by model's light.
Cloud first, then clusters in line,
Fewer swaps — the run's just fine.
A tiny map, a tidy flight.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Linked Issues check ⚠️ Warning The PR changes do not address the linked issue #39 about consolidating 'unknown' model id usage. The PR only implements model affinity-based test reordering with no changes to standardize model id identifiers. Either address the model id consolidation requirements from issue #39, or unlink that issue if it is not part of this PR's scope.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title 'perf(e2e): reorder tests by model affinity to minimize GPU swaps' clearly and specifically describes the main change: reordering tests by model affinity to reduce GPU swaps in the E2E test suite.
Out of Scope Changes check ✅ Passed All changes in the PR are directly related to the stated objective of reordering tests by model affinity to minimize GPU swaps, with no unrelated modifications to the codebase.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
  • 📝 Generate docstrings (stacked PR)
  • 📝 Generate docstrings (commit on current branch)
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch keyang/model-pool

Comment @coderabbitai help to get the list of available commands and usage tips.

@github-actions github-actions Bot added the tests Test changes label Feb 27, 2026
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request introduces significant optimizations to the test execution process by implementing a system to track and leverage GPU model affinity. The primary goal is to reduce the overhead associated with switching between different GPU models during test runs, thereby improving efficiency and potentially reducing overall test execution time. This is achieved by reordering tests to run contiguously based on their required GPU model, with cloud tests (which have no specific GPU model affinity) running first.

Highlights

  • Model Affinity Tracking: Introduced a mechanism to track the GPU model affinity for each test item, distinguishing between local GPU-dependent tests and cloud tests.
  • Test Reordering: Implemented a reordering algorithm within pytest_collection_modifyitems to group tests by their model affinity, minimizing GPU model swaps during execution.
  • State Management: Enhanced reset_collection_state to properly clear the newly introduced model affinity data, ensuring clean state for test runs.
  • Local Backend Identification: Defined a set of _LOCAL_BACKENDS to correctly identify backends that require local GPU workers for model affinity tracking.
Changelog
  • e2e_test/fixtures/hooks.py
    • Added _item_model_affinity dictionary to track test item to model ID mapping.
    • Introduced _LOCAL_BACKENDS frozenset to identify backends requiring local GPU workers.
    • Updated reset_collection_state to clear the _item_model_affinity global variable.
    • Modified pytest_collection_modifyitems to declare _item_model_affinity as global.
    • Added logic within pytest_collection_modifyitems to populate _item_model_affinity based on test parameters and backend types.
    • Integrated a call to _reorder_to_minimize_model_swaps within pytest_collection_modifyitems to apply the reordering.
    • Implemented the _reorder_to_minimize_model_swaps function to group tests by model affinity (cloud tests first, then model-specific groups).
Activity
  • No human activity (comments, reviews) has been recorded on this pull request yet.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a test reordering mechanism to minimize GPU model swaps by grouping tests based on their required model. The implementation tracks model affinity for each test and reorders them in pytest_collection_modifyitems. While the overall approach is sound, I've identified a logic issue in how model affinity is assigned, which could cause tests using cloud backends to be incorrectly grouped with local GPU tests. My review includes a specific code suggestion to correct this, ensuring that only tests utilizing local GPUs are assigned model affinity.

Comment thread e2e_test/fixtures/hooks.py Outdated
@key4ng key4ng changed the title feat(tests): implement test reordering to minimize GPU model swaps an… perf(e2e): reorder tests by model affinity to minimize GPU swaps Feb 27, 2026
@key4ng
key4ng marked this pull request as ready for review February 27, 2026 18:11
@mergify

mergify Bot commented Feb 27, 2026

Copy link
Copy Markdown
Contributor

Hi @key4ng, the DCO sign-off check has failed. All commits must include a Signed-off-by line.

To fix existing commits:

# Sign off the last N commits (replace N with the number of unsigned commits)
git rebase HEAD~N --signoff
git push --force-with-lease

To sign off future commits automatically:

  • Use git commit -s every time, or
  • VSCode: enable Git: Always Sign Off in Settings
  • PyCharm: enable Sign-off commit in the Commit tool window

… minimize GPU model swaps

Signed-off-by: key4ng <rukeyang@gmail.com>

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 2884aa7e28

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread e2e_test/fixtures/hooks.py Outdated

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@e2e_test/fixtures/hooks.py`:
- Around line 41-42: The hardcoded _LOCAL_BACKENDS frozenset can drift from the
authoritative source; change it to derive its values from the canonical
enum/constant (e.g., import and iterate over ConnectionMode or the infra
module's central constant) so the set is computed (filtering modes that require
local GPU workers) instead of being manually maintained, or at minimum add a
clear comment pointing to the exact definition of ConnectionMode/infra constant
(include module/class name) so maintainers know where to update when new
backends are added; update the reference in hooks.py to use that derived set
(symbol: _LOCAL_BACKENDS) and ensure tests still pass.

ℹ️ Review info

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between b141ba3 and 9880062.

📒 Files selected for processing (1)
  • e2e_test/fixtures/hooks.py

Comment thread e2e_test/fixtures/hooks.py Outdated
Signed-off-by: key4ng <rukeyang@gmail.com>

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: fb29df29ec

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread e2e_test/fixtures/hooks.py
Signed-off-by: key4ng <rukeyang@gmail.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@e2e_test/fixtures/hooks.py`:
- Around line 239-245: The affinity check is using the aggregated parametrize
marker values (backends) instead of the concrete backend for the current test
iteration; change the condition to look up the per-item backend from
item.callspec.params (e.g., backend = item.callspec.params.get("setup_backend")
or the real param name) and then check if that single backend is in
_LOCAL_BACKENDS, falling back to the existing `any(b in _LOCAL_BACKENDS for b in
backends)` only if callspec or the specific param is not present; update the
expression used in the if-block that references model_id, backends,
_LOCAL_BACKENDS and is_e2e to use this per-item backend variable.

ℹ️ Review info

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between fb29df2 and c35fbe5.

📒 Files selected for processing (1)
  • e2e_test/fixtures/hooks.py

Comment thread e2e_test/fixtures/hooks.py
@slin1237
slin1237 merged commit 6972f14 into main Feb 28, 2026
23 checks passed
@slin1237
slin1237 deleted the keyang/model-pool branch February 28, 2026 03:28
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

tests Test changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants