Skip to content

[CI] Reduce repeated work in vLLM interface analysis - #16364

Merged
wangxiyuan merged 2 commits into
vllm-project:mainfrom
zhao-stack:codex/optimize-vllm-interface-analysis
Sep 12, 2026
Merged

wangxiyuan merged 2 commits into
vllm-project:mainfrom
zhao-stack:codex/optimize-vllm-interface-analysis

Conversation

@zhao-stack

@zhao-stack zhao-stack commented Sep 11, 2026 •

Copy link
Copy Markdown
Collaborator

What this PR does / why we need it?

Reduce repeated work in the CI-only vLLM PR compatibility analyzer introduced by #14560, and fix import-presence checks for deleted or annotation-only names. Changes stay under tests/e2e/vllm_interface/.

The pytest entry, PR base/head resolution, complete source indexing, three analysis branches, deterministic summary, and failure policy remain in place. No persistent cache, pickle files, new environment variables, additional CI jobs, or main2main/monkey-patch analysis are introduced.

Execution flow and implementation details

  1. Resolve and verify inputs. The existing pytest entry resolves the vLLM PR base/head and records the vllm-ascend revision. Commit and checkout verification remains enabled.
  2. Index source trees. InterfaceBoundaryGenerator indexes the complete source trees and finalizes the indexes. Single-path namespace merges copy an already-normalized state instead of repeatedly rebuilding sorted keys and alternative sets. State dictionaries remain independent; immutable binding tuples can be shared.
  3. Resolve aliases and namespace states. RepositoryIndex.canonical_name() checks successively shorter dot-delimited prefixes in the live alias table instead of sorting/scanning every alias per lookup. Longest-prefix behavior and cycle protection remain. Always-bound module names are derived from the final namespace already computed. Decorator and direct-call resolution reuse identical statement-prefix states through a per-module, 256-entry LRU keyed by actual AST statements and guard names. The memo returns independent dictionaries; expression resolution and caller-specific fallback still run for each query. It is not shared across revisions or CI runs.
  4. Discover dependencies. Override and exact-call discovery retains affected call sites and inherited implementation locations. Import discovery reuses the already-parsed vllm-ascend ASTs, with a source-reading fallback for files absent from the index. Triton launch invocation kinds remain distinct.
  5. Read base/head source. Each analysis branch retains separate base/head GitSnapshot instances. Each snapshot lazily uses one git cat-file --batch process rather than launching git show per file. Binary response framing preserves empty files, CRLF and non-ASCII source. Invalid or failed reads raise analysis errors, not missing-symbol findings. An ExitStack closes processes and temporary stderr streams on success or exceptions. This is not a persistent source cache.
  6. Resolve contracts and compare usage. Snapshots reuse module/class namespace states, named bindings, owner lookup and resolved API contracts. Call-contract keys include target expression, access kind, receiver type, member and invocation kind. Argument binding and return-use checks still execute separately at every call site; one site's compatibility result is never reused for another.
  7. Check import presence. Import checks resolve presence before constructing unused signature/return details. Presence now follows the final namespace: a later del removes a binding; an annotation without a value does not create a name or erase an existing value. Constants and re-exports that remain bound are retained. A P1 requires proven presence at the base and proven absence at the head. Ambiguous bindings, star imports, dynamic module exports and unparseable source are not treated as proven removals. Package-submodule fallback is preserved. Full endpoint details are still built for findings.
  8. Merge and print. The three branches merge and sort findings deterministically. The CI entry prints the summary directly; it does not create report artifacts. Historical findings remain subject to the existing PR-only filtering policy.

Does this PR introduce any user-facing change?

Reduced analyzer runtime and more accurate detection of imports whose names were deleted or became annotation-only. This is therefore not exclusively a performance-only change. Analyzer version advances to 2.1.1. CLI options and CI log wording are unchanged; existing dynamic-dispatch and return-dataflow limitations remain.

How was this patch tested?

Local regressions

  • 123 local tests passed: Git batch-reader lifecycle/error handling, alias equivalence, final import bindings, callable/return contracts, serial/parallel fixture equivalence, pytest-entry integration, and scope-prefix isolation/eviction.
  • Included 12,000 synthetic alias comparisons and 196 generated scope-flow combinations. Deleted and annotation-only imports were exercised through actual CLI runs against temporary Git repositories.
  • The pytest integration fixtures launch the analyzer subprocess; the analyzer is not mocked. Historical replay scans described below call the analyzer core directly, not the hardware E2E suite.
  • Ruff, formatting, compileall and mypy checks passed in local validation. Full repository formatting was previously attempted; some hooks require Linux shell tools unavailable on this Windows host. Full Linux CI and NPU workloads are not claimed as locally verified.
  • Regression/replay harnesses remain outside the E2E collection path, as requested; no analyzer unit-test directory is reintroduced into vLLM PR jobs.

Latest paired performance measurement

Windows / Python 3.11, four indexing processes and three analysis threads, fixed source SHAs, fresh processes, sequential scans and no persistent analyzer cache. Times below are analyzer time, excluding network range discovery and hardware tests.

vLLM PR Before latest scope reuse After Reduction Result on both sides
#39568 73.6 s 61.9 s 15.9% 1 P1
#50685 100.3 s 82.0 s 18.2% 3 P1 impacts / 1 root cause
#50620 100.4 s 81.8 s 18.5% PASS

This baseline already includes the earlier Git batching, alias optimization and import fix; it is not the original pre-optimization implementation. Full reports matched after removing only timing metadata, stdout matched byte-for-byte, and expected exit codes were 1/1/0. These are single paired measurements, not statistical benchmarks or Linux CI speed guarantees.

Expanded accuracy replay

Eight additional cases were each run against the frozen pre-scope-reuse baseline and current code: 16 actual scans. All eight full reports matched except timing, and all eight stdout logs were byte-identical.

Fixed historical input P1 impacts P1 roots Review findings
vllm-ascend #11709 range 5 5 6
vllm-ascend #12020 range 2 1 0
vllm-ascend #12420 range 0 0 0
vllm-ascend #12502 range 33 20 0
vllm-ascend #12648 range 0 0 5
vllm-ascend #13358 range 5 4 11
vLLM #47808 2 2 5
vLLM #50504 0 0 0

The upgrade ranges are fixed CI-analyzer inputs, not a full main2main mode. #12648 uses its adapted vllm-ascend head; the other upgrade cases use their pre-adaptation baselines. The earlier three cases above were reused only after verifying current source hashes, giving 11 distinct cases overall. There were not 22 new scans in the expanded round.

29 semantic/history assertions passed. Previously confirmed P1 findings remained. #47808 matches the newer August 28 historical log (2 P1), not the older August 17 report (1 P1). #11709's six review items already existed in the frozen baseline; they are not introduced by scope reuse.

Known limitations are not hidden by these results: #12648's compute_slot_mappings(out=...) remains review-only because its dispatch requires monkey-patch/field propagation; broader return-value consumption is still outside the exact dependency model. Regression equivalence does not establish universal precision/recall.

Fresh verification of the published PR head

After pushing commit 2e31d3bcbdc82960bc0a1956abdd412f6ef1f27b, fetched refs/pull/16364/head from vllm-project/vllm-ascend and checked it out in a separate, clean detached worktree. The fetched commit and all nine analyzer source hashes matched the validated version.

Three additional fresh scans used that fetched code:

vLLM PR Result Analyzer time Process wall time Exit code
#39568 1 P1 59.8 s 61.0 s 1
#50685 3 P1 impacts / 1 root cause 82.2 s 83.7 s 1
#50620 PASS 79.8 s 81.3 s 0

All three complete reports were identical to the pre-push reports after excluding only timing metadata; raw stdout was byte-identical. Exit 1 is the expected detected-break outcome, not an analyzer crash. These runs are separate from the earlier paired measurements and expanded regression round.

The 123 local tests passed again before pushing. Ruff, formatting, spelling and Markdown checks passed; several shell-based repository hooks could not execute on Windows. The fetched code also passed mypy with the Python 3.10/Linux target. This is static type checking, not execution on Linux.

These historical scans execute analyze_range() and the production summary renderer directly. They do not run the full upstream pytest job or NPU workloads. The actual pytest entry is covered separately by the local integration fixtures.

Published-head verification details: #16364 (comment)

Reproduce a pinned case

With this PR's analyzer checked out, prepare separate clean source worktrees at the following head/revision:

vLLM PR vLLM base vLLM head vllm-ascend revision
#39568 ce29c26b31d432b1b4bc028c46bb2c3b07a667d8 c7560af42487b1570c4e6f4cea5df1605a4d59fc 60f0238b0eec4c91fe466497ae8862daf521aecc
#50685 1be36283678a9a94fc8fdaad6c95c2896d6b4015 c05d75aaa95cf89f547503044c1921625905085d f258cbbd898f2b05f38d96b20d1530d5e10f7923
#50620 c05d75aaa95cf89f547503044c1921625905085d 653cc6faca6885e36760bf35a25bb63442519b14 f258cbbd898f2b05f38d96b20d1530d5e10f7923

The analyzer checkout supplies the tool; --vllm-root and --ascend-root supply the pinned source trees being analyzed. They do not need to be the analyzer checkout itself. Existing repositories can be reused when the head/revision and required base objects match the table; no model installation or NPU is required for this static CLI scan.

For example, run #39568 from this PR's analyzer checkout after preparing the two source paths at the revisions above:

VLLM_INTERFACE_TIMINGS=1 python -u -m tests.e2e.vllm_interface.vllm_interface_contracts analyze-range \
  --vllm-root /path/to/vllm-source \
  --ascend-root /path/to/vllm-ascend-source \
  --old ce29c26b31d432b1b4bc028c46bb2c3b07a667d8 \
  --new c7560af42487b1570c4e6f4cea5df1605a4d59fc \
  --expect-ascend-sha 60f0238b0eec4c91fe466497ae8862daf521aecc \
  --index-workers 4 --analysis-workers 3 --fail-on introduced

PowerShell equivalent (replace the two source directory paths):

$env:VLLM_INTERFACE_TIMINGS = "1"
python -u -m tests.e2e.vllm_interface.vllm_interface_contracts analyze-range `
  --vllm-root "C:\path\to\vllm-source" `
  --ascend-root "C:\path\to\vllm-ascend-source" `
  --old ce29c26b31d432b1b4bc028c46bb2c3b07a667d8 `
  --new c7560af42487b1570c4e6f4cea5df1605a4d59fc `
  --expect-ascend-sha 60f0238b0eec4c91fe466497ae8862daf521aecc `
  --index-workers 4 --analysis-workers 3 --fail-on introduced
Write-Host "Analyzer exit code: $LASTEXITCODE"

Timing diagnostics are printed as phases finish, followed by the compatibility summary; this is not per-file progress. #39568 reports BREAKS FOUND: vLLM removes SchedulerInterface._get_routed_experts, still called at vllm_ascend/core/recompute_scheduler.py:907 in the pinned baseline. Expected exit: 1 for #39568/#50685 (detected P1), 0 for #50620.

These commands use the PR's CLI, not pytest: they bypass PR-range network discovery and parent conftest.py dependency loading while exercising the same analysis engine. Base objects must already exist locally to avoid Git fetching missing objects. The SHAs are replay inputs, not analyzer constants. Raw logs and full comparison evidence are retained locally; the CI entry itself still prints logs without creating report artifacts.

Signed-off-by: shenzhao <shenzhao9@huawei.com>
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request introduces performance optimizations to the vLLM PR compatibility analyzer used in CI. By implementing memoization for scopes, bindings, and API contracts, and optimizing the state merging logic, the analyzer significantly reduces redundant work without altering the analysis scope, compatibility rules, or output reports. These changes are strictly in-process and do not introduce persistent caches or external dependencies.

Highlights

  • Performance Optimization: Implemented memoization and state reuse across the vLLM interface analyzer to reduce redundant computations during CI runs, resulting in approximately 40-45% reduction in analysis runtime.
  • State Management: Optimized scope state merging by copying normalized states when only one execution path exists, avoiding unnecessary rebuilding of binding sets.
  • Memoization Tables: Introduced internal memo tables for scopes, named bindings, and endpoints within GitSnapshot to ensure computations are performed once per revision and branch.
  • Import Analysis: Refined import presence checks to resolve symbols efficiently without constructing full contracts, and reused parsed ASTs to minimize redundant file parsing.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@github-actions github-actions Bot added documentation Improvements or additions to documentation module:tests labels Sep 11, 2026
@github-actions

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces in-process optimizations to the vLLM interface compatibility analyzer to reduce redundant computations. It implements caching for scope bindings, named bindings, endpoints, and import presence within the GitSnapshot class, and allows reusing already-parsed module trees during import discovery. The review feedback correctly identifies that the PR title and summary do not adhere to the repository's style guide and provides a compliant suggestion.

Comment on lines +72 to +75
### Avoiding repeated work within a run

The scope interpreter copies an already-normalized state when there is only one execution path, instead of sorting
and merging every name again. Multiple paths still use the existing binding-alternative merge.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The Pull Request Title does not fully adhere to the Repository Style Guide format: [Branch][Module][Action] Pull Request Title. The [Action] tag (e.g., [Misc]) is missing.

Here is the suggested title and summary matching the style guide requirements:

Suggested PR Title:

[CI][Misc] Reduce repeated work in vLLM interface analysis

Suggested PR Summary:

### What this PR does / why we need it?

Reduce repeated work in the CI-only vLLM PR compatibility analyzer introduced by #14560. This PR changes only the analyzer and its README under `tests/e2e/vllm_interface/`.

The analysis scope, compatibility rules, finding locations, and printed summary are unchanged. No persistent cache, pickle files, new environment variables, additional CI jobs, or main2main/monkey-patch analysis are introduced.

### Does this PR introduce _any_ user-facing change?

Only reduced analyzer runtime. No intended change to detected breaks, review findings, CLI options, or CI log wording. Existing detection limitations remain unchanged.

### How was this patch tested?

Tested via paired real-PR replay before submission and local regression checks.
References
  1. The summary and title should follow the specific format defined in the Repository Style Guide. (link)

@zhao-stack

Copy link
Copy Markdown
Collaborator Author

Published PR code verification

Verified analyzer commit: c37215c8c34fc842cf546e7db46817baf5a978f3 (fetched from this PR's GitHub head ref).

Both replays ran the actual analyze-range CLI from a clean detached worktree of this commit. The vLLM ranges and vllm-ascend revisions were pinned to the historical fixtures listed in the PR description. These are local Windows replays, not an upstream Buildkite/NPU job run.

vLLM PR Result CLI exit code Wall time stdout vs. pre-submit optimized run
#39568 1 P1 / 1 API root cause 1 (expected) 440.0 s Byte-for-byte identical
#50620 PASS / 0 breaks / 0 review items 0 (expected) 550.0 s Byte-for-byte identical

The breaking replay still identifies removal of SchedulerInterface._get_routed_experts, called at vllm_ascend/core/recompute_scheduler.py:907. Exit code 1 is the expected --fail-on introduced result, not an analyzer crash.

The local regression suite also passed again against this published commit: 37 passed in 8.91 s. The suite includes an actual pytest E2E entry-point/subprocess fixture. Regression helpers remain outside the submitted CI directory.

No tracked files changed during verification. The original pre-optimization and optimized benchmark reports were also equivalent after removing timing metadata; this additional published-commit replay confirms the same stdout and expected exits. This establishes equivalence for the tested cases, not proof of complete detection of every possible interface change.

GitHub checks

At verification time, DCO, pre-commit, CPU unit tests and ci-gate all passed. Hardware test jobs were skipped; this is not evidence of an upstream NPU full-CI run. The local artifact bundle retains a GitHub status snapshot alongside the raw logs.

Raw replay output

stdout and stderr are shown separately below, without rewriting their content. The three analysis branches run concurrently, so phase timings must not be added together.

vLLM #39568 — raw stdout
# vLLM PR Compatibility with vllm-ascend

**Result: BREAKS FOUND**

- This PR: `ce29c26b31d432b1b4bc028c46bb2c3b07a667d8` -> `c7560af42487b1570c4e6f4cea5df1605a4d59fc`
- vllm-ascend revision: `60f0238b0eec4c91fe466497ae8862daf521aecc`
- Breaks introduced by this PR: 1
- Distinct vLLM API changes causing breaks: 1
- Items for review: 0
- Distinct vLLM API changes for review: 0

## Breaks introduced by this PR

### 1. P1 direct_call / call_target_presence

- vLLM API changed by this PR: `vllm/v1/core/sched/interface.py:SchedulerInterface._get_routed_experts`
- Affected vllm-ascend code: `vllm_ascend/core/recompute_scheduler.py:907`
- Compatibility impact: This PR removes a vLLM API called by vllm-ascend.

## Items for review

No additional interface change requires manual review.
vLLM #39568 — raw stderr
[vllm-interface] input_verification: 1.872s
[vllm-interface] repository_indexing: 59.308s
[vllm-interface] relation_comparison: 70.791s
[vllm-interface] direct_import_analysis: 261.925s
[vllm-interface] direct_call_discovery: 52.756s
[vllm-interface] direct_call_comparison: 171.091s
vLLM #50620 — raw stdout
# vLLM PR Compatibility with vllm-ascend

**Result: PASS**

- This PR: `c05d75aaa95cf89f547503044c1921625905085d` -> `653cc6faca6885e36760bf35a25bb63442519b14`
- vllm-ascend revision: `f258cbbd898f2b05f38d96b20d1530d5e10f7923`
- Breaks introduced by this PR: 0
- Distinct vLLM API changes causing breaks: 0
- Items for review: 0
- Distinct vLLM API changes for review: 0

## Breaks introduced by this PR

This PR does not introduce a detected interface break in vllm-ascend.

## Items for review

No additional interface change requires manual review.
vLLM #50620 — raw stderr
[vllm-interface] input_verification: 1.798s
[vllm-interface] repository_indexing: 80.255s
[vllm-interface] relation_comparison: 82.903s
[vllm-interface] direct_import_analysis: 279.207s
[vllm-interface] direct_call_discovery: 73.418s
[vllm-interface] direct_call_comparison: 181.927s

Batch Git blob reads, reuse module and prefix namespace states, and resolve aliases by qualified prefixes. Determine import presence from final bindings while preserving ambiguous exports for review.

Signed-off-by: shenzhao <shenzhao9@huawei.com>
@zhao-stack

Copy link
Copy Markdown
Collaborator Author

Published PR head: fresh local verification

Commit: 2e31d3bcbdc82960bc0a1956abdd412f6ef1f27b.

Fetched refs/pull/16364/head from vllm-project/vllm-ascend into a separate, clean detached worktree after pushing. All nine analyzer source hashes match the locally validated version.

Three fresh, sequential pinned-input scans used the fetched analyzer code (Windows / Python 3.11, 4 indexing processes, 3 analysis threads, no persistent analyzer cache).

vLLM PR P1 impacts Analysis time Exit Comparison with pre-push output
#39568 1 59.8 s 1 Identical report except timing; byte-identical stdout
#50685 3 82.2 s 1 Identical report except timing; byte-identical stdout
#50620 0 79.8 s 0 Identical report except timing; byte-identical stdout

Exit 1 is the expected detected-break outcome for #39568/#50685, not an analyzer crash. #50620 passed.

These real-history scans call analyze_range and the production summary renderer directly. They do not claim to execute the full upstream pytest job or NPU workloads. Before pushing, all 123 local regression tests passed again, including pytest-entry integration fixtures. Ruff, formatting, spelling and Markdown checks passed; several repository-wide shell hooks could not run on this Windows host.

Raw stdout/stderr, complete reports, commands, timings and source hashes are retained locally. No detection differences were observed in these samples; existing model limitations still apply.

@wangxiyuan
wangxiyuan merged commit 3b7cf0f into vllm-project:main Sep 12, 2026
13 checks passed
sunny-rain-63 pushed a commit to sunny-rain-63/vllm-ascend that referenced this pull request Sep 12, 2026
)

### What this PR does / why we need it?

Reduce repeated work in the CI-only vLLM PR compatibility analyzer
introduced by vllm-project#14560, and fix import-presence checks for deleted or
annotation-only names. Changes stay under `tests/e2e/vllm_interface/`.

The pytest entry, PR base/head resolution, complete source indexing,
three analysis branches, deterministic summary, and failure policy
remain in place. No persistent cache, pickle files, new environment
variables, additional CI jobs, or main2main/monkey-patch analysis are
introduced.

#### Execution flow and implementation details

1. **Resolve and verify inputs.** The existing pytest entry resolves the
vLLM PR base/head and records the vllm-ascend revision. Commit and
checkout verification remains enabled.
2. **Index source trees.** `InterfaceBoundaryGenerator` indexes the
complete source trees and finalizes the indexes. Single-path namespace
merges copy an already-normalized state instead of repeatedly rebuilding
sorted keys and alternative sets. State dictionaries remain independent;
immutable binding tuples can be shared.
3. **Resolve aliases and namespace states.**
`RepositoryIndex.canonical_name()` checks successively shorter
dot-delimited prefixes in the live alias table instead of
sorting/scanning every alias per lookup. Longest-prefix behavior and
cycle protection remain. Always-bound module names are derived from the
final namespace already computed. Decorator and direct-call resolution
reuse identical statement-prefix states through a per-module, 256-entry
LRU keyed by actual AST statements and guard names. The memo returns
independent dictionaries; expression resolution and caller-specific
fallback still run for each query. It is not shared across revisions or
CI runs.
4. **Discover dependencies.** Override and exact-call discovery retains
affected call sites and inherited implementation locations. Import
discovery reuses the already-parsed vllm-ascend ASTs, with a
source-reading fallback for files absent from the index. Triton launch
invocation kinds remain distinct.
5. **Read base/head source.** Each analysis branch retains separate
base/head `GitSnapshot` instances. Each snapshot lazily uses one `git
cat-file --batch` process rather than launching `git show` per file.
Binary response framing preserves empty files, CRLF and non-ASCII
source. Invalid or failed reads raise analysis errors, not
missing-symbol findings. An `ExitStack` closes processes and temporary
stderr streams on success or exceptions. This is not a persistent source
cache.
6. **Resolve contracts and compare usage.** Snapshots reuse module/class
namespace states, named bindings, owner lookup and resolved API
contracts. Call-contract keys include target expression, access kind,
receiver type, member and invocation kind. Argument binding and
return-use checks still execute separately at every call site; one
site's compatibility result is never reused for another.
7. **Check import presence.** Import checks resolve presence before
constructing unused signature/return details. Presence now follows the
final namespace: a later `del` removes a binding; an annotation without
a value does not create a name or erase an existing value. Constants and
re-exports that remain bound are retained. A P1 requires proven presence
at the base and proven absence at the head. Ambiguous bindings, star
imports, dynamic module exports and unparseable source are not treated
as proven removals. Package-submodule fallback is preserved. Full
endpoint details are still built for findings.
8. **Merge and print.** The three branches merge and sort findings
deterministically. The CI entry prints the summary directly; it does not
create report artifacts. Historical findings remain subject to the
existing PR-only filtering policy.

### Does this PR introduce _any_ user-facing change?

Reduced analyzer runtime and more accurate detection of imports whose
names were deleted or became annotation-only. This is therefore not
exclusively a performance-only change. Analyzer version advances to
`2.1.1`. CLI options and CI log wording are unchanged; existing
dynamic-dispatch and return-dataflow limitations remain.

### How was this patch tested?

#### Local regressions

- 123 local tests passed: Git batch-reader lifecycle/error handling,
alias equivalence, final import bindings, callable/return contracts,
serial/parallel fixture equivalence, pytest-entry integration, and
scope-prefix isolation/eviction.
- Included 12,000 synthetic alias comparisons and 196 generated
scope-flow combinations. Deleted and annotation-only imports were
exercised through actual CLI runs against temporary Git repositories.
- The pytest integration fixtures launch the analyzer subprocess; the
analyzer is not mocked. Historical replay scans described below call the
analyzer core directly, not the hardware E2E suite.
- Ruff, formatting, compileall and mypy checks passed in local
validation. Full repository formatting was previously attempted; some
hooks require Linux shell tools unavailable on this Windows host. Full
Linux CI and NPU workloads are not claimed as locally verified.
- Regression/replay harnesses remain outside the E2E collection path, as
requested; no analyzer unit-test directory is reintroduced into vLLM PR
jobs.

#### Latest paired performance measurement

Windows / Python 3.11, four indexing processes and three analysis
threads, fixed source SHAs, fresh processes, sequential scans and no
persistent analyzer cache. Times below are analyzer time, excluding
network range discovery and hardware tests.

| vLLM PR | Before latest scope reuse | After | Reduction | Result on
both sides |
| --- | ---: | ---: | ---: | --- |
| #39568 | 73.6 s | 61.9 s | 15.9% | 1 P1 |
| #50685 | 100.3 s | 82.0 s | 18.2% | 3 P1 impacts / 1 root cause |
| #50620 | 100.4 s | 81.8 s | 18.5% | PASS |

This baseline already includes the earlier Git batching, alias
optimization and import fix; it is **not** the original pre-optimization
implementation. Full reports matched after removing only timing
metadata, stdout matched byte-for-byte, and expected exit codes were
1/1/0. These are single paired measurements, not statistical benchmarks
or Linux CI speed guarantees.

#### Expanded accuracy replay

Eight additional cases were each run against the frozen pre-scope-reuse
baseline and current code: 16 actual scans. All eight full reports
matched except timing, and all eight stdout logs were byte-identical.

| Fixed historical input | P1 impacts | P1 roots | Review findings |
| --- | ---: | ---: | ---: |
| vllm-ascend vllm-project#11709 range | 5 | 5 | 6 |
| vllm-ascend vllm-project#12020 range | 2 | 1 | 0 |
| vllm-ascend vllm-project#12420 range | 0 | 0 | 0 |
| vllm-ascend vllm-project#12502 range | 33 | 20 | 0 |
| vllm-ascend vllm-project#12648 range | 0 | 0 | 5 |
| vllm-ascend vllm-project#13358 range | 5 | 4 | 11 |
| vLLM #47808 | 2 | 2 | 5 |
| vLLM #50504 | 0 | 0 | 0 |

The upgrade ranges are fixed CI-analyzer inputs, not a full main2main
mode. vllm-project#12648 uses its adapted vllm-ascend head; the other upgrade cases
use their pre-adaptation baselines. The earlier three cases above were
reused only after verifying current source hashes, giving 11 distinct
cases overall. There were not 22 new scans in the expanded round.

29 semantic/history assertions passed. Previously confirmed P1 findings
remained. #47808 matches the newer August 28 historical log (2 P1), not
the older August 17 report (1 P1). vllm-project#11709's six review items already
existed in the frozen baseline; they are not introduced by scope reuse.

Known limitations are not hidden by these results: vllm-project#12648's
`compute_slot_mappings(out=...)` remains review-only because its
dispatch requires monkey-patch/field propagation; broader return-value
consumption is still outside the exact dependency model. Regression
equivalence does not establish universal precision/recall.

#### Fresh verification of the published PR head

After pushing commit `2e31d3bcbdc82960bc0a1956abdd412f6ef1f27b`, fetched
`refs/pull/16364/head` from `vllm-project/vllm-ascend` and checked it
out in a separate, clean detached worktree. The fetched commit and all
nine analyzer source hashes matched the validated version.

Three additional fresh scans used that fetched code:

| vLLM PR | Result | Analyzer time | Process wall time | Exit code |
| --- | --- | ---: | ---: | ---: |
| #39568 | 1 P1 | 59.8 s | 61.0 s | 1 |
| #50685 | 3 P1 impacts / 1 root cause | 82.2 s | 83.7 s | 1 |
| #50620 | PASS | 79.8 s | 81.3 s | 0 |

All three complete reports were identical to the pre-push reports after
excluding only timing metadata; raw stdout was byte-identical. Exit 1 is
the expected detected-break outcome, not an analyzer crash. These runs
are separate from the earlier paired measurements and expanded
regression round.

The 123 local tests passed again before pushing. Ruff, formatting,
spelling and Markdown checks passed; several shell-based repository
hooks could not execute on Windows. The fetched code also passed mypy
with the Python 3.10/Linux target. This is static type checking, not
execution on Linux.

These historical scans execute `analyze_range()` and the production
summary renderer directly. They do not run the full upstream pytest job
or NPU workloads. The actual pytest entry is covered separately by the
local integration fixtures.

Published-head verification details:
vllm-project#16364 (comment)

#### Reproduce a pinned case

With this PR's analyzer checked out, prepare separate clean source
worktrees at the following head/revision:

| vLLM PR | vLLM base | vLLM head | vllm-ascend revision |
| --- | --- | --- | --- |
| #39568 | `ce29c26b31d432b1b4bc028c46bb2c3b07a667d8` |
`c7560af42487b1570c4e6f4cea5df1605a4d59fc` |
`60f0238b0eec4c91fe466497ae8862daf521aecc` |
| #50685 | `1be36283678a9a94fc8fdaad6c95c2896d6b4015` |
`c05d75aaa95cf89f547503044c1921625905085d` |
`f258cbbd898f2b05f38d96b20d1530d5e10f7923` |
| #50620 | `c05d75aaa95cf89f547503044c1921625905085d` |
`653cc6faca6885e36760bf35a25bb63442519b14` |
`f258cbbd898f2b05f38d96b20d1530d5e10f7923` |

The analyzer checkout supplies the tool; `--vllm-root` and
`--ascend-root` supply the pinned source trees being analyzed. They do
not need to be the analyzer checkout itself. Existing repositories can
be reused when the head/revision and required base objects match the
table; no model installation or NPU is required for this static CLI
scan.

For example, run #39568 from this PR's analyzer checkout after preparing
the two source paths at the revisions above:

```bash
VLLM_INTERFACE_TIMINGS=1 python -u -m tests.e2e.vllm_interface.vllm_interface_contracts analyze-range \
  --vllm-root /path/to/vllm-source \
  --ascend-root /path/to/vllm-ascend-source \
  --old ce29c26 \
  --new c7560af \
  --expect-ascend-sha 60f0238 \
  --index-workers 4 --analysis-workers 3 --fail-on introduced
```

PowerShell equivalent (replace the two source directory paths):

```powershell
$env:VLLM_INTERFACE_TIMINGS = "1"
python -u -m tests.e2e.vllm_interface.vllm_interface_contracts analyze-range `
  --vllm-root "C:\path\to\vllm-source" `
  --ascend-root "C:\path\to\vllm-ascend-source" `
  --old ce29c26 `
  --new c7560af `
  --expect-ascend-sha 60f0238 `
  --index-workers 4 --analysis-workers 3 --fail-on introduced
Write-Host "Analyzer exit code: $LASTEXITCODE"
```

Timing diagnostics are printed as phases finish, followed by the
compatibility summary; this is not per-file progress. #39568 reports
`BREAKS FOUND`: vLLM removes `SchedulerInterface._get_routed_experts`,
still called at `vllm_ascend/core/recompute_scheduler.py:907` in the
pinned baseline. Expected exit: 1 for #39568/#50685 (detected P1), 0 for
#50620.

These commands use the PR's CLI, not pytest: they bypass PR-range
network discovery and parent `conftest.py` dependency loading while
exercising the same analysis engine. Base objects must already exist
locally to avoid Git fetching missing objects. The SHAs are replay
inputs, not analyzer constants. Raw logs and full comparison evidence
are retained locally; the CI entry itself still prints logs without
creating report artifacts.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: shenzhao <shenzhao9@huawei.com>
Co-authored-by: shenzhao <shenzhao9@huawei.com>
johnnysluckydays pushed a commit to johnnysluckydays/vllm-ascend that referenced this pull request Sep 14, 2026
)

### What this PR does / why we need it?

Reduce repeated work in the CI-only vLLM PR compatibility analyzer
introduced by vllm-project#14560, and fix import-presence checks for deleted or
annotation-only names. Changes stay under `tests/e2e/vllm_interface/`.

The pytest entry, PR base/head resolution, complete source indexing,
three analysis branches, deterministic summary, and failure policy
remain in place. No persistent cache, pickle files, new environment
variables, additional CI jobs, or main2main/monkey-patch analysis are
introduced.

#### Execution flow and implementation details

1. **Resolve and verify inputs.** The existing pytest entry resolves the
vLLM PR base/head and records the vllm-ascend revision. Commit and
checkout verification remains enabled.
2. **Index source trees.** `InterfaceBoundaryGenerator` indexes the
complete source trees and finalizes the indexes. Single-path namespace
merges copy an already-normalized state instead of repeatedly rebuilding
sorted keys and alternative sets. State dictionaries remain independent;
immutable binding tuples can be shared.
3. **Resolve aliases and namespace states.**
`RepositoryIndex.canonical_name()` checks successively shorter
dot-delimited prefixes in the live alias table instead of
sorting/scanning every alias per lookup. Longest-prefix behavior and
cycle protection remain. Always-bound module names are derived from the
final namespace already computed. Decorator and direct-call resolution
reuse identical statement-prefix states through a per-module, 256-entry
LRU keyed by actual AST statements and guard names. The memo returns
independent dictionaries; expression resolution and caller-specific
fallback still run for each query. It is not shared across revisions or
CI runs.
4. **Discover dependencies.** Override and exact-call discovery retains
affected call sites and inherited implementation locations. Import
discovery reuses the already-parsed vllm-ascend ASTs, with a
source-reading fallback for files absent from the index. Triton launch
invocation kinds remain distinct.
5. **Read base/head source.** Each analysis branch retains separate
base/head `GitSnapshot` instances. Each snapshot lazily uses one `git
cat-file --batch` process rather than launching `git show` per file.
Binary response framing preserves empty files, CRLF and non-ASCII
source. Invalid or failed reads raise analysis errors, not
missing-symbol findings. An `ExitStack` closes processes and temporary
stderr streams on success or exceptions. This is not a persistent source
cache.
6. **Resolve contracts and compare usage.** Snapshots reuse module/class
namespace states, named bindings, owner lookup and resolved API
contracts. Call-contract keys include target expression, access kind,
receiver type, member and invocation kind. Argument binding and
return-use checks still execute separately at every call site; one
site's compatibility result is never reused for another.
7. **Check import presence.** Import checks resolve presence before
constructing unused signature/return details. Presence now follows the
final namespace: a later `del` removes a binding; an annotation without
a value does not create a name or erase an existing value. Constants and
re-exports that remain bound are retained. A P1 requires proven presence
at the base and proven absence at the head. Ambiguous bindings, star
imports, dynamic module exports and unparseable source are not treated
as proven removals. Package-submodule fallback is preserved. Full
endpoint details are still built for findings.
8. **Merge and print.** The three branches merge and sort findings
deterministically. The CI entry prints the summary directly; it does not
create report artifacts. Historical findings remain subject to the
existing PR-only filtering policy.

### Does this PR introduce _any_ user-facing change?

Reduced analyzer runtime and more accurate detection of imports whose
names were deleted or became annotation-only. This is therefore not
exclusively a performance-only change. Analyzer version advances to
`2.1.1`. CLI options and CI log wording are unchanged; existing
dynamic-dispatch and return-dataflow limitations remain.

### How was this patch tested?

#### Local regressions

- 123 local tests passed: Git batch-reader lifecycle/error handling,
alias equivalence, final import bindings, callable/return contracts,
serial/parallel fixture equivalence, pytest-entry integration, and
scope-prefix isolation/eviction.
- Included 12,000 synthetic alias comparisons and 196 generated
scope-flow combinations. Deleted and annotation-only imports were
exercised through actual CLI runs against temporary Git repositories.
- The pytest integration fixtures launch the analyzer subprocess; the
analyzer is not mocked. Historical replay scans described below call the
analyzer core directly, not the hardware E2E suite.
- Ruff, formatting, compileall and mypy checks passed in local
validation. Full repository formatting was previously attempted; some
hooks require Linux shell tools unavailable on this Windows host. Full
Linux CI and NPU workloads are not claimed as locally verified.
- Regression/replay harnesses remain outside the E2E collection path, as
requested; no analyzer unit-test directory is reintroduced into vLLM PR
jobs.

#### Latest paired performance measurement

Windows / Python 3.11, four indexing processes and three analysis
threads, fixed source SHAs, fresh processes, sequential scans and no
persistent analyzer cache. Times below are analyzer time, excluding
network range discovery and hardware tests.

| vLLM PR | Before latest scope reuse | After | Reduction | Result on
both sides |
| --- | ---: | ---: | ---: | --- |
| #39568 | 73.6 s | 61.9 s | 15.9% | 1 P1 |
| #50685 | 100.3 s | 82.0 s | 18.2% | 3 P1 impacts / 1 root cause |
| #50620 | 100.4 s | 81.8 s | 18.5% | PASS |

This baseline already includes the earlier Git batching, alias
optimization and import fix; it is **not** the original pre-optimization
implementation. Full reports matched after removing only timing
metadata, stdout matched byte-for-byte, and expected exit codes were
1/1/0. These are single paired measurements, not statistical benchmarks
or Linux CI speed guarantees.

#### Expanded accuracy replay

Eight additional cases were each run against the frozen pre-scope-reuse
baseline and current code: 16 actual scans. All eight full reports
matched except timing, and all eight stdout logs were byte-identical.

| Fixed historical input | P1 impacts | P1 roots | Review findings |
| --- | ---: | ---: | ---: |
| vllm-ascend vllm-project#11709 range | 5 | 5 | 6 |
| vllm-ascend vllm-project#12020 range | 2 | 1 | 0 |
| vllm-ascend vllm-project#12420 range | 0 | 0 | 0 |
| vllm-ascend vllm-project#12502 range | 33 | 20 | 0 |
| vllm-ascend vllm-project#12648 range | 0 | 0 | 5 |
| vllm-ascend vllm-project#13358 range | 5 | 4 | 11 |
| vLLM #47808 | 2 | 2 | 5 |
| vLLM #50504 | 0 | 0 | 0 |

The upgrade ranges are fixed CI-analyzer inputs, not a full main2main
mode. vllm-project#12648 uses its adapted vllm-ascend head; the other upgrade cases
use their pre-adaptation baselines. The earlier three cases above were
reused only after verifying current source hashes, giving 11 distinct
cases overall. There were not 22 new scans in the expanded round.

29 semantic/history assertions passed. Previously confirmed P1 findings
remained. #47808 matches the newer August 28 historical log (2 P1), not
the older August 17 report (1 P1). vllm-project#11709's six review items already
existed in the frozen baseline; they are not introduced by scope reuse.

Known limitations are not hidden by these results: vllm-project#12648's
`compute_slot_mappings(out=...)` remains review-only because its
dispatch requires monkey-patch/field propagation; broader return-value
consumption is still outside the exact dependency model. Regression
equivalence does not establish universal precision/recall.

#### Fresh verification of the published PR head

After pushing commit `2e31d3bcbdc82960bc0a1956abdd412f6ef1f27b`, fetched
`refs/pull/16364/head` from `vllm-project/vllm-ascend` and checked it
out in a separate, clean detached worktree. The fetched commit and all
nine analyzer source hashes matched the validated version.

Three additional fresh scans used that fetched code:

| vLLM PR | Result | Analyzer time | Process wall time | Exit code |
| --- | --- | ---: | ---: | ---: |
| #39568 | 1 P1 | 59.8 s | 61.0 s | 1 |
| #50685 | 3 P1 impacts / 1 root cause | 82.2 s | 83.7 s | 1 |
| #50620 | PASS | 79.8 s | 81.3 s | 0 |

All three complete reports were identical to the pre-push reports after
excluding only timing metadata; raw stdout was byte-identical. Exit 1 is
the expected detected-break outcome, not an analyzer crash. These runs
are separate from the earlier paired measurements and expanded
regression round.

The 123 local tests passed again before pushing. Ruff, formatting,
spelling and Markdown checks passed; several shell-based repository
hooks could not execute on Windows. The fetched code also passed mypy
with the Python 3.10/Linux target. This is static type checking, not
execution on Linux.

These historical scans execute `analyze_range()` and the production
summary renderer directly. They do not run the full upstream pytest job
or NPU workloads. The actual pytest entry is covered separately by the
local integration fixtures.

Published-head verification details:
vllm-project#16364 (comment)

#### Reproduce a pinned case

With this PR's analyzer checked out, prepare separate clean source
worktrees at the following head/revision:

| vLLM PR | vLLM base | vLLM head | vllm-ascend revision |
| --- | --- | --- | --- |
| #39568 | `ce29c26b31d432b1b4bc028c46bb2c3b07a667d8` |
`c7560af42487b1570c4e6f4cea5df1605a4d59fc` |
`60f0238b0eec4c91fe466497ae8862daf521aecc` |
| #50685 | `1be36283678a9a94fc8fdaad6c95c2896d6b4015` |
`c05d75aaa95cf89f547503044c1921625905085d` |
`f258cbbd898f2b05f38d96b20d1530d5e10f7923` |
| #50620 | `c05d75aaa95cf89f547503044c1921625905085d` |
`653cc6faca6885e36760bf35a25bb63442519b14` |
`f258cbbd898f2b05f38d96b20d1530d5e10f7923` |

The analyzer checkout supplies the tool; `--vllm-root` and
`--ascend-root` supply the pinned source trees being analyzed. They do
not need to be the analyzer checkout itself. Existing repositories can
be reused when the head/revision and required base objects match the
table; no model installation or NPU is required for this static CLI
scan.

For example, run #39568 from this PR's analyzer checkout after preparing
the two source paths at the revisions above:

```bash
VLLM_INTERFACE_TIMINGS=1 python -u -m tests.e2e.vllm_interface.vllm_interface_contracts analyze-range \
  --vllm-root /path/to/vllm-source \
  --ascend-root /path/to/vllm-ascend-source \
  --old ce29c26 \
  --new c7560af \
  --expect-ascend-sha 60f0238 \
  --index-workers 4 --analysis-workers 3 --fail-on introduced
```

PowerShell equivalent (replace the two source directory paths):

```powershell
$env:VLLM_INTERFACE_TIMINGS = "1"
python -u -m tests.e2e.vllm_interface.vllm_interface_contracts analyze-range `
  --vllm-root "C:\path\to\vllm-source" `
  --ascend-root "C:\path\to\vllm-ascend-source" `
  --old ce29c26 `
  --new c7560af `
  --expect-ascend-sha 60f0238 `
  --index-workers 4 --analysis-workers 3 --fail-on introduced
Write-Host "Analyzer exit code: $LASTEXITCODE"
```

Timing diagnostics are printed as phases finish, followed by the
compatibility summary; this is not per-file progress. #39568 reports
`BREAKS FOUND`: vLLM removes `SchedulerInterface._get_routed_experts`,
still called at `vllm_ascend/core/recompute_scheduler.py:907` in the
pinned baseline. Expected exit: 1 for #39568/#50685 (detected P1), 0 for
#50620.

These commands use the PR's CLI, not pytest: they bypass PR-range
network discovery and parent `conftest.py` dependency loading while
exercising the same analysis engine. Base objects must already exist
locally to avoid Git fetching missing objects. The SHAs are replay
inputs, not analyzer constants. Raw logs and full comparison evidence
are retained locally; the CI entry itself still prints logs without
creating report artifacts.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: shenzhao <shenzhao9@huawei.com>
Co-authored-by: shenzhao <shenzhao9@huawei.com>
Signed-off-by: tianming2009 <13246728590@163.com>
like-0517 pushed a commit to like-0517/vllm-ascend that referenced this pull request Sep 15, 2026
)

### What this PR does / why we need it?

Reduce repeated work in the CI-only vLLM PR compatibility analyzer
introduced by vllm-project#14560, and fix import-presence checks for deleted or
annotation-only names. Changes stay under `tests/e2e/vllm_interface/`.

The pytest entry, PR base/head resolution, complete source indexing,
three analysis branches, deterministic summary, and failure policy
remain in place. No persistent cache, pickle files, new environment
variables, additional CI jobs, or main2main/monkey-patch analysis are
introduced.

#### Execution flow and implementation details

1. **Resolve and verify inputs.** The existing pytest entry resolves the
vLLM PR base/head and records the vllm-ascend revision. Commit and
checkout verification remains enabled.
2. **Index source trees.** `InterfaceBoundaryGenerator` indexes the
complete source trees and finalizes the indexes. Single-path namespace
merges copy an already-normalized state instead of repeatedly rebuilding
sorted keys and alternative sets. State dictionaries remain independent;
immutable binding tuples can be shared.
3. **Resolve aliases and namespace states.**
`RepositoryIndex.canonical_name()` checks successively shorter
dot-delimited prefixes in the live alias table instead of
sorting/scanning every alias per lookup. Longest-prefix behavior and
cycle protection remain. Always-bound module names are derived from the
final namespace already computed. Decorator and direct-call resolution
reuse identical statement-prefix states through a per-module, 256-entry
LRU keyed by actual AST statements and guard names. The memo returns
independent dictionaries; expression resolution and caller-specific
fallback still run for each query. It is not shared across revisions or
CI runs.
4. **Discover dependencies.** Override and exact-call discovery retains
affected call sites and inherited implementation locations. Import
discovery reuses the already-parsed vllm-ascend ASTs, with a
source-reading fallback for files absent from the index. Triton launch
invocation kinds remain distinct.
5. **Read base/head source.** Each analysis branch retains separate
base/head `GitSnapshot` instances. Each snapshot lazily uses one `git
cat-file --batch` process rather than launching `git show` per file.
Binary response framing preserves empty files, CRLF and non-ASCII
source. Invalid or failed reads raise analysis errors, not
missing-symbol findings. An `ExitStack` closes processes and temporary
stderr streams on success or exceptions. This is not a persistent source
cache.
6. **Resolve contracts and compare usage.** Snapshots reuse module/class
namespace states, named bindings, owner lookup and resolved API
contracts. Call-contract keys include target expression, access kind,
receiver type, member and invocation kind. Argument binding and
return-use checks still execute separately at every call site; one
site's compatibility result is never reused for another.
7. **Check import presence.** Import checks resolve presence before
constructing unused signature/return details. Presence now follows the
final namespace: a later `del` removes a binding; an annotation without
a value does not create a name or erase an existing value. Constants and
re-exports that remain bound are retained. A P1 requires proven presence
at the base and proven absence at the head. Ambiguous bindings, star
imports, dynamic module exports and unparseable source are not treated
as proven removals. Package-submodule fallback is preserved. Full
endpoint details are still built for findings.
8. **Merge and print.** The three branches merge and sort findings
deterministically. The CI entry prints the summary directly; it does not
create report artifacts. Historical findings remain subject to the
existing PR-only filtering policy.

### Does this PR introduce _any_ user-facing change?

Reduced analyzer runtime and more accurate detection of imports whose
names were deleted or became annotation-only. This is therefore not
exclusively a performance-only change. Analyzer version advances to
`2.1.1`. CLI options and CI log wording are unchanged; existing
dynamic-dispatch and return-dataflow limitations remain.

### How was this patch tested?

#### Local regressions

- 123 local tests passed: Git batch-reader lifecycle/error handling,
alias equivalence, final import bindings, callable/return contracts,
serial/parallel fixture equivalence, pytest-entry integration, and
scope-prefix isolation/eviction.
- Included 12,000 synthetic alias comparisons and 196 generated
scope-flow combinations. Deleted and annotation-only imports were
exercised through actual CLI runs against temporary Git repositories.
- The pytest integration fixtures launch the analyzer subprocess; the
analyzer is not mocked. Historical replay scans described below call the
analyzer core directly, not the hardware E2E suite.
- Ruff, formatting, compileall and mypy checks passed in local
validation. Full repository formatting was previously attempted; some
hooks require Linux shell tools unavailable on this Windows host. Full
Linux CI and NPU workloads are not claimed as locally verified.
- Regression/replay harnesses remain outside the E2E collection path, as
requested; no analyzer unit-test directory is reintroduced into vLLM PR
jobs.

#### Latest paired performance measurement

Windows / Python 3.11, four indexing processes and three analysis
threads, fixed source SHAs, fresh processes, sequential scans and no
persistent analyzer cache. Times below are analyzer time, excluding
network range discovery and hardware tests.

| vLLM PR | Before latest scope reuse | After | Reduction | Result on
both sides |
| --- | ---: | ---: | ---: | --- |
| #39568 | 73.6 s | 61.9 s | 15.9% | 1 P1 |
| #50685 | 100.3 s | 82.0 s | 18.2% | 3 P1 impacts / 1 root cause |
| #50620 | 100.4 s | 81.8 s | 18.5% | PASS |

This baseline already includes the earlier Git batching, alias
optimization and import fix; it is **not** the original pre-optimization
implementation. Full reports matched after removing only timing
metadata, stdout matched byte-for-byte, and expected exit codes were
1/1/0. These are single paired measurements, not statistical benchmarks
or Linux CI speed guarantees.

#### Expanded accuracy replay

Eight additional cases were each run against the frozen pre-scope-reuse
baseline and current code: 16 actual scans. All eight full reports
matched except timing, and all eight stdout logs were byte-identical.

| Fixed historical input | P1 impacts | P1 roots | Review findings |
| --- | ---: | ---: | ---: |
| vllm-ascend vllm-project#11709 range | 5 | 5 | 6 |
| vllm-ascend vllm-project#12020 range | 2 | 1 | 0 |
| vllm-ascend vllm-project#12420 range | 0 | 0 | 0 |
| vllm-ascend vllm-project#12502 range | 33 | 20 | 0 |
| vllm-ascend vllm-project#12648 range | 0 | 0 | 5 |
| vllm-ascend vllm-project#13358 range | 5 | 4 | 11 |
| vLLM #47808 | 2 | 2 | 5 |
| vLLM #50504 | 0 | 0 | 0 |

The upgrade ranges are fixed CI-analyzer inputs, not a full main2main
mode. vllm-project#12648 uses its adapted vllm-ascend head; the other upgrade cases
use their pre-adaptation baselines. The earlier three cases above were
reused only after verifying current source hashes, giving 11 distinct
cases overall. There were not 22 new scans in the expanded round.

29 semantic/history assertions passed. Previously confirmed P1 findings
remained. #47808 matches the newer August 28 historical log (2 P1), not
the older August 17 report (1 P1). vllm-project#11709's six review items already
existed in the frozen baseline; they are not introduced by scope reuse.

Known limitations are not hidden by these results: vllm-project#12648's
`compute_slot_mappings(out=...)` remains review-only because its
dispatch requires monkey-patch/field propagation; broader return-value
consumption is still outside the exact dependency model. Regression
equivalence does not establish universal precision/recall.

#### Fresh verification of the published PR head

After pushing commit `2e31d3bcbdc82960bc0a1956abdd412f6ef1f27b`, fetched
`refs/pull/16364/head` from `vllm-project/vllm-ascend` and checked it
out in a separate, clean detached worktree. The fetched commit and all
nine analyzer source hashes matched the validated version.

Three additional fresh scans used that fetched code:

| vLLM PR | Result | Analyzer time | Process wall time | Exit code |
| --- | --- | ---: | ---: | ---: |
| #39568 | 1 P1 | 59.8 s | 61.0 s | 1 |
| #50685 | 3 P1 impacts / 1 root cause | 82.2 s | 83.7 s | 1 |
| #50620 | PASS | 79.8 s | 81.3 s | 0 |

All three complete reports were identical to the pre-push reports after
excluding only timing metadata; raw stdout was byte-identical. Exit 1 is
the expected detected-break outcome, not an analyzer crash. These runs
are separate from the earlier paired measurements and expanded
regression round.

The 123 local tests passed again before pushing. Ruff, formatting,
spelling and Markdown checks passed; several shell-based repository
hooks could not execute on Windows. The fetched code also passed mypy
with the Python 3.10/Linux target. This is static type checking, not
execution on Linux.

These historical scans execute `analyze_range()` and the production
summary renderer directly. They do not run the full upstream pytest job
or NPU workloads. The actual pytest entry is covered separately by the
local integration fixtures.

Published-head verification details:
vllm-project#16364 (comment)

#### Reproduce a pinned case

With this PR's analyzer checked out, prepare separate clean source
worktrees at the following head/revision:

| vLLM PR | vLLM base | vLLM head | vllm-ascend revision |
| --- | --- | --- | --- |
| #39568 | `ce29c26b31d432b1b4bc028c46bb2c3b07a667d8` |
`c7560af42487b1570c4e6f4cea5df1605a4d59fc` |
`60f0238b0eec4c91fe466497ae8862daf521aecc` |
| #50685 | `1be36283678a9a94fc8fdaad6c95c2896d6b4015` |
`c05d75aaa95cf89f547503044c1921625905085d` |
`f258cbbd898f2b05f38d96b20d1530d5e10f7923` |
| #50620 | `c05d75aaa95cf89f547503044c1921625905085d` |
`653cc6faca6885e36760bf35a25bb63442519b14` |
`f258cbbd898f2b05f38d96b20d1530d5e10f7923` |

The analyzer checkout supplies the tool; `--vllm-root` and
`--ascend-root` supply the pinned source trees being analyzed. They do
not need to be the analyzer checkout itself. Existing repositories can
be reused when the head/revision and required base objects match the
table; no model installation or NPU is required for this static CLI
scan.

For example, run #39568 from this PR's analyzer checkout after preparing
the two source paths at the revisions above:

```bash
VLLM_INTERFACE_TIMINGS=1 python -u -m tests.e2e.vllm_interface.vllm_interface_contracts analyze-range \
  --vllm-root /path/to/vllm-source \
  --ascend-root /path/to/vllm-ascend-source \
  --old ce29c26 \
  --new c7560af \
  --expect-ascend-sha 60f0238 \
  --index-workers 4 --analysis-workers 3 --fail-on introduced
```

PowerShell equivalent (replace the two source directory paths):

```powershell
$env:VLLM_INTERFACE_TIMINGS = "1"
python -u -m tests.e2e.vllm_interface.vllm_interface_contracts analyze-range `
  --vllm-root "C:\path\to\vllm-source" `
  --ascend-root "C:\path\to\vllm-ascend-source" `
  --old ce29c26 `
  --new c7560af `
  --expect-ascend-sha 60f0238 `
  --index-workers 4 --analysis-workers 3 --fail-on introduced
Write-Host "Analyzer exit code: $LASTEXITCODE"
```

Timing diagnostics are printed as phases finish, followed by the
compatibility summary; this is not per-file progress. #39568 reports
`BREAKS FOUND`: vLLM removes `SchedulerInterface._get_routed_experts`,
still called at `vllm_ascend/core/recompute_scheduler.py:907` in the
pinned baseline. Expected exit: 1 for #39568/#50685 (detected P1), 0 for
#50620.

These commands use the PR's CLI, not pytest: they bypass PR-range
network discovery and parent `conftest.py` dependency loading while
exercising the same analysis engine. Base objects must already exist
locally to avoid Git fetching missing objects. The SHAs are replay
inputs, not analyzer constants. Raw logs and full comparison evidence
are retained locally; the CI entry itself still prints logs without
creating report artifacts.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: shenzhao <shenzhao9@huawei.com>
Co-authored-by: shenzhao <shenzhao9@huawei.com>
Signed-off-by: like-0517 <ithwlike@126.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation module:tests

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants