Repository navigation
[CI] Reduce repeated work in vLLM interface analysis - #16364
wangxiyuan merged 2 commits into
Conversation
Signed-off-by: shenzhao <shenzhao9@huawei.com>
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request introduces performance optimizations to the vLLM PR compatibility analyzer used in CI. By implementing memoization for scopes, bindings, and API contracts, and optimizing the state merging logic, the analyzer significantly reduces redundant work without altering the analysis scope, compatibility rules, or output reports. These changes are strictly in-process and do not introduce persistent caches or external dependencies. Highlights
New Features🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console. Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
|
👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:
If CI fails, you can run linting and testing checks locally according Contributing and Testing. |
There was a problem hiding this comment.
Code Review
This pull request introduces in-process optimizations to the vLLM interface compatibility analyzer to reduce redundant computations. It implements caching for scope bindings, named bindings, endpoints, and import presence within the GitSnapshot class, and allows reusing already-parsed module trees during import discovery. The review feedback correctly identifies that the PR title and summary do not adhere to the repository's style guide and provides a compliant suggestion.
| ### Avoiding repeated work within a run | ||
|
|
||
| The scope interpreter copies an already-normalized state when there is only one execution path, instead of sorting | ||
| and merging every name again. Multiple paths still use the existing binding-alternative merge. |
There was a problem hiding this comment.
The Pull Request Title does not fully adhere to the Repository Style Guide format: [Branch][Module][Action] Pull Request Title. The [Action] tag (e.g., [Misc]) is missing.
Here is the suggested title and summary matching the style guide requirements:
Suggested PR Title:
[CI][Misc] Reduce repeated work in vLLM interface analysisSuggested PR Summary:
### What this PR does / why we need it?
Reduce repeated work in the CI-only vLLM PR compatibility analyzer introduced by #14560. This PR changes only the analyzer and its README under `tests/e2e/vllm_interface/`.
The analysis scope, compatibility rules, finding locations, and printed summary are unchanged. No persistent cache, pickle files, new environment variables, additional CI jobs, or main2main/monkey-patch analysis are introduced.
### Does this PR introduce _any_ user-facing change?
Only reduced analyzer runtime. No intended change to detected breaks, review findings, CLI options, or CI log wording. Existing detection limitations remain unchanged.
### How was this patch tested?
Tested via paired real-PR replay before submission and local regression checks.References
- The summary and title should follow the specific format defined in the Repository Style Guide. (link)
Published PR code verificationVerified analyzer commit: Both replays ran the actual
The breaking replay still identifies removal of The local regression suite also passed again against this published commit: 37 passed in 8.91 s. The suite includes an actual pytest E2E entry-point/subprocess fixture. Regression helpers remain outside the submitted CI directory. No tracked files changed during verification. The original pre-optimization and optimized benchmark reports were also equivalent after removing timing metadata; this additional published-commit replay confirms the same stdout and expected exits. This establishes equivalence for the tested cases, not proof of complete detection of every possible interface change. GitHub checksAt verification time, DCO, pre-commit, CPU unit tests and ci-gate all passed. Hardware test jobs were skipped; this is not evidence of an upstream NPU full-CI run. The local artifact bundle retains a GitHub status snapshot alongside the raw logs. Raw replay outputstdout and stderr are shown separately below, without rewriting their content. The three analysis branches run concurrently, so phase timings must not be added together. vLLM #39568 — raw stdoutvLLM #39568 — raw stderrvLLM #50620 — raw stdoutvLLM #50620 — raw stderr |
Batch Git blob reads, reuse module and prefix namespace states, and resolve aliases by qualified prefixes. Determine import presence from final bindings while preserving ambiguous exports for review. Signed-off-by: shenzhao <shenzhao9@huawei.com>
Published PR head: fresh local verificationCommit: Fetched Three fresh, sequential pinned-input scans used the fetched analyzer code (Windows / Python 3.11, 4 indexing processes, 3 analysis threads, no persistent analyzer cache).
Exit 1 is the expected detected-break outcome for #39568/#50685, not an analyzer crash. #50620 passed. These real-history scans call analyze_range and the production summary renderer directly. They do not claim to execute the full upstream pytest job or NPU workloads. Before pushing, all 123 local regression tests passed again, including pytest-entry integration fixtures. Ruff, formatting, spelling and Markdown checks passed; several repository-wide shell hooks could not run on this Windows host. Raw stdout/stderr, complete reports, commands, timings and source hashes are retained locally. No detection differences were observed in these samples; existing model limitations still apply. |
) ### What this PR does / why we need it? Reduce repeated work in the CI-only vLLM PR compatibility analyzer introduced by vllm-project#14560, and fix import-presence checks for deleted or annotation-only names. Changes stay under `tests/e2e/vllm_interface/`. The pytest entry, PR base/head resolution, complete source indexing, three analysis branches, deterministic summary, and failure policy remain in place. No persistent cache, pickle files, new environment variables, additional CI jobs, or main2main/monkey-patch analysis are introduced. #### Execution flow and implementation details 1. **Resolve and verify inputs.** The existing pytest entry resolves the vLLM PR base/head and records the vllm-ascend revision. Commit and checkout verification remains enabled. 2. **Index source trees.** `InterfaceBoundaryGenerator` indexes the complete source trees and finalizes the indexes. Single-path namespace merges copy an already-normalized state instead of repeatedly rebuilding sorted keys and alternative sets. State dictionaries remain independent; immutable binding tuples can be shared. 3. **Resolve aliases and namespace states.** `RepositoryIndex.canonical_name()` checks successively shorter dot-delimited prefixes in the live alias table instead of sorting/scanning every alias per lookup. Longest-prefix behavior and cycle protection remain. Always-bound module names are derived from the final namespace already computed. Decorator and direct-call resolution reuse identical statement-prefix states through a per-module, 256-entry LRU keyed by actual AST statements and guard names. The memo returns independent dictionaries; expression resolution and caller-specific fallback still run for each query. It is not shared across revisions or CI runs. 4. **Discover dependencies.** Override and exact-call discovery retains affected call sites and inherited implementation locations. Import discovery reuses the already-parsed vllm-ascend ASTs, with a source-reading fallback for files absent from the index. Triton launch invocation kinds remain distinct. 5. **Read base/head source.** Each analysis branch retains separate base/head `GitSnapshot` instances. Each snapshot lazily uses one `git cat-file --batch` process rather than launching `git show` per file. Binary response framing preserves empty files, CRLF and non-ASCII source. Invalid or failed reads raise analysis errors, not missing-symbol findings. An `ExitStack` closes processes and temporary stderr streams on success or exceptions. This is not a persistent source cache. 6. **Resolve contracts and compare usage.** Snapshots reuse module/class namespace states, named bindings, owner lookup and resolved API contracts. Call-contract keys include target expression, access kind, receiver type, member and invocation kind. Argument binding and return-use checks still execute separately at every call site; one site's compatibility result is never reused for another. 7. **Check import presence.** Import checks resolve presence before constructing unused signature/return details. Presence now follows the final namespace: a later `del` removes a binding; an annotation without a value does not create a name or erase an existing value. Constants and re-exports that remain bound are retained. A P1 requires proven presence at the base and proven absence at the head. Ambiguous bindings, star imports, dynamic module exports and unparseable source are not treated as proven removals. Package-submodule fallback is preserved. Full endpoint details are still built for findings. 8. **Merge and print.** The three branches merge and sort findings deterministically. The CI entry prints the summary directly; it does not create report artifacts. Historical findings remain subject to the existing PR-only filtering policy. ### Does this PR introduce _any_ user-facing change? Reduced analyzer runtime and more accurate detection of imports whose names were deleted or became annotation-only. This is therefore not exclusively a performance-only change. Analyzer version advances to `2.1.1`. CLI options and CI log wording are unchanged; existing dynamic-dispatch and return-dataflow limitations remain. ### How was this patch tested? #### Local regressions - 123 local tests passed: Git batch-reader lifecycle/error handling, alias equivalence, final import bindings, callable/return contracts, serial/parallel fixture equivalence, pytest-entry integration, and scope-prefix isolation/eviction. - Included 12,000 synthetic alias comparisons and 196 generated scope-flow combinations. Deleted and annotation-only imports were exercised through actual CLI runs against temporary Git repositories. - The pytest integration fixtures launch the analyzer subprocess; the analyzer is not mocked. Historical replay scans described below call the analyzer core directly, not the hardware E2E suite. - Ruff, formatting, compileall and mypy checks passed in local validation. Full repository formatting was previously attempted; some hooks require Linux shell tools unavailable on this Windows host. Full Linux CI and NPU workloads are not claimed as locally verified. - Regression/replay harnesses remain outside the E2E collection path, as requested; no analyzer unit-test directory is reintroduced into vLLM PR jobs. #### Latest paired performance measurement Windows / Python 3.11, four indexing processes and three analysis threads, fixed source SHAs, fresh processes, sequential scans and no persistent analyzer cache. Times below are analyzer time, excluding network range discovery and hardware tests. | vLLM PR | Before latest scope reuse | After | Reduction | Result on both sides | | --- | ---: | ---: | ---: | --- | | #39568 | 73.6 s | 61.9 s | 15.9% | 1 P1 | | #50685 | 100.3 s | 82.0 s | 18.2% | 3 P1 impacts / 1 root cause | | #50620 | 100.4 s | 81.8 s | 18.5% | PASS | This baseline already includes the earlier Git batching, alias optimization and import fix; it is **not** the original pre-optimization implementation. Full reports matched after removing only timing metadata, stdout matched byte-for-byte, and expected exit codes were 1/1/0. These are single paired measurements, not statistical benchmarks or Linux CI speed guarantees. #### Expanded accuracy replay Eight additional cases were each run against the frozen pre-scope-reuse baseline and current code: 16 actual scans. All eight full reports matched except timing, and all eight stdout logs were byte-identical. | Fixed historical input | P1 impacts | P1 roots | Review findings | | --- | ---: | ---: | ---: | | vllm-ascend vllm-project#11709 range | 5 | 5 | 6 | | vllm-ascend vllm-project#12020 range | 2 | 1 | 0 | | vllm-ascend vllm-project#12420 range | 0 | 0 | 0 | | vllm-ascend vllm-project#12502 range | 33 | 20 | 0 | | vllm-ascend vllm-project#12648 range | 0 | 0 | 5 | | vllm-ascend vllm-project#13358 range | 5 | 4 | 11 | | vLLM #47808 | 2 | 2 | 5 | | vLLM #50504 | 0 | 0 | 0 | The upgrade ranges are fixed CI-analyzer inputs, not a full main2main mode. vllm-project#12648 uses its adapted vllm-ascend head; the other upgrade cases use their pre-adaptation baselines. The earlier three cases above were reused only after verifying current source hashes, giving 11 distinct cases overall. There were not 22 new scans in the expanded round. 29 semantic/history assertions passed. Previously confirmed P1 findings remained. #47808 matches the newer August 28 historical log (2 P1), not the older August 17 report (1 P1). vllm-project#11709's six review items already existed in the frozen baseline; they are not introduced by scope reuse. Known limitations are not hidden by these results: vllm-project#12648's `compute_slot_mappings(out=...)` remains review-only because its dispatch requires monkey-patch/field propagation; broader return-value consumption is still outside the exact dependency model. Regression equivalence does not establish universal precision/recall. #### Fresh verification of the published PR head After pushing commit `2e31d3bcbdc82960bc0a1956abdd412f6ef1f27b`, fetched `refs/pull/16364/head` from `vllm-project/vllm-ascend` and checked it out in a separate, clean detached worktree. The fetched commit and all nine analyzer source hashes matched the validated version. Three additional fresh scans used that fetched code: | vLLM PR | Result | Analyzer time | Process wall time | Exit code | | --- | --- | ---: | ---: | ---: | | #39568 | 1 P1 | 59.8 s | 61.0 s | 1 | | #50685 | 3 P1 impacts / 1 root cause | 82.2 s | 83.7 s | 1 | | #50620 | PASS | 79.8 s | 81.3 s | 0 | All three complete reports were identical to the pre-push reports after excluding only timing metadata; raw stdout was byte-identical. Exit 1 is the expected detected-break outcome, not an analyzer crash. These runs are separate from the earlier paired measurements and expanded regression round. The 123 local tests passed again before pushing. Ruff, formatting, spelling and Markdown checks passed; several shell-based repository hooks could not execute on Windows. The fetched code also passed mypy with the Python 3.10/Linux target. This is static type checking, not execution on Linux. These historical scans execute `analyze_range()` and the production summary renderer directly. They do not run the full upstream pytest job or NPU workloads. The actual pytest entry is covered separately by the local integration fixtures. Published-head verification details: vllm-project#16364 (comment) #### Reproduce a pinned case With this PR's analyzer checked out, prepare separate clean source worktrees at the following head/revision: | vLLM PR | vLLM base | vLLM head | vllm-ascend revision | | --- | --- | --- | --- | | #39568 | `ce29c26b31d432b1b4bc028c46bb2c3b07a667d8` | `c7560af42487b1570c4e6f4cea5df1605a4d59fc` | `60f0238b0eec4c91fe466497ae8862daf521aecc` | | #50685 | `1be36283678a9a94fc8fdaad6c95c2896d6b4015` | `c05d75aaa95cf89f547503044c1921625905085d` | `f258cbbd898f2b05f38d96b20d1530d5e10f7923` | | #50620 | `c05d75aaa95cf89f547503044c1921625905085d` | `653cc6faca6885e36760bf35a25bb63442519b14` | `f258cbbd898f2b05f38d96b20d1530d5e10f7923` | The analyzer checkout supplies the tool; `--vllm-root` and `--ascend-root` supply the pinned source trees being analyzed. They do not need to be the analyzer checkout itself. Existing repositories can be reused when the head/revision and required base objects match the table; no model installation or NPU is required for this static CLI scan. For example, run #39568 from this PR's analyzer checkout after preparing the two source paths at the revisions above: ```bash VLLM_INTERFACE_TIMINGS=1 python -u -m tests.e2e.vllm_interface.vllm_interface_contracts analyze-range \ --vllm-root /path/to/vllm-source \ --ascend-root /path/to/vllm-ascend-source \ --old ce29c26 \ --new c7560af \ --expect-ascend-sha 60f0238 \ --index-workers 4 --analysis-workers 3 --fail-on introduced ``` PowerShell equivalent (replace the two source directory paths): ```powershell $env:VLLM_INTERFACE_TIMINGS = "1" python -u -m tests.e2e.vllm_interface.vllm_interface_contracts analyze-range ` --vllm-root "C:\path\to\vllm-source" ` --ascend-root "C:\path\to\vllm-ascend-source" ` --old ce29c26 ` --new c7560af ` --expect-ascend-sha 60f0238 ` --index-workers 4 --analysis-workers 3 --fail-on introduced Write-Host "Analyzer exit code: $LASTEXITCODE" ``` Timing diagnostics are printed as phases finish, followed by the compatibility summary; this is not per-file progress. #39568 reports `BREAKS FOUND`: vLLM removes `SchedulerInterface._get_routed_experts`, still called at `vllm_ascend/core/recompute_scheduler.py:907` in the pinned baseline. Expected exit: 1 for #39568/#50685 (detected P1), 0 for #50620. These commands use the PR's CLI, not pytest: they bypass PR-range network discovery and parent `conftest.py` dependency loading while exercising the same analysis engine. Base objects must already exist locally to avoid Git fetching missing objects. The SHAs are replay inputs, not analyzer constants. Raw logs and full comparison evidence are retained locally; the CI entry itself still prints logs without creating report artifacts. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: shenzhao <shenzhao9@huawei.com> Co-authored-by: shenzhao <shenzhao9@huawei.com>
) ### What this PR does / why we need it? Reduce repeated work in the CI-only vLLM PR compatibility analyzer introduced by vllm-project#14560, and fix import-presence checks for deleted or annotation-only names. Changes stay under `tests/e2e/vllm_interface/`. The pytest entry, PR base/head resolution, complete source indexing, three analysis branches, deterministic summary, and failure policy remain in place. No persistent cache, pickle files, new environment variables, additional CI jobs, or main2main/monkey-patch analysis are introduced. #### Execution flow and implementation details 1. **Resolve and verify inputs.** The existing pytest entry resolves the vLLM PR base/head and records the vllm-ascend revision. Commit and checkout verification remains enabled. 2. **Index source trees.** `InterfaceBoundaryGenerator` indexes the complete source trees and finalizes the indexes. Single-path namespace merges copy an already-normalized state instead of repeatedly rebuilding sorted keys and alternative sets. State dictionaries remain independent; immutable binding tuples can be shared. 3. **Resolve aliases and namespace states.** `RepositoryIndex.canonical_name()` checks successively shorter dot-delimited prefixes in the live alias table instead of sorting/scanning every alias per lookup. Longest-prefix behavior and cycle protection remain. Always-bound module names are derived from the final namespace already computed. Decorator and direct-call resolution reuse identical statement-prefix states through a per-module, 256-entry LRU keyed by actual AST statements and guard names. The memo returns independent dictionaries; expression resolution and caller-specific fallback still run for each query. It is not shared across revisions or CI runs. 4. **Discover dependencies.** Override and exact-call discovery retains affected call sites and inherited implementation locations. Import discovery reuses the already-parsed vllm-ascend ASTs, with a source-reading fallback for files absent from the index. Triton launch invocation kinds remain distinct. 5. **Read base/head source.** Each analysis branch retains separate base/head `GitSnapshot` instances. Each snapshot lazily uses one `git cat-file --batch` process rather than launching `git show` per file. Binary response framing preserves empty files, CRLF and non-ASCII source. Invalid or failed reads raise analysis errors, not missing-symbol findings. An `ExitStack` closes processes and temporary stderr streams on success or exceptions. This is not a persistent source cache. 6. **Resolve contracts and compare usage.** Snapshots reuse module/class namespace states, named bindings, owner lookup and resolved API contracts. Call-contract keys include target expression, access kind, receiver type, member and invocation kind. Argument binding and return-use checks still execute separately at every call site; one site's compatibility result is never reused for another. 7. **Check import presence.** Import checks resolve presence before constructing unused signature/return details. Presence now follows the final namespace: a later `del` removes a binding; an annotation without a value does not create a name or erase an existing value. Constants and re-exports that remain bound are retained. A P1 requires proven presence at the base and proven absence at the head. Ambiguous bindings, star imports, dynamic module exports and unparseable source are not treated as proven removals. Package-submodule fallback is preserved. Full endpoint details are still built for findings. 8. **Merge and print.** The three branches merge and sort findings deterministically. The CI entry prints the summary directly; it does not create report artifacts. Historical findings remain subject to the existing PR-only filtering policy. ### Does this PR introduce _any_ user-facing change? Reduced analyzer runtime and more accurate detection of imports whose names were deleted or became annotation-only. This is therefore not exclusively a performance-only change. Analyzer version advances to `2.1.1`. CLI options and CI log wording are unchanged; existing dynamic-dispatch and return-dataflow limitations remain. ### How was this patch tested? #### Local regressions - 123 local tests passed: Git batch-reader lifecycle/error handling, alias equivalence, final import bindings, callable/return contracts, serial/parallel fixture equivalence, pytest-entry integration, and scope-prefix isolation/eviction. - Included 12,000 synthetic alias comparisons and 196 generated scope-flow combinations. Deleted and annotation-only imports were exercised through actual CLI runs against temporary Git repositories. - The pytest integration fixtures launch the analyzer subprocess; the analyzer is not mocked. Historical replay scans described below call the analyzer core directly, not the hardware E2E suite. - Ruff, formatting, compileall and mypy checks passed in local validation. Full repository formatting was previously attempted; some hooks require Linux shell tools unavailable on this Windows host. Full Linux CI and NPU workloads are not claimed as locally verified. - Regression/replay harnesses remain outside the E2E collection path, as requested; no analyzer unit-test directory is reintroduced into vLLM PR jobs. #### Latest paired performance measurement Windows / Python 3.11, four indexing processes and three analysis threads, fixed source SHAs, fresh processes, sequential scans and no persistent analyzer cache. Times below are analyzer time, excluding network range discovery and hardware tests. | vLLM PR | Before latest scope reuse | After | Reduction | Result on both sides | | --- | ---: | ---: | ---: | --- | | #39568 | 73.6 s | 61.9 s | 15.9% | 1 P1 | | #50685 | 100.3 s | 82.0 s | 18.2% | 3 P1 impacts / 1 root cause | | #50620 | 100.4 s | 81.8 s | 18.5% | PASS | This baseline already includes the earlier Git batching, alias optimization and import fix; it is **not** the original pre-optimization implementation. Full reports matched after removing only timing metadata, stdout matched byte-for-byte, and expected exit codes were 1/1/0. These are single paired measurements, not statistical benchmarks or Linux CI speed guarantees. #### Expanded accuracy replay Eight additional cases were each run against the frozen pre-scope-reuse baseline and current code: 16 actual scans. All eight full reports matched except timing, and all eight stdout logs were byte-identical. | Fixed historical input | P1 impacts | P1 roots | Review findings | | --- | ---: | ---: | ---: | | vllm-ascend vllm-project#11709 range | 5 | 5 | 6 | | vllm-ascend vllm-project#12020 range | 2 | 1 | 0 | | vllm-ascend vllm-project#12420 range | 0 | 0 | 0 | | vllm-ascend vllm-project#12502 range | 33 | 20 | 0 | | vllm-ascend vllm-project#12648 range | 0 | 0 | 5 | | vllm-ascend vllm-project#13358 range | 5 | 4 | 11 | | vLLM #47808 | 2 | 2 | 5 | | vLLM #50504 | 0 | 0 | 0 | The upgrade ranges are fixed CI-analyzer inputs, not a full main2main mode. vllm-project#12648 uses its adapted vllm-ascend head; the other upgrade cases use their pre-adaptation baselines. The earlier three cases above were reused only after verifying current source hashes, giving 11 distinct cases overall. There were not 22 new scans in the expanded round. 29 semantic/history assertions passed. Previously confirmed P1 findings remained. #47808 matches the newer August 28 historical log (2 P1), not the older August 17 report (1 P1). vllm-project#11709's six review items already existed in the frozen baseline; they are not introduced by scope reuse. Known limitations are not hidden by these results: vllm-project#12648's `compute_slot_mappings(out=...)` remains review-only because its dispatch requires monkey-patch/field propagation; broader return-value consumption is still outside the exact dependency model. Regression equivalence does not establish universal precision/recall. #### Fresh verification of the published PR head After pushing commit `2e31d3bcbdc82960bc0a1956abdd412f6ef1f27b`, fetched `refs/pull/16364/head` from `vllm-project/vllm-ascend` and checked it out in a separate, clean detached worktree. The fetched commit and all nine analyzer source hashes matched the validated version. Three additional fresh scans used that fetched code: | vLLM PR | Result | Analyzer time | Process wall time | Exit code | | --- | --- | ---: | ---: | ---: | | #39568 | 1 P1 | 59.8 s | 61.0 s | 1 | | #50685 | 3 P1 impacts / 1 root cause | 82.2 s | 83.7 s | 1 | | #50620 | PASS | 79.8 s | 81.3 s | 0 | All three complete reports were identical to the pre-push reports after excluding only timing metadata; raw stdout was byte-identical. Exit 1 is the expected detected-break outcome, not an analyzer crash. These runs are separate from the earlier paired measurements and expanded regression round. The 123 local tests passed again before pushing. Ruff, formatting, spelling and Markdown checks passed; several shell-based repository hooks could not execute on Windows. The fetched code also passed mypy with the Python 3.10/Linux target. This is static type checking, not execution on Linux. These historical scans execute `analyze_range()` and the production summary renderer directly. They do not run the full upstream pytest job or NPU workloads. The actual pytest entry is covered separately by the local integration fixtures. Published-head verification details: vllm-project#16364 (comment) #### Reproduce a pinned case With this PR's analyzer checked out, prepare separate clean source worktrees at the following head/revision: | vLLM PR | vLLM base | vLLM head | vllm-ascend revision | | --- | --- | --- | --- | | #39568 | `ce29c26b31d432b1b4bc028c46bb2c3b07a667d8` | `c7560af42487b1570c4e6f4cea5df1605a4d59fc` | `60f0238b0eec4c91fe466497ae8862daf521aecc` | | #50685 | `1be36283678a9a94fc8fdaad6c95c2896d6b4015` | `c05d75aaa95cf89f547503044c1921625905085d` | `f258cbbd898f2b05f38d96b20d1530d5e10f7923` | | #50620 | `c05d75aaa95cf89f547503044c1921625905085d` | `653cc6faca6885e36760bf35a25bb63442519b14` | `f258cbbd898f2b05f38d96b20d1530d5e10f7923` | The analyzer checkout supplies the tool; `--vllm-root` and `--ascend-root` supply the pinned source trees being analyzed. They do not need to be the analyzer checkout itself. Existing repositories can be reused when the head/revision and required base objects match the table; no model installation or NPU is required for this static CLI scan. For example, run #39568 from this PR's analyzer checkout after preparing the two source paths at the revisions above: ```bash VLLM_INTERFACE_TIMINGS=1 python -u -m tests.e2e.vllm_interface.vllm_interface_contracts analyze-range \ --vllm-root /path/to/vllm-source \ --ascend-root /path/to/vllm-ascend-source \ --old ce29c26 \ --new c7560af \ --expect-ascend-sha 60f0238 \ --index-workers 4 --analysis-workers 3 --fail-on introduced ``` PowerShell equivalent (replace the two source directory paths): ```powershell $env:VLLM_INTERFACE_TIMINGS = "1" python -u -m tests.e2e.vllm_interface.vllm_interface_contracts analyze-range ` --vllm-root "C:\path\to\vllm-source" ` --ascend-root "C:\path\to\vllm-ascend-source" ` --old ce29c26 ` --new c7560af ` --expect-ascend-sha 60f0238 ` --index-workers 4 --analysis-workers 3 --fail-on introduced Write-Host "Analyzer exit code: $LASTEXITCODE" ``` Timing diagnostics are printed as phases finish, followed by the compatibility summary; this is not per-file progress. #39568 reports `BREAKS FOUND`: vLLM removes `SchedulerInterface._get_routed_experts`, still called at `vllm_ascend/core/recompute_scheduler.py:907` in the pinned baseline. Expected exit: 1 for #39568/#50685 (detected P1), 0 for #50620. These commands use the PR's CLI, not pytest: they bypass PR-range network discovery and parent `conftest.py` dependency loading while exercising the same analysis engine. Base objects must already exist locally to avoid Git fetching missing objects. The SHAs are replay inputs, not analyzer constants. Raw logs and full comparison evidence are retained locally; the CI entry itself still prints logs without creating report artifacts. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: shenzhao <shenzhao9@huawei.com> Co-authored-by: shenzhao <shenzhao9@huawei.com> Signed-off-by: tianming2009 <13246728590@163.com>
) ### What this PR does / why we need it? Reduce repeated work in the CI-only vLLM PR compatibility analyzer introduced by vllm-project#14560, and fix import-presence checks for deleted or annotation-only names. Changes stay under `tests/e2e/vllm_interface/`. The pytest entry, PR base/head resolution, complete source indexing, three analysis branches, deterministic summary, and failure policy remain in place. No persistent cache, pickle files, new environment variables, additional CI jobs, or main2main/monkey-patch analysis are introduced. #### Execution flow and implementation details 1. **Resolve and verify inputs.** The existing pytest entry resolves the vLLM PR base/head and records the vllm-ascend revision. Commit and checkout verification remains enabled. 2. **Index source trees.** `InterfaceBoundaryGenerator` indexes the complete source trees and finalizes the indexes. Single-path namespace merges copy an already-normalized state instead of repeatedly rebuilding sorted keys and alternative sets. State dictionaries remain independent; immutable binding tuples can be shared. 3. **Resolve aliases and namespace states.** `RepositoryIndex.canonical_name()` checks successively shorter dot-delimited prefixes in the live alias table instead of sorting/scanning every alias per lookup. Longest-prefix behavior and cycle protection remain. Always-bound module names are derived from the final namespace already computed. Decorator and direct-call resolution reuse identical statement-prefix states through a per-module, 256-entry LRU keyed by actual AST statements and guard names. The memo returns independent dictionaries; expression resolution and caller-specific fallback still run for each query. It is not shared across revisions or CI runs. 4. **Discover dependencies.** Override and exact-call discovery retains affected call sites and inherited implementation locations. Import discovery reuses the already-parsed vllm-ascend ASTs, with a source-reading fallback for files absent from the index. Triton launch invocation kinds remain distinct. 5. **Read base/head source.** Each analysis branch retains separate base/head `GitSnapshot` instances. Each snapshot lazily uses one `git cat-file --batch` process rather than launching `git show` per file. Binary response framing preserves empty files, CRLF and non-ASCII source. Invalid or failed reads raise analysis errors, not missing-symbol findings. An `ExitStack` closes processes and temporary stderr streams on success or exceptions. This is not a persistent source cache. 6. **Resolve contracts and compare usage.** Snapshots reuse module/class namespace states, named bindings, owner lookup and resolved API contracts. Call-contract keys include target expression, access kind, receiver type, member and invocation kind. Argument binding and return-use checks still execute separately at every call site; one site's compatibility result is never reused for another. 7. **Check import presence.** Import checks resolve presence before constructing unused signature/return details. Presence now follows the final namespace: a later `del` removes a binding; an annotation without a value does not create a name or erase an existing value. Constants and re-exports that remain bound are retained. A P1 requires proven presence at the base and proven absence at the head. Ambiguous bindings, star imports, dynamic module exports and unparseable source are not treated as proven removals. Package-submodule fallback is preserved. Full endpoint details are still built for findings. 8. **Merge and print.** The three branches merge and sort findings deterministically. The CI entry prints the summary directly; it does not create report artifacts. Historical findings remain subject to the existing PR-only filtering policy. ### Does this PR introduce _any_ user-facing change? Reduced analyzer runtime and more accurate detection of imports whose names were deleted or became annotation-only. This is therefore not exclusively a performance-only change. Analyzer version advances to `2.1.1`. CLI options and CI log wording are unchanged; existing dynamic-dispatch and return-dataflow limitations remain. ### How was this patch tested? #### Local regressions - 123 local tests passed: Git batch-reader lifecycle/error handling, alias equivalence, final import bindings, callable/return contracts, serial/parallel fixture equivalence, pytest-entry integration, and scope-prefix isolation/eviction. - Included 12,000 synthetic alias comparisons and 196 generated scope-flow combinations. Deleted and annotation-only imports were exercised through actual CLI runs against temporary Git repositories. - The pytest integration fixtures launch the analyzer subprocess; the analyzer is not mocked. Historical replay scans described below call the analyzer core directly, not the hardware E2E suite. - Ruff, formatting, compileall and mypy checks passed in local validation. Full repository formatting was previously attempted; some hooks require Linux shell tools unavailable on this Windows host. Full Linux CI and NPU workloads are not claimed as locally verified. - Regression/replay harnesses remain outside the E2E collection path, as requested; no analyzer unit-test directory is reintroduced into vLLM PR jobs. #### Latest paired performance measurement Windows / Python 3.11, four indexing processes and three analysis threads, fixed source SHAs, fresh processes, sequential scans and no persistent analyzer cache. Times below are analyzer time, excluding network range discovery and hardware tests. | vLLM PR | Before latest scope reuse | After | Reduction | Result on both sides | | --- | ---: | ---: | ---: | --- | | #39568 | 73.6 s | 61.9 s | 15.9% | 1 P1 | | #50685 | 100.3 s | 82.0 s | 18.2% | 3 P1 impacts / 1 root cause | | #50620 | 100.4 s | 81.8 s | 18.5% | PASS | This baseline already includes the earlier Git batching, alias optimization and import fix; it is **not** the original pre-optimization implementation. Full reports matched after removing only timing metadata, stdout matched byte-for-byte, and expected exit codes were 1/1/0. These are single paired measurements, not statistical benchmarks or Linux CI speed guarantees. #### Expanded accuracy replay Eight additional cases were each run against the frozen pre-scope-reuse baseline and current code: 16 actual scans. All eight full reports matched except timing, and all eight stdout logs were byte-identical. | Fixed historical input | P1 impacts | P1 roots | Review findings | | --- | ---: | ---: | ---: | | vllm-ascend vllm-project#11709 range | 5 | 5 | 6 | | vllm-ascend vllm-project#12020 range | 2 | 1 | 0 | | vllm-ascend vllm-project#12420 range | 0 | 0 | 0 | | vllm-ascend vllm-project#12502 range | 33 | 20 | 0 | | vllm-ascend vllm-project#12648 range | 0 | 0 | 5 | | vllm-ascend vllm-project#13358 range | 5 | 4 | 11 | | vLLM #47808 | 2 | 2 | 5 | | vLLM #50504 | 0 | 0 | 0 | The upgrade ranges are fixed CI-analyzer inputs, not a full main2main mode. vllm-project#12648 uses its adapted vllm-ascend head; the other upgrade cases use their pre-adaptation baselines. The earlier three cases above were reused only after verifying current source hashes, giving 11 distinct cases overall. There were not 22 new scans in the expanded round. 29 semantic/history assertions passed. Previously confirmed P1 findings remained. #47808 matches the newer August 28 historical log (2 P1), not the older August 17 report (1 P1). vllm-project#11709's six review items already existed in the frozen baseline; they are not introduced by scope reuse. Known limitations are not hidden by these results: vllm-project#12648's `compute_slot_mappings(out=...)` remains review-only because its dispatch requires monkey-patch/field propagation; broader return-value consumption is still outside the exact dependency model. Regression equivalence does not establish universal precision/recall. #### Fresh verification of the published PR head After pushing commit `2e31d3bcbdc82960bc0a1956abdd412f6ef1f27b`, fetched `refs/pull/16364/head` from `vllm-project/vllm-ascend` and checked it out in a separate, clean detached worktree. The fetched commit and all nine analyzer source hashes matched the validated version. Three additional fresh scans used that fetched code: | vLLM PR | Result | Analyzer time | Process wall time | Exit code | | --- | --- | ---: | ---: | ---: | | #39568 | 1 P1 | 59.8 s | 61.0 s | 1 | | #50685 | 3 P1 impacts / 1 root cause | 82.2 s | 83.7 s | 1 | | #50620 | PASS | 79.8 s | 81.3 s | 0 | All three complete reports were identical to the pre-push reports after excluding only timing metadata; raw stdout was byte-identical. Exit 1 is the expected detected-break outcome, not an analyzer crash. These runs are separate from the earlier paired measurements and expanded regression round. The 123 local tests passed again before pushing. Ruff, formatting, spelling and Markdown checks passed; several shell-based repository hooks could not execute on Windows. The fetched code also passed mypy with the Python 3.10/Linux target. This is static type checking, not execution on Linux. These historical scans execute `analyze_range()` and the production summary renderer directly. They do not run the full upstream pytest job or NPU workloads. The actual pytest entry is covered separately by the local integration fixtures. Published-head verification details: vllm-project#16364 (comment) #### Reproduce a pinned case With this PR's analyzer checked out, prepare separate clean source worktrees at the following head/revision: | vLLM PR | vLLM base | vLLM head | vllm-ascend revision | | --- | --- | --- | --- | | #39568 | `ce29c26b31d432b1b4bc028c46bb2c3b07a667d8` | `c7560af42487b1570c4e6f4cea5df1605a4d59fc` | `60f0238b0eec4c91fe466497ae8862daf521aecc` | | #50685 | `1be36283678a9a94fc8fdaad6c95c2896d6b4015` | `c05d75aaa95cf89f547503044c1921625905085d` | `f258cbbd898f2b05f38d96b20d1530d5e10f7923` | | #50620 | `c05d75aaa95cf89f547503044c1921625905085d` | `653cc6faca6885e36760bf35a25bb63442519b14` | `f258cbbd898f2b05f38d96b20d1530d5e10f7923` | The analyzer checkout supplies the tool; `--vllm-root` and `--ascend-root` supply the pinned source trees being analyzed. They do not need to be the analyzer checkout itself. Existing repositories can be reused when the head/revision and required base objects match the table; no model installation or NPU is required for this static CLI scan. For example, run #39568 from this PR's analyzer checkout after preparing the two source paths at the revisions above: ```bash VLLM_INTERFACE_TIMINGS=1 python -u -m tests.e2e.vllm_interface.vllm_interface_contracts analyze-range \ --vllm-root /path/to/vllm-source \ --ascend-root /path/to/vllm-ascend-source \ --old ce29c26 \ --new c7560af \ --expect-ascend-sha 60f0238 \ --index-workers 4 --analysis-workers 3 --fail-on introduced ``` PowerShell equivalent (replace the two source directory paths): ```powershell $env:VLLM_INTERFACE_TIMINGS = "1" python -u -m tests.e2e.vllm_interface.vllm_interface_contracts analyze-range ` --vllm-root "C:\path\to\vllm-source" ` --ascend-root "C:\path\to\vllm-ascend-source" ` --old ce29c26 ` --new c7560af ` --expect-ascend-sha 60f0238 ` --index-workers 4 --analysis-workers 3 --fail-on introduced Write-Host "Analyzer exit code: $LASTEXITCODE" ``` Timing diagnostics are printed as phases finish, followed by the compatibility summary; this is not per-file progress. #39568 reports `BREAKS FOUND`: vLLM removes `SchedulerInterface._get_routed_experts`, still called at `vllm_ascend/core/recompute_scheduler.py:907` in the pinned baseline. Expected exit: 1 for #39568/#50685 (detected P1), 0 for #50620. These commands use the PR's CLI, not pytest: they bypass PR-range network discovery and parent `conftest.py` dependency loading while exercising the same analysis engine. Base objects must already exist locally to avoid Git fetching missing objects. The SHAs are replay inputs, not analyzer constants. Raw logs and full comparison evidence are retained locally; the CI entry itself still prints logs without creating report artifacts. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: shenzhao <shenzhao9@huawei.com> Co-authored-by: shenzhao <shenzhao9@huawei.com> Signed-off-by: like-0517 <ithwlike@126.com>
What this PR does / why we need it?
Reduce repeated work in the CI-only vLLM PR compatibility analyzer introduced by #14560, and fix import-presence checks for deleted or annotation-only names. Changes stay under
tests/e2e/vllm_interface/.The pytest entry, PR base/head resolution, complete source indexing, three analysis branches, deterministic summary, and failure policy remain in place. No persistent cache, pickle files, new environment variables, additional CI jobs, or main2main/monkey-patch analysis are introduced.
Execution flow and implementation details
InterfaceBoundaryGeneratorindexes the complete source trees and finalizes the indexes. Single-path namespace merges copy an already-normalized state instead of repeatedly rebuilding sorted keys and alternative sets. State dictionaries remain independent; immutable binding tuples can be shared.RepositoryIndex.canonical_name()checks successively shorter dot-delimited prefixes in the live alias table instead of sorting/scanning every alias per lookup. Longest-prefix behavior and cycle protection remain. Always-bound module names are derived from the final namespace already computed. Decorator and direct-call resolution reuse identical statement-prefix states through a per-module, 256-entry LRU keyed by actual AST statements and guard names. The memo returns independent dictionaries; expression resolution and caller-specific fallback still run for each query. It is not shared across revisions or CI runs.GitSnapshotinstances. Each snapshot lazily uses onegit cat-file --batchprocess rather than launchinggit showper file. Binary response framing preserves empty files, CRLF and non-ASCII source. Invalid or failed reads raise analysis errors, not missing-symbol findings. AnExitStackcloses processes and temporary stderr streams on success or exceptions. This is not a persistent source cache.delremoves a binding; an annotation without a value does not create a name or erase an existing value. Constants and re-exports that remain bound are retained. A P1 requires proven presence at the base and proven absence at the head. Ambiguous bindings, star imports, dynamic module exports and unparseable source are not treated as proven removals. Package-submodule fallback is preserved. Full endpoint details are still built for findings.Does this PR introduce any user-facing change?
Reduced analyzer runtime and more accurate detection of imports whose names were deleted or became annotation-only. This is therefore not exclusively a performance-only change. Analyzer version advances to
2.1.1. CLI options and CI log wording are unchanged; existing dynamic-dispatch and return-dataflow limitations remain.How was this patch tested?
Local regressions
Latest paired performance measurement
Windows / Python 3.11, four indexing processes and three analysis threads, fixed source SHAs, fresh processes, sequential scans and no persistent analyzer cache. Times below are analyzer time, excluding network range discovery and hardware tests.
This baseline already includes the earlier Git batching, alias optimization and import fix; it is not the original pre-optimization implementation. Full reports matched after removing only timing metadata, stdout matched byte-for-byte, and expected exit codes were 1/1/0. These are single paired measurements, not statistical benchmarks or Linux CI speed guarantees.
Expanded accuracy replay
Eight additional cases were each run against the frozen pre-scope-reuse baseline and current code: 16 actual scans. All eight full reports matched except timing, and all eight stdout logs were byte-identical.
The upgrade ranges are fixed CI-analyzer inputs, not a full main2main mode. #12648 uses its adapted vllm-ascend head; the other upgrade cases use their pre-adaptation baselines. The earlier three cases above were reused only after verifying current source hashes, giving 11 distinct cases overall. There were not 22 new scans in the expanded round.
29 semantic/history assertions passed. Previously confirmed P1 findings remained. #47808 matches the newer August 28 historical log (2 P1), not the older August 17 report (1 P1). #11709's six review items already existed in the frozen baseline; they are not introduced by scope reuse.
Known limitations are not hidden by these results: #12648's
compute_slot_mappings(out=...)remains review-only because its dispatch requires monkey-patch/field propagation; broader return-value consumption is still outside the exact dependency model. Regression equivalence does not establish universal precision/recall.Fresh verification of the published PR head
After pushing commit
2e31d3bcbdc82960bc0a1956abdd412f6ef1f27b, fetchedrefs/pull/16364/headfromvllm-project/vllm-ascendand checked it out in a separate, clean detached worktree. The fetched commit and all nine analyzer source hashes matched the validated version.Three additional fresh scans used that fetched code:
All three complete reports were identical to the pre-push reports after excluding only timing metadata; raw stdout was byte-identical. Exit 1 is the expected detected-break outcome, not an analyzer crash. These runs are separate from the earlier paired measurements and expanded regression round.
The 123 local tests passed again before pushing. Ruff, formatting, spelling and Markdown checks passed; several shell-based repository hooks could not execute on Windows. The fetched code also passed mypy with the Python 3.10/Linux target. This is static type checking, not execution on Linux.
These historical scans execute
analyze_range()and the production summary renderer directly. They do not run the full upstream pytest job or NPU workloads. The actual pytest entry is covered separately by the local integration fixtures.Published-head verification details: #16364 (comment)
Reproduce a pinned case
With this PR's analyzer checked out, prepare separate clean source worktrees at the following head/revision:
ce29c26b31d432b1b4bc028c46bb2c3b07a667d8c7560af42487b1570c4e6f4cea5df1605a4d59fc60f0238b0eec4c91fe466497ae8862daf521aecc1be36283678a9a94fc8fdaad6c95c2896d6b4015c05d75aaa95cf89f547503044c1921625905085df258cbbd898f2b05f38d96b20d1530d5e10f7923c05d75aaa95cf89f547503044c1921625905085d653cc6faca6885e36760bf35a25bb63442519b14f258cbbd898f2b05f38d96b20d1530d5e10f7923The analyzer checkout supplies the tool;
--vllm-rootand--ascend-rootsupply the pinned source trees being analyzed. They do not need to be the analyzer checkout itself. Existing repositories can be reused when the head/revision and required base objects match the table; no model installation or NPU is required for this static CLI scan.For example, run #39568 from this PR's analyzer checkout after preparing the two source paths at the revisions above:
PowerShell equivalent (replace the two source directory paths):
Timing diagnostics are printed as phases finish, followed by the compatibility summary; this is not per-file progress. #39568 reports
BREAKS FOUND: vLLM removesSchedulerInterface._get_routed_experts, still called atvllm_ascend/core/recompute_scheduler.py:907in the pinned baseline. Expected exit: 1 for #39568/#50685 (detected P1), 0 for #50620.These commands use the PR's CLI, not pytest: they bypass PR-range network discovery and parent
conftest.pydependency loading while exercising the same analysis engine. Base objects must already exist locally to avoid Git fetching missing objects. The SHAs are replay inputs, not analyzer constants. Raw logs and full comparison evidence are retained locally; the CI entry itself still prints logs without creating report artifacts.