Repository navigation
[Misc]feat: adapt to vLLM main (54503ece) - #12420
wangxiyuan merged 3 commits into
Conversation
Signed-off-by: main2main-bot <main2main-bot@users.noreply.github.com>
Signed-off-by: main2main-bot <main2main-bot@users.noreply.github.com>
Signed-off-by: main2main-bot <main2main-bot@users.noreply.github.com>
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request synchronizes the vllm-ascend repository with the latest upstream vLLM commit. The primary focus is adapting existing model-specific logic and configuration handling to accommodate the newly introduced 'longcat_flash_ngram' model type, ensuring consistent behavior across KV cache management and speculative decoding configurations. Highlights
New Features🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console. Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
|
👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:
If CI fails, you can run linting and testing checks locally according Contributing and Testing. |
There was a problem hiding this comment.
Code Review
Suggested PR Title:
[Model][Feature] Support longcat_flash_ngram and update Hunyuan processor compatibilitySuggested PR Summary:
### What this PR does / why we need it?
This PR introduces support for the `longcat_flash_ngram` model type alongside `longcat_flash` in several components, including the Mooncake connectors, speculative config patching, and model runner. It also simplifies the `hunyuan_prompt` placeholder in tests and removes the conditional check when installing the Hunyuan VL processor compatibility layer. Additionally, the verified commit hash for `vllm-main` is updated.
### Does this PR introduce _any_ user-facing change?
No user-facing changes are introduced.
### How was this patch tested?
Tested via existing end-to-end tests and CI.I have no further feedback to provide as there are no review comments.
### What this PR does / why we need it? Upgrade vLLM commit to `54503ece` 1. Adapt `vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake_connector.py`, `vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake_layerwise_connector.py`, `vllm_ascend/patch/platform/patch_speculative_config.py`, `vllm_ascend/worker/model_runner_v1.py` due to [08dfd686](vllm-project/vllm@08dfd686) - Upstream added `longcat_flash_ngram` model_type alongside existing `longcat_flash` — vllm-ascend had 5 locations that checked `model_type == "longcat_flash"` for dual-attention module count and spec decode MTP remapping. 2. Adapt `tests/e2e/conftest.py`, `vllm_ascend/patch/hunyuan_vl_processor_compat.py` due to [0b6636cb](vllm-project/vllm@0b6636cb) - Upstream added `longcat_flash_ngram` model_type alongside existing `longcat_flash` — vllm-ascend had 5 locations that checked `model_type == "longcat_flash"` for dual-attention module count and spec decode MTP remapping. - vLLM version: v0.25.0 - vLLM main: vllm-project/vllm@85c09e9 --------- Signed-off-by: main2main-bot <main2main-bot@users.noreply.github.com> Co-authored-by: main2main-bot <main2main-bot@users.noreply.github.com> Signed-off-by: HeFangjun <hfj0219@mail.ustc.edu.cn>
### What this PR does / why we need it? Upgrade vLLM commit to `54503ece` 1. Adapt `vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake_connector.py`, `vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake_layerwise_connector.py`, `vllm_ascend/patch/platform/patch_speculative_config.py`, `vllm_ascend/worker/model_runner_v1.py` due to [08dfd686](vllm-project/vllm@08dfd686) - Upstream added `longcat_flash_ngram` model_type alongside existing `longcat_flash` — vllm-ascend had 5 locations that checked `model_type == "longcat_flash"` for dual-attention module count and spec decode MTP remapping. 2. Adapt `tests/e2e/conftest.py`, `vllm_ascend/patch/hunyuan_vl_processor_compat.py` due to [0b6636cb](vllm-project/vllm@0b6636cb) - Upstream added `longcat_flash_ngram` model_type alongside existing `longcat_flash` — vllm-ascend had 5 locations that checked `model_type == "longcat_flash"` for dual-attention module count and spec decode MTP remapping. - vLLM version: v0.25.0 - vLLM main: vllm-project/vllm@85c09e9 --------- Signed-off-by: main2main-bot <main2main-bot@users.noreply.github.com> Co-authored-by: main2main-bot <main2main-bot@users.noreply.github.com>
### What this PR does / why we need it? Upgrade vLLM commit to `54503ece` 1. Adapt `vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake_connector.py`, `vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake_layerwise_connector.py`, `vllm_ascend/patch/platform/patch_speculative_config.py`, `vllm_ascend/worker/model_runner_v1.py` due to [08dfd686](vllm-project/vllm@08dfd686) - Upstream added `longcat_flash_ngram` model_type alongside existing `longcat_flash` — vllm-ascend had 5 locations that checked `model_type == "longcat_flash"` for dual-attention module count and spec decode MTP remapping. 2. Adapt `tests/e2e/conftest.py`, `vllm_ascend/patch/hunyuan_vl_processor_compat.py` due to [0b6636cb](vllm-project/vllm@0b6636cb) - Upstream added `longcat_flash_ngram` model_type alongside existing `longcat_flash` — vllm-ascend had 5 locations that checked `model_type == "longcat_flash"` for dual-attention module count and spec decode MTP remapping. - vLLM version: v0.25.0 - vLLM main: vllm-project/vllm@85c09e9 --------- Signed-off-by: main2main-bot <main2main-bot@users.noreply.github.com> Co-authored-by: main2main-bot <main2main-bot@users.noreply.github.com>
### What this PR does / why we need it? Upgrade vLLM commit to `54503ece` 1. Adapt `vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake_connector.py`, `vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake_layerwise_connector.py`, `vllm_ascend/patch/platform/patch_speculative_config.py`, `vllm_ascend/worker/model_runner_v1.py` due to [08dfd686](vllm-project/vllm@08dfd686) - Upstream added `longcat_flash_ngram` model_type alongside existing `longcat_flash` — vllm-ascend had 5 locations that checked `model_type == "longcat_flash"` for dual-attention module count and spec decode MTP remapping. 2. Adapt `tests/e2e/conftest.py`, `vllm_ascend/patch/hunyuan_vl_processor_compat.py` due to [0b6636cb](vllm-project/vllm@0b6636cb) - Upstream added `longcat_flash_ngram` model_type alongside existing `longcat_flash` — vllm-ascend had 5 locations that checked `model_type == "longcat_flash"` for dual-attention module count and spec decode MTP remapping. - vLLM version: v0.25.0 - vLLM main: vllm-project/vllm@85c09e9 --------- Signed-off-by: main2main-bot <main2main-bot@users.noreply.github.com> Co-authored-by: main2main-bot <main2main-bot@users.noreply.github.com>
### What this PR does / why we need it? Upgrade vLLM commit to `54503ece` 1. Adapt `vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake_connector.py`, `vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake_layerwise_connector.py`, `vllm_ascend/patch/platform/patch_speculative_config.py`, `vllm_ascend/worker/model_runner_v1.py` due to [08dfd686](vllm-project/vllm@08dfd686) - Upstream added `longcat_flash_ngram` model_type alongside existing `longcat_flash` — vllm-ascend had 5 locations that checked `model_type == "longcat_flash"` for dual-attention module count and spec decode MTP remapping. 2. Adapt `tests/e2e/conftest.py`, `vllm_ascend/patch/hunyuan_vl_processor_compat.py` due to [0b6636cb](vllm-project/vllm@0b6636cb) - Upstream added `longcat_flash_ngram` model_type alongside existing `longcat_flash` — vllm-ascend had 5 locations that checked `model_type == "longcat_flash"` for dual-attention module count and spec decode MTP remapping. - vLLM version: v0.25.0 - vLLM main: vllm-project/vllm@85c09e9 --------- Signed-off-by: main2main-bot <main2main-bot@users.noreply.github.com> Co-authored-by: main2main-bot <main2main-bot@users.noreply.github.com>
### What this PR does / why we need it? Reduce repeated work in the CI-only vLLM PR compatibility analyzer introduced by #14560, and fix import-presence checks for deleted or annotation-only names. Changes stay under `tests/e2e/vllm_interface/`. The pytest entry, PR base/head resolution, complete source indexing, three analysis branches, deterministic summary, and failure policy remain in place. No persistent cache, pickle files, new environment variables, additional CI jobs, or main2main/monkey-patch analysis are introduced. #### Execution flow and implementation details 1. **Resolve and verify inputs.** The existing pytest entry resolves the vLLM PR base/head and records the vllm-ascend revision. Commit and checkout verification remains enabled. 2. **Index source trees.** `InterfaceBoundaryGenerator` indexes the complete source trees and finalizes the indexes. Single-path namespace merges copy an already-normalized state instead of repeatedly rebuilding sorted keys and alternative sets. State dictionaries remain independent; immutable binding tuples can be shared. 3. **Resolve aliases and namespace states.** `RepositoryIndex.canonical_name()` checks successively shorter dot-delimited prefixes in the live alias table instead of sorting/scanning every alias per lookup. Longest-prefix behavior and cycle protection remain. Always-bound module names are derived from the final namespace already computed. Decorator and direct-call resolution reuse identical statement-prefix states through a per-module, 256-entry LRU keyed by actual AST statements and guard names. The memo returns independent dictionaries; expression resolution and caller-specific fallback still run for each query. It is not shared across revisions or CI runs. 4. **Discover dependencies.** Override and exact-call discovery retains affected call sites and inherited implementation locations. Import discovery reuses the already-parsed vllm-ascend ASTs, with a source-reading fallback for files absent from the index. Triton launch invocation kinds remain distinct. 5. **Read base/head source.** Each analysis branch retains separate base/head `GitSnapshot` instances. Each snapshot lazily uses one `git cat-file --batch` process rather than launching `git show` per file. Binary response framing preserves empty files, CRLF and non-ASCII source. Invalid or failed reads raise analysis errors, not missing-symbol findings. An `ExitStack` closes processes and temporary stderr streams on success or exceptions. This is not a persistent source cache. 6. **Resolve contracts and compare usage.** Snapshots reuse module/class namespace states, named bindings, owner lookup and resolved API contracts. Call-contract keys include target expression, access kind, receiver type, member and invocation kind. Argument binding and return-use checks still execute separately at every call site; one site's compatibility result is never reused for another. 7. **Check import presence.** Import checks resolve presence before constructing unused signature/return details. Presence now follows the final namespace: a later `del` removes a binding; an annotation without a value does not create a name or erase an existing value. Constants and re-exports that remain bound are retained. A P1 requires proven presence at the base and proven absence at the head. Ambiguous bindings, star imports, dynamic module exports and unparseable source are not treated as proven removals. Package-submodule fallback is preserved. Full endpoint details are still built for findings. 8. **Merge and print.** The three branches merge and sort findings deterministically. The CI entry prints the summary directly; it does not create report artifacts. Historical findings remain subject to the existing PR-only filtering policy. ### Does this PR introduce _any_ user-facing change? Reduced analyzer runtime and more accurate detection of imports whose names were deleted or became annotation-only. This is therefore not exclusively a performance-only change. Analyzer version advances to `2.1.1`. CLI options and CI log wording are unchanged; existing dynamic-dispatch and return-dataflow limitations remain. ### How was this patch tested? #### Local regressions - 123 local tests passed: Git batch-reader lifecycle/error handling, alias equivalence, final import bindings, callable/return contracts, serial/parallel fixture equivalence, pytest-entry integration, and scope-prefix isolation/eviction. - Included 12,000 synthetic alias comparisons and 196 generated scope-flow combinations. Deleted and annotation-only imports were exercised through actual CLI runs against temporary Git repositories. - The pytest integration fixtures launch the analyzer subprocess; the analyzer is not mocked. Historical replay scans described below call the analyzer core directly, not the hardware E2E suite. - Ruff, formatting, compileall and mypy checks passed in local validation. Full repository formatting was previously attempted; some hooks require Linux shell tools unavailable on this Windows host. Full Linux CI and NPU workloads are not claimed as locally verified. - Regression/replay harnesses remain outside the E2E collection path, as requested; no analyzer unit-test directory is reintroduced into vLLM PR jobs. #### Latest paired performance measurement Windows / Python 3.11, four indexing processes and three analysis threads, fixed source SHAs, fresh processes, sequential scans and no persistent analyzer cache. Times below are analyzer time, excluding network range discovery and hardware tests. | vLLM PR | Before latest scope reuse | After | Reduction | Result on both sides | | --- | ---: | ---: | ---: | --- | | #39568 | 73.6 s | 61.9 s | 15.9% | 1 P1 | | #50685 | 100.3 s | 82.0 s | 18.2% | 3 P1 impacts / 1 root cause | | #50620 | 100.4 s | 81.8 s | 18.5% | PASS | This baseline already includes the earlier Git batching, alias optimization and import fix; it is **not** the original pre-optimization implementation. Full reports matched after removing only timing metadata, stdout matched byte-for-byte, and expected exit codes were 1/1/0. These are single paired measurements, not statistical benchmarks or Linux CI speed guarantees. #### Expanded accuracy replay Eight additional cases were each run against the frozen pre-scope-reuse baseline and current code: 16 actual scans. All eight full reports matched except timing, and all eight stdout logs were byte-identical. | Fixed historical input | P1 impacts | P1 roots | Review findings | | --- | ---: | ---: | ---: | | vllm-ascend #11709 range | 5 | 5 | 6 | | vllm-ascend #12020 range | 2 | 1 | 0 | | vllm-ascend #12420 range | 0 | 0 | 0 | | vllm-ascend #12502 range | 33 | 20 | 0 | | vllm-ascend #12648 range | 0 | 0 | 5 | | vllm-ascend #13358 range | 5 | 4 | 11 | | vLLM #47808 | 2 | 2 | 5 | | vLLM #50504 | 0 | 0 | 0 | The upgrade ranges are fixed CI-analyzer inputs, not a full main2main mode. #12648 uses its adapted vllm-ascend head; the other upgrade cases use their pre-adaptation baselines. The earlier three cases above were reused only after verifying current source hashes, giving 11 distinct cases overall. There were not 22 new scans in the expanded round. 29 semantic/history assertions passed. Previously confirmed P1 findings remained. #47808 matches the newer August 28 historical log (2 P1), not the older August 17 report (1 P1). #11709's six review items already existed in the frozen baseline; they are not introduced by scope reuse. Known limitations are not hidden by these results: #12648's `compute_slot_mappings(out=...)` remains review-only because its dispatch requires monkey-patch/field propagation; broader return-value consumption is still outside the exact dependency model. Regression equivalence does not establish universal precision/recall. #### Fresh verification of the published PR head After pushing commit `2e31d3bcbdc82960bc0a1956abdd412f6ef1f27b`, fetched `refs/pull/16364/head` from `vllm-project/vllm-ascend` and checked it out in a separate, clean detached worktree. The fetched commit and all nine analyzer source hashes matched the validated version. Three additional fresh scans used that fetched code: | vLLM PR | Result | Analyzer time | Process wall time | Exit code | | --- | --- | ---: | ---: | ---: | | #39568 | 1 P1 | 59.8 s | 61.0 s | 1 | | #50685 | 3 P1 impacts / 1 root cause | 82.2 s | 83.7 s | 1 | | #50620 | PASS | 79.8 s | 81.3 s | 0 | All three complete reports were identical to the pre-push reports after excluding only timing metadata; raw stdout was byte-identical. Exit 1 is the expected detected-break outcome, not an analyzer crash. These runs are separate from the earlier paired measurements and expanded regression round. The 123 local tests passed again before pushing. Ruff, formatting, spelling and Markdown checks passed; several shell-based repository hooks could not execute on Windows. The fetched code also passed mypy with the Python 3.10/Linux target. This is static type checking, not execution on Linux. These historical scans execute `analyze_range()` and the production summary renderer directly. They do not run the full upstream pytest job or NPU workloads. The actual pytest entry is covered separately by the local integration fixtures. Published-head verification details: #16364 (comment) #### Reproduce a pinned case With this PR's analyzer checked out, prepare separate clean source worktrees at the following head/revision: | vLLM PR | vLLM base | vLLM head | vllm-ascend revision | | --- | --- | --- | --- | | #39568 | `ce29c26b31d432b1b4bc028c46bb2c3b07a667d8` | `c7560af42487b1570c4e6f4cea5df1605a4d59fc` | `60f0238b0eec4c91fe466497ae8862daf521aecc` | | #50685 | `1be36283678a9a94fc8fdaad6c95c2896d6b4015` | `c05d75aaa95cf89f547503044c1921625905085d` | `f258cbbd898f2b05f38d96b20d1530d5e10f7923` | | #50620 | `c05d75aaa95cf89f547503044c1921625905085d` | `653cc6faca6885e36760bf35a25bb63442519b14` | `f258cbbd898f2b05f38d96b20d1530d5e10f7923` | The analyzer checkout supplies the tool; `--vllm-root` and `--ascend-root` supply the pinned source trees being analyzed. They do not need to be the analyzer checkout itself. Existing repositories can be reused when the head/revision and required base objects match the table; no model installation or NPU is required for this static CLI scan. For example, run #39568 from this PR's analyzer checkout after preparing the two source paths at the revisions above: ```bash VLLM_INTERFACE_TIMINGS=1 python -u -m tests.e2e.vllm_interface.vllm_interface_contracts analyze-range \ --vllm-root /path/to/vllm-source \ --ascend-root /path/to/vllm-ascend-source \ --old ce29c26 \ --new c7560af \ --expect-ascend-sha 60f0238 \ --index-workers 4 --analysis-workers 3 --fail-on introduced ``` PowerShell equivalent (replace the two source directory paths): ```powershell $env:VLLM_INTERFACE_TIMINGS = "1" python -u -m tests.e2e.vllm_interface.vllm_interface_contracts analyze-range ` --vllm-root "C:\path\to\vllm-source" ` --ascend-root "C:\path\to\vllm-ascend-source" ` --old ce29c26 ` --new c7560af ` --expect-ascend-sha 60f0238 ` --index-workers 4 --analysis-workers 3 --fail-on introduced Write-Host "Analyzer exit code: $LASTEXITCODE" ``` Timing diagnostics are printed as phases finish, followed by the compatibility summary; this is not per-file progress. #39568 reports `BREAKS FOUND`: vLLM removes `SchedulerInterface._get_routed_experts`, still called at `vllm_ascend/core/recompute_scheduler.py:907` in the pinned baseline. Expected exit: 1 for #39568/#50685 (detected P1), 0 for #50620. These commands use the PR's CLI, not pytest: they bypass PR-range network discovery and parent `conftest.py` dependency loading while exercising the same analysis engine. Base objects must already exist locally to avoid Git fetching missing objects. The SHAs are replay inputs, not analyzer constants. Raw logs and full comparison evidence are retained locally; the CI entry itself still prints logs without creating report artifacts. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: shenzhao <shenzhao9@huawei.com> Co-authored-by: shenzhao <shenzhao9@huawei.com>
) ### What this PR does / why we need it? Reduce repeated work in the CI-only vLLM PR compatibility analyzer introduced by vllm-project#14560, and fix import-presence checks for deleted or annotation-only names. Changes stay under `tests/e2e/vllm_interface/`. The pytest entry, PR base/head resolution, complete source indexing, three analysis branches, deterministic summary, and failure policy remain in place. No persistent cache, pickle files, new environment variables, additional CI jobs, or main2main/monkey-patch analysis are introduced. #### Execution flow and implementation details 1. **Resolve and verify inputs.** The existing pytest entry resolves the vLLM PR base/head and records the vllm-ascend revision. Commit and checkout verification remains enabled. 2. **Index source trees.** `InterfaceBoundaryGenerator` indexes the complete source trees and finalizes the indexes. Single-path namespace merges copy an already-normalized state instead of repeatedly rebuilding sorted keys and alternative sets. State dictionaries remain independent; immutable binding tuples can be shared. 3. **Resolve aliases and namespace states.** `RepositoryIndex.canonical_name()` checks successively shorter dot-delimited prefixes in the live alias table instead of sorting/scanning every alias per lookup. Longest-prefix behavior and cycle protection remain. Always-bound module names are derived from the final namespace already computed. Decorator and direct-call resolution reuse identical statement-prefix states through a per-module, 256-entry LRU keyed by actual AST statements and guard names. The memo returns independent dictionaries; expression resolution and caller-specific fallback still run for each query. It is not shared across revisions or CI runs. 4. **Discover dependencies.** Override and exact-call discovery retains affected call sites and inherited implementation locations. Import discovery reuses the already-parsed vllm-ascend ASTs, with a source-reading fallback for files absent from the index. Triton launch invocation kinds remain distinct. 5. **Read base/head source.** Each analysis branch retains separate base/head `GitSnapshot` instances. Each snapshot lazily uses one `git cat-file --batch` process rather than launching `git show` per file. Binary response framing preserves empty files, CRLF and non-ASCII source. Invalid or failed reads raise analysis errors, not missing-symbol findings. An `ExitStack` closes processes and temporary stderr streams on success or exceptions. This is not a persistent source cache. 6. **Resolve contracts and compare usage.** Snapshots reuse module/class namespace states, named bindings, owner lookup and resolved API contracts. Call-contract keys include target expression, access kind, receiver type, member and invocation kind. Argument binding and return-use checks still execute separately at every call site; one site's compatibility result is never reused for another. 7. **Check import presence.** Import checks resolve presence before constructing unused signature/return details. Presence now follows the final namespace: a later `del` removes a binding; an annotation without a value does not create a name or erase an existing value. Constants and re-exports that remain bound are retained. A P1 requires proven presence at the base and proven absence at the head. Ambiguous bindings, star imports, dynamic module exports and unparseable source are not treated as proven removals. Package-submodule fallback is preserved. Full endpoint details are still built for findings. 8. **Merge and print.** The three branches merge and sort findings deterministically. The CI entry prints the summary directly; it does not create report artifacts. Historical findings remain subject to the existing PR-only filtering policy. ### Does this PR introduce _any_ user-facing change? Reduced analyzer runtime and more accurate detection of imports whose names were deleted or became annotation-only. This is therefore not exclusively a performance-only change. Analyzer version advances to `2.1.1`. CLI options and CI log wording are unchanged; existing dynamic-dispatch and return-dataflow limitations remain. ### How was this patch tested? #### Local regressions - 123 local tests passed: Git batch-reader lifecycle/error handling, alias equivalence, final import bindings, callable/return contracts, serial/parallel fixture equivalence, pytest-entry integration, and scope-prefix isolation/eviction. - Included 12,000 synthetic alias comparisons and 196 generated scope-flow combinations. Deleted and annotation-only imports were exercised through actual CLI runs against temporary Git repositories. - The pytest integration fixtures launch the analyzer subprocess; the analyzer is not mocked. Historical replay scans described below call the analyzer core directly, not the hardware E2E suite. - Ruff, formatting, compileall and mypy checks passed in local validation. Full repository formatting was previously attempted; some hooks require Linux shell tools unavailable on this Windows host. Full Linux CI and NPU workloads are not claimed as locally verified. - Regression/replay harnesses remain outside the E2E collection path, as requested; no analyzer unit-test directory is reintroduced into vLLM PR jobs. #### Latest paired performance measurement Windows / Python 3.11, four indexing processes and three analysis threads, fixed source SHAs, fresh processes, sequential scans and no persistent analyzer cache. Times below are analyzer time, excluding network range discovery and hardware tests. | vLLM PR | Before latest scope reuse | After | Reduction | Result on both sides | | --- | ---: | ---: | ---: | --- | | #39568 | 73.6 s | 61.9 s | 15.9% | 1 P1 | | #50685 | 100.3 s | 82.0 s | 18.2% | 3 P1 impacts / 1 root cause | | #50620 | 100.4 s | 81.8 s | 18.5% | PASS | This baseline already includes the earlier Git batching, alias optimization and import fix; it is **not** the original pre-optimization implementation. Full reports matched after removing only timing metadata, stdout matched byte-for-byte, and expected exit codes were 1/1/0. These are single paired measurements, not statistical benchmarks or Linux CI speed guarantees. #### Expanded accuracy replay Eight additional cases were each run against the frozen pre-scope-reuse baseline and current code: 16 actual scans. All eight full reports matched except timing, and all eight stdout logs were byte-identical. | Fixed historical input | P1 impacts | P1 roots | Review findings | | --- | ---: | ---: | ---: | | vllm-ascend vllm-project#11709 range | 5 | 5 | 6 | | vllm-ascend vllm-project#12020 range | 2 | 1 | 0 | | vllm-ascend vllm-project#12420 range | 0 | 0 | 0 | | vllm-ascend vllm-project#12502 range | 33 | 20 | 0 | | vllm-ascend vllm-project#12648 range | 0 | 0 | 5 | | vllm-ascend vllm-project#13358 range | 5 | 4 | 11 | | vLLM #47808 | 2 | 2 | 5 | | vLLM #50504 | 0 | 0 | 0 | The upgrade ranges are fixed CI-analyzer inputs, not a full main2main mode. vllm-project#12648 uses its adapted vllm-ascend head; the other upgrade cases use their pre-adaptation baselines. The earlier three cases above were reused only after verifying current source hashes, giving 11 distinct cases overall. There were not 22 new scans in the expanded round. 29 semantic/history assertions passed. Previously confirmed P1 findings remained. #47808 matches the newer August 28 historical log (2 P1), not the older August 17 report (1 P1). vllm-project#11709's six review items already existed in the frozen baseline; they are not introduced by scope reuse. Known limitations are not hidden by these results: vllm-project#12648's `compute_slot_mappings(out=...)` remains review-only because its dispatch requires monkey-patch/field propagation; broader return-value consumption is still outside the exact dependency model. Regression equivalence does not establish universal precision/recall. #### Fresh verification of the published PR head After pushing commit `2e31d3bcbdc82960bc0a1956abdd412f6ef1f27b`, fetched `refs/pull/16364/head` from `vllm-project/vllm-ascend` and checked it out in a separate, clean detached worktree. The fetched commit and all nine analyzer source hashes matched the validated version. Three additional fresh scans used that fetched code: | vLLM PR | Result | Analyzer time | Process wall time | Exit code | | --- | --- | ---: | ---: | ---: | | #39568 | 1 P1 | 59.8 s | 61.0 s | 1 | | #50685 | 3 P1 impacts / 1 root cause | 82.2 s | 83.7 s | 1 | | #50620 | PASS | 79.8 s | 81.3 s | 0 | All three complete reports were identical to the pre-push reports after excluding only timing metadata; raw stdout was byte-identical. Exit 1 is the expected detected-break outcome, not an analyzer crash. These runs are separate from the earlier paired measurements and expanded regression round. The 123 local tests passed again before pushing. Ruff, formatting, spelling and Markdown checks passed; several shell-based repository hooks could not execute on Windows. The fetched code also passed mypy with the Python 3.10/Linux target. This is static type checking, not execution on Linux. These historical scans execute `analyze_range()` and the production summary renderer directly. They do not run the full upstream pytest job or NPU workloads. The actual pytest entry is covered separately by the local integration fixtures. Published-head verification details: vllm-project#16364 (comment) #### Reproduce a pinned case With this PR's analyzer checked out, prepare separate clean source worktrees at the following head/revision: | vLLM PR | vLLM base | vLLM head | vllm-ascend revision | | --- | --- | --- | --- | | #39568 | `ce29c26b31d432b1b4bc028c46bb2c3b07a667d8` | `c7560af42487b1570c4e6f4cea5df1605a4d59fc` | `60f0238b0eec4c91fe466497ae8862daf521aecc` | | #50685 | `1be36283678a9a94fc8fdaad6c95c2896d6b4015` | `c05d75aaa95cf89f547503044c1921625905085d` | `f258cbbd898f2b05f38d96b20d1530d5e10f7923` | | #50620 | `c05d75aaa95cf89f547503044c1921625905085d` | `653cc6faca6885e36760bf35a25bb63442519b14` | `f258cbbd898f2b05f38d96b20d1530d5e10f7923` | The analyzer checkout supplies the tool; `--vllm-root` and `--ascend-root` supply the pinned source trees being analyzed. They do not need to be the analyzer checkout itself. Existing repositories can be reused when the head/revision and required base objects match the table; no model installation or NPU is required for this static CLI scan. For example, run #39568 from this PR's analyzer checkout after preparing the two source paths at the revisions above: ```bash VLLM_INTERFACE_TIMINGS=1 python -u -m tests.e2e.vllm_interface.vllm_interface_contracts analyze-range \ --vllm-root /path/to/vllm-source \ --ascend-root /path/to/vllm-ascend-source \ --old ce29c26 \ --new c7560af \ --expect-ascend-sha 60f0238 \ --index-workers 4 --analysis-workers 3 --fail-on introduced ``` PowerShell equivalent (replace the two source directory paths): ```powershell $env:VLLM_INTERFACE_TIMINGS = "1" python -u -m tests.e2e.vllm_interface.vllm_interface_contracts analyze-range ` --vllm-root "C:\path\to\vllm-source" ` --ascend-root "C:\path\to\vllm-ascend-source" ` --old ce29c26 ` --new c7560af ` --expect-ascend-sha 60f0238 ` --index-workers 4 --analysis-workers 3 --fail-on introduced Write-Host "Analyzer exit code: $LASTEXITCODE" ``` Timing diagnostics are printed as phases finish, followed by the compatibility summary; this is not per-file progress. #39568 reports `BREAKS FOUND`: vLLM removes `SchedulerInterface._get_routed_experts`, still called at `vllm_ascend/core/recompute_scheduler.py:907` in the pinned baseline. Expected exit: 1 for #39568/#50685 (detected P1), 0 for #50620. These commands use the PR's CLI, not pytest: they bypass PR-range network discovery and parent `conftest.py` dependency loading while exercising the same analysis engine. Base objects must already exist locally to avoid Git fetching missing objects. The SHAs are replay inputs, not analyzer constants. Raw logs and full comparison evidence are retained locally; the CI entry itself still prints logs without creating report artifacts. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: shenzhao <shenzhao9@huawei.com> Co-authored-by: shenzhao <shenzhao9@huawei.com>
) ### What this PR does / why we need it? Reduce repeated work in the CI-only vLLM PR compatibility analyzer introduced by vllm-project#14560, and fix import-presence checks for deleted or annotation-only names. Changes stay under `tests/e2e/vllm_interface/`. The pytest entry, PR base/head resolution, complete source indexing, three analysis branches, deterministic summary, and failure policy remain in place. No persistent cache, pickle files, new environment variables, additional CI jobs, or main2main/monkey-patch analysis are introduced. #### Execution flow and implementation details 1. **Resolve and verify inputs.** The existing pytest entry resolves the vLLM PR base/head and records the vllm-ascend revision. Commit and checkout verification remains enabled. 2. **Index source trees.** `InterfaceBoundaryGenerator` indexes the complete source trees and finalizes the indexes. Single-path namespace merges copy an already-normalized state instead of repeatedly rebuilding sorted keys and alternative sets. State dictionaries remain independent; immutable binding tuples can be shared. 3. **Resolve aliases and namespace states.** `RepositoryIndex.canonical_name()` checks successively shorter dot-delimited prefixes in the live alias table instead of sorting/scanning every alias per lookup. Longest-prefix behavior and cycle protection remain. Always-bound module names are derived from the final namespace already computed. Decorator and direct-call resolution reuse identical statement-prefix states through a per-module, 256-entry LRU keyed by actual AST statements and guard names. The memo returns independent dictionaries; expression resolution and caller-specific fallback still run for each query. It is not shared across revisions or CI runs. 4. **Discover dependencies.** Override and exact-call discovery retains affected call sites and inherited implementation locations. Import discovery reuses the already-parsed vllm-ascend ASTs, with a source-reading fallback for files absent from the index. Triton launch invocation kinds remain distinct. 5. **Read base/head source.** Each analysis branch retains separate base/head `GitSnapshot` instances. Each snapshot lazily uses one `git cat-file --batch` process rather than launching `git show` per file. Binary response framing preserves empty files, CRLF and non-ASCII source. Invalid or failed reads raise analysis errors, not missing-symbol findings. An `ExitStack` closes processes and temporary stderr streams on success or exceptions. This is not a persistent source cache. 6. **Resolve contracts and compare usage.** Snapshots reuse module/class namespace states, named bindings, owner lookup and resolved API contracts. Call-contract keys include target expression, access kind, receiver type, member and invocation kind. Argument binding and return-use checks still execute separately at every call site; one site's compatibility result is never reused for another. 7. **Check import presence.** Import checks resolve presence before constructing unused signature/return details. Presence now follows the final namespace: a later `del` removes a binding; an annotation without a value does not create a name or erase an existing value. Constants and re-exports that remain bound are retained. A P1 requires proven presence at the base and proven absence at the head. Ambiguous bindings, star imports, dynamic module exports and unparseable source are not treated as proven removals. Package-submodule fallback is preserved. Full endpoint details are still built for findings. 8. **Merge and print.** The three branches merge and sort findings deterministically. The CI entry prints the summary directly; it does not create report artifacts. Historical findings remain subject to the existing PR-only filtering policy. ### Does this PR introduce _any_ user-facing change? Reduced analyzer runtime and more accurate detection of imports whose names were deleted or became annotation-only. This is therefore not exclusively a performance-only change. Analyzer version advances to `2.1.1`. CLI options and CI log wording are unchanged; existing dynamic-dispatch and return-dataflow limitations remain. ### How was this patch tested? #### Local regressions - 123 local tests passed: Git batch-reader lifecycle/error handling, alias equivalence, final import bindings, callable/return contracts, serial/parallel fixture equivalence, pytest-entry integration, and scope-prefix isolation/eviction. - Included 12,000 synthetic alias comparisons and 196 generated scope-flow combinations. Deleted and annotation-only imports were exercised through actual CLI runs against temporary Git repositories. - The pytest integration fixtures launch the analyzer subprocess; the analyzer is not mocked. Historical replay scans described below call the analyzer core directly, not the hardware E2E suite. - Ruff, formatting, compileall and mypy checks passed in local validation. Full repository formatting was previously attempted; some hooks require Linux shell tools unavailable on this Windows host. Full Linux CI and NPU workloads are not claimed as locally verified. - Regression/replay harnesses remain outside the E2E collection path, as requested; no analyzer unit-test directory is reintroduced into vLLM PR jobs. #### Latest paired performance measurement Windows / Python 3.11, four indexing processes and three analysis threads, fixed source SHAs, fresh processes, sequential scans and no persistent analyzer cache. Times below are analyzer time, excluding network range discovery and hardware tests. | vLLM PR | Before latest scope reuse | After | Reduction | Result on both sides | | --- | ---: | ---: | ---: | --- | | #39568 | 73.6 s | 61.9 s | 15.9% | 1 P1 | | #50685 | 100.3 s | 82.0 s | 18.2% | 3 P1 impacts / 1 root cause | | #50620 | 100.4 s | 81.8 s | 18.5% | PASS | This baseline already includes the earlier Git batching, alias optimization and import fix; it is **not** the original pre-optimization implementation. Full reports matched after removing only timing metadata, stdout matched byte-for-byte, and expected exit codes were 1/1/0. These are single paired measurements, not statistical benchmarks or Linux CI speed guarantees. #### Expanded accuracy replay Eight additional cases were each run against the frozen pre-scope-reuse baseline and current code: 16 actual scans. All eight full reports matched except timing, and all eight stdout logs were byte-identical. | Fixed historical input | P1 impacts | P1 roots | Review findings | | --- | ---: | ---: | ---: | | vllm-ascend vllm-project#11709 range | 5 | 5 | 6 | | vllm-ascend vllm-project#12020 range | 2 | 1 | 0 | | vllm-ascend vllm-project#12420 range | 0 | 0 | 0 | | vllm-ascend vllm-project#12502 range | 33 | 20 | 0 | | vllm-ascend vllm-project#12648 range | 0 | 0 | 5 | | vllm-ascend vllm-project#13358 range | 5 | 4 | 11 | | vLLM #47808 | 2 | 2 | 5 | | vLLM #50504 | 0 | 0 | 0 | The upgrade ranges are fixed CI-analyzer inputs, not a full main2main mode. vllm-project#12648 uses its adapted vllm-ascend head; the other upgrade cases use their pre-adaptation baselines. The earlier three cases above were reused only after verifying current source hashes, giving 11 distinct cases overall. There were not 22 new scans in the expanded round. 29 semantic/history assertions passed. Previously confirmed P1 findings remained. #47808 matches the newer August 28 historical log (2 P1), not the older August 17 report (1 P1). vllm-project#11709's six review items already existed in the frozen baseline; they are not introduced by scope reuse. Known limitations are not hidden by these results: vllm-project#12648's `compute_slot_mappings(out=...)` remains review-only because its dispatch requires monkey-patch/field propagation; broader return-value consumption is still outside the exact dependency model. Regression equivalence does not establish universal precision/recall. #### Fresh verification of the published PR head After pushing commit `2e31d3bcbdc82960bc0a1956abdd412f6ef1f27b`, fetched `refs/pull/16364/head` from `vllm-project/vllm-ascend` and checked it out in a separate, clean detached worktree. The fetched commit and all nine analyzer source hashes matched the validated version. Three additional fresh scans used that fetched code: | vLLM PR | Result | Analyzer time | Process wall time | Exit code | | --- | --- | ---: | ---: | ---: | | #39568 | 1 P1 | 59.8 s | 61.0 s | 1 | | #50685 | 3 P1 impacts / 1 root cause | 82.2 s | 83.7 s | 1 | | #50620 | PASS | 79.8 s | 81.3 s | 0 | All three complete reports were identical to the pre-push reports after excluding only timing metadata; raw stdout was byte-identical. Exit 1 is the expected detected-break outcome, not an analyzer crash. These runs are separate from the earlier paired measurements and expanded regression round. The 123 local tests passed again before pushing. Ruff, formatting, spelling and Markdown checks passed; several shell-based repository hooks could not execute on Windows. The fetched code also passed mypy with the Python 3.10/Linux target. This is static type checking, not execution on Linux. These historical scans execute `analyze_range()` and the production summary renderer directly. They do not run the full upstream pytest job or NPU workloads. The actual pytest entry is covered separately by the local integration fixtures. Published-head verification details: vllm-project#16364 (comment) #### Reproduce a pinned case With this PR's analyzer checked out, prepare separate clean source worktrees at the following head/revision: | vLLM PR | vLLM base | vLLM head | vllm-ascend revision | | --- | --- | --- | --- | | #39568 | `ce29c26b31d432b1b4bc028c46bb2c3b07a667d8` | `c7560af42487b1570c4e6f4cea5df1605a4d59fc` | `60f0238b0eec4c91fe466497ae8862daf521aecc` | | #50685 | `1be36283678a9a94fc8fdaad6c95c2896d6b4015` | `c05d75aaa95cf89f547503044c1921625905085d` | `f258cbbd898f2b05f38d96b20d1530d5e10f7923` | | #50620 | `c05d75aaa95cf89f547503044c1921625905085d` | `653cc6faca6885e36760bf35a25bb63442519b14` | `f258cbbd898f2b05f38d96b20d1530d5e10f7923` | The analyzer checkout supplies the tool; `--vllm-root` and `--ascend-root` supply the pinned source trees being analyzed. They do not need to be the analyzer checkout itself. Existing repositories can be reused when the head/revision and required base objects match the table; no model installation or NPU is required for this static CLI scan. For example, run #39568 from this PR's analyzer checkout after preparing the two source paths at the revisions above: ```bash VLLM_INTERFACE_TIMINGS=1 python -u -m tests.e2e.vllm_interface.vllm_interface_contracts analyze-range \ --vllm-root /path/to/vllm-source \ --ascend-root /path/to/vllm-ascend-source \ --old ce29c26 \ --new c7560af \ --expect-ascend-sha 60f0238 \ --index-workers 4 --analysis-workers 3 --fail-on introduced ``` PowerShell equivalent (replace the two source directory paths): ```powershell $env:VLLM_INTERFACE_TIMINGS = "1" python -u -m tests.e2e.vllm_interface.vllm_interface_contracts analyze-range ` --vllm-root "C:\path\to\vllm-source" ` --ascend-root "C:\path\to\vllm-ascend-source" ` --old ce29c26 ` --new c7560af ` --expect-ascend-sha 60f0238 ` --index-workers 4 --analysis-workers 3 --fail-on introduced Write-Host "Analyzer exit code: $LASTEXITCODE" ``` Timing diagnostics are printed as phases finish, followed by the compatibility summary; this is not per-file progress. #39568 reports `BREAKS FOUND`: vLLM removes `SchedulerInterface._get_routed_experts`, still called at `vllm_ascend/core/recompute_scheduler.py:907` in the pinned baseline. Expected exit: 1 for #39568/#50685 (detected P1), 0 for #50620. These commands use the PR's CLI, not pytest: they bypass PR-range network discovery and parent `conftest.py` dependency loading while exercising the same analysis engine. Base objects must already exist locally to avoid Git fetching missing objects. The SHAs are replay inputs, not analyzer constants. Raw logs and full comparison evidence are retained locally; the CI entry itself still prints logs without creating report artifacts. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: shenzhao <shenzhao9@huawei.com> Co-authored-by: shenzhao <shenzhao9@huawei.com> Signed-off-by: tianming2009 <13246728590@163.com>
) ### What this PR does / why we need it? Reduce repeated work in the CI-only vLLM PR compatibility analyzer introduced by vllm-project#14560, and fix import-presence checks for deleted or annotation-only names. Changes stay under `tests/e2e/vllm_interface/`. The pytest entry, PR base/head resolution, complete source indexing, three analysis branches, deterministic summary, and failure policy remain in place. No persistent cache, pickle files, new environment variables, additional CI jobs, or main2main/monkey-patch analysis are introduced. #### Execution flow and implementation details 1. **Resolve and verify inputs.** The existing pytest entry resolves the vLLM PR base/head and records the vllm-ascend revision. Commit and checkout verification remains enabled. 2. **Index source trees.** `InterfaceBoundaryGenerator` indexes the complete source trees and finalizes the indexes. Single-path namespace merges copy an already-normalized state instead of repeatedly rebuilding sorted keys and alternative sets. State dictionaries remain independent; immutable binding tuples can be shared. 3. **Resolve aliases and namespace states.** `RepositoryIndex.canonical_name()` checks successively shorter dot-delimited prefixes in the live alias table instead of sorting/scanning every alias per lookup. Longest-prefix behavior and cycle protection remain. Always-bound module names are derived from the final namespace already computed. Decorator and direct-call resolution reuse identical statement-prefix states through a per-module, 256-entry LRU keyed by actual AST statements and guard names. The memo returns independent dictionaries; expression resolution and caller-specific fallback still run for each query. It is not shared across revisions or CI runs. 4. **Discover dependencies.** Override and exact-call discovery retains affected call sites and inherited implementation locations. Import discovery reuses the already-parsed vllm-ascend ASTs, with a source-reading fallback for files absent from the index. Triton launch invocation kinds remain distinct. 5. **Read base/head source.** Each analysis branch retains separate base/head `GitSnapshot` instances. Each snapshot lazily uses one `git cat-file --batch` process rather than launching `git show` per file. Binary response framing preserves empty files, CRLF and non-ASCII source. Invalid or failed reads raise analysis errors, not missing-symbol findings. An `ExitStack` closes processes and temporary stderr streams on success or exceptions. This is not a persistent source cache. 6. **Resolve contracts and compare usage.** Snapshots reuse module/class namespace states, named bindings, owner lookup and resolved API contracts. Call-contract keys include target expression, access kind, receiver type, member and invocation kind. Argument binding and return-use checks still execute separately at every call site; one site's compatibility result is never reused for another. 7. **Check import presence.** Import checks resolve presence before constructing unused signature/return details. Presence now follows the final namespace: a later `del` removes a binding; an annotation without a value does not create a name or erase an existing value. Constants and re-exports that remain bound are retained. A P1 requires proven presence at the base and proven absence at the head. Ambiguous bindings, star imports, dynamic module exports and unparseable source are not treated as proven removals. Package-submodule fallback is preserved. Full endpoint details are still built for findings. 8. **Merge and print.** The three branches merge and sort findings deterministically. The CI entry prints the summary directly; it does not create report artifacts. Historical findings remain subject to the existing PR-only filtering policy. ### Does this PR introduce _any_ user-facing change? Reduced analyzer runtime and more accurate detection of imports whose names were deleted or became annotation-only. This is therefore not exclusively a performance-only change. Analyzer version advances to `2.1.1`. CLI options and CI log wording are unchanged; existing dynamic-dispatch and return-dataflow limitations remain. ### How was this patch tested? #### Local regressions - 123 local tests passed: Git batch-reader lifecycle/error handling, alias equivalence, final import bindings, callable/return contracts, serial/parallel fixture equivalence, pytest-entry integration, and scope-prefix isolation/eviction. - Included 12,000 synthetic alias comparisons and 196 generated scope-flow combinations. Deleted and annotation-only imports were exercised through actual CLI runs against temporary Git repositories. - The pytest integration fixtures launch the analyzer subprocess; the analyzer is not mocked. Historical replay scans described below call the analyzer core directly, not the hardware E2E suite. - Ruff, formatting, compileall and mypy checks passed in local validation. Full repository formatting was previously attempted; some hooks require Linux shell tools unavailable on this Windows host. Full Linux CI and NPU workloads are not claimed as locally verified. - Regression/replay harnesses remain outside the E2E collection path, as requested; no analyzer unit-test directory is reintroduced into vLLM PR jobs. #### Latest paired performance measurement Windows / Python 3.11, four indexing processes and three analysis threads, fixed source SHAs, fresh processes, sequential scans and no persistent analyzer cache. Times below are analyzer time, excluding network range discovery and hardware tests. | vLLM PR | Before latest scope reuse | After | Reduction | Result on both sides | | --- | ---: | ---: | ---: | --- | | #39568 | 73.6 s | 61.9 s | 15.9% | 1 P1 | | #50685 | 100.3 s | 82.0 s | 18.2% | 3 P1 impacts / 1 root cause | | #50620 | 100.4 s | 81.8 s | 18.5% | PASS | This baseline already includes the earlier Git batching, alias optimization and import fix; it is **not** the original pre-optimization implementation. Full reports matched after removing only timing metadata, stdout matched byte-for-byte, and expected exit codes were 1/1/0. These are single paired measurements, not statistical benchmarks or Linux CI speed guarantees. #### Expanded accuracy replay Eight additional cases were each run against the frozen pre-scope-reuse baseline and current code: 16 actual scans. All eight full reports matched except timing, and all eight stdout logs were byte-identical. | Fixed historical input | P1 impacts | P1 roots | Review findings | | --- | ---: | ---: | ---: | | vllm-ascend vllm-project#11709 range | 5 | 5 | 6 | | vllm-ascend vllm-project#12020 range | 2 | 1 | 0 | | vllm-ascend vllm-project#12420 range | 0 | 0 | 0 | | vllm-ascend vllm-project#12502 range | 33 | 20 | 0 | | vllm-ascend vllm-project#12648 range | 0 | 0 | 5 | | vllm-ascend vllm-project#13358 range | 5 | 4 | 11 | | vLLM #47808 | 2 | 2 | 5 | | vLLM #50504 | 0 | 0 | 0 | The upgrade ranges are fixed CI-analyzer inputs, not a full main2main mode. vllm-project#12648 uses its adapted vllm-ascend head; the other upgrade cases use their pre-adaptation baselines. The earlier three cases above were reused only after verifying current source hashes, giving 11 distinct cases overall. There were not 22 new scans in the expanded round. 29 semantic/history assertions passed. Previously confirmed P1 findings remained. #47808 matches the newer August 28 historical log (2 P1), not the older August 17 report (1 P1). vllm-project#11709's six review items already existed in the frozen baseline; they are not introduced by scope reuse. Known limitations are not hidden by these results: vllm-project#12648's `compute_slot_mappings(out=...)` remains review-only because its dispatch requires monkey-patch/field propagation; broader return-value consumption is still outside the exact dependency model. Regression equivalence does not establish universal precision/recall. #### Fresh verification of the published PR head After pushing commit `2e31d3bcbdc82960bc0a1956abdd412f6ef1f27b`, fetched `refs/pull/16364/head` from `vllm-project/vllm-ascend` and checked it out in a separate, clean detached worktree. The fetched commit and all nine analyzer source hashes matched the validated version. Three additional fresh scans used that fetched code: | vLLM PR | Result | Analyzer time | Process wall time | Exit code | | --- | --- | ---: | ---: | ---: | | #39568 | 1 P1 | 59.8 s | 61.0 s | 1 | | #50685 | 3 P1 impacts / 1 root cause | 82.2 s | 83.7 s | 1 | | #50620 | PASS | 79.8 s | 81.3 s | 0 | All three complete reports were identical to the pre-push reports after excluding only timing metadata; raw stdout was byte-identical. Exit 1 is the expected detected-break outcome, not an analyzer crash. These runs are separate from the earlier paired measurements and expanded regression round. The 123 local tests passed again before pushing. Ruff, formatting, spelling and Markdown checks passed; several shell-based repository hooks could not execute on Windows. The fetched code also passed mypy with the Python 3.10/Linux target. This is static type checking, not execution on Linux. These historical scans execute `analyze_range()` and the production summary renderer directly. They do not run the full upstream pytest job or NPU workloads. The actual pytest entry is covered separately by the local integration fixtures. Published-head verification details: vllm-project#16364 (comment) #### Reproduce a pinned case With this PR's analyzer checked out, prepare separate clean source worktrees at the following head/revision: | vLLM PR | vLLM base | vLLM head | vllm-ascend revision | | --- | --- | --- | --- | | #39568 | `ce29c26b31d432b1b4bc028c46bb2c3b07a667d8` | `c7560af42487b1570c4e6f4cea5df1605a4d59fc` | `60f0238b0eec4c91fe466497ae8862daf521aecc` | | #50685 | `1be36283678a9a94fc8fdaad6c95c2896d6b4015` | `c05d75aaa95cf89f547503044c1921625905085d` | `f258cbbd898f2b05f38d96b20d1530d5e10f7923` | | #50620 | `c05d75aaa95cf89f547503044c1921625905085d` | `653cc6faca6885e36760bf35a25bb63442519b14` | `f258cbbd898f2b05f38d96b20d1530d5e10f7923` | The analyzer checkout supplies the tool; `--vllm-root` and `--ascend-root` supply the pinned source trees being analyzed. They do not need to be the analyzer checkout itself. Existing repositories can be reused when the head/revision and required base objects match the table; no model installation or NPU is required for this static CLI scan. For example, run #39568 from this PR's analyzer checkout after preparing the two source paths at the revisions above: ```bash VLLM_INTERFACE_TIMINGS=1 python -u -m tests.e2e.vllm_interface.vllm_interface_contracts analyze-range \ --vllm-root /path/to/vllm-source \ --ascend-root /path/to/vllm-ascend-source \ --old ce29c26 \ --new c7560af \ --expect-ascend-sha 60f0238 \ --index-workers 4 --analysis-workers 3 --fail-on introduced ``` PowerShell equivalent (replace the two source directory paths): ```powershell $env:VLLM_INTERFACE_TIMINGS = "1" python -u -m tests.e2e.vllm_interface.vllm_interface_contracts analyze-range ` --vllm-root "C:\path\to\vllm-source" ` --ascend-root "C:\path\to\vllm-ascend-source" ` --old ce29c26 ` --new c7560af ` --expect-ascend-sha 60f0238 ` --index-workers 4 --analysis-workers 3 --fail-on introduced Write-Host "Analyzer exit code: $LASTEXITCODE" ``` Timing diagnostics are printed as phases finish, followed by the compatibility summary; this is not per-file progress. #39568 reports `BREAKS FOUND`: vLLM removes `SchedulerInterface._get_routed_experts`, still called at `vllm_ascend/core/recompute_scheduler.py:907` in the pinned baseline. Expected exit: 1 for #39568/#50685 (detected P1), 0 for #50620. These commands use the PR's CLI, not pytest: they bypass PR-range network discovery and parent `conftest.py` dependency loading while exercising the same analysis engine. Base objects must already exist locally to avoid Git fetching missing objects. The SHAs are replay inputs, not analyzer constants. Raw logs and full comparison evidence are retained locally; the CI entry itself still prints logs without creating report artifacts. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: shenzhao <shenzhao9@huawei.com> Co-authored-by: shenzhao <shenzhao9@huawei.com> Signed-off-by: like-0517 <ithwlike@126.com>
What this PR does / why we need it?
Upgrade vLLM commit to
54503ecevllm_ascend/distributed/kv_transfer/kv_p2p/mooncake_connector.py,vllm_ascend/distributed/kv_transfer/kv_p2p/mooncake_layerwise_connector.py,vllm_ascend/patch/platform/patch_speculative_config.py,vllm_ascend/worker/model_runner_v1.pydue to 08dfd686longcat_flash_ngrammodel_type alongside existinglongcat_flash— vllm-ascend had 5 locations that checkedmodel_type == "longcat_flash"for dual-attention module count and spec decode MTP remapping.tests/e2e/conftest.py,vllm_ascend/patch/hunyuan_vl_processor_compat.pydue to 0b6636cblongcat_flash_ngrammodel_type alongside existinglongcat_flash— vllm-ascend had 5 locations that checkedmodel_type == "longcat_flash"for dual-attention module count and spec decode MTP remapping.