Skip to content

[Feature][MRV2] expand default MRv2 architecture whitelist and add dspark - #16626

Merged
ningjingbengxiaohai merged 37 commits into
vllm-project:mainfrom
yjyang62:cursor/expand-mrv2-whitelist-9095
Sep 17, 2026
Merged

ningjingbengxiaohai merged 37 commits into
vllm-project:mainfrom
yjyang62:cursor/expand-mrv2-whitelist-9095

Conversation

@yjyang62

@yjyang62 yjyang62 commented Sep 15, 2026 •

Copy link
Copy Markdown
Collaborator

What this PR does / why we need it?
This PR expands the default Model Runner V2 model and feature whitelists so more architectures, plus dspark, default to V2.
A request uses Model Runner V2 by default only when all of the following hold:

The model architecture is on the default-V2 whitelist:

  • Qwen3ForCausalLM
  • Qwen3MoeForCausalLM
  • MiniMaxM2ForCausalLM
  • DeepseekV3ForCausalLM
  • DeepseekV32ForCausalLM
  • GlmMoeDsaForCausalLM
  • DeepseekV4ForCausalLM
  • Qwen3_5MoeForCausalLM

Enabled features are on the default-V2 feature whitelist. Static eagle3 / mtp / dflash / dspark are included. LoRA and dynamic speculative decoding (num_speculative_tokens_per_batch_size) stay on V1.
Triton is available.
Otherwise the request stays on Model Runner V1. Explicit VLLM_USE_V2_MODEL_RUNNER=0/1 still overrides the whitelist.

Does this PR introduce any user-facing change?
Yes. The architectures above now default to Model Runner V2 when the conditions above are satisfied. Other models still default to V1. Explicit VLLM_USE_V2_MODEL_RUNNER=0/1 is unchanged.

How was this patch tested?
Unit tests updated: tests/ut/test_mrv2_utils.py (architecture parametrize, dspark on the feature whitelist)
CI: cpu-ut / e2e on this PR.

@yjyang62 yjyang62 changed the title [Feat][Platform] expand default MRv2 architecture whitelist and add dspark [Feature][MRV2] expand default MRv2 architecture whitelist and add dspark Sep 15, 2026
@github-actions

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.

@yjyang62 yjyang62 added the ready-all run all e2e test for pr label Sep 15, 2026
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request expands the default Model Runner V2 (MRv2) support on Ascend platforms by updating the architecture and speculative method whitelists. It introduces a centralized utility module to manage MRv2 enablement, which decouples the logic from upstream validation and provides a more robust mechanism for applying configuration patches. Additionally, the PR improves the detection of ACL graph capturing to ensure compatibility with the V2 runner, preventing issues where GPU-specific flags were incorrectly interpreted in the NPU environment.

Highlights

  • Expanded MRv2 Whitelist: Added support for multiple new model architectures (including Qwen3, DeepseekV3/V4, and others) and the 'dspark' speculative decoding method to the default Model Runner V2 (MRv2) whitelist.
  • Centralized MRv2 Utilities: Created a new module, mrv2_utils.py, to encapsulate the logic for MRv2 enablement, decoupling it from upstream validation and providing a consistent patching mechanism.
  • Robust Graph Capturing: Implemented a new helper, is_acl_full_graph_capturing, to accurately detect ACL stream capturing, preventing conflicts between GPU-style V2 context flags and NPU graph task groups.
  • Configuration Patching: Standardized the application of MRv2 configuration overrides across the platform, worker processes, and engine core to ensure consistent runner behavior.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

Suggested PR Title:

[Misc][Feature] Implement Ascend-owned whitelist heuristics for Model Runner V2 enablement

Suggested PR Summary:

### What this PR does / why we need it?

This PR implements Ascend-owned whitelist heuristics for Model Runner V2 (MRv2) enablement instead of relying on the upstream GPU-specific defaults, which can cause crashes on unsupported configurations. It introduces `vllm_ascend/mrv2_utils.py` to manage the whitelist of architectures (e.g., Qwen3, DeepseekV3) and speculative decoding methods (e.g., eagle3, mtp, dflash, dspark), ensuring MRv2 is only enabled when the environment is ready (non-310P and Triton available).

Additionally, it resolves an issue where GPU V2's `capturing` flag was incorrectly treated as ACL stream capture, which caused hangs. It introduces `is_acl_full_graph_capturing()` to verify actual NPU stream capture status.

Feedback on the code changes:
- Safely retrieve `model_config` and `speculative_config` using `getattr` to prevent potential `AttributeError`s.
- Add a warning log when Model Runner V2 is disabled due to an unsupported speculative decoding method to improve debuggability.

### Does this PR introduce _any_ user-facing change?

Yes, Model Runner V2 will now be enabled by default for whitelisted architectures (such as Qwen3) and supported speculative decoding methods on compatible Ascend platforms.

### How was this patch tested?

The changes are covered by new unit tests in `tests/ut/test_mrv2_utils.py` and updates to existing tests in `tests/ut/test_ascend_forward_context.py` and `tests/e2e/`.

Comment thread vllm_ascend/mrv2_utils.py
Comment thread vllm_ascend/mrv2_utils.py
Comment thread vllm_ascend/mrv2_utils.py
@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

1 similar comment
@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

@cursor
cursor Bot force-pushed the cursor/expand-mrv2-whitelist-9095 branch from 4fdff6c to ae37a4e Compare September 16, 2026 04:05
@cursor
cursor Bot force-pushed the cursor/expand-mrv2-whitelist-9095 branch 3 times, most recently from eec9224 to 1355a74 Compare September 16, 2026 06:47
@yjyang62 yjyang62 added ready-precise run selected e2e test for pr and removed ready-all run all e2e test for pr labels Sep 16, 2026

@zouzy5137 zouzy5137 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@shiqiangA

shiqiangA commented Sep 16, 2026 •

Copy link
Copy Markdown
Collaborator

/nightly qwen3-30b-a3b-w8a8 kimi25_w4a8_step_3_7 kimi-k2.6_max_model_len acceptace_rate_deepseekv4-flash-w8a8-mtp Qwen3-235B-A22B-W8A8_piecewise_fullgraph_A3 Qwen3_32b_w8a8_ml_A3 qwen3-235b-a22b-w8a8 glm-5.2-w4a8-mtp glm-5.2-w4a8-dspark glm-5.2-w4a8c8-sfa-dcp compressor-metadata-cross-stream Qwen3-32B-QuaRot Qwen3_32b_W8A8_wl_A3 qwen3-32b-int8 qwen3-32b-int8-prefix-cache Qwen3-30B-A3B-W4A8-llm-compressor Qwen3-30B-QuaRot qwen3-30b-acc Minimax_m2.7_w8a8_A3 deepseek-v3-2-w8a8 Deepseek_R1_W8A8_Reasoning_output_A3 deepseek-r1-0528-w8a8-prefix-cache mtpx-deepseek-r1-0528-w8a8 --a3-560t
nightly command triggered.

@shiqiangA

shiqiangA commented Sep 16, 2026 •

Copy link
Copy Markdown
Collaborator

/nightly kimi-k2-thinking Qwq_32b_A3 MiniMax-M3-BF16-A3 Qwen3.8-27B-w8a8-A3 qwen_3_6_27b_w8a8_A3 Qwen3.5-397B-A17B-w8a8-mtp Qwen3.5-27B-w8a8-A3 Qwen3.5-122B-A10B-W8A8-A3 qwen3-vl-235b-a22b-instruct-w8a8 Qwen3-32B-W8A8C8-A3 MiniMax-M3-W8A8-A3 Kimi-K2.6-w4a8-A3 kimi-k2.5 glm-5.1-w8a8-prefill-mc2 glm-4.7-w8a8 DeepSeek-V4-Flash-W8A8-A3 Deepseek_R1_W8A8_A3 rejection-sample gemma4-31b-dense gemma4 qwen3-30b-a3b-w8a8-a2-performance qwen3-32b-int8 Qwen3-Next-80B-A3B-Instruct Qwen3-ASR-1.7B Qwen3-8B qwen3-30b-a3b-bf16-a2-performance multi-node-qwen3-235b-dp Qwen3.5-397B-A17B-w4a8-mtp Qwen3.5-27B-w8a8-A2 qwen3-vl-32b-instruct-w8a8 MiniMax-M2.5-w8a8-QuaRot-A2 multi-node-Kimi-K2.5-W4A8-A2 multi-node-GLM-5.1-w8a8-A2 --a3-560t
nightly command triggered.

@shiqiangA

shiqiangA commented Sep 16, 2026 •

Copy link
Copy Markdown
Collaborator

/nightly qwen3-30b-a3b-w8a8 kimi25_w4a8_step_3_7 kimi-k2.6_max_model_len acceptace_rate_deepseekv4-flash-w8a8-mtp Qwen3-235B-A22B-W8A8_piecewise_fullgraph_A3 Qwen3_32b_w8a8_ml_A3 qwen3-235b-a22b-w8a8 glm-5.2-w4a8-mtp glm-5.2-w4a8-dspark glm-5.2-w4a8c8-sfa-dcp compressor-metadata-cross-stream Qwen3-32B-QuaRot Qwen3_32b_W8A8_wl_A3 qwen3-32b-int8 qwen3-32b-int8-prefix-cache Qwen3-30B-A3B-W4A8-llm-compressor Qwen3-30B-QuaRot qwen3-30b-acc Minimax_m2.7_w8a8_A3 deepseek-v3-2-w8a8 Deepseek_R1_W8A8_Reasoning_output_A3 deepseek-r1-0528-w8a8-prefix-cache mtpx-deepseek-r1-0528-w8a8 --a3-560t
nightly command triggered.

@shiqiangA

shiqiangA commented Sep 16, 2026 •

Copy link
Copy Markdown
Collaborator

/nightly kimi-k2-thinking Qwq_32b_A3 MiniMax-M3-BF16-A3 Qwen3.8-27B-w8a8-A3 qwen_3_6_27b_w8a8_A3 Qwen3.5-397B-A17B-w8a8-mtp Qwen3.5-27B-w8a8-A3 Qwen3.5-122B-A10B-W8A8-A3 qwen3-vl-235b-a22b-instruct-w8a8 Qwen3-32B-W8A8C8-A3 MiniMax-M3-W8A8-A3 Kimi-K2.6-w4a8-A3 kimi-k2.5 glm-5.1-w8a8-prefill-mc2 glm-4.7-w8a8 DeepSeek-V4-Flash-W8A8-A3 Deepseek_R1_W8A8_A3 rejection-sample gemma4-31b-dense gemma4 qwen3-30b-a3b-w8a8-a2-performance qwen3-32b-int8 Qwen3-Next-80B-A3B-Instruct Qwen3-ASR-1.7B Qwen3-8B qwen3-30b-a3b-bf16-a2-performance multi-node-qwen3-235b-dp Qwen3.5-397B-A17B-w4a8-mtp Qwen3.5-27B-w8a8-A2 qwen3-vl-32b-instruct-w8a8 MiniMax-M2.5-w8a8-QuaRot-A2 multi-node-Kimi-K2.5-W4A8-A2 multi-node-GLM-5.1-w8a8-A2 --a3-560t
nightly command triggered.

@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

@cursor
cursor Bot force-pushed the cursor/expand-mrv2-whitelist-9095 branch from ba671cc to ae0186a Compare September 16, 2026 12:58
@shiqiangA

Copy link
Copy Markdown
Collaborator

/nightly QWEN3_235B_PD DeepSeek-V4-Pro-w4a8-1M-PD multi-node-deepseek-v3.2-W8A8-EP QWEN3_235B_PD_3_5K_1_5k multi-node-dpsk3.2-2node multi-node-GLM-5.1-w8a8-A3 multi-node-GLM-5.2-w8a8-A3 multi-node-qwenw8a8-2node-eplb multi-node-GLM-5.1-W8A8C8-A3_128k_90_50 multi-node-GLM-5.1-W8A8C8-MTP-A3_198k_function multi-node-deepseek-v3.1 Minimax_m2.7_in128k_1k_prefix90_tpot50 DeepSeek-V4-flash-w8a8-PD-prefix

ZhangwenTaoHW and others added 2 commits September 17, 2026 13:53
Signed-off-by: ZhangwenTaoHW <zhangwentao101@huawei.com>
Pass optional reasoning_effort through AISBench chat request generation.
Set reasoning_effort: low for the DeepSeek-V4 Flash W8A8 A3 nightly GPQA
case. Keep thinking: true and the performance case unchanged.

Cherry-picked from vllm-project#16721

Co-authored-by: GDzhu01 <116337067+GDzhu01@users.noreply.github.com>
Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com>
@cursor
cursor Bot force-pushed the cursor/expand-mrv2-whitelist-9095 branch from 29c01ae to 3c33450 Compare September 17, 2026 13:53
Qwen3_5ForConditionalGeneration hybrid VL hits MRv2 encoder graph
capture (CUDA stream assert) and hybrid KV copy. Keep it on V1 by
default. Pin dflash2 PIECEWISE acceptance to V1 after vllm-project#16726 broke
dummy propose without set_forward_context; V2 eager remains covered.

Signed-off-by: yjyang62 <yjyang62@users.noreply.github.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
@ningjingbengxiaohai
ningjingbengxiaohai merged commit c7ca0b6 into vllm-project:main Sep 17, 2026
30 checks passed
czydyy added a commit to czydyy/vllm-ascend that referenced this pull request Sep 18, 2026
…telist and add dspark (vllm-project#16626)"

This reverts commit c7ca0b6.

Signed-off-by: chenzeyu <2978509328@qq.com>
cursor Bot pushed a commit to yjyang62/vllm-ascend that referenced this pull request Sep 18, 2026
Keep DeepSeek V4.1 forward-context moe_comm_methods from main and drop
the reverted _USE_V2_EXTRA_KWARGS assertion from the unit test.

Signed-off-by: Cursor Agent <cursoragent@cursor.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
czydyy added a commit to czydyy/vllm-ascend that referenced this pull request Sep 18, 2026
…telist and add dspark (vllm-project#16626)"

This reverts commit c7ca0b6.

Signed-off-by: chenzeyu <2978509328@qq.com>
ningjingbengxiaohai pushed a commit that referenced this pull request Sep 18, 2026
…d add dspark" (#16832)

Reverts #16626

- vLLM main:
vllm-project/vllm@84030bb

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
czydyy added a commit to czydyy/vllm-ascend that referenced this pull request Sep 18, 2026
vllm-project#16544)

Reverts commit 200309d on main
(cherry-picked from main_verify 82145c7).

Conflict resolution: ascend_forward_context.py, patch/__init__.py,
worker.py and test_ascend_forward_context.py were entangled with vllm-project#16626
(merged before vllm-project#16544, already reverted on this branch); they are
restored to the pre-vllm-project#16626 state (c7ca0b6~1), the correct composition
of both reverts. dsa_cp.py and test_model_runner_v2.py keep later
upstream changes (vllm-project#16168 metadata buffer sizing fix, vllm-project#15747 spec-pp
protocol rename). All other surviving files match the pre-PR state.

Signed-off-by: chenzeyu <2978509328@qq.com>
cursor Bot pushed a commit to yjyang62/vllm-ascend that referenced this pull request Sep 18, 2026
Reapply the code removed by the upstream revert of vllm-project#16626 while retaining the newer blacklist-based runner selection from vllm-project#16848.

Signed-off-by: Cursor Agent <cursoragent@cursor.com>

Co-authored-by: yjyang62 <yjyang62@users.noreply.github.com>
cursor Bot pushed a commit to yjyang62/vllm-ascend that referenced this pull request Sep 21, 2026
Keep V2 EPLB conversion, but do not overlay the pre-vllm-project#15514 yaml snapshot
from vllm-project#16626. Restore vllm-project#16900 Kimi baseline 1296, vllm-project#16815 ROCE=1, and
vllm-project#15514 --kv-cache-dtype/--indexer_kv_dtype fp8 without the retired
enable_sparse_*_c8 additional-config keys.

Signed-off-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Sage Martin <lethamannodaiu@outlook.com>
Signed-off-by: Cursor Agent <cursoragent@cursor.com>
yjyang62 pushed a commit to yjyang62/vllm-ascend that referenced this pull request Sep 21, 2026
Keep V2 EPLB conversion, but do not overlay the pre-vllm-project#15514 yaml snapshot
from vllm-project#16626. Restore vllm-project#16900 Kimi baseline 1296, vllm-project#16815 ROCE=1, and
enable_sparse_*_c8 additional-config keys.

Signed-off-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Sage Martin <lethamannodaiu@outlook.com>
Signed-off-by: Cursor Agent <cursoragent@cursor.com>
Oseltamivir added a commit to Oseltamivir/vllm-ascend that referenced this pull request Sep 22, 2026
Restore the optional-mask access used in upstream vllm-project#16626 to resolve the shared mypy failure tracked in vllm-project#16658. Preserve configured cache routing and the all-layer fallback. Five isolated cache-routing cases pass.

Assisted-by: OpenAI Codex
Signed-off-by: Oseltamivir <58582368+Oseltamivir@users.noreply.github.com>
Oseltamivir added a commit to Oseltamivir/vllm-ascend that referenced this pull request Sep 22, 2026
Restore the optional-mask access used in upstream vllm-project#16626 to resolve the shared mypy failure tracked in vllm-project#16658. Preserve configured cache routing and the all-layer fallback. Five isolated cache-routing cases pass.

Assisted-by: OpenAI Codex
Signed-off-by: Oseltamivir <58582368+Oseltamivir@users.noreply.github.com>
cursor Bot pushed a commit to yjyang62/vllm-ascend that referenced this pull request Sep 22, 2026
Keep V2 EPLB conversion, but do not overlay the pre-vllm-project#15514 yaml snapshot
from vllm-project#16626. Restore vllm-project#16900 Kimi baseline 1296, vllm-project#16815 ROCE=1, and
enable_sparse_*_c8 additional-config keys.

Signed-off-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Sage Martin <lethamannodaiu@outlook.com>
Signed-off-by: Cursor Agent <cursoragent@cursor.com>
cursor Bot pushed a commit to yjyang62/vllm-ascend that referenced this pull request Sep 23, 2026
Keep V2 EPLB conversion, but do not overlay the pre-vllm-project#15514 yaml snapshot
from vllm-project#16626. Restore vllm-project#16900 Kimi baseline 1296, vllm-project#16815 ROCE=1, and
enable_sparse_*_c8 additional-config keys.

Signed-off-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Sage Martin <lethamannodaiu@outlook.com>
Signed-off-by: Cursor Agent <cursoragent@cursor.com>
yjyang62 pushed a commit to yjyang62/vllm-ascend that referenced this pull request Sep 23, 2026
Keep V2 EPLB conversion, but do not overlay the pre-vllm-project#15514 yaml snapshot
from vllm-project#16626. Restore vllm-project#16900 Kimi baseline 1296, vllm-project#16815 ROCE=1, and
enable_sparse_*_c8 additional-config keys.

Signed-off-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Sage Martin <lethamannodaiu@outlook.com>
Signed-off-by: Cursor Agent <cursoragent@cursor.com>
yjyang62 pushed a commit to yjyang62/vllm-ascend that referenced this pull request Sep 23, 2026
Keep V2 EPLB conversion, but do not overlay the pre-vllm-project#15514 yaml snapshot
from vllm-project#16626. Restore vllm-project#16900 Kimi baseline 1296, vllm-project#16815 ROCE=1, and
enable_sparse_*_c8 additional-config keys.

Signed-off-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Sage Martin <lethamannodaiu@outlook.com>
Signed-off-by: Cursor Agent <cursoragent@cursor.com>
cursor Bot pushed a commit to yjyang62/vllm-ascend that referenced this pull request Sep 23, 2026
Keep V2 EPLB conversion, but do not overlay the pre-vllm-project#15514 yaml snapshot
from vllm-project#16626. Restore vllm-project#16900 Kimi baseline 1296, vllm-project#16815 ROCE=1, and
enable_sparse_*_c8 additional-config keys.

Signed-off-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Sage Martin <lethamannodaiu@outlook.com>
Signed-off-by: Cursor Agent <cursoragent@cursor.com>
yjyang62 pushed a commit to yjyang62/vllm-ascend that referenced this pull request Sep 24, 2026
Keep V2 EPLB conversion, but do not overlay the pre-vllm-project#15514 yaml snapshot
from vllm-project#16626. Restore vllm-project#16900 Kimi baseline 1296, vllm-project#16815 ROCE=1, and
enable_sparse_*_c8 additional-config keys.

Signed-off-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Sage Martin <lethamannodaiu@outlook.com>
Signed-off-by: Cursor Agent <cursoragent@cursor.com>
yjyang62 pushed a commit to yjyang62/vllm-ascend that referenced this pull request Sep 24, 2026
Keep V2 EPLB conversion, but do not overlay the pre-vllm-project#15514 yaml snapshot
from vllm-project#16626. Restore vllm-project#16900 Kimi baseline 1296, vllm-project#16815 ROCE=1, and
enable_sparse_*_c8 additional-config keys.

Signed-off-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Sage Martin <lethamannodaiu@outlook.com>
Signed-off-by: Cursor Agent <cursoragent@cursor.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants