Skip to content

[Dev] Move some processing into a function so can be compiled - #3220

Merged
BestJuly merged 5 commits into
NVIDIA:devfrom
BestJuly:lit/qwen3_next_compile
Mar 5, 2026
Merged

[Dev] Move some processing into a function so can be compiled#3220
BestJuly merged 5 commits into
NVIDIA:devfrom
BestJuly:lit/qwen3_next_compile

Conversation

@BestJuly

@BestJuly BestJuly commented Feb 3, 2026

Copy link
Copy Markdown
Contributor

What does this PR do ?

Use triton to compile some ops for acceleration. PR3559 for main branch.

⚠️ For major changes (either in lines of code or in its impact), please make sure to first share a design doc with the team. If you're unsure what's the best way to do so, contact the @mcore-oncall.

Contribution process

flowchart LR
    A[Pre-checks] --> B[PR Tests]
    subgraph Code Review/Approval
        C1[Expert Review] --> C2[Final Review]
    end
    B --> C1
    C2 --> D[Merge]
Loading

Pre-checks

  • I want this PR in a versioned release and have added the appropriate Milestone (e.g., Core 0.8)
  • I have added relevant unit tests
  • I have added relevant functional tests
  • I have added proper typing to my code Typing guidelines
  • I have added relevant documentation
  • I have run the autoformatter.sh on my PR

Code review

The following process is enforced via the CODEOWNERS file for changes into megatron/core. For changes outside of megatron/core, it is up to the PR author whether or not to tag the Final Reviewer team.

For MRs into `main` branch

Feel free to message or comment the @mcore-oncall to help accelerate your merge into main. The less complex your PR is, the faster it will be approved and merged!

(Step 1): Add PR label Expert Review

(Step 2): Collect the expert reviewers reviews

  1. Attach the Expert Review label when your PR is ready for review.
  2. GitHub auto-assigns expert reviewers based on your changes. They will get notified and pick up your PR soon.

⚠️ Only proceed to the next step once all reviewers have approved, merge-conflict are resolved and the CI is passing.
Final Review might get declined if these requirements are not fulfilled.

(Step 3): Final Review

  1. Add Final Review label
  2. GitHub auto-assigns final reviewers based on your changes. They will get notified and pick up your PR soon.

(Optional Step 4): Cherry-pick into release branch

If this PR also needs to be merged into core_r* release branches, after this PR has been merged, select Cherry-pick to open a new PR into the release branch.

For MRs into `dev` branch The proposed review process for `dev` branch is under active discussion.

MRs are mergable after one approval by either eharper@nvidia.com or zijiey@nvidia.com.

Merging your PR

Any member of core-adlr and core-nemo will be able to merge your PR.

@copy-pr-bot

copy-pr-bot Bot commented Feb 3, 2026

Copy link
Copy Markdown

Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually.

Contributors can view more details about this message here.

@BestJuly
BestJuly marked this pull request as ready for review February 24, 2026 06:57
@BestJuly
BestJuly requested review from a team as code owners February 24, 2026 06:57
@ko3n1g ko3n1g added this to the Core 0.16 milestone Feb 24, 2026
@BestJuly BestJuly added the dev branch Dev branch related issues and development label Feb 24, 2026
@janEbert janEbert added Expert Review [deprecated] Apply this label to indicate that your PR is ready for expert review. complexity: low labels Feb 24, 2026
@BestJuly BestJuly changed the title Move some processing into a function so can be compiled [Dev] Move some processing into a function so can be compiled Mar 3, 2026
@BestJuly
BestJuly added this pull request to the merge queue Mar 5, 2026
@svcnvidia-nemo-ci

Copy link
Copy Markdown
Contributor

🔄 Merge queue validation started!

You can track the progress here: https://github.com/NVIDIA/Megatron-LM/actions/runs/22704504035

Merged via the queue into NVIDIA:dev with commit a268231 Mar 5, 2026
48 of 50 checks passed
@BestJuly
BestJuly deleted the lit/qwen3_next_compile branch March 5, 2026 07:56
yuzhongw-nvidia added a commit to yuzhongw-nvidia/Megatron-LM that referenced this pull request Mar 24, 2026
…NVIDIA#3040, NVIDIA#3220)

- Add context parallel (CP) support to GatedDeltaNet via all-to-all
  communication (tensor_a2a_cp2hp / tensor_a2a_hp2cp)
- Refine GDN implementation: replace causal_conv1d_fn with fla.modules.convolution,
  extract _prepare_qkv_for_gated_delta_rule and _compute_g_and_beta as @jit_fuser methods
- Update TransformerConfig to remove CP==1 assertion for gated_delta_net
  and add linear_attention_type deprecation alias
- Enable CP test cases in test_gated_delta_net.py and refactor correctness
  test to use shared _test_parallel_attention_correctness helper

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
yuzhongw-nvidia pushed a commit to yuzhongw-nvidia/Megatron-LM that referenced this pull request Mar 24, 2026
yuzhongw-nvidia pushed a commit to yuzhongw-nvidia/Megatron-LM that referenced this pull request Mar 24, 2026
yuzhongw-nvidia pushed a commit to yuzhongw-nvidia/Megatron-LM that referenced this pull request Mar 24, 2026
yuzhongw-nvidia pushed a commit to yuzhongw-nvidia/Megatron-LM that referenced this pull request Apr 7, 2026
yuzhongw-nvidia pushed a commit to yuzhongw-nvidia/Megatron-LM that referenced this pull request Apr 10, 2026
yuzhongw-nvidia pushed a commit to yuzhongw-nvidia/Megatron-LM that referenced this pull request Apr 13, 2026
yuzhongw-nvidia pushed a commit to yuzhongw-nvidia/Megatron-LM that referenced this pull request Apr 13, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

complexity: low dev branch Dev branch related issues and development Expert Review [deprecated] Apply this label to indicate that your PR is ready for expert review.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants