Use Protocols to type-check linear_proj submodules of Attention - #3434
Conversation
|
/ok to test 9db13d6 |
|
Resynced after coming back from travel, sorry for delay! |
yashaswikarnati
left a comment
There was a problem hiding this comment.
synced offline, just had a minor comment,overall lgtm!
Can you expand on this a bit? I'm guessing this is adding the row_parallel_linear_proj() function in addition to the "row_parallel_linear()" function? Don't those have the same inputs/outputs so same types? Why the need for a special one for "_proj"? |
|
@jaredcasper Sure! Fair criticism, this is sorta in a partial state so maybe I should update with a TODO for clarity or something. I'm trying to solve the following problem: backend: BackendSpecProvider = ...
submodules = SelfAttentionSubmodules(..., linear_proj=backend.get_type(), ...)
It can only do so if the thing being passed to But the return type of I don't have a great Protocol to use here for what generically a method named Thus, my proposed solution here is effectively to have |
|
/ok to test 6776b0d |
Head branch was pushed to by a user without write access
|
Sorry for delayed response! Synced and fixed import issue. |
|
/ok to test bf377a2 |
mathemakitten
left a comment
There was a problem hiding this comment.
Re-approving on behalf of @santhnm2
|
🔄 Merge queue validation started! You can track the progress here: https://github.com/NVIDIA/Megatron-LM/actions/runs/25877464990 |
Upstream main tip: f92a207 Pulled commits: - f92a207 Update copy-pr-bot.yaml [skip ci] - b3b6719 ci: tolerate git-gc race in /home/runner chown after checkout (NVIDIA#4808) - 9b4074b Inference: Optimize Prefill Engine Steps for Nemotron (NVIDIA#4764) - a53107c chore: Update nightly tests golden values (NVIDIA#4805) - e9a0930 ci: Update workflow to use same commit for build+test (NVIDIA#4787) - 266562f Update owners (NVIDIA#4794) - 98031e1 Bump nvidia-modelopt>=0.44.0 (NVIDIA#4803) - d167123 fix tokenizers in respect to newer transformers (NVIDIA#4608) - dbfc96b Use Protocols to type-check linear_proj submodules of Attention (NVIDIA#3434) Conflict resolutions: - .github/CODEOWNERS: --theirs (upstream granular team mapping) - megatron/core/transformer/attention.py: composed -- kept ours StreamBP imports, dropped dead ModuleSpec/build_module import (replaced by Protocol in NVIDIA#3434), kept ours attn_proj_manager pattern because our fine_grained_activation_offload.py keeps group_offload API while upstream split to static group_commit, but adopted apply_module(self.linear_proj) wrapper from NVIDIA#3434 - megatron/core/transformer/multi_latent_attention.py: same apply_module+attn_proj_manager composition; took ours ChunkRange parameter for StreamBP plus upstream BaseInferenceContext type annotation - pyproject.toml: bumped modelopt >=0.44 (upstream); kept transformer-engine[pytorch,core_cu13]>=2.9.0a0,<2.12.0 pin (custom for CUDA 13 / SM100) - uv.lock: --ours; modelopt bump is non-breaking. uv lock regenerate blocked by stale /home/sjpat/wheelhouse torchcomms path (pre-existing). Gates: - git diff --check: clean - conflict markers: none - py_compile (15 changed .py files): OK - attention.py + multi_latent_attention.py import OK - indexcache: 27/28 pass (same single GPU-env failure as pre-merge base) - transformer gdn/mtp/moe suite: 53 failed / 7 passed / 55 skipped / 5 errors -- identical to base (all failures are cudaErrorDevicesUnavailable from sglang occupying all H200s) - 2-rank torchrun smoke: blocked (no free GPUs) Custom preserved: StreamBP (megatron/core/transformer/streambp.py + tests/unit_tests/transformer/test_streambp.py + tools/streambp_prod_verify.py), IndexCache config + NVFP4 indexer, HISA topk1024 backward, emerging_optimizers v0.2.0 pin, mHC/MTP/MoE composition.
* origin/main: (138 commits) Refactor CUDA graph API: decompose cuda_graph_scope into full_iteration impl, inference scope, and per-layer capture modules (NVIDIA#4292) Add high-priority A2A stream and HybridEP preprocessing SMs (NVIDIA#4694) add is_torch_min_version in fsdp src (NVIDIA#4812) [Main][feat] Support A2A Overlap for Megatron-FSDP (NVIDIA#3797) Reorder mtp_post_process after attention backward in 1F1B schedule plan (NVIDIA#4695) [fix] Use MSC for checking checkpoint existence (NVIDIA#4251) Combine GEMM + SwiGLU fused MLP PRs (3890, 4071, 4095, 4219, 4311, 4324) → main (NVIDIA#4636) Strengthen test_checkpoint to verify distributed checkpoint behavior (NVIDIA#4711) Disable MSC by default; opt in via --enable-msc (NVIDIA#4629) additional tests for nvrx (NVIDIA#4522) Update copy-pr-bot.yaml [skip ci] ci: tolerate git-gc race in /home/runner chown after checkout (NVIDIA#4808) Inference: Optimize Prefill Engine Steps for Nemotron (NVIDIA#4764) chore: Update nightly tests golden values (NVIDIA#4805) ci: Update workflow to use same commit for building docker image and running tests (NVIDIA#4787) Update owners (NVIDIA#4794) Bump nvidia-modelopt>=0.44.0 (NVIDIA#4803) fix tokenizers in respect to newer transformers (NVIDIA#4608) Use Protocols to type-check linear_proj submodules of Attention (NVIDIA#3434) Fix recompute checkpointing + training CGs (NVIDIA#3919) ... # Conflicts: # megatron/core/transformer/moe/moe_utils.py
…IA#3434) Co-authored-by: gautham-kollu <gkollu@nvidia.com> Co-authored-by: Yashaswi Karnati <144376261+yashaswikarnati@users.noreply.github.com> Co-authored-by: Philip Petrakian <ppetrakian@nvidia.com>
…IA#3434) Co-authored-by: gautham-kollu <gkollu@nvidia.com> Co-authored-by: Yashaswi Karnati <144376261+yashaswikarnati@users.noreply.github.com> Co-authored-by: Philip Petrakian <ppetrakian@nvidia.com> Signed-off-by: yhgalaxy <yhgalaxy@outlook.com>
…IA#3434) Co-authored-by: gautham-kollu <gkollu@nvidia.com> Co-authored-by: Yashaswi Karnati <144376261+yashaswikarnati@users.noreply.github.com> Co-authored-by: Philip Petrakian <ppetrakian@nvidia.com> Signed-off-by: Jon Barker <jbarker@aws-cmh-slurm-1-vscode-02.cm.cluster>
…IA#3434) Co-authored-by: gautham-kollu <gkollu@nvidia.com> Co-authored-by: Yashaswi Karnati <144376261+yashaswikarnati@users.noreply.github.com> Co-authored-by: Philip Petrakian <ppetrakian@nvidia.com>
…IA#3434) Co-authored-by: gautham-kollu <gkollu@nvidia.com> Co-authored-by: Yashaswi Karnati <144376261+yashaswikarnati@users.noreply.github.com> Co-authored-by: Philip Petrakian <ppetrakian@nvidia.com> Signed-off-by: Dmytro Pykhtar <dpykhtar@nvidia.com>
What does this PR do ?
Defines Protocols representing
linear_projsubmodules, and uses them instead of ModuleSpec to enable typechecking of its construction in SelfAttention, CrossAttention, and MLA.I also updated
Backendto returnlinear_projspecifically, allowing type-checking ofRowParallelLineartypes as instances oflinear_projdirectly (otherwiseBackend"hides" the type and makes no type-checking occur).While I was in
attention, I also updated the naming conventions of the existing interfaces to match what we've finalized on.Associated design doc: Typed ModuleSpec.pdf
Contribution process
flowchart LR A[Pre-checks] --> B[PR Tests] subgraph Code Review/Approval C1[Expert Review] --> C2[Final Review] end B --> C1 C2 --> D[Merge]Pre-checks
Core 0.8)Code review
The following process is enforced via the CODEOWNERS file for changes into
megatron/core. For changes outside ofmegatron/core, it is up to the PR author whether or not to tag the Final Reviewer team.For MRs into `main` branch
Feel free to message or comment the @mcore-oncall to help accelerate your merge into main. The less complex your PR is, the faster it will be approved and merged!
(Step 1): Add PR label
Expert Review(Step 2): Collect the expert reviewers reviews
Expert Reviewlabel when your PR is ready for review.Final Review might get declined if these requirements are not fulfilled.
(Step 3): Final Review
Final Reviewlabel(Optional Step 4): Cherry-pick into release branch
If this PR also needs to be merged into
core_r*release branches, after this PR has been merged, selectCherry-pickto open a new PR into the release branch.For MRs into `dev` branch
The proposed review process for `dev` branch is under active discussion.MRs are mergable after one approval by either
eharper@nvidia.comorzijiey@nvidia.com.Merging your PR
Any member of core-adlr and
core-nemowill be able to merge your PR.