Skip to content

Fix model convert when use latest megatron - #2267

Merged
zhuzilin merged 1 commit into
THUDM:mainfrom
alexqdh:fix_model_convert_use_latest_megatron
Aug 16, 2026
Merged

zhuzilin merged 1 commit into
THUDM:mainfrom
alexqdh:fix_model_convert_use_latest_megatron

Conversation

@alexqdh

@alexqdh alexqdh commented Aug 12, 2026

Copy link
Copy Markdown
Contributor
  • Add --use-gated-attention compatibility to the HF-to-torch_dist conversion tool while retaining --attention-output-gate.
  • Default enable_gloo_process_groups to True when it is unavailable in the Megatron argument namespace.
  • Support both old and new Megatron tokenizer utility import paths.
  • Allow the custom Qwen3.5 and Qwen3-Next attention modules to accept Megatron’s optional name argument.

@zhuzilin
zhuzilin merged commit 00986d7 into THUDM:main Aug 16, 2026
60 of 61 checks passed
@alexqdh
alexqdh deleted the fix_model_convert_use_latest_megatron branch August 17, 2026 02:23
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants