Skip to content

Replace the model config shell scripts with python - #1910

Merged
fzyzcjy merged 17 commits into
mainfrom
tom/refactor-miles/op8-11
Aug 9, 2026
Merged

Replace the model config shell scripts with python#1910
fzyzcjy merged 17 commits into
mainfrom
tom/refactor-miles/op8-11

Conversation

@fzyzcjy

@fzyzcjy fzyzcjy commented Jul 28, 2026

Copy link
Copy Markdown
Collaborator

Part of #1837

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Caution

The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased.

@fzyzcjy
fzyzcjy force-pushed the tom/refactor-miles/op8-11 branch from 4583664 to cf656e3 Compare July 29, 2026 01:02
@fzyzcjy
fzyzcjy force-pushed the tom/refactor-miles/op8-12 branch from c5d2be0 to 4c327dc Compare July 29, 2026 01:02
@fzyzcjy
fzyzcjy force-pushed the tom/refactor-miles/op8-11 branch from cf656e3 to 740e031 Compare July 29, 2026 01:45
@fzyzcjy
fzyzcjy force-pushed the tom/refactor-miles/op8-12 branch 2 times, most recently from 23a508c to bc3c2f5 Compare July 29, 2026 02:06
@fzyzcjy
fzyzcjy force-pushed the tom/refactor-miles/op8-11 branch from 740e031 to cbc0762 Compare July 29, 2026 02:06
@fzyzcjy
fzyzcjy force-pushed the tom/refactor-miles/op8-12 branch from bc3c2f5 to b500137 Compare July 29, 2026 02:26
@fzyzcjy
fzyzcjy force-pushed the tom/refactor-miles/op8-11 branch from cbc0762 to 5cf6e24 Compare July 29, 2026 02:26
@fzyzcjy
fzyzcjy force-pushed the tom/refactor-miles/op8-12 branch from b500137 to dd22e48 Compare August 4, 2026 04:13
@fzyzcjy
fzyzcjy force-pushed the tom/refactor-miles/op8-11 branch from 5cf6e24 to 23a65d9 Compare August 4, 2026 04:13
@fzyzcjy
fzyzcjy force-pushed the tom/refactor-miles/op8-12 branch from dd22e48 to 5232fa7 Compare August 8, 2026 02:38
@fzyzcjy
fzyzcjy force-pushed the tom/refactor-miles/op8-11 branch from 23a65d9 to 1d07063 Compare August 8, 2026 02:38
@fzyzcjy
fzyzcjy force-pushed the tom/refactor-miles/op8-12 branch from 5232fa7 to ac2d451 Compare August 8, 2026 06:36
@fzyzcjy
fzyzcjy force-pushed the tom/refactor-miles/op8-11 branch from 1d07063 to fbefce1 Compare August 8, 2026 06:36
fzyzcjy added 6 commits August 9, 2026 17:58
Squashed from:
- Fix PYTHONBUFFERED typo in launch scripts and command utils
- Fix the same typo in the NPU docker patch
- Unbuffer the ray workers, not only the submitting client
- Unbuffer the launchers that submit ray jobs of their own
Squashed from:
- Add a shell launch script test harness
- Make the shell harness report stderr and emit shim stdout correctly
- Drop the deprecated huggingface-cli shim
- Exercise the shim behaviours the single real script never reaches
- Poll the ray cluster the way the real scripts do in the synthetic script
- Group the harness tests by what they exercise
- Make the harness record commands in fork order and refuse to be unfrozen
…r paths

Squashed from:
- Fix launch scripts whose model config path could never resolve
- Fix two launcher entrypoints that raised before issuing any command
- Cover the two regressions this op fixes
… scripts

Squashed from:
- Derive the miles checkout location instead of hardcoding it
- Quote the derived train.py path and pin the invariant
- Find the shell scripts without shelling out to git
Squashed from:
- Snapshot the external commands of every shell launch script
- Apply isort and black to the shell launch script test
- Intercept ps so the recordings do not read the host process list
- Harden the harness against host state the rollout exposed
- Share the snapshot compare-or-update step and stop running each script twice
- Keep the generated snapshots under one obvious tests/snapshots tree
- Group the launch script tests by subject
- Name the shell launcher test after what it covers
- Assert the recorded order for every script, including the concurrent ones
- Regenerate the concurrent launcher's snapshot in its true command order
ExecuteTrainConfig.num_nodes read SLURM_JOB_NUM_NODES into a class-level
default, so the value was fixed when command_utils was imported. A test that
wants a deterministic launch command cannot undo that with monkeypatch, and a
process that sets the variable after import does not see it either.

A default_factory reads it at construction instead, but dataclass_cli copied
the parameter's declared default straight into the click signature, and for a
factory field that default is dataclasses' _HAS_DEFAULT_FACTORY sentinel,
which click then type-casts:

    TypeError: int() argument must be ... not '_HAS_DEFAULT_FACTORY_CLASS'

Every scripts/run_*.py exposes this config through that bridge. Resolve the
factory when the signature is built, the way the argparse bridge already does.
fzyzcjy added 10 commits August 9, 2026 17:58
…ript

Squashed from:
- Snapshot the commands built by every python launch script
- Apply isort and black to the python launch script test
- Reuse the shell harness sanitizer and snapshot helper
- Move the python launcher snapshots into the shared tree too
- Share the command recorder with the command_utils tests
- Freeze the launcher environment that the snapshots actually depend on
- Regenerate the launcher snapshots for the ray runtime unbuffering
- Snapshot the config files a launcher generates, not just its commands
- Record the generated precision config in the deepseek-v4 snapshots
Squashed from:
- Cover the public surface of command_utils with unit tests
- Group the command_utils tests by the function under test
- Close the gaps that let the command_utils tests pass on broken behaviour
- Keep the command_utils tests in one file
Squashed from:
- Rename exec_command by the resource its command needs
- Point the nvlink and single-node conversion tests at the gpu helper
- Re-record the multi-node label the rename changed
- Rename the last two exec_command call sites the split missed
- Re-record the multi-node label in the rsync_simple test too
Squashed from:
- Move the shell exec helpers next to their only consumers
- Carry NodeAffinitySchedulingStrategy along with the moved exec helpers
- Stop patching command helpers on a module that no longer has them
…yloads

Squashed from:
- Accept inline base64 payloads for the config file arguments
- Pass config documents inline instead of through a temp file
- Make the inline config payload reach every consumer and fail loudly
- Apply pre-commit import ordering
- Regenerate the deepseek-v4 snapshots for the inline config payload
Squashed from:
- Snapshot the launchers that build their own command line
- Record what the self-executing launchers submit today
- Freeze the pid these launchers embed in their cleanup command
…ures

Squashed from:
- Let a p2p profile's rotary_base reach the model script it configures
- Test the model args command run.py actually builds, not a copy of its logic
- Flip the p2p snapshots to the rotary base each profile declares
The next ops rewrite all 62 scripts/models/*.sh into python. Once the shell
versions are gone there is no source of truth left to prove the rewrite was
faithful, so record the argv each of them expands to now. What these golden
files pin is agreement with the shell era, not merely agreement with today's
behaviour; the rewrite may only change the producer, never these files.

They also give the 18 models that no launcher snapshot reaches their first
coverage of any kind.
Squashed from:
- Expand the model args in python before building the command
- Update the launcher snapshots for the inlined model args
- Point the command_utils tests at the expanded model args
- Freeze the model-args knobs the snapshots now depend on
- Skip non-files when scanning the model scripts for environment knobs
Squashed from:
- Replace the model config shell scripts with python
- Point the run_megatron CLI tests at the model args loader
- Preserve the rotary base override in the 16-node profile launcher
- Convert the shell model configs to python and load them through one CLI
- Apply pre-commit formatting
- Keep the environment overrides and the failure path the sourced scripts had
- Point the NPU docker patch at the python model definitions
- Read each model's environment override where the shell script read it
- Match the shell mask when the model is shorter than its dense prefix
- Require keyword arguments for the moe layer frequency
- Take the model args overrides from the environment the shell already used
- Run the model args loader itself instead of a script that only forwards to it
- Let the model scripts be plain lists and read their environment at call time
- Group the model args utilities by what calls them
- Regenerate the shell launcher snapshots for the model args entry point
- Expect the TypeError an unknown model args keyword now raises
- Drop the override argument the restored environment reading made redundant
- Let a model script declare its arguments as one block of text
- Concatenate the model argument lines instead of parsing them
- Fix the callers that still joined the model args, and pin the contract
- Repair the two paths the model args conversion left behind
- Treat an explicit zero override as a value, not as a missing argument
- Let a model script reach the loader without importing the miles package
- Take the golden model args from the python loader instead of the shell
- Record the model args lookup in the two concurrent launchers' snapshots
- Point the p2p launcher at the loader and let it fail loudly
- Regenerate the self-executing launcher snapshots for the python model args
@fzyzcjy
fzyzcjy force-pushed the tom/refactor-miles/op8-12 branch from ac2d451 to 08b5d3b Compare August 9, 2026 10:00
@fzyzcjy
fzyzcjy force-pushed the tom/refactor-miles/op8-11 branch from fbefce1 to c77d60a Compare August 9, 2026 10:00
Base automatically changed from tom/refactor-miles/op8-12 to main August 9, 2026 10:48
@fzyzcjy
fzyzcjy merged commit fe30593 into main Aug 9, 2026
10 of 12 checks passed
@fzyzcjy
fzyzcjy deleted the tom/refactor-miles/op8-11 branch August 9, 2026 10:51
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants