Skip to content

miles: submission-order execution + CP-safe parallelism default; seam gate harnesses - #18

Merged
GavinZhu-GMI merged 3 commits into
mainfrom
fix/seam-parity-20260730
Jul 30, 2026
Merged

GavinZhu-GMI merged 3 commits into
mainfrom
fix/seam-parity-20260730

Conversation

@GavinZhu-GMI

@GavinZhu-GMI GavinZhu-GMI commented Jul 30, 2026 •

Copy link
Copy Markdown
Contributor

Fixes surfaced by re-validating the miles seam on the miles_tinker_seam:20260730 base (upstream multi-LoRA era). All three gates green after these changes: G1 grad-accum split-invariance bit-identical (grad_norm 178.9398358212894 for 1xfb(8) vs 2xfb(4)+step, and across reruns), G2 client-order exact (8/128 datums, incl. pipelined submission), G3 sl_basic 30-step SFT NLL 2.80->2.28 on Qwen2.5-0.5B r32.

1. dc470bd — serialize per-model backend ops. The task manager runs request handlers as concurrent asyncio tasks, so pipelined fb/optim_step broadcasts executed overlapped on one TrainGroup: optim_step consumed a later fb's gradients, and concurrent fan-outs reach the DP actors' mailboxes in inconsistent per-rank order, mispairing collectives (observed: per-sample logprobs scrambled across requests). Per-handle FIFO asyncio.Lock restores the client's submission-order contract; delete_model holds it too (delete-during-optim_step crash class).

2. 3f72f75 — engage CP only above 8K max_seq_len. The cookbook passes max_seq_len=8192 at create, which flipped auto-parallelism to CP=2; at CP>1 the seam returns per-CP-rank logprob chunks (each exactly pad-to-CP-multiple/2 long) because there is no full-sequence reassembly yet — clients see half-length per-sample logprobs. 8K fits without CP for supported model sizes; threshold raised until the seam stitches CP chunks (tracked as an open gap).

3. 6b3e1d9 — durable gate harnesses + image default. G1/G2/G3 harnesses land in scripts/gates/ (previously pod-local /tmp, lost on every pod recreate); miles profile default image bumped to miles_tinker_seam:20260730.

Companion miles-side fix (upstream get_rollout_data tuple return): GavinZhu-GMI/miles#4.

GavinZhu-GMI and others added 3 commits July 30, 2026 09:57
The task manager runs request handlers as concurrent asyncio tasks, so
pipelined fb/optim_step broadcasts interleaved on the shared TrainGroup:
optim_step could consume a later fb's gradients, and inconsistent
per-actor mailbox order mispairs DP collectives. Per-handle FIFO lock
restores the client's submission-order contract; delete_model holds it
too (delete-during-optim_step crash class).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
At CP>1 the miles seam returns per-CP-rank logprob chunks (no
full-sequence reassembly yet), so clients see half-length per-sample
logprobs. 8K fits without CP for supported model sizes; raise the
auto-parallelism CP threshold until the seam stitches CP chunks.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Durable copies of the G1/G2/G3 gate scripts (previously pod-local /tmp,
lost on pod recreate). Miles profile default image bumped to the
20260730 multi-LoRA-era bake.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@GavinZhu-GMI
GavinZhu-GMI merged commit dfcfe07 into main Jul 30, 2026
@GavinZhu-GMI
GavinZhu-GMI deleted the fix/seam-parity-20260730 branch September 16, 2026 02:59
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant