Repository navigation
Conversation
…h (operator ruling 2026-09-30; never merge) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…nch (never merge) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…eged staging) (never merge) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ion (fleet edge doesn't reach the Sparks); restore /var/lib paths (never merge) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ght (never merge) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…on a host failure (never merge) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…he toolkit hook) (never merge) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ars (never merge) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…eckpoint population; vLLM refuses /model without it (never merge; real fix is an upstream row) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…readback (image-baked env, unreadable mounts); image/rank/transport/front door still gate (never merge) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…checkpoint loads (never merge) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ait (never merge) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…evice="cpu" (vLLM builds the model under a CUDA default device, so pin_memory raised 'Only dense CPU tensors can be pinned' on every rank) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…up A fences srv5-8 in the store) (never merge) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…(run 36790534801, config ba4a953f) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…nsors; pin them to device=cpu (model init runs under a CUDA default device; every rank raised 'Expected all tensors to be on the same device' at 00:57Z, dev log read run 36803955062) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…s when the vLLM engine fails to start (never merge) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… and a text-only launch Attempt 13 died on all four ranks in CUDA-graph capture with "SM120 sparse-MLA has no decode kernel for this shape: num_tokens=8, num_heads=16, topk=1152, page_block_size=32". Two independent refusals sit behind it, and neither is a combined window+indexer list (vLLM's SM120 class already passes the window and the indexer picks as separate segments): - page_block_size=32: vllm/models/deepseek_v4_1/attention.py constructs the SWA cache with the literal block_size=32; FlashInfer's SM120 DSv4 decode kernel takes 64-token pages only, and every call with <= 64 query tokens must take it. No engine argument reaches the literal, so it is an overlay member: swa_block_size class attribute, 64 on the SM120 FlashInfer class. - topk=1152: the vision variant widens prefill SWA rows to window + vision_max_n_token. The launch now serves text only: --language-model-only --hf-overrides.vision_n_layers 0. The vision capability is dropped on purpose. The self-test member creates vllm.models.deepseek_v4_1.nvidia.sm120_decode_shape_selftest. Probe run 36857198484 (srv6/7/8, image sha256:0f6039eb...): unpatched FAILS in both modalities, patched+multimodal FAILS at 1152, patched+text-only PASSES all four shapes. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ove the claim-skip note out of the body (never merge) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 084201d135
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| let occupancy = map(hosts, h => V41HostOccupancy { | ||
| // DEV-ONLY (operator ruling 2026-09-30, never merged): stop the V4 pair units on every Group A host | ||
| // before occupancy is read, so the V4.1 launch is not refused by a stale V4 unit. Not restored. | ||
| let dev_vacate_failures = flat_map(hosts, h => flat_map(spark_serving_units(), u => unit_leg_failure(leg: host_operation_leg(run: spark_ctx_argv_run(ctx: vllm_image_target_ctx(source_ctx: ctx, host: h)), operation: SystemctlStop { unit: u }, accepted_exits: [0, 5])))) |
There was a problem hiding this comment.
Restore stopped incumbent units when launch admission fails
This stops every pair-serving unit on all four Group A hosts before image, registry, authority, and occupancy admission has completed, and there is no restoration path. If any subsequent check, staging operation, or transactional apply fails, the function returns with the incumbent V4 service still stopped despite reporting that no serving unit was touched. Use the existing bracketed vacate/restore mechanism or perform the stop only inside an authorized transaction that restores on every failure path.
Useful? React with 👍 / 👎.
| // DEV-ONLY (operator ruling 2026-09-30, branch deploy/ds41-dev-launch, never merged): the Group A | ||
| // authority conjunct is skipped for a one-off dev launch on srv5-8; hosts-held and occupancy still apply. | ||
| [] as List<String>, |
There was a problem hiding this comment.
Restore the Group A successor-authorization gate
Replacing v41_authority_causes(...) with an empty list admits the launch even when Group A remains PairServingActive, the suspension lease has expired, or the suspension names a different successor. Since the same live path now stops the incumbent units and observes the hosts as vacant, an ordinary launch dispatch can replace the active deployment without the authorization that previously guarded the cutover.
Useful? React with 👍 / 👎.
| // DEV-ONLY (operator ruling 2026-09-30, never merged): wait up to 60 min for the head's front door while the | ||
| // four ranks load the checkpoint; the modeled check asks once, immediately after the ranks join. | ||
| "for i in $(seq 1 180); do curl -fsS --max-time 30 http://127.0.0.1:", port_text, "/v1/models > ", spark_native_readback_dir, "/models.json 2>/dev/null && break; if grep -qE 'Engine core initialization failed|EngineCore failed to start' /var/log/gunbc/v41-tp4-head.log 2>/dev/null; then echo DEV-ENGINE-FAILED; grep -nE 'Error|Traceback' /var/log/gunbc/v41-tp4-head.log | grep -v 'File ' | head -n 30; break; fi; sleep 20; done\n", |
There was a problem hiding this comment.
Keep the front-door wait below the workflow timeout
When the API never starts for the existing Group B native arm, the V4.1-specific logfile condition cannot break this loop, so even with immediately failing curls the 180 sleeps consume the full 60 minutes. The spark_native_serving_apply step in .github/workflows/fleet-converge.yml is itself limited to 60 minutes (line 1381), and this readback occurs after actuation, so the runner can terminate the job before the final refusal reaches the transaction's rollback logic, leaving the arm changed and fenced.
Useful? React with 👍 / 👎.
| // DEV-ONLY (operator ruling 2026-09-30, never merged): the container-field comparison counts image-baked | ||
| // env as unplanned and cannot read mounts, so it is skipped; image, rank, transport and front door still gate. | ||
| if false && (count(disagreeing) > 0 || count(unclaimed) > 0) { |
There was a problem hiding this comment.
Reinstate complete container readback validation
This condition is in the shared spark_native_rank_readback, not a V4.1-only path, so forcing it false also affects the production Group B native apply. A container with a differing or unreadable command, network mode, shared-memory size, GPU setting, restart policy, or planned environment can now pass the remaining image/rank/transport checks and cause the transaction to commit a realization other than the one it rendered.
Useful? React with 👍 / 👎.
|
Closed without folding in the v1 closeout bankruptcy (#13641). Model-eval scratch ('DS4.1'), stale since 2026-10-01. Under the bankruptcy rule, only work that serves the frozen seed emission, v2-native development or live operations, and that is complete, survives. The branch is kept for archaeology; no follow-up obligation is created. — sent from neat-wolf-604 |
Auto-opened by session-dashboard for session
proud-deer-538.Pushing to
deploy/ds41-dev-launchadvances this PR.Worker attestation
Before flipping this PR to ready for review, confirm each item:
npm test,cargo test) and the result.Closes #Ndirective.Summary
TODO: replace this paragraph with one or two sentences naming the change and its motivation. Reviewers read this first.
Test plan