Skip to content

fix(patches): integrate graph-capture corruption guard on current main - #121

Merged
plotarmordev merged 1 commit into
mainfrom
plotarmordev/pr116-current-main
Aug 23, 2026
Merged

plotarmordev merged 1 commit into
mainfrom
plotarmordev/pr116-current-main

Conversation

@plotarmordev

Copy link
Copy Markdown
Collaborator

Summary

Current-main integration of @de1tydev's cross-repository PR #116. The original author and commit attribution are preserved; this branch exists because #116 predates merged PRs #112/#113 and cannot be updated in its fork without rewriting contributor history.

  • remove the obsolete #50004 adaptive-topk backport from the pinned pre-#50004 image
  • add upstream vLLM #52492's CUDA-capture guard to the #49486 shortcut
  • preserve current main's HEADLESS healthcheck/CI gate and scripts/loop_detector.py
  • record upstream #52823's later capture-safe re-land so the provenance statement remains current
  • remove the obsolete #50004-only benchmark field and clarify the remaining counterfactual buffer comment

Supersedes #116 for merge. Closes the integration blocker found in independent review; no changes to the contributor fork.

Verification

  • bash scripts/ci-validate.sh — PASS at 7f5538b90d5932f2fe78dc261a776ae6737d8c3d
  • scripts/test-hotfix-atomic-transaction.py — 23/23 PASS
  • all changed shell parsed; Compose rendered
  • primary-source audit: #52492 guard byte-identical; #50004 removal is the inverse of upstream #51318 for this pinned image
  • per-commit secret/source-identifier scan — 0 findings

Live two-rank recreate verification is in progress and will be posted before merge.

Cherry-pick of PR116 (c0aa088) onto current main 01df4ee (#112/#113
included); no main-side files or guards reverted.

Upstream vLLM traced intermittent silent output corruption on 2x DGX
Spark (MTP + prefix caching + CUDA graphs) to code this fork carries as
perf backports:

- #51318 reverted #50004: the C128A metadata builder writes packed rows
  at the live batch's stride while FULL-graph consumers keep the
  capture-time stride, so rows >= 1 read stale slot ids and attention
  lands on the wrong context slices. The 0.1.1 image's stock code is the
  exact pre-#50004 state, so removing the backport (compose chain,
  launcher status echo, worker sync list, test inventory) restores
  upstream's post-revert state.
- #52492 kept #49486 but barred the short-context indexer shortcut
  during stream capture: a graph captured with the shortcut baked in
  replays it against longer cached prefixes and returns candidates
  0..topk-1 unscored. The port now carries the guard verbatim and
  --status reports it; eager-step behavior is unchanged.
- Historical note: upstream re-landed adaptive top-k width capture-safely
  in #52823 on 2026-08-21; this repository still removes the obsolete
  #50004 backport because its pinned image's stock code predates #50004.

Review follow-ups folded in:
- scripts/bench-baseline-issue22-only.sh: drop the stale #50004 output
  field (and its active_topk_width probe), permanently 0 after removal.
- patches/hotfix-dsv4-mtp-buffer-50312.sh: mark the histogram width
  estimate as explicitly counterfactual post-revert (hypothetical
  headroom, not an active saving).

The third fix from the same investigation (#52836, reverting #49236's
eager scratch pool) does not apply: this fork never backported #49236.

Verification: python3 scripts/test-hotfix-atomic-transaction.py 23/23 OK;
bash -n clean on all touched shell scripts; compose YAML parses;
loop_detector.py and the #112 HEADLESS healthcheck/CI gate intact.
@plotarmordev

Copy link
Copy Markdown
Collaborator Author

Independent live validation at exact head 7f5538b90d5932f2fe78dc261a776ae6737d8c3d on the two-rank DGX Spark pair:

  • stopped the active project and recreated both candidate containers from an isolated recipe/worker directory (not a container restart)
  • head runtime: active_topk_width=0; one [PORT #52492] marker; hotfix-dsv4-skip-topk-49486.sh --status reports import/kernel/#49486/#52492 all APPLIED
  • worker runtime: identical active_topk_width=0, capture marker count 1, and all four status checks APPLIED
  • 4-way live chat smoke: 4/4 passed
  • RULER-lite at 8K/32K × {S-NIAH, variable tracking}: 4/4 passed in 55 s
  • served model remained deepseek-v4-flash-abliterated, max length 1,048,576

This validates the required container-recreate semantics and both-rank patch state. It does not claim a short run can prove the absence of an hours-later corruption event; the mechanism/source fix and exact runtime state are the evidence for that risk.

@plotarmordev

Copy link
Copy Markdown
Collaborator Author

Final combined-candidate evidence (e3d4bad7450c30ed9b5507a5f56aaf5d52121ec3, tree e26b6b2281ded34d9a7b247233e59ec362d3274e): both TP ranks report #49486 early return and #52492 capture guard APPLIED; #50004 is absent from the synchronized boot chain. Fresh-cache boot, C4 smoke 4/4, and repeated decode sentinels completed without target-kernel JIT, NCCL watchdog, or traceback lines after warmup. Independent CYBERGPT exact-head review remains APPROVE.

@plotarmordev
plotarmordev merged commit 2ff445b into main Aug 23, 2026
1 check passed
@plotarmordev
plotarmordev deleted the plotarmordev/pr116-current-main branch August 23, 2026 06:52
redcatH pushed a commit to redcatH/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark that referenced this pull request Aug 31, 2026
…rrent-main

fix(patches): integrate graph-capture corruption guard on current main
1890peachypeachy pushed a commit to 1890peachypeachy/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark that referenced this pull request Sep 6, 2026
…rrent-main

fix(patches): integrate graph-capture corruption guard on current main
SethHanford pushed a commit to SethHanford/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark that referenced this pull request Oct 7, 2026
…rrent-main

fix(patches): integrate graph-capture corruption guard on current main
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants