Skip to content

Fix context length - #757

Merged
hnyls2002 merged 1 commit into
mainfrom
fix-ctx-len
Jul 27, 2024
Merged

hnyls2002 merged 1 commit into
mainfrom
fix-ctx-len

Conversation

@hnyls2002

@hnyls2002 hnyls2002 commented Jul 27, 2024

Copy link
Copy Markdown
Collaborator

Fixed wrong context length for deepseek v2

@hnyls2002
hnyls2002 merged commit d9fccfe into main Jul 27, 2024
@hnyls2002
hnyls2002 deleted the fix-ctx-len branch July 27, 2024 01:13
timethink pushed a commit to timethink/sglang that referenced this pull request Mar 9, 2025
efschu added a commit to efschu/htsglang that referenced this pull request Aug 18, 2026
…ce in one harvest plan, with a proven runner

Consolidated from the comp4 gate run (progress.662-F4-r5),
WINDOW_TICKET_745/755, NOTE_747 par.8-9, NOTE_738/755, TICKET_727 and
the operator ledger. Two hard incompatibilities shape the plan:
WINDOW_TICKET_745 Arm 1 excludes the checkpoint interval that sgl-project#758's
anchor-cadence observable requires (-> ARM I hicache harvest, ARM II =
one flag more), and sgl-project#713's TTFT<3s needs a quiet router while every
other loaded gate needs the soak backlog (-> sgl-project#713 is the idle
sub-phase BEFORE the backlog, not a separate boot; the 06:44Z soak
driver is the load source per the Lastprobe rule).

ARM I phases: load-time (sgl-project#738 no-99G-plateau, file-backed-image
reclaim), idle-quiet (sgl-project#713, health, corridor), loaded (Gates A/B/C,
- sgl-project#757 race-holds, sgl-project#748 all three shapes, sgl-project#744/sgl-project#717 rung-funded flip,
- sgl-project#690 refill census, sgl-project#758-2 mamba host resume, WT_745's three lines,
corridor minima), teardown (image reclaim). ARM II adds
--mamba-checkpoint-interval 8192 for sgl-project#758-1 anchor cadence + NOTE_747
par.8.1-8.3. SEPARATE windows named with reasons: sgl-project#727 four-boot A/B,
WT_755 slots A/B (pool-geometry confound), sgl-project#755 metal retraction
(mechanism not built -- nothing to measure), sgl-project#709, sgl-project#735 Step-2. sgl-project#602
and sgl-project#536/sgl-project#537 carried as HONEST unresolved slots (owner-held detail /
not found with acceptance shape) rather than invented readouts.

Runner run_window_ladder.sh: PASS/FAIL/UNOBS table from boot log +
live server; never boots, never kills, never touches the soak driver;
soak-tolerant by construction (loaded checks are log observations, the
one latency check runs only in --phase idle, and choosing that phase
IS the operator's quiet-router assertion). Missing-emitter cases
(anchor cadence) report UNOBS, never FAIL -- absence of an instrument
is not absence of the property.

Mock-smoked per the desk rule, both directions: a fixture built from
the comp4 specimen lines reproduces the real run's verdicts exactly
(GATE-C crash, both sgl-project#748 shapes + vacuous relief, sgl-project#757 sentence = 5
FAIL; sgl-project#744/sgl-project#690 PASS; exit 1) and a clean fixture goes fully green
incl. the ARM II cadence line (exit 0). bash -n clean, codespell
clean. Nothing was booted.
efschu added a commit to efschu/htsglang that referenced this pull request Aug 18, 2026
…gl-project#754 retires into the 753 fold, 13-step plan re-smoked clean

The first map froze before today's second wave; new authority is
F4-r5's harvest composite 59ce2d8 (declared COMPLETE, review tip
still an ancestor). Pass 3 redefined: harvest tip = train base,
unabsorbed branches cherry-pick on top, feat/753 lands FIRST by its
owner (10 in-flight commits; carries sgl-project#749; folds sgl-project#754 at the same
seam -- distributed/utils.py:1709, its own pp_size=1 handling).

Sweep results, same git-cherry/merge-base rigor as the first map,
outputs quoted in the ledger REFRESH section:
- ABSORBED by ancestry: comp4 and its whole lineage, 915ce1b
  (F4-r5's own sgl-project#757), 57b04b2 (sgl-project#540 fix).
- ABSORBED as different commits (desk-sgl-project#752 hazard class, never merge):
  fix/748-armed-gate-scope, fix/759-arming-economy,
  feat/755-slot-reorder. The sgl-project#758 emitters need no branch -- the
  harvest TIP ITSELF is a sgl-project#758 commit.
- SUPERSEDED: my own fix/754 -- semantic-not-patch folded by 753
  (git cherry vs 082293f shows '+'); merging it after 753 lands
  guarantees a get_pp_layer_set conflict with zero gain. Retired from
  the plan without regret.
- REVIEW-never-merge: 9e56477 (independent sgl-project#757), per its
  reviewer's own in-composite note naming 915ce1b as the baseline.
- fix/706-remainder not yet visible; slot reserved.

Executor updated: HARVEST constant joins the lineage check,
DEFAULT_TIP moves to the harvest tip, the sgl-project#754 step is replaced by
the sgl-project#745 reachability pick, and the second-wave picks join (727
head-chain, sgl-project#738 verdict, sgl-project#535 tickets). Dry-run scratch smoke against
the REAL harvest tip: all 13 steps complete with ZERO conflicts
(exit 0) -- cleaner than the first wave; the sgl-project#740 pair ordering from
the previous smoke holds. Plan-mirror test updated (12 passed). Gate
unchanged: COMP4_ACCEPTED still required, nothing pushes, nothing
booted.
efschu added a commit to efschu/htsglang that referenced this pull request Aug 18, 2026
…ENT, loud review item -- and the harvest tip moved to da81871

The double-sgl-project#754 template does NOT apply. Both sgl-project#757 fixes attack the
same root (corpse-S drain disabled, rank-local disarm routes) at
DIFFERENT intervention points: 915ce1b (in harvest, regression-
verified) drains the leftover at DISARM; 9e56477 (Slot-3) re-enables
the ARMED drain with a pure 4-way demultiplex classifier and states a
liveness property the disarm-time form lacks (the upstream's blocking
commit waits on the wire being consumed AT ARRIVAL). Review question
stated sharply in the ledger: does a long armed window stall the group
under the disarm-time form? If yes, the two are COMPLEMENTARY halves,
not duplicates. Slot-3's test suite (3-process gloo repro, 4 killed
mutants, sgl-project#631-pin correction) binds to their classifier, so the tests
ride the review verdict -- not cherry-pickable bare. 9e56477 stays
OUT of the executor plan; the review joins feat/753 as a NAMED
precondition of pass 3.

Tip correction folded in: authority is now da81871 (contains the
- sgl-project#758 phase-tag 59ce2d8 AND caca352 -- F4-r5 absorbed
WINDOW_LADDER_0818 + its runner into the harvest). Executor HARVEST
constant bumped; dry-run scratch smoke re-run against the REAL new
tip: all 13 steps complete, zero conflicts, exit 0; plan-mirror tests
12 passed.
efschu added a commit to efschu/htsglang that referenced this pull request Aug 18, 2026
…fold shipped, executor gains the two steps

The review item the refresh flagged is closed with a measurement, not
an argument: gloo isend posts instantly but wait() blocks until the
peer recv (3000.3 ms stall for a 3 s armed window on a 4 KiB proxy,
size-independent -- tools/probe_757_gloo_liveness.py), and the in-tree
sender is exactly that shape. The disarm-time form alone stalls an
abandoned upstream at its commit for the rest of the armed window;
verdict (b), COMPLEMENTARY. Fold shipped as fix/757-armed-liveness
(194c3ea + probe 5e2c121): both suites + corrected sgl-project#631 pin
green together, managers selection 633 vs 625 passed with identical 2
pre-existing failures, 0 new. 9e56477 superseded by the fold.

Executor: two steps added (the fold pick + the probe pick), pass-3
precondition list drops the review and keeps only feat/753. Dry-run
scratch smoke re-run against the real harvest tip: 15 steps, zero
conflicts, exit 0; plan-mirror 12 passed.
efschu pushed a commit to efschu/htsglang that referenced this pull request Aug 20, 2026
…ion, not before the exchange

A boot reached serving, one admissible queued request was never served, and the
process died. The wedge detector reported "1 queued, 0 running, NO first token
for 340-370s and no prefill chunk either" and named its own gap in the same
line -- "no phase-policy corroboration seen -- the wedge class is broader than
that path". It was right: this is not an admission defect at all.

EVIDENCE OF RECORD. py-spy on the specimen (WEDGE_788_specimen.log:732/915/1083)
catches all three ranks in a circular wait:

  PP0, PP1  gloo waitSend   _pp_commit_comm_work (scheduler_pp_mixin.py:2465)
                            from _pp_forward_and_process_input_requests:1007
  PP2       gloo waitRecv   _pp_recv_typed_dict:2709
                            from _pp_recv_proxy_tensors:2803

The request chain was flushed at the TOP of the pass, before the rank had
posted anything else it owed its peers that iteration -- the proxy tensor-dict
send and the output-ring send both come later in _event_loop_pp_body. A
downstream whose progress needs one of those was therefore waiting on a rank
that was itself waiting on that downstream. The chain recv is posted only at
the top of a pass (:420) and nowhere else in the loop body, so a rank parked in
the proxy receive cannot drain the message its upstream is blocked flushing.

THE FIX IS ONE ADDITION. _pp_forward_and_process_input_requests is byte-identical
to before: it still commits send_req_work, then posts the async forward, then
runs process_input_requests. What is new is _pp_commit_pending_req_work, called
once per iteration from _event_loop_pp_body after the proxy isend and the output
commit and BEFORE the phase-flip round hook. By the time that loop returns to
the top-of-pass commit, the handle has already been waited on and cleared, so
the old call site is a no-op there -- and the two disaggregation loops (:618,
:765), which do not call the new method, keep their previous behaviour exactly.

Deleting rather than moving the old call site is deliberate on both counts: it
keeps the sgl-project#633 ordering contract literally intact (commit, then forward, then
process, with send_req_work holding this pass's handle on return -- pinned by
test_scheduler_pp_request_order_633), and it avoids a second staging slot. An
earlier shape of this fix did stage the send separately; it broke that contract
and left the two disagg loops overwriting a live P2PWork handle every pass.

Precedent: a7ff250 "[sgl-project#753] Flush the output sends after the exchange, not
before it" made the same move for the OUTPUT channel. This is that fix for the
REQUEST channel, ungated -- the hazard is not gapped-specific.

NEGATIVE FINDING, RECORDED RATHER THAN BURIED. A deterministic 3-process gloo
reproduction of the specimen deadlock was attempted and NOT achieved under a
faithful model of the loop. Two candidate mechanisms were tested and refuted:

  - Eager-send asymmetry. Hypothesis: an empty request list puts one 8-byte
    tensor on the wire and completes eagerly, so only a large first request can
    block. MEASURED FALSE: a 2-process probe with a 2.0s-delayed matching recv
    blocked for the full delay at every payload size from 8 B to 1 MiB in this
    build. No eager completion at any size.
  - Schedule asymmetry. A linear chain+proxy relay that flushes the old handle
    before dispatching the new one is deadlock-free by induction (each rank's
    flush needs only its immediate downstream one pass behind), and dozens of
    real 3-process runs across depths and offset combinations all completed.

So the reduced two-channel model cannot close a cycle. The production graph has
an edge that model lacks: the output ring closes last-rank -> rank 0
(_pp_send_recv_and_preprocess_output_tensors:3038, last-rank branch :3012-3014,
via next_first_rank_mb_id at :416/:608/:754). That, and real GPU compute skew,
are the two named unconfirmed candidates for the missing ingredient. Neither is
built into the test, and the gap is stated in its docstring rather than papered
over. The mechanism's evidence of record therefore remains the py-spy trace
above plus the sgl-project#753 precedent, and the survival boot is the integration proof.

RECOVERY RUNG. The detector gains one bounded action: after a wedge stays
continuously alarming past a threshold well above the report threshold (default
3x, env SGLANG_ADMISSION_WEDGE_RECOVERY_SECONDS; non-positive values fall back
to the default rather than firing every poll), it makes ONE forced-admission
attempt per episode through the existing corridor_admission actuator and logs it
loudly either way. Its own docstring is explicit that it relieves VRAM pressure
at the admission site and will NOT move a comms deadlock like this one, and that
calling an actuator from the watchdog thread is a new cross-thread shape over
CUDA allocator state.

TESTS
  test_pp_chain_flush_deadlock_788.py (new; 3 real gloo processes, CPU only)
    - load-bearing: runtime ordering pin. The SHIPPED _pp_commit_pending_req_work
      is wrapped and the observed sequence asserted to be
      proxy_send -> chain_flush -> round_hook for every pass and every non-last
      rank, so the flush cannot drift back past the collective and cost it the
      "last blocking op of the iteration" property its own comment relies on.
    - can-fail: the same run with the flush relocated to the top records a
      different sequence, so the pin discriminates placement.
    - the honest negative case described above, in place of the deleted
      hang-repro. A test that is green or red for the wrong reason is worse
      than no test.
  test_pp_flip_leftover_proxy_757.py: harness-only repair, no assertion touched.
    It was silently DEAD on this branch -- its _GlooWire predated
    rank_in_group/world_size and the src positional recv_typed_tensor_dict now
    passes, so it errored 3/3 before reaching an assertion. Now 3 passed and a
    working regression check again.

  Verified green with this diff: test_pp_chain_flush_deadlock_788 and
  test_scheduler_pp_request_order_633 (8 passed), test_phase_policy (90 passed),
  test_pp_flip_leftover_proxy_757 (3 passed).

PRE-EXISTING FAILURES, MEASURED NOT ASSUMED. The managers slice reads 43F/2707P
here against 39F/2705P at a pristine detached ff5651a worktree. Every delta
is accounted for: -3 (the 757 repair above now passes), +5 stub drift and +1
ordering-contract break, both introduced by the earlier staging-slot shape and
both gone with it, +1 test_pp_drain_completeness_787 which is INTENTIONALLY RED
here -- that file's fix lives on fix/787-drain-completeness and it flips green
at the merge, which makes it a free integration can-fail. The residual 39 are
the same interface-drift family (rank_in_group, _pp_gapped_wire) plus one stale
constant, all present at base with no part of this diff.

DEBTS
  sgl-project#789                the two disaggregation PP loops keep the top-of-pass
                      flush and therefore still carry this hazard
  sgl-project#757-harness-drift  repaired here; the same drift still blinds
                      test_pp_proxy_stamp_631 and test_pp_slot_last_batch_631
  sgl-project#753-stale-constant 10.923 != 11.923 in the gapped entry protocol test
  sgl-project#766-pointer-stale  the register points sgl-project#766 at ARM_defaultfull7.log, which
                      is a 26-line host-ledger snapshot with no
                      Bar1CollectiveAborted and no proxy/drain lines
efschu pushed a commit to efschu/htsglang that referenced this pull request Aug 21, 2026
…eriving it (slice 1)

Ten fixes on this branch (sgl-project#757 sgl-project#789 sgl-project#790 #791b #791c sgl-project#792 sgl-project#795 sgl-project#797 #797b #797c) all have
ONE form: each rank RE-DERIVES the pass schedule -- rid set, chunk length, prefix length --
locally from its own state, and a phase flip invalidates that state non-atomically. Every
"new root" was the next consumer of the same re-derivation, so the list grew instead of
converging. This stops patching consumers.

THE DATUM WAS ALREADY ON THE WIRE AND NOBODY READ IT. PPAdmissionEntry.extend_len
(pp_admission_congruence.py:177) has crossed the wire since sgl-project#791's first commit with exactly
three consumers: to_wire (:223), from_wire (:244), one log string (:824). Nothing ever built
a batch from it; reconcile_pp_admission_decision returns Dict[rid, prefix_len] and drops the
second number on the floor.

THE MECHANISM, CORRECTED AGAINST THE LOG rather than assumed. boot_instr20.log:5171,5181-5183:
  PP0 ADMIT rid=6cbe2733 prefix_lens=0   chunked=1  -> 512-row chunk
  PP1 ADMIT rid=6cbe2733 prefix_lens=512 chunked=0  -> 333-token remainder
PP1 DID receive prefix_len=0 and DID clamp prefix_indices to it -- scheduler.py:7026-7027
worked. Then add_one_req's HOST LOAD-BACK put the 512 back: needs_host_load_back() went true
when the HiCache prefetch landed and schedule_policy.py:1539-1549 concatenates the recovered
indices. "MAMBA-HOST-RESUME ... triggers load_back" appears on PP1 and PP2 and is ABSENT on
PP0 -- that asymmetry is the bug. 845-512=333 then fitted rem_chunk_tokens whole, so the
NON-chunked branch fired. The re-derivation lives INSIDE THE ADDER, after the schedule was
already applied, and it re-derived BOTH numbers.

DESIGN. forwarded_schedule() (pp_admission_congruence.py:500) is the pass geometry as a
value: rid -> (prefix_len, extend_len) for exactly the rids `effective` names; None or a
voided decision yields {}. _add_scheduled_req (schedule_policy.py:1237) EXECUTES both
numbers: no rem_chunk_tokens, no page/align rounding, no host load-back, no budget veto --
the budget is still charged. The gate sits above every local veto (:1600), and
add_chunked_req (:1328) gains the gate it never had at all (it is entered from
scheduler.py:7004, BEFORE the admission loop). The three membership vetoes that silently
narrowed -- batch_is_full (:7035), the HiCache prefetch_done skip (:7047, the instr20 race
itself) and the LoRA gate (:7024) -- become refusals (scheduler.py:7190).

REFUSAL IS CONTROL FLOW, NOT A RESULT CODE. PPScheduleRefused (:163) is an exception because
every AddReqResult means "build a batch without this request", which is precisely the
corruption. A refusal reuses sgl-project#797 end to end (_pp_refuse_forwarded_schedule, scheduler.py:6333
/:6369): sets _pp_admission_pass_voided, voids the forwarded decision, re-notes the slot
expectation. No new mechanism. Inside the loop a refusal is CARRIED, not thrown, so
alloc_group_end() still runs (:7025, :7118, :7169).

WHY THE GUARDS CAN NO LONGER FIRE, each owed a reason: sgl-project#631's _want is extend_num_tokens,
which now comes only from the forwarded extend_len with load-back suppressed, so it IS the
upstream's row count by construction rather than by agreement (green arm: rows=512,
batch_tokens=512). sgl-project#757/sgl-project#787 stamp and sgl-project#795 epoch were already structural; what changes is
that they can no longer be correct-but-insufficient, as instr17 and instr20 both were --
every identity right, only the width wrong. Width is now an identity too. sgl-project#789 needs a
membership divergence, which is now identical-or-refused. #791c's tripwire detects a
self-narrowed batch, and no path creates one.

HONEST CORRECTION TO THE SUBSUMPTION CLAIM: retraction does NOT become unnecessary. Physical
impossibility is real -- a rank genuinely lacking KV for [local, told) cannot execute. What
becomes structurally impossible is the NARROW-THEN-DETECT shape: no code path is left that
builds a batch of a geometry the upstream did not name.

NOT COVERED BY THIS SLICE, stated so nobody assumes otherwise:
 - BATCH ORDER. can_run_list follows the local waiting_queue order, the decision follows
   PP0's. Same rid set in a different order gives EQUAL WIDTHS and permuted rows -- silent.
   Covered by the full design (execute in decision order), not by this slice.
 - Decode batches: retract_decode (scheduler.py:7489/:7517) mutates long-lived state on a
   tp_cpu_group reduce that is world=1 under TP=1/PP=3. Different root, filed by sgl-project#797.
 - Radix eviction divergence, KV pool sizing, spec-decode draft schedules.

NAMED RESIDUAL, filed at the site (scheduler.py:7118) in sgl-project#797's practice: a request admitted
earlier in a loop that later refuses has taken a persistent inc_lock_ref, released on batch
completion. Undoing it needs the exact IncLockRefResult (SWA/Mamba tombstone params) the
adder does not keep, and a blind release makes the one thing a mismatched release worsens.
Bounded: reaching that line takes a genuinely unexecutable geometry and kills the pass.

DEFAULT PATH UNTOUCHED: _pp_scheduled_extents() returns None on PP0 and on every pp_size<=1
boot, so scheduled_extent_for returns None and both adders take the pre-existing arithmetic
unentered; PPScheduleRefused is unraisable there. Pinned by
test_no_mapping_is_the_untouched_default_path and corroborated by 56/56 on the
schedule_policy neighbours.

TESTS. test_pp_forwarded_schedule_791.py: 17 passed (3 live gloo arms + 14 pure), 92 s.
  test_red_without_the_forwarded_geometry_instr20_reappears -- can-fail, rebinding ONLY the
    fix's return value in the child: batch=(512,333) rows=512 mismatch=True, byte-identical
    to instr20 PP1 09:40:30
  test_green_the_forwarded_geometry_survives_the_mid_pass_prefetch -- batch=(0,512) rows=512
    mismatch=False, load_back_calls=0
  test_an_impossible_geometry_raises_rather_than_narrowing -- the architectural property
Neighbour set (791/791b/791c/797/631 x3) re-measured AT HEAD: 11 failed / 75 passed; after:
11 failed / 92 passed, failure names byte-identical, zero regressions, +17.
schedule_policy neighbours 56 passed / 0 failed. ruff 119 = 119 at HEAD (parity),
ruff format clean on everything authored, codespell identical.
efschu pushed a commit to efschu/htsglang that referenced this pull request Aug 22, 2026
…n a send nobody owes

The PP output ring wedged with all three ranks alive and none of them
able to move (specimen /spinning/evidence-665-f1/wedge_802f_1712/, PP=3,
--enable-phase-flip, py-spy of all three schedulers):

    PP0  _pp_commit_comm_work <- _pp_commit_pending_req_work  (:4071/:2262)
    PP1  _pp_recv_dict_from_prev_stage <- _do_recv            (:5608/:6069)
    PP2  PpChainReceiver.recv <- recv_requests   (pp_chain_receiver.py:329)

PP0 is flushing the request chain, which PP1 can only take at the top of
its next pass; PP1 never reaches that top because it is blocked in the
output receive; PP2 waits for the chain send PP1 has not reached. Three
arcs, one cycle, no timeout anywhere.

THE ASYMMETRY. The two ends of the intermediate hop apply unrelated
predicates. `_do_recv` decides to receive from THIS rank's own slot
state; the non-last sender in `_pp_send_output_to_next_stage` decides to
forward on `if pp_outputs:`, which is whatever it received LAST
iteration. The last-rank hop is matched by construction because it
consults `_pp_output_expected_for_slot` -- but that flag is the FIRST
rank's verdict, published for PP0's arc. Nothing publishes the same
thing for the intermediate hop.

WHY THE EXISTING VOID CONTRACT DOES NOT COVER IT. `pp_void_forward_
payload` (sgl-project#797) already forwards a void along this hop and stops at
`pp_first_retracting_rank`, on the argument that rank r and everything
after it has an empty slot whose receive early-returns. That holds for
the VOIDED slot -- `_pp_void_own_batch` empties it. It does not hold for
a slot the resident decode path still occupies, and both void paths keep
resident requests on purpose (`_pp_absorb_void_output` refuses to release
them as a double-free; `_pp_void_own_batch` deliberately leaves
`running_mbs` alone). The specimen shows exactly that: PP0's last act was
absorbing a void for slot 0, leaving `pp_outputs` None, while PP1 logged
`running=1 chunked=1` and re-entered the receive for a slot PP0 would
never send to. PP0 cannot know PP1's resident set, so closing this needs
a per-slot expectation every non-first rank publishes to its predecessor
-- a protocol extension, and not this commit.

WHAT THIS COMMIT DOES. The sgl-project#789 readiness gate is parameterised by wire
kind and the output receive now passes through it. The CHAN_DICT counter
was never proxy-specific -- one counter per wire, demultiplexed by
`__msg_type__` after it comes off -- so the gate reads the same true
statement about the same wire either way; `kind` selects only the
per-(src, kind) inbox peek. The ring is cut at its one cuttable arc: PP1
waits boundedly and refuses by name instead of for ever. It does not make
the missing send appear, and the error text says so.

AN ALIAS, NOT A WRAPPER. `_pp_wait_for_proxy_readiness` is now a
class-level alias for the same function object. About ten stand-in
holders across the sgl-project#631/sgl-project#757/sgl-project#787/sgl-project#789/sgl-project#791/sgl-project#795/sgl-project#797/sgl-project#798 test family
bind that name one method at a time; a delegating wrapper resolved the
second name on the HOLDER and turned 9 green tests into 5 failures.
Measured, then fixed.

Tests: test/registered/unit/managers/test_pp_output_readiness_ring_802.py
-- three real gloo processes, real PhaseFlipCounters, neutering done in
the child (spawn re-executes the module, not the test body).
  * red arm, gate neutered: stuck_ranks == [0, 1, 2], specimen reproduced
  * green arm: stuck == [], PP1 reports "sgl-project#789 OUTPUT READINESS TIMEOUT"
    and "sgl-project#802-ring", PP0 chain-flushed, PP2 requests-received
  * false-positive direction: upstream really posts -> receive succeeds
  * no-op without counters, so the non-phase-flip default path is unchanged
Mutants, both die: call site removed -> green arm red (all ranks stuck);
gate raises unconditionally -> false-positive test red.
Regression: baseline 787+789 = 9 passed / 0 failed; with this change
791b + 789 + 787 + 802 = 17 passed / 0 failed.
efschu pushed a commit to efschu/htsglang that referenced this pull request Aug 23, 2026
…he ring with the existing drain

W4a made the PP chain receive bounded, named and resumable, but left its
automatic abort off (SGLANG_PP_CHAIN_RECV_STALL_S=0), so on metal it could
diagnose nothing and recover nothing. This wires it.

THE DEFAULT STAYS OFF, AND THAT IS THE POINT. An idle PP rank legitimately
blocks in this receive until a request arrives, so there is no duration
that separates "idle" from "wedged"; a wall-clock default would SIGQUIT a
healthy idle server, the same trap sgl-project#821's marker documents for its own arm.
The recovery is therefore armed by EVIDENCE.

THE PREDICATE. bump_attempted publishes that a rank has ENTERED a send
BEFORE it posts it (phase_flip_counters.py:220-226) -- the only counter
whose timing can witness a peer parked INSIDE a send rather than one that
has finished. When this rank's CHAN_DICT upstream has entered more dict
sends than this rank has taken off that wire, the peer is parked in a send
only this rank can drain while this rank is parked in a receive that peer
will never feed. That is boot_827's ring stated in counters: PP0 in
_pp_commit_admission_send_work on the typed-dict channel, PP1 and PP2 in
the request-relay chain receive. abort_check is consulted only at the
yield, i.e. only once a receive is already overdue, so a healthy pass never
reaches it.

THE RING-CUT REUSES sgl-project#757 RATHER THAN INVENTING A CONSUMPTION PATH.
pp_flip_drain_leftover_dicts already demultiplexes, stashes a wrong-kind
message in _pp_tensor_dict_inbox where its real consumer looks, and
discards only a provably void proxy -- which is exactly what makes taking
the dict off the wire out of the pass's normal order safe. It runs on every
disarm route already. request_receiver catches PpChainRecvStalled, runs one
drain turn, and resumes the SAME posted receive, which is what ParkedWait's
resumability is for. Servicing that does not clear the stall re-raises
rather than spins: at that point the ring is closed for a reason this code
does not model, and retrying would turn a diagnosable wedge into an
invisible one.

Follows sgl-project#789's shape deliberately. _pp_wait_for_dict_readiness argues the
false-positive direction is the safe one for the mirror gate, and it is
here too: a spurious fire costs one drain turn and a resumed receive, and
the receive stays posted and framed throughout, so the late message still
arrives intact. Missing a real one costs the boot. sgl-project#789 also declines to
invent a new protocol for its case; this declines likewise and cuts the
ring on the arc waiting for a message nobody posted.

Also: _pp_flip_pass_tick publishes _pp_live_mb_id, which the drain needs to
tell an owed proxy from a leftover one.

TESTS (hermetic, CVD="", CPU only, no gloo, no scheduler construction)
  test/registered/unit/managers/test_pp_chain_abort_check_824.py 9 passed;
  all 9 fail pre-fix. Covers both directions: the boot_827 counter state
  aborts, and an idle rank is NOT aborted at 0 s or at 3600 s.

  Mutants killed, one per edge:
    entered >= taken           -> the idle rank is aborted (1 failed)
    service-turn cap raised    -> an unclearable stall stops being re-raised
    service hook never called  -> the ring is never cut (2 failed)

  Run under /spinning/htsglang-gpu/.venv (torch 2.11.0+cu130, datasets
  5.0.0) with PYTHONPATH leading to this worktree, verified by
  `import sglang; sglang.__file__`. My earlier runs used /usr/bin/python3,
  which lacks datasets and silently collected only part of the suite; see
  the corrected note in WINDOW-QUEUE.md.

No boot was run. This is desk work.
efschu pushed a commit to efschu/htsglang that referenced this pull request Aug 23, 2026
Wave 3, stage 3 of the batched-window tree. WINDOW-QUEUE tickets W4 (PP ring
wedges immediately after an abandoned flip) and W5 (sgl-project#821 did not instrument
the receive path that actually wedged), both preflight_pass=Y.

Four commits, base 500be7e -- which is EXACTLY this tree's stage-1 branch,
so this stage sits directly on top of W1/W2 with no divergence to reconcile:

  2efb933  [sgl-project#824 W4a] Bound the PP chain receive without tearing the
              stream it guards
  dd92de4  [sgl-project#824 W5] Name the arm that tripped the watchdog, and
              instrument the path that wedged
  430e468  [sgl-project#824 W4b] A zero-iteration armed window must not move the slot
  ceb79d6  [sgl-project#824 W4a] Arm the chain-recv recovery on STATE, and cut the
              ring with the existing drain

GATED AGAINST ceb79d6, NOT 430e468. The branch head moved while this
stage's first battery was already running, so that run was measuring a
superseded commit and was DISCARDED rather than reported -- a gate that
certifies a commit the tree does not carry is worse than no gate. The
in-flight battery was stopped by explicit PID (no pkill), py-spy dumped first
per standing rule and found healthy mid-test rather than wedged, and its
partial log kept aside as w3c_stale430.log. A foreign battery belonging to
another session was running on the box at the same time and was NOT touched.
The final W4a commit arms the recovery on STATE (`bump_attempted`, entered >=
taken as the evidence predicate) and cuts the ring through the existing sgl-project#757
drain, with the wall-clock default left at 0.

W4b is the ROOT of the pair, and it lands in the same mechanism as this
branch's own known-red neighbours: an armed window that ran ZERO slot
iterations must not advance mb_id. The boot_827 specimen recorded exactly that
-- "rank 0 ran 0 slot iteration(s) (armed at mb_id=0, disarmed at mb_id=1)" --
one line before the ring stopped dead and stayed silent for 31 s until the
health check noticed. That is the sgl-project#631 defect class with the spread reduced to
a single rank: arm on one slot, leave on another, then re-enter the pipeline
wherever you happened to stop.

TOUCH SURFACES, computed rather than assumed. This stage shares
scheduler_pp_mixin.py with both earlier wave-3 stages and with wave 2, and it
extends sgl-project#821's own suite (test_pp_wedge_watchdog_is_honest_821.py) rather than
duplicating it -- W5 is the continuation of sgl-project#821, by the same argument sgl-project#821
made, so touching that file is intended and not a collision. It also adds to
mem_cache/hicache_collective.py, the file carrying sgl-project#734's
`waited < timeout_s * 0.95` discriminator that strand 17c's 622 falsification
identified as the fragile part.

That last point is what makes this gate non-vacuous rather than merely green:
`test_pp_sync_rendezvous_630.py` -- the suite that went red under the 622 merge
for precisely this threshold, and the reason 622 is still off the line -- sits
INSIDE the battery. If this stage disturbed that discriminator, the gate would
say so in the same way it said so for 622.

NOTE ON A PATH, because it looks like a contradiction and is not: this stage
modifies python/sglang/srt/utils/watchdog.py, while commit 9f1af20 fixed a
catalog citation that resolved against python/sglang/srt/watchdog.py. Those are
different files; the catalog's intended target was turnkey/watchdog.py:88 (the
generation probe retired by user order), which is where it now points.

GATE. Battery test/registered/unit/{managers,planner,server_args,mem_cache},
hermetic under CUDA_VISIBLE_DEVICES="", one battery at a time.

    baseline (tip 4f2072a)   7 failed, 8585 passed, 1852 skipped, 887 s
    stage 1 (W1+W2)             7 failed, 8638 passed, 1852 skipped, 868 s
    stage 2 (W3)                7 failed, 8653 passed, 1852 skipped, 922 s
    this stage                  7 failed, 8673 passed, 1852 skipped, 915 s
    NEW failure ids             NONE (comm against the baseline list is empty)
    fixed vs baseline           0

THE +20 IS ACCOUNTED FOR EXACTLY, not assumed. Four new suites adding only 20
passes is low enough to be worth checking, because a suite that silently skips
looks identical to a suite that passes cheaply. Two facts settle it: the
skipped count is 1852 in every run of this train -- baseline, stage 1, stage 2
and this one -- so nothing new is being skipped; and running the four suites
standalone reports exactly 20 passed, matching the battery's delta to the
case. They run, and they are collected.

COVERAGE, and this stage is the reason to state it explicitly. The W4/W5
strand reported that its OWN worktree cannot collect the managers suite --
92 modules import `datasets`, which is not installed there. That limitation is
NOT present in this gate, verified in the positive direction rather than
inferred from a missing error line: datasets 5.0.0 is importable in this venv,
the battery logs carry zero collection errors and zero ModuleNotFoundError,
and an explicit --collect-only over test/registered/unit/managers reports 3436
tests collected. So this stage's no-regression claim rests on a fully
collected managers suite, which the strand's own local run could not provide.

Its four new suites land inside the battery directories and are therefore
actually collected: test_pp_ring_abort_recovery.py,
test_pp_wedge_watchdog_names_the_arm_824.py,
test_pp_zero_iteration_window_slot_824.py and
test_pp_chain_abort_check_824.py, all under managers.

ON PROVENANCE, if this stage ever shows a new failure id: the branch is
PRE-sgl-project#815 (its base 500be7e sits on 21ff075), so the last_batch-631
failures visible in the strand's own worktree are expected to DISAPPEAR on
this line, which carries fix/815. A new id here would therefore have to be
checked against that provenance before being attributed to the strand's
change. In this run none appeared, so the question stays hypothetical.

No boot. This tree is the input to ONE batched window.
efschu pushed a commit to efschu/htsglang that referenced this pull request Aug 24, 2026
…s on

sgl-project#829's root -- an arm slot outliving the ring it names -- is already closed on
integ/round5 by d507e19, and test_pp_arm_slot_outlives_ring_829.py pins the
behaviour with 9 green tests. This adds no runtime code and changes no source
line. It pins the PREMISES that make the closure a proof rather than an
observation.

THE ROOT, restated with file:line so the pins have a referent:

  :2606  rising edge records self._pp_flip_arm_mb_id = mb_id
  :2616  and self._pp_flip_arm_epoch alongside it
         cutover commits with _pp_flip_armed_passes == 0
  :3905  init_pp_loop_state REBUILDS mbs/last_mbs/running_mbs from zero, and
         pp_loop_size may CHANGE (TP phase runs pp_size=1), so a carried index
         can be out of range
  :2776  the falling edge jumps the pass loop to the recorded slot

Pre-sgl-project#829 that put PP0 on slot 2 of a fresh ring while its downstream sat on 0
(boot_window2_0823_1554, epoch 6, "RESUME SLOTS [2, 1, 1] -- DIVERGED"), and
PP1 raised sgl-project#631 PROXY LEFTOVER REFUSED one second later.

TWO GUARDS CLOSE IT, and they are COMPLEMENTARY rather than redundant:

  A  :3903  init_pp_loop_state calls pp_flip_forget_ring_scoped_slots(self)
            unconditionally, retiring all three ring-scoped carriers.
  B  :2755  the falling edge refuses a restore across a generation
            (ring_rebuilt = arm_epoch != now_epoch); the only jump is the
            elif arm at :2776.

init_pp_loop_state has THREE callers and only the cutover advances the epoch,
so guard B is blind to the other two rebuilds and guard A covers them. Neither
may be dropped.

WHAT WAS MISSING. The proof rests on three uniqueness facts that no test
pinned: the ring is rebuilt in exactly one place, the arm slot is recorded in
exactly one place, and the pass loop can be jumped from exactly one place. A
second rebuild site or a second jump writer would bypass both guards WITHOUT
failing any existing test, silently restoring the state that killed the boot.

Verified by exhaustive enumeration over python/sglang/: one rebuild site
(:3905-3910), one arm writer (:2606), one jump site (:2776).

test_pp_ring_rebuild_choke_point_829.py -- 5 AST-level uniqueness pins, ~2 s,
CPU only, no torch and no process group. A failure there is not automatically a
bug; it is a demand that a newly added site be checked against the epoch gate,
and the failure messages say so and name the offending function and line.

Deliberately NOT pinned: the order of the clear against the array assignments.
Both orders are correct and pinning one would make a harmless refactor red.

DANGER DIRECTION. Nothing here enforces uniformity and no runtime code was
added, so a healthy ring cannot be stalled by it. The existing danger tests
still pass: an abandon in the same epoch still restores the arm slot, and
sgl-project#757's leftover drain still runs on a committed falling edge.

Tests: 14 passed (9 existing + 5 new).
Mutants: mutants_829_chokepoint.sh, 4 written, 4 KILLED, tree restored green --
second rebuild site, rebuild stops clearing, second resume-slot writer, arm slot
recorded without its epoch.
test/registered/unit/managers/ collects 3670 tests cleanly.
ruff check clean, ruff format clean, codespell clean.

Merge compatibility: two new test files, zero source lines, so this cannot
conflict with the round-6 source merge. If round-6 adds a ring-rebuild or jump
site the pins go red by design, naming the site.

sgl-project#829 remains open only on METAL VERIFICATION, which is a window question.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant