Fix context length - #757
Merged
Merged
Conversation
5 tasks
timethink
pushed a commit
to timethink/sglang
that referenced
this pull request
Mar 9, 2025
efschu
added a commit
to efschu/htsglang
that referenced
this pull request
Aug 18, 2026
…ce in one harvest plan, with a proven runner Consolidated from the comp4 gate run (progress.662-F4-r5), WINDOW_TICKET_745/755, NOTE_747 par.8-9, NOTE_738/755, TICKET_727 and the operator ledger. Two hard incompatibilities shape the plan: WINDOW_TICKET_745 Arm 1 excludes the checkpoint interval that sgl-project#758's anchor-cadence observable requires (-> ARM I hicache harvest, ARM II = one flag more), and sgl-project#713's TTFT<3s needs a quiet router while every other loaded gate needs the soak backlog (-> sgl-project#713 is the idle sub-phase BEFORE the backlog, not a separate boot; the 06:44Z soak driver is the load source per the Lastprobe rule). ARM I phases: load-time (sgl-project#738 no-99G-plateau, file-backed-image reclaim), idle-quiet (sgl-project#713, health, corridor), loaded (Gates A/B/C, - sgl-project#757 race-holds, sgl-project#748 all three shapes, sgl-project#744/sgl-project#717 rung-funded flip, - sgl-project#690 refill census, sgl-project#758-2 mamba host resume, WT_745's three lines, corridor minima), teardown (image reclaim). ARM II adds --mamba-checkpoint-interval 8192 for sgl-project#758-1 anchor cadence + NOTE_747 par.8.1-8.3. SEPARATE windows named with reasons: sgl-project#727 four-boot A/B, WT_755 slots A/B (pool-geometry confound), sgl-project#755 metal retraction (mechanism not built -- nothing to measure), sgl-project#709, sgl-project#735 Step-2. sgl-project#602 and sgl-project#536/sgl-project#537 carried as HONEST unresolved slots (owner-held detail / not found with acceptance shape) rather than invented readouts. Runner run_window_ladder.sh: PASS/FAIL/UNOBS table from boot log + live server; never boots, never kills, never touches the soak driver; soak-tolerant by construction (loaded checks are log observations, the one latency check runs only in --phase idle, and choosing that phase IS the operator's quiet-router assertion). Missing-emitter cases (anchor cadence) report UNOBS, never FAIL -- absence of an instrument is not absence of the property. Mock-smoked per the desk rule, both directions: a fixture built from the comp4 specimen lines reproduces the real run's verdicts exactly (GATE-C crash, both sgl-project#748 shapes + vacuous relief, sgl-project#757 sentence = 5 FAIL; sgl-project#744/sgl-project#690 PASS; exit 1) and a clean fixture goes fully green incl. the ARM II cadence line (exit 0). bash -n clean, codespell clean. Nothing was booted.
efschu
added a commit
to efschu/htsglang
that referenced
this pull request
Aug 18, 2026
…gl-project#754 retires into the 753 fold, 13-step plan re-smoked clean The first map froze before today's second wave; new authority is F4-r5's harvest composite 59ce2d8 (declared COMPLETE, review tip still an ancestor). Pass 3 redefined: harvest tip = train base, unabsorbed branches cherry-pick on top, feat/753 lands FIRST by its owner (10 in-flight commits; carries sgl-project#749; folds sgl-project#754 at the same seam -- distributed/utils.py:1709, its own pp_size=1 handling). Sweep results, same git-cherry/merge-base rigor as the first map, outputs quoted in the ledger REFRESH section: - ABSORBED by ancestry: comp4 and its whole lineage, 915ce1b (F4-r5's own sgl-project#757), 57b04b2 (sgl-project#540 fix). - ABSORBED as different commits (desk-sgl-project#752 hazard class, never merge): fix/748-armed-gate-scope, fix/759-arming-economy, feat/755-slot-reorder. The sgl-project#758 emitters need no branch -- the harvest TIP ITSELF is a sgl-project#758 commit. - SUPERSEDED: my own fix/754 -- semantic-not-patch folded by 753 (git cherry vs 082293f shows '+'); merging it after 753 lands guarantees a get_pp_layer_set conflict with zero gain. Retired from the plan without regret. - REVIEW-never-merge: 9e56477 (independent sgl-project#757), per its reviewer's own in-composite note naming 915ce1b as the baseline. - fix/706-remainder not yet visible; slot reserved. Executor updated: HARVEST constant joins the lineage check, DEFAULT_TIP moves to the harvest tip, the sgl-project#754 step is replaced by the sgl-project#745 reachability pick, and the second-wave picks join (727 head-chain, sgl-project#738 verdict, sgl-project#535 tickets). Dry-run scratch smoke against the REAL harvest tip: all 13 steps complete with ZERO conflicts (exit 0) -- cleaner than the first wave; the sgl-project#740 pair ordering from the previous smoke holds. Plan-mirror test updated (12 passed). Gate unchanged: COMP4_ACCEPTED still required, nothing pushes, nothing booted.
efschu
added a commit
to efschu/htsglang
that referenced
this pull request
Aug 18, 2026
…ENT, loud review item -- and the harvest tip moved to da81871 The double-sgl-project#754 template does NOT apply. Both sgl-project#757 fixes attack the same root (corpse-S drain disabled, rank-local disarm routes) at DIFFERENT intervention points: 915ce1b (in harvest, regression- verified) drains the leftover at DISARM; 9e56477 (Slot-3) re-enables the ARMED drain with a pure 4-way demultiplex classifier and states a liveness property the disarm-time form lacks (the upstream's blocking commit waits on the wire being consumed AT ARRIVAL). Review question stated sharply in the ledger: does a long armed window stall the group under the disarm-time form? If yes, the two are COMPLEMENTARY halves, not duplicates. Slot-3's test suite (3-process gloo repro, 4 killed mutants, sgl-project#631-pin correction) binds to their classifier, so the tests ride the review verdict -- not cherry-pickable bare. 9e56477 stays OUT of the executor plan; the review joins feat/753 as a NAMED precondition of pass 3. Tip correction folded in: authority is now da81871 (contains the - sgl-project#758 phase-tag 59ce2d8 AND caca352 -- F4-r5 absorbed WINDOW_LADDER_0818 + its runner into the harvest). Executor HARVEST constant bumped; dry-run scratch smoke re-run against the REAL new tip: all 13 steps complete, zero conflicts, exit 0; plan-mirror tests 12 passed.
efschu
added a commit
to efschu/htsglang
that referenced
this pull request
Aug 18, 2026
…fold shipped, executor gains the two steps The review item the refresh flagged is closed with a measurement, not an argument: gloo isend posts instantly but wait() blocks until the peer recv (3000.3 ms stall for a 3 s armed window on a 4 KiB proxy, size-independent -- tools/probe_757_gloo_liveness.py), and the in-tree sender is exactly that shape. The disarm-time form alone stalls an abandoned upstream at its commit for the rest of the armed window; verdict (b), COMPLEMENTARY. Fold shipped as fix/757-armed-liveness (194c3ea + probe 5e2c121): both suites + corrected sgl-project#631 pin green together, managers selection 633 vs 625 passed with identical 2 pre-existing failures, 0 new. 9e56477 superseded by the fold. Executor: two steps added (the fold pick + the probe pick), pass-3 precondition list drops the review and keeps only feat/753. Dry-run scratch smoke re-run against the real harvest tip: 15 steps, zero conflicts, exit 0; plan-mirror 12 passed.
efschu
pushed a commit
to efschu/htsglang
that referenced
this pull request
Aug 20, 2026
…ion, not before the exchange
A boot reached serving, one admissible queued request was never served, and the
process died. The wedge detector reported "1 queued, 0 running, NO first token
for 340-370s and no prefill chunk either" and named its own gap in the same
line -- "no phase-policy corroboration seen -- the wedge class is broader than
that path". It was right: this is not an admission defect at all.
EVIDENCE OF RECORD. py-spy on the specimen (WEDGE_788_specimen.log:732/915/1083)
catches all three ranks in a circular wait:
PP0, PP1 gloo waitSend _pp_commit_comm_work (scheduler_pp_mixin.py:2465)
from _pp_forward_and_process_input_requests:1007
PP2 gloo waitRecv _pp_recv_typed_dict:2709
from _pp_recv_proxy_tensors:2803
The request chain was flushed at the TOP of the pass, before the rank had
posted anything else it owed its peers that iteration -- the proxy tensor-dict
send and the output-ring send both come later in _event_loop_pp_body. A
downstream whose progress needs one of those was therefore waiting on a rank
that was itself waiting on that downstream. The chain recv is posted only at
the top of a pass (:420) and nowhere else in the loop body, so a rank parked in
the proxy receive cannot drain the message its upstream is blocked flushing.
THE FIX IS ONE ADDITION. _pp_forward_and_process_input_requests is byte-identical
to before: it still commits send_req_work, then posts the async forward, then
runs process_input_requests. What is new is _pp_commit_pending_req_work, called
once per iteration from _event_loop_pp_body after the proxy isend and the output
commit and BEFORE the phase-flip round hook. By the time that loop returns to
the top-of-pass commit, the handle has already been waited on and cleared, so
the old call site is a no-op there -- and the two disaggregation loops (:618,
:765), which do not call the new method, keep their previous behaviour exactly.
Deleting rather than moving the old call site is deliberate on both counts: it
keeps the sgl-project#633 ordering contract literally intact (commit, then forward, then
process, with send_req_work holding this pass's handle on return -- pinned by
test_scheduler_pp_request_order_633), and it avoids a second staging slot. An
earlier shape of this fix did stage the send separately; it broke that contract
and left the two disagg loops overwriting a live P2PWork handle every pass.
Precedent: a7ff250 "[sgl-project#753] Flush the output sends after the exchange, not
before it" made the same move for the OUTPUT channel. This is that fix for the
REQUEST channel, ungated -- the hazard is not gapped-specific.
NEGATIVE FINDING, RECORDED RATHER THAN BURIED. A deterministic 3-process gloo
reproduction of the specimen deadlock was attempted and NOT achieved under a
faithful model of the loop. Two candidate mechanisms were tested and refuted:
- Eager-send asymmetry. Hypothesis: an empty request list puts one 8-byte
tensor on the wire and completes eagerly, so only a large first request can
block. MEASURED FALSE: a 2-process probe with a 2.0s-delayed matching recv
blocked for the full delay at every payload size from 8 B to 1 MiB in this
build. No eager completion at any size.
- Schedule asymmetry. A linear chain+proxy relay that flushes the old handle
before dispatching the new one is deadlock-free by induction (each rank's
flush needs only its immediate downstream one pass behind), and dozens of
real 3-process runs across depths and offset combinations all completed.
So the reduced two-channel model cannot close a cycle. The production graph has
an edge that model lacks: the output ring closes last-rank -> rank 0
(_pp_send_recv_and_preprocess_output_tensors:3038, last-rank branch :3012-3014,
via next_first_rank_mb_id at :416/:608/:754). That, and real GPU compute skew,
are the two named unconfirmed candidates for the missing ingredient. Neither is
built into the test, and the gap is stated in its docstring rather than papered
over. The mechanism's evidence of record therefore remains the py-spy trace
above plus the sgl-project#753 precedent, and the survival boot is the integration proof.
RECOVERY RUNG. The detector gains one bounded action: after a wedge stays
continuously alarming past a threshold well above the report threshold (default
3x, env SGLANG_ADMISSION_WEDGE_RECOVERY_SECONDS; non-positive values fall back
to the default rather than firing every poll), it makes ONE forced-admission
attempt per episode through the existing corridor_admission actuator and logs it
loudly either way. Its own docstring is explicit that it relieves VRAM pressure
at the admission site and will NOT move a comms deadlock like this one, and that
calling an actuator from the watchdog thread is a new cross-thread shape over
CUDA allocator state.
TESTS
test_pp_chain_flush_deadlock_788.py (new; 3 real gloo processes, CPU only)
- load-bearing: runtime ordering pin. The SHIPPED _pp_commit_pending_req_work
is wrapped and the observed sequence asserted to be
proxy_send -> chain_flush -> round_hook for every pass and every non-last
rank, so the flush cannot drift back past the collective and cost it the
"last blocking op of the iteration" property its own comment relies on.
- can-fail: the same run with the flush relocated to the top records a
different sequence, so the pin discriminates placement.
- the honest negative case described above, in place of the deleted
hang-repro. A test that is green or red for the wrong reason is worse
than no test.
test_pp_flip_leftover_proxy_757.py: harness-only repair, no assertion touched.
It was silently DEAD on this branch -- its _GlooWire predated
rank_in_group/world_size and the src positional recv_typed_tensor_dict now
passes, so it errored 3/3 before reaching an assertion. Now 3 passed and a
working regression check again.
Verified green with this diff: test_pp_chain_flush_deadlock_788 and
test_scheduler_pp_request_order_633 (8 passed), test_phase_policy (90 passed),
test_pp_flip_leftover_proxy_757 (3 passed).
PRE-EXISTING FAILURES, MEASURED NOT ASSUMED. The managers slice reads 43F/2707P
here against 39F/2705P at a pristine detached ff5651a worktree. Every delta
is accounted for: -3 (the 757 repair above now passes), +5 stub drift and +1
ordering-contract break, both introduced by the earlier staging-slot shape and
both gone with it, +1 test_pp_drain_completeness_787 which is INTENTIONALLY RED
here -- that file's fix lives on fix/787-drain-completeness and it flips green
at the merge, which makes it a free integration can-fail. The residual 39 are
the same interface-drift family (rank_in_group, _pp_gapped_wire) plus one stale
constant, all present at base with no part of this diff.
DEBTS
sgl-project#789 the two disaggregation PP loops keep the top-of-pass
flush and therefore still carry this hazard
sgl-project#757-harness-drift repaired here; the same drift still blinds
test_pp_proxy_stamp_631 and test_pp_slot_last_batch_631
sgl-project#753-stale-constant 10.923 != 11.923 in the gapped entry protocol test
sgl-project#766-pointer-stale the register points sgl-project#766 at ARM_defaultfull7.log, which
is a 26-line host-ledger snapshot with no
Bar1CollectiveAborted and no proxy/drain lines
efschu
pushed a commit
to efschu/htsglang
that referenced
this pull request
Aug 21, 2026
…eriving it (slice 1) Ten fixes on this branch (sgl-project#757 sgl-project#789 sgl-project#790 #791b #791c sgl-project#792 sgl-project#795 sgl-project#797 #797b #797c) all have ONE form: each rank RE-DERIVES the pass schedule -- rid set, chunk length, prefix length -- locally from its own state, and a phase flip invalidates that state non-atomically. Every "new root" was the next consumer of the same re-derivation, so the list grew instead of converging. This stops patching consumers. THE DATUM WAS ALREADY ON THE WIRE AND NOBODY READ IT. PPAdmissionEntry.extend_len (pp_admission_congruence.py:177) has crossed the wire since sgl-project#791's first commit with exactly three consumers: to_wire (:223), from_wire (:244), one log string (:824). Nothing ever built a batch from it; reconcile_pp_admission_decision returns Dict[rid, prefix_len] and drops the second number on the floor. THE MECHANISM, CORRECTED AGAINST THE LOG rather than assumed. boot_instr20.log:5171,5181-5183: PP0 ADMIT rid=6cbe2733 prefix_lens=0 chunked=1 -> 512-row chunk PP1 ADMIT rid=6cbe2733 prefix_lens=512 chunked=0 -> 333-token remainder PP1 DID receive prefix_len=0 and DID clamp prefix_indices to it -- scheduler.py:7026-7027 worked. Then add_one_req's HOST LOAD-BACK put the 512 back: needs_host_load_back() went true when the HiCache prefetch landed and schedule_policy.py:1539-1549 concatenates the recovered indices. "MAMBA-HOST-RESUME ... triggers load_back" appears on PP1 and PP2 and is ABSENT on PP0 -- that asymmetry is the bug. 845-512=333 then fitted rem_chunk_tokens whole, so the NON-chunked branch fired. The re-derivation lives INSIDE THE ADDER, after the schedule was already applied, and it re-derived BOTH numbers. DESIGN. forwarded_schedule() (pp_admission_congruence.py:500) is the pass geometry as a value: rid -> (prefix_len, extend_len) for exactly the rids `effective` names; None or a voided decision yields {}. _add_scheduled_req (schedule_policy.py:1237) EXECUTES both numbers: no rem_chunk_tokens, no page/align rounding, no host load-back, no budget veto -- the budget is still charged. The gate sits above every local veto (:1600), and add_chunked_req (:1328) gains the gate it never had at all (it is entered from scheduler.py:7004, BEFORE the admission loop). The three membership vetoes that silently narrowed -- batch_is_full (:7035), the HiCache prefetch_done skip (:7047, the instr20 race itself) and the LoRA gate (:7024) -- become refusals (scheduler.py:7190). REFUSAL IS CONTROL FLOW, NOT A RESULT CODE. PPScheduleRefused (:163) is an exception because every AddReqResult means "build a batch without this request", which is precisely the corruption. A refusal reuses sgl-project#797 end to end (_pp_refuse_forwarded_schedule, scheduler.py:6333 /:6369): sets _pp_admission_pass_voided, voids the forwarded decision, re-notes the slot expectation. No new mechanism. Inside the loop a refusal is CARRIED, not thrown, so alloc_group_end() still runs (:7025, :7118, :7169). WHY THE GUARDS CAN NO LONGER FIRE, each owed a reason: sgl-project#631's _want is extend_num_tokens, which now comes only from the forwarded extend_len with load-back suppressed, so it IS the upstream's row count by construction rather than by agreement (green arm: rows=512, batch_tokens=512). sgl-project#757/sgl-project#787 stamp and sgl-project#795 epoch were already structural; what changes is that they can no longer be correct-but-insufficient, as instr17 and instr20 both were -- every identity right, only the width wrong. Width is now an identity too. sgl-project#789 needs a membership divergence, which is now identical-or-refused. #791c's tripwire detects a self-narrowed batch, and no path creates one. HONEST CORRECTION TO THE SUBSUMPTION CLAIM: retraction does NOT become unnecessary. Physical impossibility is real -- a rank genuinely lacking KV for [local, told) cannot execute. What becomes structurally impossible is the NARROW-THEN-DETECT shape: no code path is left that builds a batch of a geometry the upstream did not name. NOT COVERED BY THIS SLICE, stated so nobody assumes otherwise: - BATCH ORDER. can_run_list follows the local waiting_queue order, the decision follows PP0's. Same rid set in a different order gives EQUAL WIDTHS and permuted rows -- silent. Covered by the full design (execute in decision order), not by this slice. - Decode batches: retract_decode (scheduler.py:7489/:7517) mutates long-lived state on a tp_cpu_group reduce that is world=1 under TP=1/PP=3. Different root, filed by sgl-project#797. - Radix eviction divergence, KV pool sizing, spec-decode draft schedules. NAMED RESIDUAL, filed at the site (scheduler.py:7118) in sgl-project#797's practice: a request admitted earlier in a loop that later refuses has taken a persistent inc_lock_ref, released on batch completion. Undoing it needs the exact IncLockRefResult (SWA/Mamba tombstone params) the adder does not keep, and a blind release makes the one thing a mismatched release worsens. Bounded: reaching that line takes a genuinely unexecutable geometry and kills the pass. DEFAULT PATH UNTOUCHED: _pp_scheduled_extents() returns None on PP0 and on every pp_size<=1 boot, so scheduled_extent_for returns None and both adders take the pre-existing arithmetic unentered; PPScheduleRefused is unraisable there. Pinned by test_no_mapping_is_the_untouched_default_path and corroborated by 56/56 on the schedule_policy neighbours. TESTS. test_pp_forwarded_schedule_791.py: 17 passed (3 live gloo arms + 14 pure), 92 s. test_red_without_the_forwarded_geometry_instr20_reappears -- can-fail, rebinding ONLY the fix's return value in the child: batch=(512,333) rows=512 mismatch=True, byte-identical to instr20 PP1 09:40:30 test_green_the_forwarded_geometry_survives_the_mid_pass_prefetch -- batch=(0,512) rows=512 mismatch=False, load_back_calls=0 test_an_impossible_geometry_raises_rather_than_narrowing -- the architectural property Neighbour set (791/791b/791c/797/631 x3) re-measured AT HEAD: 11 failed / 75 passed; after: 11 failed / 92 passed, failure names byte-identical, zero regressions, +17. schedule_policy neighbours 56 passed / 0 failed. ruff 119 = 119 at HEAD (parity), ruff format clean on everything authored, codespell identical.
efschu
pushed a commit
to efschu/htsglang
that referenced
this pull request
Aug 22, 2026
…n a send nobody owes
The PP output ring wedged with all three ranks alive and none of them
able to move (specimen /spinning/evidence-665-f1/wedge_802f_1712/, PP=3,
--enable-phase-flip, py-spy of all three schedulers):
PP0 _pp_commit_comm_work <- _pp_commit_pending_req_work (:4071/:2262)
PP1 _pp_recv_dict_from_prev_stage <- _do_recv (:5608/:6069)
PP2 PpChainReceiver.recv <- recv_requests (pp_chain_receiver.py:329)
PP0 is flushing the request chain, which PP1 can only take at the top of
its next pass; PP1 never reaches that top because it is blocked in the
output receive; PP2 waits for the chain send PP1 has not reached. Three
arcs, one cycle, no timeout anywhere.
THE ASYMMETRY. The two ends of the intermediate hop apply unrelated
predicates. `_do_recv` decides to receive from THIS rank's own slot
state; the non-last sender in `_pp_send_output_to_next_stage` decides to
forward on `if pp_outputs:`, which is whatever it received LAST
iteration. The last-rank hop is matched by construction because it
consults `_pp_output_expected_for_slot` -- but that flag is the FIRST
rank's verdict, published for PP0's arc. Nothing publishes the same
thing for the intermediate hop.
WHY THE EXISTING VOID CONTRACT DOES NOT COVER IT. `pp_void_forward_
payload` (sgl-project#797) already forwards a void along this hop and stops at
`pp_first_retracting_rank`, on the argument that rank r and everything
after it has an empty slot whose receive early-returns. That holds for
the VOIDED slot -- `_pp_void_own_batch` empties it. It does not hold for
a slot the resident decode path still occupies, and both void paths keep
resident requests on purpose (`_pp_absorb_void_output` refuses to release
them as a double-free; `_pp_void_own_batch` deliberately leaves
`running_mbs` alone). The specimen shows exactly that: PP0's last act was
absorbing a void for slot 0, leaving `pp_outputs` None, while PP1 logged
`running=1 chunked=1` and re-entered the receive for a slot PP0 would
never send to. PP0 cannot know PP1's resident set, so closing this needs
a per-slot expectation every non-first rank publishes to its predecessor
-- a protocol extension, and not this commit.
WHAT THIS COMMIT DOES. The sgl-project#789 readiness gate is parameterised by wire
kind and the output receive now passes through it. The CHAN_DICT counter
was never proxy-specific -- one counter per wire, demultiplexed by
`__msg_type__` after it comes off -- so the gate reads the same true
statement about the same wire either way; `kind` selects only the
per-(src, kind) inbox peek. The ring is cut at its one cuttable arc: PP1
waits boundedly and refuses by name instead of for ever. It does not make
the missing send appear, and the error text says so.
AN ALIAS, NOT A WRAPPER. `_pp_wait_for_proxy_readiness` is now a
class-level alias for the same function object. About ten stand-in
holders across the sgl-project#631/sgl-project#757/sgl-project#787/sgl-project#789/sgl-project#791/sgl-project#795/sgl-project#797/sgl-project#798 test family
bind that name one method at a time; a delegating wrapper resolved the
second name on the HOLDER and turned 9 green tests into 5 failures.
Measured, then fixed.
Tests: test/registered/unit/managers/test_pp_output_readiness_ring_802.py
-- three real gloo processes, real PhaseFlipCounters, neutering done in
the child (spawn re-executes the module, not the test body).
* red arm, gate neutered: stuck_ranks == [0, 1, 2], specimen reproduced
* green arm: stuck == [], PP1 reports "sgl-project#789 OUTPUT READINESS TIMEOUT"
and "sgl-project#802-ring", PP0 chain-flushed, PP2 requests-received
* false-positive direction: upstream really posts -> receive succeeds
* no-op without counters, so the non-phase-flip default path is unchanged
Mutants, both die: call site removed -> green arm red (all ranks stuck);
gate raises unconditionally -> false-positive test red.
Regression: baseline 787+789 = 9 passed / 0 failed; with this change
791b + 789 + 787 + 802 = 17 passed / 0 failed.
efschu
pushed a commit
to efschu/htsglang
that referenced
this pull request
Aug 23, 2026
…he ring with the existing drain W4a made the PP chain receive bounded, named and resumable, but left its automatic abort off (SGLANG_PP_CHAIN_RECV_STALL_S=0), so on metal it could diagnose nothing and recover nothing. This wires it. THE DEFAULT STAYS OFF, AND THAT IS THE POINT. An idle PP rank legitimately blocks in this receive until a request arrives, so there is no duration that separates "idle" from "wedged"; a wall-clock default would SIGQUIT a healthy idle server, the same trap sgl-project#821's marker documents for its own arm. The recovery is therefore armed by EVIDENCE. THE PREDICATE. bump_attempted publishes that a rank has ENTERED a send BEFORE it posts it (phase_flip_counters.py:220-226) -- the only counter whose timing can witness a peer parked INSIDE a send rather than one that has finished. When this rank's CHAN_DICT upstream has entered more dict sends than this rank has taken off that wire, the peer is parked in a send only this rank can drain while this rank is parked in a receive that peer will never feed. That is boot_827's ring stated in counters: PP0 in _pp_commit_admission_send_work on the typed-dict channel, PP1 and PP2 in the request-relay chain receive. abort_check is consulted only at the yield, i.e. only once a receive is already overdue, so a healthy pass never reaches it. THE RING-CUT REUSES sgl-project#757 RATHER THAN INVENTING A CONSUMPTION PATH. pp_flip_drain_leftover_dicts already demultiplexes, stashes a wrong-kind message in _pp_tensor_dict_inbox where its real consumer looks, and discards only a provably void proxy -- which is exactly what makes taking the dict off the wire out of the pass's normal order safe. It runs on every disarm route already. request_receiver catches PpChainRecvStalled, runs one drain turn, and resumes the SAME posted receive, which is what ParkedWait's resumability is for. Servicing that does not clear the stall re-raises rather than spins: at that point the ring is closed for a reason this code does not model, and retrying would turn a diagnosable wedge into an invisible one. Follows sgl-project#789's shape deliberately. _pp_wait_for_dict_readiness argues the false-positive direction is the safe one for the mirror gate, and it is here too: a spurious fire costs one drain turn and a resumed receive, and the receive stays posted and framed throughout, so the late message still arrives intact. Missing a real one costs the boot. sgl-project#789 also declines to invent a new protocol for its case; this declines likewise and cuts the ring on the arc waiting for a message nobody posted. Also: _pp_flip_pass_tick publishes _pp_live_mb_id, which the drain needs to tell an owed proxy from a leftover one. TESTS (hermetic, CVD="", CPU only, no gloo, no scheduler construction) test/registered/unit/managers/test_pp_chain_abort_check_824.py 9 passed; all 9 fail pre-fix. Covers both directions: the boot_827 counter state aborts, and an idle rank is NOT aborted at 0 s or at 3600 s. Mutants killed, one per edge: entered >= taken -> the idle rank is aborted (1 failed) service-turn cap raised -> an unclearable stall stops being re-raised service hook never called -> the ring is never cut (2 failed) Run under /spinning/htsglang-gpu/.venv (torch 2.11.0+cu130, datasets 5.0.0) with PYTHONPATH leading to this worktree, verified by `import sglang; sglang.__file__`. My earlier runs used /usr/bin/python3, which lacks datasets and silently collected only part of the suite; see the corrected note in WINDOW-QUEUE.md. No boot was run. This is desk work.
efschu
pushed a commit
to efschu/htsglang
that referenced
this pull request
Aug 23, 2026
Wave 3, stage 3 of the batched-window tree. WINDOW-QUEUE tickets W4 (PP ring wedges immediately after an abandoned flip) and W5 (sgl-project#821 did not instrument the receive path that actually wedged), both preflight_pass=Y. Four commits, base 500be7e -- which is EXACTLY this tree's stage-1 branch, so this stage sits directly on top of W1/W2 with no divergence to reconcile: 2efb933 [sgl-project#824 W4a] Bound the PP chain receive without tearing the stream it guards dd92de4 [sgl-project#824 W5] Name the arm that tripped the watchdog, and instrument the path that wedged 430e468 [sgl-project#824 W4b] A zero-iteration armed window must not move the slot ceb79d6 [sgl-project#824 W4a] Arm the chain-recv recovery on STATE, and cut the ring with the existing drain GATED AGAINST ceb79d6, NOT 430e468. The branch head moved while this stage's first battery was already running, so that run was measuring a superseded commit and was DISCARDED rather than reported -- a gate that certifies a commit the tree does not carry is worse than no gate. The in-flight battery was stopped by explicit PID (no pkill), py-spy dumped first per standing rule and found healthy mid-test rather than wedged, and its partial log kept aside as w3c_stale430.log. A foreign battery belonging to another session was running on the box at the same time and was NOT touched. The final W4a commit arms the recovery on STATE (`bump_attempted`, entered >= taken as the evidence predicate) and cuts the ring through the existing sgl-project#757 drain, with the wall-clock default left at 0. W4b is the ROOT of the pair, and it lands in the same mechanism as this branch's own known-red neighbours: an armed window that ran ZERO slot iterations must not advance mb_id. The boot_827 specimen recorded exactly that -- "rank 0 ran 0 slot iteration(s) (armed at mb_id=0, disarmed at mb_id=1)" -- one line before the ring stopped dead and stayed silent for 31 s until the health check noticed. That is the sgl-project#631 defect class with the spread reduced to a single rank: arm on one slot, leave on another, then re-enter the pipeline wherever you happened to stop. TOUCH SURFACES, computed rather than assumed. This stage shares scheduler_pp_mixin.py with both earlier wave-3 stages and with wave 2, and it extends sgl-project#821's own suite (test_pp_wedge_watchdog_is_honest_821.py) rather than duplicating it -- W5 is the continuation of sgl-project#821, by the same argument sgl-project#821 made, so touching that file is intended and not a collision. It also adds to mem_cache/hicache_collective.py, the file carrying sgl-project#734's `waited < timeout_s * 0.95` discriminator that strand 17c's 622 falsification identified as the fragile part. That last point is what makes this gate non-vacuous rather than merely green: `test_pp_sync_rendezvous_630.py` -- the suite that went red under the 622 merge for precisely this threshold, and the reason 622 is still off the line -- sits INSIDE the battery. If this stage disturbed that discriminator, the gate would say so in the same way it said so for 622. NOTE ON A PATH, because it looks like a contradiction and is not: this stage modifies python/sglang/srt/utils/watchdog.py, while commit 9f1af20 fixed a catalog citation that resolved against python/sglang/srt/watchdog.py. Those are different files; the catalog's intended target was turnkey/watchdog.py:88 (the generation probe retired by user order), which is where it now points. GATE. Battery test/registered/unit/{managers,planner,server_args,mem_cache}, hermetic under CUDA_VISIBLE_DEVICES="", one battery at a time. baseline (tip 4f2072a) 7 failed, 8585 passed, 1852 skipped, 887 s stage 1 (W1+W2) 7 failed, 8638 passed, 1852 skipped, 868 s stage 2 (W3) 7 failed, 8653 passed, 1852 skipped, 922 s this stage 7 failed, 8673 passed, 1852 skipped, 915 s NEW failure ids NONE (comm against the baseline list is empty) fixed vs baseline 0 THE +20 IS ACCOUNTED FOR EXACTLY, not assumed. Four new suites adding only 20 passes is low enough to be worth checking, because a suite that silently skips looks identical to a suite that passes cheaply. Two facts settle it: the skipped count is 1852 in every run of this train -- baseline, stage 1, stage 2 and this one -- so nothing new is being skipped; and running the four suites standalone reports exactly 20 passed, matching the battery's delta to the case. They run, and they are collected. COVERAGE, and this stage is the reason to state it explicitly. The W4/W5 strand reported that its OWN worktree cannot collect the managers suite -- 92 modules import `datasets`, which is not installed there. That limitation is NOT present in this gate, verified in the positive direction rather than inferred from a missing error line: datasets 5.0.0 is importable in this venv, the battery logs carry zero collection errors and zero ModuleNotFoundError, and an explicit --collect-only over test/registered/unit/managers reports 3436 tests collected. So this stage's no-regression claim rests on a fully collected managers suite, which the strand's own local run could not provide. Its four new suites land inside the battery directories and are therefore actually collected: test_pp_ring_abort_recovery.py, test_pp_wedge_watchdog_names_the_arm_824.py, test_pp_zero_iteration_window_slot_824.py and test_pp_chain_abort_check_824.py, all under managers. ON PROVENANCE, if this stage ever shows a new failure id: the branch is PRE-sgl-project#815 (its base 500be7e sits on 21ff075), so the last_batch-631 failures visible in the strand's own worktree are expected to DISAPPEAR on this line, which carries fix/815. A new id here would therefore have to be checked against that provenance before being attributed to the strand's change. In this run none appeared, so the question stays hypothetical. No boot. This tree is the input to ONE batched window.
efschu
pushed a commit
to efschu/htsglang
that referenced
this pull request
Aug 24, 2026
…s on sgl-project#829's root -- an arm slot outliving the ring it names -- is already closed on integ/round5 by d507e19, and test_pp_arm_slot_outlives_ring_829.py pins the behaviour with 9 green tests. This adds no runtime code and changes no source line. It pins the PREMISES that make the closure a proof rather than an observation. THE ROOT, restated with file:line so the pins have a referent: :2606 rising edge records self._pp_flip_arm_mb_id = mb_id :2616 and self._pp_flip_arm_epoch alongside it cutover commits with _pp_flip_armed_passes == 0 :3905 init_pp_loop_state REBUILDS mbs/last_mbs/running_mbs from zero, and pp_loop_size may CHANGE (TP phase runs pp_size=1), so a carried index can be out of range :2776 the falling edge jumps the pass loop to the recorded slot Pre-sgl-project#829 that put PP0 on slot 2 of a fresh ring while its downstream sat on 0 (boot_window2_0823_1554, epoch 6, "RESUME SLOTS [2, 1, 1] -- DIVERGED"), and PP1 raised sgl-project#631 PROXY LEFTOVER REFUSED one second later. TWO GUARDS CLOSE IT, and they are COMPLEMENTARY rather than redundant: A :3903 init_pp_loop_state calls pp_flip_forget_ring_scoped_slots(self) unconditionally, retiring all three ring-scoped carriers. B :2755 the falling edge refuses a restore across a generation (ring_rebuilt = arm_epoch != now_epoch); the only jump is the elif arm at :2776. init_pp_loop_state has THREE callers and only the cutover advances the epoch, so guard B is blind to the other two rebuilds and guard A covers them. Neither may be dropped. WHAT WAS MISSING. The proof rests on three uniqueness facts that no test pinned: the ring is rebuilt in exactly one place, the arm slot is recorded in exactly one place, and the pass loop can be jumped from exactly one place. A second rebuild site or a second jump writer would bypass both guards WITHOUT failing any existing test, silently restoring the state that killed the boot. Verified by exhaustive enumeration over python/sglang/: one rebuild site (:3905-3910), one arm writer (:2606), one jump site (:2776). test_pp_ring_rebuild_choke_point_829.py -- 5 AST-level uniqueness pins, ~2 s, CPU only, no torch and no process group. A failure there is not automatically a bug; it is a demand that a newly added site be checked against the epoch gate, and the failure messages say so and name the offending function and line. Deliberately NOT pinned: the order of the clear against the array assignments. Both orders are correct and pinning one would make a harmless refactor red. DANGER DIRECTION. Nothing here enforces uniformity and no runtime code was added, so a healthy ring cannot be stalled by it. The existing danger tests still pass: an abandon in the same epoch still restores the arm slot, and sgl-project#757's leftover drain still runs on a committed falling edge. Tests: 14 passed (9 existing + 5 new). Mutants: mutants_829_chokepoint.sh, 4 written, 4 KILLED, tree restored green -- second rebuild site, rebuild stops clearing, second resume-slot writer, arm slot recorded without its epoch. test/registered/unit/managers/ collects 3670 tests cleanly. ruff check clean, ruff format clean, codespell clean. Merge compatibility: two new test files, zero source lines, so this cannot conflict with the round-6 source merge. If round-6 adds a ring-rebuild or jump site the pins go red by design, naming the site. sgl-project#829 remains open only on METAL VERIFICATION, which is a window question.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixed wrong context length for deepseek v2