Skip to content

[Router] Gate readiness on peer bootstrap, and make it observable (13/13) - #40699

Merged
ShangmingCai merged 11 commits into
mainfrom
router-peer-bootstrap-13-readiness-gate
Oct 1, 2026
Merged

ShangmingCai merged 11 commits into
mainfrom
router-peer-bootstrap-13-readiness-gate

Conversation

@Kangyan-Zhou

@Kangyan-Zhou Kangyan-Zhou commented Sep 22, 2026 •

Copy link
Copy Markdown
Collaborator

Stack 13 of 13. Base: router-peer-bootstrap-12-k8s-e2e. 9 files changed, 837 insertions(+), 32 deletions(-)

The stack

# PR lines what it adds
1 #40687 962 the sharded tree gains export_snapshot / restore_snapshot; nothing calls them yet
2 #40688 2213 GET /internal/kv_snapshot — the producer half; nothing consumes it yet
3 #40689 536 --kv-peer-selector, the peer registry, the _peers gauges; no behaviour change
4 #40690 1150 the EndpointSlice watch that fills the registry; still no consumer
5 #40691 1177 BootstrapTracker + the fetch client; the consumer's state and transport
6 #40692 933 VettedSnapshot::from_wire — the only bridge from wire bytes to the tree
7 #40693 2750 the pump holds a Pending rank's batches and grafts a snapshot handed to it
8 #40694 2478 the sweep: ask siblings, vet, deliver — one sweep per discovered worker
9 #40695 1211 fold a discovery burst into one fleet-wide fetch, plus the gap retry it enables
10 #40696 844 a graft nothing witnessed asks the fleet instead of guessing
11 #40697 646 component test over the real transport: same match_prefix answers as the source
12 #40698 943 kind-cluster proof, the KV-publishing fake worker, and the RBAC/downward-API manifest
13 #40699 869 /readyz holds until bootstrap settles; --kv-bootstrap-seed-required; the metrics ← this PR

Each PR's base is the branch below it, so every diff shown here is that
PR's own change. Review bottom-up; GitHub retargets each child to main as
its parent merges.

What the series does

A router replica subscribes to each worker's KV topic mid-stream, so every
block already resident in the engine's radix cache is invisible to it — and
engines publish BlockStored only as they insert, so a prefix cached hours
ago is never re-announced. A cache-blind replica then scatters the prefixes the
warm replicas were keeping consolidated, degrading the engines' locality for
the whole fleet; a rolling update does that to every replica in turn. This
series makes a booting replica pull a tree snapshot from a warm sibling over
HTTP and graft it beneath its live delta stream. Off unless --kv-peer-selector
is set.

Supersedes #39750, which carried the same work as one branch on a stale base.

What this change does

A replica that boots with an empty KV tree does not just route cache-blind
itself — its dispatches scatter the prefixes the warm replicas were keeping
consolidated, so it degrades the engines' locality for the whole fleet. A
rolling update does that to every replica in turn, and with nothing in the
metrics to attribute the hit-rate drop to.

/readyz grows two conditions. The third holds 503 until peer bootstrap has
settled: a snapshot grafted, or every sibling having proved it has nothing, or
--kv-bootstrap-timeout-ms giving up. It latches, so a later scale-up can
never drag an already-serving replica back out of the Service. It is inert
unless both cache-aware routing and a peer selector are configured, so nothing
changes for a fleet that has not opted in.

The fourth is --kv-bootstrap-seed-required, off by default, and it exists
because condition 3 settles either way once the deadline expires — that escape
is load-bearing, and it is also why settled() cannot tell "seeded" from "gave
up". This one can: it holds readiness when a sweep proved siblings were there
and their tree could not be pulled. A NotReady pod stays out of the Service, so
a failed seed DELAYS that replica's readiness — and with it the rolling update,
the previous generation still serving — by up to
max(3× --kv-bootstrap-timeout-ms, 60s) — 30 minutes at the default — after
which the replica goes ready
and serves cache-blind. Nothing re-sweeps the failed ranks while the gate holds
(only a sweep started by a newly discovered worker can clear it), so the flag
buys a bounded delay and a louder signal, not a guarantee: on its own it does
not stall a rollout indefinitely, and a rollout whose seeds keep failing still
completes, with cache-blind replicas.

Which verdict may gate is the entire safety argument. NoPeers (first deploy,
single replica) and FleetCold (every sibling PROVED it holds nothing) are
legitimate nothing-to-inherit outcomes; gating on either bricks a boot that has
nothing to inherit — and because an unready replica leaves its own
EndpointSlice, every sibling would then see an empty peer set: a fleet-wide
deadlock with no automatic exit. Only TimedOut over a non-empty candidate set
counts. Two further rails cover the case the verdict genuinely cannot
distinguish, a simultaneous fleet-wide restart where every peer is booting: the
hold is bounded at max(3× the bootstrap budget, 60s), measured from first
registration, and then opens with a WARN; and it latches open the first time
/readyz answers 200. It cannot latch open early: while ranks are registered
and the boot sweep has not reported, the gate stays closed, because every rank
can go terminal (overflow, publisher reset, the deadline) before that verdict
arrives and a 200 in that window would latch it unchecked. The latch holds in
the other direction too — a pool can go ready on a worker that publishes no KV
events, and the first KV rank registering afterwards does not un-ready the
serving replica.

The tests assert over the verdict class rather than by example, because the
wrong answer in either direction is silent — gate too little and a bad rollout
ships, gate too much and the fleet deadlocks.

The metrics make the whole path attributable. bootstrap_state per rank, so a
rank stuck pending (holding its events back, heading for overflow) is visible;
peer_snapshot_total by fetch outcome, where a fleet pinned at unreachable
with a warm tree means the per-fetch bound is too small for the body rather
than a network fault; bootstrap_rank_total once per rank at its final
verdict; bootstrap_sweep_total per sweep, which is what separates a healthy
early settle on a cold fleet from burning the whole deadline — the rank
outcomes fold both into abandoned. Each counter family emits a zero row for
every label in its closed set, so increase() has a baseline for a series that
moves once per boot. seed_failed is emitted whether or not
the gate is armed: with it off, that gauge is the only thing that says the seed
did not land.

Fresh-engine fix

  • Restacked onto the current [Router] k8s e2e for cache-aware peer bootstrap (12/13) #40698 (it was on an older copy of the stack).
  • The seed gate and skipped sweeps. With --kv-bootstrap-seed-required, the gate treats "ranks registered, no sweep verdict yet" as a boot sweep still in flight. A router booting alongside fresh engines now resolves every rank from origin, often before the coordinator takes the batch. retain_graftable then skips the sweep, so no verdict was ever recorded and /readyz held until the gate's hard bound. The coordinator now records ranks_resolved for a batch skipped this way, but only when no rank is still Pending. A Pending rank's own sweep is the one that reports, and an early verdict would open the gate ahead of it. Tests: coordinator::tests::a_batch_every_rank_left_before_its_sweep_still_records_a_verdict and an_emptied_batch_leaves_the_verdict_to_a_rank_still_pending.
  • RankOutcome::ALL / SweepOutcome::ALL, the label sets the metrics surface pre-registers, gain from_origin and ranks_resolved.

Tests

cargo fmt --check, cargo clippy --all-targets -- -D warnings, and the lib +
component + proxy suites all pass on this branch on its own, not only on the
tip of the stack.

Review pass

Reviewed with /code-review --fix and /simplify; fixes were folded into this PR's own commit and the stack was re-verified tier by tier (cargo fmt --check, cargo clippy --all-targets -D warnings, lib + component + proxy tests on every branch).

  • Fixed: the seed gate can no longer latch open before the boot sweep reports, and a first KV registration no longer un-readies a replica that is already serving.
  • Bootstrap budgets raised: the whole sync defaults to 10 minutes and each step (per-fetch cap/floor, connect/read timeouts, per-peer retry backoff) is sized for it.
  • File layout: this PR's code lives in server/routes/health.rs and server/routes/metrics.rs, plus state/kv_events/bootstrap/tracker.rs and bootstrap.rs; no file the stack creates exceeds ~1,300 lines.

🤖 Generated with Claude Code

https://claude.ai/code/session_016HmJvHV7QDPk3qAjQYzthd


CI States

Latest PR Test (Base): ✅ Run #36806942090
Latest PR Test (Extra): ❌ Run #36806941607
Latest PR Test (AMD ROCm 10): ➖ No AMD PR run found for this commit.

@Kangyan-Zhou
Kangyan-Zhou force-pushed the router-peer-bootstrap-12-k8s-e2e branch from 8bda96f to 5e5c384 Compare September 27, 2026 02:28
@Kangyan-Zhou
Kangyan-Zhou force-pushed the router-peer-bootstrap-13-readiness-gate branch from 088a6ca to 71f1211 Compare September 27, 2026 02:28
@Kangyan-Zhou
Kangyan-Zhou force-pushed the router-peer-bootstrap-12-k8s-e2e branch from 5e5c384 to f6d90bb Compare September 27, 2026 06:15
@Kangyan-Zhou
Kangyan-Zhou force-pushed the router-peer-bootstrap-13-readiness-gate branch 2 times, most recently from 2a3059b to d6f9086 Compare September 27, 2026 06:25
@Kangyan-Zhou
Kangyan-Zhou force-pushed the router-peer-bootstrap-12-k8s-e2e branch from f6d90bb to c51ab53 Compare September 27, 2026 07:04
@Kangyan-Zhou
Kangyan-Zhou force-pushed the router-peer-bootstrap-13-readiness-gate branch from d6f9086 to 8e63f1d Compare September 27, 2026 07:05
@Kangyan-Zhou
Kangyan-Zhou force-pushed the router-peer-bootstrap-12-k8s-e2e branch from c51ab53 to bd6a6b5 Compare September 27, 2026 09:07
@Kangyan-Zhou
Kangyan-Zhou force-pushed the router-peer-bootstrap-13-readiness-gate branch from 8e63f1d to 6dbda21 Compare September 27, 2026 09:07
@Kangyan-Zhou
Kangyan-Zhou force-pushed the router-peer-bootstrap-12-k8s-e2e branch from bd6a6b5 to 232520a Compare September 28, 2026 21:54
@Kangyan-Zhou
Kangyan-Zhou force-pushed the router-peer-bootstrap-13-readiness-gate branch 2 times, most recently from 6dbda21 to aba475c Compare September 28, 2026 21:54
kzhou-radixark and others added 5 commits September 29, 2026 14:35
…the pump

The producer half now has a consumer. A booting replica asks its siblings for a
cache-aware tree, vets what comes back, and hands it to the pump — which has
been able to graft one since the previous change but had nothing to graft.

The sweep retries rather than asking once. Worker discovery regularly completes
before the peer watch has delivered its first EndpointSlice list, so a single
pass sees zero candidates and abandons, and the joining replica boots cold
beside warm siblings. The bootstrap deadline bounds the whole search, not each
request, so a fleet of slow peers cannot outlast the readiness gate.

The per-fetch timeout is a strict fraction of that deadline, not the deadline
itself. `reqwest`'s total timeout would otherwise let the first unresponsive
peer starve every other candidate — the outer deadline then cancels the sweep
mid-fetch and the replica boots cold having tallied no peer outcome at all.
`--kv-bootstrap-fetch-timeout-cap-ms` bounds the derivation from above, so a
deliberately generous readiness budget still cannot park on one hung peer; the
router warns at startup when the derived value lands below the floor a
multi-megabyte body needs, because that failure otherwise looks like every peer
being unreachable with the configured cap looking blameless.

Accepting a snapshot is not the same as accepting a peer. A body can vet
cleanly and still know nothing about the ranks being bootstrapped, so the sweep
keeps looking rather than ending on it; a peer whose snapshot is permanently
incompatible is never re-fetched; and one that simply has nothing right now
sits out a few passes, because retrying every 250ms means re-downloading a
multi-megabyte tree from a replica that is itself serving traffic. The sit-out
doubles on each repeat miss by the same peer, up to 30s, so the sweep keeps
its 250ms pickup of new siblings without hammering one that keeps failing.

Every exit path — found, no peers, fleet cold, timed out — sends exactly one
`PumpControl`, which is what releases the ranks from `Pending`. The cold-fleet
verdict is discarded if the candidate set changed mid-pass: a peer that was
never consulted holds state that, unlike post-subscription events, cannot be
recovered later.

One sweep per discovered worker for now. A fleet's worth of workers therefore
fetches the same fleet-wide body once per worker; the coordinator that folds a
discovery burst into a single fetch is the next change.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JBP3reyKKk4TPeZuppMPmW
…for it

A new engine's ranks sit Pending on every router at once, and while a
rank is Pending its deltas stay off the tree. No sibling's tree ever
holds nodes for it, `covers_any` never succeeds, and each router waits
out the deadline on siblings that are waiting too; the whole-tree
`fleet_is_cold` verdict never fires because every sibling is warm.

- `PeerSnapshot` gains `empty_ranks`: live ranks the export holds no
  node for, including ranks the producer is itself still bootstrapping.
  Additive on the wire: it defaults to empty and is omitted when empty,
  so an older producer reads as "no evidence".
- A rank settles cold mid-sweep once every candidate names it there, or
  is hopeless as a whole. A silent, unanswered or unreachable peer
  vetoes. The sweep goes on for its other ranks.
- A re-fetch demands an export newer than that peer's last answer. A
  floor fixed at sweep start was met by the producer's cached export for
  the whole sweep, so a retry was served the same non-covering answer
  until the deadline. The producer already pins the demanded instant on
  arrival, so a herd still shares one build.
- Each pass sweeps only obligations still Pending on their own
  incarnation, and a sweep whose every rank left Pending (resolved from
  its origin, settled, forgotten) ends `ranks_resolved` and sends
  nothing.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0169aRN6guH335zF1pG456FC
A peer snapshot is fleet-wide: one body carries the whole worker table, every
cursor, and the whole tree, so a single fetch already contains everything every
pending rank needs. `add_worker` fires once per discovered engine, so sweeping
from there re-downloads that same document once per engine — and the document
grows with the fleet while the fetch count grows with it too. On a 168-engine
fleet that is 168 fetches of a multi-megabyte body per booting replica, most of
which time out and leave their ranks cold.

`add_worker` now hands its obligations to a coordinator that serialises
bootstrap into one sweep at a time and lets every rank pending when that sweep
lands share its snapshot. The fetch is not delayed to collect a batch first
because it does not need to be: the sweep IS the collection window — a
multi-megabyte transfer takes far longer than an EndpointSlice watch event
takes to deliver a fleet, so ranks discovered while it is in flight merge into
the same delivery at no latency cost.

Sharing is safe because the invariant is a SEQUENCE condition, not a wall-clock
one. Vetting runs against the live-worker set at the moment the body arrives, so
a worker discovered during the fetch is already covered; and the watermark check
adjudicates every rank independently, so a rank whose publisher did advance in
between is discarded to `Gap` rather than spliced over a hole.

Sharing one snapshot only helps if it is fresh enough for every rank sharing it,
and ranks in a batch begin holding at different moments — so the sweep asks with
the strictest floor among them. That guarantee covers the ranks it STARTED with;
a batch merged in afterwards was not represented in the request. Discovery
accepts that trade, a gap retry does not, which is what `LateJoin` distinguishes.

That retry is the other half of this change. A gap is the costliest failure —
a snapshot fetched, grafted, then thrown away because the live stream did not
join its watermark — and a fresher snapshot usually splices. It needs somewhere
to put the rank back, which is the queue this change introduces, so until now a
gapped rank was terminal. The tracker caps it at one retry per rank, and the
retry is stamped NOW and refuses to ride an in-flight sweep, or it would be
adjudicated against a snapshot taken before the gap and re-gap by construction.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JBP3reyKKk4TPeZuppMPmW
…olved

The coordinator folds ranks queued while a sweep runs into that sweep's
delivery. A sweep whose own ranks all left Pending without it (resolved
from their origin, settled cold per rank) ends `ranks_resolved` and
sends no control message, so a late joiner folded into it would stay
Pending with nothing left to release it, holding its batches and
/readyz. Hand whatever the sweep never spoke for back for a sweep of
its own.

A rank the sweep settled cold mid-sweep is owed nothing either, but it
can still read Pending until the pump drains that release, so the sweep
now reports the obligations it settled and the coordinator drops them
by obligation rather than by state. Otherwise such a rank would buy a
second sweep here, or a coverage retry after a `found` verdict.

Also pins that a rank back in Pending for a gap retry catches an
in-place engine restart against its held queue: the retry keeps the
first attempt's batches and no cursor.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0169aRN6guH335zF1pG456FC
A graft whose held queue was empty leaves its watermark unchecked until a live
batch arrives, and nothing guarantees one ever does. An idle rank would serve
grafted state forever with its continuity unproven, and a `BlockRemoved` lost in
the subscribe window would survive as a permanent false cache hit that no later
event corrects.

Silence is not evidence, so the bounded wait ends in a QUESTION, not a verdict.
Reaching the deferred path means nothing arrived between subscribing and
grafting, and the subscriber is live before the fetch — so silence is far more
often "this rank published nothing" than "we lost a delta". Discarding on a timer
would cost every quiet fleet its warm tree, which is the regression this feature
exists to prevent.

So the pump asks the fleet whether the publisher moved past the watermark. Any
peer's cursor is admissible: sequence numbers are the publisher's, so a peer
reporting one above ours proves a batch was emitted that we never saw. A peer too
cold to bootstrap from is still a valid witness, which is why the probe reads the
wire cursor directly instead of vetting. It asks for the cursor table alone, not
a snapshot — the question is answered completely by one integer per rank, and
fetching a tree to read it made the proof cost scale with the tree, so a fleet
large enough to need bootstrap was also the fleet that could not afford to prove
it.

Only a witness ABOVE the watermark discards. A graft no batch and no peer can
speak to is KEPT, tallied `warm_unwitnessed` rather than `warm`, after
`MAX_UNKNOWN_PROBES` unanswerable rounds — without a stop, a single-replica
deployment would re-probe forever and its verdict would never resolve in the
metrics.

The probe runs off-pump because it does network I/O; the verdict comes back over
the pump's own control channel so the tree write stays on the single writer. An
outstanding probe blocks relaunch inside its timeout and stops blocking past it:
a verdict that never lands (a panicked probe task, say) would otherwise freeze
the rank in `Recovered` with no outcome ever recorded.

Also re-attaches `demote_unproven_rank`'s doc comment, which sat above
`requeue_gapped_rank` describing the wrong function.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YZorgAox1CpLNHSjzpcxdb
kzhou-radixark and others added 5 commits September 29, 2026 14:43
…robe

A splice probe answers for the graft it was launched about. Once a
batch 0 has replaced that graft with a restarted engine's stream, a
verdict still in flight, whose witnesses count in the dead numbering,
must not demote the new stream's state.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0169aRN6guH335zF1pG456FC
…t copied

Component-level proof over the real transport: an axum server serving
`/internal/kv_snapshot`, a reqwest client fetching it, JSON over a loopback
socket. The interesting failure modes live in the wire format and the vetting
step, not in the tree algebra — that is already covered by the unit tests on
`export_snapshot` / `restore_snapshot`.

The equivalence assertion is deliberately behavioural: for a large query set,
`match_prefix` must return the same matched length and the same carrier set on
both replicas. That is the property routing actually depends on; comparing node
counts alone would pass while routing diverged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YZorgAox1CpLNHSjzpcxdb
A kind-cluster proof of the two questions the feature exists to answer: a
replica joining a warm fleet ends up with the same cache-aware view as the
replicas it bootstrapped from, and it keeps up with new engine events afterwards
— the snapshot is spliced UNDER the live stream, not substituted for it. Plus
the rolling-update hazard: new replicas must not bootstrap from each other and
inherit an empty tree.

Ships the KV-publishing fake worker the test drives, and the manifest that wires
`POD_NAME`/`POD_NAMESPACE`/`POD_IP` from the downward API and the EndpointSlice
RBAC the peer watch needs — deferred from the discovery commit to here, where
there is something for it to exercise.

Three anti-flake decisions, because the naive version of this test is flaky:
events are driven by a POST rather than a timer, so nothing is ever in flight
the test did not ask for; views are compared only after every replica's cursor
has provably caught up to the worker's `last_seq`, never on a fixed sleep; and
nothing asserts on a transient mid-rollout state — the scale-up case is asserted
directly and the rollout case only after `rollout status` completes. Views are
compared canonically, since snapshot node order follows per-shard hash-map
iteration.

Not yet run: this needs a kind cluster, so it is verified by construction only.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YZorgAox1CpLNHSjzpcxdb
…vable

A replica that boots with an empty KV tree does not just route cache-blind
itself — its dispatches scatter the prefixes the warm replicas were keeping
consolidated, so it degrades the engines' locality for the whole fleet. A
rolling update does that to every replica in turn, and with nothing in the
metrics to attribute the hit-rate drop to.

`/readyz` grows two conditions. The third holds 503 until peer bootstrap has
settled: a snapshot grafted, or every sibling having proved it has nothing, or
`--kv-bootstrap-timeout-ms` giving up. It latches, so a later scale-up can
never drag an already-serving replica back out of the Service. It is inert
unless both cache-aware routing and a peer selector are configured, so nothing
changes for a fleet that has not opted in.

The fourth is `--kv-bootstrap-seed-required`, off by default, and it exists
because condition 3 settles either way once the deadline expires — that escape
is load-bearing, and it is also why `settled()` cannot tell "seeded" from "gave
up". This one can: it refuses readiness when a sweep proved siblings were there
and their tree could not be pulled. A NotReady pod leaves the Service, so a
failed seed then STALLS a rolling update with the previous generation still
serving, instead of completing it with cache-blind replicas.

Which verdict may gate is the entire safety argument. `NoPeers` (first deploy,
single replica) and `FleetCold` (every sibling PROVED it holds nothing) are
legitimate nothing-to-inherit outcomes; gating on either bricks a boot that has
nothing to inherit — and because an unready replica leaves its own
EndpointSlice, every sibling would then see an empty peer set: a fleet-wide
deadlock with no automatic exit. Only `TimedOut` over a non-empty candidate set
counts. Two further rails cover the case the verdict genuinely cannot
distinguish, a simultaneous fleet-wide restart where every peer is booting: the
hold is bounded at max(3x --kv-bootstrap-timeout-ms, 60s) and then opens with a WARN,
and it latches open the first time `/readyz` answers 200.

The tests assert over the verdict class rather than by example, because the
wrong answer in either direction is silent — gate too little and a bad rollout
ships, gate too much and the fleet deadlocks.

The metrics make the whole path attributable. `bootstrap_state` per rank, so a
rank stuck pending (holding its events back, heading for overflow) is visible;
`peer_snapshot_total` by fetch outcome, where a fleet pinned at `unreachable`
with a warm tree means the per-fetch bound is too small for the body rather
than a network fault; `bootstrap_rank_total` once per rank at its final
verdict; `bootstrap_sweep_total` per sweep, which is what separates a healthy
early settle on a cold fleet from burning the whole deadline — the rank
outcomes fold both into `abandoned`. `seed_failed` is emitted whether or not
the gate is armed: with it off, that gauge is the only thing that says the seed
did not land.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YZorgAox1CpLNHSjzpcxdb
Under --kv-bootstrap-seed-required the gate reads "ranks registered, no
sweep verdict yet" as a boot sweep still in flight. A router booting
alongside fresh engines now resolves every rank from its stream's
origin, often before the coordinator takes the batch, and
`retain_graftable` then skips the sweep entirely: no verdict is ever
recorded and /readyz holds until the gate's hard bound. Record
`ranks_resolved` for a batch skipped that way, the verdict a sweep
started a moment later would reach before its first fetch.

Only when no rank is still Pending, though: such a rank's own sweep is
the one that reports, and an early verdict here would open the gate
ahead of it, letting a probe latch the replica ready before that
sweep's `timed_out` marks the seed failed.

Also lists `from_origin` and `ranks_resolved` in the outcome label sets
the metrics surface pre-registers.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0169aRN6guH335zF1pG456FC
@Kangyan-Zhou
Kangyan-Zhou force-pushed the router-peer-bootstrap-12-k8s-e2e branch from 232520a to b8ab8e6 Compare September 29, 2026 21:48
@Kangyan-Zhou
Kangyan-Zhou force-pushed the router-peer-bootstrap-13-readiness-gate branch from aba475c to b4d3ebf Compare September 29, 2026 21:48
Base automatically changed from router-peer-bootstrap-12-k8s-e2e to main October 1, 2026 01:43
@ShangmingCai
ShangmingCai marked this pull request as ready for review October 1, 2026 03:48
@ShangmingCai
ShangmingCai merged commit 3b6d37a into main Oct 1, 2026
94 of 98 checks passed
@ShangmingCai
ShangmingCai deleted the router-peer-bootstrap-13-readiness-gate branch October 1, 2026 03:55
ShangmingCai added a commit that referenced this pull request Oct 1, 2026
Follow-ups from the oracle review of #40687-#40699, the mechanical ones:

- The snapshot client no longer follows redirects. A sibling router never
  redirects /internal/kv_snapshot, so a 3xx is a misconfigured or hostile
  peer steering the fetch - and its multi-gigabyte buffering budget - at an
  arbitrary in-cluster URL. It now lands as FetchAnswer::NoBody and the
  target is never contacted; asserted against the index's own client, since
  a test-local client would prove nothing about the one the sweep uses.
- The cursors-only route answers 503 on an encode failure, like the full
  export already did, instead of a 200 carrying an empty body the caller
  cannot decode. Unreachable in practice; kept so the paths cannot drift.
- PeerRegistry::replace logs the peer count at info and the full URL list
  only at debug - on a large fleet every rolling update re-lists every
  sibling.
- monitoring/README.md documents the seven bootstrap series the readiness
  gate added, with their label sets.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
ShangmingCai added a commit that referenced this pull request Oct 1, 2026
Resolves the test-module conflict in src/server/routes/health.rs against
#40699 (readiness gated on peer bootstrap): keeps this branch's
`readiness_rejects_portless_prefill` alongside main's new `readyz_status`,
`ctx_with_tracker`, `ctx_with_seed_gate` and `add_test_worker` helpers.
`readyz` still resolves through `PdPoolResolver` -> `healthy_workers_for`,
so the portless-prefill exclusion keeps gating readiness.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
ShangmingCai added a commit that referenced this pull request Oct 1, 2026
Resolves ten conflicts left by #41610, #41611 and #41612 landing on main while
this branch still carried their pre-squash commits. Two needed judgement rather
than a side:

- `readyz`: main gained the peer-bootstrap seed gate (#40699) while this branch
  rewrote the pool check for version groups. The hunks overlap, and taking
  either side whole drops the other, so the resolution keeps this branch's
  group-aware `pool_ready` *and* main's `&& ctx.kv_bootstrap_admit_ready()`.
- `chat/forward.rs`: this branch carries #41612's pre-fix copy; main has the
  post-fix one. The version-group feature does not touch this file at all, so
  it takes main's version, keeping the `cancelled` outcome #41612 records for a
  decode dispatch that prefill's failure abandoned.
- `pd_bootstrap_injection/reliability.rs` (add/add): keeps this branch's
  `..Default::default()` spec and main's two #41612 assertions.

The rest are the mechanical `..Default::default()` churn this branch introduced
meeting fields main spells out, plus a `readiness_rejects_portless_prefill` body
git aligned against its own merged copy; resolved so each test and helper is
defined once.

Verified on the merge result: 1017 lib + 129 component + 157 proxy tests pass,
clippy -D warnings and rustfmt clean. The two fixes a careless resolution would
have reverted are covered by tests that still pass -- the five
`pd_bootstrap_injection::reliability` cases and the four seed-gate `readyz`
cases.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants