Repository navigation
fix(runtime): fix QUIC reassembly and reduce response stalls - #14876
Conversation
Signed-off-by: jthomson04 <jwillthomson19@gmail.com>
Signed-off-by: jthomson04 <jwillthomson19@gmail.com>
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Enterprise Run ID: ⛔ Files ignored due to path filters (4)
📒 Files selected for processing (2)
Included review availability: Your plan provides up to 12 included reviews per hour; 10 remain after this review. WalkthroughThe workspace updates Quinn from 0.11.9 to 0.11.12. QUIC server and client failure logs now use alternate error formatting. Server lane warnings also include the connection close reason. ChangesQUIC updates
Priority: ⬇️ Low Estimated code review effort: 1 (Trivial) | ~5 minutes Merge Risk: ⚪ Minimal · up to The dependency versions are consistent, and the diagnostic changes preserve QUIC cleanup and registration-error behavior without demonstrated production impact. 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
Full details: Docstring CoverageExplanation Docstring coverage is 50.00% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 2 functions across 1 files. (1 skipped: 1 unsupported.)
Comment |
jthomson04
left a comment
There was a problem hiding this comment.
Automated review (Claude Opus 5, xhigh effort). Every claim below was checked against the quinn 0.11.9 / quinn-proto 0.11.17 sources in the local cargo registry and against the published quinn-proto 0.11.18 crate; line references are to head 2bcd9ae.
Six inline comments. The two I'd want answered before merge:
- The new
close_reasonfield is overwritten by our own teardown before most lanes read it, so it usually reports our close rather than the cause it was added to expose. - 0.11.18 carries behavior changes beyond the reassembly fix —
send_windowaccounting and ACK bundling — that the Validation section doesn't cover.
One item with no diff line to anchor to: no regression test pins either half of this PR. The fix lives entirely in a transitive crate version, and the existing quic_response::tests already ran real local QUIC connections without catching the original failure — they don't push enough small contiguous frames through one lane to reach COMPACT_THRESHOLD = 2 * MAX_CHUNKS = 2048 retained chunks on the receiver (quinn-proto src/connection/assembler.rs:358-364). A test that streams >2048 small frames through a single lane with a slow reader and asserts the stream survives would fail on 0.11.17, pass on 0.11.18, and would also catch a future Cargo.toml relaxation from 0.11.12 back to 0.11.
Signed-off-by: jthomson04 <jwillthomson19@gmail.com>
|
Review follow-up in For the review-summary request for a >2,048-frame regression test: application frames do not map one-to-one to retained Quinn assembler chunks. Packetization, buffering, and consumption affect that state, so the proposed count alone does not make the test deterministic. The existing upstream reproduction inserts small chunks directly: it fails with The CodeRabbit docstring-coverage warning is advisory for existing private functions. No documentation expansion is needed for this change. |
dmitry-tokarev-nv
left a comment
There was a problem hiding this comment.
Reviewed at 24001ac. Intent, as stated in the description and readable from the diff: move quinn 0.11.9 -> 0.11.12 / quinn-proto 0.11.17 -> 0.11.18 to pick up the stream-assembler coalescing fix, and widen QUIC failure reporting from a single-line Display to the full cause chain plus bundle / lane / peer identity.
What I checked by execution, on macOS ARM64 with the pinned toolchain (rust-toolchain.toml = 1.96.1):
Lockfiles resolved, not hand-merged. find . -name Cargo.lock returns exactly the four files this PR touches, and all four move both crates together. The only manifest naming quinn is the root Cargo.toml; lib/runtime/Cargo.toml takes it from the workspace, so there is no second pin to drift. I downloaded the four .crate files from crates.io and the SHA-256s match the lockfile lines byte for byte (quinn 0.11.12 4051e23e…d9d11, quinn-proto 0.11.18 a9746dbd…c9fc). Diffing the published manifests: quinn-proto 0.11.18 changes no dependencies at all, and quinn 0.11.12 changes only dev-dependencies and a test-only feature, its one runtime change being proto 0.11.12 -> 0.11.18. That is exactly the shape the lock diff shows — no dependency-list churn, which is what a hand-merge would have gotten wrong.
The fix is actually in the bump. quinn-proto-0.11.18/src/connection/assembler.rs introduces MIN_RETAINED_CHUNK_SIZE = 128 and coalesces any chunk below max(buffered.div_ceil(MAX_CHUNKS), MIN_RETAINED_CHUNK_SIZE) during defragment(), and ships ordered_lossless_stream_is_not_rejected as its regression test. That is the change the description points at.
MSRV moves. quinn's rust-version goes 1.74.1 -> 1.85 (quinn-proto was already 1.85). No impact at 1.96.1, but it is a real constraint change.
Tests. cargo test --locked -p dynamo-runtime --lib pipeline::network::quic_response -- --test-threads=8 -> 21 passed here against the description's 22; the gap is the #[cfg(target_os = "linux")]-gated test at line 3106, not the diff.
One P2 inline below. I also left replies with independent verification on the five open threads, including one measurement that retracts part of an earlier claim.
dmitry-tokarev-nv
left a comment
There was a problem hiding this comment.
Approving at 24001ac. Nothing I found is blocking: zero P0/P1, two non-blocking items, both named below.
Verified by execution, not by reading (macOS ARM64, toolchain pinned at 1.96.1):
- Lockfiles are re-resolved, not hand-merged. All four
Cargo.lockfiles in the repo are the four touched here, and they move both crates together. The rootCargo.tomlis the only manifest naming quinn (lib/runtimetakes the workspace dep), so there is no second pin to drift. I fetched the four.cratefiles from crates.io and their SHA-256s match the lockfile lines exactly. Diffing the published manifests: quinn-proto 0.11.18 changes no dependencies at all, and quinn 0.11.12 changes only dev-dependencies plus a test-only feature, its one runtime change beingproto0.11.12->0.11.18— which is precisely why the lock diff has no dependency-list churn. - The fix is in the bump.
assembler.rsin 0.11.18 addsMIN_RETAINED_CHUNK_SIZE = 128, coalesces sub-threshold chunks indefragment(), and shipsordered_lossless_stream_is_not_rejected. - Transport-parameter strictness is receive-side only (
transport_parameters.rs: aremaining_before - r.remaining() != lencheck rejectingError::Malformed). Both 0.11.17 and 0.11.18 write well-formed parameters, so a mixed-version fleet during a rolling upgrade does not break on it. - MSRV moves 1.74.1 -> 1.85 on quinn; no impact at 1.96.1.
cargo test --locked -p dynamo-runtime --lib pipeline::network::quic_response -- --test-threads=8-> 21 passed. The description's 22 is the#[cfg(target_os = "linux")]test at line 3106, not a discrepancy.
Outstanding, non-blocking:
- P2 — line 1972:
{error:#}now carries a peer-supplied CONNECTION_CLOSE reason phrase intoreason, and%reasonis written unescaped by the console formatter, so a peer can inject newlines and ANSI into the worker's log. Measured over a real connection on the pinned crates: base string 15 bytes, head string 1201 bytes with an embedded newline, from a 65 KiB reason the peer offered. Size-bounded, so this is a log-integrity issue rather than a resource one. One-token fix (%reason->reason), verified to compile, test and lint clean. - P3 — the description covers quinn-proto 0.11.18 thoroughly but not quinn's own 0.11.9 -> 0.11.12 span, which changes connection-handle refcounting,
RecvStreamstop/drop, and — relevant to this file's cooperative-budget comment —drive_timerno longer trustingSleep::pollfor expiry. Details on theCargo.lockthread.
Deliberately deferred, as decisions rather than oversights — I have replied on each with independent verification:
- Benign-vs-invariant classification of lane exits and the unreachable
close_reason().is_none()guard in the accept loop. Confirmed unreachable on 0.11.12 too, and confirmed pre-existing. - Coordinated first-cause capture for
close_reason; the description now labels it best-effort. Keeping the explicit close on 938-939 is correct —fail_server_connectionreturns early when no bundle is registered, which is exactly the pre-registration failure case. send_prologue'sto_string(): I followedenqueue_onand every arm is a flat error with nosource(), so{error:#}there would be inert. Correct to leave alone.send_windowaccounting: theunacked()->buffered() = unacked_len + front_trimmedchange and the 8 x 1.25 MB priority-connection arithmetic both check out, but nothing I have says it bites in practice. Accepting the deferral.
I authored no commits on this PR.
There was a problem hiding this comment.
Previously reported defects still present:
- Original discussion: Still present: the new
%reasonlog field renders peer-controlled QUIC close-reason text without escaping. A peer can inject newlines or terminal escapes into worker logs; recordreasonas a normal string field instead. - Original discussion: Verified still present: the new full cause chain can contain a peer-controlled QUIC CONNECTION_CLOSE reason, and
%reasonwrites it unescaped to the worker log. A peer can inject physical log lines or ANSI control sequences. Recordreasonas a normal string field instead of a Display field so tracing escapes it. - Original discussion: Verified still present: the newly expanded Quinn error chain can contain a peer-supplied CONNECTION_CLOSE reason, and
%reasonlogs it without escaping. Newlines or ANSI escapes can forge physical log records. Recordreasonas a normal string field instead of a Display field.
SummaryTyche follow-up: 8 versus 32 QUIC receive endpointsThe focused low-concurrency comparison found no repeatable throughput, total-p99 latency, or serving-CPU regression above the agreed limits from increasing the frontend receive endpoint count from 8 to 32. There is a material memory cost. Scope: these measurements use separate campaign commits, both with the reader scheduling change and other common frontend campaign changes. They isolate the endpoint count. PR commit Validation
Across all 18 pairs, the worst observed regressions were 3.10% in throughput, 1.60% in total p99, and 0.17% in serving CPU per completed request. All six case medians satisfy the 5% throughput/total-p99 and 10% serving-CPU limits. Percentages in the table are the median of three paired changes (
Costs and remaining limits
Client “inter-token latency” is each request's average Source records and exclusions
The source difference between these arms is the Linux Five attempts are outside the 36-run comparison: two initial checks; one completed unpaired run retained after correcting the endpoint-order schedule; one infrastructure-interrupted run; and that run's completed partner. The interruption was an expired Kerberos ticket that prevented writing result files. Both arms were rerun together on the same four nodes after access was restored. All 40 completed runs were audited; the failed attempt and its artifacts remain retained. Conclusion: this supports the low-concurrency performance of 32 endpoints for the tested short-response workload, subject to its memory cost. The earlier high-concurrency QUIC/TCP gaps and other qualification gaps remain. It does not qualify a default transport change. All 36 selected runs: per-run CSV with sample counts, latency tails, resource costs, and UDP drops
workers,concurrency,endpoints,repetition,completed,errors,requests_s,output_tokens_s,request_latency_p50_ms,request_latency_p95_ms,request_latency_p99_ms,time_to_first_token_p50_ms,time_to_first_token_p95_ms,time_to_first_token_p99_ms,inter_token_latency_p50_ms,inter_token_latency_p95_ms,inter_token_latency_p99_ms,serving_cpu_ms,node_cpu_ms,frontend_rss_mib,frontend_fds,udp_drops,resource_window_start_ns,resource_window_end_ns,resource_window_successes
1,1,8,0,40301,0,672.311206,21513.958584,1.314816,1.501280,1.551775,0.702560,0.764863,0.811392,0.019755,0.025838,0.027447,3.001865,2.890298,148.187500,56,0,1789524970647513639,1789525016757260612,30999.634341
1,1,32,0,40267,0,671.745733,21495.863468,1.302560,1.488000,1.539936,0.700128,0.757868,0.797259,0.019445,0.025479,0.027123,2.979117,2.872005,146.562500,80,0,1789525037670960228,1789525083786712182,30998.911076
1,16,32,0,256448,0,4276.562195,136849.990231,3.602303,4.702996,5.274306,2.766943,3.888020,4.480336,0.025517,0.055019,0.067382,2.036449,1.997514,171.125000,187,0,1789525106113746602,1789525152238365358,197437.518299
1,16,8,0,259502,0,4326.298040,138441.537288,3.559071,4.642141,5.191230,2.633840,3.679807,4.239807,0.028940,0.057896,0.070608,2.038616,2.000395,165.750000,166,0,1789525174317828656,1789525220434865409,199735.280685
1,64,8,0,270900,0,4516.840764,144538.904450,13.887567,19.599431,22.278692,12.386603,18.139878,20.802150,0.009709,0.201083,0.278205,1.679423,1.648040,164.687500,277,38,1789525243894771257,1789525290011193244,208443.980915
1,64,32,0,274871,0,4582.152633,146628.884270,13.676829,19.590874,22.398548,11.885019,17.669115,20.372152,0.010770,0.219389,0.300558,1.671521,1.644911,170.562500,302,2383,1789525312966130233,1789525359078514759,211455.335348
8,1,32,0,39663,0,661.658387,21173.068373,1.285312,1.524736,1.578476,0.749215,0.833312,0.903898,0.017229,0.025110,0.026734,3.186790,3.101935,152.312500,108,0,1789525385117292616,1789525431228932918,30387.677396
8,1,8,0,38677,0,645.208886,20646.684361,1.299584,1.551136,1.605919,0.756992,0.837600,0.907872,0.017380,0.025709,0.027426,3.220085,3.126404,129.812500,84,0,1789525454562775261,1789525500677299696,29667.821223
8,16,8,0,247138,0,4121.144947,131876.638300,3.749264,4.958208,5.535793,3.158575,4.512517,5.096689,0.012764,0.046610,0.060100,2.836586,2.721312,158.562500,189,0,1789525521666971215,1789525567779095854,190248.120734
8,16,32,0,244968,0,4083.918142,130685.380533,3.794207,4.967295,5.542291,3.349759,4.619604,5.186207,0.008530,0.039648,0.053376,2.820904,2.704909,188.875000,216,0,1789525591318031893,1789525637414974224,188486.099893
8,64,32,0,250630,0,4178.710296,133718.729476,14.912157,22.231661,25.482072,13.045870,20.392311,23.848720,0.035061,0.191769,0.278920,2.466876,2.380057,219.750000,321,0,1789525661560566534,1789525707674020928,192608.081587
8,64,8,0,249972,0,4168.327930,133386.493761,14.984862,22.159806,25.346898,12.785951,20.067240,23.446861,0.046643,0.212636,0.296616,2.462665,2.384964,169.437500,298,1352,1789525732755701131,1789525778872724925,192300.736080
8,64,8,1,253245,0,4222.834455,135130.702564,14.698816,22.140304,25.937699,12.544800,20.055930,23.659683,0.045221,0.210811,0.302523,2.464152,2.381258,165.750000,297,4885,1789525922560536202,1789525968679523619,194855.611945
8,64,32,1,253326,0,4223.104168,135139.333377,14.740077,21.999684,25.336211,12.919038,20.173437,23.735292,0.033187,0.189772,0.270029,2.459981,2.368423,220.312500,320,0,1789525991697547465,1789526037808301911,194831.474849
8,16,32,1,250114,0,4169.777690,133432.886080,3.709952,4.854378,5.407658,3.244703,4.458294,5.008162,0.008884,0.041495,0.055708,2.802952,2.682367,206.562500,217,0,1789526061784590687,1789526107902390942,192493.403266
8,16,8,1,248481,0,4142.414403,132557.260891,3.736768,4.905343,5.459027,3.229184,4.533663,5.084434,0.010330,0.042712,0.056151,2.837454,2.722922,161.312500,191,0,1789526132994797351,1789526179114950946,191429.079378
8,1,8,1,39263,0,654.984956,20959.518594,1.297536,1.542655,1.598144,0.755392,0.833152,0.894801,0.017473,0.025484,0.027187,3.225118,3.116084,164.562500,84,0,1789526201096146197,1789526247187165877,30081.274166
8,1,32,1,39795,0,663.850242,21243.207739,1.280160,1.526889,1.580168,0.755807,0.843562,0.890279,0.016850,0.025040,0.026741,3.188375,3.074904,169.687500,108,0,1789526268780017007,1789526314889850208,30460.734749
1,64,32,1,274538,0,4577.825831,146490.426604,13.708306,19.293765,21.972770,12.307070,17.880617,20.505783,0.009221,0.191760,0.267103,1.659412,1.628712,184.625000,302,105,1789526336840429501,1789526382960702289,211347.537946
1,64,8,1,272025,0,4534.878021,145116.096681,13.860543,19.303199,21.888111,12.445120,17.915653,20.505857,0.009561,0.189069,0.264055,1.666420,1.633931,165.375000,275,20,1789526405931084367,1789526452041951087,209265.171030
1,16,8,1,259752,0,4331.652805,138612.889761,3.554431,4.646431,5.219216,2.631936,3.683118,4.256702,0.028815,0.058002,0.070938,2.019204,1.978417,165.437500,163,0,1789526476010036041,1789526522120914065,199976.520228
1,16,32,1,261023,0,4351.516771,139248.536661,3.540767,4.606080,5.160792,2.673311,3.711196,4.252056,0.026982,0.055362,0.067807,2.019060,1.977315,150.625000,186,0,1789526545092500040,1789526591206156032,200822.750462
1,1,32,1,39002,0,650.644496,20820.623863,1.273824,1.495069,1.549792,0.721824,0.791840,0.845886,0.017821,0.025076,0.026805,3.010175,2.888805,147.937500,80,0,1789526998193199566,1789527044309763637,30064.153997
1,1,8,1,38992,0,650.475336,20815.210744,1.306736,1.508750,1.561311,0.724448,0.822240,1.305579,0.018230,0.025413,0.027002,3.035208,2.916143,171.625000,56,0,1789527066217924158,1789527112343419093,30021.419966
1,1,8,2,37984,0,633.662498,20277.199945,1.319664,1.516922,1.567472,0.729088,0.789664,0.840188,0.019069,0.025469,0.027015,3.026477,2.983596,156.437500,56,0,1789527134822856736,1789527180945402012,29072.038292
1,1,32,2,40501,0,675.646467,21620.686957,1.312960,1.483168,1.533696,0.691872,0.751328,0.787072,0.020077,0.025542,0.027160,2.975070,2.843094,183.500000,80,0,1789527202854765442,1789527248979964297,31416.716741
1,16,32,2,257981,0,4300.807151,137625.828839,3.580607,4.673791,5.244589,2.764672,3.875806,4.469285,0.024809,0.054324,0.067249,2.022255,1.982738,169.562500,187,0,1789527270377336869,1789527316500389103,198556.619224
1,16,8,2,258024,0,4302.441356,137678.123389,3.581246,4.671870,5.242378,2.701312,3.758681,4.302959,0.027361,0.056237,0.069034,2.044957,2.007965,148.250000,167,0,1789527338468775932,1789527384584470228,198645.529631
1,64,8,2,266144,0,4437.824237,142010.375570,14.149773,19.937841,22.646613,12.544574,18.316874,20.997775,0.010073,0.209471,0.288597,1.701626,1.676853,154.062500,272,989,1789527406486319642,1789527452598044520,204953.179433
1,64,32,2,274944,0,4583.504132,146672.132236,13.686848,19.449773,22.140856,12.175792,17.986961,20.724008,0.009302,0.199603,0.270842,1.669362,1.636541,165.375000,296,0,1789527477240352523,1789527523356660803,211611.338672
8,1,32,2,38984,0,650.330330,20810.570564,1.303824,1.535872,1.596229,0.750080,0.826912,0.880086,0.017801,0.025363,0.027034,3.178618,3.111227,182.562500,108,0,1789527549414116300,1789527595527476906,29819.398335
8,1,8,2,40229,0,671.106516,21475.408523,1.278016,1.531091,1.593942,0.736256,0.814656,1.042342,0.017284,0.025689,0.027418,3.230697,3.109662,174.312500,84,0,1789527619113114235,1789527665224543473,30970.068348
8,16,8,2,245750,0,4096.869864,131099.835646,3.764352,5.044736,5.719106,3.206271,4.599008,5.267683,0.011700,0.046265,0.060583,2.824211,2.702484,188.000000,187,0,1789527686209400726,1789527732328565657,189126.237809
8,16,32,2,248120,0,4137.448090,132398.338865,3.729951,4.942338,5.550816,3.229855,4.522273,5.124303,0.009879,0.043277,0.057767,2.800715,2.684331,193.437500,221,0,1789527754471179452,1789527800591548422,191143.085422
8,64,32,2,250905,0,4183.217128,133862.948085,14.881435,22.272586,25.603769,12.976189,20.304303,23.940368,0.035182,0.199322,0.285460,2.479758,2.394273,219.812500,325,0,1789527825658066777,1789527871784134645,193036.548396
8,64,8,2,249199,0,4155.983807,132991.481831,15.071646,22.183645,25.257951,12.927231,20.240748,23.497567,0.044935,0.207177,0.292127,2.475827,2.383605,183.437500,304,1211,1789527896798410249,1789527942915525522,191616.682584 |
There was a problem hiding this comment.
Previously reported defects still present:
- Original discussion: Still present: the new
%reasonfield logs a peer-controlled QUIC close reason without escaping. Embedded newlines or terminal escapes can forge physical worker-log records; recordreasonas a normal string field instead. - Original discussion: Verified still present:
fail_client_connection_bundlestill logs the peer-influenced full QUIC close reason as%reason, which renders without string escaping in the worker warning. Recordreasonas a normal string field instead. - Original discussion: Verified still present: the new full QUIC cause chain can include a peer-controlled CONNECTION_CLOSE reason, and
%reasonemits it unescaped in the worker warning. Newlines or ANSI escape sequences can forge log lines or alter terminal output; recordreasonas a normal string field instead. - Original discussion: Verified still present: the new full cause chain can include a peer-controlled QUIC close reason, and
%reasonwrites it unescaped to the worker log. Embedded newlines or terminal escapes can forge log records; recordreasonas a normal string field instead.
Questions for the author:
- Cargo.toml: For the supported QUIC response path, can you provide a baseline-versus-head measurement using this PR's exact dependency update and the shipped 5 ms batch interval? The supplied Tyche results show a higher short-response p99 but intentionally do not isolate the Quinn update or test the shipped configuration.
Drive the reverse-control writer in the receive task and charge the cooperative budget at 16-frame or 4 KiB thresholds. Spread Linux frontend receive traffic across 32 reuse-port endpoints. Port the measured receive-path changes from 2c41ea8476, 901cb517bc, and f969ff855f. Retain the existing PR diagnostics and batching configuration. Signed-off-by: jthomson04 <jwillthomson19@gmail.com>
There was a problem hiding this comment.
Previously reported defects still present:
- Original discussion: Still present: the new full Quinn cause chain can include a peer-controlled CONNECTION_CLOSE reason, and
%reasonrecords it without string escaping. Embedded newlines or terminal escapes can forge worker log records; recordreasonas a normal string field instead. - Original discussion: The worker warning still logs the peer-influenced full QUIC close reason with
%reason, which permits embedded newlines or terminal escapes to forge log records. Recordreasonas a normal string field. - Original discussion: Verified still present:
fail_client_connection_bundlestill records the peer-influenced full QUIC close reason as%reason, which renders without string escaping in the worker warning. Recordreasonas a normal string field instead. - Original discussion: Verified still present: the new full Quinn cause chain can include a peer-controlled CONNECTION_CLOSE reason, and
%reasonlogs it without string escaping. Embedded newlines or terminal control sequences can forge worker log records; recordreasonas a normal string field instead.
Questions for the author:
- Cargo.toml: For the supported QUIC response path, please provide a baseline-versus-head measurement using this PR's exact Quinn update and the shipped 5 ms batch interval. The supplied comparisons use a zero batch interval and additional campaign changes, so they do not isolate the changed transport behavior or assess the observed short-response tail increase under the shipped configuration.
Signed-off-by: jthomson04 <jwillthomson19@gmail.com>
|
Review follow-up is complete in 1ec03db608: the worker log field now escapes peer-supplied control characters, and the description covers Quinn's own release changes. All 22 QUIC tests, Clippy (no warnings), formatting, diff checks, and the local log-escaping check passed. The request for a baseline-versus-head benchmark with the shipped 5 ms batch interval is explicitly deferred for this PR. No new performance benchmark was run. The retained Tyche results use a zero batch interval and common campaign changes; they do not validate this exact PR with shipped settings or establish latency parity. The performance-validation gap remains open. TCP remains the default. |
There was a problem hiding this comment.
Previously reported defects still present:
- Original discussion: Verified still present: when
accept_bi()returns an error, Quinn has already setclose_reason, so theclose_reason().is_none()guard prevents the frontend accept-loop warning from ever being emitted. Transport failures therefore remain silent at this point.
* feat: KV DC Relay file based source mode (ai-dynamo#14807) Add live-reloaded file sources for KV DC Relay namespace selection and expose readiness and source revisions through /engine/state. Preserve applied membership on invalid updates, coalesce discovery refreshes, and isolate native integration tests in forked processes. Signed-off-by: Nikita Sukharev <kaonael@gmail.com> * feat(sglang): expose cross-encoder reranking through /v1/rerank (ai-dynamo#14032) Signed-off-by: xianlubird <xianlubird@gmail.com> * fix(profiler): explain inaccessible model paths during trust checks (ai-dynamo#14860) Signed-off-by: hongkuanz <hongkuanz@nvidia.com> * fix(sglang): sync discovery from native pause state (ai-dynamo#13951) Signed-off-by: William Arnold <7565007+Aphoh@users.noreply.github.com> Co-authored-by: Zero Rains <57100978+zeroRains@users.noreply.github.com> * feat(recipes): add Solar Open2 250B NVFP4 aggregated and disaggregated recipes for B200 (ai-dynamo#14376) Signed-off-by: Sandhya Rani Narravula <snarravula@nvidia.com> * refactor(agents): session_id reader from AgentContext + forward to vLLM (ai-dynamo#14428) Signed-off-by: Karen Chung <karenc@nvidia.com> Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com> * fix(discovery): allow served aliases for the same model source (ai-dynamo#14857) Signed-off-by: jthomson04 <jwillthomson19@gmail.com> * fix(router): reject unknown explicit worker targets (ai-dynamo#14858) Signed-off-by: jthomson04 <jwillthomson19@gmail.com> * fix(xpu): stabilize XPU test workers (ai-dynamo#14539) Signed-off-by: Wenxin Zhang <wenxin.zhang@intel.com> Signed-off-by: VincyZhang <wenxin.zhang@intel.com> * feat(mm-routing): add Nemotron 3 Nano Omni video routing (ai-dynamo#14653) Signed-off-by: krishung5 <krish@nvidia.com> * fix(sglang): validate diffusion input_reference and bound media fetches (ai-dynamo#14435) The sglang image-diffusion and video-generation handlers passed the client-supplied input_reference through to the generator's image_path after only a non-empty check. Validate it first, and for remote references materialize it locally before the generator sees it, so the generator is always handed a trusted local path. This brings the sglang diffusion path in line with the vLLM/omni and trtllm backends, which already validate the same field. Behavior change: local I2I/I2V references now require DYN_MM_LOCAL_PATH to be set to the allowed directory; previously any path was accepted. common/http: - validate_media_reference() returns a plain filesystem path for local references; local_media_reference() is an async context manager that fetches a remote one through fetch_bytes(policy=...), which revalidates every redirect hop, into a temp file removed on exit. data: is rejected -- a URI is not a path. - fetch_bytes() gained max_bytes, streaming through collect_capped at an explicit read granularity so the cap is an allocation bound and not only a rejection: a 128 MiB-decoded gzip body against the 64 MiB cap peaks at 68,032,217 bytes rather than the whole decompressed body. Content-Length is caller-controlled and absent when chunked, and aiohttp's read(n) returns at most n bytes, so neither a header check nor a single capped read suffices. Defaults to None, leaving existing callers unchanged. - DYN_MM_MAX_FILE_SIZE_MB makes that cap operator-tunable, in megabytes, as the SGLang arg it replaces was. Read per call; empty, unparseable or non-positive falls back to 64 with a warning, so a malformed value neither takes the worker down nor reads as unlimited. - Messages built from caller input are bounded via describe_media_source, moved from multimodal/media_source.py (it pulls in torch) into url_validator.py and re-exported from its old home; a no-op below 120 characters. - HttpStatusError bounds its .message attribute, not only the rendered string: errors.rs::extract_http_like_error reads .status and .message off this class by name and forwards .message on a 4xx without calling str(). Backend exception text is bounded head-and-tail, since aiohttp renders the host before the errno. - validate_local_path uses exc.strerror rather than the raw OSError, whose text repeats the filename, and now catches the ValueError that Path.resolve() raises on an embedded NUL so callers keep their 4xx-vs-5xx decision. Rebased onto ai-dynamo#14563 (single aiohttp backend); the httpx-side half of the max_bytes plumbing went with that backend. Signed-off-by: nnshah1 <neelays@nvidia.com> Signed-off-by: Dmitry Tokarev <dtokarev@nvidia.com> Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore(deps): upgrade fastokens to 0.3.2 (ai-dynamo#14798) Signed-off-by: jthomson04 <jwillthomson19@gmail.com> * fix(vllm): ship codec-free OpenCV for image inputs (ai-dynamo#14361) Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com> Signed-off-by: GLAMR <svc-glamr@nvidia.com> Co-authored-by: Anant Sharma <anants@nvidia.com> Co-authored-by: yunzhoul-nv <232973175+yunzhoul-nv@users.noreply.github.com> * docs: refresh community events Automated refresh from the public Dynamo Google Calendar. Generated by .github/workflows/community-events-refresh.yml. Signed-off-by: dynamo-ops <170655669+dynamo-ops@users.noreply.github.com> * ci: refresh the compliance baseline in auto-upgrade pipeline (ai-dynamo#14206) Signed-off-by: Anant Sharma <anants@nvidia.com> * feat(triton): honor KServe classification on tensor outputs (ai-dynamo#14783) Signed-off-by: Yingge He <yinggeh@nvidia.com> * docs(rl): stop the verl guide sending readers to a vLLM version it cannot run on (ai-dynamo#14571) Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com> * feat(mocker): publish native KV events from the vLLM gRPC server (ai-dynamo#14737) Signed-off-by: jthomson04 <jwillthomson19@gmail.com> * fix(kv-router): release unowned radix branches after eviction (ai-dynamo#14878) Signed-off-by: jthomson04 <jwillthomson19@gmail.com> * fix: show correct backend versions in the install selectors (ai-dynamo#13599) Signed-off-by: Anant Sharma <anants@nvidia.com> * build(vllm): prepare v0.29.0 bump (ai-dynamo#14543) Signed-off-by: Julien Darve <jdarve@NVIDIA.com> * ci(xpu): validation PR for the re-applied XPU workflows and Dockerfile Throwaway PR to prove the CI merged in #22 actually runs end to end on XPU hardware. Adds only a comment to container/templates/vllm_runtime.Dockerfile, which matches the `vllm` path filter (container/templates/vllm_*) and so makes changed-files set vllm=true, which is what gates build-xpu and the heterog-test-px-dn / heterog-test-pn-dx jobs. What this exercises: - .github/workflows/pr-xpu.yaml (push to pull-request/[0-9]+, needs the xpu label) - .github/workflows/pr-xpu-heterogeneous.yaml (push; its guard deliberately skips the label gate) - .github/workflows/epd-test-template.yml (workflow_call, from the heterog jobs) - .github/scripts/test-filters.js (the brace fix from #22) - container/templates/vllm_runtime.Dockerfile rendered and built for device=xpu Not exercised: .github/workflows/xpu-heterogeneous-dispatch.yaml is workflow_dispatch only and has to be run by hand from the Actions tab. The marker comment must be removed before this branch is ever merged. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(triton): Update Triton Base Image to 26.08 (ai-dynamo#14854) Signed-off-by: J Wyman <jwyman@nvidia.com> Co-authored-by: Rini Gupta <rinig@nvidia.com> * fix(operator): normalize equivalent worker hash inputs (ai-dynamo#14721) Signed-off-by: bzsuni <bingzhe.sun@daocloud.io> * test(sglang): exercise NIXL in embedding cache E/PD test (ai-dynamo#14795) Signed-off-by: Sai Kiran Polisetty <spolisetty@nvidia.com> * fix(sglang): stop the elastic-EP scale-up worker crash-looping at startup (ai-dynamo#14568) Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com> Co-authored-by: yunzhoul-nv <232973175+yunzhoul-nv@users.noreply.github.com> * fix(responses): honor tool_choice when parsing tool calls from text (ai-dynamo#14843) Signed-off-by: xianlubird <xianlubird@gmail.com> * ci: accept trusted full-CI request comments (ai-dynamo#14868) Signed-off-by: Matej Kosec <mkosec@nvidia.com> * docs: clarify EPP mode boundary and single-replica Dynamo mode fixes [DYN-4310] (ai-dynamo#14756) Signed-off-by: Anna Tchernych <atchernych@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * ci(docs): move the generated-tables determinism gate out of link checking (ai-dynamo#14135) Signed-off-by: Dan Gil <dagil@nvidia.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> * ci(docs): generate the Kubernetes API reference at publish time (ai-dynamo#14122) Signed-off-by: Dan Gil <dagil@nvidia.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> * fix(operator): discover pull secrets for init containers (ai-dynamo#14922) Signed-off-by: bojiang-li <327132355+bojiang-li@users.noreply.github.com> * fix(sglang): stop an unusable mooncake backend crashing workers after model load (ai-dynamo#14461) Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com> Signed-off-by: glamr-agent <glamr-agent@users.noreply.github.com> Signed-off-by: GLAMR <svc-glamr@nvidia.com> * fix(sglang): emit prefill handoff before completion in sidecar (ai-dynamo#14260) Signed-off-by: jain-ria <riajain@NVIDIA.com> Co-authored-by: jain-ria <riajain@NVIDIA.com> Co-authored-by: Connor Carpenter <connorc@nvidia.com> Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com> * test(trtllm): enable fault tolerance coverage (ai-dynamo#14609) Signed-off-by: tanmayv25 <tanmay2592@gmail.com> * fix(frontend): evict async tokenizer executors when the tokenizer is retired (ai-dynamo#13368) Signed-off-by: Peter Pan <Peter.Pan@daocloud.io> * fix(llm): report KServe datatypes by their wire names, not protobuf variants (ai-dynamo#14957) `ModelMetadata` reported each Triton-registered tensor's `datatype` using `inference::DataType::as_str_name()`, which returns the `model_config.proto` variant name (`TYPE_FP32`, `TYPE_STRING`, ...) instead of the KServe v2 wire names (`FP32`, `BYTES`, ...). Every datatype was wrong, so spec-conforming clients cannot parse any tensor the RPC describes. Adds `oip_name()` next to `tensor::DataType::to_kserve` covering all fifteen proto variants (incl. FP16 and BF16) and mapping `TYPE_STRING → BYTES`. Original PR by @ayaangazali: ai-dynamo#14770. Reissued under a signed commit to unblock the copy-pr-bot signature gate; diff is byte-identical. Closes ai-dynamo#14520. Signed-off-by: ayaangazali <ayaangazali@users.noreply.github.com> Signed-off-by: ayaangazali <ayaangazali.work@gmail.com> Signed-off-by: Vinya Kestur <vinyak@nvidia.com> Co-authored-by: ayaangazali <ayaangazali.work@gmail.com> * docs(mm-routing): document video KV routing (ai-dynamo#14958) Signed-off-by: krishung5 <krish@nvidia.com> * fix(sidecar): honor worker namespace suffix (ai-dynamo#14955) Signed-off-by: Biswa Panda <biswa.panda@gmail.com> * fix(bindings): drain bridge tasks before interpreter finalization (ai-dynamo#14813) Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com> Signed-off-by: GLAMR <svc-glamr@nvidia.com> Co-authored-by: Tushar Sharma <tusharma@nvidia.com> * fix(discovery): stop a Qwen3-VL worker from serving video with another worker's contract (ai-dynamo#14624) Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com> Signed-off-by: GLAMR <svc-glamr@nvidia.com> * fix(gms): honor configured timeout during initial weights admission (ai-dynamo#14877) Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com> Signed-off-by: Schwinn Saereesitthipitak <schwinns@nvidia.com> Signed-off-by: GLAMR <svc-glamr@nvidia.com> Co-authored-by: Schwinn Saereesitthipitak <schwinns@nvidia.com> * feat(kv-router): add construction-time indexer delegates (ai-dynamo#14945) * fix(sglang): support min_tokens on tokenizer-free decode workers (ai-dynamo#14276) Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com> Signed-off-by: GLAMR <svc-glamr@nvidia.com> Signed-off-by: jain-ria <riajain@NVIDIA.com> Co-authored-by: jain-ria <riajain@NVIDIA.com> Co-authored-by: MatejKosec <mkosec@nvidia.com> * feat(router): add SessionPrefixIndexer for session-block lineage (ai-dynamo#13807) Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com> Signed-off-by: Karen Chung <karenc@nvidia.com> Signed-off-by: Matej Kosec <mkosec@nvidia.com> Co-authored-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com> Co-authored-by: Matej Kosec <mkosec@nvidia.com> * fix(vllm): settle kvwarm stages through a per-step round on every attention-DP rank (ai-dynamo#14728) Signed-off-by: Yiming Liu <yimingl@nvidia.com> * feat(vllm): benchmark hybrid caches with random KDA state (ai-dynamo#14900) Signed-off-by: hongkuanz <hongkuanz@nvidia.com> * fix(runtime): fix QUIC reassembly and reduce response stalls (ai-dynamo#14876) Signed-off-by: jthomson04 <jwillthomson19@gmail.com> * feat(router): unify frontend and standalone selection core (ai-dynamo#14570) Signed-off-by: Ishan Dhanani <ishandhanani@gmail.com> Signed-off-by: Thomas Montfort <tjmontfort12@gmail.com> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: Thomas Montfort <tjmontfort12@gmail.com> * fix(planner): keep control APIs responsive during Prometheus collection (ai-dynamo#14377) Signed-off-by: xianlubird <xianlubird@gmail.com> Co-authored-by: Hongkuan Zhou <tedzhouhk@gmail.com> * fix(router): record SGLang prefill completion after stream ends (ai-dynamo#14968) Signed-off-by: jain-ria <riajain@NVIDIA.com> * fix(frontend): send inline media once on the TCP request plane (ai-dynamo#14801) Signed-off-by: Sumit Mishra <sah299610@gmail.com> Co-authored-by: Indrajit Bhosale <iamindrajitb@gmail.com> * docs: refresh community events Automated refresh from the public Dynamo Google Calendar. Generated by .github/workflows/community-events-refresh.yml. Signed-off-by: dynamo-ops <170655669+dynamo-ops@users.noreply.github.com> * fix(vllm): initialize synchronizer in KV warmup capacity test (ai-dynamo#14984) Signed-off-by: Alec Flowers <aflowers@nvidia.com> * fix(recipes): make the Solar Open2 250B benchmark and docs link usable (ai-dynamo#14956) Signed-off-by: Sandhya Rani Narravula <snarravula@nvidia.com> * feat(recipes): add K-EXAONE 2.0 750B-A37B NVFP4 vLLM recipes for B200 (ai-dynamo#14822) Signed-off-by: Cheng Wang <chengwa@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat: KVCR Resiliency Deployment Example (ai-dynamo#14695) Add two-node DynamoGraphDeployment examples for process-local KVCR and the KVCR memory service. Run one vLLM worker per GPU node, use stable Grove ordinals for cache-owner slots, and request GPU-local RDMA resources for engines and Guard services. Provide a deployment helper for rendering and selecting either variant. Run the KV state agent alongside vLLM for process-local host memory. In memory-service mode, keep KVCR and the state agent in a separate container so its Guard and shared-memory pool survive engine restarts. Document that restarting the services sidecar invalidates the MVP recovery contract and requires deployment-level replacement. Add manifest coverage and an opt-in two-host lifecycle test. Kill the source EngineCore, hold it offline, and verify that the promoted Guard serves its preserved cache to the surviving target. Correlate response equality and KVCR transfer metrics with transmit and receive counters from the selected active HCA to prove RDMA transport. Pin compatible KVCR and vLLM revisions and document the runtime, discovery, compatibility-digest, and recovery prerequisites. Signed-off-by: Adit Ranadive <aranadive@nvidia.com> * feat(omni): add Nemotron Audex speech synthesis to /v1/audio/speech (ai-dynamo#12788) Signed-off-by: Thanaji Rao Thakkalapelli <thanaji.rao.thakkalapelli@intel.com> * ci: allow glamr-agent to request CI on its own unsigned PRs (ai-dynamo#14964) Signed-off-by: Matej Kosec <mkosec@nvidia.com> * fix(vllm): isolate multimodal worker ports (ai-dynamo#14751) Signed-off-by: Keiven Chang <keivenchang@users.noreply.github.com> Co-authored-by: Keiven Chang <keivenchang@users.noreply.github.com> * fix(runtime): reject invalid DYN_REQUEST_PLANE values (ai-dynamo#12612) Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com> Signed-off-by: Matej Kosec <mkosec@nvidia.com> Signed-off-by: Coding Agent <svc-glamr@nvidia.com> Signed-off-by: GLAMR <svc-glamr@nvidia.com> Co-authored-by: MatejKosec <mkosec@nvidia.com> * fix(responses): preserve text instead of inferring tool calls (ai-dynamo#14846) Signed-off-by: xianlubird <xianlubird@gmail.com> Co-authored-by: Ryan McCormick <rmccormick@nvidia.com> * chore: temporarily increase frontend build time limit 45 --> 90 min (ai-dynamo#15019) Signed-off-by: Dmitry Tokarev <dtokarev@nvidia.com> * test(operator): cover scoped CA injection ownership (ai-dynamo#14961) Signed-off-by: Julien Mancuso <jmancuso@nvidia.com> * feat(frontend): map semantic errors to HTTP responses (ai-dynamo#14396) Signed-off-by: Biswa Panda <biswa.panda@gmail.com> * docs: correct fault-tolerance architecture details (ai-dynamo#14880) Signed-off-by: Elizabeth Thomas <email2eliza@gmail.com> * build(deps): bump nats-server to v2.14.7 (ai-dynamo#14919) Signed-off-by: Dan Gil <dagil@nvidia.com> * build(deps): bump AISimulate to 0.12.0 (ai-dynamo#15012) Signed-off-by: Harrison King Saturley-Hall <hsaturleyhal@nvidia.com> * remove oneAPI env for XPU detection * feat(backends): expose native LoRA capacity in model registration (ai-dynamo#14754) Signed-off-by: Julien Darve <jdarve@NVIDIA.com> Signed-off-by: bzsuni <bingzhe.sun@daocloud.io> Co-authored-by: bzsuni <86399306+bzsuni@users.noreply.github.com> * fix(planner): handle pending decisions in virtual connector wait (ai-dynamo#14841) Signed-off-by: bzsuni <bingzhe.sun@daocloud.io> Co-authored-by: Hongkuan Zhou <tedzhouhk@gmail.com> * feat(vllm): add sidecar LoRA lifecycle (ai-dynamo#13068) Signed-off-by: Julien Darve <jdarve@NVIDIA.com> Signed-off-by: bzsuni <bingzhe.sun@daocloud.io> Co-authored-by: Julien Darve <jdarve@NVIDIA.com> Co-authored-by: bzsuni <86399306+bzsuni@users.noreply.github.com> * fix(vllm/omni): pass response_format into video EngineInputs (ai-dynamo#14667) (ai-dynamo#14844) * chore: bump version to 1.6.0 post 1.5.0 branch cut (ai-dynamo#15009) Signed-off-by: pvijayakrish <pvijayakrish@nvidia.com> Signed-off-by: Pavithra Vijayakrishnan <160681768+pvijayakrish@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(ci): Use `pytest --ignore` to Skip Tests Based on Framework (ai-dynamo#14815) Signed-off-by: J Wyman <jwyman@nvidia.com> * feat(sidecar): add e2e CI testing for sidecar launch scripts (ai-dynamo#14508) Signed-off-by: tanmayv25 <tanmay2592@gmail.com> Signed-off-by: Julien Darve <jdarve@NVIDIA.com> Co-authored-by: Julien Darve <jdarve@NVIDIA.com> * chore(xpu): upgrade vllm and omni to 0.29.0 Signed-off-by: wenxin.zhang <wenxin.zhang@intel.com> * docs(operator): document the DGDR workload-creation trust boundary (ai-dynamo#14429) Signed-off-by: nnshah1 <neelays@nvidia.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> * fix(xpu): use released vllm-omni prerelease Signed-off-by: Wenxin Zhang <wenxin.zhang@intel.com> * test(efa): add the EFA disaggregated deploy test for sglang (ai-dynamo#13893) Signed-off-by: Jie Hao <jihao@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(runtime): support IPv6-only IP resolution (ai-dynamo#13126) Signed-off-by: jthomson04 <jwillthomson19@gmail.com> * docs(fault-tolerance): clarify migration after shutdown grace expires (ai-dynamo#14872) Signed-off-by: Jacky <18255193+kthui@users.noreply.github.com> * feat(vllm-omni): preserve generated video audio (ai-dynamo#13707) Signed-off-by: Guan Luo <gluo@nvidia.com> Co-authored-by: Guan Luo <gluo@nvidia.com> * feat(vllm-omni): pass model-specific video parameters (ai-dynamo#13708) Signed-off-by: Guan Luo <gluo@nvidia.com> Co-authored-by: Guan Luo <gluo@nvidia.com> * feat(vllm-omni): qualify MiniMax-H3 T2VA on B200 (ai-dynamo#13589) Signed-off-by: Guan Luo <gluo@nvidia.com> Signed-off-by: GuanLuo <41310872+GuanLuo@users.noreply.github.com> Co-authored-by: Guan Luo <gluo@nvidia.com> Co-authored-by: GuanLuo <41310872+GuanLuo@users.noreply.github.com> Co-authored-by: Ryan McCormick <rmccormick@nvidia.com> * fix(vllm): remove obsolete Omni compatibility guard Signed-off-by: Wenxin Zhang <wenxin.zhang@intel.com> * fix(vllm): retain Omni compatibility guard Signed-off-by: Wenxin Zhang <wenxin.zhang@intel.com> * .github/workflows/pr-xpu-heterogeneous.yaml; pin GPU_TAG to latest * .github/workflows/; add post-merge and nightly XPU heterogeneous CI Extract the XPU heterogeneous P/D pipeline out of pr-xpu-heterogeneous.yaml into xpu-heterogeneous-run.yml, a workflow_call reusable workflow, and call it from three thin trigger workflows so all three merge phases run the identical pipeline instead of drifting copies. xpu-heterogeneous-run.yml new, reusable. guard, changed-files, build-xpu, build-nvidia, resolve-images and both heterog tests, unchanged, plus 7 inputs. pr-xpu-heterogeneous.yaml reduced to the pre-merge trigger, the slash-command gate and the reaction. post-merge-xpu-heterogeneous.yaml new. push to main. nightly-xpu-heterogeneous.yaml new file, but the cron is MOVED, not added: it is the 0 23 * * * schedule that was already in pr-xpu-heterogeneous.yaml. No behaviour change per phase. force_all_tests replaces the old github.event_name == 'schedule' || github.event_name == 'issue_comment' expression with the same truth table: pre-merge passes github.event_name == 'issue_comment', nightly passes true. Post-merge also passes true, because a push to main has no PR base for .github/actions/changed-files to diff against, and post-merge exists to catch what per-PR gating missed. xpu-status-check stays a TOP-LEVEL job in each caller rather than moving into the reusable workflow. A job contributed by a reusable workflow reports to the Checks API as "run / xpu-status-check", so hosting it there would rename the context and leave any branch protection rule requiring xpu-status-check waiting forever on a check that no longer reports. The concurrency mapping stays byte-identical across all four workflows that touch this hardware, now including xpu-heterogeneous-dispatch.yaml. Three files do NOT get three slots: the cluster, the dynamo-system namespace and the onexpu-/onenvidia-rdma-kueue ResourceClaimTemplates are one global resource. The reusable workflow deliberately carries no concurrency block of its own, which would deadlock against the slot the caller's run already holds. Parameterised gpu_tag, model, tensor_parallel and runner as inputs so the callers can diverge; all default to the previously hardcoded values. Added workflow_dispatch to the nightly, without which a schedule-only workflow cannot be exercised before it reaches the default branch. Verified: all files parse; the four concurrency mappings are byte-identical; the reusable workflow declares no concurrency; every input each caller passes exists and every required input is supplied; nesting is depth 3 of the 4 GitHub allows. actionlint was not available to run, and will report queue:max as an unknown key in all four files, a known false positive. --------- Signed-off-by: Nikita Sukharev <kaonael@gmail.com> Signed-off-by: xianlubird <xianlubird@gmail.com> Signed-off-by: hongkuanz <hongkuanz@nvidia.com> Signed-off-by: William Arnold <7565007+Aphoh@users.noreply.github.com> Signed-off-by: Sandhya Rani Narravula <snarravula@nvidia.com> Signed-off-by: Karen Chung <karenc@nvidia.com> Signed-off-by: jthomson04 <jwillthomson19@gmail.com> Signed-off-by: Wenxin Zhang <wenxin.zhang@intel.com> Signed-off-by: VincyZhang <wenxin.zhang@intel.com> Signed-off-by: krishung5 <krish@nvidia.com> Signed-off-by: nnshah1 <neelays@nvidia.com> Signed-off-by: Dmitry Tokarev <dtokarev@nvidia.com> Signed-off-by: svc-glamr@nvidia.com <svc-glamr@nvidia.com> Signed-off-by: GLAMR <svc-glamr@nvidia.com> Signed-off-by: dynamo-ops <170655669+dynamo-ops@users.noreply.github.com> Signed-off-by: Anant Sharma <anants@nvidia.com> Signed-off-by: Yingge He <yinggeh@nvidia.com> Signed-off-by: Julien Darve <jdarve@NVIDIA.com> Signed-off-by: J Wyman <jwyman@nvidia.com> Signed-off-by: bzsuni <bingzhe.sun@daocloud.io> Signed-off-by: Sai Kiran Polisetty <spolisetty@nvidia.com> Signed-off-by: Matej Kosec <mkosec@nvidia.com> Signed-off-by: Anna Tchernych <atchernych@nvidia.com> Signed-off-by: Dan Gil <dagil@nvidia.com> Signed-off-by: bojiang-li <327132355+bojiang-li@users.noreply.github.com> Signed-off-by: glamr-agent <glamr-agent@users.noreply.github.com> Signed-off-by: jain-ria <riajain@NVIDIA.com> Signed-off-by: tanmayv25 <tanmay2592@gmail.com> Signed-off-by: Peter Pan <Peter.Pan@daocloud.io> Signed-off-by: ayaangazali <ayaangazali@users.noreply.github.com> Signed-off-by: ayaangazali <ayaangazali.work@gmail.com> Signed-off-by: Vinya Kestur <vinyak@nvidia.com> Signed-off-by: Biswa Panda <biswa.panda@gmail.com> Signed-off-by: Schwinn Saereesitthipitak <schwinns@nvidia.com> Signed-off-by: Yiming Liu <yimingl@nvidia.com> Signed-off-by: Ishan Dhanani <ishandhanani@gmail.com> Signed-off-by: Thomas Montfort <tjmontfort12@gmail.com> Signed-off-by: Sumit Mishra <sah299610@gmail.com> Signed-off-by: Alec Flowers <aflowers@nvidia.com> Signed-off-by: Cheng Wang <chengwa@nvidia.com> Signed-off-by: Adit Ranadive <aranadive@nvidia.com> Signed-off-by: Thanaji Rao Thakkalapelli <thanaji.rao.thakkalapelli@intel.com> Signed-off-by: Keiven Chang <keivenchang@users.noreply.github.com> Signed-off-by: Coding Agent <svc-glamr@nvidia.com> Signed-off-by: Julien Mancuso <jmancuso@nvidia.com> Signed-off-by: Elizabeth Thomas <email2eliza@gmail.com> Signed-off-by: Harrison King Saturley-Hall <hsaturleyhal@nvidia.com> Signed-off-by: pvijayakrish <pvijayakrish@nvidia.com> Signed-off-by: Pavithra Vijayakrishnan <160681768+pvijayakrish@users.noreply.github.com> Signed-off-by: wenxin.zhang <wenxin.zhang@intel.com> Signed-off-by: Jie Hao <jihao@nvidia.com> Signed-off-by: Jacky <18255193+kthui@users.noreply.github.com> Signed-off-by: Guan Luo <gluo@nvidia.com> Signed-off-by: GuanLuo <41310872+GuanLuo@users.noreply.github.com> Co-authored-by: Nikita Sukharev <kaonael@gmail.com> Co-authored-by: Xianlu Bird <xianlubird@gmail.com> Co-authored-by: Hongkuan Zhou <tedzhouhk@gmail.com> Co-authored-by: William Arnold <7565007+Aphoh@users.noreply.github.com> Co-authored-by: Zero Rains <57100978+zeroRains@users.noreply.github.com> Co-authored-by: snarravula-dl <snarravula@nvidia.com> Co-authored-by: Karen Chung <karenc@nvidia.com> Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com> Co-authored-by: jthomson04 <jwillthomson19@gmail.com> Co-authored-by: VincyZhang <wenxin.zhang@intel.com> Co-authored-by: Kris Hung <krish@nvidia.com> Co-authored-by: Neelay Shah <neelays@nvidia.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: GLAMR <svc-glamr@nvidia.com> Co-authored-by: Anant Sharma <anants@nvidia.com> Co-authored-by: yunzhoul-nv <232973175+yunzhoul-nv@users.noreply.github.com> Co-authored-by: dynamo-ops <170655669+dynamo-ops@users.noreply.github.com> Co-authored-by: Yingge He <157551214+yinggeh@users.noreply.github.com> Co-authored-by: JulienDarve <86800349+JulienDarve@users.noreply.github.com> Co-authored-by: J Wyman <jwyman@nvidia.com> Co-authored-by: Rini Gupta <rinig@nvidia.com> Co-authored-by: bzsuni <86399306+bzsuni@users.noreply.github.com> Co-authored-by: Sai Kiran Polisetty <spolisetty@nvidia.com> Co-authored-by: MatejKosec <mkosec@nvidia.com> Co-authored-by: atchernych <atchernych@nvidia.com> Co-authored-by: Dan Gil <dagil@nvidia.com> Co-authored-by: Bojiang Li <327132355+bojiang-li@users.noreply.github.com> Co-authored-by: Connor Carpenter <connorcarpenter15@gmail.com> Co-authored-by: jain-ria <riajain@NVIDIA.com> Co-authored-by: Connor Carpenter <connorc@nvidia.com> Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com> Co-authored-by: Tanmay Verma <tanmayv@nvidia.com> Co-authored-by: Peter Pan <peter.pan@daocloud.io> Co-authored-by: Vinya Kestur Tumakuru Arun Kumar <vinyak@nvidia.com> Co-authored-by: ayaangazali <ayaangazali.work@gmail.com> Co-authored-by: Biswa Panda <biswa.panda@gmail.com> Co-authored-by: Tushar Sharma <tusharma@nvidia.com> Co-authored-by: Schwinn Saereesitthipitak <schwinns@nvidia.com> Co-authored-by: Ryan Olson <ryanolson@users.noreply.github.com> Co-authored-by: Yimingl_Nvidia <yimingl@nvidia.com> Co-authored-by: Thomas Montfort <tjmontfort12@gmail.com> Co-authored-by: Sumit884-byte <sah299610@gmail.com> Co-authored-by: Indrajit Bhosale <iamindrajitb@gmail.com> Co-authored-by: Alec <35311602+alec-flowers@users.noreply.github.com> Co-authored-by: chw001 <chengwa@nvidia.com> Co-authored-by: Adit Ranadive <aranadive@nvidia.com> Co-authored-by: Thanaji Rao Thakkalapelli <thanaji.rao.thakkalapelli@intel.com> Co-authored-by: Keiven C <213854356+keivenchang@users.noreply.github.com> Co-authored-by: Keiven Chang <keivenchang@users.noreply.github.com> Co-authored-by: Ryan McCormick <rmccormick@nvidia.com> Co-authored-by: Dmitry Tokarev <dtokarev@nvidia.com> Co-authored-by: Julien Mancuso <161955438+julienmancuso@users.noreply.github.com> Co-authored-by: Elizabeth Thomas <email2eliza@gmail.com> Co-authored-by: Harrison Saturley-Hall <hsaturleyhal@nvidia.com> Co-authored-by: Julien Darve <jdarve@NVIDIA.com> Co-authored-by: Jasim Kareem <mj9034812@gmail.com> Co-authored-by: Pavithra Vijayakrishnan <160681768+pvijayakrish@users.noreply.github.com> Co-authored-by: Jie Hao <jihao@nvidia.com> Co-authored-by: Jacky <18255193+kthui@users.noreply.github.com> Co-authored-by: Qi Wang <qiwa@nvidia.com> Co-authored-by: Guan Luo <gluo@nvidia.com> Co-authored-by: GuanLuo <41310872+GuanLuo@users.noreply.github.com>
Summary
QUIC response connections can fail with
too many gaps in stream bufferwhen quinn-proto 0.11.17 retains many small, contiguous chunks. Its stream assembler can reach the chunk limit even with lossless, ordered data. Update Quinn from 0.11.9 to 0.11.12 and quinn-proto from 0.11.17 to 0.11.18 in all four lockfiles to include the upstream fix that coalesces contiguous chunks during defragmentation (quinn-rs/quinn#2826).Include the full error cause chain when a QUIC response lane or worker connection bundle fails. Frontend lane warnings also include Quinn's connection-close reason as a best-effort observation; another task can initiate teardown before that reason is read. The once-per-bundle worker warning includes the bundle ID, failure path (
writerorcontrol_reader), and remote frontend address. Worker registration failures receive the richer cause text through the existing failure handler. The worker warning records the failure reason as a string field so console logs escape embedded newlines and terminal control characters. Added formatting and address lookup run on error paths.Reduce receive-side scheduling stalls by driving each lane's reader and reverse-control writer as pinned futures in one task. Charge Tokio's cooperative budget at 16-frame or 4 KiB thresholds so buffered token frames can advance without charging each frame separately, while still yielding for control traffic. Both futures stay pinned across polls so partial reads are preserved.
Increase Linux frontend receive endpoints from 8 to 32, using the existing reuse-port group and advertised address. This adds 24 sockets per response server. The endpoint count remains one on other platforms. The wire format, worker connection/lane counts, failure counters, and public configuration are unchanged. TCP remains the default; the shipped QUIC batch interval remains 5 ms.
This is a dependency release update. In addition to reassembly, quinn-proto 0.11.18 changes send-window accounting, ACK bundling, blocked-frame emission, datagram/path handling, and transport-parameter validation. The evidence below does not isolate the performance effects of those changes.
The Quinn 0.11.9 → 0.11.12 update also changes connection-handle reference counting to use atomics, cleans up receive-stream stop/drop state, and uses the clock to detect timer expiry when Tokio cooperative-budget limits can delay
Sleep::poll. Quinn's minimum supported Rust version increases from 1.74.1 to 1.85; the repository's pinned Rust 1.96.1 toolchain meets that requirement. The measured performance differences cannot be assigned to an individual Quinn or quinn-proto change.Validation
Checks on the PR branch
1ec03db608ae462282cdb8747f24cfa210ea54ab,cargo test --locked -p dynamo-runtime --lib pipeline::network::quic_response::tests -- --test-threads=8— 22 passed, including 1,000 ordered responses, cancellation, connection failure/reconnection, failure before prologue, and endpoint shutdown.cargo clippy --locked -p dynamo-runtime --lib— passed with no warnings.cargo fmt --all -- --checkandgit diff --check— passed.%reason, but one physical line with escaped payload characters usingreason. The check and output remain local artifacts, outside the commit.cargo tree --locked --manifest-path <manifest> -i quinn-proto --prefix none --depth 1— all four workspaces resolve Quinn 0.11.12 and quinn-proto 0.11.18 without lockfile changes.ordered_lossless_stream_is_not_rejectedtest fails withTooManyChunkson 0.11.17; all 26 assembler tests pass on 0.11.18, including the excessive-gap limit tests. The new crate checksum matches the PR lockfiles. Application-frame count alone does not guarantee the assembler retention needed to reproduce this defect, so the deterministic reproduction remains in the upstream assembler tests.Deferred validation
The requested baseline-versus-head benchmark with the shipped 5 ms batch interval is deferred for this PR. No new performance benchmark was run for this follow-up. The existing Tyche results use a zero batch interval and common campaign changes; they do not validate this exact PR under the shipped configuration. That performance-validation gap remains open.
Existing Tyche comparison: Quinn update
The two measured QUIC arms differ only in the Quinn manifest/lockfile update: old
829f64238cfe7cf06c52d4d52087a60ea5d13e62, updatedfee3f7408a9a754245a18ed611020f8cb44d7997. Both include other campaign changes and use a zero batch interval, rather than the exact PR code with the shipped 5 ms interval.Cross-node mock-worker runs used eight worker processes, a separate load-generator node, fresh services, consistent warmup, identical request fixtures, and two 90-second repetitions per case with reversed arm order. Coverage was 128-input/32-output requests at concurrency 512, 1,024, and 4,096, plus mixed requests at concurrency 1,024. These are existing campaign results; they precede the reader-budget and 32-endpoint changes below.
The updated arm's exported request lengths also matched, with no bundle-failure counter increments in the measured resource windows. The previous arm included one response with 24 output tokens instead of the requested 32.
For the short workload at concurrency 512:
Latency percentiles exclude failed requests, so these are not equivalent successful cohorts. The connection-failure fix is supported in this tested scope, but p99 increased and frontend UDP receive drops remain unresolved. This evidence does not establish latency parity, isolate ACK or send-window effects, or qualify QUIC as the default. TCP remains the default.
Tyche receive-path evidence
The receive-task change and the 16-frame / 4 KiB budget thresholds match the measured campaign implementation; the Linux endpoint count is 32. These are focused ports of campaign commits
2c41ea8476,901cb517bc, andf969ff855f.In the earlier single-run short-response screens (128 input / 32 output tokens, concurrency 512), previous QUIC had 192.571 ms total p99 at 38,765.47 requests/s. Reader batching plus 32 endpoints had 24.398 ms p99 at 40,637.41 requests/s. Its paired TCP run had 10.635 ms p99 at 36,176.16 requests/s. The approximately 87% QUIC p99 reduction is a single-run observation; it does not establish repeatability or TCP latency parity.
The full 8-versus-32 endpoint comparison has 36 runs, three alternating pairs per case, concurrency 1/16/64, one/eight mock workers, and 6,651,658 successful requests with no client errors. Worst paired regressions were 3.10% throughput, 1.60% total p99, and 0.17% serving CPU per completed request.
There is a memory cost: with eight workers at concurrency 64, median frontend peak RSS increased from 169.4 to 219.8 MiB. The extra endpoints add 69 MiB of Quinn receive-buffer allocation capacity under the checked Tyche Linux GRO settings. UDP-drop variation and isolated TTFT/token-time tail regressions are retained in the linked results.
Both campaigns used other common frontend/QUIC changes and explicitly set the batch interval to zero. The campaign also used a 32-frame sender batch cap and reorder-limit backpressure; this PR retains the existing 512-frame cap and reorder behavior. These measurements are not a benchmark of this exact PR commit or its shipped 5 ms interval. The tests above validate the port. The remaining high-concurrency and workload coverage gaps still block a default transport change.
Related Issues
Summary by CodeRabbit
Bug Fixes
Chores