Skip to content

Preserve and validate score-centering candidate distributions - #3622

Merged
guapisolo merged 16 commits into
shi/score-centering-mathfrom
shi/score-centering-transport
Oct 3, 2026
Merged

guapisolo merged 16 commits into
shi/score-centering-mathfrom
shi/score-centering-transport

Conversation

@Shi-Dong

@Shi-Dong Shi-Dong commented Sep 23, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Carry rollout candidate IDs and probabilities through samples, session serialization, trimming, reset, and training-batch conversion. Validate shapes, padding, uniqueness, and sampled-token agreement.

Accept unfiltered sampling and bounded top-p/top-k sampling in the configuration contract. For filtered requests, ask SGLang for post-filter probabilities over the complete realized support. Reject support larger than the saved candidate count or missing probability metadata. This uses the API from SGLang #40932, incorporated into sglang-miles by #41047. The minimum source revision on sglang-miles is merge commit ae04cb14046896b6d453758c5769d639deedd353, or a descendant containing it. Session/OpenAI transport also needs a compatible router build forwarding sampling_logprobs_mode; the corresponding router change #41373 is still open, so this engine revision is not a released-router guarantee.

Stack part 2/6. Depends on #3621. Merge in dependency order.

Validation: 80 focused tests passed and 6 GPU-dependent cases skipped on this intermediate branch.

@Shi-Dong

Copy link
Copy Markdown
Collaborator Author

@claude review always

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This repository is configured for manual code reviews. Comment @claude review for a one-time review, or @claude review always to subscribe this PR to a review on every future push.

Tip: disable this comment in your organization's Code Review settings.

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code review found no issues

No high-confidence issues detected in this change.

@Shi-Dong
Shi-Dong force-pushed the shi/score-centering-transport branch from 24fb3aa to ba1b261 Compare September 23, 2026 19:47

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code review found no issues

No high-confidence issues detected in this change.

@Shi-Dong
Shi-Dong force-pushed the shi/score-centering-transport branch from ba1b261 to e70c84f Compare September 24, 2026 08:07

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Beyond the reported findings, I also checked whether validate_score_centering_args (the new config-contract validator with all its MIS/TIS/advantage-estimator checks) is invoked anywhere outside its own test file — it isn't; nothing in miles/utils/arguments.py or elsewhere calls it, so it's currently dead code pending a later stack PR.

Extended reasoning...

Reviewed miles/utils/score_centering.py and miles/rollout/generate_utils/score_centering.py, which add candidate-distribution capture/validation for the new score-centering feature; no injection/auth/data-exposure surface is touched, this is internal RL training plumbing. Confirmed both ruled-out items from the candidate list: (1) validate_score_centering_sample's cross-check of sampled-token vs. candidate logprobs only fires when rollout_sampling_mask is set, since the match array is empty otherwise; (2) validate_score_centering_args is unreferenced outside tests. Two prior reviews on earlier pushes already said no issues; this run's diff is one incremental commit consistent with the PR's stated "part 2/6, wiring comes later" scope, and no third-party changes-requested review is outstanding.

@Shi-Dong
Shi-Dong force-pushed the shi/score-centering-transport branch from e70c84f to 4d6bf6c Compare September 24, 2026 17:08

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nothing blocking. The comments below are optional suggestions. There is no need to push a fix for them before merging.

Beyond the inline finding, I checked two other candidate issues from this run's hunt: (1) --loss-type choices in miles/utils/arguments.py still omit "score_centering", so the feature can't be enabled via CLI yet — consistent with this being an intermediate stack PR (2/6), not a bug in this diff; (2) the multi-turn merge path in sample_utils.py/score_centering.py for mismatched candidate fields across turns raises a clear ValueError (merge_score_centering_field) rather than crashing unguarded, so that's deliberate fail-fast behavior, not a silent defect.

Extended reasoning...

This diff adds a score-centering feature (candidate-probability capture, validation, serialization, and merge/conversion wiring) across rollout generation, sample types, and training-batch conversion — no auth/crypto/data-exposure surface, but plenty of validation logic where edge cases matter. A prior-run-confirmed bug (unhandled TypeError on zero-generated-tokens in support/filtered sampling mode) is being posted as an inline finding, so this warrants defer rather than approve. I independently verified that finding by re-reading append_score_centering_topk, and separately re-checked two of the four ruled-out candidates (CLI reachability gating and cross-turn merge error handling), confirming both are intentional/non-issues rather than additional bugs.

Comment thread miles/rollout/generate_utils/score_centering.py Outdated
@Shi-Dong
Shi-Dong force-pushed the shi/score-centering-transport branch from 4d6bf6c to e83351e Compare September 24, 2026 18:32

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code review found no issues

No high-confidence issues detected in this change.

Still open from earlier reviews (1):

  • Unresolved: 1 minor or pre-existing.

@Shi-Dong
Shi-Dong force-pushed the shi/score-centering-transport branch from e83351e to 704485c Compare September 24, 2026 19:09

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nothing blocking. The comments below are optional suggestions. There is no need to push a fix for them before merging.

This automated review ran and found the wiring gap noted inline (validate_score_centering_args is never invoked from miles_validate_args, so the new config-contract checks don't run at startup); given that, a human look at this PR is still worthwhile before merging.

Extended reasoning...

The diff adds a score-centering rollout/training pipeline touching sampling request construction, sample serialization, and training-batch conversion; no auth/crypto/injection surface is involved. The decisive fact is the confirmed inline finding that the new validate_score_centering_args function is dead code in production (only called from tests), so misconfigured score-centering runs (e.g. via --custom-config-path bypassing argparse choices) would not fail fast as intended.

Comment thread miles/utils/score_centering.py
@Shi-Dong
Shi-Dong force-pushed the shi/score-centering-transport branch from 704485c to e40d41f Compare September 24, 2026 19:43

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reviewed this PR and didn't find any bugs. My earlier findings — the TypeError on zero-new-token completions in support mode, and the missing validate_score_centering_args startup call — are both fixed in the latest commits (b566b61, e40d41f). Because this is a large, numerically intricate change to core rollout/training-data plumbing, a human look would still be worthwhile.

What was reviewed: sample merge/truncate/reset handling of the new top-k fields, wire encode/decode, and validate_score_centering_sample's shape/duplicate/probability-mass checks. Also checked: a trained token missing from its recorded top-k candidates (score_centering.py:176) — the row's consistency check is vacuously skipped rather than a false accept in the expected case; and enabling score_centering without a populated rollout path (train_data_conversion.py:124) — raises a clear ValueError rather than crashing.

Extended reasoning...

The diff adds a new score-centering training path spanning sample lifecycle (merge/trim/reset), wire serialization, CLI-arg validation, and training-batch conversion; it touches no auth/crypto/data-exposure surface. Both prior inline findings from earlier runs (empty-completion TypeError at score_centering.py, and the unwired validate_score_centering_args) were fixed by the new commits b566b61 and e40d41f pushed since the last review. No new findings were reported this run, and two additional candidate issues (silent accept of an out-of-top-k sampled token, and a missing-candidates crash path) were examined and ruled out as non-issues. Given the size and numeric-correctness intricacy of the remaining code (probability-mass invariants, support-set matching) and that this is one part of a larger stacked series, a human review is still warranted.

@Shi-Dong
Shi-Dong force-pushed the shi/score-centering-transport branch from e40d41f to 637bce7 Compare September 24, 2026 20:06

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reviewed this PR and found no bugs. The two issues I flagged in earlier reviews of this PR — the zero-new-token crash in append_score_centering_topk and validate_score_centering_args never being invoked from miles_validate_args — have both been fixed by later commits (7748886, 637bce7), each with a regression test. Because this is a large, numerically intricate feature spanning rollout generation, session serialization, and training-batch conversion, a human look would still be worthwhile.

What was reviewed: candidate-array validation/support-membership logic in score_centering.py, the new field merge/pad/reset handling in sample_utils.py and types.py, and the new CLI args/startup validation wiring in arguments.py. Checked whether validate_score_centering_sample's final check could miss an unrecorded sampled token, and whether merge_samples crashes (rather than truncating) when a turn lacks score-centering candidates — both are pre-existing/design-level observations, not new regressions.

Extended reasoning...

The diff adds a large score-centering feature: new modules for config validation and per-token candidate-distribution capture, threaded through Sample dataclass fields, session codec, rollout merge/pad logic, and the training-batch conversion pipeline. No injection/auth/data-exposure surface; the sensitive surface is numerical correctness of probability/validation logic feeding training data. Both previously-flagged inline concerns (crash on zero-token completions; missing startup validation call) were fixed with real code changes plus tests since the last review, and this run found no new issues, but the change's size and complexity (six-part stack, heavy vectorized validation) still warrant a human pass rather than an autonomous approval.

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nothing blocking. The comments below are optional suggestions. There is no need to push a fix for them before merging.

Beyond the inline finding, I checked whether the new score-centering request/response wiring (configure_score_centering_request, append_score_centering_topk in miles/rollout/generate_utils/score_centering.py) is invoked from any production rollout path (sglang_rollout.py's generate, generate_endpoint_utils.py, session merge.py) — it isn't yet, which would make train_data_conversion.py's validate_score_centering_sample raise on a real rollout, but I'm not flagging that as a bug here since this PR is explicitly stack part 2/6 (plumbing only) and later parts likely add those call sites.

Extended reasoning...

This PR wires score-centering candidate-distribution plumbing through Sample, codec, sampling-mask, and train-data conversion across 12 files; the confirmed inline finding is a dead validation check (use_sampling_support_replay is forced before validate_score_centering_args runs, so its filtered-rollout guard can never trigger). I additionally confirmed configure_score_centering_request/append_score_centering_topk are never called from any real generate() call site, meaning a real rollout with --loss-type score_centering would currently fail sample validation, but treated this as intended given the PR's explicit stack-part-2/6, dependency-ordered-merge framing rather than a standalone bug.

Comment thread miles/utils/score_centering.py Outdated

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reviewed this PR again and found no new bugs. The commits pushed since my last review (2886558 through aa584fc) address all three issues I previously flagged inline: the empty-completion TypeError in append_score_centering_topk is now guarded on n, validate_score_centering_args is now wired into miles_validate_args, and the vacuous use_sampling_support_replay check was removed in favor of the already-derived flag. Given this is a substantial new RL loss feature with several interacting validation paths, a human look is still worthwhile before merging.

Extended reasoning...

The diff adds score-centering candidate-distribution plumbing across sample types, session codec, sampling-mask logic, argument validation, and training-batch conversion; the security-relevant surface is limited to input validation (no auth/crypto/injection). All three previously-raised inline findings (a TypeError-vs-ValueError edge case, a missing validator call site, and a dead-code check) are now fixed by the latest commits, and no third-party CHANGES_REQUESTED review is outstanding. The change remains large and touches training correctness, which is why I'm deferring rather than approving outright.

guapisolo added a commit that referenced this pull request Oct 2, 2026
Follow the #3622 split in test_score_centering_filtered.py: import the
helpers from miles.rollout.generate_utils.rollout_topk_logprobs under their new
names, and select --rollout-sampling-logprobs-mode support where the
tests build filtered score-centering args, since score centering now
requires the mode that matches its sampling. The unnormalized-support and
top-p-without-top-k cases go: the rollout args and replay own those checks
after #3622 trimmed the candidate helpers.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code review found no issues

No high-confidence issues detected in this change.

Still open from earlier reviews (1):

  • Unresolved: 1 minor or pre-existing.

@guapisolo

Copy link
Copy Markdown
Collaborator

codex cmt

This history rewrite from 82d7a8f9df36a77b6373ba01d522bd86addd660c to 4980d27e982ba7131e0f777ed1e107db54e6a689 will replay this PR onto the math parent with typed score-centering inputs and metrics, retaining its transport commits so the stack stays aligned. The final stack passed 213 tests with 7 GPU-dependent skips and 72 exact baseline comparisons of losses, metrics, and gradients.

@guapisolo
guapisolo force-pushed the shi/score-centering-transport branch from 82d7a8f to 4980d27 Compare October 2, 2026 09:18

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code review found no issues

No high-confidence issues detected in this change.

Still open from earlier reviews (1):

  • Unresolved: 1 minor or pre-existing.

@guapisolo

Copy link
Copy Markdown
Collaborator

codex cmt

This history rewrite from 4980d27e982ba7131e0f777ed1e107db54e6a689 to ae02067b81e66158b791505b0421f717b41da638 will replay this PR onto the updated math parent, retaining its transport commits so the stack incorporates the unified score-centering input record. The final stack passed 213 tests with 7 GPU-dependent skips and 144 exact baseline comparisons covering default and custom options, losses, metrics, and gradients.

@guapisolo
guapisolo force-pushed the shi/score-centering-transport branch from 4980d27 to ae02067 Compare October 2, 2026 17:47

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code review found no issues

No high-confidence issues detected in this change.

Still open from earlier reviews (1):

  • Unresolved: 1 minor or pre-existing.

@guapisolo

Copy link
Copy Markdown
Collaborator

codex cmt

This history rewrite from ae02067b81e66158b791505b0421f717b41da638 to 1ee4332eee53a82d6b3bdf9fe08620369fb4eb0e will rebase the transport PR onto its updated math parent and the new shared training layout, retaining its candidate-data commits and behavior. The final stack passed 219 focused tests, including native-actor/FSDP checks, real Ray transport and TP2/CP2 Gloo, with 7 GPU-dependent skips; all 75 original commits retain their raw author and committer metadata.

@guapisolo
guapisolo force-pushed the shi/score-centering-transport branch from ae02067 to 1ee4332 Compare October 2, 2026 18:27
@guapisolo

Copy link
Copy Markdown
Collaborator

codex cmt

This history rewrite from 1ee4332eee53a82d6b3bdf9fe08620369fb4eb0e to 8741fcbb60f0200cb77da9121317d3f7a48cdc1b will advance this PR onto main 53f5d94224b7fc584857eb2aba68ffdd364dc139 through the updated stack parent, incorporating the concurrent CUDA 12 build cleanup while retaining the score-centering commits and the verified shared-training adaptations. The new final stack passed 219 focused tests with 7 GPU-dependent skips; this replay had no conflicts and preserved all 77 source commits with their raw author and committer metadata.

@guapisolo
guapisolo force-pushed the shi/score-centering-transport branch from 1ee4332 to 8741fcb Compare October 2, 2026 18:33
guapisolo added a commit that referenced this pull request Oct 2, 2026
Follow the #3622 split: the native payload builders and response paths
use configure_rollout_topk_logprobs_request / append_rollout_topk_logprobs /
pad_rollout_topk_logprobs from miles.rollout.generate_utils.rollout_topk_logprobs, and
size the candidates with --rollout-top-logprobs-num instead of
score_centering_top_k(args).

The tests and the manual live probe move to the new args; the filtered
native test selects --rollout-sampling-logprobs-mode support
explicitly. The response-side gates are unchanged: they still tell a
training request that asked for candidates from an evaluation request.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
guapisolo added a commit that referenced this pull request Oct 2, 2026
SessionServerConfig carries rollout_top_logprobs_num and
rollout_sampling_logprobs_mode instead of loss_type and
score_centering_top_k, so the session server applies the same rollout
args as native generation. prepare_chat_request still calls
configure_rollout_topk_logprobs_request after the session sampling defaults and
model rules, because the helper's own sampling defaults must not
pre-empt them; merge.py sizes the recorded candidates with
args.rollout_top_logprobs_num.

The two fields are read from args directly; rollout_temperature stays
in the config unchanged, and the older-args test now covers only its
default. Tests, fixtures and the overhead bench move to the new fields,
and the sampling-rule test matches the "Rollout top-k logprobs" wording from
#3622; an injected top_p is now rejected by the replay request check, which
owns top-p/top-k bounds.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
guapisolo added a commit that referenced this pull request Oct 2, 2026
Follow the #3622 split in test_score_centering_filtered.py: import the
helpers from miles.rollout.generate_utils.rollout_topk_logprobs under their new
names, and select --rollout-sampling-logprobs-mode support where the
tests build filtered score-centering args, since score centering now
requires the mode that matches its sampling. The unnormalized-support and
top-p-without-top-k cases go: the rollout args and replay own those checks
after #3622 trimmed the candidate helpers.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code review found no issues

No high-confidence issues detected in this change.

Still open from earlier reviews (1):

  • Unresolved: 1 minor or pre-existing.

@guapisolo
guapisolo merged commit 9b061ba into main Oct 3, 2026
22 checks passed
@guapisolo
guapisolo deleted the shi/score-centering-transport branch October 3, 2026 07:23
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants