feat(validation): add paired rating range evidence - #579
Conversation
|
Important Review skippedDraft detected. Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Pro Plus Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
@opencode-agent address Operate only on exact current Draft head Implement the minimal GREEN slice from the committed design and plan without scope expansion:
Keep the essay validation-report schema unchanged in this PR. Add no provider SDK, feedback generation, new likelihood, generalized rMFRM, latent-space term, workflow/dependency/version change, or Python numerical fallback. Keep Draft; do not mark Ready, enable auto-merge, approve, merge, release, or close issue #397. A final unchanged head still requires full Python, Rust/PyO3, package, GPU-no-skip, fuzz, Security Scan, SAST, exact-head review, qualifying independent approval, and branch protection. |
|
@opencode-agent address Refetch exact current head/base before writing. This Draft already has RED public-contract tests and an implementation plan for paired automated/reference rating-range evidence. Implement only the bounded first slice described in the branch plan: Rust-owned Add Rust oracle/branch tests and Python delegation/parity tests as required by the committed plan. Keep this PR Draft until GREEN and full gates. Do not add essay-report schema wiring, provider SDKs, generated feedback, a new likelihood, rMFRM, latent-space terms, universal thresholds/pass-fail policy, or external-repository changes. Add doctoring/changelog material only after the numerical contract is GREEN and render the authoritative CHANGELOG. If head/base moves before writing, refetch/reconcile and do not use stale assumptions. |
|
Exact-head review finding for |
|
@opencode-agent address Operate only on exact current Draft head Make the smallest architecture-correct GREEN change:
|
|
@opencode-agent address Operate only on exact current Draft head Exact-head CI Complete the smallest GREEN repair without scope expansion:
Keep Draft. Do not change essay-report schemas, providers, likelihoods, generalized rMFRM, workflows, dependencies, versioning, branch protection, approval, merge, release, or issue #397 state. |
|
@opencode-agent Exact-head review follow-up on current The Rust→PyO3→Python delegation and explicit Boolean/category-count validation are now present and the direct delegation test compares every public field. Two bounded hardening gaps remain before final GREEN:
Do not expand into generalized rMFRM/rater inference or change the descriptive semantics. After these narrow repairs, focused tests plus exact-head CI/Security/SAST/fuzz and current-head review still govern Draft→Ready. |
|
@jules address Take a bounded writer lease on this Draft branch only. Exact head is now Do not change numerical behavior unless a focused current-head test proves a defect. First run the repository changelog renderer so Commit only rendered |
|
@opencode-agent address Operate only on exact current Draft head Make one bounded deterministic repair only: run |
|
@opencode-agent address Exact-current source head at handoff: Before writing, refetch the branch head and abort if it moved. If unchanged, run |
|
@opencode-agent address Operate only on exact current Draft head Perform only the release-note synchronization repair:
Keep Draft. Do not alter rating-range arithmetic, PyO3 ownership, public API semantics, unrelated AGENTS/CLAUDE content, dependencies, workflows, review identities, version, release state, or issue #397. |
|
@opencode-agent address Reacquire one bounded integration writer lease only if a final refetch still shows Draft source head The feature implementation itself already has strong predecessor evidence: exact CI previously reached Perform the smallest current-main integration repair:
Preserve Rust-owned descriptive arithmetic, thin PyO3 marshalling, Python validation-only behavior, descriptive-not-inferential interpretation, and no universal acceptance threshold. No essay-report schema, provider SDK, generalized rMFRM, dependencies/workflows/credentials/version/release or canonical #604 docs changes. |
|
@opencode-agent address The prior current-main integration handoff has had no receipt/reaction and no source movement for more than two hours. Reacquire one bounded writer lease only if a final refetch still shows Draft head Do not change the validated paired-rating-range arithmetic unless current integration tests prove a source defect. Perform only the smallest integration repair: reconcile protected main non-destructively, preserving #590 and the unique rating-range Rust/PyO3/Python-validation/tests/doctoring/fragment slice; run the authoritative changelog renderer |
|
@opencode-agent address Take a bounded writer lease only if a final refetch still shows Draft head The paired rating-range branch has already been reconciled non-destructively with current protected main as a two-parent merge commit and now retains only the 14 Rust/PyO3/Python/tests/doctoring/fragment paths belonging to this diagnostic. Preserve Rust ownership and current-main report/scaling behavior. Complete only deterministic acceptance cleanup:
Keep Draft. Do not change essay-report schema, add provider SDK/feedback/new likelihood/generalized rMFRM/latent-space terms, alter dependencies/workflows/version/release, modify canonical #604, mark Ready, approve, merge, or close #397. Stop after exact-head deterministic evidence. |
|
@opencode-agent address Take a bounded writer lease on PR #579 only if a final refetch still shows exact Draft head The Rust-owned paired rating-range implementation, PyO3 marshalling, Python validation/transport, direct-delegation tests and scientific boundary are already present. Exact-head Security, SAST and ClusterFuzzLite are green; the known remaining repository integration failure is managed changelog parity, and protected main has since advanced through #618. Execute only deterministic integration cleanup:
Keep Draft. Do not change the essay report schema, add provider/feedback generation, generalized rMFRM or latent-space arithmetic, modify dependencies/workflows/version/release, or touch canonical architecture PR #604. |
|
Superseded bookkeeping note: the exact-current writer handoff for head |
|
@opencode-agent address Take a bounded writer lease only if a final refetch still shows Draft head Current exact-head Security Scan Non-destructively reconcile protected main, retaining #618 and all accepted-main behavior, then run the focused paired-rating-range Rust/PyO3/Python delegation/parity tests, render/check authoritative changelog, formatting/lint and |
|
@opencode-agent address Take a bounded writer lease only if a final refetch still shows Draft head Do not change the diagnostic semantics. Run focused Rust rating-range tests, PyO3/delegation tests, Python evidence tests, then synchronize only managed release notes using the authoritative changelog renderer ( |
|
@jules address Fallback bounded writer handoff for exact Draft head Do not change paired rating-range diagnostic semantics. Exact-head Rust/PyO3, package/reinstall, GPU, fuzz, ClusterFuzzLite, Security Scan and SAST are green; the feature tests are green and the remaining Python failure is deterministic managed changelog parity. Run focused Rust rating-range, PyO3/delegation and Python evidence tests, then only |
|
@jules address Superseding paired-rating-range integration handoff after both CodeQL dependency merges. Fresh identities: Draft #579 exact head Do not change diagnostic semantics. Predecessor exact-head evidence has Rust/PyO3/package/GPU/fuzz/ClusterFuzzLite/Security/SAST and feature tests GREEN with only deterministic managed changelog parity outstanding. Reconcile protected main non-destructively, preserving both CodeQL Keep Draft. No essay-report/provider/likelihood/latent-space/rMFRM/dependency/workflow/version/release/canonical-docs #604/Ready/approval/merge/issue-closure expansion. Stop source writes after one coherent verified update; fresh exact-head full CI/Security/SAST/review returns to the maintainer loop. |
c4a290b to
0d3ba94
Compare
0d3ba94 to
c4a290b
Compare
|
@opencode-agent address Fresh current-main reconciliation handoff for Draft #579. Immediately refetch exact source head, protected Fresh compare is Revalidate the feature itself on the reconciled head: Rust remains the sole descriptive-arithmetic owner; Python performs bounded validation/marshalling only; paired observations remain exact same-case evidence; population-divisor SD, span/distinct-category ratios, endpoint gaps and degenerate-reference unavailable semantics remain unchanged; no universal range-compression threshold or validity claim is introduced. Re-run direct field-by-field Rust↔Python delegation, caller immutability, degenerate/error tests and realistic paired range-compression fixtures, then full same-head gates and changelog render/check. Keep Draft; do not widen to generalized rMFRM, essay report schema/provider/likelihood work, workflows/dependencies/version/release or canonical #604/#621 docs. |
|
@opencode-agent address Fresh exact-current reconciliation/replacement handoff for Draft #579 / issue #397. Before any source write, refetch exact branch head Fresh compare is Preserve the intended scientific contract: descriptive paired-category support/dispersion evidence only; Rust owns min/max/distinct counts/spans/population-divisor SDs/ratios/endpoint gaps and conservative range-compression flags; degenerate reference span/SD returns unavailable ratios rather than NaN/Inf; Python rejects invalid/Boolean/non-integral labels before coercion, marshals immutable arrays, and never recomputes statistics; no universal threshold, generalized rMFRM, construct-validity/fairness claim, provider SDK or generated feedback. Reuse the current package's canonical PyO3 registration pattern rather than resurrecting obsolete extension-module registration. Re-establish focused RED→GREEN/parity against the integrated/replacement branch, then run Rust unit/integration, PyO3/public delegation, Python rating-range tests, meaningful statement/branch/docstrings, repository-authoritative changelog render/check, |
|
Superseded by surgical #748 (or next number) on current main. |
Buyer-visible gap
Issue #397 requires evidence that an automated essay scorer is not merely agreeing on average while using a materially narrower portion of the ordinal rating scale. Existing validation exposes agreement, association, fairness/SMD, severity-related evidence and human-human degradation, but did not expose explicit paired category-range evidence.
Implemented bounded slice
The branch contains a Rust-owned descriptive diagnostic over the same paired automated/reference validation cases and a thin PyO3/Python product path.
mlsirm_core::rating_range::paired_rating_range_evidenceowns all descriptive arithmetic and returns paired sample size, automated/reference minima/maxima, distinct-category counts, spans, population-divisor empirical SDs, relative ratios where identified, signed endpoint gaps, conservativenarrower_observed_support, and strictercentral_tendency_signal. Degenerate reference span/SD returns unavailable relative evidence rather than NaN/Inf; no universal acceptance threshold is encoded.Rust/PyO3/Python ownership
crates/mlsirm-core/src/rating_range.rs;crates/fast-mlsirm-py/src/rating_range_bindings.rs;python/fast_mlsirm/rating_range.py;Python does not recompute the diagnostic.
Scientific boundary
This is descriptive paired-sample evidence, not a generalized many-facet range-restriction parameter. Rater severity, agreement, central-category support and inferential rater range restriction remain different constructs. A future MFRM/rMFRM range-restriction model requires a separate Rust likelihood/identification contract, connected-design tests, true-parameter bias/MAE/RMSE/coverage/convergence and uncertainty evidence, relation-safe model comparison and CPU/GPU parity where material.
Exact-current evidence
Freshly revalidated:
main:8db4bf358b0a469915d6c5e336054f4a4f9c6b46;f11f46c3a4945c76779c545ba9d9cf9207410de0;31309352764: Rust/PyO3, package/reinstall/release acceptance, enterprise sales-readiness smoke, explicit GPU no-skip and fuzz succeed; Python reaches the full suite with the paired-rating-range feature GREEN and fails only deterministic managed-CHANGELOG.mdrender parity;31309352732: success;31309352750: success;31309352734: success;docs/doctoring/paired-rating-range-evidence.mdanddocs/changelog.d/397-paired-rating-range-evidence.mdare authoritative branch evidence. Older body identities such asd8c06ace...are predecessor state only.Remaining Draft gate
Keep Draft. Reconcile current protected main non-destructively, preserving accepted-main behavior and only this unique rating-range Rust/PyO3/Python/tests/doctoring/fragment slice; render/check authoritative changelog; then require one unchanged exact head with focused/full Python coverage/docstrings, Rust workspace/all-target/clippy/fmt, PyO3/wheel/package/reinstall, explicit GPU-no-skip, fuzz/ClusterFuzzLite, Security Scan, SAST, fresh current-head automated review, zero valid unresolved findings and repository approval/branch-protection policy.
No essay-report schema change, provider SDK, generated feedback, new likelihood, latent-space term, generalized rMFRM or release bump belongs here.
Advances #397.