GPGPU-in-Rust: device sub-option for the Rust backend (supersedes Python compute_backend) - #109
Conversation
Resolve GPU support by doing GPGPU inside the Rust core instead of adding a
competing Python compute_backend axis. The backend axis stays {numpy, rust,
auto}; the Rust backend gains a device sub-option (rust_device = cpu|gpu|auto,
default auto).
Rust core (crates/mlsirm-core):
- Add a wgpu (MIT/Apache-2.0) GPGPU implementation of the penalized
neg-loglik + gradient hot path in gpu.rs, using a race-free per-output-thread
reduction (no atomics; WGSL has no f64/atomic-float). Kernels run in f32; the
L2 penalty and final objective/tau reductions are done in f64 on the host,
mirroring the CPU reference.
- Add Device{Cpu,Gpu,Auto} and neg_loglik_and_grad_device with a runtime CPU
fallback: Gpu/Auto use the GPU when an adapter is present and otherwise fall
back to the identical CPU path, so CI and GPU-less machines pass unchanged.
- Feature-gate wgpu behind a default `gpu` feature (build --no-default-features
for a CPU-only core).
PyO3 binding: thread an optional device argument (default "cpu") into
neg_loglik_and_grad.
Python: add FitConfig.rust_device (validated), thread it through objective/fit,
record it on FitResult and fit_summary.json, and expose
`fast-mlsirm fit --rust-device`. No cupy/mlx/opencl code and no top-level
compute_backend field are introduced.
Tests: Rust device-parity test (runs the real GPU kernels when present),
Python parity test asserting the rust device paths match numpy within tolerance,
plus config/CLI coverage. cargo build, cargo test, the maturin wheel build, and
pytest are all green; the GPU path gracefully falls back to CPU without a GPU.
docs/papers: add Wu et al. (2021, arXiv:2108.11579, CC BY 4.0) grounding fast,
accelerator-friendly IRT estimation.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RTAMs4bpSZS77Xe3RQjv9P
The coverage-evidence gate runs `cargo llvm-cov --workspace --all-features --fail-under-lines 100`. The wgpu GPGPU path could not reach 100% lines in a single run: on a GPU-equipped host the CPU fallback branch was dead, and on a GPU-less host the GPU-success path was dead, so neither state covered both sides of `neg_loglik_and_grad_device`. Split the GPU/CPU resolution into a pure `finish_device` helper and unit-test both the GPU-succeeded and GPU-unavailable branches directly, so both are exercised regardless of whether the test host has a GPU adapter. Add a device-path test with `mask: None` to cover the dense-matrix host branch in `gpu.rs` (previously line 408). Verified locally: `cargo llvm-cov --workspace --all-features --fail-under-lines 100 --show-missing-lines` reports gpu.rs 100% / lib.rs 100% lines and exits 0. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RTAMs4bpSZS77Xe3RQjv9P
|
Prepared current head What I changed on top of the GPGPU branch:
Local verification on this head:
Local Rust verification note:
I am leaving this PR behind lower-numbered PRs and will not merge it until the current-head required checks and OpenCode review pass. |
OpenCode Review Overview
Pull request overviewOpenCode reviewed the current-head bounded evidence and found no blocking issues. FindingsNo blocking findings. SummaryApproval sufficiency: bounded evidence supplied affirmative approval evidence for changed files, coverage/docstring posture, risk surfaces, and current-head verification; approval is not based merely on the absence of known blockers.
Changed-File Evidence Mapflowchart LR
PR["PR changed files"] --> Evidence["OpenCode bounded evidence"]
Evidence --> S1["Changed file (15 files)"]
S1 --> I1["repository behavior"]
I1 --> R1["Review risk: Changed file (15 files)"]
R1 --> V1["required checks"]
Evidence --> S2["Docs: README.md"]
S2 --> I2["operator or user guidance"]
I2 --> R2["Review risk: Docs: README.md"]
R2 --> V2["docs review"]
Evidence --> S3["Test (3 files)"]
S3 --> I3["regression suite"]
I3 --> R3["Review risk: Test (3 files)"]
R3 --> V3["targeted test run"]
|
There was a problem hiding this comment.
Pull request overview
OpenCode reviewed the current-head bounded evidence and found no blocking issues.
Findings
No blocking findings.
Summary
Approval sufficiency: bounded evidence supplied affirmative approval evidence for changed files, coverage/docstring posture, risk surfaces, and current-head verification; approval is not based merely on the absence of known blockers.
Verification posture: CodeGraph evidence was initialized and bounded current-head evidence reviewed for changed-file evidence including CHANGELOG.md, Cargo.lock, README.md, crates/fast-mlsirm-py/Cargo.lock, crates/fast-mlsirm-py/src/lib.rs, and 14 more.
Linter/static: workflow/static review evidence is bounded by the current-head GitHub Checks gate and changed-file evidence.
TDD/regression: coverage execution evidence and focused changed hunks were reviewed from bounded-review-evidence.md.
Coverage: coverage execution evidence reports supported repository test suites passed.
Docstring coverage: coverage execution evidence reports configured repository docstring gates passed or docstring coverage was advisory.
DAG: CodeGraph/source-backed behavior map connects CHANGELOG.md to the affected review, runtime, or workflow path and required checks.
PoC/execution: coverage-evidence job executed on the current head and reported PASS.
DDD/domain: workflow and repository-governance invariants were reviewed against changed files in bounded evidence.
CDD/context: CodeGraph evidence, changed-file history, and focused hunks were reviewed from bounded-review-evidence.md.
Similar issues: changed-file history evidence was reviewed for comparable local precedents.
Claim/concept check: bounded evidence, repository source, current-head workflow evidence, and, where numeric, scientific, statistical, or literature-backed claims are affected, original-paper/formula evidence and parameter-recovery expectations were used for claims.
Standards search: standards and external-source checks are delegated to configured OpenCode web_search/Context7/DeepWiki sources when applicable; no evidence-backed standards blocker is present in bounded evidence.
Compatibility/convention: changed workflow/script conventions, object naming, and reserved-word safety for schema/API/config/code surfaces were checked in bounded evidence.
Breaking-change/backcompat: deployment evidence and changed-file history were checked for backward-compatibility risk.
Performance: changed surfaces were checked for performance risk in bounded evidence.
Developer experience: changed automation, review, test, setup, and maintenance surfaces were checked for helpful or obstructive DX impact in bounded evidence.
User experience: connected user, operator, API, CLI, documentation, review-comment, status-check, rendering, and workflow-reader behavior was checked for contradictions against code, docs, and tests in bounded evidence.
Visual/DOM: Playwright visual, DOM locator, ARIA snapshot, console, and responsive evidence were checked when a web UI surface was present; for non-web surfaces, API/CLI/log/docs/workflow interaction evidence was reviewed instead.
Accessibility/i18n: accessibility, localization, and human-readable text surfaces were checked where UI, CLI, API message, docs, logs, or review text changed.
Supply-chain/license: dependency, package, model, container, and external-tool changes were checked in bounded evidence.
Packaging: package, build, test, lint, and security contracts were checked in bounded evidence.
Security/privacy: workflow-token, review-gate, and repository-automation security/privacy boundaries were checked in bounded evidence.
- Result: APPROVE
- Reason: The changes introduce a GPGPU-in-Rust backend with a device sub-option, unifying the backend abstraction. All tests pass, and the implementation is well-documented and structured.
- Head SHA:
01b40bbcafa0b2f04119eabde2215a1ce1e4de7d - Workflow run: 29081551170
- Workflow attempt: 1
Changed-File Evidence Map
flowchart LR
PR["PR changed files"] --> Evidence["OpenCode bounded evidence"]
Evidence --> S1["Changed file (15 files)"]
S1 --> I1["repository behavior"]
I1 --> R1["Review risk: Changed file (15 files)"]
R1 --> V1["required checks"]
Evidence --> S2["Docs: README.md"]
S2 --> I2["operator or user guidance"]
I2 --> R2["Review risk: Docs: README.md"]
R2 --> V2["docs review"]
Evidence --> S3["Test (3 files)"]
S3 --> I3["regression suite"]
I3 --> R3["Review risk: Test (3 files)"]
R3 --> V3["targeted test run"]
|
Current-head fix/update for PR #109:
Remaining blocker, if GitHub still shows blocked: org/repo native branch ruleset approvals are still required ( |
|
@copilot resolve the merge conflicts on this branch. |
Head branch was pushed to by a user without write access
Conflicts are resolved on this branch in commit |
There was a problem hiding this comment.
Pull request overview
OpenCode cannot approve yet because required coverage evidence did not pass.
Review outcome
1. HIGH .github/workflows/opencode-review.yml:1 - Coverage evidence did not prove required test/docstring evidence
-
Problem: The required coverage-evidence job result was
failure, so OpenCode cannot establish approval sufficiency for this head. -
Root cause: Automated approval is only valid when the same-head coverage-evidence job proves supported repository test suites passed and configured docstring gates passed or were advisory, or reports not applicable because no supported source files or package manifests exist. Missing, failed, skipped, unavailable, or unsupported-tooling test evidence is a blocker.
-
Fix: Install or configure the repository test/docstring evidence tooling when source files or package manifests exist, rerun the current-head coverage-evidence job, and approve only after it reports
successwith required evidence or explicit no-source not-applicable evidence. -
Regression test: Keep the approval branch checking
needs.coverage-evidence.result == successbefore posting APPROVE, and publish REQUEST_CHANGES when coverage-evidence blocker states such as cancelled, skipped, failed, unsupported-tooling, or below-100 evidence are present. -
Result: REQUEST_CHANGES
-
Reason: coverage-evidence result was
failure, so required test/docstring evidence was not proven for current head8a12774f47ca0d2d6367c40648085d35dcd330f6. -
Head SHA:
8a12774f47ca0d2d6367c40648085d35dcd330f6 -
Workflow run: 29090880398
-
Workflow attempt: 2
Coverage evidence
Coverage Evidence
- Head SHA:
8a12774f47ca0d2d6367c40648085d35dcd330f6 - Required test evidence: supported repository test suites must pass.
- Required docstring evidence: repository-owned docstring gates must pass when configured; otherwise docstring coverage is advisory.
Python project dependencies (.)
$ uv sync --project . --extra dev
Using CPython 3.12.3 interpreter at: /usr/bin/python3
Creating virtual environment at: .venv
Resolved 13 packages in 130ms
Building fast-mlsirm @ file:///home/runner/work/fast-mlsirm/fast-mlsirm/pr-head
Downloading pygments (1.2MiB)
Downloading numpy (15.9MiB)
Downloaded pygments
Downloaded numpy
Built fast-mlsirm @ file:///home/runner/work/fast-mlsirm/fast-mlsirm/pr-head
Prepared 7 packages in 1m 05s
Installed 7 packages in 20ms
+ fast-mlsirm==0.1.0 (from file:///home/runner/work/fast-mlsirm/fast-mlsirm/pr-head)
+ iniconfig==2.3.0
+ numpy==2.5.1
+ packaging==26.2
+ pluggy==1.6.0
+ pygments==2.20.0
+ pytest==9.1.1
- Result: PASS
Python coverage with missing-line report (.)
$ bash -c cd\ \"\$1\"\ \&\&\ PYTHONPATH=.\ uv\ run\ --with\ coverage\ --with\ pytest\ coverage\ run\ -m\ pytest\ tests\ \&\&\ uv\ run\ --with\ coverage\ coverage\ report\ --show-missing bash .
Installed 6 packages in 16ms
============================= test session starts ==============================
platform linux -- Python 3.12.3, pytest-9.1.1, pluggy-1.6.0
rootdir: /home/runner/work/fast-mlsirm/fast-mlsirm/pr-head
configfile: pyproject.toml
collected 201 items
tests/test_backend.py . [ 0%]
tests/test_benchmark_report.py .. [ 1%]
tests/test_buyer_evidence_packet.py .... [ 3%]
tests/test_cli.py .................... [ 13%]
tests/test_commercial_release_builder.py .... [ 15%]
tests/test_config.py ..................... [ 25%]
tests/test_diagnostics.py ............... [ 33%]
tests/test_estimator_mmle.py ....... [ 36%]
tests/test_figma_evidence_sync.py ... [ 38%]
tests/test_fit_dos.py . [ 38%]
tests/test_fit_pipeline.py ..... [ 41%]
tests/test_initial_params.py .. [ 42%]
tests/test_io.py .. [ 43%]
tests/test_irt_stability.py ....... [ 46%]
tests/test_math.py ....... [ 50%]
tests/test_objective.py .............. [ 57%]
tests/test_pr_queue_governance.py .... [ 59%]
tests/test_procurement_due_diligence.py ... [ 60%]
tests/test_release_evidence_index.py .. [ 61%]
tests/test_report.py .......F... [ 67%]
tests/test_rust_parity.py ......................................... [ 87%]
tests/test_sales_readiness.py ....................... [ 99%]
tests/test_simulation.py .. [100%]
=================================== FAILURES ===================================
_____________ test_render_table_section_charts_later_numeric_rows ______________
tmp_path = PosixPath('/tmp/pytest-of-runner/pytest-0/test_render_table_section_char0')
def test_render_table_section_charts_later_numeric_rows(tmp_path):
source = tmp_path / "fit_diagnostics.json"
out = tmp_path / "report.html"
item_ids = list(range(13))
outfit = [None] * 12 + [1.2]
source.write_text(
json.dumps(
{
"model_fit": {"loglik": -3.2},
"itemfit": {"item_id": item_ids, "outfit_mnsq": outfit, "observed_count": [4] * 13},
}
),
encoding="utf-8",
)
render_diagnostics_report(source, out)
html = out.read_text(encoding="utf-8")
> assert '<div class="bar-chart" role="img" aria-label="Compact diagnostics bar chart">' in html
E assert '<div class="bar-chart" role="img" aria-label="Compact diagnostics bar chart">' in '<!doctype html>\n<html lang="en">\n<head>\n<meta charset="utf-8">\n<meta name="viewport" content="width=device-width,...>\n</tbody>\n</table>\n</div>\n<p class="table-note">Showing 12 of 13 rows.</p>\n</section>\n</main>\n</body>\n</html>'
tests/test_report.py:213: AssertionError
=========================== short test summary info ============================
FAILED tests/test_report.py::test_render_table_section_charts_later_numeric_rows - assert '<div class="bar-chart" role="img" aria-label="Compact diagnostics bar chart">' in '<!doctype html>\n<html lang="en">\n<head>\n<meta charset="utf-8">\n<meta name="viewport" content="width=device-width,...>\n</tbody>\n</table>\n</div>\n<p class="table-note">Showing 12 of 13 rows.</p>\n</section>\n</main>\n</body>\n</html>'
======================== 1 failed, 200 passed in 2.74s =========================
- Result: FAIL (exit 1)
Python docstring coverage advisory
$ bash -c python3\ -m\ interrogate\ .\ \|\|\ true
RESULT: PASSED (minimum: 0.0%, actual: 2.9%)
- Result: PASS
Rust coverage tooling (cargo-llvm-cov)
$ cargo install cargo-llvm-cov --locked
Updating crates.io index
Downloading crates ...
Downloaded cargo-llvm-cov v0.8.7
Installing cargo-llvm-cov v0.8.7
Updating crates.io index
Updating crates.io index
Downloading crates ...
Downloaded cargo-config2 v0.1.44
Downloaded itoa v1.0.18
Downloaded lcov2cobertura v1.0.9
Downloaded shared_thread v0.2.0
Downloaded glob v0.3.3
Downloaded fs-err v3.3.0
Downloaded filetime v0.2.29
Downloaded shell-escape v0.1.5
Downloaded os_pipe v1.2.3
Downloaded shared_child v1.1.1
Downloaded serde_spanned v1.1.1
Downloaded bitflags v2.11.1
Downloaded same-file v1.0.6
Downloaded xattr v1.6.1
Downloaded zmij v1.0.21
Downloaded rustc-demangle v0.1.27
Downloaded toml_parser v1.1.2+spec-1.1.0
Downloaded camino v1.2.2
Downloaded walkdir v2.5.0
Downloaded toml_datetime v1.1.1+spec-1.1.0
Downloaded toml v1.1.2+spec-1.1.0
Downloaded tar v0.4.45
Downloaded ruzstd v0.8.3
Downloaded memchr v2.8.0
Downloaded quote v1.0.45
Downloaded opener v0.8.4
Downloaded serde_json v1.0.149
Downloaded regex v1.12.3
Downloaded winnow v1.0.2
Downloaded aho-corasick v1.1.4
Downloaded quick-xml v0.39.4
Downloaded duct v1.1.1
Downloaded lexopt v0.3.2
Downloaded anyhow v1.0.102
Downloaded syn v2.0.117
Downloaded autocfg v1.5.0
Downloaded bstr v1.12.1
Downloaded errno v0.3.14
Downloaded regex-syntax v0.8.10
Downloaded rustix v1.1.4
Downloaded regex-automata v0.4.14
Downloaded linux-raw-sys v0.12.1
Compiling serde_core v1.0.228
Compiling memchr v2.8.0
Compiling libc v0.2.186
Compiling proc-macro2 v1.0.106
Compiling regex-syntax v0.8.10
Compiling aho-corasick v1.1.4
Compiling quote v1.0.45
Compiling unicode-ident v1.0.24
Compiling rustix v1.1.4
Compiling anyhow v1.0.102
Compiling zmij v1.0.21
Compiling bitflags v2.11.1
Compiling regex-automata v0.4.14
Compiling linux-raw-sys v0.12.1
Compiling serde v1.0.228
Compiling winnow v1.0.2
Compiling autocfg v1.5.0
Compiling fs-err v3.3.0
Compiling toml_parser v1.1.2+spec-1.1.0
Compiling serde_spanned v1.1.1
Compiling toml_datetime v1.1.1+spec-1.1.0
Compiling syn v2.0.117
Compiling cfg-if v1.0.4
Compiling camino v1.2.2
Compiling serde_json v1.0.149
Compiling filetime v0.2.29
Compiling toml v1.1.2+spec-1.1.0
Compiling xattr v1.6.1
Compiling serde_derive v1.0.228
Compiling regex v1.12.3
Compiling bstr v1.12.1
Compiling shared_child v1.1.1
Compiling os_pipe v1.2.3
Compiling quick-xml v0.39.4
Compiling shared_thread v0.2.0
Compiling itoa v1.0.18
Compiling rustc-demangle v0.1.27
Compiling same-file v1.0.6
Compiling walkdir v2.5.0
Compiling cargo-config2 v0.1.44
Compiling lcov2cobertura v1.0.9
Compiling duct v1.1.1
Compiling opener v0.8.4
Compiling tar v0.4.45
Compiling termcolor v1.4.1
Compiling ruzstd v0.8.3
Compiling glob v0.3.3
Compiling shell-escape v0.1.5
Compiling lexopt v0.3.2
Compiling cargo-llvm-cov v0.8.7
Finished `release` profile [optimized] target(s) in 56.59s
Installing /home/runner/.cargo/bin/cargo-llvm-cov
Installed package `cargo-llvm-cov v0.8.7` (executable `cargo-llvm-cov`)
- Result: PASS
Software Vulkan adapter (Mesa lavapipe) for GPGPU coverage
$ bash -c sudo\ apt-get\ update\ \&\&\ sudo\ apt-get\ install\ -y\ --no-install-recommends\ mesa-vulkan-drivers\ libvulkan1\ vulkan-tools\ \|\|\ true
Get:1 file:/etc/apt/apt-mirrors.txt Mirrorlist [144 B]
Get:6 https://packages.microsoft.com/repos/azure-cli noble InRelease [3564 B]
Hit:2 http://azure.archive.ubuntu.com/ubuntu noble InRelease
Get:7 https://packages.microsoft.com/ubuntu/24.04/prod noble InRelease [3600 B]
Get:3 http://azure.archive.ubuntu.com/ubuntu noble-updates InRelease [126 kB]
Get:4 http://azure.archive.ubuntu.com/ubuntu noble-backports InRelease [126 kB]
Get:8 https://dl.google.com/linux/chrome-stable/deb stable InRelease [1825 B]
Get:5 http://azure.archive.ubuntu.com/ubuntu noble-security InRelease [126 kB]
Get:9 https://packages.microsoft.com/repos/azure-cli noble/main amd64 Packages [2314 B]
Get:10 https://packages.microsoft.com/ubuntu/24.04/prod noble/main arm64 Packages [175 kB]
Get:11 https://packages.microsoft.com/ubuntu/24.04/prod noble/main armhf Packages [11.7 kB]
Get:12 https://packages.microsoft.com/ubuntu/24.04/prod noble/main amd64 Packages [208 kB]
Get:13 http://azure.archive.ubuntu.com/ubuntu noble-updates/main amd64 Packages [1092 kB]
Get:14 http://azure.archive.ubuntu.com/ubuntu noble-updates/main Translation-en [269 kB]
Get:15 http://azure.archive.ubuntu.com/ubuntu noble-updates/main amd64 Components [181 kB]
Get:16 http://azure.archive.ubuntu.com/ubuntu noble-updates/universe amd64 Packages [1660 kB]
## Changed-File Evidence Map
```mermaid
flowchart LR
PR["PR changed files"] --> Evidence["OpenCode bounded evidence"]
Evidence --> S1["Changed file (15 files)"]
S1 --> I1["repository behavior"]
I1 --> R1["Review risk: Changed file (15 files)"]
R1 --> V1["required checks"]
Evidence --> S2["Docs: README.md"]
S2 --> I2["operator or user guidance"]
I2 --> R2["Review risk: Docs: README.md"]
R2 --> V2["docs review"]
Evidence --> S3["Test (3 files)"]
S3 --> I3["regression suite"]
I3 --> R3["Review risk: Test (3 files)"]
R3 --> V3["targeted test run"]
There was a problem hiding this comment.
Pull request overview
OpenCode reviewed the current-head bounded evidence and found no blocking issues.
Findings
No blocking findings.
Summary
Approval sufficiency: bounded evidence supplied affirmative approval evidence for changed files, coverage/docstring posture, risk surfaces, and current-head verification; approval is not based merely on the absence of known blockers.
Verification posture: CodeGraph evidence was initialized and bounded current-head evidence reviewed for changed-file evidence including CHANGELOG.md, Cargo.lock, README.md, crates/fast-mlsirm-py/Cargo.lock, crates/fast-mlsirm-py/src/lib.rs, and 14 more.
Linter/static: workflow/static review evidence is bounded by the current-head GitHub Checks gate and changed-file evidence.
TDD/regression: coverage execution evidence and focused changed hunks were reviewed from bounded-review-evidence.md.
Coverage: coverage execution evidence reports supported repository test suites passed.
Docstring coverage: coverage execution evidence reports configured repository docstring gates passed or docstring coverage was advisory.
DAG: CodeGraph/source-backed behavior map connects CHANGELOG.md to the affected review, runtime, or workflow path and required checks.
PoC/execution: coverage-evidence job executed on the current head and reported PASS.
DDD/domain: workflow and repository-governance invariants were reviewed against changed files in bounded evidence.
CDD/context: CodeGraph evidence, changed-file history, and focused hunks were reviewed from bounded-review-evidence.md.
Similar issues: changed-file history evidence was reviewed for comparable local precedents.
Claim/concept check: bounded evidence, repository source, current-head workflow evidence, and, where numeric, scientific, statistical, or literature-backed claims are affected, original-paper/formula evidence and parameter-recovery expectations were used for claims.
Standards search: standards and external-source checks are delegated to configured OpenCode web_search/Context7/DeepWiki sources when applicable; no evidence-backed standards blocker is present in bounded evidence.
Compatibility/convention: changed workflow/script conventions, object naming, and reserved-word safety for schema/API/config/code surfaces were checked in bounded evidence.
Breaking-change/backcompat: deployment evidence and changed-file history were checked for backward-compatibility risk.
Performance: changed surfaces were checked for performance risk in bounded evidence.
Developer experience: changed automation, review, test, setup, and maintenance surfaces were checked for helpful or obstructive DX impact in bounded evidence.
User experience: connected user, operator, API, CLI, documentation, review-comment, status-check, rendering, and workflow-reader behavior was checked for contradictions against code, docs, and tests in bounded evidence.
Visual/DOM: Playwright visual, DOM locator, ARIA snapshot, console, and responsive evidence were checked when a web UI surface was present; for non-web surfaces, API/CLI/log/docs/workflow interaction evidence was reviewed instead.
Accessibility/i18n: accessibility, localization, and human-readable text surfaces were checked where UI, CLI, API message, docs, logs, or review text changed.
Supply-chain/license: dependency, package, model, container, and external-tool changes were checked in bounded evidence.
Packaging: package, build, test, lint, and security contracts were checked in bounded evidence.
Security/privacy: workflow-token, review-gate, and repository-automation security/privacy boundaries were checked in bounded evidence.
- Result: APPROVE
- Reason: All tests pass, coverage is sufficient, and no unresolved issues remain.
- Head SHA:
5c82c39c353cb866b13b9191bd1275546703dbc3 - Workflow run: 29101095089
- Workflow attempt: 1
Changed-File Evidence Map
flowchart LR
PR["PR changed files"] --> Evidence["OpenCode bounded evidence"]
Evidence --> S1["Changed file (15 files)"]
S1 --> I1["repository behavior"]
I1 --> R1["Review risk: Changed file (15 files)"]
R1 --> V1["required checks"]
Evidence --> S2["Docs: README.md"]
S2 --> I2["operator or user guidance"]
I2 --> R2["Review risk: Docs: README.md"]
R2 --> V2["docs review"]
Evidence --> S3["Test (3 files)"]
S3 --> I3["regression suite"]
I3 --> R3["Review risk: Test (3 files)"]
R3 --> V3["targeted test run"]
Summary
This PR delivers GPU support for the estimator by doing GPGPU inside the Rust core, unifying on a single backend axis. It supersedes PR #51, which added a competing Python
compute_backendaxis (cpu/cuda/mlx/opencl) with a cupy/mlx reimplementation of the objective that collided head-on with the existing backend abstraction inbackend.py/config.py/types.py/objective.py/fit.py.Resolution (per maintainer direction): keep the single backend axis
{numpy, rust, auto}; the Rust backend gains GPU execution (CPU and GPGPU) via a permissive Rust GPU library, exposed as a device sub-option (rust_device = cpu|gpu|auto, defaultauto) rather than a new top-level Python axis. No cupy/mlx/opencl code is introduced.What changed
Rust core (
crates/mlsirm-core)gpu.rs: a wgpu (MIT/Apache-2.0) GPGPU implementation of the penalized negative-log-likelihood and gradient hot path. Each gradient slot is owned by exactly one GPU invocation that reduces over its contributing axis, so there are no write races and no atomics (WGSL has neitherf64nor atomic-float). Kernels run inf32; the L2 penalty and the objective/taureductions are done inf64on the host, mirroring the CPU reference exactly.Device { Cpu, Gpu, Auto }+neg_loglik_and_grad_device(...)with a runtime CPU fallback:Gpu/Autouse the GPU when an adapter is present and otherwise fall back to the identical CPU path. No GPU is ever required.gpufeature (--no-default-featuresyields a CPU-only core with no wgpu graph).PyO3 binding (
crates/fast-mlsirm-py): optionaldeviceargument (default"cpu") threaded intoneg_loglik_and_grad.Python
FitConfig.rust_device(validated:cpu/gpu/auto), threaded throughobjective.pyandfit.py.FitResult.rust_deviceprovenance, persisted infit_summary.json.fast-mlsirm fit --backend rust --rust-device {auto,cpu,gpu}; the resolved device is reported in the JSON output.main), keeping the single{numpy, rust}axis intact.Tests
cpu/gpu/auto) match the NumPy reference within tolerance — the estimator core is not silently wrong.Docs:
docs/papers/adds Wu et al. (2021, arXiv:2108.11579, CC BY 4.0), which grounds fast, accelerator-friendly IRT estimation; README + CHANGELOG document the device option.Verification
cargo fmt --check,cargo clippy --workspace— clean.cargo test --workspace(6 tests, incl. GPU parity on Metal) — green.cargo teston the PyO3 crate (3 tests) — green.maturinwheel build — succeeds.pytest— 137 passed.gpu/autogracefully fall back to thef64CPU implementation and all tests pass without a GPU.Licensing
Permissive only: wgpu is MIT/Apache-2.0. No GPL/AGPL. No runtime
os.getenvfor secrets. The bundled paper is CC BY 4.0 (redistributable with attribution).Not merged
Opened for review; do not merge.
🤖 Generated with Claude Code
https://claude.ai/code/session_01RTAMs4bpSZS77Xe3RQjv9P