⚡ Bolt: [성능 개선] MMLE-EM M-step 벡터화 - #185
Conversation
Python for loop를 사용하던 MMLE-EM의 M-step을 `active_mask`를 활용한 벡터화 연산으로 변경하여 상당한 성능 향상(~6x)을 달성했습니다.
|
👋 Jules, reporting for duty! I'm here to lend a hand with this pull request. When you start a review, I'll add a 👀 emoji to each comment to let you know I've read it. I'll focus on feedback directed at me and will do my best to stay out of conversations between you and other bots or reviewers to keep the noise down. I'll push a commit with your requested changes shortly after. Please note there might be a delay between these steps, but rest assured I'm on the job! For more direct control, you can switch me to Reactive Mode. When this mode is on, I will only act on comments where you specifically mention me with New to Jules? Learn more at jules.google/docs. For security, I will only act on instructions from the user who triggered this task. |
There was a problem hiding this comment.
Pull request overview
OpenCode cannot approve yet because required coverage evidence did not pass.
Review outcome
1. HIGH .github/workflows/opencode-review.yml:1 - Coverage evidence did not prove required test/docstring evidence
-
Problem: The required coverage-evidence job result was
failure, so OpenCode cannot establish approval sufficiency for this head. -
Root cause: Automated approval is only valid when the same-head coverage-evidence job proves supported repository test suites passed and configured docstring gates passed or were advisory, or reports not applicable because no supported source files or package manifests exist. Missing, failed, skipped, unavailable, or unsupported-tooling test evidence is a blocker.
-
Fix: Install or configure the repository test/docstring evidence tooling when source files or package manifests exist, rerun the current-head coverage-evidence job, and approve only after it reports
successwith required evidence or explicit no-source not-applicable evidence. -
Regression test: Keep the approval branch checking
needs.coverage-evidence.result == successbefore posting APPROVE, and publish REQUEST_CHANGES when coverage-evidence blocker states such as cancelled, skipped, failed, unsupported-tooling, or below-100 evidence are present. -
Result: REQUEST_CHANGES
-
Reason: coverage-evidence result was
failure, so required test/docstring evidence was not proven for current head9642ea624c655848899a2c94427819586c03cf10. -
Head SHA:
9642ea624c655848899a2c94427819586c03cf10 -
Workflow run: 29660926874
-
Workflow attempt: 1
Coverage evidence
Coverage Decision
- Result: FAIL
- Test evidence: not proven passing
- Docstring evidence: not proven passing when configured
- Failure count: 1
Changed-File Evidence Map
flowchart LR
PR["PR changed files"] --> Evidence["OpenCode bounded evidence"]
Evidence --> S1["Changed file (2 files)"]
S1 --> I1["repository behavior"]
I1 --> R1["Review risk: Changed file (2 files)"]
R1 --> V1["required checks"]
OpenCode Review Overview
Pull request overviewOpenCode cannot approve yet because required coverage evidence did not pass. Review outcome1. HIGH .github/workflows/opencode-review.yml:1 - Coverage evidence did not prove required test/docstring evidence
Coverage evidenceCoverage Decision
Changed-File Evidence Mapflowchart LR
PR["PR changed files"] --> Evidence["OpenCode bounded evidence"]
Evidence --> S1["Changed file (2 files)"]
S1 --> I1["repository behavior"]
I1 --> R1["Review risk: Changed file (2 files)"]
R1 --> V1["required checks"]
|
|
중복 정리: 배경: 2026-07-14 이후 조직 coverage-evidence 인프라 문제로 모든 PR이 REQUEST_CHANGES 상태였습니다(인프라 수정: ContextualWisdomLab/.github#611). 필요 시 재오픈 가능합니다. Generated by Claude Code |
Understood. Acknowledging that this PR is a duplicate of #162 and will be closed as obsolete. |
💡 What (무엇을 변경했는가)
python/fast_mlsirm/estimators/mmle.py의fit_mmle_2pl함수 내 M-step에서 아이템 차원(n_items)을 순회하던 Pythonfor루프를 제거하고, 전체 아이템에 대한 Newton-Raphson 단계를active_mask와 행렬 곱셈을 사용하여 완전히 벡터화했습니다.🎯 Why (왜 변경했는가)
큰 차원(아이템 수)을 순회하는 Python의 내장
for루프는 성능 병목의 주요 원인입니다. 특히 반복적인 Newton-Raphson 업데이트가 진행되는 경우 연산 지연이 심하게 발생합니다.📊 Impact (어떤 효과가 있는가)
수행된 벤치마크 테스트 결과(1000 persons x 1000 items), 기존 구현 대비 약 6배(6.11x)의 성능 향상을 달성했습니다. 또한 루프 내부의 상태를 추적하기 위해
active_mask를 사용하여 이미 수렴했거나 행렬식이 유효하지 않은 항목의 추가 연산을 건너뛰도록 하여 최적의 리소스 활용을 유지했습니다.🔬 Measurement (어떻게 확인할 수 있는가)
python -m pytest tests/test_estimator_mmle.py를 통해 정확성에 문제가 없음을 확인했습니다.PR created automatically by Jules for task 11535402146991342864 started by @seonghobae