OSAC-1273: acquire row lock in updatePoolCapacity to prevent counter drift - #647
Conversation
…drift
updatePoolCapacity reads the pool record then writes back the entire proto
with updated allocated/available counters. Without a row lock on the read,
a concurrent pool reconciler update can interleave and overwrite the counter
change, leaving the pool with a stale allocated count. This makes the pool
undeletable ("N public IP(s) are still allocated") even when no IPs exist.
Adding SetLock(true) to the DAO Get issues a SELECT FOR UPDATE, serializing
the read-modify-write cycle against any concurrent pool writers.
Signed-off-by: akshaynadkarni <25892229+akshaynadkarni@users.noreply.github.com>
Assisted-by: Cursor/Claude
Signed-off-by: akshaynadkarni <25892229+akshaynadkarni@users.noreply.github.com>
|
Skipping CI for Draft Pull Request. |
|
@akshaynadkarni: This pull request references OSAC-1273 which is a valid jira issue. Warning: The referenced jira issue has an invalid target version for the target branch this PR targets: expected the sub-task to target the "5.0.0" version, but no target version was set. DetailsIn response to this:
Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the openshift-eng/jira-lifecycle-plugin repository. |
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Repository: osac-project/coderabbit/.coderabbit.yaml Review profile: ASSERTIVE Plan: Enterprise Run ID: 📒 Files selected for processing (1)
WalkthroughThis PR adds pessimistic locking to the pool capacity update operation in ChangesPool Capacity Update Locking
Estimated code review effort🎯 2 (Simple) | ⏱️ ~10 minutes Risk AssessmentConcurrency Risk Mitigated: This change addresses a race condition in the pool capacity update path. Without locking, concurrent calls to Lock Implementation Risk: The effectiveness of this fix depends on:
Validation Required: Reviewers should verify that Possibly related PRs
Suggested labels
Suggested reviewers
🚥 Pre-merge checks | ✅ 11✅ Passed checks (11 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
@coderabbitai review |
✅ Action performedReview finished.
|
|
[APPROVALNOTIFIER] This PR is APPROVED This pull-request has been approved by: akshaynadkarni, jhernand The full list of commands accepted by this bot can be found here. The pull request process is described here DetailsNeeds approval from an approver in each of these files:
Approvers can indicate their approval by writing |
Summary
OSAC-1273: Fix a race condition where the
PublicIPPool.status.allocatedcounter becomes stale after deleting a PublicIP, making the pool undeletable.Why
updatePoolCapacity()reads the pool record without a row lock, then writes the entire proto back with updated counters. A concurrent pool reconciler update can interleave and overwrite the counter change, leaving the pool with a staleallocatedcount. The pool delete precondition then fails with "N public IP(s) are still allocated" even when no IPs exist.Testing
Deployed to edge22 and verified the full lifecycle:
Unit tests: 853/853 pass (
ginkgo run ./internal/servers/).Ticket
OSAC-1273
Signed-off-by: akshaynadkarni 25892229+akshaynadkarni@users.noreply.github.com
Assisted-by: Cursor/Claude
Summary by CodeRabbit