Skip to content

OSAC-2135: Design — CaaS Bare-Metal Worker Node Provisioning - #198

Merged
openshift-merge-bot[bot] merged 8 commits into
osac-project:mainfrom
rccrdpccl:design/OSAC-2135
Aug 17, 2026
Merged

openshift-merge-bot[bot] merged 8 commits into
osac-project:mainfrom
rccrdpccl:design/OSAC-2135

Conversation

@rccrdpccl

@rccrdpccl rccrdpccl commented Aug 7, 2026 •

Copy link
Copy Markdown
Contributor

Design: CaaS Bare-Metal Worker Node Provisioning

Jira: OSAC-2135
PRD: prd.md

Summary

Replaces the static BareMetalPool-based agent pre-boot pool with on-demand bare-metal worker provisioning. The ClusterOrder controller creates individual BareMetalInstances via the fulfillment-service private API, each booting a pre-registered RHCOS DiskImage with cluster-specific discovery ignition. Agents register, are correlated to BMIs via MAC address, and join the HyperShift-managed cluster as worker nodes. CaaS-managed BMIs are hidden from tenant APIs via label-based visibility filtering.

Requesting Review On

  • DiskImage integration (OSAC-2540/OSAC-1270): The design proposes a CaaS-specific label (osac.openshift.io/ocp-version) on DiskImage resources for version-based lookup. This label is not part of the DiskImage PRD — alignment needed.
  • MAC address dependency (OSAC-2308/OSAC-3254): Exact field path for MAC in BMI status is TBD. Agent-to-BMI correlation is blocked until this ships.
  • Visibility filtering mechanism: Public API excludes CaaS-managed BMIs via managed-by: caas label. Requires fulfillment-service changes (baremetal_instances_server.go).
  • BareMetalPool removal: The cluster_infra AAP step and the existing pool-based flow are removed entirely — no coexistence period. Upgrade procedure includes AAP job queue drain.
  • New source_type value (disk_image): Proto schema change on BareMetalInstanceImage — version skew implications documented.

Documents

  • design.md — technical design document

How to Review

  • Comment inline on specific sections
  • Approve when the design accurately reflects a viable implementation approach

Summary by CodeRabbit

  • New Features

    • Added on-demand bare-metal worker provisioning for CaaS clusters.
    • Added worker status and lifecycle tracking in cluster order details.
    • Added support for worker scaling, upgrades, recovery, and cleanup.
    • Added visibility into provisioning progress, failures, retries, and readiness.
  • Documentation

    • Documented provisioning workflows, failure handling, monitoring, security, operational procedures, and testing coverage.

@openshift-ci-robot

openshift-ci-robot commented Aug 7, 2026 •

Copy link
Copy Markdown

@rccrdpccl: This pull request references OSAC-2135 which is a valid jira issue.

Warning: The referenced jira issue has an invalid target version for the target branch this PR targets: expected the feature to target the "5.0.0" version, but no target version was set.

Details

In response to this:

Design: CaaS Bare-Metal Worker Node Provisioning

Jira: OSAC-2135
PRD: prd.md

Summary

Replaces the static BareMetalPool-based agent pre-boot pool with on-demand bare-metal worker provisioning. The ClusterOrder controller creates individual BareMetalInstances via the fulfillment-service private API, each booting a pre-registered RHCOS DiskImage with cluster-specific discovery ignition. Agents register, are correlated to BMIs via MAC address, and join the HyperShift-managed cluster as worker nodes. CaaS-managed BMIs are hidden from tenant APIs via label-based visibility filtering.

Requesting Review On

  • DiskImage integration (OSAC-2540/OSAC-1270): The design proposes a CaaS-specific label (osac.openshift.io/ocp-version) on DiskImage resources for version-based lookup. This label is not part of the DiskImage PRD — alignment needed.
  • MAC address dependency (OSAC-2308/OSAC-3254): Exact field path for MAC in BMI status is TBD. Agent-to-BMI correlation is blocked until this ships.
  • Visibility filtering mechanism: Public API excludes CaaS-managed BMIs via managed-by: caas label. Requires fulfillment-service changes (baremetal_instances_server.go).
  • BareMetalPool removal: The cluster_infra AAP step and the existing pool-based flow are removed entirely — no coexistence period. Upgrade procedure includes AAP job queue drain.
  • New source_type value (disk_image): Proto schema change on BareMetalInstanceImage — version skew implications documented.

Documents

  • design.md — technical design document

How to Review

  • Comment inline on specific sections
  • Approve when the design accurately reflects a viable implementation approach

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the openshift-eng/jira-lifecycle-plugin repository.

@openshift-ci
openshift-ci Bot requested review from eranco74 and mennyaboush August 7, 2026 13:32
@openshift-ci openshift-ci Bot added the approved label Aug 7, 2026
@coderabbitai

coderabbitai Bot commented Aug 7, 2026 •

Copy link
Copy Markdown

Review Change Stack

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review

Walkthrough

The design specifies on-demand CaaS bare-metal worker provisioning through the fulfillment-service private API. It defines resource creation, image selection, Agent correlation, NodePool binding, scaling, cleanup, tenancy, status, recovery, and validation.

Changes

CaaS bare-metal worker provisioning

Layer / File(s) Summary
Resource and image contracts
enhancements/OSAC-2135-caas-bare-metal-worker-provisioning/design.md
Defines worker status fields, shared InfraEnv use, BareMetalInstance visibility, and release- and architecture-matched disk_image provisioning.
Worker lifecycle reconciliation
enhancements/OSAC-2135-caas-bare-metal-worker-provisioning/design.md
Defines provisioning, scaling, deletion, Agent correlation, NodePool binding, retries, status updates, and restart-safe reconciliation.
Tenancy and operational controls
enhancements/OSAC-2135-caas-bare-metal-worker-provisioning/design.md
Specifies ignition and credential handling, private API boundaries, tenant isolation, metrics, events, dependencies, and rejected alternatives.
Validation and lifecycle recovery
enhancements/OSAC-2135-caas-bare-metal-worker-provisioning/design.md
Defines unit, integration, and E2E coverage, rollout procedures, version compatibility, diagnostics, cleanup, and recovery.
Design and implementation prerequisites
enhancements/OSAC-2135-caas-bare-metal-worker-provisioning/design.md
Records infrastructure, documentation, MCE, and fulfillment-service prerequisites for the design.

Estimated code review effort: 3 (Moderate) | ~25 minutes

Sequence Diagram(s)

sequenceDiagram
  participant ClusterOrderController
  participant InfraEnv
  participant FulfillmentServicePrivateAPI
  participant Agents
  participant HyperShiftNodePools
  ClusterOrderController->>InfraEnv: create shared discovery configuration
  ClusterOrderController->>FulfillmentServicePrivateAPI: create BareMetalInstance with disk_image
  InfraEnv->>Agents: discover worker hardware
  Agents->>ClusterOrderController: provide MAC address
  ClusterOrderController->>HyperShiftNodePools: bind matching worker Agent
Loading

Possibly related PRs

Suggested reviewers: adriengentil, alonakaplan, carbonin, avishayt

🚥 Pre-merge checks | ✅ 11
✅ Passed checks (11 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly identifies the OSAC-2135 design and the main change: CaaS bare-metal worker node provisioning.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
No-Hardcoded-Secrets ✅ Passed The PR changes only design.md; scans found no private-key material, embedded-credential URLs, long base64/hex blobs, or named secret/token/password literals, only credential references.
No-Weak-Crypto ✅ Passed The changed design and PRD contain no MD5, SHA-1, DES, 3DES, RC4, Blowfish, ECB, or custom crypto/comparison implementation; Vault is only named as the storage backend.
No-Injection-Vectors ✅ Passed The PR changes only a Markdown design document; scans found no SQL concatenation, shell=True, eval/exec, unsafe pickle/YAML loading, os.system, or dangerouslySetInnerHTML.
Container-Privileges ✅ Passed The PR changes only a Markdown design document; no container/Kubernetes manifest or listed privilege setting appears in the diff or document.
No-Sensitive-Data-In-Logs ✅ Passed The PR changes only design.md; tenant-safe messages omit BMI names, MACs, and Ironic errors, and no proposed log or event includes secrets, tokens, ignition, or raw error payloads.
Ai-Attribution ✅ Passed Both OSAC-2135 commits include Assisted-by: Claude Code <noreply@anthropic.com>; no AI Co-Authored-By trailer appears in the PR commit range.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Comment @coderabbitai help to get the list of available commands.

@rccrdpccl
rccrdpccl requested review from AlonaKaplan, adriengentil, avishayt, carbonin and vladikr and removed request for eranco74 and mennyaboush August 7, 2026 13:33
@github-actions

github-actions Bot commented Aug 7, 2026

Copy link
Copy Markdown

AI Design Review: EP-198

Score: 8/8 | Verdict: PASS

Criterion Score Notes
Feasibility 2/2 Deep technical detail throughout: proto schema snippets for BareMetalInstanceSpec and BareMetalInstanceImage, Go struct for ClusterOrderStatus extension, 4 idempotent controller phases (ensureInfraEnv, reconcileWorkers, correlateAgents, reconcileNodePoolReplicas), all lifecycle operations covered (provisioning, scale-up, scale-down, cluster deletion) with step-by-step flows and sequence diagrams. Failure handling table enumerates 9 specific failure modes with recovery strategies. Risks are concr
Testability 2/2 Test plan specifies concrete scenarios at all three levels: 10 unit tests (reconcileWorkers count logic, correlateAgents MAC matching, ensureInfraEnv owner references, phase transitions, idempotency, visibility filtering), 4 integration tests with kind cluster and mocked private API, 6 E2E scenarios (full provisioning, scale-up, scale-down, deletion, visibility filtering, pool removal verification). Honestly notes E2E dependency on BMaaS-capable test environment. Graduation criteria N/A is justi
Scope 2/2 Summary is concise (3 sentences covering what/why/capabilities). Non-goals are specific with deferral rationale and Jira links (autoscaling, VM workers, static IP, network boot). PRD referenced via frontmatter. Three real alternatives analyzed with clear rejection rationales (BareMetalPool, direct CR creation, AAP orchestration). Cross-cutting dimensions addressed: provisioning (core), installation (upgrade/downgrade strategy), E2E testing, documentation, networking (network attachments pass-thr
Architecture 2/2 Follows OSAC controller patterns: finalizer-based lifecycle, idempotent reconciliation, conditions for state (WorkersFailed, InfraEnvReady, RHCOSImageNotFound). Tenant isolation maintained via osac.openshift.io/managed-by label with visibility filtering at the fulfillment-service public API layer. Dependencies clearly enumerated with Jira links and blocking-impact analysis. Integration described across osac-operator, fulfillment-service (private API), assisted-service (InfraEnv/Agent), HyperShif

Verdict: Exceptionally thorough design document that follows all OSAC architectural patterns, provides deep implementation detail with proto schemas and controller phase decomposition, covers all lifecycle operations with specific failure handling, and includes a concrete multi-level test plan — one of the strongest OSAC designs reviewed.

Feedback: Fix the RBAC section contradiction: the implementation details require patch permission on agents in the agent-install.openshift.io API group, but the RBAC section claims no new permissions are needed — enumerate the specific ClusterRole changes required. Reframe implementation-focused goals (e.g., 'Reuse the existing ClusterOrder controller reconciliation pattern') as user-visible outcomes. Consider adding more detail on how in-memory worker phase state is rebuilt from live CR and Agent state on controller restart — the current description says it happens but not how ambiguous states (e.g., a BMI in Running but no matching Agent yet) are resolved.

Critical (0)

None.

Important (2)

  1. RBAC section contradiction: Implementation details state 'The osac-operator's RBAC must include patch on agents in the agent-install.openshift.io API group' (line 127), but the RBAC/Tenancy section claims 'No new RBAC roles are introduced. The osac-operator's service account already has permissions to create CRs in cluster namespaces and call the private API' (line 389). The new Agent patch permission is a real RBAC change that should be explicitly enumerated with the ClusterRole diff.
  2. Some goals are implementation-focused rather than user-visible: 'Reuse the existing ClusterOrder controller reconciliation pattern and the private gRPC API for BMI lifecycle management' (line 41) describes an implementation strategy, not a user outcome. Goals should describe what users or operators observe.

Suggestions (3)

  1. The in-memory worker phase state rebuild on controller restart (line 217: 'rebuilt from live CR and Agent state on restart') would benefit from more specificity: how does the controller resolve ambiguous states like a BMI in Running with no matching Agent (could be WaitingForAgent or a missed timeout)?
  2. Consider documenting the expected InfraEnv ignition generation latency and whether the 30s requeue interval (line 123 area) is sufficient, given that InfraEnv creation in assisted-service can take longer in degraded environments.
  3. The visibility filtering implementation mentions CEL filter clause (line 327) but the fulfillment-service uses SQL-backed generic servers — clarify whether the filter is applied at the CEL layer, the SQL query layer, or both, to avoid a gap where one layer filters but the other doesn't.

Review cost

Model: claude-opus-4-6
Cost: $0.9285
Tokens: 10 in / 6.0k out
Cache: 202.2k read
Active time: 2m 17s
API calls: 0

@github-actions github-actions Bot added the rfe-creator-auto-reviewed EP was reviewed by AI label Aug 7, 2026

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 13

🧹 Nitpick comments (1)
enhancements/OSAC-2135-caas-bare-metal-worker-provisioning/design.md (1)

444-475: 🗄️ Data Integrity & Integration | 🔵 Trivial | ⚡ Quick win

Add failure-window and contract tests.

The test plan does not cover a lost Create response, a BMI stuck in Deleting, ambiguous or cross-cluster MAC matches, zero or multiple DiskImages, reserved-label writes, URL limits, or old-service capability gating. Add these cases before rollout.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@enhancements/OSAC-2135-caas-bare-metal-worker-provisioning/design.md` around
lines 444 - 475, Add failure-window and contract coverage to the Test Plan for
lost Create responses, BMIs stuck in Deleting, ambiguous or cross-cluster MAC
matches, zero and multiple DiskImages, reserved-label writes, URL-length limits,
and old-service capability gating. Place the cases under the appropriate Unit,
Integration, or E2E sections and require them before rollout.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@enhancements/OSAC-2135-caas-bare-metal-worker-provisioning/design.md`:
- Line 428: Update the MAC-based Agent correlation algorithm to require a unique
candidate matching the current worker’s Agent namespace, InfraEnv or
ClusterDeployment, and osac.openshift.io/cluster-order ownership. Reject zero or
multiple matches and perform no Delete call unless exactly one candidate
satisfies all scope constraints; add tests covering these failure cases.
- Around line 126-134: The scale-up capacity calculation must not count workers
in `Failed` phase toward `current`; update the controller’s
desired-versus-current worker logic so failed slots are replaced and the desired
`Ready` capacity is restored, while preserving the existing treatment of
`Deleting` workers.
- Around line 319-327: Update the public BareMetalInstances Create and Update
validation to reject any request containing the reserved
osac.openshift.io/managed-by label key, regardless of its value. Keep visibility
filtering in BareMetalInstances.List consistent with this reserved-key contract
so tenant-owned resources cannot be hidden by assigning another value.
- Around line 352-362: The security design must not leave discovery ignition
pull secrets exposed in immutable, plaintext user_data. Update the user_data
handling to use an approved secret reference or envelope-encryption mechanism,
restrict and redact reads and logs, and expire or remove the credential material
after provisioning where supported; revise the surrounding Security
Considerations to describe the resulting protections.
- Around line 114-117: Update the discovery ignition fetch in the InfraEnv
polling flow after status.bootArtifacts.discoveryIgnitionURL is populated:
validate that the URL uses HTTPS and an allowed external destination, configure
connection, read, and total timeouts, disable unsafe redirects, and stream-read
with a strict 64 KiB maximum. Reject invalid schemes, destinations, redirects,
timeouts, or oversized responses, and do not log the fetched ignition content.
- Around line 163-171: Update the worker deletion flow around
BareMetalInstances.Delete and status.workers so the worker record remains in a
Deleting state after the delete request, allowing reconciliation retries while
deletion is pending or fails. Remove the record only after the private API
confirms terminal deletion and BMaaS cleanup; apply the same retention and retry
behavior to the break-glass procedure.
- Around line 329-333: Before implementing the MAC correlation workflow, define
the canonical BMI status MAC field and its type, normalization rules, readiness
semantics, and API version gate, replacing the TBD status.host.mac_address
reference in the design. Gate controller matching on that contract and add
producer-consumer coverage alongside the controller matching test.
- Around line 199-211: Replace the proposed ClusterOrderStatus.Workers
ObjectReference-only design with a durable, authoritative worker record
containing the fulfillment-service BMI ID and reconciliation/lifecycle state, or
explicitly define the authoritative BMI CR, ID mapping, ownership, and watch
contract. Align the workers[] description and schema so they do not conflict,
and ensure reconciliation can rebuild identical worker mappings and state after
restart from persisted status and authoritative resources rather than in-memory
state.
- Around line 499-503: The Version Skew Strategy must require
fulfillment-service capability validation before the controller creates BMIs
using source_type "disk_image". Replace the cosmetic upgrade-window guidance
with a fail-closed requirement: upgrade fulfillment-service first or
simultaneously, and prevent BMI creation and tenant operations when the required
capability or visibility filtering is unavailable.
- Around line 277-289: Update the RHCOS DiskImage resolution design to require
exactly one matching provider-global image: filter for AVAILABLE lifecycle,
empty tenant metadata, Linux guest OS family, amd64 architecture, and the target
OCP version label; fail when zero or multiple images match. Define and test the
NodePool.spec.release.image parser against every supported release-image format,
including z-stream variants.
- Around line 291-317: Update the BMI creation flow around
BareMetalInstances.Create to be idempotent when responses or status updates are
lost: use a deterministic idempotency key or enforce the generated worker name
as unique, and reconcile AlreadyExists by retrieving and reusing the existing
BMI. Define the gRPC deadline and bounded retry behavior, then add a test
covering a successful Create followed by a lost response/status update and
reconciliation of the existing BMI.
- Around line 381-385: Update the RBAC / Tenancy design section to define the
exact Role or ClusterRole and RoleBinding for the target namespace, including
patch access to agents in the agent-install.openshift.io API group and the
permissions needed to watch Agent resources. Add an RBAC test covering these
permissions and bindings.
- Around line 335-337: Update the “Minimum MCE Version” section to replace the
inconsistent post-v2.55.0 and assisted-service 5.0.0+ prerequisites with one
exact released MCE bundle version known to contain both MGMT-24903 fixes. Add an
installation compatibility check that validates the deployed MCE version meets
this minimum before proceeding.

---

Nitpick comments:
In `@enhancements/OSAC-2135-caas-bare-metal-worker-provisioning/design.md`:
- Around line 444-475: Add failure-window and contract coverage to the Test Plan
for lost Create responses, BMIs stuck in Deleting, ambiguous or cross-cluster
MAC matches, zero and multiple DiskImages, reserved-label writes, URL-length
limits, and old-service capability gating. Place the cases under the appropriate
Unit, Integration, or E2E sections and require them before rollout.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository: osac-project/coderabbit/.coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: ae07ffdd-d4a2-4357-bd1b-198661d88b86

📥 Commits

Reviewing files that changed from the base of the PR and between 6e2722a and 385d2b4.

📒 Files selected for processing (1)
  • enhancements/OSAC-2135-caas-bare-metal-worker-provisioning/design.md

Comment thread enhancements/OSAC-2135-caas-bare-metal-worker-provisioning/design.md Outdated
Comment thread enhancements/OSAC-2135-caas-bare-metal-worker-provisioning/design.md Outdated
Comment thread enhancements/OSAC-2135-caas-bare-metal-worker-provisioning/design.md Outdated
Comment thread enhancements/OSAC-2135-caas-bare-metal-worker-provisioning/design.md Outdated
Comment on lines +335 to +337
#### Minimum MCE Version

The MGMT-24903 fix (persistent-boot day-2 installs) is merged to assisted-service master ([PR #10717](https://github.com/openshift/assisted-service/pull/10717), 2026-07-29) and assisted-installer-agent master ([PR #1568](https://github.com/openshift/assisted-installer-agent/pull/1568), 2026-07-30). The fix is not yet in a tagged release (post-v2.55.0). The design requires a MCE version shipping these commits. Without them, workers fail to install because `osImageURL` is stripped from the ignition config. The controller does not implement a workaround — the deployment prerequisite documentation must specify the minimum MCE version.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🩺 Stability & Availability | 🟠 Major | 🏗️ Heavy lift

🧩 Analysis chain

🏁 Script executed:

sed -n '320,345p;510,535p' enhancements/OSAC-2135-caas-bare-metal-worker-provisioning/design.md
printf '\n--- related MCE/version references ---\n'
rg -n -i 'MCE|assisted-service|assisted-installer-agent|5\.0\.0|v2\.55\.0|minimum' enhancements/OSAC-2135-caas-bare-metal-worker-provisioning

Repository: osac-project/enhancement-proposals

Length of output: 10000


🌐 Web query:

OpenShift Multi-Cluster Engine release containing assisted-service PR 10717 and assisted-installer-agent PR 1568, and the corresponding MCE bundle/version

💡 Result:

Based on the provided search results, there is no evidence of a specific OpenShift Multi-Cluster Engine (MCE) release that explicitly bundles or corresponds to an assisted-service pull request (PR) 10717 or an assisted-installer-agent PR 1568. The current PR numbering in the openshift/assisted-service and openshift/assisted-installer-agent repositories does not indicate that these specific PR numbers have been merged or associated with a released MCE version in the context of the provided data [1][2][3][4]. The assisted-installer-agent and assisted-service repositories utilize continuous integration and development workflows where PRs are merged independently [5][6][7][8][9][10][11][12][13][14]. MCE bundles are managed through standard Red Hat release processes, and while the assisted-installer-agent is a component of the MCE operator (often referred to as Central Infrastructure Management or CIM) [15][16], specific PR numbers are rarely used to identify MCE bundle versions in public documentation [15][17]. The provided information does not link the requested PR numbers to any MCE bundle or version.

Citations:


🏁 Script executed:

set -eu
for spec in \
  'openshift/assisted-service/pulls/10717' \
  'openshift/assisted-installer-agent/pulls/1568'
do
  echo "--- $spec ---"
  curl -fsSL "https://api.github.com/repos/$spec" |
    jq '{number,state,title,merged_at,merge_commit_sha,head:.head.sha,base:.base.ref,html_url}'
done

for repo in openshift/assisted-service openshift/assisted-installer-agent; do
  echo "--- tags: $repo ---"
  curl -fsSL -H 'Accept: application/vnd.github+json' \
    "https://api.github.com/repos/$repo/tags?per_page=100" |
    jq -r '.[] | [.name,.commit.sha] | `@tsv`' | head -40
done

echo "--- MCE repositories matching repository metadata ---"
curl -fsSL 'https://api.github.com/search/repositories?q=multicluster-engine+org:redhat-open-cluster-management&per_page=20' |
  jq -r '.items[] | [.full_name,.html_url,.archived] | `@tsv`'

Repository: osac-project/enhancement-proposals

Length of output: 662


🌐 Web query:

site:github.com/openshift/assisted-service/pull/10717 MGMT-24903

💡 Result:

I couldn’t retrieve the contents of PR #10717 or confirm its association with MGMT-24903 from indexed results. The pull request is here:

https://github.com/openshift/assisted-service/pull/10717

GitHub’s PR API documents the endpoint format for retrieving PR details. [1] (docs.github.com)

Citations:


🌐 Web query:

site:github.com/openshift/assisted-installer-agent/pull/1568 MGMT-24903

💡 Result:

PR #1568, “MGMT-24903: Preserve encapsulated MachineConfig in filtered ignition,” was merged on July 30, 2026. It preserves /etc/ignition-machine-config-encapsulated.json so the service can obtain spec.osImageURL for persistent-boot day-2 installations when the release image is unavailable. It added unit and regression tests; all 17 checks passed. [1] (github.com)

Citations:


🌐 Web query:

OpenShift MCE release notes assisted-service 5.0.0 MGMT-24903 osImageURL

💡 Result:

The issue identifier MGMT-24903 in the OpenShift assisted-service (related to OpenShift MCE/Agent-based installation) is associated with two specific technical improvements implemented in recent releases: 1. Preserve encapsulated MachineConfig in filtered ignition: This fix ensures that MachineConfig data is preserved during ignition filtering [1][2][3][4][5]. 2. Fall back to CoreOS image from worker ignition for day-2 persistent-boot: This improvement addresses persistent boot issues during day-2 operations by enabling a fallback mechanism to the CoreOS image provided in the worker ignition [3][6][7]. The query for osImageURL appears in related OpenShift technical contexts (such as OCPNODE-4518, which involves blocking runc on RHEL 10 via OSImageURL stream class inspection), but it is not the primary subject of MGMT-24903 [8][7][9][10]. The MGMT-24903 identifier specifically tracks the ignition and boot-related fixes mentioned above in the assisted-installer and agent-installer components [1][3][7].

Citations:


🏁 Script executed:

set -u
for repo in \
  https://github.com/openshift/assisted-service.git \
  https://github.com/openshift/assisted-installer-agent.git
do
  echo "--- $repo PR refs ---"
  git ls-remote "$repo" \
    'refs/pull/10717/*' 'refs/pull/1568/*' 2>&1 || true
  echo "--- $repo version tags ---"
  git ls-remote --tags "$repo" 'refs/tags/v2.55*' 'refs/tags/v5.0*' 2>&1 |
    tail -30
done

Repository: osac-project/enhancement-proposals

Length of output: 888


Pin the minimum released MCE bundle.

The document uses inconsistent prerequisites (post-v2.55.0 and assisted-service 5.0.0+) and does not identify an MCE bundle containing both MGMT-24903 fixes. Specify the exact supported MCE version and add an installation compatibility check.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@enhancements/OSAC-2135-caas-bare-metal-worker-provisioning/design.md` around
lines 335 - 337, Update the “Minimum MCE Version” section to replace the
inconsistent post-v2.55.0 and assisted-service 5.0.0+ prerequisites with one
exact released MCE bundle version known to contain both MGMT-24903 fixes. Add an
installation compatibility check that validates the deployed MCE version meets
this minimum before proceeding.

Comment thread enhancements/OSAC-2135-caas-bare-metal-worker-provisioning/design.md Outdated
Comment thread enhancements/OSAC-2135-caas-bare-metal-worker-provisioning/design.md Outdated
Comment thread enhancements/OSAC-2135-caas-bare-metal-worker-provisioning/design.md Outdated
Comment thread enhancements/OSAC-2135-caas-bare-metal-worker-provisioning/design.md Outdated

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@enhancements/OSAC-2135-caas-bare-metal-worker-provisioning/design.md`:
- Line 130: Align the failed-worker behavior described in the controller
capacity calculation with the terminal failure policy: either exclude automatic
replacement for workers in Failed phase, or define a persisted retry-attempt
limit with bounded backoff and repeated-failure tests. Update the lifecycle and
failure table consistently, including the cleanup and status.workers behavior.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository: osac-project/coderabbit/.coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: 9581afa6-51c7-423f-b207-372ee82f5fd0

📥 Commits

Reviewing files that changed from the base of the PR and between 385d2b4 and 8363973.

📒 Files selected for processing (1)
  • enhancements/OSAC-2135-caas-bare-metal-worker-provisioning/design.md

Comment thread enhancements/OSAC-2135-caas-bare-metal-worker-provisioning/design.md Outdated
@github-actions

github-actions Bot commented Aug 7, 2026 •

Copy link
Copy Markdown

AI Design Review: EP-198

Score: 8/8 | Verdict: PASS

Criterion Score Notes
Feasibility 2/2 Deep technical detail throughout. Go struct for ClusterOrder status extension with kubebuilder annotations, proto fields for BMI creation (only change: new source_type value 'disk_image'), YAML status examples showing mixed ready/failed states. Controller reconciliation structure names 4 idempotent phases (ensureInfraEnv, reconcileWorkers, correlateAgents, reconcileNodePoolReplicas). MAC correlation algorithm specified with 3-dimension matching (namespace, ownership label, MAC). Visibility filte
Testability 2/2 Test plan specifies concrete scenarios at all three levels. Unit tests: 10 specific cases covering worker reconciliation (create/delete counts), agent correlation (MAC match, namespace scoping), InfraEnv creation (ClusterDeployment ref, owner ref), phase transitions (full lifecycle + timeout failure), idempotency (no duplicate API calls), visibility filtering (List exclusion, Get returns NotFound). Integration tests: 4 scenarios using kind cluster with mocked private API — InfraEnv creation, Age
Scope 2/2 Summary is 3 sentences covering what/why/how with PRD reference. Non-goals are specific with deferral targets: autoscaling (future CaaS feature), VM workers (VMaaS integration), static IP (not PoC-validated), network boot acceleration (OSAC-2134). Three real alternatives with detailed rejection rationale (BareMetalPool, direct CR creation, AAP-orchestrated). PRD referenced in frontmatter and body. Relevant cross-cutting dimensions addressed: Provisioning (core topic), Installation (MCE version,
Architecture 2/2 All OSAC patterns followed: owner-reference annotation on BMIs (ClusterOrder/), tenant metadata present via private API, controller reconciliation follows established ensureX/reconcileX phase pattern. No new CRDs introduced — extends ClusterOrder status with workers[] using corev1.ObjectReference list (not maps). Dependencies explicitly gated with Jira links and impact analysis. Cross-component changes enumerated: fulfillment-service (visibility filtering), osac-operator (controller phases, RBAC

Verdict: A strong, implementation-ready design that follows all OSAC patterns, provides deep technical detail validated by a PoC, covers all lifecycle operations with comprehensive failure handling, and includes a concrete test plan at all levels — scoring 8/8.

Feedback: The design is well-structured and thorough. Two minor improvements: (1) Reword goals 1 and 3 to be user-visible outcomes rather than implementation choices — e.g., 'Workers provision on-demand when a cluster requests bare-metal nodes' instead of 'Reuse the existing ClusterOrder controller reconciliation pattern.' (2) Consider adding a brief Terminology section defining CaaS, BMI, InfraEnv, and Agent upfront for reviewers less familiar with the assisted-service stack — the networking EP's terminology section is the reference pattern.

Critical (0)

None.

Important (2)

  1. Goals 1 and 3 are implementation-focused ('Reuse the existing ClusterOrder controller reconciliation pattern', 'Support both initial provisioning and manual scale-up/scale-down through the same controller logic') rather than user-visible outcomes. Design documents should still frame goals in terms of the capability delivered, even when the document is about the HOW.
  2. The MAC address field path in BareMetalInstance status is noted as TBD (depends on OSAC-2308/OSAC-3254). While the dependency is correctly gated, specifying the expected field name (e.g., status.host.mac_address as mentioned in step 7) in the proto schema section would make the design more concrete and give the OSAC-2308 authors a clear target to implement.

Suggestions (3)

  1. Add a Terminology section defining CaaS, BMI, InfraEnv, Agent, ClusterDeployment, NodePool, and CAPI — following the pattern established by the networking EP. Several of these terms are assisted-service or HyperShift concepts that may be unfamiliar to reviewers focused on fulfillment-service or networking.
  2. corev1.ObjectReference in the workers[] field is functional but considered semi-deprecated in newer Kubernetes API patterns. A typed reference struct (e.g., WorkerReference with Kind, Name, Namespace) would be more self-documenting and avoid carrying unused ObjectReference fields (UID, ResourceVersion, FieldPath). This is cosmetic — the current approach works.
  3. The Upgrade/Downgrade Strategy section thoroughly covers the pool removal migration but could note whether existing ClusterOrder CRs (created before the upgrade) need any status backfill for the new workers[] field, or whether the controller's 'rebuild from live state' logic handles this transparently.

Review cost

Model: claude-opus-4-6
Cost: $0.9522
Tokens: 12 in / 6.8k out
Cache: 598.6k read
Active time: 2m 32s
API calls: 0

@github-actions

github-actions Bot commented Aug 7, 2026

Copy link
Copy Markdown

AI Design Review: EP-198

Score: 8/8 | Verdict: PASS

Criterion Score Notes
Feasibility 2/2 Deeply specific implementation: names data structures, defines MAC correlation algorithm with three-dimensional scoping (namespace, ownership, MAC), specifies error conditions with named conditions, describes state rebuild on restart. Proto schemas included for existing fields and new source_type value. All lifecycle operations covered with step-by-step detail (create 10 steps, scale-up 5, scale-down 9). Failure table covers 10 specific modes with recovery paths. Risks are concrete with specific
Testability 2/2 Test plan specifies 12 unit test cases (controller phases, visibility filtering, RBAC), 4 integration scenarios (kind cluster with mocked private API), and 6 E2E scenarios (full provisioning, scale-up/down, deletion, visibility, pool removal). E2E tests honestly acknowledge BMaaS environment requirement. Graduation criteria is N/A with valid explanation (OSAC unreleased).
Scope 2/2 Clear 4-sentence summary. Goals are user-visible outcomes. Non-goals are specific with deferral targets (autoscaling, VM workers, static IP, network boot with Jira links). PRD referenced in frontmatter. Three real alternatives with specific rejection rationale. All relevant cross-cutting dimensions addressed or explicitly N/A (UI hidden, provisioning core, installation MCE version, documentation updates listed).
Architecture 2/2 All OSAC patterns followed: owner references and tenant isolation on resources, controller reconciliation follows standard finalizer/status/provisioning lifecycle, spec/status ownership correct (workers is status-only), conditions used for lifecycle state (WorkersFailed, InfraEnvReady, RHCOSImageNotFound, RHCOSImageAmbiguous). Dependencies clearly enumerated with cross-component impacts. Maps avoided in CRD. Terminology consistent throughout.

Verdict: Exceptionally well-crafted design document that follows all OSAC patterns, provides deep implementation specificity across all lifecycle operations, thoroughly addresses failure modes and risks, and includes a comprehensive multi-level test plan.

Feedback: This design is ready for merge. Minor polish opportunities: (1) the worker_type metric label dimension could specify its value space (e.g., 'bare_metal', 'compute_instance') to help SREs set up dashboards; (2) consider noting whether the managed-by label filtering pattern could be generalized for future system-managed resources beyond BMIs; (3) the Support Procedures section is an excellent addition not required by the template — other designs could learn from this.

Critical (0)

None.

Important (0)

None.

Suggestions (3)

  1. The worker_type metric label on observability metrics should specify its allowed values (e.g., 'bare_metal', 'compute_instance') to help operators configure dashboards and alerts.
  2. Consider documenting whether the osac.openshift.io/managed-by label filtering pattern is intended as a reusable convention for future system-managed resources, or is specific to this CaaS use case.
  3. The Support Procedures section is a strong addition beyond template requirements — consider proposing it as a recommended section in the design template for operationally impactful features.

Review cost

Model: claude-opus-4-6
Cost: $0.8465
Tokens: 10 in / 6.0k out
Cache: 219.0k read
Active time: 2m 19s
API calls: 0

@github-actions

github-actions Bot commented Aug 7, 2026

Copy link
Copy Markdown

AI Design Review: EP-198

Score: 8/8 | Verdict: PASS

Criterion Score Notes
Feasibility 2/2 Deep technical detail throughout: proto field changes specified (source_type: disk_image), all lifecycle operations covered (provision, scale-up, scale-down, deletion), failure handling table with 9 specific failure modes and recovery strategies, retry behavior differentiated by failure type with concrete backoff values, risks are specific (MAC dependency, DiskImage dependency, ignition size limits) with concrete mitigations. Drawbacks section steel-mans the coupling argument. One concern: retry
Testability 2/2 Test plan specifies concrete scenarios at all three levels: 12 specific unit tests (reconcileWorkers creates correct BMI count, correlateAgents matches by MAC, visibility filtering excludes managed-by label), 4 integration tests with kind cluster and mocked private API, 6 E2E tests covering the full lifecycle plus visibility verification. Acknowledges E2E infrastructure dependency honestly. Graduation criteria is N/A with valid justification (OSAC not yet released).
Scope 2/2 Clear boundaries: summary is concise, goals are user-visible outcomes (tenant invisibility, reuse of existing reconciliation patterns), non-goals are specific with deferrals (autoscaling, VM workers, static IP, network boot acceleration — each with rationale or Jira link). Three real alternatives with substantive rejection rationale. PRD referenced in frontmatter and body. Relevant cross-cutting dimensions (provisioning, installation, documentation, E2E testing) addressed; UI explicitly N/A.
Architecture 2/2 All OSAC patterns followed: tenant isolation metadata present (metadata.tenant, osac.openshift.io/owner-reference annotation, managed-by label), standard controller reconciliation pattern (idempotent phases), conditions for lifecycle state (WorkersFailed, InfraEnvReady, RHCOSImageNotFound), workers field correctly in status not spec. Dependencies enumerated with Jira links and blocking impact. Cross-component changes well-described: osac-operator controller extension, fulfillment-service visibil

Verdict: Exceptionally thorough design that follows all OSAC conventions, covers all lifecycle operations in detail, specifies concrete proto changes and failure handling, and includes a well-structured test plan — the strongest area is the depth of workflow description and failure mode analysis.

Feedback: The retry attempt count reconstruction from Kubernetes events (counting WorkerFailed events per worker index on restart) is fragile: events have a default TTL and can be garbage-collected, which would reset the retry counter and potentially allow unlimited retries for a persistently failing worker slot. Consider persisting the attempt count in the ClusterOrder status (e.g., per-worker annotation or a count field in the workers list) or documenting the accepted risk with a fallback behavior. Additionally, define the values for the worker_type metric label (e.g., bare_metal, compute_instance) to avoid ambiguity for dashboard authors.

Critical (0)

None.

Important (1)

  1. Retry count state is rebuilt from Kubernetes events on controller restart (line 277: 'rebuilt on restart from the ClusterOrder's Kubernetes events counting WorkerFailed events per worker index'). Kubernetes events are best-effort with a configurable TTL (default 1 hour). If the controller is down longer than the event TTL, retry counts are lost, potentially allowing unlimited retries for persistently failing worker slots. Consider persisting attempt counts in ClusterOrder status or an annotation

Suggestions (3)

  1. The worker_type metric label (lines 429-435) is used on all worker metrics but its allowed values are not defined. Clarify whether it maps to resource kind (BareMetalInstance vs ComputeInstance), resource class name, or another dimension — this affects dashboard design and cardinality.
  2. The visibility filtering change to the public BareMetalInstances.List API (lines 352-358) adds implicit behavior to a shared server implementation. Consider noting whether this requires documentation or a changelog entry for other fulfillment-service contributors who may not expect system-reserved label semantics.
  3. The design could mention whether osac-installer Helm charts need any configuration changes (e.g., RBAC for the new agent-install.openshift.io API group patch permission, MCE version constraint) or if these are purely deployment-time admin concerns outside the chart.

Review cost

Model: claude-opus-4-6
Cost: $0.7262
Tokens: 7 in / 4.5k out
Cache: 258.6k read
Active time: 1m 56s
API calls: 0

Comment thread enhancements/OSAC-2135-caas-bare-metal-worker-provisioning/design.md Outdated
Comment thread enhancements/OSAC-2135-caas-bare-metal-worker-provisioning/design.md Outdated
Comment thread enhancements/OSAC-2135-caas-bare-metal-worker-provisioning/design.md Outdated
Comment thread enhancements/OSAC-2135-caas-bare-metal-worker-provisioning/design.md Outdated
Comment thread enhancements/OSAC-2135-caas-bare-metal-worker-provisioning/design.md Outdated
Comment thread enhancements/OSAC-2135-caas-bare-metal-worker-provisioning/design.md Outdated
Assisted-by: Claude Code <noreply@anthropic.com>
Signed-off-by: Riccardo Piccoli <rpiccoli@redhat.com>
@github-actions

Copy link
Copy Markdown

AI Design Review: EP-198

Score: 8/8 | Verdict: PASS

Criterion Score Notes
Feasibility 2/2 Deep technical detail throughout. Implementation names specific API calls (BareMetalInstances.Create, Secrets.Create), specific fields (status.bootArtifacts.discoveryIgnitionURL), specific conditions and error codes. Go struct for ClusterOrder status extensions is complete with kubebuilder annotations and a worked YAML example. Proto usage references existing fields (no new CRDs). All lifecycle operations covered: create, scale-up, scale-down, cluster deletion — each with step-by-step flow and s
Testability 2/2 Test plan covers all three levels with concrete scenarios. Unit tests: 11 specific cases covering controller logic (reconcileWorkers, correlateAgents, ensureInfraEnv), phase transitions, idempotency, tenant isolation, and RBAC. Integration tests: 4 scenarios using kind cluster with shared InfraEnv and mocked private API, covering BMI creation, agent correlation, scale-down, and deletion cleanup. E2E tests: 6 scenarios covering full provisioning, scale-up, scale-down, cluster deletion, tenant iso
Scope 2/2 Summary is 3 concise sentences. PRD referenced in frontmatter and inline link. Non-goals are specific: autoscaling deferred, VM workers deferred to VMaaS, static IP deferred (not validated by PoC), network boot acceleration tracked as OSAC-2134. Three real alternatives rejected with specific rationale (BareMetalPool, direct CR creation, AAP orchestration). Cross-cutting dimensions covered: Provisioning (core), Networking (BMI network attachments, subnet mapping), Installation (upgrade/downgrade)
Architecture 2/2 All OSAC patterns followed. New BMIs carry owner-reference annotation and system tenant isolation. ClusterOrder status extension uses a list of named WorkerStatus subobjects (not maps). Controller reconciliation follows the standard finalizer/status/lifecycle pattern with four idempotent phases. Conditions (WorkersFailed, InfraEnvReady, RHCOSImageNotFound) used for aggregate lifecycle state; per-worker Phase enum is justified for linear single-worker lifecycle. Dependencies clearly enumerated wi

Verdict: Exceptionally thorough design document that follows all OSAC architectural patterns, provides deep implementation detail with specific error codes and backoff values, clearly scopes the work with specific non-goals and three rejected alternatives, and includes a concrete multi-level test plan.

Feedback: The design is ready for merge. Two minor improvements: (1) Reframe Goal 1 as a user-visible outcome rather than an implementation choice — e.g., 'Provision and manage bare-metal workers through the existing fulfillment-service BMI lifecycle' instead of 'Reuse the existing ClusterOrder controller reconciliation pattern.' (2) Consider adding a brief note to the test plan about success criteria for the E2E tests (e.g., 'all lifecycle operations complete within X minutes, no orphaned BMIs after deletion') to give reviewers confidence in what 'passing' looks like at the E2E level.

Critical (0)

None.

Important (2)

  1. Goal 1 ('Reuse the existing ClusterOrder controller reconciliation pattern and the private gRPC API for BMI lifecycle management') is an implementation task, not a user-visible outcome. Reframe as the user/operator benefit it delivers.
  2. MAC address field path in BMI status is TBD (depends on OSAC-2308/OSAC-3254). While honestly declared as a blocking dependency, the correlation algorithm's correctness cannot be fully verified until the field shape is known — consider documenting the minimum contract (e.g., 'a single MAC string per BMI') so implementation can proceed once the dependency lands.

Suggestions (3)

  1. The E2E test plan could include approximate timing expectations (e.g., 'provisioning completes within 10 minutes based on PoC measurements of ~6 min BMI provisioning') to help set CI timeout values and catch performance regressions.
  2. Consider noting what happens during a rolling cluster upgrade where NodePools at different OCP versions coexist — the DiskImage resolution uses the NodePool's release image, but the scale-up flow should clarify whether workers added to a partially-upgraded cluster use the old or new version's boot image.
  3. The Graduation Criteria 'N/A' is valid but could include a forward-looking note about what production-readiness signals would look like (e.g., 'sustained zero orphaned BMIs across 100+ cluster lifecycles') to guide future E2E coverage expansion.

Review cost

Model: claude-opus-4-6
Cost: $0.9455
Tokens: 10 in / 5.7k out
Cache: 209.3k read
Active time: 2m 22s
API calls: 0

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 8

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@enhancements/OSAC-2135-caas-bare-metal-worker-provisioning/design.md`:
- Line 125: Update the RBAC section associated with the Agent correlation and
watch flow to define the exact Role or ClusterRole and corresponding RoleBinding
for agent-install.openshift.io Agents. Grant get, list, watch, and patch on
agents, and ensure the design specifies testing the complete binding rather than
only patch access.
- Line 431: Update the BMI creation phase and its reconciliation logic to remain
idempotent when BareMetalInstances.Create succeeds but its response or status
update is lost: define and consistently use an idempotency key or unique worker
name, handle AlreadyExists by retrieving and adopting the existing BMI, and add
a test covering this lost-response failure path.
- Around line 418-420: Update the “Minimum MCE Version” section to name one
released MCE bundle that contains both assisted-service PR `#10717` and
assisted-installer-agent PR `#1568`, rather than separate component versions.
Define an installation-time compatibility check in the provisioning flow and
block worker provisioning when the deployed MCE bundle is below that exact
supported release.
- Around line 168-175: Separate teardown from replacement by introducing a
teardown-specific worker phase instead of using Failed when Agent unbinding
times out, and ensure automatic BMI replacement only handles genuine
provisioning failures. Update the scale-down and deletion flows around the
Failed handling, unbinding-timeout assignment, and status.workers removal so
teardown requires a drained node and unbound Agent before calling
BareMetalInstances.Delete, then retains the status entry until the BMI reaches
terminal deletion; apply the same guard to manual removal.
- Around line 258-270: Update the YAML example’s currentWorkers value from 5 to
3 so it counts only the three Ready workers and excludes the two Failed workers,
while leaving desiredWorkers and readyWorkers unchanged.
- Around line 411-416: Update the Agent-to-BMI correlation algorithm to require
an authoritative match for the expected ClusterDeployment and InfraEnv in
addition to namespace, ClusterOrder ownership, and MAC. Reject zero or multiple
candidates before any binding or BMI deletion side effect, including the flow at
the referenced deletion logic, and add coverage for stale, cross-installation,
and ambiguous matches.
- Around line 120-121: The design must define how BareMetalInstances.Create
handles spec.user_data: either specify a supported Secret-reference format and
the API/controller resolution path before BMI creation, or pass the ignition
content inline while enforcing the 64 KB limit. Update the Secret and BMI
provisioning flow to match the chosen contract, and add an integration test
covering host boot with the resulting user data.
- Line 119: The design must define how worker architecture is handled
consistently with DiskImage resolution: either restrict provisioning to amd64
and explicitly reject other resourceClass architectures, or derive the worker
architecture from nodeRequests[].resourceClass, resolve the matching
architecture-specific DiskImage, and add tests for each supported architecture.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository: osac-project/coderabbit/.coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: dd6dc1e1-773f-4246-a974-1d00db382576

📥 Commits

Reviewing files that changed from the base of the PR and between d562722 and 4322c8a.

📒 Files selected for processing (1)
  • enhancements/OSAC-2135-caas-bare-metal-worker-provisioning/design.md

Comment thread enhancements/OSAC-2135-caas-bare-metal-worker-provisioning/design.md Outdated
Comment thread enhancements/OSAC-2135-caas-bare-metal-worker-provisioning/design.md Outdated
Comment thread enhancements/OSAC-2135-caas-bare-metal-worker-provisioning/design.md Outdated
Comment thread enhancements/OSAC-2135-caas-bare-metal-worker-provisioning/design.md Outdated
Comment thread enhancements/OSAC-2135-caas-bare-metal-worker-provisioning/design.md Outdated
@github-actions

github-actions Bot commented Aug 10, 2026 •

Copy link
Copy Markdown

AI Design Review: EP-198

Score: 8/8 | Verdict: PASS

Criterion Score Notes
Feasibility 2/2 Deep technical specificity throughout. Go struct definitions with kubebuilder annotations for ClusterOrderStatus/WorkerStatus. Proto field references showing existing fields consumed (not new definitions). Retry backoff tables with exact values per failure type (30s/60s/120s cap 5m for transient; 5m/15m/30m cap 30m for resource availability). MAC correlation algorithm specified with three-dimension scoping (namespace, ownership, MAC match). Idempotency addressed explicitly — controller checks fo
Testability 2/2 Strong three-level test plan. Unit tests: 11 specific scenarios covering worker reconciliation (create/delete counts), agent correlation (MAC match, namespace isolation), InfraEnv verification, phase transitions, idempotency (re-run produces no new API calls), system tenant isolation (List excludes, Get returns NotFound), RBAC verification. Integration tests: 4 scenarios using kind cluster with mocked private API — BMI creation verification, Agent correlation simulation, scale-down flow, deletio
Scope 2/2 Summary is 4 sentences covering what, why, and key capabilities. Non-goals are specific with deferral targets (autoscaling, VM-based workers, static IP/NMStateConfig, network boot acceleration) and Jira references. Three substantive alternatives evaluated (BareMetalPool, Direct CR creation, AAP-orchestrated) with clear rejection rationale. PRD referenced in frontmatter and linked in summary. Cross-cutting dimensions addressed: provisioning (core), networking (network_attachments + OSAC-1437), in
Architecture 2/2 All OSAC patterns followed: owner-reference annotation and system tenant isolation on CaaS-managed BMIs, ClusterOrder status extension uses list of named subobjects (not maps), controller reconciliation follows the standard finalizer-status-provisioning lifecycle, conditions (WorkersFailed, InfraEnvReady, RHCOSImageNotFound) used for lifecycle signaling. Dependencies between components are exceptionally clear — three blocking dependencies enumerated in a table with Jira links and impact analysis

Verdict: Exceptionally thorough design document that follows all OSAC architectural patterns, provides deep technical specificity (Go structs, proto references, backoff tables, failure mode matrix), has clear scope boundaries with real alternatives, and includes a concrete three-level test plan — scoring 8/8 with no weaknesses severe enough to reduce any criterion.

Feedback: Two minor improvements: (1) Add a formal Terminology section defining key terms (shared InfraEnv, system tenant, discovery ignition, MAC correlation, MinHealthyDuration) — the networking EP sets the bar here, and while your usage is consistent, an upfront glossary helps reviewers orient faster. (2) Consider replacing the N/A graduation criteria with a concrete readiness checklist (e.g., 'all CRUD lifecycle operations pass E2E, worker retry converges within 3 attempts for transient failures, no orphaned BMIs after cluster deletion') since the test plan already implies these conditions.

Critical (0)

None.

Important (2)

  1. Two goals in the Goals section are implementation-oriented rather than user-visible outcomes: 'Reuse the existing ClusterOrder controller reconciliation pattern and the private gRPC API' and 'Ensure host cleanup flows through BMaaS's existing deprovision pipeline.' These describe HOW, not WHAT — consider reframing as user/operator outcomes (e.g., 'On scale-down and cluster deletion, all bare-metal hosts are fully cleaned up before being returned to inventory').
  2. Graduation criteria is N/A with justification, but the test plan implies measurable readiness conditions that would serve as graduation criteria. Adding them would strengthen the testability score and give implementation teams a clear definition of done.

Suggestions (3)

  1. Add a Terminology section defining key terms (shared InfraEnv, system tenant, discovery ignition, MAC correlation, MinHealthyDuration, worker slot) to match the precedent set by the networking EP and help reviewers orient quickly.
  2. The MAC address field path is noted as 'TBD — depends on OSAC-2308/OSAC-3254.' Consider adding a concrete placeholder field path (e.g., status.host.interfaces[0].mac_address) to give the dependency team a target contract, even if the final path may differ.
  3. The UX Alignment section states 'This section does not apply' but doesn't confirm whether a temp-api file was checked. Adding 'No matching temp-api file exists at osac-ux/libs/ui-components/src/api/v1/' would satisfy the template requirement more explicitly.

Review cost

Model: claude-opus-4-6
Cost: $0.7707
Tokens: 8 in / 5.1k out
Cache: 331.6k read
Active time: 2m 17s
API calls: 0

@carbonin carbonin left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Mostly I think we need to ensure this stays decoupled from ironic/BMO. I don't think there's any reason to be talking about those concepts here.


## Proposal

The ClusterOrder controller in osac-operator gains a new reconciliation phase for bare-metal worker management. When a ClusterOrder's `nodeRequests` reference bare-metal resource classes, the controller:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should it check for a BareMetalInstanceType matching the resource class? Or are these different things?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

reworded to clarify

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Where did you change this? I'm still not sure what this sentence as written means exactly.

3. The controller reads the InfraEnv's `status.bootArtifacts.discoveryIgnitionURL` and fetches the discovery ignition content. The ignition is architecture-neutral (the `assisted-installer-agent` image is a multi-arch manifest), so the same InfraEnv serves hosts of any architecture.
4. For each bare-metal worker requested, the controller calls `BareMetalInstances.Create` on the private API with: `spec.catalog_item` resolved from the resource class, `spec.image` set to the resolved RHCOS DiskImage ID (see RHCOS DiskImage Resolution), `spec.user_data` set to the fetched ignition content inline (the `user_data` field accepts raw first-boot data up to 64KB; the PoC measured 15KB), `spec.network_attachments` built from the Cluster's `ClusterNetworkAttachment` (subnet + security groups) and the node set's HostType (fabric interface), and `metadata.tenant = "system"` (see System Tenant Isolation). The network attachment mapping is a pass-through: the controller reads the Cluster's `ClusterNetworkAttachment` for the subnet and security group references, resolves the fabric interface name from the node set's HostType definition (first interface with role `fabric`), and constructs a `BareMetalNetworkAttachment` with `primary: true`. BMaaS handles the physical networking — moving the host to the tenant subnet VLAN and assigning an IP via fabric DHCP — as part of BMI provisioning (dependency: OSAC-1437). If the host fails to join the tenant network, the agent will not register on the expected subnet, and the existing `AgentRegistrationTimeout` handles this failure mode. API and ingress VIPs are provisioned by the existing AAP template (MetalLB LoadBalancer Services) and are not managed by this controller.
5. The controller updates ClusterOrder status with the BMI references in `workers[]`.
6. BMaaS allocates a host, writes the qcow2 to disk via Ironic, and boots with the discovery ignition. The host registers as an Agent with assisted-service.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think you should mention Ironic here. In theory this should be backend agnostic.

2. The controller removes `Failed` workers first — deletes their dead BMIs and removes their `status.workers` entries. If more removals are needed after clearing all failed slots, the controller decreases NodePool `.spec.replicas` by the remaining excess.
3. CAPI's MachineDeployment controller (used by HyperShift's default Replace upgrade type) manages MachineSets, which select Machines for deletion. CaaS does not control the selection order.
4. CAPI drains each selected node, then the AgentMachine controller unbinds the Agent (clears `ClusterDeploymentName`, removes labels and ignition refs).
5. Because BMH resources exist, the Agent enters `UnbindingPendingUserAction`. The BMH agent controller triggers Ironic deprovision (clears `bmh.Spec.Image`, removes the `detached` annotation).

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This will not necessarily be true and the BMH, if it does exist, won't be related to the agent directly (it won't have the infraenv lable).

The BMH is an implementation detail of BMaaS. I think this whole section needs to be reworked.

When the host is unbound by CAPI agent controller it will sit in unbinding-pending-user-action and it will remain there. So the next step needs to be deleting the BareMetalInstance.

I don't remember, but does this allow CAPI to properly drain and remove the node? I didn't test this in the PcC.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

So it's possible that assisted will move through "reclaim" in this case, but this all needs to be investigated more. Generally I'd look closely anywhere you're making assumptions about ironic, BMH, etc. This should work with any backend we implement behind the BMaaS API.


The InfraEnv has no `clusterRef` and no `sshAuthorizedKey` — it generates unbound discovery ignition that is not scoped to any cluster. Agent-to-cluster binding happens explicitly in the correlation phase (step 9), where the controller sets `clusterDeploymentName` on each Agent after MAC-based matching.

**Why a shared InfraEnv:** The discovery ignition is architecture-neutral (`assisted-installer-agent` is a multi-arch manifest) and does not vary by cluster or tenant. The `cpuArchitecture` field on InfraEnv only affects ISO/kernel/rootfs URLs in `status.bootArtifacts`, not the ignition content — and this design uses the ignition-only flow, not ISO download. The InfraEnv's pull secret is a platform-level credential (Cloud Provider Admin's registry credentials for pulling the discovery agent image), separate from the per-cluster pull secret used for OCP release images. The existing OSAC deployment already uses a single `infraenv` in the `hardware-inventory` namespace with a platform-level `pull-secret`.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

and does not vary by cluster or tenant.

This is only true if we never have to provide static networking, right? Are we sure that's the case?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

good call, we changed approach to 1 infraenv per cluster to simplify all this


#### RHCOS DiskImage Resolution

The controller resolves the RHCOS boot image via a pre-registered DiskImage resource (dependency: OSAC-2540 DiskImage, OSAC-1270 BMI DiskImage integration). The Cloud Infrastructure Admin registers RHCOS qcow2 images as provider-global DiskImages with guest OS family (`linux`) and architecture (`amd64`), and applies a CaaS-specific label `osac.openshift.io/ocp-version: "4.22"` to enable version-based lookup. This label is a CaaS convention — the DiskImage resource itself (OSAC-2540) has no OCP version field, since version-based lookup is a CaaS-specific need. If richer metadata is needed (e.g., multiple image variants per version, automated registration), a dedicated `ClusterDiskImage` resource could wrap DiskImage with CaaS-specific fields. Labeling is sufficient for this design.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The Cloud Infrastructure Admin registers RHCOS qcow2 images

@adriengentil we agreed to only support OCI images, right? Do we know how are they getting an RHCOS OCI image to use with this flow before BMO supports bootc?


The controller resolves the RHCOS boot image via a pre-registered DiskImage resource (dependency: OSAC-2540 DiskImage, OSAC-1270 BMI DiskImage integration). The Cloud Infrastructure Admin registers RHCOS qcow2 images as provider-global DiskImages with guest OS family (`linux`) and architecture (`amd64`), and applies a CaaS-specific label `osac.openshift.io/ocp-version: "4.22"` to enable version-based lookup. This label is a CaaS convention — the DiskImage resource itself (OSAC-2540) has no OCP version field, since version-based lookup is a CaaS-specific need. If richer metadata is needed (e.g., multiple image variants per version, automated registration), a dedicated `ClusterDiskImage` resource could wrap DiskImage with CaaS-specific fields. Labeling is sufficient for this design.

The controller reads `NodePool.spec.release.image`, extracts the OCP major.minor version (e.g., `4.22` from `ocp-release:4.22.5-x86_64`), and resolves the worker architecture from the node set's `BareMetalInstanceType` (which defines the HostType and its CPU architecture). It then queries for provider-global DiskImages matching both the `osac.openshift.io/ocp-version` label and the resolved architecture. This design targets `amd64` only; other architectures require Cloud Infrastructure Admin to register the corresponding RHCOS DiskImages.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Will this always be a tag? What if it's a digest in a mirror or something? Or is it not an image at all?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

good call, addressed in the design proposal


CaaS-managed BMIs are created under the builtin `system` tenant via the private API. The `system` tenant is excluded from `DetermineVisibleTenants`, so these BMIs are invisible to all regular users without any additional filtering. The private API bypasses tenant-scoped OPA policies because it operates with system-level credentials. Ownership is traceable via the `osac.openshift.io/owner-reference` annotation linking each BMI to its parent ClusterOrder (which belongs to the real tenant).

The discovery ignition contains the InfraEnv's pull secret and the assisted-service endpoint URL (both platform-level, not cluster-specific). It is passed inline as `user_data` on each BMI (max 64KB; PoC measured 15KB). The `user_data` field is immutable (enforced by the proto `IMMUTABLE` field behavior annotation). The ignition comes from the shared platform-level InfraEnv and is not cluster-scoped — agent-to-cluster binding is enforced by the MAC correlation algorithm, not by the ignition content.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The discovery ignition contains the InfraEnv's pull secret and the assisted-service endpoint URL (both platform-level, not cluster-specific)

Is the pull-secret platform level? How does that work in practice? Does the cloud admin use their pull secret? Won't that get into the cluster in a way the tenant user can see?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, the cloud admin should use their pull secret for the discovery phase to pull the initial image. The discovery ignition is in the BMI's user_data, which is system owned so not readable.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sure, the user can't see the user data, but they will be able to pull the secret out of the cluster once it is finished installing. In openshift-config namespace (or something like that)

@github-actions

github-actions Bot commented Aug 11, 2026 •

Copy link
Copy Markdown

AI Design Review: EP-198

Score: 8/8 | Verdict: PASS

Criterion Score Notes
Feasibility 2/2 Deep technical detail throughout: Go struct definitions with kubebuilder markers for ClusterOrder status, concrete retry backoff tables (30s/60s/120s transient, 5m/15m/30m availability), specific timeout values (30m agent registration), idempotency guards (list-before-create by label), MAC correlation algorithm with three-dimension matching. All lifecycle operations covered (provisioning, scale-up, scale-down, cluster deletion) with step-by-step flows. Six specific risks with concrete mitigation
Testability 2/2 Test plan specifies 11 unit test cases (reconcileWorkers count logic, correlateAgents MAC matching, ensureInfraEnv creation, phase transitions, idempotency, system tenant isolation, RBAC), 4 integration scenarios (kind cluster with mocked private API), and 6 E2E scenarios (full provisioning, scale-up, scale-down, deletion, tenant isolation, pool removal). Honest about E2E environment constraints (BMaaS-capable hosts needed). Graduation criteria marked N/A with valid justification (OSAC not yet r
Scope 2/2 Clear boundaries: extends ClusterOrder controller, removes BareMetalPool, no new CRDs. Four specific non-goals with justification (autoscaling deferred, VM workers deferred to VMaaS, static IP not validated by PoC, network boot acceleration tracked as OSAC-2134). Three real alternatives with specific rejection rationale. PRD referenced in frontmatter. Relevant dimensions addressed: CaaS/BMaaS services, personas (Cloud Infra Admin, Tenant User, Cloud Provider Admin), provisioning lifecycle, netwo
Architecture 2/2 All OSAC patterns followed: system tenant isolation via builtin tenant (migration 48, DetermineVisibleTenants exclusion), owner-reference annotations on BMIs linking to ClusterOrder, controller reconciliation phases (ensureInfraEnv, reconcileWorkers, correlateAgents, reconcileNodePoolReplicas) following the standard finalizer -> status update -> provisioning lifecycle. Conditions used for lifecycle state (WorkersFailed, InfraEnvReady, RHCOSImageNotFound, RHCOSImageAmbiguous). Dependencies clearl

Verdict: A thorough, well-structured design that follows all OSAC architectural patterns, provides deep implementation detail (Go structs, retry tables, correlation algorithms), covers all lifecycle operations with concrete failure handling, and includes a specific multi-level test plan — one of the stronger designs in the enhancement-proposals corpus.

Feedback: Two minor dimension gaps worth addressing: (1) explicitly state whether osac-installer Helm charts need changes for the new controller configuration or InfraEnv RBAC, or mark Installation as N/A with a sentence explaining why; (2) add a brief note on inventory capacity planning — even if sizing is the admin's responsibility, a recommendation or link to guidance helps operators. The MAC address field path (TBD pending OSAC-2308/OSAC-3254) is properly flagged as a dependency but should be updated with the concrete field path once that work lands, before this design is considered implementation-ready.

Critical (0)

None.

Important (2)

  1. Installation dimension not explicitly addressed: the design removes the cluster_infra AAP step and adds InfraEnv RBAC (patch on agents), but does not state whether osac-installer Helm charts or values need changes for the new controller behavior or RBAC. Even if no Helm changes are needed, an explicit statement would close the gap.
  2. MAC address field path marked TBD (depends on OSAC-2308/OSAC-3254): the correlation algorithm is well-specified but the exact BMI status field for MAC is unresolved. The design correctly gates on this dependency, but implementation planning requires this to be pinned before work begins.

Suggestions (3)

  1. Define key terminology upfront (CaaS, BMI, InfraEnv, Agent, BMaaS, system tenant) in a short glossary — terms are used consistently but readers unfamiliar with the domain would benefit from explicit definitions.
  2. Add a brief inventory capacity planning note: even a sentence like 'Cloud Infrastructure Admins should maintain at least 2x headroom in host inventory for the largest expected concurrent scale-up' would help operators plan.
  3. Consider adding a sequence diagram or state machine for the worker phase lifecycle (Provisioning -> WaitingForAgent -> Binding -> Ready, with Failed as a side state and Unbinding -> Deleting for teardown) — the text describes it well but a visual would aid reviewers.

Review cost

Model: claude-opus-4-6
Cost: $0.9414
Tokens: 10 in / 5.3k out
Cache: 210.9k read
Active time: 2m 7s
API calls: 0

@github-actions

github-actions Bot commented Aug 13, 2026 •

Copy link
Copy Markdown

AI Design Review: EP-198

Score: 8/8 | Verdict: PASS

Criterion Score Notes
Feasibility 2/2 The design is grounded in a validated PoC (OSAC-2817) that confirmed end-to-end BMI provisioning (~6min), agent registration, and worker join. All three blocking dependencies (MAC address in BMI status, DiskImage integration, BMaaS networking) are explicitly identified with Jira links and impact-if-not-delivered analysis. The design consumes existing proto fields and private API methods rather than inventing new ones, and the system tenant isolation leverages an existing builtin (migration 48).
Testability 2/2 The test plan is concrete and stratified across three levels. Unit tests enumerate 11 specific test cases covering reconciliation logic, MAC correlation, phase transitions, idempotency, and tenant isolation. Integration tests describe 4 scenarios using kind clusters with a mocked private API. E2E tests cover 6 flows including provisioning, scale-up, scale-down, deletion, tenant isolation verification, and pool removal validation. The design honestly notes E2E tests require BMaaS-capable environm
Scope 2/2 Scope is tightly defined: on-demand BMI provisioning replaces the static BareMetalPool, with explicit non-goals (autoscaling, VM workers, static IP, network boot acceleration). The design correctly identifies what it owns vs. what it consumes (DiskImage integration, MAC address field, BMaaS networking are consumed dependencies, not owned). The BareMetalPool removal is clean with no coexistence period, and the upgrade/downgrade strategy addresses mid-provisioning clusters and AAP job queue draini
Architecture 2/2 The design follows OSAC architectural patterns faithfully. The separate BareMetalWorkerReconciler keeps the ClusterOrder controller generic and extensible for future VM workers — a clean separation of concerns. Shared status.workers[] partitioned by kind with optimistic concurrency is the standard controller-runtime approach. System tenant isolation via the existing builtin is architecturally superior to label-based filtering (the PR description's approach), avoiding any public API changes. The

Verdict: A thorough, well-structured design that replaces a fragile static pool with on-demand BMI provisioning, grounded in a validated PoC, with clear dependency tracking, clean architectural separation, and comprehensive test coverage across all three levels.

Feedback: The design is strong and implementation-ready once its three blocking dependencies land. Two minor improvements: (1) The Catalog Item Resolution section introduces a ClusterTemplate metadata addition (catalog_item field in meta/osac.yaml) but doesn't show the schema or validate that ClusterTemplate metadata is extensible today — add a brief proto/schema snippet showing the exact change needed so implementers don't discover a blocker at code time. (2) The scale-down flow delegates Machine selection to CAPI ('CaaS does not control the selection order') which is correct, but consider documenting whether CaaS should prefer removing newest-provisioned workers to preserve longer-running workloads, or whether this is explicitly left to CAPI's default behavior as a design choice.

Critical (0)

None.

Important (2)

  1. PR description says 'Public API excludes CaaS-managed BMIs via managed-by: caas label' but the design itself uses system tenant isolation — the PR description is stale and should be updated to match the design to avoid reviewer confusion.
  2. Catalog Item Resolution requires a ClusterTemplate metadata extension (catalog_item field in meta/osac.yaml) but no schema snippet or validation that ClusterTemplate metadata supports this extension is provided — this could surface as a hidden blocker during implementation.

Suggestions (3)

  1. Consider documenting an explicit position on Machine selection order during scale-down (newest-first vs. CAPI default) since workload disruption patterns differ.
  2. The MinHealthyDuration of 1 hour for attemptCount reset is stated but not justified — consider whether this should be configurable or document the rationale (e.g., based on typical BMaaS transient failure windows).
  3. The 'future optimization — shared InfraEnv for agent pooling' note could benefit from a brief analysis of the pull secret isolation trade-off it would require solving, to help future designers evaluate it.

Review cost

Model: claude-opus-4-6
Cost: $0.6164
Tokens: 14 in / 2.6k out
Cache: 751.0k read
Active time: 1m 19s
API calls: 0

@github-actions

github-actions Bot commented Aug 13, 2026 •

Copy link
Copy Markdown

AI Design Review: EP-198

Score: 8/8 | Verdict: PASS

Criterion Score Notes
Feasibility 2/2 The design is grounded in a validated PoC (OSAC-2817) that confirmed end-to-end flow: BMI provisioning in ~6 minutes, agent registration, and worker joining the HyperShift cluster. Uses existing proto fields and API surfaces with no new CRDs. Three blocking dependencies (MAC address OSAC-2308/OSAC-3254, DiskImage OSAC-2540/OSAC-1270, BMaaS networking OSAC-1437) are explicitly gated rather than assumed. Controller patterns (watch-based reconciliation, idempotent phases, state rebuild on restart)
Testability 2/2 Comprehensive three-level test plan. Unit tests cover specific controller methods (reconcileWorkers, correlateAgents, ensureInfraEnv) with concrete assertions including idempotency, phase transitions, timeout handling, and system tenant isolation. Integration tests use kind clusters with mocked private API to validate controller behavior end-to-end. E2E tests cover the full lifecycle (provisioning, scale-up, scale-down, deletion) plus tenant isolation and BareMetalPool removal verification. The
Scope 2/2 Goals and non-goals are sharply defined. Autoscaling, VM workers, static IP, and network boot caching are explicitly deferred. The old BareMetalPool flow is removed entirely with no coexistence period, which is bold but reduces long-term complexity. Upgrade/downgrade strategy is detailed and practical, including AAP job queue drain, handling of mid-provisioning clusters, and CRD status field backward compatibility. Dependencies are enumerated with concrete Jira links and impact analysis. Minor s
Architecture 2/2 Excellent separation of concerns: a dedicated BareMetalWorkerReconciler keeps BM-specific logic out of the ClusterOrder controller, with a clear extension point for future VMWorkerReconciler. Shared status.workers[] partitioned by kind avoids separate status fields per worker type while maintaining controller independence via optimistic concurrency. System tenant isolation (builtin system tenant, DetermineVisibleTenants exclusion) is architecturally superior to label-based filtering — no public

Verdict: A thorough, well-architected design validated by a working PoC, with clean separation of concerns, sound isolation decisions, and comprehensive failure handling — strong across all dimensions.

Feedback: The PR description is stale relative to the design document: it describes label-based visibility filtering ('managed-by: caas') and CaaS-specific DiskImage labels ('osac.openshift.io/ocp-version'), while the actual design uses system tenant isolation and typed ClusterVersion references, which are architecturally superior — update the PR body to match the current design to avoid confusing reviewers. Consider whether the future DiskImage automation path (extracting RHCOS from release payloads) can be scoped into the same release to reduce the day-1 operational burden of the manual 5-step registration process. The dependency on the exact MAC field path in BMI status (OSAC-2308/OSAC-3254) is well-documented but consider defining a placeholder interface now to unblock controller development in parallel.

Critical (0)

None.

Important (2)

  1. PR description describes a stale design (label-based visibility filtering via 'managed-by: caas' label, CaaS-specific 'osac.openshift.io/ocp-version' DiskImage labels) while the design document uses system tenant isolation and typed ClusterVersion disk_image references. The design document is correct and architecturally superior, but reviewers reading the PR body will get a misleading picture of the approach.
  2. Three hard-blocking dependencies (MAC address in BMI status, DiskImage resource + BMI integration, BMaaS networking for subnet attachment) must all ship before implementation can begin. Each dependency has Jira links and impact analysis, but the cumulative delivery risk of this dependency chain warrants explicit coordination with the BMaaS and DiskImage teams on timeline alignment.

Suggestions (3)

  1. Update the PR description to match the current design document — specifically the visibility mechanism (system tenant, not labels) and DiskImage resolution (ClusterVersion typed reference, not label lookup).
  2. Consider defining an interface/mock for the MAC address field path now (even before OSAC-2308/OSAC-3254 ships) to enable parallel controller development and testing.
  3. The manual DiskImage registration process (5 steps including qcow2-to-OCI repackaging) is operationally heavy for day-1. Evaluate whether the automated extraction path (introspecting ClusterVersion release payloads for RHCOS references) can be scoped into the initial release to reduce admin burden.

Review cost

Model: claude-opus-4-6
Cost: $0.4075
Tokens: 7 in / 4.5k out
Cache: 282.3k read
Active time: 1m 38s
API calls: 0

@github-actions

github-actions Bot commented Aug 13, 2026 •

Copy link
Copy Markdown

AI Design Review: EP-198

Score: 8/8 | Verdict: PASS

Criterion Score Notes
Feasibility 2/2 PoC (OSAC-2817) validated the end-to-end flow (6-min BMI provisioning, agent registration, worker join). Three hard dependencies (MAC address OSAC-2308/3254, DiskImage OSAC-2540/1270, BMaaS networking OSAC-1437) are clearly documented as blocking. All proto fields and API patterns used already exist or are introduced by those dependencies. Idempotent reconciliation, state rebuild on restart, and escalating retry backoff are well-specified. Manual DiskImage registration is operationally heavy but
Testability 2/2 Comprehensive three-tier test plan: unit tests cover all key behaviors (worker CRUD, MAC correlation, InfraEnv creation, phase transitions, idempotency, tenant isolation, RBAC), integration tests use kind clusters with mocked private API, and E2E tests cover full provisioning/scale/deletion flows. The design honestly acknowledges E2E limitations (requires BMaaS-capable environment) and proposes API-level mocking as initial coverage. Test cases are specific and trace back to design behaviors.
Scope 2/2 Well-bounded with explicit goals and non-goals (autoscaling, VM workers, static IP, network boot acceleration all deferred). Clean removal of BareMetalPool with no coexistence period. Changes span osac-operator (new controller), fulfillment-service (ClusterVersion proto, CLI --disk-image flag), AAP (remove cluster_infra step), and deployment config (disable assisted image service), but each change is minimal and justified. Dependencies are tracked with Jira tickets.
Architecture 2/2 Separate BareMetalWorkerReconciler keeps BM-specific logic out of ClusterOrder controller — extensible for future VMWorkerReconciler. System tenant isolation is elegant: zero public API changes, leverages existing tenancy exclusion logic. MAC correlation scoped by namespace + ownership + MAC prevents cross-cluster contamination. Late-binding InfraEnv gives full control over agent lifecycle. Shared status.workers[] partitioned by kind handles multi-type workers cleanly. Metrics use bounded tenant

Verdict: A thorough, well-validated design that replaces static BareMetalPool pre-boot with on-demand BMI provisioning via a clean controller separation, system tenant isolation, and MAC-based agent correlation — all backed by PoC validation and comprehensive failure handling.

Feedback: The design is strong across all dimensions. Two items worth addressing before implementation: (1) the PR description still references 'managed-by: caas label-based visibility filtering' while the design correctly uses system tenant isolation — update the PR body to match the final design to avoid reviewer confusion. (2) Consider making the MinHealthyDuration (1 hour) for attemptCount reset configurable via a controller flag or ConfigMap, since different environments may have different stability baselines, and a hardcoded value will require a code change to tune in production.

Critical (0)

None.

Important (2)

  1. PR description is stale: mentions 'managed-by: caas label-based visibility filtering' but the design document correctly uses system tenant isolation (builtin system tenant excluded from DetermineVisibleTenants). The PR body should be updated to match the final design to prevent reviewers from evaluating the wrong isolation mechanism.
  2. MAC address field path in BMI status is explicitly TBD (depends on OSAC-2308/OSAC-3254). Implementation cannot proceed without this being resolved. The design correctly gates the feature on this dependency but should specify a fallback or validation step to confirm the field path once the dependency lands.

Suggestions (3)

  1. Consider making MinHealthyDuration (1 hour for attemptCount reset) and agent registration timeout (30 minutes) configurable via controller flags or ConfigMap rather than hardcoded constants, to allow tuning in environments with different infrastructure response characteristics.
  2. The concurrent multi-cluster provisioning scenario (multiple ClusterOrders competing for limited host inventory) is noted in Risks but could benefit from a brief discussion of ordering/priority semantics — e.g., whether FIFO ordering is preserved or whether higher-priority tenants could preempt.
  3. Consider adding a worker retry count histogram metric (osac_clusterorder_worker_retry_count) to help operators distinguish persistent infrastructure issues from transient failures without inspecting individual ClusterOrder statuses.

Review cost

Model: claude-opus-4-6
Cost: $0.4935
Tokens: 7 in / 3.3k out
Cache: 261.2k read
Active time: 1m 18s
API calls: 0

@github-actions

github-actions Bot commented Aug 13, 2026 •

Copy link
Copy Markdown

AI Design Review: EP-198

Score: 8/8 | Verdict: PASS

Criterion Score Notes
Feasibility 2/2 The design is grounded in a successful PoC (OSAC-2817) that validated the end-to-end flow with concrete measurements (6-minute BMI provisioning, 15KB ignition). It uses existing proto fields and private API endpoints — no new proto changes are owned by this design. Three blocking dependencies (MAC address in BMI status, DiskImage resource, BMaaS networking) are explicitly identified with Jira tickets and correctly gated. The ClusterVersion proto extension (disk_image field) is minimal and the sy
Testability 2/2 Comprehensive three-tier test plan: 11 specific unit tests covering reconciler logic, MAC correlation, phase transitions, idempotency, and tenant isolation; 4 integration test scenarios using kind clusters with mocked private API; 6 E2E test scenarios covering full provisioning, scale-up/down, deletion, and tenant visibility. Observability is well-covered with 5 Prometheus metrics aggregated by tenant (avoiding unbounded cardinality) and 8 Kubernetes event types for per-instance diagnostics. The
Scope 2/2 Scope is tightly defined with 6 goals and 4 explicit non-goals (autoscaling, VM workers, static IP, network boot acceleration). The design stays focused on replacing the static BareMetalPool with on-demand BMI provisioning. Dependencies are clearly separated from owned work. Upgrade/downgrade strategy is thorough, covering AAP job queue draining, mid-provisioning cluster handling, and status field compatibility. The BareMetalPool removal is clean with no coexistence period, and the rationale for
Architecture 2/2 The separate BareMetalWorkerReconciler provides clean separation of concerns and is explicitly extensible for future VM workers via the same pattern. System tenant isolation is elegant — it reuses the existing builtin tenant exclusion rather than adding label-based filtering to every API endpoint. The 4-phase reconciliation structure (ensureInfraEnv, reconcileWorkers, correlateAgents, reconcileNodePoolReplicas) is idempotent with state rebuild on restart. MAC correlation uses 3-dimension scoping

Verdict: A thorough, well-structured design backed by PoC validation, with clean architectural separation (dedicated reconciler, system tenant isolation, categorized retry), comprehensive test coverage across all tiers, and tightly scoped goals with explicitly gated dependencies.

Feedback: The catalog item resolution mechanism is the weakest link — describing it as a 'template parameter' passthrough without specifying the actual plumbing (how it reaches ClusterOrder.nodeRequests, whether it needs CRD changes or is carried as an annotation) leaves an implementation gap that could cause surprises. Consider adding a concrete example of the data flow from ClusterTemplate YAML through to the BareMetalWorkerReconciler's access path. The indefinite retry with no circuit breaker is a deliberate choice but could benefit from an operator-facing annotation (e.g., osac.openshift.io/pause-retries) to stop burning hosts during investigation of persistent misconfigurations, rather than relying solely on capped backoff.

Critical (0)

None.

Important (2)

  1. Catalog item resolution plumbing is underspecified: the design says catalogItem is 'passed through to the ClusterOrder as a template parameter' but does not trace the data flow through proto/CRD fields, leaving ambiguity about whether NodeRequest needs a new field or if it's carried as metadata/annotation. This could block implementation if the passthrough mechanism doesn't exist.
  2. The PR description mentions 'managed-by: caas' label-based visibility filtering, but the design document uses system tenant isolation instead. This discrepancy should be resolved to avoid confusion during review — confirm the PR description is stale and the system tenant approach is canonical.

Suggestions (3)

  1. Consider adding an operator-facing pause mechanism (annotation or condition) to halt retries on a specific worker slot during investigation, rather than relying solely on capped backoff. Operators diagnosing persistent failures (wrong DiskImage, broken network) would benefit from stopping the retry loop without scaling down.
  2. The future optimization note about shared InfraEnv for agent pooling could reference OSAC-2134 (network boot acceleration) since pre-booting hosts addresses the same latency concern from a different angle — linking them helps future planning.
  3. Document the expected ClusterOrder finalizer behavior more explicitly: the design says 'the ClusterOrder's finalizer holds until all status.workers[] entries are cleaned up' but doesn't specify which controller adds/removes this finalizer or whether it's a new finalizer or the existing one.

Review cost

Model: claude-opus-4-6
Cost: $0.6462
Tokens: 7 in / 3.9k out
Cache: 240.7k read
Active time: 1m 44s
API calls: 0

@vladikr

vladikr commented Aug 13, 2026

Copy link
Copy Markdown
Contributor

Sorry @vladikr for not answering this part.
Any part that involves cluster knowledge is on the CaaS side (interacting with Hypershift and CAP Agent). The provisioning of the machine themselves is part of BMaaS, and we're just going to interact with BMIs.
In short, BMaaS boots the host with the image we point towards, and what the image will do is our problem to handle (in this case it'll spin an Agent for the CAP Agent Hypershift provider)
What parts are you concerned about, or you think it's a BM detail that we're handling within CaaS?

Yeah, thanks.
That totally makes sense.
I was only concerned about the part that's treating InfraEnv / ignition / MAC / Agent lifecycle as an install plane not being in the generic ClusterOrder/topology logic.
The separate reconciler already does that.
I just thought that it might live somewhere else, mainly so VM workers won't inherit BM install details, and ClusterOrder stays about desired counts, resource class, NodePool, and aggregates.
Not asking to move that into BMaaS :)

Looks good from my side. Thank you @rccrdpccl !

@github-actions

github-actions Bot commented Aug 13, 2026 •

Copy link
Copy Markdown

AI Design Review: EP-198

Score: 8/8 | Verdict: PASS

Criterion Score Notes
Feasibility 2/2 PoC (OSAC-2817) validated the end-to-end flow with measured 6-minute BMI provisioning and successful agent registration. Three blocking dependencies (MAC address OSAC-2308/3254, DiskImage OSAC-2540/1270, BMaaS networking OSAC-1437) are honestly identified with Jira tickets and impact-if-not-delivered analysis. The design uses existing infrastructure (no new CRDs, no new services), existing private API patterns, and existing system tenant isolation. Idempotency is carefully handled: BMI creation
Testability 2/2 Test plan enumerates 11 specific unit test scenarios, 4 integration test scenarios, and 6 E2E test scenarios with concrete expected behaviors. Unit tests cover the controller phases (ensureInfraEnv, reconcileWorkers, correlateAgents), phase transitions, idempotency, and system tenant isolation at both List and Get levels. Integration tests target kind cluster with mocked private API. E2E tests cover full lifecycle (provision, scale-up, scale-down, deletion) and tenant isolation. The limitation t
Scope 2/2 Scope is large but well-bounded. Clear non-goals defer autoscaling, VM workers, static IP support, and network boot acceleration with specific Jira references. The BareMetalPool removal is clean with no coexistence period. The ClusterNodeSet proto redesign from host_type to oneof catalog_item is significant cross-component scope, but it is necessary for the feature and well-justified (eliminates ambiguous HostType-to-CatalogItem reverse lookup). The design touches osac-operator, fulfillment-serv
Architecture 2/2 Clean separation of concerns: BareMetalWorkerReconciler is a dedicated controller that watches ClusterOrder CRs alongside the existing ClusterOrder controller, partitioned by concern (generic lifecycle vs BM-specific phases). System tenant isolation is elegant - reuses the existing builtin system tenant and DetermineVisibleTenants exclusion without any public API changes. MAC-based correlation is scoped across three dimensions (namespace, ownership label, MAC match) preventing cross-cluster inte

Verdict: A thorough, well-structured design document validated by a PoC, with clean architectural patterns (separate controller, system tenant isolation, MAC-scoped correlation), comprehensive failure handling, and honestly identified blocking dependencies.

Feedback: The PR description is inconsistent with the design: it describes label-based visibility filtering via a 'managed-by: caas' label requiring fulfillment-service changes, but the design actually uses system tenant isolation requiring no public API changes. Update the PR description to match the design's approach to avoid reviewer confusion. Consider adding a brief discussion of rate limiting for BMI creation API calls in large-cluster scenarios (e.g., 50+ workers) and whether there are practical upper bounds on worker count per cluster.

Critical (0)

None.

Important (2)

  1. PR description claims 'CaaS-managed BMIs are hidden from tenant APIs via label-based visibility filtering' and lists 'Visibility filtering mechanism: Public API excludes CaaS-managed BMIs via managed-by: caas label. Requires fulfillment-service changes' as a review point, but the actual design uses system tenant isolation (metadata.tenant = system) with zero fulfillment-service changes needed. The PR description should be updated to match the design to avoid misleading reviewers.
  2. MAC address field path in BMI status is TBD (stated as 'depends on OSAC-2308/OSAC-3254, which add inventory metadata to BareMetalInstance status; the field does not exist in the current proto'). While the dependency is tracked, the correlation algorithm specification is incomplete without knowing the exact field structure - if the field ends up as a list of MACs per interface rather than a single MAC, the matching logic changes.

Suggestions (3)

  1. Add a brief discussion of rate limiting or concurrency control for BMI creation calls when provisioning large clusters. If a tenant creates a cluster with 50 bare-metal workers, the controller loops through all of them calling BareMetalInstances.Create - consider whether this should be throttled or batched.
  2. The ClusterNodeSet redesign (host_type to oneof catalog_item) is a significant cross-component breaking change. Consider whether it merits a brief API migration note in the Upgrade/Downgrade section, since existing ClusterOrder CRs with the old host_type field would need handling.
  3. Testing of concurrent controller access to status.workers[] (BareMetalWorkerReconciler and future VMWorkerReconciler writing to the same field) should be mentioned in the integration test plan, even if VM support is deferred, since the architectural decision to share the field is made now.

Review cost

Model: claude-opus-4-6
Cost: $0.9925
Tokens: 14 in / 5.8k out
Cache: 723.6k read
Active time: 2m 27s
API calls: 0

- Add OSAC-1437 as blocking dependency with network attachment
  pass-through (ClusterNetworkAttachment → BareMetalNetworkAttachment)
- Switch to shared platform-level InfraEnv (architecture-neutral
  ignition, platform-level pull secret, no clusterRef)
- Replace managed-by label with system tenant isolation
- Replace corev1.ObjectReference with WorkerStatus struct
  (phase, attemptCount, failure details, nextRetryTime)
- Add aggregate counts: desiredWorkers, currentWorkers, readyWorkers
- Infinite retry with escalating backoff, no PermanentlyFailed state
- Clarify DiskImage source_type owned by OSAC-1270
- Add OSAC-1604 alignment note for tenant-visible status

Assisted-by: Claude Code <noreply@anthropic.com>
Signed-off-by: Riccardo Piccoli <rpiccoli@redhat.com>
- Clarify DiskImage architecture resolved from BareMetalInstanceType
- Fix RBAC section to enumerate Agent permissions (get/list/watch/patch)
- Fix currentWorkers example (3 not 5, excludes Failed)
- Simplify ignition flow: inline passthrough via user_data, no
  intermediate Secret
- Separate Failed (provisioning retry) from Unbinding/Deleting
  (teardown) — teardown failures do not trigger replacement
- Add BMI creation idempotency: list-before-create with ownership
  label, note unique constraint alternative
- Renumber provisioning steps (1-9) after Secret removal

Assisted-by: Claude Code <noreply@anthropic.com>
Signed-off-by: Riccardo Piccoli <rpiccoli@redhat.com>
- Remove all Ironic/BMH/Metal3 references — design treats BMaaS as
  a black box, interacts only via private API
- Rework scale-down: watch for unbinding-pending-user-action
  (not *-unbound), then delete BMI directly
- Revert to per-cluster InfraEnv (clusterRef, ownerReference) for
  simplicity and to avoid pull secret scope / static networking
  concerns; note shared InfraEnv as future pooling optimization
- Clarify DiskImage resolution: driven by ClusterVersion.spec.version,
  no release image pullspec parsing. Document manual OCI packaging
  steps and future automation path via release image introspection
- Clarify resourceClass → HostType resolution
- Fix step numbering throughout

Assisted-by: Claude Code <noreply@anthropic.com>
Signed-off-by: Riccardo Piccoli <rpiccoli@redhat.com>
- Separate BM-specific logic into BareMetalWorkerReconciler, keeping
  ClusterOrder controller generic (ready for future VM worker support)
- Both controllers watch the same ClusterOrder CR; shared
  status.workers[] partitioned by kind field
- Add BMaaS-owned worker lifecycle as evaluated alternative
- Update goals, test plan, and references throughout

Assisted-by: Claude Code <noreply@anthropic.com>
Signed-off-by: Riccardo Piccoli <rpiccoli@redhat.com>
- Replace label-based DiskImage lookup with typed DiskImageReference
  on ClusterVersionSpec (per OSAC-1330 type-safe references)
- Reference validation at ClusterVersion creation time, deletion
  protection, no ambiguity (eliminates RHCOSImageAmbiguous condition)
- Update manual registration flow: link DiskImage to ClusterVersion
  via --disk-image flag on CLI
- Flag cluster upgrade dependency as a risk (PRD stage)

Assisted-by: Claude Code <noreply@anthropic.com>
Signed-off-by: Riccardo Piccoli <rpiccoli@redhat.com>
@github-actions

github-actions Bot commented Aug 14, 2026 •

Copy link
Copy Markdown

AI Design Review: EP-198

Score: 8/8 | Verdict: PASS

Criterion Score Notes
Feasibility 2/2 PoC (OSAC-2817) validates the end-to-end flow. Dependencies are explicitly documented with Jira links and impact-if-not-delivered analysis. Implementation details are thorough: gRPC call behavior (30s deadline, 3-error backoff), idempotency guards (list-before-create by label), state rebuild on restart, and retry behavior with categorized backoff. The controller uses established osac-operator patterns (controller-runtime, private API client, finalizers). The only uncertainty is the TBD MAC addre
Testability 2/2 Comprehensive test plan with 11 unit test scenarios, 4 integration scenarios, and 6 E2E scenarios. Unit tests cover reconciliation logic, MAC correlation, idempotency, tenant isolation, and RBAC. Integration tests use kind cluster with mocked private API. E2E tests cover full provisioning, scale-up/down, deletion, and tenant isolation. The observability section provides 5 Prometheus metrics and 8 Kubernetes events that support operational verification. Test environment constraints (BMaaS-capable
Scope 2/2 Well-bounded: clear Goals (6) and Non-Goals (4) with explicit deferrals (autoscaling, VM workers, static IP, network boot). Four blocking dependencies are documented with Jira links. The BareMetalPool removal is a clean cut (no coexistence period) with a thorough upgrade procedure including AAP job queue drain. Cross-component changes are fully mapped (proto, CRD, CLI, CatalogItem, AAP steps). The ClusterNodeSet redesign adds significant scope but is necessary and well-justified by the HostType
Architecture 2/2 Strong separation of concerns: BareMetalWorkerReconciler isolated from ClusterOrder controller. System tenant isolation is elegant — uses the existing builtin mechanism (migration 48, DetermineVisibleTenants exclusion) requiring zero API changes. Late binding InfraEnv with MAC correlation is well-reasoned for agent lifecycle control. DiskImage reference on ClusterVersion (not Cluster) is architecturally sound for the upgrade path. Worker lifecycle phases are clearly defined with proper state mac

Verdict: A thorough, well-structured design that replaces a fragile static pool with on-demand provisioning using validated patterns (PoC-backed, controller-runtime, system tenant isolation); the multi-node-set NodePool mapping and shared-status patch strategy are the main areas needing clarification before implementation.

Feedback: The design should clarify how multiple node sets (e.g., 'compute' + 'gpu' with different BareMetalInstanceTypes) map to NodePools — the current text references a single NodePool and CAPI's random Machine selection, which won't correctly handle per-node-set scale-down if each node set should scale independently. Commit to a specific strategy for the shared status.workers[] patch (server-side apply with field ownership would be cleanest) rather than leaving it as 'either/or'. The management-state annotation check (osac.openshift.io/management-state → skip reconciliation when Unmanaged) is an established osac-operator pattern that the BareMetalWorkerReconciler section does not mention — ensure it is included for consistency with other controllers.

Critical (0)

None.

Important (4)

  1. Multi-node-set to NodePool mapping is unspecified: the design shows examples with both 'compute' and 'gpu' node sets but references a single NodePool with a single agentBareMetal label. Per-node-set scale-down requires per-node-set NodePools (or equivalent CAPI selection control), otherwise CAPI's random Machine deletion could remove workers from the wrong node set. The reconcileNodePoolReplicas phase and scale-down flow need to address this.
  2. Shared status.workers[] concurrency strategy is uncommitted: the design acknowledges two controllers (ClusterOrder + BareMetalWorkerReconciler) writing to the same status subresource and proposes two alternatives (extend patchStatusWithRetry vs. separate patch paths) without committing. Under high reconciliation frequency (e.g., multiple clusters scaling simultaneously), this could cause excessive optimistic concurrency retries. A committed approach (server-side apply with field ownership is the
  3. MAC address field path is TBD (depends on OSAC-2308/OSAC-3254): the correlation algorithm's implementation is blocked until the exact BMI status field is defined. The design should specify the expected field structure (even if provisional) so the controller interface can be designed without coupling to the final field path.
  4. Management-state annotation check missing: all osac-operator controllers check osac.openshift.io/management-state and skip reconciliation when set to Unmanaged. The BareMetalWorkerReconciler does not mention this pattern — it should be included for consistency and operational control.

Suggestions (4)

  1. Source markers are absent throughout the document. The section guidance requires traceability markers ([PRD: ...], [Assumption], [Codebase: ...]) for non-obvious decisions. Key assumptions (e.g., 'the boot image is ephemeral', 'any Z stream within the same minor version is acceptable', 'CAPA does not delete the Agent') should be marked as [Assumption] to help reviewers validate them.
  2. The PR body/description is stale — it references 'label-based visibility filtering' and 'CaaS-specific label (osac.openshift.io/ocp-version)' which the design has correctly evolved away from in favor of system tenant isolation and DiskImageReference on ClusterVersion. Consider updating the PR description to match the current design to avoid reviewer confusion.
  3. Consider adding a readiness/feature gate for the BareMetalWorkerReconciler so it can be selectively disabled in environments where CaaS bare-metal provisioning is not needed, without disabling the entire osac-operator. This aligns with the Support Procedures section's 'disabling the feature' guidance.
  4. The cross-API-group ownerReference (ClusterOrder apiVersion osac.openshift.io → InfraEnv apiVersion agent-install.openshift.io) for garbage collection should be explicitly validated — Kubernetes GC handles cross-group ownerReferences but the owning controller needs appropriate permissions and the GC controller must be aware of the cross-group relationship. Document this as a deployment validation step.

Review cost

Model: claude-opus-4-6
Cost: $1.3706
Tokens: 22 in / 9.0k out
Cache: 822.3k read
Active time: 3m 51s
API calls: 0

@github-actions

github-actions Bot commented Aug 14, 2026 •

Copy link
Copy Markdown

AI Design Review: EP-198

Score: 8/8 | Verdict: PASS

Criterion Score Notes
Feasibility 2/2 The design is grounded in a validated PoC (OSAC-2817) that confirmed the end-to-end flow. All four blocking dependencies are explicitly identified with Jira links and impact-if-not-delivered analysis. Proto field changes, CRD status extensions, and Go type definitions are concrete and use existing fulfillment-service patterns (private API, system tenant, controller-runtime). The gRPC call behavior (30s deadline, 3-consecutive-error backoff) and idempotency guard (label-based dedup before Create)
Testability 2/2 Test plan covers all three tiers with specific, implementable scenarios. Unit tests enumerate 11 concrete behaviors (phase transitions, MAC correlation scoping, idempotency, tenant isolation). Integration tests specify kind-cluster scenarios with mocked private API for InfraEnv creation, Agent correlation, scale-down, and deletion cleanup. E2E tests cover the full provisioning, scale-up, scale-down, deletion, and tenant isolation flows. The design honestly notes that full E2E requires BMaaS-capa
Scope 2/2 Scope is well-bounded: one new controller (BareMetalWorkerReconciler), CRD status extensions (no new CRDs), one new proto field (disk_image on ClusterVersionSpec), and removal of the BareMetalPool workflow. Non-goals are clear (autoscaling, VM workers, static IP, network boot caching). The design explicitly delineates what it owns vs. what it consumes from dependencies (OSAC-2540 owns DiskImage, OSAC-1270 owns source_type disk_image, OSAC-1330 owns typed references, OSAC-1437 owns BMaaS networki
Architecture 2/2 The design follows established osac-operator patterns: separate reconciler watching ClusterOrder CRs, idempotent phase-based reconciliation, status-persisted worker lifecycle, controller-runtime requeue for retry. System tenant isolation reuses the existing builtin tenant (migration 48) and DetermineVisibleTenants exclusion — no new filtering mechanism needed. The controller separation (ClusterOrder controller for generic lifecycle, BareMetalWorkerReconciler for BM-specific phases) keeps concern

Verdict: A thorough, implementation-ready design document that covers all required template sections with concrete technical detail, validated by a PoC, with explicit dependency tracking and well-reasoned architectural decisions.

Feedback: The PR description mentions label-based visibility filtering ('managed-by: caas' label) for tenant isolation, but the design document correctly uses the system tenant approach instead — update the PR description to match the final design to avoid reviewer confusion. The MAC address field path is marked 'TBD' pending OSAC-2308/OSAC-3254; consider adding an Open Questions section for this since it affects the controller's core correlation logic. The shared status.workers[] between two controllers warrants a brief note on whether a Server-Side Apply field manager strategy would be preferable to the described patchStatusWithRetry approach, as SSA provides stronger field-ownership guarantees for multi-controller status updates.

Critical (0)

None.

Important (2)

  1. PR description is stale: it describes label-based visibility filtering ('managed-by: caas') and label injection in the fulfillment-service server, but the design document uses the system tenant (migration 48) approach which requires no fulfillment-service changes. This mismatch will confuse reviewers who read the PR body before the design.
  2. The MAC address field path on BareMetalInstance status is marked 'TBD' (depends on OSAC-2308/OSAC-3254) but there is no Open Questions section documenting this. Since MAC correlation is the core mechanism for Agent-to-BMI matching and the exact field path affects the controller implementation, this should be tracked as an explicit open question with an owner.

Suggestions (3)

  1. Consider documenting whether Server-Side Apply with distinct field managers (one for ClusterOrder controller, one for BareMetalWorkerReconciler) would be a cleaner approach than the described patchStatusWithRetry for the shared status.workers[] field, as SSA provides stronger field-ownership guarantees and avoids the 'must not clobber' concern mentioned in the Controller Reconciliation Structure section.
  2. The retry backoff table distinguishes three failure types with different escalation strategies, but the controller's classification logic (how it determines which backoff curve to apply) is not described. A brief note on how the controller maps BMI/Agent error conditions to the failure type taxonomy would help implementers.
  3. The InfraEnv creation section mentions future NMStateConfig support as a reason for per-cluster InfraEnvs, but static IP is listed as a non-goal. Consider removing the NMStateConfig forward-reference or noting it is a secondary justification, to avoid implying scope creep.

Review cost

Model: claude-opus-4-6
Cost: $0.7444
Tokens: 14 in / 2.9k out
Cache: 682.3k read
Active time: 1m 19s
API calls: 0

…sistency fixes

ClusterNodeSet: replace HostType with BareMetalInstanceTypeReference.
Add nodeSet field to WorkerStatus for multi-node-set scaling.
Own disk_image field on ClusterVersionSpec (with upgrade path).
Add OSAC-1330 to dependency table.

User flow: unified Setup Flow with before/after CLI commands and
explanations for DiskImage, BareMetalInstanceType, and template changes.
BMI Create field-source table in provisioning step 4.

Networking: use ClusterNetworkAttachment, add Network Attachment
Enrichment section, read subnet from ClusterOrder CRD.

Installer/Enclave: document impact on importAgents, assisted image
service, pool playbooks, and Wizard configuration.

Fixes: stale InfraEnv reference, gRPC timeout pattern, RBAC for
infraenvs, worker_type metric values, patchStatusWithRetry note,
InfraEnv deletion recovery, OSAC-3266 name uniqueness for BMI
creation idempotency.

Assisted-by: Claude Code <noreply@anthropic.com>
Signed-off-by: Riccardo Piccoli <rpiccoli@redhat.com>
@github-actions

github-actions Bot commented Aug 14, 2026 •

Copy link
Copy Markdown

AI Design Review: EP-198

Score: 7/8 | Verdict: PASS

Criterion Score Notes
Feasibility 2/2 The design is technically feasible with clear implementation paths. All blocking dependencies are identified with Jira links and impact statements. The PoC (OSAC-2817) validated the end-to-end flow including BMI provisioning (~6 min), agent registration, and worker join. The design uses existing infrastructure (private API, system tenant, controller-runtime) and introduces no novel unproven patterns. Proto changes, CRD extensions, and controller phases are all specified with enough detail to imp
Testability 2/2 The test plan is concrete and well-structured across all three tiers. Unit tests cover specific reconciler behaviors (reconcileWorkers creates correct BMI count, correlateAgents matches MACs, phase transitions, idempotency). Integration tests exercise real controller behavior in kind clusters with mocked private API. E2E tests cover the full provisioning, scale-up, scale-down, and deletion flows. The design also honestly calls out that full E2E requires BMaaS-capable environments and initial cov
Scope 2/2 The scope is well-defined and appropriately sized. Goals and non-goals are clear — autoscaling, VM workers, static IP, and network boot acceleration are explicitly deferred. The BareMetalPool removal is a clean break (no coexistence period) with a documented upgrade/downgrade strategy. The design owns exactly one new field (disk_image on ClusterVersionSpec) and carefully delineates what it owns vs. what it consumes from dependencies. The UX Alignment section correctly identifies installer/Enclav
Architecture 1/2 The architecture is sound overall: dedicated BareMetalWorkerReconciler with clean separation from the ClusterOrder controller, four well-scoped idempotent phases, system tenant isolation via existing tenancy logic, and list-before-create plus DB uniqueness for idempotency. However, the dual-controller status write strategy is underspecified — the design acknowledges two controllers writing to the same ClusterOrder status and notes it 'must be extended' but does not specify a concrete field-owner

Verdict: A thorough, well-structured design that demonstrates strong feasibility (PoC-validated), clear scope boundaries, and concrete testability, held back from a top score only by an underspecified dual-controller status write strategy and the lack of a circuit breaker on infinite retries.

Feedback: Concretize the dual-controller status write strategy: specify whether the BareMetalWorkerReconciler uses Server-Side Apply with a distinct field manager or a dedicated status patch path that avoids clobbering ClusterOrder controller fields — the current 'must be extended' language leaves a real race condition unresolved. Consider adding an operator-configurable maximum attempt count (or circuit breaker) for worker retries to prevent resource waste on persistent misconfigurations; indefinite retry with 30m cap is reasonable for transient failures but burns hosts repeatedly when the root cause is a wrong DiskImage or broken network. Finally, define a contract or interface boundary for the MAC address field on BMI status (even if the exact proto field is TBD) so integration tests can be written against a stable expectation.

Critical (0)

None.

Important (2)

  1. Dual-controller status writes to ClusterOrder are acknowledged but the resolution is vague ('must be extended' or 'must use its own status patch path'). Two controllers writing to the same status subresource without a concrete field-ownership strategy (SSA field managers, partitioned status patches) risks field clobbering under concurrent reconciliation.
  2. Infinite retry with no circuit breaker: the design explicitly rejects PermanentlyFailed and retries indefinitely with capped backoff (30m). For persistent misconfigurations (wrong DiskImage, broken network), this burns hosts repeatedly over days. An operator-configurable max attempt count would limit resource waste while preserving declarative semantics.

Suggestions (3)

  1. Define a contract for the MAC address field on BMI status (OSAC-2308/OSAC-3254) so the correlation algorithm and integration tests can be written against a stable interface, even if the exact proto field path is finalized later.
  2. The PR description mentions 'CaaS-managed BMIs hidden from tenant APIs via label-based visibility filtering' and a 'managed-by: caas' label, but the design document uses system tenant isolation instead. The PR description should be updated to match the final design to avoid reviewer confusion.
  3. Consider documenting the expected behavior when multiple ClusterOrders compete for the same BareMetalInstanceType hosts simultaneously — the Risks section mentions inventory exhaustion but does not describe ordering/priority semantics.

Review cost

Model: claude-opus-4-6
Cost: $1.2363
Tokens: 62 in / 4.8k out
Cache: 751.6k read
Active time: 2m 5s
API calls: 0

@carbonin carbonin left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just some minor comments. Seems most of the real issues are resolved.


```bash
osac-admin create diskimage rhcos-4.18 \
--source-ref quay.io/osac/rhcos:4.18.0 \

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm guessing this is a placeholder or are we actually going to publish these?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yep, it's just to give an idea but we're not planning to publish these in this Design at least


#### Cluster Deletion

On ClusterOrder deletion, the worker reconciler does not need to explicitly orchestrate scale-down. Deleting the HostedCluster cascades through HyperShift (deletes all NodePools) → CAPI (drains nodes, deletes Machines) → CAPA (unbinds Agents). The worker reconciler reacts to Agents entering `unbinding-pending-user-action` and cleans up Agent CRs and BMIs through the normal scale-down watch (steps 5-8). The ClusterOrder's finalizer holds until all `status.workers[]` entries are cleaned up. The InfraEnv CR is garbage collected via its ownerReference to the ClusterOrder.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

BMI's should still be deleted though, right?

}
```

The system-owned catalog item is a deployment prerequisite — created during OSAC installation with unlocked parameters so the CaaS controller can set image, user_data, and network_attachments freely. The `BareMetalInstanceType` referenced in the `ClusterNodeSet` determines which host hardware profile is allocated by BMaaS.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

created during OSAC installation by who? The install scripts or one of the admins?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

explicitly added a proposal on how we can add it

- Backend decoupling: drop Metal3/Ironic examples in host labeling;
  state CaaS does not assume a specific BMaaS backend (carbonin)
- Cluster deletion: clarify BMIs are actively deleted by the ClusterOrder
  finalizer via BareMetalInstances.Delete — the HyperShift/CAPI/CAPA
  cascade cleans up K8s objects and Agent CRs only, BMIs are unknown to
  HyperShift (carbonin r3786801576)
- System-owned catalog item: specify it is created by automation
  (installer seed + controller reconcile), not by an admin, gated on
  CaaS deployed + BMaaS integrated + a BareMetalInstanceType registered
  (carbonin r3786851230)

Assisted-by: Claude Code <noreply@anthropic.com>
Signed-off-by: Riccardo Piccoli <rpiccoli@redhat.com>
@github-actions

github-actions Bot commented Aug 17, 2026 •

Copy link
Copy Markdown

AI Design Review: EP-198

Score: 8/8 | Verdict: PASS

Criterion Score Notes
Feasibility 2/2 Strong feasibility grounded in a validated PoC (OSAC-2817, ~6 min BMI provisioning, agent registration, worker join). Four hard blocking dependencies are clearly enumerated with Jira links and impact analysis. The design reuses existing infrastructure (fulfillment-service private API, assisted-service, HyperShift) and follows established osac-operator controller patterns. Idempotency is handled via list-before-create plus a DB uniqueness constraint (OSAC-3266). Retry/backoff strategy differentia
Testability 2/2 Comprehensive three-tier test plan covering unit tests (reconciler methods, phase transitions, idempotency, tenant isolation, RBAC), integration tests (kind cluster with mocked private API, agent registration simulation, scale-down flow, deletion cleanup), and E2E tests (full lifecycle including scale-up/down, cluster deletion, tenant isolation, pool removal verification). The design explicitly notes E2E environment constraints (BMaaS-capable hosts required) and acknowledges initial coverage may
Scope 2/2 Scope is large but well-defined. Non-goals are clearly stated (autoscaling, VM workers, static IP, network boot). The changes overview table provides a clear before/after mapping across all affected components. Cross-component impact is thoroughly documented: proto changes, CRD extensions, CLI additions, installer Helm values, AAP template changes, and Enclave Wizard updates. The no-coexistence decision for BareMetalPool removal increases deployment risk but is explicitly justified with a detail
Architecture 2/2 Sound architectural decisions throughout. Separation of BareMetalWorkerReconciler from the ClusterOrder controller keeps BM-specific logic isolated and the generic controller extensible. System tenant isolation via the existing builtin 'system' tenant is cleaner than label-based filtering (the design evolved past the approach described in the PR body). The DiskImage-to-ClusterVersion link is architecturally sound for the upgrade path. Controller reconciliation follows established patterns (final

Verdict: A thorough, well-structured design with strong PoC validation, clear dependency tracking, sound architectural separation, and comprehensive test coverage — the two unresolved implementation details (catalog item seeding mechanism, status patch strategy) are minor and appropriate to defer to implementation.

Feedback: Two open items should be resolved before implementation begins: (1) Decide whether the system-owned BareMetalInstanceCatalogItem is seeded by the installer chart or reconciled by the controller — the 'either way' framing is fine for design review but the implementation PR will need a clear answer, and the choice affects the installer's bare-metal integration toggle. (2) Resolve the status patch strategy for shared status.workers[] between the two controllers — optimistic concurrency handles correctness, but deciding between extending patchStatusWithRetry vs a separate patch path determines the code structure and testing approach. Additionally, consider adding explicit test cases for retry/backoff behavior (e.g., verify escalating backoff caps, verify attemptCount reset after MinHealthyDuration) since the retry logic is a key correctness property of the design.

Critical (0)

None.

Important (2)

  1. Two controllers (ClusterOrder controller and BareMetalWorkerReconciler) write to shared status.workers[] field — the design offers two options (extend patchStatusWithRetry or separate status patch path) but does not decide. This should be resolved before implementation to avoid structural rework.
  2. System-owned BareMetalInstanceCatalogItem seeding mechanism is explicitly left as an open item ('whether the seed lives in the installer chart or is reconciled entirely by the controller is an implementation choice'). The choice affects the installer Helm chart and the controller's self-healing behavior.

Suggestions (3)

  1. Add unit/integration test cases for retry/backoff behavior: verify escalating backoff caps per failure type, verify attemptCount persists across controller restarts, verify attemptCount resets after MinHealthyDuration (1 hour). The retry logic is a key correctness property that the current test plan does not explicitly cover.
  2. The PR description mentions 'label-based visibility filtering' (managed-by: caas) while the design uses system tenant isolation — consider updating the PR description to match the design's final approach to avoid reviewer confusion.
  3. Consider documenting the expected behavior when a ClusterVersion's disk_image reference is changed while workers are actively provisioning (e.g., mid-scale-up). The design covers the upgrade path (existing workers are not reprovisioned) but not a reference change during active provisioning.

Review cost

Model: claude-opus-4-6
Cost: $0.6946
Tokens: 10 in / 4.0k out
Cache: 392.6k read
Active time: 1m 38s
API calls: 0

@openshift-ci openshift-ci Bot added the lgtm label Aug 17, 2026
@openshift-ci

openshift-ci Bot commented Aug 17, 2026

Copy link
Copy Markdown
Contributor

[APPROVALNOTIFIER] This PR is APPROVED

This pull-request has been approved by: carbonin, rccrdpccl

The full list of commands accepted by this bot can be found here.

The pull request process is described here

Details Needs approval from an approver in each of these files:

Approvers can indicate their approval by writing /approve in a comment
Approvers can cancel approval by writing /approve cancel in a comment

@openshift-merge-bot
openshift-merge-bot Bot merged commit e7ea5f9 into osac-project:main Aug 17, 2026
5 checks passed
vladikr added a commit to vladikr/osac-enhancement-proposals that referenced this pull request Aug 24, 2026
- Use ClusterNetworkAttachment (proto-aligned) instead of ComputeInstance's
  NetworkAttachment type — different semantics for cluster vs per-NIC
- Add hook stability contract (document variables, compatibility test)
- CUDN naming uses ClusterOrder UID (not cluster_name) for uniqueness
- Reject non-empty SecurityGroupRefs in Phase 1
- Add NAD readiness wait and VM sizing validation
- Clarify Phase 1/2 boundary: triggers, migration path, production-grade
- Add status.workers[] for CAPK VM inventory visibility (same pattern
  as BMI workers in PR osac-project#198)

Assisted-by: Claude Code <noreply@anthropic.com>
Signed-off-by: Vladik Romanovsky <vromanso@redhat.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants