Skip to content

OSAC-1029: Per-service networking EPs and unified EP restructuring - #107

Merged
openshift-merge-bot[bot] merged 58 commits into
osac-project:mainfrom
danmanor:OSAC-1029/restructure-unified-networking
Jul 27, 2026
Merged

openshift-merge-bot[bot] merged 58 commits into
osac-project:mainfrom
danmanor:OSAC-1029/restructure-unified-networking

Conversation

@danmanor

@danmanor danmanor commented Jul 8, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • Restructure unified networking EP/PRD (rename files, fix contradictions, add HostType)
  • Create per-service networking EPs with detailed flows for VMaaS, CaaS, and BMaaS
  • Create default-networking EP for tenant defaults and auto ExternalIP/NATGateway
  • All PRDs are user-facing only (no implementation details — no proto fields, component names, or internal mechanisms)

Changes

Unified networking EP/PRD (restructured):

  • Two-phase auto-provisioning lifecycle with ExternalIPAttachment controller preconditions per target type
  • DHCP-based host networking: fabric DHCP server assigns IPs, operator discovers them after assignment
  • IP discovery per service: VMaaS from KubeVirt VMI feedback controller, BMaaS mechanism is open question (OQ#4), CaaS from Agent CR
  • Resource Status section with ComputeNetworkAttachmentStatus and BareMetalNetworkAttachmentStatus protos for IP visibility
  • Per-resource attachment types: ComputeNetworkAttachment, BareMetalNetworkAttachment, ClusterNetworkAttachment
  • Proto field numbers aligned with actual codebase (ComputeInstance: 18/19/20, Cluster: 9, BaremetalInstance: 8)
  • NATGateway controller preconditions (ExternalIP must be Allocated)
  • Auto-provisioned resource cleanup: phased requeue with auto-provisioned-for label
  • Fabric manager is switch-side only (add port to V-Net, no IPAM, no host config)
  • "Per deployment" model (not "per region")

VMaaS networking (new):

  • ComputeNetworkAttachment with primary field for multi-NIC (field 18 on ComputeInstanceSpec)
  • ComputeNetworkAttachmentStatus with ip_address for IP visibility (from KubeVirt VMI feedback)
  • Optional compute_network_attachments with tenant defaults; dual-field migration from old field 14
  • Auto ExternalIP / NATGateway (separate DB transaction, best-effort for NATGateway)
  • ExternalIPAttachment precondition uses compute_network_attachment_statuses IP (not VirtualMachineReference)

CaaS networking (new):

  • Single subnet per cluster — one ClusterNetworkAttachment with subnet + security_groups only (field 9 on ClusterSpec)
  • fabricInterface resolved per node set from HostType, stored on node set definition (NOT on attachment)
  • Operator handles agent selection + reconcileNetworking (switch-side config) before provisioning
  • MetalLB dynamically allocates VIPs from IPAddressPool, template discovers and writes to ClusterOrder status
  • MetalLB IPAddressPool created by k8s_manager at subnet creation time (shared prerequisite)
  • VIP feedback loop for ExternalIPAttachment activation
  • NodeSetStatus/AgentStatus CRD for per-agent data
  • DHCP for agent host networking (no NMState/static config needed)
  • NATGateway NOT cleaned up on cluster deletion (shared per-VN resource)
  • Open questions: MetalLB/DHCP CIDR partitioning (OQ#7), DNS record target IP (OQ#8)

BMaaS networking (new):

  • BareMetalNetworkAttachment with tenant-specified interface + primary (field 8 on BaremetalInstanceSpec)
  • BareMetalNetworkAttachmentStatus with ip_address (field name: network_attachment_statuses)
  • Interface validated against HostType's NetworkInterface list; lifecycle interfaces rejected
  • DHCP for host networking; runtime IP discovery mechanism is open question (OQ#4 — Metal3 inspection IP is stale after network reconfiguration)
  • Auto-provisioned resources labeled with osac.openshift.io/auto-provisioned and auto-provisioned-for
  • Default interface selection: first interface with role fabric from HostType (when attachments omitted)

Default networking (new):

  • NetworkClass defaults field for tenant onboarding
  • fulfillment-service creates default VN/Subnet/SG at tenant onboarding via its own API
  • DefaultNetworkingReady condition on Tenant — set to True with reason DefaultsNotConfigured when no defaults configured (tenant still becomes READY)
  • Auto ExternalIP pool selection and auto NATGateway reuse logic
  • NATGateway NOT auto-cleaned on resource deletion (shared per-VN)

Key architectural decisions

Decision Resolution
Host-side IP assignment DHCP from fabric DHCP server; OSAC discovers IP after assignment
Host-side networking config DHCP handles IP, gateway, prefix, DNS automatically — no template config needed
IP discovery VMaaS: KubeVirt VMI feedback controller; BMaaS: open question OQ#4; CaaS: Agent CR status
IP status visibility VMaaS: ComputeNetworkAttachmentStatus; BMaaS: BareMetalNetworkAttachmentStatus; CaaS: api_endpoint/ingress_endpoint
CaaS VIPs MetalLB dynamic allocation from IPAddressPool, template discovers
Fabric manager role Switch-side only (add port to V-Net, no IPAM, no host config)
MetalLB IPAddressPool k8s_manager creates at subnet creation time (shared prerequisite)
CaaS subnet model One subnet per cluster, all node sets share it
CaaS fabricInterface System-resolved per node set from HostType, stored on node set definition
NATGateway cleanup NOT auto-cleaned (shared per-VN resource)
NATGateway auto ExternalIP Separate ExternalIP, separate DB transaction, best-effort
Default networking creation fulfillment-service creates defaults via its own API (not operator)
Auto-provisioned cleanup Phased requeue with auto-provisioned-for label for orphan detection
Proto field numbers ComputeInstance: 18/19/20; Cluster: 9; BaremetalInstance: 8 (aligned with codebase)

Open questions (unresolved)

OQ EP Issue
OQ#4 BMaaS Runtime IP discovery after network reconfiguration — Metal3 inspection IP is stale
OQ#7 CaaS MetalLB/DHCP CIDR partitioning — IP collision risk between VIPs and agent IPs
OQ#8 CaaS DNS record target IP — MetalLB VIP vs ExternalIP affects bootstrap sequencing
OQ (all) All NATGateway reuse of Failed/Deleting gateway — should Deleting be treated as "does not exist"?

Jira

@openshift-ci-robot

openshift-ci-robot commented Jul 8, 2026 •

Copy link
Copy Markdown

@danmanor: This pull request references OSAC-1029 which is a valid jira issue.

Warning: The referenced jira issue has an invalid target version for the target branch this PR targets: expected the feature to target the "5.0.0" version, but no target version was set.

Details

In response to this:

Summary

  • Restructure unified networking EP/PRD (rename files, fix contradictions, add HostType)
  • Create per-service networking EPs with detailed flows for VMaaS, CaaS, and BMaaS
  • All PRDs are user-facing only (no implementation details)

Changes

Unified networking EP/PRD (restructured):

  • README.md → design.md, PRD moved into same directory as prd.md
  • Fixed contradictions: HostType interface validation (not "from BMaaS"), DHCP-or-static, CaaS BM-only for v0.2, dispatcher scope clarification
  • Added HostType + NetworkInterface to API Extensions, fabric_interface on ClusterNetworkAttachment
  • PRD converted to ai-workflows template format (FR-1 through FR-7, acceptance criteria checkboxes)

VMaaS networking (new):

  • ComputeNetworkAttachment with primary field for multi-NIC
  • Optional network_attachments with tenant defaults
  • Auto ExternalIP / NATGateway
  • Operator resolves subnet → namespace, template creates VM (no networking logic in template)

CaaS networking (new):

  • ClusterNetworkAttachment with node_set + system-resolved fabric_interface from HostType
  • Operator handles agent selection + network attachment before provisioning
  • Step collections (cluster_infra, external_access) removed
  • CaaS BM-only for v0.2 (VM node sets deferred)
  • VIP feedback loop for ExternalIPAttachment Pending → Ready

BMaaS networking (new):

  • BareMetalNetworkAttachment with tenant-specified interface + primary
  • Two-operator architecture: bare-metal-fulfillment-operator adds reconcileNetworking phase
  • Interface validated against HostType's NetworkInterface list
  • IP address feedback for ExternalIPAttachment DNAT

Jira

🤖 Generated with Claude Code

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the openshift-eng/jira-lifecycle-plugin repository.

@openshift-ci
openshift-ci Bot requested review from jhernand and mennyaboush July 8, 2026 12:15
@openshift-ci openshift-ci Bot added the approved label Jul 8, 2026
@coderabbitai

coderabbitai Bot commented Jul 8, 2026 •

Copy link
Copy Markdown

Review Change Stack

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review

Walkthrough

The PR adds BMaaS, CaaS, and VMaaS networking design and PRD documents, revises unified networking contracts and lifecycle descriptions, and defines tenant default networking with automatic external access and NAT behavior.

Changes

Unified Networking Base Updates

Layer / File(s) Summary
Unified networking design contracts
enhancements/unified-networking/design.md
Updates API scope, deployment-level resource ownership, HostType interface resolution, cluster attachment shape, VIP feedback, and auto-provisioning lifecycle behavior.
Unified networking PRD structure
enhancements/unified-networking/prd.md
Reorganizes metadata, gaps, goals, requirements, acceptance criteria, and dependencies.

BMaaS Networking Proposal

Layer / File(s) Summary
BMaaS attachment contracts
enhancements/bmaas-networking/design.md
Defines HostType-based interface selection, BareMetalNetworkAttachment, validation, immutability, and fulfillment-service propagation.
BMaaS networking reconciliation
enhancements/bmaas-networking/design.md
Defines switch-side attachment configuration, networking readiness gating, IP discovery, automatic external access and NAT, and deletion ordering.
BMaaS operations and validation
enhancements/bmaas-networking/design.md
Documents security, failures, observability, testing, lifecycle criteria, support procedures, infrastructure, and dependencies.
BMaaS PRD requirements
enhancements/bmaas-networking/prd.md
Defines requirements, acceptance criteria, assumptions, risks, and open questions.

CaaS Networking Proposal

Layer / File(s) Summary
CaaS networking contracts
enhancements/caas-networking/design.md
Defines HostType metadata, cluster API and operator fields, validation, endpoint status, and template wiring.
CaaS provisioning and VIP feedback
enhancements/caas-networking/design.md
Defines agent selection, bare-metal attachment provisioning, template endpoint allocation, status feedback, cleanup, and failure handling.
CaaS PRD requirements
enhancements/caas-networking/prd.md
Defines requirements, acceptance criteria, assumptions, dependencies, risks, and open questions.

VMaaS Networking Proposal

Layer / File(s) Summary
VMaaS attachment and template contracts
enhancements/vmaas-networking/design.md
Defines ComputeNetworkAttachment, primary validation, automatic-access modes, compatibility handling, and multi-NIC template wiring.
VMaaS provisioning workflow
enhancements/vmaas-networking/design.md
Defines controller annotation, KubeVirt interface creation, automatic external access and NAT, and deletion semantics.
VMaaS operations and compatibility
enhancements/vmaas-networking/design.md
Documents failure handling, observability, testing, graduation, version skew, support, and infrastructure prerequisites.
VMaaS PRD requirements
enhancements/vmaas-networking/prd.md
Defines requirements, acceptance criteria, assumptions, risks, and open questions.

Default Networking Proposal

Layer / File(s) Summary
Default networking onboarding
enhancements/default-networking/design.md
Defines NetworkClass-driven defaults, tenant readiness gating, omitted attachment resolution, and automatic external access fields.
Default and auto-provisioned lifecycle
enhancements/default-networking/design.md
Defines labels, validation, cleanup ordering, NATGateway reuse, failure handling, testing, upgrade behavior, support, and infrastructure deliverables.

Estimated code review effort: 4 (Complex) | ~60 minutes

Possibly related PRs

Suggested reviewers: AlonaKaplan, rgolangh, wgordon17, avishayt, zszabo-rh

🚥 Pre-merge checks | ✅ 11
✅ Passed checks (11 passed)
Check name Status Explanation
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
No-Hardcoded-Secrets ✅ Passed Touched docs contain no hardcoded credentials, embedded-auth URLs, private keys, or secret/token/password string literals.
No-Weak-Crypto ✅ Passed Scanned the networking docs and repo text; no MD5/SHA1/DES/RC4/3DES/Blowfish/ECB, custom crypto, or secret-comparison code was introduced.
No-Injection-Vectors ✅ Passed Only markdown design/PRD docs changed; scans found no unsafe code patterns like eval/exec, yaml.load, shell=True, or DOM injection.
Container-Privileges ✅ Passed HEAD only changes a markdown design doc; no YAML/manifests or privilege settings (privileged, hostNetwork, SYS_ADMIN, allowPrivilegeEscalation) are present.
No-Sensitive-Data-In-Logs ✅ Passed Touched files are design/PRD docs; observability sections list generic event names only, with no secrets, PII, or hostnames in log payloads.
Ai-Attribution ✅ Passed Recent PR commits use Assisted-by: Claude Code trailers; no Co-Authored-By appears in the OSAC-1029 commit series.
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately summarizes the main change: per-service networking epics plus the unified networking restructuring.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions

github-actions Bot commented Jul 8, 2026

Copy link
Copy Markdown

AI EP Review: EP-107

Score: 9/10 | Verdict: PASS

Criterion Score Notes
What 2/2 All four PRDs clearly describe user-observable outcomes. Personas are identified across all PRDs (Tenant User, Tenant Admin, Cloud Infrastructure Admin, Cloud Provider Admin). Services are named (BMaa
Why 2/2 Each PRD provides concrete justification. BMaaS: 'Provisioning bare-metal servers requires manual switch configuration outside the OSAC API' with consequences (tenants forced to out-of-band documentat
How 1/2 Acceptance criteria are specific and testable across all PRDs (checkbox format, user-observable). BMaaS has NFR targets (synchronous allocation, 2 min/interface). However, design leakage weakens some
Task 2/2 All four documents are proper enhancements adding new networking capabilities to the OSAC platform — not bug fixes or tasks. Each PRD is structured as a product requirements document with problem stat
Size 2/2 Each PRD is individually well-scoped around a coherent set of tightly coupled capabilities. BMaaS: network attachments + auto external access + cleanup — provisioning without cleanup creates orphaned

Verdict: Strong set of PRDs with clear user outcomes, concrete justification, well-scoped per service type, and testable acceptance criteria; weakened by design leakage in some functional requirements (BMaaS FR-8/FR-11) and inconsistent NFR coverage across PRDs.

Feedback: Remove implementation details from PRD requirements: BMaaS FR-8 should say 'network connectivity is configured for each attachment before provisioning begins' rather than describing switch port configuration, fabric server identification, and DHCP/static allocation. BMaaS FR-11 describes an internal configuration parameter that belongs in the design doc, not the PRD. Add quantitative success metrics to the CaaS and VMaaS PRDs (BMaaS sets a good example with provisioning time and success rate targets), and add at least one NFR to the unified networking PRD to establish baseline performance expectations for the cross-service networking model.

Critical (0)

None.

Important (3)

  1. BMaaS PRD FR-8 contains design leakage: 'identifies the physical server in the network fabric, adds the server's interface to the subnet's network segment, and allocates an IP address (DHCP or static)' describes implementation behavior not observable by a PM using the product. Rewrite as user outcome: 'network connectivity is established for each attachment before OS provisioning begins; the server can communicate on the attached subnets.'
  2. BMaaS PRD FR-11 describes an internal configuration concern ('distinct configuration parameter for identifying which network fabric automation to use') that belongs in the design doc. The PRD should state the user-observable outcome: 'the system avoids naming confusion between fabric management configuration and networking resources.'
  3. Unified Networking PRD lacks any non-functional requirements ('No non-functional requirements were specified'), making it impossible to verify performance, scalability, or reliability expectations for the foundational networking layer.

Suggestions (3)

  1. Add quantitative success metrics to CaaS and VMaaS PRDs to match BMaaS's example (provisioning time targets, success rate thresholds).
  2. VMaaS PRD is missing Cloud Infrastructure Admin and Tenant Admin persona stories — consider adding stories for admin visibility into multi-NIC configurations and network troubleshooting.
  3. CaaS PRD FR-7 and FR-8 could be more user-focused: 'the system prepares hosts and network connectivity before provisioning begins' rather than detailing host selection/reservation mechanics.

Review cost

Model: claude-opus-4-6
Cost: $1.1013
Tokens: 6.1k in / 10.5k out
Cache: 202.2k read
Active time: 3m 27s
API calls: 0

@github-actions github-actions Bot added the rfe-creator-auto-reviewed EP was reviewed by AI label Jul 8, 2026
@github-actions

github-actions Bot commented Jul 8, 2026

Copy link
Copy Markdown

AI Design Review: EP-107

Score: 8/8 | Verdict: PASS

Criterion Score Notes
Feasibility 2/2 All three designs are grounded in existing architecture (two-operator model, AAP integration, CR patterns). Proto definitions, Go structs, CEL validation rules, and component responsibility matrices a
Testability 2/2 Each design includes unit tests (validation logic, primary resolution, pool selection, interface resolution), integration/E2E tests (full lifecycle including create, auto-provision, delete, cleanup),
Scope 2/2 Each design has clear non-goals separating VMaaS/CaaS/BMaaS concerns. Explicit deferrals (DNS API, VM-based CaaS nodes, multi-NIC CaaS, NIC bonding) prevent scope creep. Dependencies are enumerated wi
Architecture 2/2 The unified networking API maintains consistent resource model across all three service types. Resource-specific attachment messages (ComputeNetworkAttachment, BareMetalNetworkAttachment, ClusterNetwo

Verdict: A comprehensive, well-structured set of design documents that introduces consistent networking support across VMaaS, CaaS, and BMaaS with sound architectural principles, thorough test plans, and clearly defined scope boundaries.

Feedback: The NATGateway reuse-regardless-of-state behavior is an open question across all three designs and should be resolved centrally before implementation begins, since it directly affects user experience in failure scenarios. Consider consolidating the duplicated HostType proto definition (repeated in BMaaS and CaaS designs) by referencing the unified-networking HostType section to reduce drift risk. Adding mermaid sequence diagrams for the more complex cross-component flows (especially the CaaS VIP feedback loop spanning template, ClusterOrder, feedback controller, fulfillment-service, and ExternalIPAttachment controller) would significantly improve reviewability.

Critical (0)

None.

Important (3)

  1. NATGateway auto-reuse 'regardless of state' is documented as an open question in all three designs (VMaaS FR-5, CaaS FR-4/FR-7, BMaaS FR-5) but affects core user experience — a tenant creating a resource with nat_gateway_mode=AUTO on a VN with a Failed NATGateway gets silently broken outbound connectivity with no system-level recovery. This should be resolved as a cross-cutting decision before implementation.
  2. Several work items appear as 'GAP' (not tracked in Jira) in the dependency tables: BMaaS has 5 GAPs (CRD update, mutateBMI, IP feedback, networkClass rename, HostType interfaces), CaaS has 5 GAPs (HostType, step collection removal, NETWORK_STEPS_COLLECTION removal, agent selection, interface resolution). These represent real implementation work that could be missed during planning.
  3. The CaaS VIP feedback loop (template writes VIPs to ClusterOrder status -> feedback controller fires Signal RPC -> fulfillment-service syncs to Cluster -> ExternalIPAttachment controller reads endpoints -> creates DNAT) spans 5 components with failure possible at each hop. The design acknowledges this complexity as a drawback but the support procedure for 'VIPs not synced' only covers partial diagnosis. Consider adding a reconciliation or polling fallback mechanism.

Suggestions (3)

  1. The HostType proto definition and interface role convention table are duplicated verbatim in BMaaS design.md, CaaS design.md, and unified-networking design.md. Consider having BMaaS and CaaS reference the canonical definition in unified-networking to avoid drift.
  2. The open questions about capacity exhaustion behavior (return API error vs. create Failed resource) and NATGateway state-aware reuse appear identically across all three service-specific designs. These are API-level design decisions that should be resolved once in a shared location rather than independently per service type.
  3. Consider adding mermaid sequence diagrams for the CaaS VIP feedback loop and the BMaaS IP address feedback flow — these multi-component interactions are the hardest parts of the design to follow from prose alone.

Review cost

Model: claude-opus-4-6
Cost: $1.0359
Tokens: 6.0k in / 3.3k out
Cache: 159.8k read
Active time: 1m 19s
API calls: 0

@github-actions

github-actions Bot commented Jul 8, 2026

Copy link
Copy Markdown

AI EP Review: EP-107

Score: 9/10 | Verdict: PASS

Criterion Score Notes
What 2/2 All four PRDs clearly describe user-observable outcomes with specific personas (Tenant User, Tenant Admin, Cloud Infrastructure Admin, Cloud Provider Admin) and affected services. User stories are con
Why 2/2 Problem statements name specific pains (manual switch configuration, out-of-band documentation, zero tenant control, sequential API calls requiring infra admin coordination). Success metrics are prese
How 1/2 Mostly user-facing but some design leakage. BMaaS PRD FR-8 references 'switch port configuration' (implementation term; user-facing would be 'network connectivity configuration'). FR-11 describes 'a s
Task 2/2 All acceptance criteria across all four PRDs are user-verifiable: create resources with specific options, verify errors on invalid inputs, verify cleanup on deletion, verify status visibility. NFRs in
Size 2/2 Each PRD is well-scoped to networking for a single service type. Capabilities within each PRD are tightly coupled (network attachments + validation + auto external access + cleanup). The unified PRD s

Verdict: Strong set of PRDs with clear user-facing requirements, concrete justification, well-scoped features, and testable acceptance criteria; minor design leakage in BMaaS PRD (switch port terminology, fabric manager config detail) prevents a perfect score.

Feedback: Replace implementation-specific terminology in BMaaS PRD: FR-8's 'switch port configuration' should be 'network connectivity configuration' and FR-11 should describe the user-facing outcome (avoiding name collision with networking resources) rather than the mechanism (static configuration string). Also resolve the auto NAT gateway reuse-regardless-of-state design decision that appears as both a risk and open question in all three service-specific PRDs — either commit to a direction or explicitly mark it as a blocking open question. Finally, VMaaS Non-Goals incorrectly states bare-metal multi-interface support is 'out of scope' when the BMaaS PRD explicitly covers it — reword to 'covered in BMaaS Networking PRD'.

Critical (0)

None.

Important (2)

  1. BMaaS PRD FR-8 and FR-11 contain design leakage: 'switch port configuration' is implementation terminology (should be 'network connectivity configuration'), and FR-11 describes a 'static configuration string set at deployment time' which exposes internal mechanics rather than user-observable behavior.
  2. CaaS PRD NFR-1 states 'Endpoint addresses are available in cluster status during provisioning, not minutes later' but the design's VIP feedback loop (template -> ClusterOrder status -> Signal RPC -> fulfillment-service -> Cluster) is inherently asynchronous — the NFR may be inaccurate or misleading about the actual timing.

Suggestions (3)

  1. VMaaS Non-Goals section says 'Multiple network interfaces for bare-metal servers (bare-metal multi-interface support is out of scope)' which contradicts the BMaaS PRD that explicitly supports multi-NIC BM. Reword to 'covered in BMaaS Networking PRD' to avoid confusion.
  2. The auto NAT gateway reuse-regardless-of-state design appears in all three service-specific PRDs as both a risk and an open question — consider resolving this cross-cutting design decision centrally before finalizing the PRDs.
  3. BMaaS PRD Goal Document the enhancement proposal process #6 in section 2.1 ('The system configures switch ports for each network attachment') should use user-facing language consistent with the other PRDs (e.g., 'The system configures network connectivity for each network attachment').

Review cost

Model: claude-opus-4-6
Cost: $1.1865
Tokens: 6.1k in / 7.0k out
Cache: 169.0k read
Active time: 2m 15s
API calls: 0

@github-actions

github-actions Bot commented Jul 8, 2026

Copy link
Copy Markdown

AI EP Review: EP-107

Score: 8/10 | Verdict: PASS

Criterion Score Notes
What 2/2 All four PRDs clearly describe user-observable capabilities with well-structured user stories (Tenant User, Tenant Admin, Cloud Infrastructure Admin, Cloud Provider Admin). Each PRD names concrete act
Why 1/2 Each PRD describes concrete pain: BMaaS requires 'manual switch configuration outside the OSAC API' and 'manual coordination with infrastructure administrators'; CaaS has 'zero tenant control' and 'co
How 2/2 Requirements are specific and measurable. BMaaS has 13 functional requirements plus 2 NFRs with numeric targets (95% switch port success rate, 2 min per interface). CaaS has 12 FRs. VMaaS has 7 FRs pl
Task 2/2 These are clearly platform enhancements adding new networking capabilities: multi-NIC support, auto external IP/NAT gateway provisioning, unified networking API across service types, host type interfa
Size 1/2 Each individual PRD is well-scoped and internally cohesive — within each service-type PRD, capabilities are tightly coupled (auto external IP needs basic attachment, cleanup needs creation, multi-NIC

Verdict: A well-structured set of networking PRDs with clear user outcomes, specific requirements, and thorough acceptance criteria, held back slightly by gap-descriptive (rather than impact-quantified) business justification and the bundling of three independently deliverable service-type PRDs into a single submission.

Feedback: Strengthen the WHY in each PRD by quantifying the impact of the current gaps — e.g., how many tenants are blocked, how much manual coordination time is required per BM provisioning, or tie explicitly to a strategic goal like 'CaaS adoption is blocked for tenants requiring network isolation.' Consider splitting the per-service PRDs (BMaaS, CaaS, VMaaS) into separate PRs since they are independently deliverable and serve different service types, which would make each easier to review, prioritize, and track independently.

Critical (0)

None.

Important (2)

  1. WHY sections across all PRDs describe gaps but don't quantify business impact or tie to strategic goals — e.g., BMaaS says 'manual switch configuration outside the OSAC API' but doesn't say why that matters (tenant churn? blocked adoption? SLA violations?). Adding one sentence connecting the gap to a measurable consequence would elevate each from Y=1 to Y=2.
  2. PR bundles 3 independently deliverable service-type PRDs (BMaaS, CaaS, VMaaS) that could each ship on their own. The unified networking PRD is the shared foundation, but CaaS networking doesn't depend on VMaaS networking being delivered simultaneously. Splitting into separate PRs per service type would enable independent prioritization and review.

Suggestions (3)

  1. Some functional requirements contain implementation-flavored language that approaches design leakage: BMaaS FR-8 mentions 'identifies the physical server in the network fabric, adds the server's interface to the subnet's network segment'; CaaS FR-9 mentions 'performs DNS record creation'. Reframing these as pure user outcomes (e.g., 'network connectivity is available before provisioning begins') would keep the PRD firmly in WHAT territory.
  2. The unified networking PRD's non-functional requirements section says 'No non-functional requirements were specified in the original document.' — consider adding NFRs for the shared resource model (e.g., subnet creation latency, cross-service consistency SLAs) since the per-service PRDs already have their own NFRs.
  3. All three service-specific PRDs share the same open question about auto NAT gateway state checking (should it check existing state before reusing?). Consider resolving this cross-cutting question once in the unified PRD rather than leaving it open in three separate places.

Review cost

Model: claude-opus-4-6
Cost: $1.2584
Tokens: 6.1k in / 6.6k out
Cache: 329.6k read
Active time: 2m 22s
API calls: 0

@openshift-ci-robot

openshift-ci-robot commented Jul 8, 2026 •

Copy link
Copy Markdown

@danmanor: This pull request references OSAC-1029 which is a valid jira issue.

Warning: The referenced jira issue has an invalid target version for the target branch this PR targets: expected the feature to target the "5.0.0" version, but no target version was set.

Details

In response to this:

Summary

  • Restructure unified networking EP/PRD (rename files, fix contradictions, add HostType)
  • Create per-service networking EPs with detailed flows for VMaaS, CaaS, and BMaaS
  • All PRDs are user-facing only (no implementation details)

Changes

Unified networking EP/PRD (restructured):

  • README.md → design.md, PRD moved into same directory as prd.md
  • Fixed contradictions: HostType interface validation (not "from BMaaS"), DHCP-or-static, CaaS BM-only for v0.2, dispatcher scope clarification
  • Added HostType + NetworkInterface to API Extensions, fabric_interface on ClusterNetworkAttachment
  • PRD converted to ai-workflows template format (FR-1 through FR-7, acceptance criteria checkboxes)

VMaaS networking (new):

  • ComputeNetworkAttachment with primary field for multi-NIC
  • Optional network_attachments with tenant defaults
  • Auto ExternalIP / NATGateway
  • Operator resolves subnet → namespace, template creates VM (no networking logic in template)

CaaS networking (new):

  • ClusterNetworkAttachment with node_set + system-resolved fabric_interface from HostType
  • Operator handles agent selection + network attachment before provisioning
  • Step collections (cluster_infra, external_access) removed
  • CaaS BM-only for v0.2 (VM node sets deferred)
  • VIP feedback loop for ExternalIPAttachment Pending → Ready

BMaaS networking (new):

  • BareMetalNetworkAttachment with tenant-specified interface + primary
  • Two-operator architecture: bare-metal-fulfillment-operator adds reconcileNetworking phase
  • Interface validated against HostType's NetworkInterface list
  • IP address feedback for ExternalIPAttachment DNAT

Jira

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features

  • Added unified networking support for multi-interface provisioning across VM, cluster, and bare-metal workflows.

  • Introduced optional network attachments with primary interface selection, auto-filled defaults, and validation rules.

  • Added automatic external IP and NAT gateway setup for inbound and outbound access, plus status visibility for allocated addresses.

  • Expanded host/interface discovery so networking can be matched to available hardware interfaces before provisioning.

  • Documentation

  • Added and updated product and design docs covering networking behavior, lifecycle, cleanup, and acceptance criteria.

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the openshift-eng/jira-lifecycle-plugin repository.

@github-actions

github-actions Bot commented Jul 8, 2026

Copy link
Copy Markdown

AI Design Review: EP-107

Score: 8/8 | Verdict: PASS

Criterion Score Notes
Feasibility 2/2 All three designs specify concrete API extensions (protobuf, Go CRD structs, CEL validation), reconciliation phase ordering, and component responsibility tables. Dependencies are tracked with Jira IDs
Testability 2/2 Each design includes unit test strategy with specific cases, E2E integration scenarios covering happy path and error conditions, and a 'tricky test cases' section for edge cases (multi-NIC primary on
Scope 2/2 Each document has explicit goals and non-goals with clear service-type boundaries. Version scoping is disciplined (CaaS v0.2 BM-only, single attachment per node set; VMaaS dual-field migration timelin
Architecture 2/2 Uniform networking API across all service types with per-resource-type attachment messages (ComputeNetworkAttachment, BareMetalNetworkAttachment, ClusterNetworkAttachment) providing type-specific vali

Verdict: A comprehensive, well-structured set of per-service networking designs that build cleanly on the unified networking architecture, with honest dependency tracking, thorough test plans, and clear graduation criteria.

Feedback: The auto NATGateway reuse-regardless-of-state design is the weakest point across all three documents -- consider resolving open question #1 (check state before reusing) before merging, as it affects user experience in all service types identically. The 5 untracked GAPs in BMaaS and 5 in CaaS should get Jira tickets before implementation begins to prevent scope creep. The CaaS VIP feedback loop (template -> ClusterOrder status -> Signal RPC -> fulfillment-service -> Cluster -> ExternalIPAttachment controller) spans 5 components and is the highest-risk integration path; consider adding a sequence diagram and explicit timeout/retry semantics for each hop.

Critical (0)

None.

Important (3)

  1. Auto NATGateway reuse-regardless-of-state is an open question in all 3 service designs but has real user impact (VMs/clusters/BM servers silently get broken outbound connectivity). This should be resolved before implementation, not deferred.
  2. 10 work items across BMaaS (5) and CaaS (5) are marked as untracked GAPs with no Jira tickets, risking scope underestimation during planning.
  3. The CaaS VIP feedback loop spans 5 components (template -> ClusterOrder -> feedback controller -> fulfillment-service -> ExternalIPAttachment controller) with no explicit timeout, retry, or circuit-breaker semantics documented for each hop.

Suggestions (3)

  1. Add a sequence diagram (mermaid) for the CaaS VIP feedback loop and the BMaaS IP address feedback flow to make the multi-component coordination visually clear.
  2. Consider documenting what happens when the lifecycle interface appears in network_attachments -- FR-3 in BMaaS PRD says validation should catch invalid interfaces but the design only documents that lifecycle 'should not appear', without specifying whether validation explicitly rejects it or relies on convention.
  3. The unified-networking PRD restructuring (moving from separate directory to prd.md inside unified-networking/) loses the YAML frontmatter -- consider whether downstream tooling depends on that metadata.

Review cost

Model: claude-opus-4-6
Cost: $0.8882
Tokens: 6.0k in / 3.0k out
Cache: 287.3k read
Active time: 1m 13s
API calls: 0

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 8

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@enhancements/bmaas-networking/design.md`:
- Around line 126-141: The tenant-facing interface validation for
BareMetalNetworkAttachment needs to explicitly reject lifecycle/BMC ports, not
just verify the name exists in HostType.interfaces. Update the validation logic
that resolves the HostType from the catalog_item/template so it filters out or
errors on interfaces marked as lifecycle, and ensure network_attachments only
accept tenant-attachable interfaces such as fabric, management, and storage.

In `@enhancements/bmaas-networking/prd.md`:
- Around line 114-117: FR-12 only defines cleanup for auto-provisioned external
IP resources, but it does not state what happens to NAT gateways when a server
is deleted. Update the deletion cleanup section in the PRD to explicitly define
ownership for NAT gateway resources referenced by FR-7 and the networking flow,
using the same AUTO vs reused/shared distinction as the external IP rules. Make
it clear that server deletion deletes only system-created NAT gateways and
leaves reused/shared gateways intact, or specify the alternative policy
consistently across the networking requirements.

In `@enhancements/caas-networking/design.md`:
- Around line 124-127: Persist the parent ClusterOrder record before allocating
dependent networking resources, or add rollback for any allocations made first.
Update the provisioning flow described around external_ip_mode and
nat_gateway_mode so the Cluster record/ClusterOrder exists before creating
ExternalIPs, ExternalIPAttachments, or NATGateway, and use the existing
ClusterOrder/network_attachments flow to keep ownership clear if a later step
fails.
- Around line 403-407: The Auto-Provisioned Resource Lifecycle summary is
missing NATGateway from the cleanup contract, so update the lifecycle
description to include NATGateway alongside the other auto-provisioned
resources. Make sure the cleanup order and permanent-failure behavior in the
design section explicitly reference NATGateway so it is removed with the cluster
and not left orphaned; use the existing lifecycle terminology in the
Auto-Provisioned Resource Lifecycle section to keep it consistent.

In `@enhancements/caas-networking/prd.md`:
- Around line 66-67: Update the FR-1 requirement in the networking PRD so
nodeSet is required for v0.2 instead of optional, and align the wording with
FR-11’s one-attachment-per-node-set, bare-metal-only scope. Make sure the
cluster creation/network attachment description and any related references to
nodeSet in the same requirements block are consistent so interface selection is
no longer ambiguous for mixed HostTypes.
- Around line 78-79: The NAT gateway reuse behavior in FR-4 should be narrowed
so cluster creation only reuses a healthy gateway. Update the requirement around
automatic NAT gateway provisioning to reference the relevant cluster creation
flow and NAT gateway handling so that only Ready resources are reused, and add
an explicit rejection or creation-failure path for existing Failed or Deleting
gateways.

In `@enhancements/unified-networking/prd.md`:
- Around line 337-338: Update the PRD checklist in this section to match the
design’s per-resource attachment contracts: replace the shared
`NetworkAttachment` wording with the specific `ComputeNetworkAttachment`,
`ClusterNetworkAttachment`, and `BareMetalNetworkAttachment` types, and refer to
`HostType` instead of `BaremetalInstanceTemplate`. Keep the language aligned
with the resource-specific API shapes so the requirements mirror the design
exactly.

In `@enhancements/vmaas-networking/prd.md`:
- Around line 71-74: Clarify the Auto NAT Gateway failure contract in the PRD:
FR-5 currently says `--nat-gateway=auto` reuses an existing gateway regardless
of state, which conflicts with the outbound connectivity promise. Update the
`Auto NAT Gateway` requirement so the behavior is explicit when an existing
gateway is failed, deleting, or otherwise unusable—either require the
create/retry flow to fail in that case, or state that degraded connectivity is
acceptable. Use the `FR-5` requirement text and the `--nat-gateway=auto`
acceptance language as the places to revise.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository: osac-project/coderabbit/.coderabbit.yaml

Review profile: ASSERTIVE

Plan: Enterprise

Run ID: 4148b28a-cd27-4c4f-9598-fd73bed371ed

📥 Commits

Reviewing files that changed from the base of the PR and between e60803c and c2693a3.

📒 Files selected for processing (8)
  • enhancements/bmaas-networking/design.md
  • enhancements/bmaas-networking/prd.md
  • enhancements/caas-networking/design.md
  • enhancements/caas-networking/prd.md
  • enhancements/unified-networking/design.md
  • enhancements/unified-networking/prd.md
  • enhancements/vmaas-networking/design.md
  • enhancements/vmaas-networking/prd.md

Comment thread enhancements/bmaas-networking/design.md Outdated
Comment thread enhancements/OSAC-1437-bmaas-networking/prd.md
Comment thread enhancements/caas-networking/design.md Outdated
Comment thread enhancements/OSAC-1436-caas-networking/design.md
Comment thread enhancements/caas-networking/prd.md Outdated
Comment thread enhancements/caas-networking/prd.md Outdated
Comment thread enhancements/unified-networking/prd.md Outdated
Comment thread enhancements/vmaas-networking/prd.md Outdated
@github-actions

github-actions Bot commented Jul 8, 2026

Copy link
Copy Markdown

AI EP Review: EP-107

Score: 9/10 | Verdict: PASS

Criterion Score Notes
What 2/2 All PRDs clearly describe user-facing capabilities with specific personas (Tenant User, Tenant Admin, Cloud Infrastructure Admin, Cloud Provider Admin) and services (VMaaS, CaaS, BMaaS) identified. Us
Why 2/2 Strong business justification across all PRDs. The unified PRD identifies concrete pain: CaaS and BMaaS have no networking API, tenants must choose networking backends they shouldn't understand, Exter
How 1/2 Requirements are numbered (FR-1 through FR-13), measurable (NFR-1: synchronous allocation, NFR-2: <2 min per interface), and have detailed acceptance criteria checklists. However, some design leakage
Task 2/2 All documents are clearly enhancements adding new platform capabilities (unified networking API, multi-NIC support, auto external IP allocation, default networking at tenant onboarding). Not tasks or
Size 2/2 Each PRD is well-scoped to a coherent set of capabilities. The unified PRD covers the foundational networking model. Per-service PRDs (VMaaS, CaaS, BMaaS) each cover their service type's networking in

Verdict: Strong set of PRDs with clear user-facing needs, compelling justification, well-identified personas, and testable acceptance criteria; minor design leakage in BMaaS and CaaS requirements (FR-8, FR-11, FR-9) keeps the 'how' score at 1.

Feedback: Move implementation-specific language out of PRD requirements into design docs: BMaaS FR-8 should describe the user-observable outcome ('network connectivity is configured for each interface before OS provisioning begins') without specifying HOW ('identifies the physical server in the network fabric, adds the server's interface to the subnet's network segment'). BMaaS FR-11 should describe the user-observable effect of fabric manager configuration, not its internal representation ('static configuration string set at deployment time'). CaaS FR-9 should separate 'endpoint addresses are available in cluster status' (user-observable) from 'performs DNS record creation' (implementation detail).

Critical (0)

None.

Important (3)

  1. BMaaS PRD FR-8 contains design leakage: 'the system identifies the physical server in the network fabric, adds the server's interface to the subnet's network segment, and allocates an IP address (DHCP or static)' — these are implementation details that would change if the implementation changed. Rewrite as user-observable outcome: 'network connectivity is configured for each interface-to-subnet mapping before OS provisioning begins'.
  2. BMaaS PRD FR-11 describes an internal configuration parameter: 'a static configuration string set at deployment time (e.g., openstack), not a reference to a networking resource' — this is a system internal not verifiable by a PM using the product. Rewrite to describe the user-observable effect or move to design doc.
  3. CaaS PRD FR-9 mixes user outcomes with implementation: 'performs DNS record creation' is an implementation detail. The user-facing requirement is that endpoint addresses are available in cluster status.

Suggestions (2)

  1. Consider adding success metrics with baselines to the per-service PRDs (VMaaS, CaaS) — BMaaS has provisioning time and success rate metrics, but VMaaS and CaaS PRDs lack comparable metrics.
  2. The open question about auto NAT gateway state-aware reuse (appears in all 4 per-service PRDs) should be resolved before implementation — it affects user experience when a tenant encounters a failed NAT gateway and the system silently reuses it.

Review cost

Model: claude-opus-4-6
Cost: $1.2819
Tokens: 6.1k in / 5.2k out
Cache: 283.7k read
Active time: 1m 55s
API calls: 0

@github-actions

github-actions Bot commented Jul 8, 2026

Copy link
Copy Markdown

AI Design Review: EP-107

Score: 8/8 | Verdict: PASS

Criterion Score Notes
Feasibility 2/2 All four designs build on proven OSAC patterns (CRDs, operator reconciliation, dispatcher, AAP templates) and extend existing working infrastructure. Dependencies are explicitly tracked with Jira tick
Testability 2/2 Each design has a comprehensive test plan covering unit tests, integration/E2E tests, and tricky edge cases. Test scenarios are specific and enumerable: multi-NIC with primary on non-first interface,
Scope 2/2 Each document covers exactly one service type's networking with clear non-goals preventing scope creep (e.g., CaaS excludes VM-based node sets, DNS API, multi-NIC per node set; BMaaS excludes CaaS/VMa
Architecture 2/2 Designs follow consistent architectural patterns: component responsibility tables, reconciliation phase ordering (inventory -> networking -> provisioning -> power), auto-provisioned resource lifecycle

Verdict: A comprehensive, well-structured set of per-service networking enhancement proposals that consistently extend the unified networking architecture across VMaaS, CaaS, BMaaS, and default networking with thorough test plans, clear dependency tracking, and sound architectural patterns.

Feedback: The NATGateway reuse-regardless-of-state decision appears as both a risk and an open question in all four designs -- resolve this cross-cutting concern in one place (e.g., the unified EP or default networking EP) and reference it from the others to avoid divergent implementations. The 'Not tracked' GAPs in dependency tables (CRD updates, mutateBMI, IP feedback, HostType interfaces, agent selection logic) should be tracked as Jira tickets before implementation begins to prevent discovery-phase delays. Consider adding a sequence diagram (mermaid) for the VIP feedback loop in CaaS (template -> ClusterOrder status -> Signal RPC -> fulfillment-service -> Cluster -> ExternalIPAttachment controller) since it spans 5+ components and is called out as a complexity risk.

Critical (0)

None.

Important (2)

  1. NATGateway reuse-regardless-of-state is an unresolved open question duplicated across all four design documents (VMaaS, CaaS, BMaaS, Default Networking) with no centralized resolution plan -- this cross-cutting decision should be resolved once and referenced consistently
  2. Multiple implementation items are listed as 'Not tracked' / GAP in dependency tables across BMaaS (5 gaps), CaaS (5 gaps), and unified EP updates -- these represent unplanned work that could delay implementation

Suggestions (3)

  1. Add mermaid sequence diagrams for the most complex cross-component flows: CaaS VIP feedback loop (5+ components), BMaaS IP address feedback path, and auto ExternalIP prerequisite ordering for clusters
  2. Consider consolidating the shared patterns (auto ExternalIP pool selection algorithm, auto NATGateway reuse logic, auto-provisioned resource cleanup lifecycle) into the default networking or unified networking EP and referencing from per-service EPs to reduce duplication and ensure consistency
  3. The VMaaS PRD Non-Goals section states 'Multiple network interfaces for bare-metal servers (bare-metal multi-interface support is out of scope)' which contradicts the BMaaS design -- clarify that this means multi-NIC is out of scope for the VMaaS PRD specifically, not for the platform

Review cost

Model: claude-opus-4-6
Cost: $1.0152
Tokens: 6.0k in / 3.9k out
Cache: 303.5k read
Active time: 1m 32s
API calls: 0

@github-actions

github-actions Bot commented Jul 9, 2026

Copy link
Copy Markdown

AI Design Review: EP-107

Score: 8/8 | Verdict: PASS

Criterion Score Notes
Feasibility 2/2 Each design document provides concrete protobuf and Go type definitions, clear component responsibility tables, and implementation details. Dependencies are tracked with Jira tickets and status indica
Testability 2/2 All four design documents include comprehensive Test Plan sections with unit tests, integration tests, and tricky test cases. Unit tests cover validation logic (primary designation, interface validati
Scope 2/2 The decomposition is logical: one unified networking architecture doc, four per-service designs (VMaaS, CaaS, BMaaS, Default Networking), each with its own PRD. Goals and non-goals are clearly separat
Architecture 2/2 Consistent architectural patterns across all services: network_attachments field, auto ExternalIP/NATGateway modes, dispatcher-based fabric manager calls, label-based resource lifecycle management (os

Verdict: A comprehensive and well-structured set of per-service networking design documents that extend a unified architecture with consistent patterns, concrete API definitions, thorough test plans, and honest dependency tracking.

Feedback: The open questions about NATGateway state-aware reuse appear in all four designs and the PRDs — resolving this cross-cutting decision before implementation would eliminate ambiguity for all service teams simultaneously. Consider adding idempotency guarantees for dispatcher calls (create_network_attachment, delete_network_attachment) since reconciliation loops will retry on transient failures. The 5 untracked GAPs in both BMaaS and CaaS dependency tables should be converted to Jira tickets before work begins to avoid items falling through the cracks.

Critical (0)

None.

Important (2)

  1. NATGateway state-aware reuse is an unresolved open question repeated across all 4 designs and 3 PRDs — reusing a Failed NATGateway silently breaks outbound connectivity with no automated recovery path. This should be resolved as a cross-cutting architectural decision before implementation begins.
  2. 10 untracked dependency GAPs across BMaaS (5) and CaaS (5) designs — items like 'BareMetalInstance CRD: add NetworkAttachments', 'Agent selection logic in operator', 'mutateBMI: copy network_attachments to K8s CR' are substantive work items that risk being overlooked without Jira tracking.

Suggestions (3)

  1. Add idempotency requirements for dispatcher calls (create_network_attachment, delete_network_attachment) — reconciliation loops will retry failed calls, and the fabric manager roles need to handle duplicate invocations gracefully to avoid resource leaks or double-allocation.
  2. The IP/VIP feedback loops (fabric manager -> CR status -> feedback controller -> fulfillment-service -> ExternalIPAttachment controller) cross 4-5 components — consider adding an end-to-end sequence diagram (Mermaid) to each design to make the full chain visually clear for implementors and reviewers.
  3. Consider whether the lifecycle interface exclusion from tenant-attachable interfaces (documented as a convention) should be enforced in validation rather than just documented — accidental attachment to a lifecycle interface could disrupt provisioning infrastructure.

Review cost

Model: claude-opus-4-6
Cost: $1.0128
Tokens: 6.1k in / 3.7k out
Cache: 305.6k read
Active time: 1m 25s
API calls: 0

@github-actions

github-actions Bot commented Jul 9, 2026

Copy link
Copy Markdown

AI Design Review: EP-107

Score: 8/8 | Verdict: PASS

Criterion Score Notes
Feasibility 2/2 Each design builds incrementally on existing K8s operator patterns (CRDs, controllers, finalizers, CEL validation) and proven OSAC infrastructure (AAP integration, dispatcher, feedback controllers). H
Testability 2/2 All four design documents include comprehensive test plans covering unit tests (validation logic, pool selection, reuse behavior), integration/E2E tests (multi-NIC provisioning, auto ExternalIP/NATGat
Scope 2/2 Each per-service design document is self-contained with clear boundaries: VMaaS covers multi-NIC VMs and auto external access, BMaaS covers switch port configuration and physical interface mapping, Ca
Architecture 2/2 The design follows sound architectural principles: unified API model across three service types, pluggable manager architecture via NetworkClass/dispatcher, clean separation of concerns (fulfillment-s

Verdict: A comprehensive, well-structured set of per-service networking design documents that decompose the OSAC unified networking architecture into clear, implementable per-service flows with thorough API definitions, failure handling, test plans, and operational procedures.

Feedback: The NATGateway reuse-regardless-of-state design is flagged as an open question in all four documents — resolve this cross-cutting concern once (perhaps in the unified EP) rather than leaving identical open questions in each per-service doc. Consider adding mermaid sequence diagrams for the multi-component flows (especially CaaS VIP feedback loop: template -> ClusterOrder status -> Signal RPC -> fulfillment-service -> Cluster -> ExternalIPAttachment controller) as the text-based workflow descriptions are harder to follow for complex interactions. The 5 untracked GAPs in BMaaS and 5 in CaaS dependency tables should be converted to Jira tickets before implementation begins to avoid work falling through the cracks.

Critical (0)

None.

Important (2)

  1. NATGateway reuse-regardless-of-state is an unresolved open question repeated across all 4 design docs — reusing a Failed NATGateway silently breaks outbound connectivity with no automatic recovery. This should be resolved as a cross-cutting decision before implementation, ideally with at least a clear warning event when reusing a non-Ready NATGateway.
  2. Multiple dependency items marked as 'GAP' (not tracked in Jira) across BMaaS (5 items: CRD update, mutateBMI, IP feedback, networkClass rename, HostType extension) and CaaS (5 items: HostType, step collection removal, NETWORK_STEPS_COLLECTION removal, agent selection logic, interface resolution). These represent real implementation work that could be missed without tracking tickets.

Suggestions (3)

  1. The HostType proto definition is repeated verbatim across BMaaS design, CaaS design, and the updated unified networking design. Consider defining it once in the unified EP and referencing it from per-service docs to avoid divergence during implementation.
  2. Add mermaid sequence diagrams for the CaaS VIP feedback loop and the BMaaS IP address feedback flow — these multi-component async flows involve 5+ components and are the hardest parts to follow in text form.
  3. The Default Networking design specifies that tenant onboarding failure recovery is 'delete and re-create tenant' — consider whether a more graceful recovery (retry provisioning of failed default resources) would be operationally preferable, especially if the tenant already has resources.

Review cost

Model: claude-opus-4-6
Cost: $1.0235
Tokens: 6.1k in / 4.0k out
Cache: 305.8k read
Active time: 1m 37s
API calls: 0

@github-actions

github-actions Bot commented Jul 9, 2026

Copy link
Copy Markdown

AI Design Review: EP-107

Score: 7/8 | Verdict: PASS

Criterion Score Notes
Feasibility 2/2 Each of the five enhancements (VMaaS, CaaS, BMaaS, Default Networking, plus Unified Networking updates) is technically sound. Proto messages, CRD structs, CEL validation rules, operator reconciliation
Testability 2/2 All five enhancements include structured test plans with unit tests, integration/E2E tests, and explicit tricky test cases. Acceptance criteria are concrete and checkable (e.g., 'creating a BM with an
Scope 1/2 The scope is well-defined per-enhancement — each document clearly states goals, non-goals, and what is deferred (e.g., VM-based CaaS node sets, DNS API, multi-NIC failover/bonding). However, the overa
Architecture 2/2 The architecture is strong. The unified networking model (NetworkClass, dispatcher, infrastructure-agnostic subnets) provides a clean foundation. Per-service-type attachment messages (ComputeNetworkAt

Verdict: A comprehensive and architecturally sound set of enhancements that extend the OSAC unified networking API to all three service types (VMaaS, CaaS, BMaaS) with default networking, auto external access, and multi-NIC support; the main weakness is the sheer breadth of the PR and several unresolved cross-cutting design questions.

Feedback: The repeated open question about NATGateway reuse semantics (reuse regardless of state vs. state-aware reuse) appears identically in all four service-specific enhancements and the default-networking EP — resolve this once in the unified networking EP and reference it from the others to avoid divergence during implementation. The dependency tables show 5+ untracked GAP items per enhancement (CRD updates, mutateBMI, IP feedback, HostType extensions, agent selection logic); these should be tracked in Jira before merging to prevent implementation surprises. Consider splitting this PR into the unified networking updates (including HostType) as one PR and the per-service EPs as follow-ups to make review more tractable.

Critical (0)

None.

Important (3)

  1. NATGateway reuse-regardless-of-state design decision is duplicated as an open question across all 5 enhancements (BMaaS OQ#1, CaaS OQ#4, VMaaS OQ#1, Default Networking OQ#1, unified PRD) — this should be resolved once in the unified EP and referenced, not left open in parallel documents that could diverge during implementation.
  2. Multiple untracked dependency GAPs per enhancement (e.g., BMaaS has 5 GAPs: CRD update, mutateBMI, IP feedback, networkClass rename, HostType extension; CaaS has 5 GAPs: HostType, step collection removal, NETWORK_STEPS_COLLECTION removal, agent selection, interface resolution). These represent real implementation work with no Jira tracking, risking scope creep or missed deliverables.
  3. The auto NATGateway reuse logic explicitly reuses Failed or Deleting NATGateways, which the document itself flags as a risk with only 'document expected behavior' as mitigation. This is a known-bad UX path that could be resolved at design time rather than deferred to implementation.

Suggestions (3)

  1. Consider extracting the HostType NetworkInterface extension into its own small enhancement since it is a shared dependency across BMaaS and CaaS — this would make the dependency graph clearer and allow it to land independently.
  2. The capacity-exhaustion-vs-failed-resource open question (OQ#2 in BMaaS, OQ#2 in VMaaS, OQ#2 in Default Networking) should also be resolved once centrally rather than left open in three places.
  3. Add a sequence diagram (mermaid) for the CaaS VIP feedback loop (template -> ClusterOrder -> Signal -> fulfillment-service -> Cluster -> ExternalIPAttachment) — the text description spans 4 steps across 3 components and would benefit from visual representation.

Review cost

Model: claude-opus-4-6
Cost: $1.0381
Tokens: 6.1k in / 2.5k out
Cache: 422.5k read
Active time: 1m 20s
API calls: 0

@github-actions

github-actions Bot commented Jul 9, 2026

Copy link
Copy Markdown

AI Design Review: EP-107

Score: 7/8 | Verdict: PASS

Criterion Score Notes
Feasibility 2/2 The design is technically feasible and builds incrementally on proven patterns (existing two-operator architecture, existing ExternalIPAttachment requeue pattern, existing AAP integration). Each per-s
Testability 2/2 Each design document includes a structured test plan with unit tests, integration/E2E tests, and explicitly identified tricky test cases. The test plans cover golden paths (multi-NIC provisioning, aut
Scope 1/2 The scope is ambitious but coherent — four per-service EPs plus default networking, all consuming the unified networking architecture. Each EP is self-contained with clear boundaries. However, the tot
Architecture 2/2 The architecture is sound and well-layered. The unified networking API provides a clean abstraction (VirtualNetwork/Subnet/SecurityGroup/ExternalIP/NATGateway) consumed uniformly by all three service

Verdict: A comprehensive, well-architected networking enhancement suite that cleanly decomposes a unified networking API into per-service flows with consistent patterns; the main concern is the large scope with many untracked implementation gaps and several identical open questions that need resolution before implementation begins.

Feedback: Resolve the NATGateway state-aware reuse open question once and reference it from all four EPs rather than repeating the same open question and risk in each document — this is a cross-cutting design decision that should be settled at the unified networking level. Create Jira tickets for the 10+ untracked gaps (marked GAP in dependency tables) before implementation starts, as these represent real work that could surprise the schedule. Consider adding a sequence diagram or state machine for the auto-provisioning two-phase lifecycle, as the text description spans multiple documents and the interaction between ExternalIP Pending->Allocated and ExternalIPAttachment Pending->Ready across different target types (VM/BM/Cluster) would benefit from a visual representation.

Critical (0)

None.

Important (3)

  1. NATGateway reuse-regardless-of-state is listed as both a design decision AND an open question in all four per-service EPs (VMaaS, CaaS, BMaaS, default-networking) plus both the BMaaS and CaaS PRDs. This inconsistency (is it decided or open?) should be resolved at the unified networking level before implementation.
  2. 10+ dependency items are marked as GAP (not tracked in Jira) across BMaaS and CaaS dependency tables — including CRD changes, mutateBMI updates, IP address feedback flow, HostType extensions, agent selection logic, and step collection removal. These represent significant unplanned work.
  3. The IP writeback mechanism in BMaaS (AAP role writes to CR status via kubeconfig) is stated as 'Recommended' (Option A) but also listed as Open Question adds Apache 2.0 LICENSE file #3 needing confirmation. The design should commit to one option or explicitly gate implementation on resolving this question.

Suggestions (3)

  1. Add a cross-reference section to the unified networking EP that provides a single source of truth for shared design decisions (NATGateway reuse policy, capacity exhaustion behavior, IP feedback mechanism) rather than repeating them in each per-service EP.
  2. The VMaaS EP mentions 'BMaaS has no multi-NIC concept' in the alternatives section, but the BMaaS EP is entirely about multi-NIC support — this appears to be a stale statement from an earlier draft that should be corrected.
  3. Consider adding an explicit ordering or phasing recommendation across the four per-service EPs — which should be implemented first, which has the fewest dependencies, and which can be parallelized — to help the team plan implementation sprints.

Review cost

Model: claude-opus-4-6
Cost: $1.0561
Tokens: 6.1k in / 2.4k out
Cache: 428.4k read
Active time: 1m 5s
API calls: 0

@github-actions

github-actions Bot commented Jul 9, 2026

Copy link
Copy Markdown

AI Design Review: EP-107

Score: 8/8 | Verdict: PASS

Criterion Score Notes
Feasibility 2/2 Each design provides concrete implementation paths with proto definitions, CRD structs, CEL validation rules, and reconciliation phase ordering. Dependencies are tracked with Jira tickets and status.
Testability 2/2 All four designs include comprehensive test plans with unit tests (validation logic, pool selection, reuse logic), integration/E2E tests (multi-NIC, auto ExternalIP, auto NATGateway, VIP feedback, cle
Scope 2/2 Each document is clearly scoped to a single service type with explicit non-goals. Cross-cutting concerns (HostType, auto-provisioning lifecycle, NATGateway reuse logic) are defined once in the Unified
Architecture 2/2 The designs follow sound architectural principles: consistent resource hierarchy across all service types, per-resource-type attachment messages with type-specific validation, two-operator architectur

Verdict: A thorough and well-structured set of per-service networking enhancement proposals that consistently extend the unified networking architecture with clear implementation paths, comprehensive test plans, and explicit dependency tracking — the main gaps are untracked Jira items and a few unresolved open questions that should be closed before implementation begins.

Feedback: Resolve the three recurring open questions before implementation: (1) NATGateway Deleting-state reuse semantics — the current 'reuse regardless of state' behavior is a documented weakness across all four designs; at minimum, treat Deleting as 'does not exist' to avoid silently attaching to a disappearing resource. (2) Capacity exhaustion behavior (API error vs Failed resource) — pick one and apply consistently. (3) IP address feedback mechanism for BMaaS (Option A is recommended but still listed as an open question in design.md OQ#3 — promote to a decision). Additionally, create Jira tickets for all items marked as 'GAP' in the dependency tables (5 in BMaaS, 5 in CaaS) — these represent real implementation work that is invisible to planning without tracking.

Critical (0)

None.

Important (3)

  1. Dependency tables in BMaaS (design.md lines 776-780) and CaaS (design.md lines 1782-1786) list 5+ items each as 'Not tracked / GAP' — these represent real implementation work (CRD updates, mutateBMI changes, IP feedback flow, HostType extensions, agent selection logic) that needs Jira tickets for planning visibility.
  2. The NATGateway 'reuse regardless of state' behavior (including Failed and Deleting states) is acknowledged as problematic in all four designs' open questions and risk sections, but no resolution is proposed. A Deleting NATGateway will soon cease to exist — silently reusing it guarantees broken outbound connectivity with no recovery path except manual delete-and-retry.
  3. BMaaS design OQ#3 (IP address feedback mechanism) is listed as an open question despite Option A being recommended in the design body. This is a core architectural decision affecting component responsibilities, the ExternalIPAttachment controller, and the reconciliation phase ordering — it should be promoted to a decision before implementation begins.

Suggestions (3)

  1. Consider adding mermaid sequence diagrams for the two most complex flows: the CaaS VIP feedback loop (template -> ClusterOrder status -> Signal RPC -> fulfillment-service -> Cluster -> ExternalIPAttachment controller) and the auto-provisioning lifecycle (Phase 1 synchronous DB transaction -> Phase 2 async controller reconciliation). These multi-hop flows are described textually but would benefit from visual representation.
  2. CaaS open question Bump actions/setup-python from 5 to 6 #1 (how does the operator select agents — K8s API vs AAP job) is fundamental to the reconcileAgentSelection implementation but has no recommendation. Adding a recommended approach would help the implementing team.
  3. CaaS open question adds Apache 2.0 LICENSE file #3 (MetalLB IPAddressPool CR ownership — subnet creation vs cluster provisioning) affects the boundary between k8s_manager and CaaS template responsibilities. Clarifying this would prevent integration issues during implementation.

Review cost

Model: claude-opus-4-6
Cost: $1.0554
Tokens: 6.1k in / 3.8k out
Cache: 311.8k read
Active time: 1m 30s
API calls: 0

@github-actions

github-actions Bot commented Jul 9, 2026

Copy link
Copy Markdown

AI Design Review: EP-107

Score: 8/8 | Verdict: PASS

Criterion Score Notes
Feasibility 2/2 All five designs are grounded in the existing OSAC architecture with detailed proto definitions, CRD changes, Go struct layouts, and CEL validation rules. Component responsibilities are clearly assign
Testability 2/2 Each design includes structured test plans: unit tests covering validation rules (primary designation, interface validation, pool selection, NATGateway reuse), integration/E2E tests covering full prov
Scope 2/2 Each EP is cleanly scoped to a single service type (VMaaS, CaaS, BMaaS) or cross-cutting concern (Default Networking), with clear non-goals preventing scope creep. The decomposition from the parent Un
Architecture 2/2 The designs follow sound architectural principles: consistent patterns across all service types (auto ExternalIP/NATGateway lifecycle, auto-provisioned cleanup via finalizers, operator IPAM), clear se

Verdict: A comprehensive, well-structured suite of networking enhancement designs that decompose a complex cross-service feature into modular, per-service EPs with consistent patterns, detailed API definitions, thorough test plans, and honest dependency/risk tracking.

Feedback: The designs are strong overall. Three areas for improvement: (1) The NATGateway 'reuse regardless of state' decision should be resolved before implementation rather than deferred — silently attaching to a Deleting NATGateway is a user-facing bug, and the open question about treating Deleting as 'does not exist' should be closed in this design phase. (2) The CaaS design's agent selection migration from template to operator (reconcileAgentSelection) is the highest-risk change and could benefit from a more detailed specification of the selection algorithm rather than 'port existing logic.' (3) Consider adding a dependency tracking summary table across all EPs showing the critical path — the individual dependency tables are thorough but the cross-EP ordering (e.g., which EP blocks which) would help with implementation planning.

Critical (0)

None.

Important (3)

  1. NATGateway reuse of Deleting state is an unresolved open question appearing in all 4 service designs — this design decision should be closed before implementation to avoid user-facing issues where auto-provisioned resources silently reference a disappearing NATGateway
  2. Five untracked dependency GAPs in BMaaS (CRD update, mutateBMI, IP feedback, networkClass rename, HostType extension) and five in CaaS (HostType, step collection removal, NETWORK_STEPS_COLLECTION removal, agent selection, interface resolution) need Jira tickets before implementation begins
  3. CaaS reconcileAgentSelection open question (Bump actions/setup-python from 5 to 6 #1) has no specification for the agent selection algorithm — 'port existing logic' defers design to implementation phase for a high-risk component migration

Suggestions (3)

  1. Add a cross-EP dependency graph or critical path summary to help implementation planning across the 4 new EPs plus the unified networking updates
  2. The capacity exhaustion open question (API error vs Failed resource) appears in all 4 designs — consider resolving it once in the unified networking EP rather than leaving it open in each per-service design
  3. Consider adding a concurrency/race condition section to the BMaaS and CaaS designs covering operator IPAM — two operators allocating IPs from the same subnet CIDR simultaneously could produce conflicts without a locking mechanism

Review cost

Model: claude-opus-4-6
Cost: $1.0366
Tokens: 6.1k in / 3.2k out
Cache: 312.3k read
Active time: 1m 22s
API calls: 0

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 10

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (3)
enhancements/unified-networking/design.md (1)

835-839: 🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Clarify node_set vs fabric_interface for cluster attachments. node_set is optional here, but fabric_interface is a single immutable value derived from a node set’s HostType. If one attachment can cover multiple node sets, the spec needs to say how that field is resolved for differing HostTypes; otherwise make node_set required for v0.2.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@enhancements/unified-networking/design.md` around lines 835 - 839, The
cluster attachment spec is ambiguous about how an optional node_set interacts
with the single immutable fabric_interface value. Update the unified-networking
design text to either require node_set for v0.2 or explicitly define how
fulfillment-service resolves fabric_interface when one attachment applies to
multiple node sets with different HostTypes. Refer to the cluster attachment
entry, node_set, fabric_interface, and fulfillment-service resolution rules so
the behavior is unambiguous.
enhancements/bmaas-networking/prd.md (1)

96-96: 🩺 Stability & Availability | 🟠 Major | ⚡ Quick win

Restrict AUTO NAT gateway reuse to ready resources. Reusing an existing NAT gateway while it’s deleting or otherwise unavailable can leave the new server with broken outbound connectivity. Define a READY/ACTIVE gate, or fail/create a replacement when the current gateway isn’t usable.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@enhancements/bmaas-networking/prd.md` at line 96, Update the NAT gateway AUTO
behavior in the BMAAS networking spec so reuse only applies to gateways that are
actually ready/active; if the existing gateway is deleting or otherwise
unavailable, the server should not attach to it. Adjust the FR-7 wording in the
NAT gateway mode section to reference the readiness gate and clarify that
create-or-replace behavior is required when the current gateway is unusable,
while keeping the existing AUTO/NONE semantics in the same spec area.
enhancements/vmaas-networking/design.md (1)

304-305: 🩺 Stability & Availability | 🟠 Major | 🏗️ Heavy lift

Keep the finalizer until cleanup succeeds.

Dropping the finalizer after N retries can delete the parent while auto-provisioned ExternalIP/Attachment records are still allocated, which leaks capacity and leaves public endpoints orphaned. Keep the finalizer and surface a terminal cleanup condition, or hand off orphan cleanup to a separate reconciler.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@enhancements/vmaas-networking/design.md` around lines 304 - 305, The cleanup
flow currently removes the finalizer after N retries in the auto-provisioned
resource cleanup path, which can let the parent resource delete while
ExternalIP/ExternalIPAttachment records are still allocated. Update the
finalizer/reconciliation design so the finalizer stays until cleanup actually
succeeds, and use a terminal cleanup condition or a separate reconciler to
handle orphaned ExternalIP/Attachment cleanup; refer to the auto-provisioned
resource cleanup transient/permanent failure behavior in the design doc.
♻️ Duplicate comments (2)
enhancements/caas-networking/prd.md (2)

78-78: 🩺 Stability & Availability | 🟠 Major | ⚡ Quick win

Restrict NAT gateway reuse to healthy gateways.

“Reuse any existing NAT gateway” still allows Failed/Deleting gateways to be reused, which can leave new clusters without outbound connectivity. Limit reuse to Ready gateways or fail creation explicitly.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@enhancements/caas-networking/prd.md` at line 78, Restrict NAT gateway reuse
in the cluster creation flow so only healthy gateways are reused: update the
FR-4 behavior around automatic NAT gateway provisioning to check the NAT
gateway’s status before reusing it, and treat Failed or Deleting gateways as
unusable. In the cluster creation logic that implements the NAT gateway
lookup/reuse path, ensure only Ready gateways are selected; otherwise fail
creation explicitly instead of silently reusing the unhealthy gateway.

66-66: 🎯 Functional Correctness | 🟠 Major | ⚡ Quick win

Define precedence when nodeSet is omitted.

The attachment scope is still ambiguous: FR-1 allows nodeSet to be optional, but the PRD never says whether an omitted value applies cluster-wide or how it interacts with node-set-specific attachments. Make that explicit for v0.2.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@enhancements/caas-networking/prd.md` at line 66, Clarify FR-1’s network
attachment scope by explicitly stating the default behavior when nodeSet is
omitted, since it is currently ambiguous. Update the FR-1 requirement so it says
whether an omitted nodeSet applies cluster-wide or only to a default node pool,
and define its precedence relative to nodeSet-specific attachments in v0.2.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@enhancements/bmaas-networking/design.md`:
- Around line 204-205: The NATGateway auto-provisioning rule is too broad
because `nat_gateway_mode == AUTO` currently says to reuse an existing
NATGateway regardless of state, which can bind a new BaremetalInstance to a
gateway that is deleting or otherwise unstable. Update the `nat_gateway_mode`
flow in the design so reuse is limited to stable, ready-like states only, and
explicitly state that non-stable states must not be reused; keep the wording
aligned with the NATGateway state lifecycle sections referenced later in the
same document.
- Around line 244-247: The workflow description is inconsistent about who owns
primary-IP writeback: `ExternalIPAttachment`/`reconcileNetworking` says the
fabric manager writes `status.networkAttachments[].ipAddress`, while the “Subnet
IPAM — Resolved” section implies the operator does it. Update the design so a
single component is clearly responsible for allocating and writing back the
primary IP, and then make both the step 6b text and the later subnet/IPAM
section match that owner consistently.
- Around line 216-217: The current subnet IPAM flow in the design is race-prone
because `list existing allocations → pick next free IP` can let concurrent
reconciles assign the same address. Update the operator-managed allocation path
for `status.networkAttachments[].ipAddress` to use an atomic reservation
mechanism keyed on the Subnet CR before writing status back. Keep the
switch-side dispatch in `osac.templates.{{ fabric_manager
}}.create_network_attachment` unchanged, but ensure the IP reservation happens
first and only one BaremetalInstance can claim a given subnet IP at a time.

In `@enhancements/bmaas-networking/prd.md`:
- Around line 145-148: The defaulted-server networking requirements in the PRD
still assume a fabric-role interface exists when attachments are omitted, but
the host-type interface list may be empty. Update the default networking
requirement around the tenant onboarding/defaults and server creation flow to
explicitly cover the no-fabric-interface case for omitted attachments, and
define the same clear failure behavior used for explicit attachments when the
host type has no fabric interface. Reference the FR-5/defaulted server wording
and the host type physical network interface list requirement so the rule is
consistent in both paths.

In `@enhancements/caas-networking/design.md`:
- Around line 138-140: The IP/VIP allocation flow in the operator-managed IPAM
section is subject to a race because Subnet CR reads and status writes are not
atomic. Update the provisioning logic around the operator allocation path to use
a transactional claim/lease or other conflict-safe reservation mechanism when
assigning `status.nodeSets[].agents[].ipAddress`, `status.apiVIP`, and
`status.ingressVIP`, and make the reconcile retry/handle conflicts cleanly in
the same allocation routine that currently computes the next available IPs.
- Line 125: The AUTO NATGateway flow currently allocates a separate ExternalIP
before creating the NATGateway in a later transaction, but it does not define
cleanup if that second step fails. Update the design around the NATGateway AUTO
path to either make the ExternalIP and NATGateway creation atomic or add a
compensating delete/release for the allocated ExternalIP on failure, and ensure
the logic tied to nat_gateway_mode == AUTO and the auto-provisioned
ExternalIP/NATGateway handling preserves pool capacity.

In `@enhancements/default-networking/design.md`:
- Around line 566-570: The Tenant readiness story is inconsistent: the risk note
says onboarding can still complete with Tenant becoming READY even when
NetworkClass defaults are missing, while the main flow around
DefaultNetworkingReady implies READY should be gated on default networking being
configured. Align the design by choosing one contract and updating the relevant
sections in the document, especially the DefaultNetworkingReady readiness gating
and the deployment risk/mitigation text, so implementers do not wire conflicting
readiness checks.
- Around line 459-466: Update the NATGateway reuse behavior in the
default-networking design so `nat_gateway_mode=AUTO` only reuses an existing
NATGateway when it is READY. The current reuse rule in the NATGateway reuse
logic incorrectly allows Failed or Deleting gateways to be selected; change the
described flow to filter by ready state before reusing the first matching
NATGateway by name, and if no READY gateway exists, describe the
create-or-replace/recovery path explicitly instead of reusing unhealthy
resources.

In `@enhancements/unified-networking/design.md`:
- Around line 630-637: The Cluster row in the ExternalIPAttachment preconditions
uses the wrong status owner and should be updated to match the Cluster-backed
contract. In the target-type table, replace references to
ClusterOrder.status.apiVIP/ingressVIP with
Cluster.status.api_endpoint/ingress_endpoint, and keep the description aligned
with the CaaS flow and API ownership model so implementers read the correct
source of target IP.

In `@enhancements/vmaas-networking/design.md`:
- Around line 159-163: Update the NATGateway reuse logic in the
fulfillment-service flow described under the NATGateway creation/reuse step so
it only reuses an existing NATGateway when it is Ready. Treat Failed, Deleting,
or any non-Ready state as absent and create a replacement NATGateway instead,
while keeping the separate ExternalIP and separate-transaction behavior intact.
Use the existing NATGateway state checks in the fulfillment-service path and the
dispatcher/fabric-manager create_nat_gateway flow to locate the change.

---

Outside diff comments:
In `@enhancements/bmaas-networking/prd.md`:
- Line 96: Update the NAT gateway AUTO behavior in the BMAAS networking spec so
reuse only applies to gateways that are actually ready/active; if the existing
gateway is deleting or otherwise unavailable, the server should not attach to
it. Adjust the FR-7 wording in the NAT gateway mode section to reference the
readiness gate and clarify that create-or-replace behavior is required when the
current gateway is unusable, while keeping the existing AUTO/NONE semantics in
the same spec area.

In `@enhancements/unified-networking/design.md`:
- Around line 835-839: The cluster attachment spec is ambiguous about how an
optional node_set interacts with the single immutable fabric_interface value.
Update the unified-networking design text to either require node_set for v0.2 or
explicitly define how fulfillment-service resolves fabric_interface when one
attachment applies to multiple node sets with different HostTypes. Refer to the
cluster attachment entry, node_set, fabric_interface, and fulfillment-service
resolution rules so the behavior is unambiguous.

In `@enhancements/vmaas-networking/design.md`:
- Around line 304-305: The cleanup flow currently removes the finalizer after N
retries in the auto-provisioned resource cleanup path, which can let the parent
resource delete while ExternalIP/ExternalIPAttachment records are still
allocated. Update the finalizer/reconciliation design so the finalizer stays
until cleanup actually succeeds, and use a terminal cleanup condition or a
separate reconciler to handle orphaned ExternalIP/Attachment cleanup; refer to
the auto-provisioned resource cleanup transient/permanent failure behavior in
the design doc.

---

Duplicate comments:
In `@enhancements/caas-networking/prd.md`:
- Line 78: Restrict NAT gateway reuse in the cluster creation flow so only
healthy gateways are reused: update the FR-4 behavior around automatic NAT
gateway provisioning to check the NAT gateway’s status before reusing it, and
treat Failed or Deleting gateways as unusable. In the cluster creation logic
that implements the NAT gateway lookup/reuse path, ensure only Ready gateways
are selected; otherwise fail creation explicitly instead of silently reusing the
unhealthy gateway.
- Line 66: Clarify FR-1’s network attachment scope by explicitly stating the
default behavior when nodeSet is omitted, since it is currently ambiguous.
Update the FR-1 requirement so it says whether an omitted nodeSet applies
cluster-wide or only to a default node pool, and define its precedence relative
to nodeSet-specific attachments in v0.2.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository: osac-project/coderabbit/.coderabbit.yaml

Review profile: ASSERTIVE

Plan: Enterprise

Run ID: b30acebd-4acc-492f-8199-2a86e5ddedbf

📥 Commits

Reviewing files that changed from the base of the PR and between c2693a3 and 28309b7.

📒 Files selected for processing (10)
  • enhancements/bmaas-networking/design.md
  • enhancements/bmaas-networking/prd.md
  • enhancements/caas-networking/design.md
  • enhancements/caas-networking/prd.md
  • enhancements/default-networking/design.md
  • enhancements/default-networking/prd.md
  • enhancements/unified-networking/design.md
  • enhancements/unified-networking/prd.md
  • enhancements/vmaas-networking/design.md
  • enhancements/vmaas-networking/prd.md

Comment thread enhancements/bmaas-networking/design.md Outdated
Comment thread enhancements/bmaas-networking/design.md Outdated
Comment thread enhancements/bmaas-networking/design.md Outdated
Comment thread enhancements/OSAC-1437-bmaas-networking/prd.md
Comment thread enhancements/caas-networking/design.md Outdated
Comment thread enhancements/default-networking/design.md Outdated
Comment thread enhancements/default-networking/design.md Outdated
Comment thread enhancements/OSAC-1433-unified-networking/design.md
Comment thread enhancements/vmaas-networking/design.md Outdated
@github-actions

github-actions Bot commented Jul 9, 2026

Copy link
Copy Markdown

AI Design Review: EP-107

Score: 8/8 | Verdict: PASS

Criterion Score Notes
Feasibility 2/2 All four per-service designs specify concrete protobuf messages, Go CRD structs, CEL validation rules, reconciliation phase ordering, and dispatcher call signatures. The operator IPAM approach (alloca
Testability 2/2 Each design document includes unit tests (validation logic, pool selection, NATGateway reuse), integration/E2E tests (full lifecycle for multi-NIC, auto ExternalIP, auto NATGateway, cleanup), and 'tri
Scope 2/2 Each per-service design is well-bounded: VMaaS covers multi-NIC + auto external access, BMaaS covers switch port config + interface mapping, CaaS covers cluster networking + VIP feedback loop, Default
Architecture 2/2 The unified networking model is architecturally coherent: a single resource hierarchy (VirtualNetwork -> Subnet -> SecurityGroup) shared across all three service types, with per-service attachment mes

Verdict: A comprehensive, architecturally sound set of design documents that define a unified networking model across VMaaS, CaaS, and BMaaS with consistent patterns for auto-provisioning, IPAM, tenant isolation, and resource lifecycle management.

Feedback: The designs are strong overall. Two areas to tighten before implementation: (1) The operator IPAM approach allocates IPs by scanning existing allocations on the subnet, but the concurrency model isn't specified — clarify how two concurrent reconcileNetworking runs on different resources sharing the same subnet avoid allocating the same IP (K8s optimistic concurrency on the Subnet CR, or a dedicated IPAllocation sub-resource?). (2) The NATGateway 'reuse regardless of state' decision is flagged as an open question in all four designs — resolve this before implementation, as silently attaching to a Deleting NATGateway is a user-facing footgun that will generate support tickets.

Critical (0)

None.

Important (2)

  1. Operator IPAM concurrency model is underspecified: if two BareMetalInstance or ClusterOrder reconcilers allocate IPs from the same subnet CIDR simultaneously, the 'list existing allocations, pick next available' approach could produce duplicate IPs without an explicit locking or optimistic concurrency mechanism on the allocation state.
  2. NATGateway reuse-regardless-of-state is flagged as an open question in all four per-service designs plus the default networking design, but no resolution timeline is given — this cross-cutting design decision should be resolved once (in the unified networking EP) rather than left open in five places.

Suggestions (3)

  1. Consider adding an automated orphan detection mechanism (e.g., a periodic reconciler that scans for auto-provisioned resources with no parent) rather than relying solely on manual cleanup when finalizer cleanup fails permanently.
  2. The HostType NetworkInterface list is defined identically in BMaaS and CaaS designs — consider a single canonical definition in the unified networking EP with cross-references, to avoid drift if the schema evolves.
  3. The 'GAP' items in the dependency tables (5+ per design, untracked in JIRA) should be converted to tracked work items before implementation begins to avoid scope creep and missed deliverables.

Review cost

Model: claude-opus-4-6
Cost: $1.0607
Tokens: 6.1k in / 3.8k out
Cache: 313.4k read
Active time: 1m 31s
API calls: 0

@github-actions

github-actions Bot commented Jul 9, 2026

Copy link
Copy Markdown

AI Design Review: EP-107

Score: 7/8 | Verdict: PASS

Criterion Score Notes
Feasibility 2/2 The design is technically feasible and grounded in existing infrastructure. Each service type (VMaaS, CaaS, BMaaS) extends proven patterns — existing CRDs, operator reconciliation loops, AAP workflows
Testability 2/2 Each design document includes a structured test plan with unit tests, integration/E2E tests, and tricky test cases. The test plans are specific and actionable — e.g., 'create BaremetalInstance with in
Scope 1/2 The scope is well-defined per service type, with clear non-goals and explicit deferral of items (VM-based CaaS node sets, DNS API, NIC bonding). However, the aggregate scope across all 6 documents is
Architecture 2/2 The architecture is sound and well-structured. The unified networking model (NetworkClass, dispatcher, infrastructure-agnostic subnets) provides clean separation of concerns. Each service type has a d

Verdict: A comprehensive, architecturally sound design suite for unified networking across VMaaS, CaaS, and BMaaS — well-specified with clear API contracts, test plans, and dependency tracking — held back slightly by aggregate scope breadth and several unresolved open questions that could affect implementation.

Feedback: The individual designs are strong, but consider staging the delivery more explicitly — the dependency tables show most CaaS and BMaaS work items are 'New' while prerequisites like dispatcher core and NATGateway full stack are only partially in progress. A phased delivery plan (e.g., VMaaS auto-access first, then BMaaS networking, then CaaS with VIP feedback) with explicit gates would reduce integration risk. Resolve the NATGateway Deleting-state open question before implementation — silently attaching to a disappearing resource is a user-facing bug, not a design tradeoff to defer. Finally, the cross-operator IPAM concurrency section mentions optimistic locking on Subnet CR status but lacks detail on the allocation tracking structure; specify whether allocations are tracked as a list in Subnet status or via separate IPAllocation CRs, as this affects conflict behavior at scale.

Critical (0)

None.

Important (4)

  1. NATGateway Deleting-state reuse is flagged as an open question across all 4 service designs but deferred — this creates a real user-facing issue where auto-provisioned resources silently reference a disappearing NATGateway, and the workaround (manual delete + retry) is poor UX. This should be resolved before implementation begins.
  2. Cross-operator IPAM concurrency (osac-operator and bare-metal-fulfillment-operator both allocating from the same subnet CIDR) mentions optimistic concurrency on Subnet CR status but does not specify the allocation tracking data structure or conflict resolution behavior under load. Two operators writing to the same status field need a defined protocol.
  3. MetalLB IPAddressPool ownership is called out as a risk in the CaaS design but left unresolved — without an IPAddressPool CR covering the pre-allocated VIPs, MetalLB cannot announce them. Whether this is created at subnet creation (k8s_manager) or cluster provisioning (template) affects the component responsibility matrix and must be decided before implementation.
  4. 16+ untracked work items (gaps) across BMaaS and CaaS designs represent real implementation work with no Jira tracking — these should be tracked to avoid scope surprises during delivery.

Suggestions (3)

  1. Consider adding a single cross-cutting 'phased delivery plan' section to the unified networking design that sequences the four service-type implementations and defines explicit gates between phases.
  2. The FR-12 auto-cleanup for CaaS says NATGateways ARE cleaned up (contradicting the per-VN shared resource design stated elsewhere) — verify consistency of NATGateway cleanup semantics across all documents.
  3. The PRD risk 8.3 in the unified networking PRD mentions 'fabric manager does not write the allocated IP to server status' but the design explicitly says the operator (not fabric manager) allocates IPs — this PRD risk is stale and should be updated to match the resolved design.

Review cost

Model: claude-opus-4-6
Cost: $1.0201
Tokens: 6.1k in / 2.4k out
Cache: 313.7k read
Active time: 1m 11s
API calls: 0

@github-actions

github-actions Bot commented Jul 9, 2026

Copy link
Copy Markdown

AI Design Review: EP-107

Score: 8/8 | Verdict: PASS

Criterion Score Notes
Feasibility 2/2 All five designs provide concrete protobuf/Go struct definitions, reconciliation phase ordering, operator IPAM strategy, and component responsibility matrices. Hard dependencies (dispatcher core, fabr
Testability 2/2 Each design includes unit test plans (validation logic, pool selection, reuse logic), E2E integration tests (full lifecycle including multi-NIC, auto ExternalIP, auto NATGateway, cleanup), and tricky
Scope 2/2 Each service type (VMaaS, CaaS, BMaaS) has its own EP with clear non-goals. Default networking is separated as its own concern. Dependencies and GAPs are explicitly tracked. Graduation criteria define
Architecture 2/2 Consistent architecture across all service types: same NetworkAttachment pattern, same auto-provisioning two-phase lifecycle, same spec/status/feedback data flow. Clean separation of concerns (fulfill

Verdict: A comprehensive, well-structured design family that extends unified networking across VMaaS, CaaS, and BMaaS with consistent architectural patterns, concrete API definitions, and thorough test plans.

Feedback: Resolve the three shared open questions (NATGateway Deleting-state behavior, capacity exhaustion error-vs-Failed-resource, CaaS agent selection mechanism) before implementation begins, as they affect cross-cutting behavior in all service types. Create JIRA tickets for the 11 untracked GAP items identified in the dependency tables (CRD updates, mutateBMI, operator IPAM, HostType NetworkInterface list, etc.) to avoid planning gaps. Consider adding explicit metrics for IPAM allocation failures and auto-provisioning latency rather than relying solely on existing provisioning metrics.

Critical (0)

None.

Important (3)

  1. Eleven dependency items marked as GAP (not tracked) across BMaaS and CaaS designs need JIRA tickets before implementation planning: CRD updates, mutateBMI copy logic, operator IPAM implementation, dispatcher RBAC for Subnet/NetworkClass CRs, networkClass rename, HostType NetworkInterface list, step collection removal, NETWORK_STEPS_COLLECTION removal, agent selection logic, and fulfillment-service interface resolution.
  2. NATGateway reuse of Failed or Deleting instances is a known design weakness documented in all four service-type designs and the default-networking design. The current behavior silently attaches to a non-functional or disappearing resource. The open question about treating Deleting as 'does not exist' should be resolved before implementation, not deferred.
  3. The cross-operator IPAM concurrency model (osac-operator and bare-metal-fulfillment-operator both allocating from the same subnet CIDR) mentions optimistic concurrency on Subnet CR status but does not specify the tracking structure, conflict resolution strategy, or retry behavior. This needs more detail to avoid double-allocation bugs in production.

Suggestions (3)

  1. Add a sequence diagram to the CaaS design for the VIP feedback loop (template -> ClusterOrder status -> Signal RPC -> fulfillment-service -> Cluster -> ExternalIPAttachment controller), as this is the most complex cross-component flow and the drawbacks section already flags it as a complexity risk.
  2. Consider adding new metrics for IPAM allocation failures, auto-provisioning latency, and NATGateway reuse counts rather than relying solely on existing provisioning duration and failure rate metrics. The observability sections state 'no new metrics or alerts' but operator IPAM and auto-provisioning are new subsystems worth monitoring independently.
  3. The three open questions shared across all designs (NATGateway Deleting state, capacity exhaustion model, CaaS agent selection) should be consolidated into a single design decision document or tracked in the unified networking EP to avoid divergent resolutions across service types.

Review cost

Model: claude-opus-4-6
Cost: $0.8517
Tokens: 6.1k in / 3.3k out
Cache: 347.6k read
Active time: 1m 18s
API calls: 0

@github-actions

github-actions Bot commented Jul 9, 2026

Copy link
Copy Markdown

AI Design Review: EP-107

Score: 8/8 | Verdict: PASS

Criterion Score Notes
Feasibility 2/2 All five designs are thoroughly grounded in existing infrastructure — operator architectures, CRD patterns, proto definitions, AAP workflows, and dispatcher mechanisms are concretely specified. Depend
Testability 2/2 Each design includes comprehensive test plans with unit tests (validation logic, pool selection, reuse logic), E2E integration tests (full lifecycle per feature), and 'tricky test cases' (edge cases l
Scope 2/2 Each document has well-defined boundaries: clear non-goals (per-service scoping prevents scope creep), version-scoped limitations (e.g., 'v0.2: CaaS supports BM node sets only', 'v0.2: one attachment
Architecture 2/2 Sound architectural principles throughout: clean separation of concerns (fulfillment-service writes spec, operators write status, feedback syncs to DB), consistent dispatcher pattern for pluggable fab

Verdict: An exceptionally thorough set of five interconnected networking design documents that define a unified, consistent networking architecture across VMaaS, CaaS, and BMaaS with well-justified design decisions, comprehensive test plans, and clear dependency tracking.

Feedback: The cross-operator IPAM concurrency section (Unified Networking) states 'implementation should use optimistic concurrency on a shared tracking resource' but doesn't specify the concrete mechanism — clarify which resource (Subnet CR status with allocated IPs list?), what fields, and how conflicts are resolved, since osac-operator and bare-metal-fulfillment-operator allocating from the same subnet simultaneously is a critical correctness concern. The NATGateway Deleting-state open question appears in all four per-service documents and the Default Networking document — consider resolving it now (treating Deleting as 'does not exist' seems clearly correct since the SNAT rule is being removed) or at minimum consolidating it into the Unified Networking document with cross-references to reduce redundancy. Several implementation work items are marked 'Not tracked | GAP' across documents (CRD updates, mutateBMI, operator IPAM, agent selection logic) — track these in Jira before implementation b

Critical (0)

None.

Important (3)

  1. Cross-operator IPAM concurrency mechanism is underspecified — 'optimistic concurrency on a shared tracking resource' needs a concrete design (which resource, which fields, conflict resolution strategy) since osac-operator and bare-metal-fulfillment-operator may allocate from the same subnet simultaneously, risking IP collisions
  2. Several implementation work items are 'Not tracked' as Jira GAPs across all documents (BareMetalInstance CRD update, mutateBMI, operator IPAM, dispatcher RBAC, HostType NetworkInterface list, agent selection logic, step collection removal) — these represent unplanned work that could delay implementation
  3. NATGateway reuse of Failed/Deleting state is a known design weakness deferred to implementation — reusing a Failed NATGateway silently breaks outbound connectivity with no automatic recovery, and reusing a Deleting NATGateway attaches to a disappearing resource

Suggestions (3)

  1. Resolve the NATGateway Deleting-state open question before implementation — treating Deleting as 'does not exist' avoids silently attaching to a disappearing resource and is safe since the SNAT rule is being removed; consolidate this decision into the Unified Networking document rather than repeating the open question in all five documents
  2. Specify the concrete IPAM tracking mechanism in the Unified Networking document — e.g., 'Subnet CR status.allocatedIPs list with resourceVersion-based optimistic concurrency and retry on conflict' — so implementers across both operators have a shared contract
  3. Consider adding a reconciliation timeout or circuit breaker for the CaaS VIP feedback loop (template -> ClusterOrder status -> Signal RPC -> fulfillment-service -> Cluster -> ExternalIPAttachment controller) — failure in any step silently breaks the flow, and the current design only mentions requeue without a max-retry or alerting strategy

Review cost

Model: claude-opus-4-6
Cost: $1.0799
Tokens: 6.1k in / 3.9k out
Cache: 316.2k read
Active time: 1m 33s
API calls: 0

@github-actions

github-actions Bot commented Jul 9, 2026

Copy link
Copy Markdown

AI Design Review: EP-107

Score: 8/8 | Verdict: PASS

Criterion Score Notes
Feasibility 2/2 The design builds on the existing two-operator architecture, feedback controllers, and networking API with additive API extensions and backward-compatible field additions. Reconciliation phase orderin
Testability 2/2 Each EP includes a thorough test plan with unit tests (validation logic, IPAM, pool selection, NATGateway reuse), integration tests (E2E for multi-NIC, auto ExternalIP, auto NATGateway, VIP feedback l
Scope 2/2 Clean decomposition: one EP per service type (VMaaS, CaaS, BMaaS) plus shared Unified Networking EP and Default Networking EP. Each document clearly states goals, non-goals, and what is deferred (VM-b
Architecture 2/2 Sound separation of concerns: fulfillment-service handles API validation/DB/auto-provisioning, operators handle reconciliation/IPAM/switch-side dispatch, templates handle OS-level provisioning, fabric

Verdict: A comprehensive, well-architected set of enhancement proposals that cleanly decompose networking for VMaaS, CaaS, and BMaaS into per-service EPs built on a shared Unified Networking foundation, with consistent patterns for auto-provisioning, IPAM, cleanup, and feedback loops.

Feedback: Three areas to strengthen: (1) The cross-operator IPAM concurrency design (osac-operator vs bare-metal-fulfillment-operator allocating from the same subnet) needs more detail — specify the tracking resource schema, conflict detection mechanism, and retry behavior rather than deferring to implementation. (2) The duplicated open questions (NATGateway Deleting state, capacity exhaustion behavior) appear in 4 separate EPs — resolve these at the Unified Networking EP level and reference the decision from service EPs to avoid divergent implementations. (3) Track the 10+ untracked GAP dependencies in Jira now to prevent implementation surprises — several (CRD updates, mutateBMI changes, operator RBAC for Subnet/NetworkClass CRs) are on the critical path.

Critical (0)

None.

Important (3)

  1. Cross-operator IPAM concurrency between osac-operator and bare-metal-fulfillment-operator is under-specified — 'should use optimistic concurrency on a shared tracking resource' is insufficient for a correctness-critical component. Define the Subnet status schema for tracking allocated IPs, the conflict detection/retry protocol, and behavior under sustained contention.
  2. 10+ implementation tasks are listed as GAP (not tracked in Jira) across BMaaS and CaaS EPs, including critical-path items like CRD updates, mutateBMI changes, operator RBAC, HostType extensions, and agent selection logic. These should be tracked before the design is approved to ensure complete implementation planning.
  3. Auto NATGateway reusing a Deleting NATGateway is an unresolved open question appearing in all 4 service EPs. The current 'reuse regardless of state' design will cause silent failures when a NATGateway is being deleted concurrently. This should be resolved at the Unified Networking EP level before per-service implementation begins.

Suggestions (3)

  1. Consider adding a sequence diagram (mermaid) for the CaaS VIP feedback loop (template -> ClusterOrder status -> Signal RPC -> fulfillment-service -> Cluster -> ExternalIPAttachment controller) since it spans 5 components and is the most complex cross-component flow in the design.
  2. The cleanup pattern 'after N retries, finalizer is removed, parent deleted, orphaned resources left' should specify what N is and whether it is configurable, to help operators plan for orphan cleanup at scale.
  3. PRD 8.3 for BMaaS still references 'fabric manager writes the allocated IP to server status' as a risk, but the design resolved this with operator IPAM. Update the PRD risk section to match the resolved design.

Review cost

Model: claude-opus-4-6
Cost: $1.0772
Tokens: 6.1k in / 3.8k out
Cache: 317.6k read
Active time: 1m 30s
API calls: 0

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (2)
enhancements/default-networking/design.md (1)

196-198: 🗄️ Data Integrity & Integration | 🟠 Major | 🏗️ Heavy lift

Add rollback for the split NATGateway create path.

The separate ExternalIP/NATGateway transaction has no compensating cleanup if NATGateway creation fails, so pool capacity and the reserved IP can leak. Make the pair atomic or define the release path explicitly.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@enhancements/default-networking/design.md` around lines 196 - 198, The split
NATGateway creation path needs compensating cleanup because the separate
ExternalIP and NATGateway transaction can fail after reserving capacity, leaking
both the IP and pool usage. Update the NATGateway provisioning flow in the
design around the separate transaction so that the ExternalIP reservation is
either made atomic with NATGateway creation or explicitly rolled back on any
failure, and ensure the release path covers the auto-provisioned resource labels
and pool decrement/release logic.
enhancements/vmaas-networking/design.md (1)

161-161: 🗄️ Data Integrity & Integration | 🟠 Major | 🏗️ Heavy lift

Add rollback for the split NATGateway create path.

If the NATGateway write fails after the ExternalIP is allocated, the design leaks the IP and pool capacity. Make the pair atomic or spell out the compensating delete/release path.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@enhancements/vmaas-networking/design.md` at line 161, The split NATGateway
creation path in the design needs a rollback/compensation plan because
`NATGateway` and its auto-provisioned `ExternalIP` are created separately and
can leak pool capacity if the second step fails. Update the `NATGateway` create
flow to either make the `ExternalIP` allocation and `NATGateway` persist atomic,
or explicitly document and implement a compensating delete/release path that
cleans up the allocated IP and restores pool capacity when the `NATGateway`
write fails. Refer to the
`external_ip_mode=AUTO`/`osac.openshift.io/auto-provisioned` path and the
separate DB transaction described for the NATGateway creation.
♻️ Duplicate comments (1)
enhancements/bmaas-networking/design.md (1)

205-205: 🩺 Stability & Availability | 🟠 Major | ⚡ Quick win

Gate NATGateway reuse on Ready state.

nat_gateway_mode == AUTO still reuses a gateway regardless of state, so a deleting or failed NATGateway can strand outbound connectivity for the next BaremetalInstance.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@enhancements/bmaas-networking/design.md` at line 205, Update the
`nat_gateway_mode == AUTO` design in the NATGateway provisioning flow so reuse
only happens when an existing NATGateway is in Ready state, and treat
deleting/failed gateways as non-reusable. Adjust the `AUTO` branch description
around NATGateway lookup/reuse to reference the Ready condition explicitly,
while keeping the separate ExternalIP and separate DB transaction behavior
unchanged.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@enhancements/caas-networking/design.md`:
- Line 125: Update the NATGateway reuse logic in the AUTO branch so it only
reuses an existing gateway when it is in Ready state; if an existing NATGateway
is deleting, failed, or otherwise not Ready, treat it as unavailable and proceed
with creating a new one after the parent resource is persisted. Adjust the
decision point in the NATGateway provisioning flow to check the gateway’s status
before reuse, and keep the separate DB transaction and auto-provisioned labeling
behavior unchanged.

In `@enhancements/unified-networking/design.md`:
- Around line 565-569: Update the ExternalIPAttachment flow to read the VIPs
from the Cluster object, not from ClusterOrder status, so ownership stays with
the Cluster-backed source of truth. Adjust the design text around the
ExternalIPAttachment controller to reference the synced `api_endpoint` and
`ingress_endpoint` fields on Cluster, and remove any mention of `apiVIP` /
`ingressVIP` being consumed from ClusterOrder. Keep the wording aligned with the
Feedback controller and the ExternalIPAttachment controller to avoid implying
the wrong status contract.

---

Outside diff comments:
In `@enhancements/default-networking/design.md`:
- Around line 196-198: The split NATGateway creation path needs compensating
cleanup because the separate ExternalIP and NATGateway transaction can fail
after reserving capacity, leaking both the IP and pool usage. Update the
NATGateway provisioning flow in the design around the separate transaction so
that the ExternalIP reservation is either made atomic with NATGateway creation
or explicitly rolled back on any failure, and ensure the release path covers the
auto-provisioned resource labels and pool decrement/release logic.

In `@enhancements/vmaas-networking/design.md`:
- Line 161: The split NATGateway creation path in the design needs a
rollback/compensation plan because `NATGateway` and its auto-provisioned
`ExternalIP` are created separately and can leak pool capacity if the second
step fails. Update the `NATGateway` create flow to either make the `ExternalIP`
allocation and `NATGateway` persist atomic, or explicitly document and implement
a compensating delete/release path that cleans up the allocated IP and restores
pool capacity when the `NATGateway` write fails. Refer to the
`external_ip_mode=AUTO`/`osac.openshift.io/auto-provisioned` path and the
separate DB transaction described for the NATGateway creation.

---

Duplicate comments:
In `@enhancements/bmaas-networking/design.md`:
- Line 205: Update the `nat_gateway_mode == AUTO` design in the NATGateway
provisioning flow so reuse only happens when an existing NATGateway is in Ready
state, and treat deleting/failed gateways as non-reusable. Adjust the `AUTO`
branch description around NATGateway lookup/reuse to reference the Ready
condition explicitly, while keeping the separate ExternalIP and separate DB
transaction behavior unchanged.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository: osac-project/coderabbit/.coderabbit.yaml

Review profile: ASSERTIVE

Plan: Enterprise

Run ID: 04dc9314-b038-4793-b612-11abfcc47e2d

📥 Commits

Reviewing files that changed from the base of the PR and between 28309b7 and 16f23e2.

📒 Files selected for processing (5)
  • enhancements/bmaas-networking/design.md
  • enhancements/caas-networking/design.md
  • enhancements/default-networking/design.md
  • enhancements/unified-networking/design.md
  • enhancements/vmaas-networking/design.md

Comment thread enhancements/unified-networking/design.md Outdated
@github-actions

github-actions Bot commented Jul 9, 2026

Copy link
Copy Markdown

AI Design Review: EP-107

Score: 8/8 | Verdict: PASS

Criterion Score Notes
Feasibility 2/2 The design is grounded in existing, proven architecture (two-operator model, dispatcher pattern, feedback controller). API definitions are concrete in both protobuf and Go with CEL validation rules. R
Testability 2/2 Comprehensive test plan with unit tests (validation logic, pool selection, NATGateway reuse), integration/E2E tests (multi-NIC provisioning, auto ExternalIP, auto NATGateway, full connectivity, error
Scope 2/2 Cleanly scoped to BMaaS-only networking. Non-goals explicitly exclude CaaS/VMaaS, dispatcher infrastructure, and fabric manager implementation. The boundary between this EP and the unified networking
Architecture 2/2 Sound separation of concerns: fulfillment-service validates and creates CRs, operator handles IPAM and switch-side dispatch, AAP templates handle host-side config, feedback controller syncs status. Th

Verdict: A thorough, well-structured BMaaS networking enhancement with concrete APIs, clear operator responsibilities, honest dependency tracking, and comprehensive test coverage — ready for implementation pending resolution of three well-scoped open questions.

Feedback: The IPAM concurrency model needs explicit treatment: when multiple BaremetalInstances allocate IPs from the same subnet concurrently in reconcileNetworking, how are races prevented (optimistic locking on Subnet CR, lease-based allocation, etc.)? This is a real operational concern for multi-tenant environments. Second, PRD risk 8.3 ('IP address feedback mechanism fails — if the fabric manager does not write the allocated IP to server status') contradicts the design which has the operator managing IPAM, not the fabric manager — the PRD risk should be updated to match the resolved design. Third, consider resolving open question #1 (Deleting NATGateway reuse) before implementation, as silently attaching to a disappearing resource will create user-facing confusion that's hard to diagnose.

Critical (0)

None.

Important (2)

  1. IPAM concurrency not addressed: concurrent reconcileNetworking runs for multiple BaremetalInstances on the same subnet could allocate duplicate IPs. The design says 'reads the Subnet CR to obtain the CIDR, computes available IPs, picks the next available' but does not describe how concurrent allocations are serialized. This needs an explicit mechanism (e.g., optimistic locking on a Subnet IP tracking resource, or atomic IPAM allocation).
  2. PRD risk 8.3 contradicts resolved design: the PRD states 'If the fabric manager does not write the allocated IP to server status, external IP attachment cannot configure inbound NAT' but the design explicitly resolved this — the operator allocates IPs, not the fabric manager. The PRD risk should be updated to reflect the actual resolved approach (operator IPAM failure, not fabric manager writeback failure).

Suggestions (3)

  1. Open question Bump actions/setup-python from 5 to 6 #1 (should Deleting NATGateway be treated as 'does not exist') should be resolved before implementation rather than deferred — the current behavior of reusing a Deleting NATGateway creates a race condition where the BM's outbound connectivity silently breaks as the SNAT rule is removed mid-deletion.
  2. Consider explicitly rejecting lifecycle-role interfaces in server-side validation (FR-3) rather than relying on documentation convention. The design mentions 'Interfaces with role lifecycle are rejected' in server validation rules but the HostType interface list does not enforce this — a validation rule checking interface.role != 'lifecycle' would prevent misconfiguration.
  3. The 6 untracked dependency GAPs (CRD update, mutateBMI, operator IPAM, dispatcher RBAC, networkClass rename, HostType NetworkInterface list) should be tracked in Jira before implementation begins to avoid scope creep and ensure nothing is missed during sprint planning.

Review cost

Model: claude-opus-4-6
Cost: $0.6544
Tokens: 793 in / 4.0k out
Cache: 293.1k read
Active time: 1m 41s
API calls: 0

@github-actions

github-actions Bot commented Jul 27, 2026 •

Copy link
Copy Markdown

AI Design Review: EP-107

Score: 8/8 | Verdict: PASS

Criterion Score Notes
Feasibility 2/2 Deep technical detail with complete proto schemas for all new messages (NetworkDefaults, SecurityGroupRule, ComputeNetworkAttachment, ClusterNetworkAttachment, BareMetalNetworkAttachment, attachment statuses). All CRUD lifecycle operations covered per resource type with specific error codes and messages. Risks are specific and technical (ExternalIPPool exhaustion, fabric_manager implementation blocked, multi-job tracking dependency) with concrete mitigations and named reviewers. Drawbacks sectio
Testability 2/2 Each EP includes unit, integration/E2E, and tricky test case sections with specific scenarios. Unit tests specify validation logic (CIDR parsing, primary validation, interface validation, pool selection). Integration tests describe user-observable workflows (create Cluster with --external-ip-attachment, verify DNAT functional). Tricky tests cover edge cases (multi-NIC with primary on second attachment, ExternalIPPool exhaustion, cleanup failure with finalizer retry). Graduation criteria are conc
Scope 2/2 Five inter-related EPs decomposed cleanly: unified networking parent with per-service children (VMaaS, CaaS, BMaaS) plus default networking. Each EP has clear boundaries with specific non-goals (VM-based cluster node sets deferred, DNS API deferred, UI deferred). All PRDs referenced via frontmatter 'prd:' field. Alternatives sections include real alternatives with clear rejection rationale (per-tenant CIDR allocation, capacity exhaustion as Failed resource, single-operator architecture, keep net
Architecture 2/2 All OSAC patterns consistently followed across 5 EPs. Tenant isolation via osac.openshift.io/tenant annotation enforced on all new resources with OPA policies. Standard object shape (Spec/Status) with clear spec=desired-state, status=observed-state separation throughout. Controller patterns follow finalizer → status update → provisioning lifecycle with conditions-based status (DefaultNetworkingReady, NetworkingConfigured, InventoryAssigned, ProvisioningComplete, IPDiscoveryComplete). Phased requ

Verdict: Exceptionally thorough set of five inter-related networking design documents with complete proto schemas, detailed per-service workflows, specific risks with concrete mitigations, and comprehensive test plans — a strong pass across all four dimensions.

Feedback: The designs are well-structured and remarkably consistent across all five EPs. Two actionable improvements: (1) The tenant onboarding failure recovery ('delete and re-create tenant') is destructive and worth exploring a repair/retry mechanism that preserves tenant data. (2) The 'No new metrics or alerts' stance across all EPs is insufficient for a feature set this large — add at least default networking provisioning duration, auto-ExternalIP allocation rate, and pool utilization metrics to enable proactive capacity management.

Critical (0)

None.

Important (4)

  1. All five EPs state 'No new metrics or alerts (existing provisioning duration and failure rate metrics apply)' in their Observability sections. For a feature set introducing default networking at tenant onboarding, auto-ExternalIP allocation, and multi-phase controller reconciliation, dedicated metrics are needed: default networking provisioning duration, auto-ExternalIP allocation success/failure rate, ExternalIPPool utilization percentage, and VIP feedback loop latency. Without these, operators
  2. Tenant onboarding failure recovery is 'delete and re-create tenant' (default-networking design.md, Failure Handling section). This is destructive — it discards any tenant configuration, RBAC bindings, or resources created between tenant creation and default networking failure. Consider a repair-and-retry mechanism (e.g., reconcile default networking on next Tenant status update) instead of requiring full tenant re-creation.
  3. CaaS design has 5 items and BMaaS design has 6 items marked as 'Not tracked / GAP' in their dependency tables (e.g., HostType NetworkInterface fields, agent selection logic in operator, mutateBMI network_attachments copy, query_dhcp_lease role). These represent significant untracked implementation work that risks being overlooked during planning. Create Jira tickets for all GAP items before merging.
  4. Permanent cleanup failure threshold is unspecified across all EPs — 'after N retries, finalizer is removed' without defining N or the backoff strategy. This makes the behavior non-deterministic across environments and difficult to tune. Specify a concrete default (e.g., 10 retries with exponential backoff, max 5 minutes) and make it configurable.

Suggestions (3)

  1. SecurityGroupRule uses string fields for direction ('ingress'/'egress') and protocol ('tcp'/'udp'/'icmp'). Consider using proto enums for type safety and cleaner validation — this follows the API.md convention of preferring constrained types over open strings.
  2. The dual-field migration timeline for ComputeInstance (old field 14 vs new field 18) is 'TBD (OSAC-1471)'. Setting a concrete deprecation target (e.g., 'removed in v0.3') would help avoid indefinite dual-field maintenance burden.
  3. Consider adding a UX Alignment section proactively, even with UI deferred. The API shape (especially field names, label conventions, and error messages) directly impacts future UI implementation. Documenting the mapping now prevents retrofit work later.

Review cost

Model: claude-opus-4-6
Cost: $1.5298
Tokens: 16 in / 7.4k out
Cache: 1.3M read
Active time: 3m 9s
API calls: 0

Unified and default networking directories use OSAC-1433 prefix but
Jira tracking links pointed to OSAC-1029. Updated to OSAC-1433 for
consistency.

Assisted-by: Claude Code <noreply@anthropic.com>
Signed-off-by: Dan Manor <dmanor@redhat.com>
@github-actions

Copy link
Copy Markdown

AI EP Review: EP-107

Score: 10/10 | Verdict: PASS

Criterion Score Notes
WHAT (clear need) 2/2 All five PRDs clearly describe user-facing capabilities with affected personas and per-persona user stories. The unified PRD identifies VMaaS/CaaS/BMaaS services and Tenant/Provider personas. Each per-service PRD covers Tenant User, Tenant Admin, Cloud Infrastructure Admin, and Cloud Provider Admin with concrete 'As a...' stories. The default networking PRD identifies exactly what tenants can do (single-call resource creation, auto ExternalIP, inspect/customize defaults). OSAC dimensions (servic
WHY (justification) 2/2 Strong business justification across all PRDs. The unified PRD identifies 9 specific gaps in the current design (CaaS/BMaaS have no networking API, tenants must choose backends, ExternalIPAttachment only supports VMs, etc.). The default networking PRD quantifies friction ('6+ sequential API calls') and names consequences ('slows onboarding, increases misconfiguration, makes OSAC harder to adopt'). Per-service PRDs name concrete pain: VMaaS cannot create multi-NIC VMs; CaaS has zero tenant contro
User-Facing Focus 2/2 PRDs describe user-observable outcomes without prescribing implementation. Requirements use platform vocabulary (VirtualNetwork, Subnet, SecurityGroup, ExternalIP, ClusterOrder) which are user-facing API resources, not internal components. No controllers, reconcilers, playbooks, or internal conditions are named. The CaaS PRD's FR-6/FR-7 describe system behavior ordering (host selection before provisioning) but frame it as observable preconditions, not controller logic. The BMaaS PRD's FR-10 desc
Right-Sized 2/2 Each PRD is individually well-scoped. The unified PRD covers the foundational resource model that all services need. Each per-service PRD (VMaaS, CaaS, BMaaS) is scoped to one service type's networking requirements. The default networking PRD bundles default resources, optional attachments, auto ExternalIP, and NATGateway under 'simplified creation' — these capabilities depend on each other (optional attachments require defaults, auto ExternalIP builds on the simplified flow). No PRD bundles unr
Testability 2/2 Every requirement and acceptance criterion across all five PRDs can be verified by a PM or QA engineer using the product. Acceptance criteria are structured as testable checklist items: 'A Tenant User can create a VM with --external-ip-attachment and no explicit network attachments — the VM is created on the default subnet with an auto-provisioned external IP.' The BMaaS PRD includes measurable NFRs (NFR-2: 'Network attachment provisioning completes within 2 minutes per interface'). No requireme

Verdict: Exceptionally well-structured PRD suite that cleanly separates a large networking initiative into five focused, individually coherent PRDs with clear user-facing requirements, concrete business justification, and testable acceptance criteria across all four OSAC personas.

Feedback: This is a strong PRD suite. Minor improvement opportunities: (1) The CaaS PRD's FR-6 ('The system selects and reserves suitable bare-metal hosts...to prevent allocation conflicts') borders on describing internal orchestration — consider reframing as 'The system ensures sufficient hosts are available before cluster provisioning begins' to keep focus on the user-observable outcome. (2) Consider adding explicit 'In Scope / Out of Scope' sections to each per-service PRD to make scope boundaries even clearer, since each PRD's non-goals reference the others. (3) The BMaaS PRD's success metrics (FR-7 < 5min, NFR-2 < 2min) are excellent — consider adding similar measurable targets to the other per-service PRDs for consistency.

Critical (0)

None.

Important (0)

None.

Suggestions (3)

  1. CaaS PRD FR-6/FR-7/FR-8 describe internal orchestration ordering (host selection → network config → provisioning). While not naming specific controllers, the ordering constraints are implementation concerns a PM wouldn't directly test. Consider reframing as user-observable outcomes: 'When a cluster is created, the system ensures hosts have network connectivity before cluster setup begins.'
  2. The BMaaS PRD has explicit success metrics (provisioning time < 5min, per-interface config < 2min) that the other per-service PRDs lack. Adding similar measurable targets to VMaaS and CaaS PRDs would strengthen testability consistency across the suite.
  3. Each per-service PRD references the unified and default networking PRDs as dependencies but doesn't specify which version or milestone is required. Adding milestone-level dependency tracking (e.g., 'requires unified networking GA' vs 'requires unified networking tech preview') would clarify sequencing for planning.

Review cost

Model: claude-opus-4-6
Cost: $1.7092
Tokens: 17 in / 7.1k out
Cache: 1.2M read
Active time: 2m 48s
API calls: 0

@github-actions

github-actions Bot commented Jul 27, 2026 •

Copy link
Copy Markdown

AI Design Review: EP-107

Score: 8/8 | Verdict: PASS

Criterion Score Notes
Feasibility 2/2 Deeply technical across all five design documents. Full proto schemas with field types and validation annotations for every new resource (NetworkDefaults, SecurityGroupRule, ComputeNetworkAttachment, BareMetalNetworkAttachment, ClusterNetworkAttachment, per-service status messages). All CRUD lifecycle operations covered with specific error codes (e.g., 'ExternalIPPool exhaustion: no available capacity in any READY pool for IPv4'). Auto-provisioning lifecycle described in granular two-phase detai
Testability 2/2 Each of the five design documents includes a structured Test Plan with unit tests (validation logic, state transitions, pool selection), integration/E2E tests (specific scenarios: multi-NIC provisioning, auto ExternalIP, cleanup, pool exhaustion, BM-only region rejection), and a 'Tricky Test Cases' section identifying edge cases (primary on second attachment, cleanup failure, VIP feedback loop failure, IP feedback latency). Graduation criteria provide measurable Tech Preview and GA checklists wi
Scope 2/2 Well-decomposed into five cohesive design documents: unified networking (shared architecture), default networking (tenant onboarding), and three per-service designs (VMaaS, CaaS, BMaaS). Each design has clear boundaries with specific non-goals (e.g., 'VM-based cluster node sets deferred', 'DNS API stays inline', 'per-tenant CIDR allocation rejected'). All PRD references present in frontmatter. Alternatives sections include real alternatives with substantive rejection rationale. Cross-cutting dim
Architecture 2/2 All OSAC patterns followed consistently across the five designs. Tenant isolation metadata (osac.openshift.io/tenant, osac.openshift.io/owner-reference) on all new resources. API conventions per fulfillment-service/docs/API.md: standard object shape with spec/status separation, spec for desired state (user-controlled), status for observed state (system-controlled). Controller patterns: finalizer → status update → provisioning lifecycle. Conditions used for lifecycle state (DefaultNetworkingReady

Verdict: Exceptionally thorough five-document networking design suite covering unified architecture, default networking, and per-service flows for VMaaS/CaaS/BMaaS — deeply technical with proto schemas, controller reconciliation patterns, IP discovery mechanisms, and comprehensive test plans across all dimensions.

Feedback: The main actionable items are: (1) assign concrete proto field numbers for placeholder 'N' fields before implementation (ComputeNetworkAttachmentStatus, BareMetalNetworkAttachmentStatus) to avoid field number conflicts; (2) track the GAP items in the CaaS and BMaaS dependency tables as Jira tickets — untracked work is invisible to planning; (3) establish the dual-field migration timeline for VMaaS (OSAC-1471 is 'TBD') to give consumers a deprecation horizon.

Critical (0)

None.

Important (3)

  1. Proto field numbers use placeholder 'N' in several status messages (ComputeInstanceStatus.compute_network_attachment_statuses, BareMetalInstanceStatus.network_attachment_statuses, ClusterStatus.api_endpoint/ingress_endpoint) — these must be assigned concrete numbers before implementation to avoid field number conflicts across proto versions.
  2. CaaS and BMaaS dependency tables include 5+ items marked GAP (not tracked in Jira): HostType NetworkInterface fields, agent selection logic in operator, fulfillment-service interface resolution, step collection removal, mutateBMI copy, query_dhcp_lease role, BareMetalInstanceType network ports, etc. Untracked dependencies risk being forgotten during implementation.
  3. The query_dhcp_lease IP discovery mechanism for BMaaS introduces a novel dependency on the fabric manager's DHCP lease API (querying by subnet + matching port MAC to lease). This is a new integration surface not validated in any existing flow — ensure the fabric manager actually exposes this API and test the MAC-to-IP matching in integration tests before relying on it.

Suggestions (3)

  1. The dual-field migration timeline for VMaaS (network_attachments field 14 → compute_network_attachments field 18) is marked 'TBD (OSAC-1471)' — establishing even a rough deprecation horizon (e.g., 'removed in v0.3') would help tenants and CLI consumers plan migration.
  2. Consider adding an explicit IP discovery comparison table in the unified networking design's auto-provisioning lifecycle section — the per-target-type precondition table exists but a reader must cross-reference three separate per-service designs to understand the full picture.
  3. BMaaS design introduces reconcileIPDiscovery as a distinct reconciliation phase (inventory → networking → provisioning → IP discovery → power). Consider documenting whether this phase ordering pattern should become standard for all resource types or is BMaaS-specific, to guide future controller authors.

Review cost

Model: claude-opus-4-6
Cost: $1.2361
Tokens: 1.5k in / 5.3k out
Cache: 711.0k read
Active time: 1m 59s
API calls: 0

BMaaS PRD: "primary-traffic" → "fabric" role example in FR-2.
CaaS PRD: removed k8s_manager/MetalLB implementation details from
resolved OQ 9.3.
Unified PRD: removed proto field name from Gap osac-project#1, removed proto
jargon (oneof, compute_instance) from Gap osac-project#4.

Assisted-by: Claude Code <noreply@anthropic.com>
Signed-off-by: Dan Manor <dmanor@redhat.com>
@github-actions

Copy link
Copy Markdown

AI EP Review: EP-107

Score: 10/10 | Verdict: PASS

Criterion Score Notes
WHAT (clear need) 2/2 Clear, specific user-facing capabilities across five well-decomposed PRDs. All four OSAC personas (Tenant User, Tenant Admin, Cloud Infrastructure Admin, Cloud Provider Admin) have dedicated user story headings with concrete stories in each per-service PRD. Services (VMaaS, CaaS, BMaaS) explicitly identified. Cross-cutting dimensions covered: tenant onboarding, networking, provisioning, and installation. The capabilities are unambiguously user-observable: single-call resource creation, auto exte
WHY (justification) 2/2 Concrete, quantified justification. The default networking PRD names the specific pain ('6+ sequential API calls for a reachable resource'), describes the consequence ('friction slows onboarding, increases misconfiguration, makes OSAC harder to adopt'), and ties it to a strategic goal (competitive parity with platforms offering single-command provisioning). The unified networking PRD identifies 9 specific gaps with detailed impact. Each per-service PRD has a clear problem statement naming what t
User-Facing Focus 2/2 PRDs consistently describe user-observable outcomes without prescribing implementation. Platform vocabulary (NetworkClass, ExternalIPPool, SecurityGroup, Subnet, VirtualNetwork) is used appropriately per the rubric. No controllers, reconcilers, playbooks, finalizers, or internal conditions are named in the PRD requirements or acceptance criteria. Design documents contain implementation details, while PRDs stay focused on what users can do and observe. Minor note: BMaaS FR-10 mentions 'network au
Right-Sized 2/2 Excellent decomposition. Each PRD covers a coherent set of tightly coupled capabilities: the default networking PRD bundles defaults-at-onboarding + optional attachments + auto ExternalIP (these require each other for the 'single API call' value proposition). Each per-service PRD is scoped to one service type's specific networking needs. The unified networking PRD covers the shared foundation. Capabilities within each PRD cannot ship independently and provide value — e.g., optional attachments w
Testability 2/2 Every acceptance criterion across all five PRDs is verifiable by a PM or QA engineer using the product. Examples: 'create a ComputeInstance with --external-ip-attachment and no explicit network attachments — the VM is created on the default subnet with an auto-provisioned ExternalIP' (CLI test), 'Creating a VM in a bare-metal-only region returns an error with a clear message' (error path test), 'VM status shows the allocated IP address for each network attachment after provisioning completes' (A

Verdict: Exceptionally well-structured set of PRDs with clear user-facing needs, concrete business justification, clean separation from design, well-scoped decomposition, and fully testable acceptance criteria across all five documents.

Feedback: The unified networking PRD groups provider stories under a generic 'Provider Stories' heading rather than splitting by OSAC persona (Cloud Infrastructure Admin vs Cloud Provider Admin) — align with the per-service PRDs for consistency. Consider cleaning up resolved open questions (marked with strikethrough) across all PRDs, as they add length without informational value in the published document. The CaaS PRD FR-8 mentions 'performs DNS record creation' while DNS API is listed as a non-goal — add a clarifying note that DNS remains template-based (not tenant-controlled) to avoid reviewer confusion.

Critical (0)

None.

Important (0)

None.

Suggestions (3)

  1. Unified networking PRD: split 'Provider Stories' into 'Cloud Infrastructure Admin Stories' and 'Cloud Provider Admin Stories' headings to match per-service PRDs and OSAC persona model
  2. All PRDs: remove resolved open questions (strikethrough sections) before publishing — they served their purpose during drafting but add noise in the final document
  3. CaaS PRD FR-8: clarify that DNS record creation is template-managed (not tenant-facing API), since DNS API is explicitly a non-goal — avoids apparent contradiction

Review cost

Model: claude-opus-4-6
Cost: $1.2983
Tokens: 17 in / 9.0k out
Cache: 1.1M read
Active time: 3m 25s
API calls: 0

@github-actions

Copy link
Copy Markdown

AI Design Review: EP-107

Score: 8/8 | Verdict: PASS

Criterion Score Notes
Feasibility 2/2 Exceptionally deep implementation detail across all 5 designs. Full proto schemas with field numbers and types, Go CRD types, CEL validation rules, SQL migration references, AAP template changes, dispatcher call patterns, and reconciliation phase ordering. All lifecycle operations (create/get/list/update/delete) covered. Error handling is specific with concrete error messages (e.g., 'ExternalIPPool exhaustion: no available capacity in any READY pool for IPv4'). Risks are specific technical risks
Testability 2/2 Each design has structured test plans with Unit Tests, Integration Tests, and Tricky Test Cases. Unit tests specify concrete scenarios (e.g., 'primary validation: reject >1 primary, accept single implicit primary', 'fabric_interface resolution per node set'). E2E tests describe user-observable scenarios with specific verification points. Graduation criteria are measurable checklists with concrete conditions (e.g., 'Auto ExternalIP attachment provisioning functional', 'Integration tests pass'). M
Scope 2/2 All designs reference their PRDs via frontmatter ('prd: prd.md'). Non-goals are specific: 'Custom default configurations per tenant', 'VM-based cluster node sets (deferred)', 'DNS API', 'Multi-NIC cluster nodes'. Each design has real alternatives with concrete rejection rationale. Cross-cutting dimensions are well-covered: tenant onboarding (core focus of default-networking), networking (primary domain), provisioning (extensive per-service coverage), installation (osac-installer changes), E2E te
Architecture 2/2 All OSAC patterns followed consistently across all 5 designs. Tenant isolation via osac.openshift.io/tenant annotation on all new resources, enforced by OPA. Auto-provisioned resources labeled with osac.openshift.io/auto-provisioned and auto-provisioned-for. Controller patterns follow finalizer → status update → provisioning lifecycle. Conditions used for lifecycle (DefaultNetworkingReady, NetworkingConfigured, InventoryAssigned). Dependencies between components clearly identified with Component

Verdict: An exceptionally comprehensive family of 5 interrelated networking designs that follow all OSAC patterns, provide deep implementation detail with proto schemas and Go types, clearly delineate scope with specific non-goals and real alternatives, and include concrete test plans with measurable graduation criteria — scoring a perfect 8/8.

Feedback: The designs are production-quality and ready for implementation. Two minor items worth addressing: (1) Replace placeholder proto field numbers ('= N') in ComputeInstanceStatus and BareMetalInstanceStatus with concrete field numbers to avoid ambiguity during implementation. (2) The CaaS and BMaaS dependency tables list several items as 'Not tracked / GAP' — consider creating Jira tickets for these before merge to avoid implementation blind spots (HostType NetworkInterface fields, agent selection logic in operator, mutateBMI network_attachments copy, etc.).

Critical (0)

None.

Important (3)

  1. Proto field numbers placeholder '= N' in ComputeInstanceStatus.compute_network_attachment_statuses and BareMetalInstanceStatus.network_attachment_statuses — should be concrete field numbers to avoid ambiguity during implementation (default-networking design line ~560, VMaaS design, BMaaS design)
  2. CaaS design Dependencies table lists 5 items as 'Not tracked / GAP' including HostType NetworkInterface fields, agent selection logic in operator, and step collection removal — these are prerequisites with no Jira tracking
  3. BMaaS design Dependencies table lists 6 items as 'Not tracked / GAP' including BareMetalInstance CRD NetworkAttachments, mutateBMI changes, IP discovery query_dhcp_lease role, operator RBAC, BareMetalInstanceType network ports, and unused networkClass field removal

Suggestions (3)

  1. Consider adding explicit Documentation dimension coverage in each per-service design (currently only mentioned in graduation criteria)
  2. Test plan sections label all tests as 'E2E:' under 'Integration Tests' heading — clarifying which are kind-cluster integration tests vs. full-stack E2E would help QE planning
  3. Per-service designs reference the unified networking terminology implicitly — consider adding a brief Terminology section or explicit cross-reference to the unified networking Terminology section for standalone readability

Review cost

Model: claude-opus-4-6
Cost: $1.6461
Tokens: 18 in / 5.9k out
Cache: 1.5M read
Active time: 2m 21s
API calls: 0

…uture

BMaaS design: reverted to HostType for interface validation (v0.2).
Added "Future: BareMetalInstanceType Integration" section documenting
the migration plan when PR osac-project#119 lands with enhanced network_ports.

HostType is the system-level resource that exists today and works for
both CaaS and BMaaS. BareMetalInstanceType will be the tenant-facing
catalog once it's available.

Assisted-by: Claude Code <noreply@anthropic.com>
Signed-off-by: Dan Manor <dmanor@redhat.com>
@github-actions

Copy link
Copy Markdown

AI EP Review: EP-107

Score: 9/10 | Verdict: PASS

Criterion Score Notes
WHAT (clear need) 2/2 All five PRDs clearly describe user-facing capabilities with concrete user stories grouped by persona (Tenant User, Tenant Admin, Cloud Infrastructure Admin, Cloud Provider Admin). Each affected persona has explicit 'As a ...' stories. OSAC dimensions well-covered: all three services (VMaaS, CaaS, BMaaS) in scope, all four personas addressed, cross-cutting dimensions (networking, provisioning, tenant onboarding) addressed, UI explicitly deferred as non-goal. API resources (VirtualNetwork, Subnet
WHY (justification) 2/2 Strong justification across all PRDs. Default networking: '6+ sequential API calls... friction slows onboarding, increases the chance of misconfiguration, and makes OSAC harder to adopt.' Unified networking: thorough 9-gap analysis identifying concrete deficiencies (CaaS/BMaaS have no networking API, tenants must choose backends, ExternalIPAttachment only supports VMs, etc.). Per-service PRDs each name specific pain: VMaaS cannot do multi-NIC, CaaS has zero tenant control over cluster networking
User-Facing Focus 2/2 PRDs consistently describe user-observable outcomes. Requirements are framed as what users can do ('A tenant can create a VM with multiple network interfaces'), not how the system implements it internally. Platform vocabulary (NetworkClass, VirtualNetwork, ExternalIP, etc.) is used appropriately per the rubric's OSAC Platform Vocabulary list. No controllers, reconcilers, playbooks, or internal conditions appear in the functional requirements. The few system-level mentions (e.g., 'the system conf
Right-Sized 1/2 Each individual PRD is well-scoped and coherent: default networking (provisioning + defaults + auto ExternalIP), VMaaS (multi-NIC + auto external access), CaaS (cluster networking + VIP feedback), BMaaS (interface mapping + switch port config). However, the PR bundles 5 PRDs + 5 design documents — the per-service PRDs (VMaaS, CaaS, BMaaS) are independent from each other and could ship as separate PRs with their own review cycles. This is closer to an epic than a single feature scope.
Testability 2/2 All PRDs have detailed, user-verifiable acceptance criteria. Examples: 'A Tenant User can create a ComputeInstance with --external-ip-attachment and no explicit network attachments', 'Creating a bare-metal server with an invalid interface returns an error', 'Cluster status exposes API server and ingress endpoint addresses after provisioning completes'. Every criterion can be verified by a PM or QA engineer using the product — no internal-only observations required. BMaaS PRD includes success met

Verdict: A strong set of PRDs with clear user-facing capabilities, concrete justification, and verifiable acceptance criteria across all five documents. The only weakness is that the PR bundles five independent PRDs (unified, default, VMaaS, CaaS, BMaaS) that could be reviewed and shipped as separate PRs.

Feedback: Consider splitting the per-service PRDs (VMaaS, CaaS, BMaaS) into separate PRs to simplify review and allow independent prioritization — each is self-contained with its own user stories and acceptance criteria. The unified networking and default networking PRDs could remain together as they form the foundation. The unified networking PRD's requirements section (Section 4.1) lost its per-requirement acceptance criteria during the restructuring from the old format — adding inline testability markers (e.g., 'verified by...') to each FR would strengthen traceability between requirements and the acceptance criteria in Section 5.

Critical (0)

None.

Important (2)

  1. PR bundles 5 independent PRDs (unified, default, VMaaS, CaaS, BMaaS) plus 5 design documents. The per-service PRDs are independent from each other and should ideally be separate PRs for easier review and independent shipping. Recommend splitting into: (1) unified networking + its design, (2) default networking + its design, (3-5) one PR per service type.
  2. Unified networking PRD Section 4.1 functional requirements (FR-1 through FR-7) lost their per-requirement acceptance criteria during restructuring. The old format had inline 'Acceptance criteria' under each requirement; the new format moves all criteria to Section 5 without explicit traceability. A reviewer cannot easily verify which acceptance criteria cover which requirement.

Suggestions (3)

  1. BMaaS PRD FR-5 references 'the host type's default interface' for when network attachments are omitted, but does not define what 'default interface' means. The design document clarifies it as 'first interface with role fabric,' but the PRD should be self-contained — add 'using the first interface with the primary traffic role' or similar.
  2. CaaS PRD NFR-1 states 'Endpoint addresses are available in cluster status during provisioning, not minutes later' — this is imprecise. The endpoints only appear after MetalLB allocates VIPs, which happens during provisioning. Consider rewording to 'Endpoint addresses appear in cluster status as soon as they are allocated during provisioning.'
  3. Consider adding a success metrics table to the default networking PRD (the unified and BMaaS PRDs have one). Metrics like 'API calls to create a reachable resource: target 1, baseline 6+' would strengthen the justification.

Review cost

Model: claude-opus-4-6
Cost: $1.7046
Tokens: 17 in / 7.5k out
Cache: 1.3M read
Active time: 3m 14s
API calls: 0

@github-actions

Copy link
Copy Markdown

AI Design Review: EP-107

Score: 8/8 | Verdict: PASS

Criterion Score Notes
Feasibility 2/2 Deep technical detail across all five design documents. Proto schemas included with field types and validation annotations for every new resource and modification. All CRUD lifecycle operations covered per resource type. Error handling and failure modes described exhaustively (pool exhaustion, cleanup failures, feedback loop failures, two-operator synchronization). Risks are specific technical risks with concrete mitigations (e.g., 'concurrent subnet allocation may cause CIDR overlap' → 'fabric-
Testability 2/2 Each of the five designs includes a three-level test plan: unit tests specifying what's tested (CIDR validation, primary validation, interface validation, pool selection, phase ordering), integration tests with concrete E2E scenarios (create Tenant → verify defaults, create ComputeInstance with --external-ip-attachment → verify DNAT functional, delete → verify cleanup), and tricky test cases covering edge cases (multi-NIC primary on second attachment, pool exhaustion, cleanup failure, VIP feedba
Scope 2/2 Each design has clear boundaries with specific non-goals (e.g., 'VM-based cluster node sets deferred', 'DNS API stays inline', 'UI support deferred'). All designs reference their PRDs via frontmatter. Non-goals are specific about what's excluded and why. Alternatives sections compare real approaches with trade-offs (e.g., per-tenant CIDR allocation vs. single default, Failed resource vs. API error on exhaustion). The decomposition into unified EP + four per-service EPs + default networking EP pr
Architecture 2/2 All OSAC architectural patterns followed consistently: tenant isolation via osac.openshift.io/tenant annotations enforced by OPA, owner references via osac.openshift.io/owner-reference, controller patterns (finalizer → status update → provisioning lifecycle), conditions for lifecycle state (DefaultNetworkingReady, NetworkingConfigured, etc.), spec/status ownership separation. Cross-repo impacts enumerated in component responsibility tables across fulfillment-service, osac-operator, bare-metal-fu

Verdict: Exceptionally thorough design suite covering unified networking across VMaaS, CaaS, and BMaaS with consistent architectural patterns, detailed proto schemas, comprehensive failure handling, and concrete test plans — minor gaps around unspecified field number placeholders and missing observability metrics do not materially weaken the design.

Feedback: Three areas to tighten before implementation: (1) Replace all '= N' proto field number placeholders with actual assigned numbers in ComputeNetworkAttachmentStatus, BareMetalNetworkAttachmentStatus, and ClusterStatus messages — implementers need concrete field assignments. (2) Specify the retry count and backoff strategy for 'permanent failure' during auto-provisioned resource cleanup instead of 'after N retries' — this is a production-critical parameter that affects orphan cleanup behavior. (3) All five designs state 'No new metrics or alerts' despite introducing significant new failure modes (pool exhaustion, default networking provisioning failure, cleanup orphans); add at least pool capacity gauge metrics and default networking provisioning failure counters to enable operational alerting.

Critical (0)

None.

Important (4)

  1. Proto field number placeholders: Several status messages use '= N' instead of assigned field numbers (ComputeInstanceStatus.compute_network_attachment_statuses, BareMetalInstanceStatus.network_attachment_statuses, ClusterStatus.api_endpoint/ingress_endpoint). These must be assigned before implementation to avoid wire-format conflicts.
  2. Unspecified cleanup retry parameters: All designs reference 'after N retries, finalizer is removed' for permanent cleanup failure, but the retry count, backoff interval, and timeout are never specified. This is a production-critical parameter that determines orphan creation rate.
  3. No new metrics despite new failure modes: All five designs declare 'No new metrics or alerts (existing provisioning duration and failure rate metrics apply)' but introduce ExternalIPPool exhaustion, default networking provisioning failure, cleanup orphans, and VIP feedback loop failures — these need dedicated metrics for operational alerting.
  4. CaaS and BMaaS designs list multiple untracked dependency GAPs (HostType NetworkInterface fields, agent selection logic in operator, interface resolution in fulfillment-service, unused networkClass field removal) — these should be tracked in Jira before implementation begins to avoid discovery-phase delays.

Suggestions (3)

  1. SecurityGroupRule proto uses string fields for direction and protocol — consider adding validation documentation or server-side enum checks to prevent typos like 'inbound' instead of 'ingress'.
  2. BMaaS reconcileIPDiscovery should specify a timeout for DHCP lease query (similar to CaaS's NetworkingIPDiscoveryTimeout condition at 5 minutes) to prevent indefinite blocking if the host never receives a DHCP lease.
  3. Consider adding a cross-design sequence diagram showing the end-to-end flow from tenant onboarding through resource creation to external access — the individual workflows are clear but the full system interaction across all five designs would benefit reviewers.

Review cost

Model: claude-opus-4-6
Cost: $1.6334
Tokens: 21 in / 5.6k out
Cache: 1.6M read
Active time: 2m 28s
API calls: 0

…on non-goal

Default design: clarified that NATGateway creation also fails without
ExternalIPPool (not just auto external access).

Default PRD: added non-goal about existing tenant migration matching
the design's upgrade section.

Assisted-by: Claude Code <noreply@anthropic.com>
Signed-off-by: Dan Manor <dmanor@redhat.com>
@github-actions

Copy link
Copy Markdown

AI EP Review: EP-107

Score: 9/10 | Verdict: PASS

Criterion Score Notes
WHAT (clear need) 2/2 Each PRD clearly describes user-facing capabilities: default networking at tenant onboarding, optional network attachments, auto ExternalIP, multi-NIC VMs, cluster networking, BM server networking. All four OSAC personas (Tenant User, Tenant Admin, Cloud Infrastructure Admin, Cloud Provider Admin) are identified and have per-persona user stories across all five PRDs. Services (VMaaS, CaaS, BMaaS) are explicitly scoped. Cross-cutting dimensions (tenant onboarding, provisioning, networking) are ad
WHY (justification) 2/2 Strong justification across all PRDs. The default networking PRD names specific pain: '6+ sequential API calls' to create a reachable resource, friction that 'slows onboarding, increases misconfiguration, and makes OSAC harder to adopt.' The unified networking PRD identifies 9 concrete gaps with specific consequences. Each per-service PRD has a targeted problem statement (e.g., CaaS: 'zero tenant control' over cluster networking). BMaaS PRD includes quantified success metrics (provisioning time
User-Facing Focus 1/2 The per-service PRDs (VMaaS, CaaS, BMaaS) and the default networking PRD are clean — they describe user-observable outcomes without prescribing controllers, reconcilers, or playbooks. However, the unified networking PRD's acceptance criteria include some architectural statements that are not directly user-observable: 'A single networking backend handles all physical networking operations' (a PM cannot verify how many backends handle operations), 'VM networking is integrated into the same network
Right-Sized 2/2 Each PRD is individually well-scoped. The unified networking PRD covers the shared resource model (tightly coupled — VirtualNetwork, Subnet, SecurityGroup, ExternalIP form one coherent model). Each per-service PRD focuses on how one service type consumes that model (multi-NIC for VMs, host selection for clusters, interface mapping for BM). The default networking PRD covers defaults + auto ExternalIP (provisioning without visibility is incomplete). The multi-PRD structure with explicit dependency
Testability 2/2 Every acceptance criterion across all five PRDs can be verified by using the product. Examples: 'A Tenant User can create a ComputeInstance with --external-ip-attachment and no explicit network attachments' (testable via CLI), 'Creating a bare-metal server with an invalid interface returns an error' (testable), 'When no ExternalIPPool has available capacity, the create API call returns an error and the resource is not persisted' (testable). No acceptance criteria describe internal system state o

Verdict: A strong, well-structured set of PRDs that clearly define user-facing networking capabilities across VMaaS, CaaS, and BMaaS with concrete justification and testable requirements; minor design leakage in the unified networking PRD's acceptance criteria prevents a perfect score.

Feedback: The unified networking PRD's acceptance criteria should be rewritten to describe user-observable outcomes rather than architectural properties. Replace 'A single networking backend handles all physical networking operations' with something a PM could verify (e.g., 'A provider can add a new networking backend by deploying configuration without modifying the API or operator code'). Similarly, 'VM networking is integrated into the same networking layer' should describe the user-visible consequence (e.g., 'A VM and a bare-metal server on the same subnet can communicate at their subnet IPs'). The per-service PRDs are exemplary — apply the same discipline to the unified networking PRD's acceptance criteria.

Critical (0)

None.

Important (2)

  1. Unified networking PRD acceptance criteria contain architectural statements ('A single networking backend handles all physical networking operations', 'VM networking is integrated into the same networking layer as bare-metal servers') that describe system architecture rather than user-observable outcomes. Rewrite to describe what a PM would verify by using the product.
  2. The unified networking PRD's problem statement section (Gap adds Apache 2.0 LICENSE file #3, Gap Document the enhancement proposal process #6) references internal components (K8s manager, ConfigMap, controller-runtime) without clarifying these are existing-state context. While acceptable as problem description, it sets a precedent that bleeds into acceptance criteria.

Suggestions (3)

  1. Consider adding explicit success metrics to the VMaaS and CaaS PRDs, as the BMaaS PRD does (provisioning time, success rate). The unified networking PRD has success metrics, but the per-service PRDs (except BMaaS) lack quantified targets.
  2. The default networking PRD's FR-1 mentions 'status condition describing the failure' — consider specifying what the user sees (e.g., 'the tenant status shows the reason for the failure') rather than the mechanism ('status condition').
  3. CaaS PRD's NFR-1 ('Endpoint addresses are available in cluster status during provisioning, not minutes later') is somewhat vague about timing. Consider specifying a measurable target like the BMaaS PRD's NFR-2 ('within 2 minutes').

Review cost

Model: claude-opus-4-6
Cost: $1.3435
Tokens: 20 in / 8.3k out
Cache: 1.3M read
Active time: 3m 5s
API calls: 0

@github-actions

Copy link
Copy Markdown

AI Design Review: EP-107

Score: 8/8 | Verdict: PASS

Criterion Score Notes
Feasibility 2/2 Deep technical detail across all five EPs: full proto schemas with field numbers aligned to codebase, Go CRD types with CEL validation rules, component responsibility tables, reconciliation phase ordering, specific error codes and validation rules. All CRUD lifecycle operations covered for each resource type. Risks are specific technical risks (k8s_manager blocked, ExternalIPPool exhaustion, two-operator synchronization) with concrete mitigations. Drawbacks sections steel-man the trade-offs. Som
Testability 2/2 Each EP has structured test plans at three levels: unit tests specifying validation logic (primary validation, interface validation, pool selection, phase ordering), integration/E2E tests with user-observable scenarios (create with --external-ip-attachment, verify DNAT functional, verify cleanup on deletion), and tricky test cases covering edge cases (multi-NIC primary on second attachment, pool exhaustion, cleanup failure, VIP feedback loop failure). Graduation criteria are concrete checklists
Scope 2/2 Excellent decomposition: one unified EP defining shared architecture plus four per-service EPs with focused scope. Each EP has specific non-goals with deferrals tracked (e.g., 'VM-based cluster node sets — deferred, HyperShift/CUDN integration not in scope'). Alternatives sections include 2+ real alternatives with rationale for rejection. All designs reference their PRD via frontmatter. Personas (Tenant User, Tenant Admin, Cloud Infrastructure Admin, Cloud Provider Admin) are addressed in workfl
Architecture 2/2 All OSAC patterns followed consistently across five EPs: tenant isolation via osac.openshift.io/tenant annotation with OPA enforcement, owner references, standard object shapes (spec/status separation), controller patterns (finalizer -> status update -> provisioning lifecycle), conditions for lifecycle state (DefaultNetworkingReady, NetworkingConfigured, etc.). Cross-repo impacts enumerated with Jira-tracked dependency tables (fulfillment-service, osac-operator, bare-metal-fulfillment-operator,

Verdict: Exceptionally thorough design submission covering five EPs (unified + per-service + default networking) with deep technical detail, consistent OSAC patterns, comprehensive test plans, and well-defined scope boundaries — one of the strongest design submissions in terms of completeness and cross-component coordination.

Feedback: Track all 'GAP' dependency items in Jira (CaaS has 5 untracked items, BMaaS has 6) — these represent real implementation work that should be visible in project planning. The default networking EP's failure recovery strategy ('delete and re-create tenant') is a blunt instrument; consider adding a reconciliation retry mechanism or partial re-creation flow so admins don't lose tenant state on transient failures. Check whether matching temp-api files exist in osac-ux for affected resources — if they do, a UX Alignment section with field-by-field mapping is required by the design template.

Critical (0)

None.

Important (2)

  1. CaaS and BMaaS dependency tables list 5-6 items each as 'Not tracked / GAP' — these represent real implementation tasks (e.g., 'HostType: add NetworkInterface fields', 'Agent selection logic in operator', 'Remove unused BareMetalInstance spec.networkClass field') that should be tracked in Jira before implementation begins.
  2. Default networking EP recovery flow requires deleting and re-creating the entire tenant when default networking provisioning fails (line 450: 'Cloud Provider Admin inspects failure, fixes root cause, deletes tenant, re-creates tenant'). This destroys any existing tenant state. A reconciliation retry or partial re-creation mechanism would be less disruptive.

Suggestions (3)

  1. Consider adding a UX Alignment section if matching temp-api files exist in osac-ux/libs/ui-components/src/api/v1/ for any of the affected resources (ComputeInstance, Cluster, BaremetalInstance, NetworkClass) — the design template conditionally requires this.
  2. The 'No new metrics or alerts' statement appears in all five EPs — while existing provisioning metrics may cover the basics, auto ExternalIP pool capacity exhaustion and default networking provisioning failures could benefit from dedicated Prometheus metrics for operational visibility.
  3. Several EPs repeat the same auto-provisioned resource cleanup pattern verbatim — consider whether a shared reference section in the unified EP could reduce duplication and make future updates easier to maintain.

Review cost

Model: claude-opus-4-6
Cost: $1.8521
Tokens: 24 in / 6.1k out
Cache: 2.0M read
Active time: 3m 1s
API calls: 0

@vladikr

vladikr commented Jul 27, 2026

Copy link
Copy Markdown
Contributor

@danmanor Thank you for the clarification :)
It makes sense to me!
/lgmt

@vladikr

vladikr commented Jul 27, 2026

Copy link
Copy Markdown
Contributor

/lgtm

@openshift-ci openshift-ci Bot added the lgtm label Jul 27, 2026
@openshift-merge-bot
openshift-merge-bot Bot merged commit bb96154 into osac-project:main Jul 27, 2026
5 checks passed
empovit pushed a commit to empovit/osac-enhancement-proposals that referenced this pull request Aug 2, 2026
networking/README.md's superseded-by field was N/A despite
unified-networking/README.md already declaring
replaces: /enhancements/networking. Points at the current
unified-networking path (pre-restructure) rather than a
post-osac-project#107/post-merge path, since that work hasn't landed yet.

Signed-off-by: Tommy Hughes <tohughes@redhat.com>
@openshift-ci

openshift-ci Bot commented Aug 26, 2026

Copy link
Copy Markdown
Contributor

@aminhussainbarbhuiya17-art: changing LGTM is restricted to collaborators

Details

In response to this:

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes-sigs/prow repository.

@openshift-ci

openshift-ci Bot commented Aug 26, 2026

Copy link
Copy Markdown
Contributor

[APPROVALNOTIFIER] This PR is APPROVED

This pull-request has been approved by: aminhussainbarbhuiya17-art, danmanor

The full list of commands accepted by this bot can be found here.

The pull request process is described here

Details Needs approval from an approver in each of these files:

Approvers can indicate their approval by writing /approve in a comment
Approvers can cancel approval by writing /approve cancel in a comment

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants