Repository navigation
Add Cluster-as-a-Service enhancement proposal - #33
Conversation
Define the CaaS API for provisioning and managing OpenShift clusters via pre-defined templates and bare-metal hosts, following the same fulfillment workflow established by VMaaS. Jira: [MGMT-23417](https://redhat.atlassian.net/browse/MGMT-23417) Generated with [Claude Code](https://claude.com/claude-code)
|
Note Reviews pausedIt looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the Use the following commands to manage reviews:
Use the checkboxes below for quick actions:
WalkthroughAdds a Cluster-as-a-Service enhancement proposal at enhancements/caas/README.md describing a tenant self-service workflow for provisioning and managing OpenShift clusters. Introduces provider-facing resources: Estimated code review effort🎯 1 (Trivial) | ⏱️ ~2 minutes 🚥 Pre-merge checks | ✅ 3✅ Passed checks (3 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches🧪 Generate unit tests (beta)
Warning Review ran into problems🔥 ProblemsGit: Failed to clone repository. Please run the Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Actionable comments posted: 2
🧹 Nitpick comments (4)
enhancements/caas/README.md (4)
407-461: Track completion of TBD sections before finalization.Several important sections are marked as TBD:
- Risks and Mitigations (line 409)
- Test Plan (line 441)
- Graduation Criteria (line 445)
- Support Procedures (line 461)
Before finalizing this proposal, especially Risks and Mitigations should be completed given the complexity of multi-tenant cluster provisioning, bare-metal allocation, and credential management.
Would you like guidance on what to include in the Risks and Mitigations section? I can suggest common risks for CaaS architectures (e.g., resource exhaustion, noisy neighbor, quota enforcement, disaster recovery).
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@enhancements/caas/README.md` around lines 407 - 461, Fill in all "TBD" sections before merging: complete the "Risks and Mitigations" section by enumerating concrete risks (resource exhaustion, noisy neighbor, host-class allocation failures, provisioning/credential failures, security/Isolate tenant boundaries, upgrade rollback) and corresponding mitigations; provide a detailed "Test Plan" with acceptance, integration, scalability and failure/recovery tests (include HyperShift control plane, bare-metal allocation, multi-tenant isolation, node set changes); define measurable "Graduation Criteria" (SLOs, stability thresholds, field trial results, API stability, upgrade success rates) and list exit criteria; and document "Support Procedures" (runbooks, escalation paths, observability/alerts, debugging steps for DEGRADED condition). Update the README headings "Risks and Mitigations", "Test Plan", "Graduation Criteria", and "Support Procedures" accordingly.
344-349: Add API examples for credential retrieval endpoints.The document mentions
GetKubeconfigandGetPasswordendpoints but doesn't provide API request/response examples. Including these would help implementers understand the expected formats, especially for security-sensitive operations.Consider adding examples showing:
- Request format (path parameters, authentication)
- Response format (credential structure, encoding)
- Error responses (e.g., cluster not ready, unauthorized)
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@enhancements/caas/README.md` around lines 344 - 349, Add concrete API request/response examples for the GetKubeconfig and GetPassword endpoints in the README: show sample HTTP requests (method, path with path/query parameters and required auth header), a sample successful response body for GetKubeconfig (e.g., kubeconfig YAML as base64 or plain string) and GetPassword (e.g., JSON with password and expiry metadata and encoding), and representative error responses (cluster not ready, 401/403 unauthorized, 404 not found) including status codes and error JSON shapes; reference the endpoint names GetKubeconfig and GetPassword and include notes about authentication required and any encoding (base64) or size/expiry considerations.
208-218: Consider template versioning strategy.The template management workflow describes periodic synchronization but doesn't address versioning. Consider these scenarios:
- A tenant initiates cluster creation using template version 1, but the template is updated to version 2 during provisioning
- A cluster created with template v1 needs troubleshooting after the template has been updated to v2
Without version tracking, it will be difficult to:
- Ensure reproducible cluster deployments
- Troubleshoot issues with older clusters
- Safely update templates without affecting in-flight operations
Consider adding template versioning (semantic versioning or timestamps) and storing the template version in the Cluster resource's spec or metadata.
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@enhancements/caas/README.md` around lines 208 - 218, The README's "Cluster template management" section lacks a template versioning strategy; update the design to require and persist a template version identifier (semantic version or timestamp) whenever the Ansible Automation Platform sync job updates the central templates and when the Fulfillment Service applies a template; record that identifier on the Cluster resource (e.g., add a templateVersion field in the Cluster spec/metadata), ensure the Fulfillment Service copies the templateVersion from the repo into the Cluster during provisioning, and document how to lock or reference specific template versions to enable reproducible deployments, safe updates, and post-creation troubleshooting.
110-147: Clarify that each cluster gets its own dedicated HostPool resource.Line 133 states "Creates a HostPool to allocate the required bare-metal hosts." Clarify that a new HostPool is created specifically for this cluster (not allocated from pre-existing pools), and that this HostPool serves as a dedicated resource managing the cluster's bare-metal hosts throughout the cluster's lifecycle.
Reference to the bare-metal-fulfillment proposal would help clarify: HostPools are created on-demand as custom resources, and each contains a collection of requested bare metal machines with their network configuration.
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@enhancements/caas/README.md` around lines 110 - 147, Clarify the existing step that creates a HostPool by stating that the Operator creates a new, dedicated HostPool custom resource for each Cluster (not drawn from pre-existing pools): update the line "Creates a HostPool to allocate the required bare-metal hosts" to say the Operator creates an on-demand HostPool CR specifically for that ClusterOrder/Cluster, which owns the requested bare-metal machines and network configuration for the cluster's lifecycle; mention that the HostPool persists and is managed for the cluster until deletion and add a brief reference to the bare-metal-fulfillment proposal for HostPool details.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Inline comments:
In `@enhancements/caas/README.md`:
- Around line 344-349: Add a "Security Considerations" subsection that documents
how GetKubeconfig and GetPassword requests are authenticated and authorized
(e.g., tenant-scoped tokens, OIDC, RBAC checks in the Fulfillment service
ensuring tenants can only access their own cluster credentials), describe
credential storage and protection (encrypted-at-rest storage for kubeconfigs and
passwords, KMS/secret manager usage, in-memory handling safeguards), define
rotation and expiry policy for admin credentials (rotation intervals, automated
rotation workflow and revocation), specify auditing and monitoring controls
(immutable access logs for every credential retrieval, alerting on anomalous
access, and retention), and note rate-limiting/throttling to mitigate credential
harvesting plus a recommendation to offer least-privilege kubeconfigs
(optionally generate time-scoped, limited-RBAC kubeconfigs instead of always
returning full admin credentials).
- Around line 227-251: The README shows inconsistent cluster ID usage between
the request example ("object.id": "mycluster") and the response UUID; clarify
the ID policy by documenting whether object.id is tenant-provided or
server-generated, and when each applies (e.g., allow tenant-specified
unique-per-tenant IDs, otherwise server assigns a UUID and returns it). Update
the examples so they align with the policy (either change the request to omit id
and show the server-generated UUID in the response, or keep "mycluster" in both
request and response and state uniqueness constraints), and add a short note
referencing "object.id" and the API response ID format to remove ambiguity.
---
Nitpick comments:
In `@enhancements/caas/README.md`:
- Around line 407-461: Fill in all "TBD" sections before merging: complete the
"Risks and Mitigations" section by enumerating concrete risks (resource
exhaustion, noisy neighbor, host-class allocation failures,
provisioning/credential failures, security/Isolate tenant boundaries, upgrade
rollback) and corresponding mitigations; provide a detailed "Test Plan" with
acceptance, integration, scalability and failure/recovery tests (include
HyperShift control plane, bare-metal allocation, multi-tenant isolation, node
set changes); define measurable "Graduation Criteria" (SLOs, stability
thresholds, field trial results, API stability, upgrade success rates) and list
exit criteria; and document "Support Procedures" (runbooks, escalation paths,
observability/alerts, debugging steps for DEGRADED condition). Update the README
headings "Risks and Mitigations", "Test Plan", "Graduation Criteria", and
"Support Procedures" accordingly.
- Around line 344-349: Add concrete API request/response examples for the
GetKubeconfig and GetPassword endpoints in the README: show sample HTTP requests
(method, path with path/query parameters and required auth header), a sample
successful response body for GetKubeconfig (e.g., kubeconfig YAML as base64 or
plain string) and GetPassword (e.g., JSON with password and expiry metadata and
encoding), and representative error responses (cluster not ready, 401/403
unauthorized, 404 not found) including status codes and error JSON shapes;
reference the endpoint names GetKubeconfig and GetPassword and include notes
about authentication required and any encoding (base64) or size/expiry
considerations.
- Around line 208-218: The README's "Cluster template management" section lacks
a template versioning strategy; update the design to require and persist a
template version identifier (semantic version or timestamp) whenever the Ansible
Automation Platform sync job updates the central templates and when the
Fulfillment Service applies a template; record that identifier on the Cluster
resource (e.g., add a templateVersion field in the Cluster spec/metadata),
ensure the Fulfillment Service copies the templateVersion from the repo into the
Cluster during provisioning, and document how to lock or reference specific
template versions to enable reproducible deployments, safe updates, and
post-creation troubleshooting.
- Around line 110-147: Clarify the existing step that creates a HostPool by
stating that the Operator creates a new, dedicated HostPool custom resource for
each Cluster (not drawn from pre-existing pools): update the line "Creates a
HostPool to allocate the required bare-metal hosts" to say the Operator creates
an on-demand HostPool CR specifically for that ClusterOrder/Cluster, which owns
the requested bare-metal machines and network configuration for the cluster's
lifecycle; mention that the HostPool persists and is managed for the cluster
until deletion and add a brief reference to the bare-metal-fulfillment proposal
for HostPool details.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: CHILL
Plan: Pro
Run ID: 9f39b584-53e9-4b76-bff3-42935be6788c
📒 Files selected for processing (1)
enhancements/caas/README.md
The id field should be server-generated (UUID), not tenant-provided. This aligns with REST API conventions where the server assigns resource identifiers on creation. Generated with [Claude Code](https://claude.com/claude-code)
There was a problem hiding this comment.
Actionable comments posted: 1
🧹 Nitpick comments (1)
enhancements/caas/README.md (1)
408-408: Consider completing TBD sections before finalization.Several sections are marked TBD:
- Risks and Mitigations (line 408)
- Test Plan (line 440)
- Graduation Criteria (line 444)
- Support Procedures (line 460)
While TBD placeholders are acceptable for early-stage proposals, having these sections completed would strengthen the proposal, especially "Risks and Mitigations" and "Test Plan" which inform implementation planning.
Would you like assistance outlining any of these sections? For example, potential risks might include:
- HyperShift control plane resource contention on the hub cluster under high tenant load
- Bare-metal host inventory exhaustion leading to cluster creation failures
- Template version compatibility issues during upgrades
Also applies to: 440-440, 444-444, 460-460
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@enhancements/caas/README.md` at line 408, Complete the TBD sections by adding concrete content for the headings "Risks and Mitigations", "Test Plan", "Graduation Criteria", and "Support Procedures" in the README: for Risks and Mitigations list likely failure modes (e.g., HyperShift control plane contention, bare-metal host inventory exhaustion, template/version compatibility) and corresponding mitigations; for Test Plan provide acceptance, load, upgrade, and failure-recovery test cases with pass/fail criteria and test data/setup steps; for Graduation Criteria enumerate measurable signals for moving from alpha to beta/stable (e.g., X clusters created, Y% success under load, documented upgrade path); and for Support Procedures describe on-call escalation, runbooks, monitoring/alert thresholds, and rollback steps — edit the sections titled exactly "Risks and Mitigations", "Test Plan", "Graduation Criteria", and "Support Procedures" to replace the TBD placeholders with these concrete items.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Inline comments:
In `@enhancements/caas/README.md`:
- Around line 324-341: The README is ambiguous about READY vs DEGRADED; change
the CLUSTER_STATE_READY description (symbol CLUSTER_STATE_READY) to "The cluster
API is operational and accessible" (remove "fully operational") and add an
explicit sentence stating that CLUSTER_CONDITION_TYPE_DEGRADED may be TRUE while
the cluster state is READY to indicate reduced capacity (e.g., missing worker
nodes); alternatively, if you prefer a distinct state, add
CLUSTER_STATE_DEGRADED as a new state value and update references to use it—pick
one approach and make sure CLUSTER_CONDITION_TYPE_DEGRADED semantics and the
example about nodes not being allocated (mentioned elsewhere) are updated to
match.
---
Nitpick comments:
In `@enhancements/caas/README.md`:
- Line 408: Complete the TBD sections by adding concrete content for the
headings "Risks and Mitigations", "Test Plan", "Graduation Criteria", and
"Support Procedures" in the README: for Risks and Mitigations list likely
failure modes (e.g., HyperShift control plane contention, bare-metal host
inventory exhaustion, template/version compatibility) and corresponding
mitigations; for Test Plan provide acceptance, load, upgrade, and
failure-recovery test cases with pass/fail criteria and test data/setup steps;
for Graduation Criteria enumerate measurable signals for moving from alpha to
beta/stable (e.g., X clusters created, Y% success under load, documented upgrade
path); and for Support Procedures describe on-call escalation, runbooks,
monitoring/alert thresholds, and rollback steps — edit the sections titled
exactly "Risks and Mitigations", "Test Plan", "Graduation Criteria", and
"Support Procedures" to replace the TBD placeholders with these concrete items.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: CHILL
Plan: Pro
Run ID: 13414535-f657-430b-a925-88ec69705572
📒 Files selected for processing (1)
enhancements/caas/README.md
READY state now describes control plane accessibility, not full capacity. DEGRADED condition explicitly documented as compatible with READY state. Generated with [Claude Code](https://claude.com/claude-code)
There was a problem hiding this comment.
Actionable comments posted: 2
🧹 Nitpick comments (1)
enhancements/caas/README.md (1)
399-401: Add an explicit dependency contract to Bare Metal Fulfillment.This section references HostPool behavior but doesn’t pin the expected API contract/version. A short compatibility note (e.g., “targets HostPool schema from Bare Metal Fulfillment revision X”) will prevent spec drift between proposals.
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@enhancements/caas/README.md` around lines 399 - 401, The README mentions HostPool but lacks a pinned compatibility contract with the Bare Metal Fulfillment proposal; add a short explicit dependency note stating the expected HostPool API/schema and revision (e.g., “This specification targets the HostPool schema from the Bare Metal Fulfillment proposal, revision X”) near the paragraph referencing HostPool so readers know which version to implement against; reference the HostPool term and the Bare Metal Fulfillment proposal name so the contract is unambiguous and include a link or anchor to the specific revision.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Inline comments:
In `@enhancements/caas/README.md`:
- Around line 409-447: Replace the "TBD" placeholders in the release-critical
sections by providing minimal, actionable content: under the "Risks and
Mitigations" header list the top 3 operational risks and one mitigation per
risk; under "Test Plan" describe at least one verification scenario and success
criteria (e.g., cluster creation, scaling, and failure recovery tests); under
"Graduation Criteria" enumerate clear pass/fail gates (e.g., stable provisioning
X times, supported upgrade path, monitoring/alerting present); and under
"Support Procedures" describe who owns support, escalation steps, and required
runbook items. Update the specific headers "Risks and Mitigations", "Test Plan",
"Graduation Criteria", and "Support Procedures" in the README.md with those
brief, actionable items so implementers have acceptance and operability
criteria.
- Around line 186-200: Clarify that deletion is handled via Kubernetes
finalizers: instead of deleting the ClusterOrder CR immediately, the Fulfillment
Service or caller should set the CR's deletionTimestamp (triggering deletion)
and the O-SAC Operator's reconcile loop (watching the ClusterOrder resource)
must perform cleanup (delete HostedCluster, release hosts, etc.) while the
finalizer is present; once cleanup completes the Operator must remove the
finalizer from the ClusterOrder so Kubernetes can garbage-collect the CR and
then update deletion status/conditions appropriately (or leave status indicating
Deleting on failure). Reference the ClusterOrder CR, its finalizers field,
deletionTimestamp, and the Operator reconcile/cleanup flow in the updated text.
---
Nitpick comments:
In `@enhancements/caas/README.md`:
- Around line 399-401: The README mentions HostPool but lacks a pinned
compatibility contract with the Bare Metal Fulfillment proposal; add a short
explicit dependency note stating the expected HostPool API/schema and revision
(e.g., “This specification targets the HostPool schema from the Bare Metal
Fulfillment proposal, revision X”) near the paragraph referencing HostPool so
readers know which version to implement against; reference the HostPool term and
the Bare Metal Fulfillment proposal name so the contract is unambiguous and
include a link or anchor to the specific revision.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: CHILL
Plan: Pro
Run ID: 17411b87-a729-45af-8873-2c35b48dc834
📒 Files selected for processing (1)
enhancements/caas/README.md
- Explain finalizer-based deletion mechanism (deletionTimestamp + finalizer removal) - Clarify HostPool is deleted during cluster cleanup - Add CLUSTER_STATE_DELETING state - Clarify template node_sets are defaults, tenants can override sizes at creation - Add networking integration subsection referencing bare-metal-fulfillment and networking proposals Generated with [Claude Code](https://claude.com/claude-code)
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Inline comments:
In `@enhancements/caas/README.md`:
- Around line 391-395: Clarify whether tenants may change the template
node_sets' host_class at cluster creation: update the README's paragraph about
node_sets to explicitly state whether host_class can be overridden during
initial cluster creation (in addition to size) or not; if overrides are allowed,
add a sentence like "Tenants may override the template's host_class for default
node_sets when creating a cluster" and document any constraints, and if
overrides are not allowed, replace or append with "Tenants can override the size
of default node_sets (and optionally add new node_sets with different
host_class), but the template's host_class for default node_sets must be used at
creation." Also cross-reference the existing scaling validation wording (the
rule that "The host class for existing node sets cannot be changed" used for
updates) so readers know creation-time behavior differs from update-time
constraints.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: CHILL
Plan: Pro
Run ID: 9d79f9c4-3ed3-4e07-bb7e-533f6936da9f
📒 Files selected for processing (1)
enhancements/caas/README.md
- Tenants create clusters via ClusterOrder, not the Clusters API directly - POST /clusters is a server-internal operation - Remove gRPC-style "object" wrapper from JSON examples (show HTTP format) - Add endpoint reference for status retrieval Generated with [Claude Code](https://claude.com/claude-code)
|
First thought; let's clarify the scope of this. I was confused because we already have cluster-aaS. But it looks like this is actually about proposing a more explicit API. |
Right. The CaaS capability already exists. This proposal promotes cluster configuration from the opaque This follows the same approach already taken for VMaaS. I've updated the summary and motivation to make this clear. |
Update Summary and Motivation to clarify that this proposal defines the tenant-facing API contract for the existing CaaS capability, not introducing CaaS itself. Generated with [Claude Code](https://claude.com/claude-code)
|
I've updated the Summary and Motivation sections to make this distinction clearer. |
Deletion uses deletion_timestamp with finalizers (standard Kubernetes pattern) rather than a separate CLUSTER_STATE_DELETING enum value. Generated with [Claude Code](https://claude.com/claude-code)
Replace the generic map<string, Any> template_parameters approach with explicit typed fields (pull_secret, ssh_public_key) in ClusterSpec. This provides type safety, discoverability, and proto-level validation. Updated all examples and workflow descriptions to reflect the new API contract. Generated with [Claude Code](https://claude.com/claude-code)
Promote pull_secret, ssh_public_key (from template_parameters), and release_image, cluster_network_cidr, service_network_cidr (from hardcoded values) to explicit ClusterSpec fields. All fields are optional with sensible defaults. CIDRs use plain string notation, consistent with VirtualNetwork/Subnet protos and Kubernetes conventions. Generated with [Claude Code](https://claude.com/claude-code)
The purpose of this proposal is moving cluster configuration from templates/playbooks into the tenant-facing API. Updated Summary and Motivation to clearly state this: parameters are currently hidden in template_parameters or hardcoded in Ansible roles, and this proposal makes them explicit API fields. Generated with [Claude Code](https://claude.com/claude-code)
User stories now focus on tenant control over cluster configuration (pull secret, SSH key, OCP version, networking CIDRs). Goals now focus on defining explicit ClusterSpec fields to replace the opaque template_parameters mechanism. Generated with [Claude Code](https://claude.com/claude-code)
Non-goals now clarify what stays out: template-specific parameters, provider-managed settings, workflow changes, and removal of template_parameters. Generated with [Claude Code](https://claude.com/claude-code)
Replace the full workflow descriptions with a clear current-state vs proposed-change structure. Node scaling, deletion, and template management workflows are unchanged and noted as such. Generated with [Claude Code](https://claude.com/claude-code)
Remove repeated documentation of existing cluster states, conditions, credentials, and status examples — these are unchanged. API section now only shows the new ClusterSpec fields. Drawbacks and Risks updated to be specific to this proposal. Generated with [Claude Code](https://claude.com/claude-code)
Every section from the enhancement template is present. Each concept appears exactly once. The proposal focuses solely on promoting template_parameters and hardcoded values to explicit ClusterSpec fields. Generated with [Claude Code](https://claude.com/claude-code)
adriengentil
left a comment
There was a problem hiding this comment.
Looks good overall!
|
BTW, it would be interesting to learn from ROSA's API:
|
Mark pull_secret as write-only (redacted in GET responses) per adriengentil's feedback. Add note that templates can override system defaults for any field. Generated with [Claude Code](https://claude.com/claude-code)
Move pod_cidr and service_cidr into a nested network message for cleaner API organization as the networking config grows. Updated examples and CLI flags. Generated with [Claude Code](https://claude.com/claude-code)
Good points. Grouped the networking CIDRs into a Agreed on keeping |
Summary
Promote cluster configuration from opaque
template_parametersand hardcoded Ansible values to explicit typed fields inClusterSpec, following the same approach taken for VMaaS.New
ClusterSpecfieldspull_secret— image registry credentials (currently intemplate_parameters)ssh_public_key— SSH key for worker nodes (currently intemplate_parameters)release_image— OCP version (currently hardcoded in template defaults)cluster_network_cidr— pod network CIDR (currently hardcoded in Ansible role)service_network_cidr— service network CIDR (currently hardcoded in Ansible role)All fields are optional with sensible defaults.
Jira
MGMT-23417