Skip to content

Add Cluster-as-a-Service enhancement proposal - #33

Merged
tzvatot merged 17 commits into
osac-project:mainfrom
tzvatot:enhancement/caas-api
Apr 16, 2026
Merged

tzvatot merged 17 commits into
osac-project:mainfrom
tzvatot:enhancement/caas-api

Conversation

@tzvatot

@tzvatot tzvatot commented Mar 31, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Promote cluster configuration from opaque template_parameters and hardcoded Ansible values to explicit typed fields in ClusterSpec, following the same approach taken for VMaaS.

New ClusterSpec fields

  • pull_secret — image registry credentials (currently in template_parameters)
  • ssh_public_key — SSH key for worker nodes (currently in template_parameters)
  • release_image — OCP version (currently hardcoded in template defaults)
  • cluster_network_cidr — pod network CIDR (currently hardcoded in Ansible role)
  • service_network_cidr — service network CIDR (currently hardcoded in Ansible role)

All fields are optional with sensible defaults.

Jira

MGMT-23417

Define the CaaS API for provisioning and managing OpenShift clusters
via pre-defined templates and bare-metal hosts, following the same
fulfillment workflow established by VMaaS.

Jira: [MGMT-23417](https://redhat.atlassian.net/browse/MGMT-23417)

Generated with [Claude Code](https://claude.com/claude-code)
@coderabbitai

coderabbitai Bot commented Mar 31, 2026 •

Copy link
Copy Markdown

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review

Walkthrough

Adds a Cluster-as-a-Service enhancement proposal at enhancements/caas/README.md describing a tenant self-service workflow for provisioning and managing OpenShift clusters. Introduces provider-facing resources: Cluster (tenant-visible representation) and ClusterTemplate (provider-defined, Ansible-role-backed blueprints with parameterization and default node_sets). Describes control flow where tenants create/delete via Fulfillment Service APIs/CLI producing ClusterOrder CRs; an O-SAC Operator reconciles ClusterOrder; O-SAC AAP automation provisions/deprovisions HyperShift HostedClusters using HostPools, updates statuses (api_url, console_url), documents node-scaling rules, credential endpoints, lifecycle conditions, GitOps template sync, goals, non-goals, and placeholders for risks and test plans.

Estimated code review effort

🎯 1 (Trivial) | ⏱️ ~2 minutes

🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately and concisely summarizes the main change: adding a Cluster-as-a-Service enhancement proposal document.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Warning

Review ran into problems

🔥 Problems

Git: Failed to clone repository. Please run the @coderabbitai full review command to re-trigger a full review. If the issue persists, set path_filters to include or exclude specific files.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🧹 Nitpick comments (4)
enhancements/caas/README.md (4)

407-461: Track completion of TBD sections before finalization.

Several important sections are marked as TBD:

  • Risks and Mitigations (line 409)
  • Test Plan (line 441)
  • Graduation Criteria (line 445)
  • Support Procedures (line 461)

Before finalizing this proposal, especially Risks and Mitigations should be completed given the complexity of multi-tenant cluster provisioning, bare-metal allocation, and credential management.

Would you like guidance on what to include in the Risks and Mitigations section? I can suggest common risks for CaaS architectures (e.g., resource exhaustion, noisy neighbor, quota enforcement, disaster recovery).

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@enhancements/caas/README.md` around lines 407 - 461, Fill in all "TBD"
sections before merging: complete the "Risks and Mitigations" section by
enumerating concrete risks (resource exhaustion, noisy neighbor, host-class
allocation failures, provisioning/credential failures, security/Isolate tenant
boundaries, upgrade rollback) and corresponding mitigations; provide a detailed
"Test Plan" with acceptance, integration, scalability and failure/recovery tests
(include HyperShift control plane, bare-metal allocation, multi-tenant
isolation, node set changes); define measurable "Graduation Criteria" (SLOs,
stability thresholds, field trial results, API stability, upgrade success rates)
and list exit criteria; and document "Support Procedures" (runbooks, escalation
paths, observability/alerts, debugging steps for DEGRADED condition). Update the
README headings "Risks and Mitigations", "Test Plan", "Graduation Criteria", and
"Support Procedures" accordingly.

344-349: Add API examples for credential retrieval endpoints.

The document mentions GetKubeconfig and GetPassword endpoints but doesn't provide API request/response examples. Including these would help implementers understand the expected formats, especially for security-sensitive operations.

Consider adding examples showing:

  • Request format (path parameters, authentication)
  • Response format (credential structure, encoding)
  • Error responses (e.g., cluster not ready, unauthorized)
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@enhancements/caas/README.md` around lines 344 - 349, Add concrete API
request/response examples for the GetKubeconfig and GetPassword endpoints in the
README: show sample HTTP requests (method, path with path/query parameters and
required auth header), a sample successful response body for GetKubeconfig
(e.g., kubeconfig YAML as base64 or plain string) and GetPassword (e.g., JSON
with password and expiry metadata and encoding), and representative error
responses (cluster not ready, 401/403 unauthorized, 404 not found) including
status codes and error JSON shapes; reference the endpoint names GetKubeconfig
and GetPassword and include notes about authentication required and any encoding
(base64) or size/expiry considerations.

208-218: Consider template versioning strategy.

The template management workflow describes periodic synchronization but doesn't address versioning. Consider these scenarios:

  1. A tenant initiates cluster creation using template version 1, but the template is updated to version 2 during provisioning
  2. A cluster created with template v1 needs troubleshooting after the template has been updated to v2

Without version tracking, it will be difficult to:

  • Ensure reproducible cluster deployments
  • Troubleshoot issues with older clusters
  • Safely update templates without affecting in-flight operations

Consider adding template versioning (semantic versioning or timestamps) and storing the template version in the Cluster resource's spec or metadata.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@enhancements/caas/README.md` around lines 208 - 218, The README's "Cluster
template management" section lacks a template versioning strategy; update the
design to require and persist a template version identifier (semantic version or
timestamp) whenever the Ansible Automation Platform sync job updates the central
templates and when the Fulfillment Service applies a template; record that
identifier on the Cluster resource (e.g., add a templateVersion field in the
Cluster spec/metadata), ensure the Fulfillment Service copies the
templateVersion from the repo into the Cluster during provisioning, and document
how to lock or reference specific template versions to enable reproducible
deployments, safe updates, and post-creation troubleshooting.

110-147: Clarify that each cluster gets its own dedicated HostPool resource.

Line 133 states "Creates a HostPool to allocate the required bare-metal hosts." Clarify that a new HostPool is created specifically for this cluster (not allocated from pre-existing pools), and that this HostPool serves as a dedicated resource managing the cluster's bare-metal hosts throughout the cluster's lifecycle.

Reference to the bare-metal-fulfillment proposal would help clarify: HostPools are created on-demand as custom resources, and each contains a collection of requested bare metal machines with their network configuration.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@enhancements/caas/README.md` around lines 110 - 147, Clarify the existing
step that creates a HostPool by stating that the Operator creates a new,
dedicated HostPool custom resource for each Cluster (not drawn from pre-existing
pools): update the line "Creates a HostPool to allocate the required bare-metal
hosts" to say the Operator creates an on-demand HostPool CR specifically for
that ClusterOrder/Cluster, which owns the requested bare-metal machines and
network configuration for the cluster's lifecycle; mention that the HostPool
persists and is managed for the cluster until deletion and add a brief reference
to the bare-metal-fulfillment proposal for HostPool details.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@enhancements/caas/README.md`:
- Around line 344-349: Add a "Security Considerations" subsection that documents
how GetKubeconfig and GetPassword requests are authenticated and authorized
(e.g., tenant-scoped tokens, OIDC, RBAC checks in the Fulfillment service
ensuring tenants can only access their own cluster credentials), describe
credential storage and protection (encrypted-at-rest storage for kubeconfigs and
passwords, KMS/secret manager usage, in-memory handling safeguards), define
rotation and expiry policy for admin credentials (rotation intervals, automated
rotation workflow and revocation), specify auditing and monitoring controls
(immutable access logs for every credential retrieval, alerting on anomalous
access, and retention), and note rate-limiting/throttling to mitigate credential
harvesting plus a recommendation to offer least-privilege kubeconfigs
(optionally generate time-scoped, limited-RBAC kubeconfigs instead of always
returning full admin credentials).
- Around line 227-251: The README shows inconsistent cluster ID usage between
the request example ("object.id": "mycluster") and the response UUID; clarify
the ID policy by documenting whether object.id is tenant-provided or
server-generated, and when each applies (e.g., allow tenant-specified
unique-per-tenant IDs, otherwise server assigns a UUID and returns it). Update
the examples so they align with the policy (either change the request to omit id
and show the server-generated UUID in the response, or keep "mycluster" in both
request and response and state uniqueness constraints), and add a short note
referencing "object.id" and the API response ID format to remove ambiguity.

---

Nitpick comments:
In `@enhancements/caas/README.md`:
- Around line 407-461: Fill in all "TBD" sections before merging: complete the
"Risks and Mitigations" section by enumerating concrete risks (resource
exhaustion, noisy neighbor, host-class allocation failures,
provisioning/credential failures, security/Isolate tenant boundaries, upgrade
rollback) and corresponding mitigations; provide a detailed "Test Plan" with
acceptance, integration, scalability and failure/recovery tests (include
HyperShift control plane, bare-metal allocation, multi-tenant isolation, node
set changes); define measurable "Graduation Criteria" (SLOs, stability
thresholds, field trial results, API stability, upgrade success rates) and list
exit criteria; and document "Support Procedures" (runbooks, escalation paths,
observability/alerts, debugging steps for DEGRADED condition). Update the README
headings "Risks and Mitigations", "Test Plan", "Graduation Criteria", and
"Support Procedures" accordingly.
- Around line 344-349: Add concrete API request/response examples for the
GetKubeconfig and GetPassword endpoints in the README: show sample HTTP requests
(method, path with path/query parameters and required auth header), a sample
successful response body for GetKubeconfig (e.g., kubeconfig YAML as base64 or
plain string) and GetPassword (e.g., JSON with password and expiry metadata and
encoding), and representative error responses (cluster not ready, 401/403
unauthorized, 404 not found) including status codes and error JSON shapes;
reference the endpoint names GetKubeconfig and GetPassword and include notes
about authentication required and any encoding (base64) or size/expiry
considerations.
- Around line 208-218: The README's "Cluster template management" section lacks
a template versioning strategy; update the design to require and persist a
template version identifier (semantic version or timestamp) whenever the Ansible
Automation Platform sync job updates the central templates and when the
Fulfillment Service applies a template; record that identifier on the Cluster
resource (e.g., add a templateVersion field in the Cluster spec/metadata),
ensure the Fulfillment Service copies the templateVersion from the repo into the
Cluster during provisioning, and document how to lock or reference specific
template versions to enable reproducible deployments, safe updates, and
post-creation troubleshooting.
- Around line 110-147: Clarify the existing step that creates a HostPool by
stating that the Operator creates a new, dedicated HostPool custom resource for
each Cluster (not drawn from pre-existing pools): update the line "Creates a
HostPool to allocate the required bare-metal hosts" to say the Operator creates
an on-demand HostPool CR specifically for that ClusterOrder/Cluster, which owns
the requested bare-metal machines and network configuration for the cluster's
lifecycle; mention that the HostPool persists and is managed for the cluster
until deletion and add a brief reference to the bare-metal-fulfillment proposal
for HostPool details.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 9f39b584-53e9-4b76-bff3-42935be6788c

📥 Commits

Reviewing files that changed from the base of the PR and between 39cebc6 and d64fb69.

📒 Files selected for processing (1)
  • enhancements/caas/README.md

Comment thread enhancements/caas/README.md
Comment thread enhancements/caas/README.md Outdated
The id field should be server-generated (UUID), not tenant-provided.
This aligns with REST API conventions where the server assigns
resource identifiers on creation.

Generated with [Claude Code](https://claude.com/claude-code)

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🧹 Nitpick comments (1)
enhancements/caas/README.md (1)

408-408: Consider completing TBD sections before finalization.

Several sections are marked TBD:

  • Risks and Mitigations (line 408)
  • Test Plan (line 440)
  • Graduation Criteria (line 444)
  • Support Procedures (line 460)

While TBD placeholders are acceptable for early-stage proposals, having these sections completed would strengthen the proposal, especially "Risks and Mitigations" and "Test Plan" which inform implementation planning.

Would you like assistance outlining any of these sections? For example, potential risks might include:

  • HyperShift control plane resource contention on the hub cluster under high tenant load
  • Bare-metal host inventory exhaustion leading to cluster creation failures
  • Template version compatibility issues during upgrades

Also applies to: 440-440, 444-444, 460-460

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@enhancements/caas/README.md` at line 408, Complete the TBD sections by adding
concrete content for the headings "Risks and Mitigations", "Test Plan",
"Graduation Criteria", and "Support Procedures" in the README: for Risks and
Mitigations list likely failure modes (e.g., HyperShift control plane
contention, bare-metal host inventory exhaustion, template/version
compatibility) and corresponding mitigations; for Test Plan provide acceptance,
load, upgrade, and failure-recovery test cases with pass/fail criteria and test
data/setup steps; for Graduation Criteria enumerate measurable signals for
moving from alpha to beta/stable (e.g., X clusters created, Y% success under
load, documented upgrade path); and for Support Procedures describe on-call
escalation, runbooks, monitoring/alert thresholds, and rollback steps — edit the
sections titled exactly "Risks and Mitigations", "Test Plan", "Graduation
Criteria", and "Support Procedures" to replace the TBD placeholders with these
concrete items.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@enhancements/caas/README.md`:
- Around line 324-341: The README is ambiguous about READY vs DEGRADED; change
the CLUSTER_STATE_READY description (symbol CLUSTER_STATE_READY) to "The cluster
API is operational and accessible" (remove "fully operational") and add an
explicit sentence stating that CLUSTER_CONDITION_TYPE_DEGRADED may be TRUE while
the cluster state is READY to indicate reduced capacity (e.g., missing worker
nodes); alternatively, if you prefer a distinct state, add
CLUSTER_STATE_DEGRADED as a new state value and update references to use it—pick
one approach and make sure CLUSTER_CONDITION_TYPE_DEGRADED semantics and the
example about nodes not being allocated (mentioned elsewhere) are updated to
match.

---

Nitpick comments:
In `@enhancements/caas/README.md`:
- Line 408: Complete the TBD sections by adding concrete content for the
headings "Risks and Mitigations", "Test Plan", "Graduation Criteria", and
"Support Procedures" in the README: for Risks and Mitigations list likely
failure modes (e.g., HyperShift control plane contention, bare-metal host
inventory exhaustion, template/version compatibility) and corresponding
mitigations; for Test Plan provide acceptance, load, upgrade, and
failure-recovery test cases with pass/fail criteria and test data/setup steps;
for Graduation Criteria enumerate measurable signals for moving from alpha to
beta/stable (e.g., X clusters created, Y% success under load, documented upgrade
path); and for Support Procedures describe on-call escalation, runbooks,
monitoring/alert thresholds, and rollback steps — edit the sections titled
exactly "Risks and Mitigations", "Test Plan", "Graduation Criteria", and
"Support Procedures" to replace the TBD placeholders with these concrete items.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 13414535-f657-430b-a925-88ec69705572

📥 Commits

Reviewing files that changed from the base of the PR and between d64fb69 and 3e0b7c5.

📒 Files selected for processing (1)
  • enhancements/caas/README.md

Comment thread enhancements/caas/README.md Outdated
READY state now describes control plane accessibility, not full capacity.
DEGRADED condition explicitly documented as compatible with READY state.

Generated with [Claude Code](https://claude.com/claude-code)

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🧹 Nitpick comments (1)
enhancements/caas/README.md (1)

399-401: Add an explicit dependency contract to Bare Metal Fulfillment.

This section references HostPool behavior but doesn’t pin the expected API contract/version. A short compatibility note (e.g., “targets HostPool schema from Bare Metal Fulfillment revision X”) will prevent spec drift between proposals.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@enhancements/caas/README.md` around lines 399 - 401, The README mentions
HostPool but lacks a pinned compatibility contract with the Bare Metal
Fulfillment proposal; add a short explicit dependency note stating the expected
HostPool API/schema and revision (e.g., “This specification targets the HostPool
schema from the Bare Metal Fulfillment proposal, revision X”) near the paragraph
referencing HostPool so readers know which version to implement against;
reference the HostPool term and the Bare Metal Fulfillment proposal name so the
contract is unambiguous and include a link or anchor to the specific revision.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@enhancements/caas/README.md`:
- Around line 409-447: Replace the "TBD" placeholders in the release-critical
sections by providing minimal, actionable content: under the "Risks and
Mitigations" header list the top 3 operational risks and one mitigation per
risk; under "Test Plan" describe at least one verification scenario and success
criteria (e.g., cluster creation, scaling, and failure recovery tests); under
"Graduation Criteria" enumerate clear pass/fail gates (e.g., stable provisioning
X times, supported upgrade path, monitoring/alerting present); and under
"Support Procedures" describe who owns support, escalation steps, and required
runbook items. Update the specific headers "Risks and Mitigations", "Test Plan",
"Graduation Criteria", and "Support Procedures" in the README.md with those
brief, actionable items so implementers have acceptance and operability
criteria.
- Around line 186-200: Clarify that deletion is handled via Kubernetes
finalizers: instead of deleting the ClusterOrder CR immediately, the Fulfillment
Service or caller should set the CR's deletionTimestamp (triggering deletion)
and the O-SAC Operator's reconcile loop (watching the ClusterOrder resource)
must perform cleanup (delete HostedCluster, release hosts, etc.) while the
finalizer is present; once cleanup completes the Operator must remove the
finalizer from the ClusterOrder so Kubernetes can garbage-collect the CR and
then update deletion status/conditions appropriately (or leave status indicating
Deleting on failure). Reference the ClusterOrder CR, its finalizers field,
deletionTimestamp, and the Operator reconcile/cleanup flow in the updated text.

---

Nitpick comments:
In `@enhancements/caas/README.md`:
- Around line 399-401: The README mentions HostPool but lacks a pinned
compatibility contract with the Bare Metal Fulfillment proposal; add a short
explicit dependency note stating the expected HostPool API/schema and revision
(e.g., “This specification targets the HostPool schema from the Bare Metal
Fulfillment proposal, revision X”) near the paragraph referencing HostPool so
readers know which version to implement against; reference the HostPool term and
the Bare Metal Fulfillment proposal name so the contract is unambiguous and
include a link or anchor to the specific revision.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 17411b87-a729-45af-8873-2c35b48dc834

📥 Commits

Reviewing files that changed from the base of the PR and between 3e0b7c5 and df79311.

📒 Files selected for processing (1)
  • enhancements/caas/README.md

Comment thread enhancements/caas/README.md Outdated
Comment thread enhancements/caas/README.md
Comment thread enhancements/caas/README.md Outdated
Comment thread enhancements/caas/README.md Outdated
Comment thread enhancements/caas/README.md Outdated
Comment thread enhancements/caas/README.md
Comment thread enhancements/caas/README.md Outdated
- Explain finalizer-based deletion mechanism (deletionTimestamp + finalizer removal)
- Clarify HostPool is deleted during cluster cleanup
- Add CLUSTER_STATE_DELETING state
- Clarify template node_sets are defaults, tenants can override sizes at creation
- Add networking integration subsection referencing bare-metal-fulfillment and networking proposals

Generated with [Claude Code](https://claude.com/claude-code)

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@enhancements/caas/README.md`:
- Around line 391-395: Clarify whether tenants may change the template
node_sets' host_class at cluster creation: update the README's paragraph about
node_sets to explicitly state whether host_class can be overridden during
initial cluster creation (in addition to size) or not; if overrides are allowed,
add a sentence like "Tenants may override the template's host_class for default
node_sets when creating a cluster" and document any constraints, and if
overrides are not allowed, replace or append with "Tenants can override the size
of default node_sets (and optionally add new node_sets with different
host_class), but the template's host_class for default node_sets must be used at
creation." Also cross-reference the existing scaling validation wording (the
rule that "The host class for existing node sets cannot be changed" used for
updates) so readers know creation-time behavior differs from update-time
constraints.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 9d79f9c4-3ed3-4e07-bb7e-533f6936da9f

📥 Commits

Reviewing files that changed from the base of the PR and between df79311 and b7a21c2.

📒 Files selected for processing (1)
  • enhancements/caas/README.md

Comment thread enhancements/caas/README.md Outdated
- Tenants create clusters via ClusterOrder, not the Clusters API directly
- POST /clusters is a server-internal operation
- Remove gRPC-style "object" wrapper from JSON examples (show HTTP format)
- Add endpoint reference for status retrieval

Generated with [Claude Code](https://claude.com/claude-code)
@mhrivnak

Copy link
Copy Markdown
Member

First thought; let's clarify the scope of this. I was confused because we already have cluster-aaS. But it looks like this is actually about proposing a more explicit API.

@tzvatot

tzvatot commented Apr 13, 2026 •

Copy link
Copy Markdown
Contributor Author

First thought; let's clarify the scope of this. I was confused because we already have cluster-aaS. But it looks like this is actually about proposing a more explicit API.

Right. The CaaS capability already exists. This proposal promotes cluster configuration from the opaque template_parameters map and hardcoded Ansible values into explicit typed ClusterSpec fields (pull_secret, ssh_public_key, release_image, and a network sub-message with pod_cidr/service_cidr).

This follows the same approach already taken for VMaaS.

I've updated the summary and motivation to make this clear.

Update Summary and Motivation to clarify that this proposal defines
the tenant-facing API contract for the existing CaaS capability,
not introducing CaaS itself.

Generated with [Claude Code](https://claude.com/claude-code)
@tzvatot

tzvatot commented Apr 13, 2026

Copy link
Copy Markdown
Contributor Author

I've updated the Summary and Motivation sections to make this distinction clearer.

tzvatot added 9 commits April 15, 2026 10:52
Deletion uses deletion_timestamp with finalizers (standard Kubernetes
pattern) rather than a separate CLUSTER_STATE_DELETING enum value.

Generated with [Claude Code](https://claude.com/claude-code)
Replace the generic map<string, Any> template_parameters approach with
explicit typed fields (pull_secret, ssh_public_key) in ClusterSpec.
This provides type safety, discoverability, and proto-level validation.

Updated all examples and workflow descriptions to reflect the new API
contract.

Generated with [Claude Code](https://claude.com/claude-code)
Promote pull_secret, ssh_public_key (from template_parameters), and
release_image, cluster_network_cidr, service_network_cidr (from
hardcoded values) to explicit ClusterSpec fields. All fields are
optional with sensible defaults. CIDRs use plain string notation,
consistent with VirtualNetwork/Subnet protos and Kubernetes conventions.

Generated with [Claude Code](https://claude.com/claude-code)
The purpose of this proposal is moving cluster configuration from
templates/playbooks into the tenant-facing API. Updated Summary and
Motivation to clearly state this: parameters are currently hidden in
template_parameters or hardcoded in Ansible roles, and this proposal
makes them explicit API fields.

Generated with [Claude Code](https://claude.com/claude-code)
User stories now focus on tenant control over cluster configuration
(pull secret, SSH key, OCP version, networking CIDRs). Goals now
focus on defining explicit ClusterSpec fields to replace the opaque
template_parameters mechanism.

Generated with [Claude Code](https://claude.com/claude-code)
Non-goals now clarify what stays out: template-specific parameters,
provider-managed settings, workflow changes, and removal of
template_parameters.

Generated with [Claude Code](https://claude.com/claude-code)
Replace the full workflow descriptions with a clear current-state vs
proposed-change structure. Node scaling, deletion, and template
management workflows are unchanged and noted as such.

Generated with [Claude Code](https://claude.com/claude-code)
Remove repeated documentation of existing cluster states, conditions,
credentials, and status examples — these are unchanged. API section now
only shows the new ClusterSpec fields. Drawbacks and Risks updated to
be specific to this proposal.

Generated with [Claude Code](https://claude.com/claude-code)
Every section from the enhancement template is present. Each concept
appears exactly once. The proposal focuses solely on promoting
template_parameters and hardcoded values to explicit ClusterSpec fields.

Generated with [Claude Code](https://claude.com/claude-code)

@adriengentil adriengentil left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good overall!

Comment thread enhancements/caas/README.md
Comment thread enhancements/caas/README.md Outdated
Comment thread enhancements/caas/README.md
Comment thread enhancements/caas/README.md
@avishayt

Copy link
Copy Markdown
Contributor

BTW, it would be interesting to learn from ROSA's API:

  1. Consider how they separate out the networking configuration - I think it will lead to a cleaner API once it grows:
  {               
    "network": {
      "pod_cidr": "10.128.0.0/14",
      "service_cidr": "172.30.0.0/16",
      "machine_cidr": "10.0.0.0/16",                                                                                                                                                                                
      "host_prefix": 23                                                                                                                                                                                             
    }                                                                                                                                                                                                               
  }
  1. They specify openshift_version rather than release_image, but release_image is better for sovereign so let's leave it as is.
  2. Pull secret - they take it from OCM, we don't have that luxury.
  3. SSH key - they reference AWS key pairs, we should follow a similar model in the future by creating a separate resource (not for now)
  4. ROSA has a nice grouping of related fields (similar to the first point above):
  {                                                                                                                                                                                                                 
    "name": "my-cluster",
    "cloud_provider": {...},
    "aws": {  // Cloud-specific nested
      "access_key_id": "...",                                                                                                                                                                                       
      "subnet_ids": [...],                                                                                                                                                                                          
      "tags": {...}                                                                                                                                                                                                 
    },                                                                                                                                                                                                              
    "network": {...},                                                                                                                                                                                               
    "nodes": {...},                                                                                                                                                                                                 
    "version": {"id": "openshift-v4.17.0"},                                                                                                                                                                         
    "state": "ready",  // Read-only status 
    "api": {...},      // Read-only endpoints                                                                                                                                                                       
    "console": {...}   // Read-only endpoints                                                                                                                                                                       
  }

tzvatot added 2 commits April 16, 2026 10:25
Mark pull_secret as write-only (redacted in GET responses) per
adriengentil's feedback. Add note that templates can override system
defaults for any field.

Generated with [Claude Code](https://claude.com/claude-code)
Move pod_cidr and service_cidr into a nested network message for
cleaner API organization as the networking config grows. Updated
examples and CLI flags.

Generated with [Claude Code](https://claude.com/claude-code)
@tzvatot

tzvatot commented Apr 16, 2026

Copy link
Copy Markdown
Contributor Author

BTW, it would be interesting to learn from ROSA's API:
...

Good points. Grouped the networking CIDRs into a ClusterNetwork sub-message with pod_cidr and service_cidr - easier to extend later (e.g., machine_cidr, host_prefix).

Agreed on keeping release_image. SSH key as separate resource noted for the future (see discussion on the inline thread about sync concerns with shared resources).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants