Virtual machines as a service - #9
Conversation
5efe03c to
3ba0dc3
Compare
This enhancement proposes a virtualization-as-a-service feature.
3ba0dc3 to
2680ca6
Compare
|
I was going to put in a smaller enahncement request as an alternative, but I think better to write comments here after looking at the existing demo. At a fundamental level, I think before we flesh out an enhancement request, we need quick alignment on the goals and MVP here. I see two use cases that we can go after here:
For completness a third use case which I believe is separate is a virtual OpenShift cluster as a service; where we basically take the current bare metal stuff and have it work on VMs. This is super valuable for developers that want a single GPU... and want to spin up and down clusters. I think we should write up this use case seperately, so ignoring for now. I think that the current enhancement request is neither fish nor fowel, i.e., it includes some discussion of features needed for the second use case, but we really want to instantiate the entire virtual data center as a single thing.... My inclination is that we should target the first use case, which is basically the current spike adding external networking and a catalog of templates with a way to add templates. I would cut a large part of the current enahcement request that is really focused on the VDC model. I think we can have this running very quickly, and it will help people that find our user experience a barrier for openshift... This means we have something real and in production quickly with VMs. Personally, since its almost a trival addition, I would have a container as a service, where its the same, but we deploy containers instead of VMs. In parallel I think we should write up a VDCaaS enhancement request, that will be much more complicated, and start fleshing it out over a few sprints. I think its worth alignment of the team first on if we want something super simple that proves VMs are in scope, or if the value all comes from the VDCaaS. I think its also worth writing up a virtual openshift cluster as a service. It would be good to understand how easy the latter is, and if its worth prioritizing before VDCaaS because we can get it running quickly. I am personally baised towards the two extremes; expsing individual VMs or k8s, since I think most workloads will be on k8s. I do get that there may be reasons why tenants may want a whole virtual data center that is not running k8s, e.g. one of the reasons we are doing the bare metal layer is to support SLURM. So, a super simple VDC environment could be a SLURM cluster; but they really do want bare metal. From my biased perspective, I think we shold |
|
Thanks for your feedback @okrieg, it helps a lot! I agree with you, this design tries to include the networking capabilities that we don't have, and it make it hard to move forward. If you see a value going to the simple VM service, then I think we could add latter an enhancement to integrate VDCaaS into this solution. I like this proposal because, as you said, it's something we can do now. I still have questions:
I think these questions apply to container-as-a-service as well. About Virtual OpenShift cluster as a service should be trivial to implement through the current Cluster-as-a-service, as hypershift abstracts OCP Virt away. I wonder if it can even be designed as a simple template in our current flow. |
|
Do you really mean only 1 external floating IP address Orran? We usually get requests for more than one from the MOC users who want VMs (for example to make a database accessible separately from an http interface). Re: the registry of images, we already support quay.io for that in MOC. There is not a lot of duplication in the images that VM users need currently, based on MOC demand, so adding this in to O-SAC adds a lot of maintenance and upgrading debt without a big customer advantage. For managing the VMs, we would need to listen closely to the ops folks who already do this with OpenStack. They value simplicity and reliability highly. I am sure they will prefer the cluster(s) supporting simple VMs to be separate from the service clusters we are deploying. They already support secrets and will need to continue doing so. I support the suggestion to submit the enhancement request only to support the simple VM case and not the other services yet at this point. |
Cool!
Definately, never run anything on HUB; HUB won't have GPUs... I think we want a single cluster for most VMs, and a ticky box where they can say they don't want to share, and in that case have VMs on a tenant specific cluster.
The actual registry should be I assume quay.io, but we want tenant and provider to be able to upload images right?
Yes, but I don't have good understanding of this.
Really, I think we should think seriously then about doing this ASAP.
Am tempted to start with the easiest thing possible, and then add more... runpod, as example, just has one IP address. For VDC as a service, all DB access... are internal to the VDC, so the floating IP is just for external access I think.
Is that on-premis? I assume so. Yes, don't see adding another registry, just adding interfaces to get to it.
YES!!!
|
|
|
||
| ### OS base image management | ||
|
|
||
| Tenants will have the ability to import their own base images. Since OSC backend relies on multiple OCP clusters, and tenants’ VMs may be distributed on multiple ones, base images be centralized on an OCI container registry before being consumed by KubeVirt on the destination cluster. |
There was a problem hiding this comment.
Could we save "import their own base images" for a future enhancement, and start by relying on the templates to define their own fixed images?
There was a problem hiding this comment.
yes, has discussed, we can assume the images live in quay.io
|
|
||
| - The main point is to provide a public API that will select a management cluster to provision a VM (given requested resources CPU/Mem/GPU)? => not the priority | ||
| - What is the added value on top of kube virt? => multi-tenancy, multiple virt clusters, higher level networking primitives | ||
| - Since we plan to rely on the OCP stack (ACM/KubeVirt/Hypershift), is there something we want to make pluggable using Ansible? (access network configuration to the VM?) |
There was a problem hiding this comment.
There are Ansible roles for kubevirt and creating VMs and getting VM info:
- https://docs.ansible.com/ansible/latest/collections/kubevirt/core/index.html
And the more official Ansible collection from Red Hat OpenShift Virtualization: - https://catalog.redhat.com/en/software/collection/redhat/openshift_virtualization#documentation
| - What is the added value on top of kube virt? => multi-tenancy, multiple virt clusters, higher level networking primitives | ||
| - Since we plan to rely on the OCP stack (ACM/KubeVirt/Hypershift), is there something we want to make pluggable using Ansible? (access network configuration to the VM?) | ||
| - Do we need to provision infra on-demand to run kubevirt workload? | ||
| - What about networking isolation, how kubevirt works? What model to prioritize? |
There was a problem hiding this comment.
There is Red Hat documentation for User defined networks in Red Hat OpenShift Virtualization, and setting up a UserDefinedNetwork resource that provides network isolation between tenants within the same namespace.
|
I updated the doc following our discussion yesterday, I think there is still some things that are not clear but I like to push my work before the end of the day. While working on it, I think that without the concept of regions (mapped on HUB clusters?) in the Fulfillment Service, we can't move much forward with:
I'm saying that, because we talked about the ability to attach/de-attach additional storage, and have some private networking feature. At the moment I propose a simple VM service where each VM is isolated from each other (in a UDN L2), and all VMs have a floating IP assigned to them. Do you think its enough to get a useful service? Or should I explore a bit more? |
aab7064 to
c4d2138
Compare
| 4. The O-SAC Operator detects the new VirtualMachine CR and triggers the reconciliation process. | ||
| 5. The Operator, via AAP (Ansible Automation Platform), performs the following automation steps: | ||
| - Creates a dedicated namespace for the VM (if not already present) | ||
| - Provisions the required network resources (e.g., UDN L2 network) |
There was a problem hiding this comment.
Yes, I will add few words about it in the proposal section
| - Assigns a floating IP to the VM using ESI APIs | ||
| - ...other operations depending on the selected virtual machine template | ||
| 6. The Operator monitors the status of the VM and updates the VirtualMachine CR status accordingly. | ||
| 7. The tenant can query the status of the VM via the Fulfillment CLI or API, and access the VM using the assigned floating IP. |
There was a problem hiding this comment.
Will we block other tenant watch/edit VMs that don't belong to the tenant?
There was a problem hiding this comment.
Yes, VMs are scoped to the tenant
| 2. The Fulfillment Service receives the request and validates: | ||
| - The existence and availability of the specified template | ||
| - The correctness and completeness of the provided parameters | ||
| 3. The Fulfillment Service creates a new VirtualMachine custom resource (CR) in the appropriate namespace. |
There was a problem hiding this comment.
We have to find another name, as VirtualMachine is already used name by CNV. Maybe ComputeInstance?
There was a problem hiding this comment.
Looks good to me! As it's a name commonly used in cloud providers
okrieg
left a comment
There was a problem hiding this comment.
I think this is a good start, lets get it in there. Don't think this should be limtied to Hub cluster, but think that is a reasonable start.
|
|
||
| #### Virtual machines on HUB cluster | ||
|
|
||
| Virtual machines will be created on the HUB cluster that was selected by Fulfillment Service, it was discussed to create a dedicated cluster to handle VM workloads using Cluster-as-a-Service API, but since they are HostedCluster it won't increase the reliability of the solution, as their reliability are tied to the same HUB cluster. |
There was a problem hiding this comment.
I am not a fan of creating this on the HUB cluster long term because I don't think we want to put any GPUs in the hub cluster, they should all be in compute clsuters. I think its fine for the first MVP to put in the Hub cluster.
There was a problem hiding this comment.
I agree that we should be able to manage VMs remotely. I think it should be part of another enhancement as is implies some challenges of its own, and I would like to understand the constraints we have on such clusters (can we deploy our stack on them, or should they stay vanilla as possible, and just manage kubevirt CRs).
18cafd9 to
d44dcf6
Compare
|
Before going forward with this design, I would like to check 2 assumptions:
|
knikolla
left a comment
There was a problem hiding this comment.
Proposal seems good and agree with general direction.
| Because O-SAC does not yet provide a VDCaaS (Virtual Data Center as a Service) | ||
| layer, and the Fulfillment Service cannot guarantee that two virtual machines | ||
| will be provisioned on the same HUB cluster, each virtual machine must be | ||
| assigned a floating IP. This ensures that every VM is accessible regardless of |
There was a problem hiding this comment.
This comment is from the perspective of the MOC/NERC.
- I worry about the scalability of assigning one floating IP per each VM that is running. IIRC, the floating IP pool that we have on the NERC's OpenStack is a /23 and for that we have generally been stingy about handing too many out.
- Maybe this isn't a concern scope-wise in the short term and future features that may be implemented (eg. VDCaaS) will provide an off ram before we hit scalability problems.
There was a problem hiding this comment.
I agree, this current design will need to be reviewed as soon as VDCaaS is delivered. I'll add your concerns in the document!
There was a problem hiding this comment.
Agreed. I think we can describe this as just a starting point until we have a more complete networking architecture agreed upon.
There was a problem hiding this comment.
Added in the "drawbacks" section
mhrivnak
left a comment
There was a problem hiding this comment.
I left some feedback, but this is a good proposal. I think we should merge it soon once we get all the comments resolved.
| } | ||
| ``` | ||
|
|
||
| ### Implementation Details/Notes/Constraints |
There was a problem hiding this comment.
nit: A lot of the content above should probably be in this "implementation details" section. That helps keep the Proposal succinct and focused.
There was a problem hiding this comment.
can you be more specific? I guess the last 2 paragraphs in the proposal section? Are the workflows too detailed too ?
There was a problem hiding this comment.
I move some of the details in this section
| ### Implementation Details/Notes/Constraints | ||
|
|
||
|
|
||
| #### Virtual machines on HUB cluster |
There was a problem hiding this comment.
I agree with using the hub cluster as a starting point. But longer-term, I don't think reliability is the issue. The reasons for using a separate cluster to host VMs would include:
- further isolates customer workloads for the cloud provider's management tooling
- enables the management cluster to be lifecycled (upgraded, fixed, migrated, etc) at a different pace than the cluster that customers depend on.
- the cluster hosting VMs needs different uptime characteristics, and so it might be monitored and managed differently even if its control plane is on shared hardware with the management cluster.
- makes it easier to have per-tenant virt clusters if/when that becomes a need for tenants who need stronger isolation
There was a problem hiding this comment.
I added you points in the doc !
| Because O-SAC does not yet provide a VDCaaS (Virtual Data Center as a Service) | ||
| layer, and the Fulfillment Service cannot guarantee that two virtual machines | ||
| will be provisioned on the same HUB cluster, each virtual machine must be | ||
| assigned a floating IP. This ensures that every VM is accessible regardless of |
There was a problem hiding this comment.
Agreed. I think we can describe this as just a starting point until we have a more complete networking architecture agreed upon.
|
|
||
| ## Alternatives (Not Implemented) | ||
|
|
||
| TBD |
There was a problem hiding this comment.
Putting all of a tenant's machines on the same isolated network might be another reasonable starting point. It's probably worth some commentary on why you're not proposing that direction.
There was a problem hiding this comment.
While investigating Openshift networking in the last few days, I saw that it is recommended to not have more then 100 UDNs per cluster. So I actually think that having a UDN/tenant would be beneficial. VMs will still be required to talk to each others using the floating IP, as it doesn't solve the VM placement by the fulfillment-service.
There was a problem hiding this comment.
I updated to a UDN per tenant highlighting the limitation of 100 UDNs
|
|
||
| TBD | ||
|
|
||
| ## Open Questions [optional] |
There was a problem hiding this comment.
how to influence which cluster a VM would land on?
There was a problem hiding this comment.
I think it's a non-goal in the scope of this proposal. It should probably be covered by VDC proposal, as the VM placements will likely be linked to the way the networks are defined. WDYT?
The information I got was not right, this is not working. We'll need to let the tenant specify ports they want to expose to the outside
This is working |
355effe to
184d82e
Compare
|
@okrieg can the PR be merged? Adrien answered all the comments. |
|
I think this is in good shape! Unfortunately I cannot approve it because I'm the one who originally created it :). We can merge once we have @mhrivnak's approval, since he owns the "requested changes" flag. |
- Resolve AdminNetworksPage topology view wording contradiction (issue osac-project#3) - Specify IPv4/IPv6 CIDRs explicitly in FR-6 (issue osac-project#4) - Pick side drawer pattern for subnet detail display in FR-12 (issue osac-project#5) - Add Priority field to SecurityGroup rule specification in FR-18 (issue osac-project#6) - Standardize PublicIP action terminology to 'Release' in FR-28 (issue osac-project#7) - Document multi-NIC same-VN constraint rationale in FR-34 (issue osac-project#8) - Align wizard empty-state flow with inline overlay pattern in FR-38 (issue osac-project#9) - Define Retry action API contract in FR-42 (issue osac-project#10) - Clarify Subnet endpoints are create/delete only in FR-45 (issue osac-project#11) - Remove redundant NFR-9 (issue osac-project#12) - Update Open Question 8.2 wording to match Non-Goals Assisted-by: Claude Code <noreply@anthropic.com> Signed-off-by: Elay Aharoni <elayaha@gmail.com>
BMaaS design: updated 8 stale "host/inventory system" IP discovery references to match resolved OQ#4 (create_network_attachment queries DHCP lease table, returns IP as role output). Unified design: updated BMaaS IP discovery from "open question" to resolved mechanism in preconditions and IP discovery tables. Default design: fixed ExternalIPAttachment precondition from stale VirtualMachineReference to compute_network_attachment_statuses IP. CaaS design: added AgentStatus.IPAddress discovery mechanism to reconcileNetworking step and component responsibility table. Unified PRD: updated Gap osac-project#9 — hairpin NAT superseded by MetalLB VIP direct access on same subnet. VMaaS PRD: added NATGateway non-cleanup statement to FR-4. Step numbering: fixed gaps in VMaaS (8→10), BMaaS (9→11), Default (9→13) designs. OQ numbering: restored missing OQ#5 (CaaS) and OQ#1 (BMaaS) as resolved entries. Assisted-by: Claude Code <noreply@anthropic.com> Signed-off-by: Dan Manor <dmanor@redhat.com>
BMaaS design: updated 8 stale "host/inventory system" IP discovery references to match resolved OQ#4 (create_network_attachment queries DHCP lease table, returns IP as role output). Unified design: updated BMaaS IP discovery from "open question" to resolved mechanism in preconditions and IP discovery tables. Default design: fixed ExternalIPAttachment precondition from stale VirtualMachineReference to compute_network_attachment_statuses IP. CaaS design: added AgentStatus.IPAddress discovery mechanism to reconcileNetworking step and component responsibility table. Unified PRD: updated Gap osac-project#9 — hairpin NAT superseded by MetalLB VIP direct access on same subnet. VMaaS PRD: added NATGateway non-cleanup statement to FR-4. Step numbering: fixed gaps in VMaaS (8→10), BMaaS (9→11), Default (9→13) designs. OQ numbering: restored missing OQ#5 (CaaS) and OQ#1 (BMaaS) as resolved entries. Assisted-by: Claude Code <noreply@anthropic.com> Signed-off-by: Dan Manor <dmanor@redhat.com>
BMaaS design: updated 8 stale "host/inventory system" IP discovery references to match resolved OQ#4 (create_network_attachment queries DHCP lease table, returns IP as role output). Unified design: updated BMaaS IP discovery from "open question" to resolved mechanism in preconditions and IP discovery tables. Default design: fixed ExternalIPAttachment precondition from stale VirtualMachineReference to compute_network_attachment_statuses IP. CaaS design: added AgentStatus.IPAddress discovery mechanism to reconcileNetworking step and component responsibility table. Unified PRD: updated Gap osac-project#9 — hairpin NAT superseded by MetalLB VIP direct access on same subnet. VMaaS PRD: added NATGateway non-cleanup statement to FR-4. Step numbering: fixed gaps in VMaaS (8→10), BMaaS (9→11), Default (9→13) designs. OQ numbering: restored missing OQ#5 (CaaS) and OQ#1 (BMaaS) as resolved entries. Assisted-by: Claude Code <noreply@anthropic.com> Signed-off-by: Dan Manor <dmanor@redhat.com>
BMaaS design: updated 8 stale "host/inventory system" IP discovery references to match resolved OQ#4 (create_network_attachment queries DHCP lease table, returns IP as role output). Unified design: updated BMaaS IP discovery from "open question" to resolved mechanism in preconditions and IP discovery tables. Default design: fixed ExternalIPAttachment precondition from stale VirtualMachineReference to compute_network_attachment_statuses IP. CaaS design: added AgentStatus.IPAddress discovery mechanism to reconcileNetworking step and component responsibility table. Unified PRD: updated Gap osac-project#9 — hairpin NAT superseded by MetalLB VIP direct access on same subnet. VMaaS PRD: added NATGateway non-cleanup statement to FR-4. Step numbering: fixed gaps in VMaaS (8→10), BMaaS (9→11), Default (9→13) designs. OQ numbering: restored missing OQ#5 (CaaS) and OQ#1 (BMaaS) as resolved entries. Assisted-by: Claude Code <noreply@anthropic.com> Signed-off-by: Dan Manor <dmanor@redhat.com>
This enhancement proposes a virtualization-as-a-service feature.
This is an extract of the "VM fulfillment" section of the design doc.
This document is going to need a lot of massaging, both to address all the questions from the google doc and to fill out the missing sections of the proposal.