From 2680ca6136a3e8eb7977a239827333effa6b05c8 Mon Sep 17 00:00:00 2001 From: Lars Kellogg-Stedman Date: Mon, 15 Sep 2025 16:50:51 -0400 Subject: [PATCH 1/7] Virtual machines as a service This enhancement proposes a virtualization-as-a-service feature. --- enhancements/vmaas/README.md | 472 +++++++++++++++++++++++++++++++++++ 1 file changed, 472 insertions(+) create mode 100644 enhancements/vmaas/README.md diff --git a/enhancements/vmaas/README.md b/enhancements/vmaas/README.md new file mode 100644 index 000000000..555c583ff --- /dev/null +++ b/enhancements/vmaas/README.md @@ -0,0 +1,472 @@ +--- +title: Virtualization as a service +authors: + - Adrien Gentil +creation-date: 2025-09-15 +last-updated: 2025-09-15 +tracking-link: + - TBD +see-also: +replaces: +superseded-by: +--- + +# Virtualization as a service + + +## Summary + +We want to provide a service that allows a tenant to create and manage VMs. + +## Motivation + +TBD + +### User Stories + +#### As a tenant + +- I want to list a pre-defined VM templates +- I want to create a VM based on a pre-defined template +- I want a VM that has access to specialized hardware (e.g.: GPU) +- I want to customize the pre-defined template for my usage +- I want to attach and detach block storage on my VM that transcends the lifecycle of the VM +- I want to manage the lifecycle of my VM (start/stop/terminate) +- I want to import my OS base image +- I want to list available base OS images +- I want to manage (create, list, delete) my private network +- I want to attach my VMs in a virtual private network +- I want my VM to be exposed outside of my private network +- I want my VMs to be attached to my own physical L2 network +- I want to connect on my VM through serial console +- I want to keep running my VM when a node goes out of service + +#### As a service provider: + +- I want to define VM templates +- I want to manually trigger VM migration + +### Goals + +TBD + +### Non-Goals + +TBD + +## Proposal + +### Network isolation + +This design relies on UDN (User Defined Networking, it is the networking component provided with Openshift, it allows the definition of virtual private subnets, and provides the ability to connect on physical networks (localnet). + +Moverover, Kubevirt supports UDN layer 2 mode which offers the ability to live-migrate VMs across OCP nodes. + +### Ingress + +We can expose services running on VM(s) using a “LoadBalancer” service (e.g.: relying on MetalLB). + +### APIs + +Virtual private subnet creation: + +``` +apiVersion: cloudkit.openshift.io/v1alpha1 +kind: VirtualPrivateSubnet +metadata: + name: example + finalizers: + - cloudkit.openshift.io/finalizer +annotations: + cloudkit.openshift.io/ref-count: 1 # +1 when a VM is created, -1 when a VM is deleted   +spec: + ipv4Cidr: 192.168.100.0/24 +``` + +We will create a new, unique namespace, and the UDN definition to create a L2 layer and the corresponding subnet. + +Since several VMs can be attached to the same network, we introduce an annotation that will be used to count the number of VMs attached to this network, and a finalizer to prevent its deletion before all the VMs are deleted. + +Load balancer: + +``` +apiVersion: cloudkit.openshift.io/v1alpha1 +kind: LoadBalancer +metadata: + name: example +spec: + hostRefs: + - kind: VirtualMachineHost + name: exampleVM + portMapping: + - protocol: TCP + externalPort: 80 + internalPort: 80 +``` + +All hosts must be associated with the same network. + +### Template management + +Like cluster fulfillment, there will be a mechanism to define templates as Ansible roles. Templates contain predefined VM classes which define default OS base image, CPU architecture, number of cores, amount of memory, storage, user data, … + +A template VM might define the following properties, and let a tenant to customize some of them (should be a 1:1 mapping to KubeVirt resources): + +- Name +- Tenant +- OS base image +- Instance class, which specify the family (translating into general/compute/memory/gpu/arch), and the amount of CPU and memory (TBD, e.g.: general.large) +- User-data (cloud-init) +- Networking, namespace + UDN at first, and localnet later? + * Layer 3 only for Networking. +- SSH keys +- Storage, size of the boot volume (+ I/O performance, encryption) +- State (started/stopped) + +There will be an API to publish these templates in the fulfillment service, and list these templates from it. + +KubeVirt UI provides a functionality based on OCP templates that generates a VirtualMachine resource, it is safe to assume that we can use any tool we want to, and rely on Ansible/AAP to define our own. + +### VM management + +VMs will be managed using KuberVirt, basically the user will have to provide 3 informations to create a VM: + +- Template name +- Template overrides +- A network + +### OS base image management + +Tenants will have the ability to import their own base images. Since OSC backend relies on multiple OCP clusters, and tenants’ VMs may be distributed on multiple ones, base images be centralized on an OCI container registry before being consumed by KubeVirt on the destination cluster. + +### APIs + +``` +apiVersion: cloudkit.openshift.io/v1alpha1 +kind: VirtualMachineHost +metadata: + name: example + finalizers: + - cloudkit.openshift.io/finalizer # to update ref-count in VirtualPrivateNetwork +spec + networkRef: + kind: VirtualPrivateNetwork + name: exampleNetwork + template: rhel-10 + templateParameters: + instanceType: o1.small + SSHPublicKey: ... + powerState: stopped +``` + +### Workflow Description + +TBD + +### API Extensions + +TBD + +### Implementation Details/Notes/Constraints + +#### VM creation minimal flow + +1. Tenant creates a VM Host, the automation: + 1. Creates a namespace + 2. Creates a UDN L2 network + 3. Creates a Kubevirt CR + +#### Tasks + +- Fulfilment-cli + * CRUD on VirtualMachineHost + * List VM templates +- Fulfillment-service: + * [Create VirtualMachineHost management API in fulfillment-service](https://github.com/innabox/issues/issues/201&sa=D&source=editors&ust=1757970010188331&usg=AOvVaw07B9NvURxHW07nF4X-1X41) + * [Publish VM templates endpoint in fulfillment-service](https://github.com/innabox/issues/issues/203&sa=D&source=editors&ust=1757970010188598&usg=AOvVaw35ilTLz7wwMg4_W6YdNeMB) + * [List VM templates](https://github.com/innabox/issues/issues/202&sa=D&source=editors&ust=1757970010188846&usg=AOvVaw2vg_K25W8xerGfSgYxrI4z) +- Cloudkit-operator: + * [Create VirtualMachineHost controller](https://github.com/innabox/issues/issues/204&sa=D&source=editors&ust=1757970010189157&usg=AOvVaw3brz_MVBsWOSWzjjv6Mv9P) +* AAP: + - [Create/Delete VM EDA and playbooks](https://github.com/innabox/issues/issues/207&sa=D&source=editors&ust=1757970010189421&usg=AOvVaw3xVLsU3dXhvFOKs1ZxGPzf + - [Create VM templates](https://github.com/innabox/issues/issues/205&sa=D&source=editors&ust=1757970010189609&usg=AOvVaw0buvglSiCKKLOSh-VBrxJH) + - [Publish VM templates to fulfillment-service](https://github.com/innabox/issues/issues/206&sa=D&source=editors&ust=1757970010189850&usg=AOvVaw1JCATg3Xsg0QA93zUyQUZf) + +Development: + +- Public API to manage networks + operator + AAP (?) +- Public API to manage VMs + operator + AAP +- Public API to publish template +- Create VM template in AAP + publication +- OS base image management + operator + +#### UDN layer 2 definition + +``` +apiVersion: v1 +kind: Namespace +metadata: + name: udn-dev + labels: + k8s.ovn.org/primary-user-defined-network: "" +--- +apiVersion: k8s.ovn.org/v1 +kind: UserDefinedNetwork +metadata: + name: udn-dev + namespace: udn-dev +spec: + layer2: + ipam: + lifecycle: Persistent + role: Primary + subnets: + - 10.200.0.0/16 + topology: Layer2 +``` + +#### UDN localnet definition + +Requires NMState operator: + +``` +apiVersion: nmstate.io/v1 +kind: NodeNetworkConfigurationPolicy +metadata: + name: mapping +spec: + nodeSelector: + node-role.kubernetes.io/worker: '' + desiredState: + ovn: + bridge-mappings: + - localnet: localnet1 + bridge: br-ex + state: present +--- +apiVersion: k8s.ovn.org/v1 +kind: ClusterUserDefinedNetwork +metadata: + name: cudn-localnet +spec: + namespaceSelector: + matchExpressions: + - key: kubernetes.io/metadata.name + operator: In + values: ["red", "blue"] + network: + topology: Localnet + localnet: + role: Secondary + physicalNetworkName: localnet1 + ipam: + mode: Disabled +``` + +#### Live Migrations using UDN and KubeVirt + +- Kubevirt only allows/supports UDN L2 +- Localnet is not ideal due to the static nature of this approach and the VLAN is not necessarily the same on the destination node. (IE VxLAN endpoints) +- Complex hardware dependencies are hit and miss.  It all depends on the hardware, but there is the ability to migrate things like GPUs and NICs (SR-IOV) in specific circumstances. +- Here is an instance where we can use templates to avoid the complexity. +- Things to be aware of which could cause problems: + * Localnet networks which tie to specific h/w i/f., + * SR-IOV, + * DPDK, + * Node selecting or filtering logic, + * Possible DHCP issues. + +- Live migrations describe the type of migration and effected changes. + + * L2 UDN: Updates MAC learning tables, & ARP caches. + * L3 UDN: Updates node IP routing tables pointing to the new hosting node. + * L3 migrations should not be interpreted as allowing an IP address to change on the fly. + +#### UDN Topology Types + +UDN has topology types (L2/L3) with different block types highlighting the differences between one or the other. + +Layer2 (layer2: block) + +- Bridged networking - VMs appear on same broadcast domain +- DHCP/Static IP assignment - Needs IPAM (IP Address Management) for IP allocation +- East-west traffic flows directly between VMs without routing +- Flat network model - All VMs share the same subnet + +Layer3 (layer3: block): + +- Routed networking - Each node gets its own subnet slice +- Distributed subnets - hostSubnet defines how the CIDR is carved up per node +- Inter-node routing required for VM-to-VM communication +- Scalable addressing - Prevents IP exhaustion on large clusters + + Configuration Differences: + +- Layer2 needs IPAM for IP management + + ``` + layer2: + ipam: + lifecycle: Persistent # How long IPs are reserved + subnets: [10.200.0.0/16] # Flat subnet + ``` + +- Layer3 needs subnet distribution + + ``` + layer3: + subnets: + - cidr: 10.200.0.0/16 # Overall range + hostSubnet: 24 # Each node gets /24 slice + ``` + +#### Tasks + +- List options in kubevirt that prevents Live VM migration +- Try UDN + kubevirt + - Describe how Egress IP works + - Default Egress + - Why do we need an Egress IP + - Look at the manifests, will help to design API +- Describe Template and MVP + - VM Host / HostPool + - Disk + - Memory + - Storage + - Persistent vs Ephemeral + - SSH keys + - Cloud-init + - Network + - OS base image (Registry Service) + - Test: + - Create network + - Create VM in the network + - Communicate between Nodes (2 VMs) + - Floating IP / LoadBalancer +- Create a sequence diagram, showing service/operator/AAP responsibilities + - VM management + - Base image management + - Network management +- Get information about Metal LB and UDN’s floating IP is supposed to work in order to grab a public IP (and other related networking stuff) + +#### External resources + +- [Red Hat Sovereignty Technical Architecture](https://docs.google.com/presentation/d/1g23omSQ46NZ6qyDi-I4Q8mGNt-NQpKk0vOuaZf8Usqw/edit?slide%3Did.g344ca2ed691_0_2780%23slide%3Did.g344ca2ed691_0_2780&sa=D&source=editors&ust=1757970010209722&usg=AOvVaw0KkWOZaTJyX2nac50mzfW) +- [RHSovCloud-Network-Scenarios-0.1](https://docs.google.com/presentation/d/1_wYAbfoCcmuIOx-esCcOoSFJFR7tgkrypyyIHa7S4AY/edit?slide%3Did.g30e84f14e68_0_6149%23slide%3Did.g30e84f14e68_0_6149&sa=D&source=editors&ust=1757970010209955&usg=AOvVaw3oLbSpaYriqUYAxZ1TQlWX) +- [Red Hat Sovereignty Technical Architecture](https://docs.google.com/presentation/d/1g23omSQ46NZ6qyDi-I4Q8mGNt-NQpKk0vOuaZf8Usqw/edit?slide%3Did.g344ca2ed691_0_2780%23slide%3Did.g344ca2ed691_0_2780&sa=D&source=editors&ust=1757970010210162&usg=AOvVaw1NbKJtheHnk47tN4Ot5KVy) + +UDN:   + +- [User Defined Networks on OpenShift](https://docs.google.com/presentation/d/1Hx1Fzm1F9EkmqrmTjbMHPBAuIls-2-oK1IVrzFnW3L4/edit?slide%3Did.g31966f3f64c_0_855%23slide%3Did.g31966f3f64c_0_855&sa=D&source=editors&ust=1757970010210475&usg=AOvVaw3lrbnP41fTOVcfwnbXAuwW) +- +- [Chapter 2. Primary networks | Multiple networks | OpenShift Container Platform | 4.19 | Red Hat Documentation](https://docs.redhat.com/en/documentation/openshift_container_platform/4.19/html/multiple_networks/primary-networks%23about-user-defined-networks&sa=D&source=editors&ust=1757970010211384&usg=AOvVaw22g82g1WsGtCpIV_OOnpnt) +- [Chapter 3. Secondary networks | Multiple networks | OpenShift Container Platform | 4.19 | Red Hat Documentation](https://docs.redhat.com/en/documentation/openshift_container_platform/4.19/html/multiple_networks/secondary-networks%23configuration-localnet-switched-topology_configuring-additional-network-ovnk&sa=D&source=editors&ust=1757970010211903&usg=AOvVaw0q1VmdCCcREjz1O3xKPmt) +- [User defined networks in Red Hat OpenShift Virtualization](https://www.redhat.com/en/blog/user-defined-networks-red-hat-openshift-virtualization&sa=D&source=editors&ust=1757970010212242&usg=AOvVaw2KBAek4qp-jL8eF_I9gN5) + +### Risks and Mitigations + +TBD + +### Drawbacks + +TBD + +## Alternatives (Not Implemented) + +TBD + +## Open Questions [optional] + +- The main point is to provide a public API that will select a management cluster to provision a VM (given requested resources CPU/Mem/GPU)? => not the priority +- What is the added value on top of kube virt? => multi-tenancy, multiple virt clusters, higher level networking primitives +- Since we plan to rely on the OCP stack (ACM/KubeVirt/Hypershift), is there something we want to make pluggable using Ansible? (access network configuration to the VM?) +- Do we need to provision infra on-demand to run kubevirt workload? +- What about networking isolation, how kubevirt works? What model to prioritize? +- Representation of resources a la AWS? `.` (e.g.: t3.xlarge (general purpose), c3.small (compute), m3.medium (memory)) +- Storage options? +- Should we focus on cluster provisioning or VM provisioning? +- Which networking model should we prioritize? The ability to connect on a physical network or the ability to create a virtual private network? +- How to make sure we co-collocate network and VM definitions together? +- From what I understand, there will be different ways to handle networking depending if we want VM workloads (UDN), BM workloads (Openstack), or BM+VM workloads (UDN localnet + Openstack). Would it make sense to create a templating system where the service provider would define pre-defined network setups, and have only one “Network” CR? + +- Our existing template concept should continue to be valuable for defining different kinds of VMs or even virtualized applications delivered as VMs. + +- That seem like an advanced use case to tackle later. It seems reasonable to start with an assumption that cloud providers will divide their compute hosts into logical virtualization clusters based largely on physical location and network topology. If they have 900 virt host, maybe they divide into 3 clusters of 300 nodes each. I wouldn't worry about "rebalancing" among different virt clusters at this point. + + Conversation from google doc comment: + + > * yes, makes sense + > + > * Are we thinking VMs will be served by the HCP hub worker nodes or the HCP spoke clusters? + > + > * A provider-run cluster will host VMs. For simplicity, we can start by making that the same cluster that the rest of the management tooling runs on. But there could be good reason to separate the virt-hosting cluster later. + > + > * Will networking become complex if we have a tenant cluster and a tenant VM that want to be on the same L2?  Technically doesn't the VM need to run on the tenant cluster for this? + > + > * No. There is no requirement that a VM needs to be located on the tenant cluster in order to access a particular L2 network. L2 networks are managed external to the clusters. + > + > * A UDN can be connected to a VLAN on the physical network. I assume that'll be a primary mechanism for establishing a "VPC" experience. A tenant's VMs will be connected to the UDN, whether those are individual VMs or openshift nodes in VMs. If the tenant also has a bare metal openshift cluster, then we'd depend on the physical network fabric to put those nodes onto a VLAN, which can be bridged to the UDN. + > + > * Basic UDN+virt info: https://www.redhat.com/en/blog/user-defined-networks-red-hat-openshift-virtualization + > + > * Lots of detail about UDNs: https://docs.redhat.com/en/documentation/openshift\_container\_platform/4.19/html/multiple\_networks/primary-networks#about-user-defined-networks + +- We need to start fitting this into a broader concept of a VPC. But generally, we should do a quick survey of what the options are for connecting a VM to other networks besides the pod network. + + > - We talked about the concept of network in the case of the baremetal fulfillment, so I was thinking that we could extend it to mkae it work for the VM case as well. + + > - Yes, though I think that comes with downsides IIRC, for example there won't be health checks on a VM if it's not on the pod network. Not a deal-breaker, but that's the sort of thing we need to understand. + + > - Also we could potentially simplify the network isolation story even if we're using real VLANs on the physical network. If we put the host OS in charge of connecting a VM to a particular VLAN, and give each host access to all the VLANs, that means we don't have to worry about about dynamically managing the network fabric like we do in the BM case. + + > - @mhrivnak@redhat.com Is the definition of VPC you're referring to the AWS definition? + + > - That's the inspiration, but we need to have our own definition. Private network, isolated from other tenants, controls around ingress and egress, perhaps continuity across management clusters, etc. Lots of things to consider. + + > - Sure.  Fundamentally VPC only allows layer 3 and they do some funky stuff that is hidden in layer 2.  For v1.0, I would say we stick to what is supported by OCP Virtualization. + +- I think I need more context to help me understand this question. + + > It's what we discussed in our last meeting.  + + > In case of UDN L2, VM and network definition need to be on the same HUB cluster. Knowing that network must be defined before the VM, what should condition the selection of the HUB cluster. The network or the VM? + + > Thinking about it, it's maybe not an interesting question, because a tenant may want to spawn VMs in that same network latter, and then face a issue getting the resources they want. + + > For localnet, we should not have this constraint, as the network is defined globally, and not local to a HUB cluster. + +- openstack? Is this just referencing that ESI uses some openstack parts? + +- I think we need some formality around defining networks, similar to a VPC construct found in many clouds. We should have some more focused design discussion on that topic. + +- Is there another option supported by openshift virt? + +- Could you add a bit of context for each of these explaining why it might cause a problem that we should be aware of? + +## Test Plan + +TBD + +## Graduation Criteria + +TBD + +### Removing a deprecated feature + +N/A + +## Upgrade / Downgrade Strategy + +N/A + +## Version Skew Strategy + +N/A + +## Support Procedures + +TBD + +## Infrastructure Needed [optional] + +N/A From c4d2138f61db0d0951a79894d92c08a1c4072905 Mon Sep 17 00:00:00 2001 From: Adrien Gentil Date: Fri, 26 Sep 2025 18:22:06 +0200 Subject: [PATCH 2/7] Re-vamp and reduce scope of this proposal --- enhancements/vmaas/README.md | 583 ++++++++++++----------------------- 1 file changed, 204 insertions(+), 379 deletions(-) diff --git a/enhancements/vmaas/README.md b/enhancements/vmaas/README.md index 555c583ff..3c9e50bea 100644 --- a/enhancements/vmaas/README.md +++ b/enhancements/vmaas/README.md @@ -11,357 +11,246 @@ replaces: superseded-by: --- -# Virtualization as a service +# Virtualization-as-a-Service ## Summary -We want to provide a service that allows a tenant to create and manage VMs. +This document proposes a service enabling tenants to easily create, manage, and operate virtual machines (VMs) within a self-service environment. The service will provide user-friendly APIs for provisioning, customizing, and controlling the lifecycle of VMs attaching storage, and accessing specialized hardware (GPUs). + +This design is focused at delivering VM to tenants with an simple networking model, advanced features related to Virtual Data Center-as-a-Service (VDCaaS) will be part of another design. ## Motivation -TBD +Virtualization-as-a-Service (VMaaS) addresses the need for flexible, on-demand compute resources within a multi-tenant environments. This is a need identified in the scope of O-SAC, and will benefit to the MOC. ### User Stories -#### As a tenant - -- I want to list a pre-defined VM templates -- I want to create a VM based on a pre-defined template -- I want a VM that has access to specialized hardware (e.g.: GPU) -- I want to customize the pre-defined template for my usage -- I want to attach and detach block storage on my VM that transcends the lifecycle of the VM -- I want to manage the lifecycle of my VM (start/stop/terminate) -- I want to import my OS base image -- I want to list available base OS images -- I want to manage (create, list, delete) my private network -- I want to attach my VMs in a virtual private network -- I want my VM to be exposed outside of my private network -- I want my VMs to be attached to my own physical L2 network -- I want to connect on my VM through serial console -- I want to keep running my VM when a node goes out of service - -#### As a service provider: - -- I want to define VM templates -- I want to manually trigger VM migration +- As a provider, I want to define VM templates that my tenants will be able to use +- As a tenant, I want to list a pre-defined VM templates +- As a tenant, I want to create a VM based on a pre-defined template +- As a tenant, I want a VM that has access to specialized hardware (e.g.: GPU) +- As a tenant, I want to manage the lifecycle of my VM (start/stop/terminate) +- As a tenant, I want to connect on my VM through serial console +- As a tenant, I want to be able to expose network services to an external network ### Goals -TBD +- Provide a self-service API for tenants to create, manage, and operate virtual machines (VMs) with minimal operational overhead +- Support a catalog of pre-defined VM templates +- Offer access to specialized hardware (e.g., GPUs) for tenants that require it trhough the catalog of tamplates +- Use ESI to assign floating IPs to created VM so they can be accessed +- Achieve high availability by supporting live migration of VMs ### Non-Goals -TBD - -## Proposal - -### Network isolation - -This design relies on UDN (User Defined Networking, it is the networking component provided with Openshift, it allows the definition of virtual private subnets, and provides the ability to connect on physical networks (localnet). - -Moverover, Kubevirt supports UDN layer 2 mode which offers the ability to live-migrate VMs across OCP nodes. - -### Ingress - -We can expose services running on VM(s) using a “LoadBalancer” service (e.g.: relying on MetalLB). - -### APIs - -Virtual private subnet creation: - -``` -apiVersion: cloudkit.openshift.io/v1alpha1 -kind: VirtualPrivateSubnet -metadata: - name: example - finalizers: - - cloudkit.openshift.io/finalizer -annotations: - cloudkit.openshift.io/ref-count: 1 # +1 when a VM is created, -1 when a VM is deleted   -spec: - ipv4Cidr: 192.168.100.0/24 -``` - -We will create a new, unique namespace, and the UDN definition to create a L2 layer and the corresponding subnet. - -Since several VMs can be attached to the same network, we introduce an annotation that will be used to count the number of VMs attached to this network, and a finalizer to prevent its deletion before all the VMs are deleted. - -Load balancer: - -``` -apiVersion: cloudkit.openshift.io/v1alpha1 -kind: LoadBalancer -metadata: - name: example -spec: - hostRefs: - - kind: VirtualMachineHost - name: exampleVM - portMapping: - - protocol: TCP - externalPort: 80 - internalPort: 80 -``` - -All hosts must be associated with the same network. - -### Template management +The following are explicitly out of scope for this proposal: -Like cluster fulfillment, there will be a mechanism to define templates as Ansible roles. Templates contain predefined VM classes which define default OS base image, CPU architecture, number of cores, amount of memory, storage, user data, … +- Implementing advanced VM orchestration features such as auto-scaling, and region placement +- Automatically add physical resources to cope with the demand of VMs +- Offering built-in backup, restore, or disaster recovery solutions for VMs or attached storage +- Delivering a marketplace for third-party VM images or applications, we expect tenant to rely on en external image registry to distribute they OS base images -A template VM might define the following properties, and let a tenant to customize some of them (should be a 1:1 mapping to KubeVirt resources): - -- Name -- Tenant -- OS base image -- Instance class, which specify the family (translating into general/compute/memory/gpu/arch), and the amount of CPU and memory (TBD, e.g.: general.large) -- User-data (cloud-init) -- Networking, namespace + UDN at first, and localnet later? - * Layer 3 only for Networking. -- SSH keys -- Storage, size of the boot volume (+ I/O performance, encryption) -- State (started/stopped) - -There will be an API to publish these templates in the fulfillment service, and list these templates from it. - -KubeVirt UI provides a functionality based on OCP templates that generates a VirtualMachine resource, it is safe to assume that we can use any tool we want to, and rely on Ansible/AAP to define our own. +## Proposal -### VM management +The implementation of virtual machine fulfillment relies on two concepts: -VMs will be managed using KuberVirt, basically the user will have to provide 3 informations to create a VM: +* VirtualMachine: allows the management of a virtual machine by a tenant +* VirtualMachineTemplate: created by the provider, a template represent a pre-configuration for virtual machines that is made available to tenants. This pre-configuration is exposed to tenant though a template ID, and a set of parameters, parameters may be required or optional. -- Template name -- Template overrides -- A network +At a high level, tenants will request a VirtualMachine by specifying: -### OS base image management +* the ID of the template they want to use +* input parameters for the selected template +* the state of the virtual machine (started, stopped) -Tenants will have the ability to import their own base images. Since OSC backend relies on multiple OCP clusters, and tenants’ VMs may be distributed on multiple ones, base images be centralized on an OCI container registry before being consumed by KubeVirt on the destination cluster. +We expect virtual machine fulfillment to follow the same workflow used for other fulfillment workflows; as such, we expect to update the +following existing O-SAC components: -### APIs - -``` -apiVersion: cloudkit.openshift.io/v1alpha1 -kind: VirtualMachineHost -metadata: - name: example - finalizers: - - cloudkit.openshift.io/finalizer # to update ref-count in VirtualPrivateNetwork -spec - networkRef: - kind: VirtualPrivateNetwork - name: exampleNetwork - template: rhel-10 - templateParameters: - instanceType: o1.small - SSHPublicKey: ... - powerState: stopped -``` +* Fulfillment Service: Define the API for VirtualMachine and VirtualMachineTemplate +* Fulfillment CLI: Give the tenant access to the API +* O-SAC Operator: Manage and reconcile the Custom Resources for VirtualMachine +* O-SAC AAP: Use Ansible playbooks to perform the requested reconciliation operations of VirtualMachine by calling KubeVirt and ESI APIs (to assign floating IPs). ### Workflow Description -TBD +#### Virtual machine creation and update + +1. The tenant uses the Fulfillment CLI to request the creation of a new VirtualMachine, specifying: + - The desired VirtualMachineTemplate ID + - Any required or optional parameters for the template (e.g., CPU, memory, disk size, network configuration) + - The initial state of the VM (started or stopped) +2. The Fulfillment Service receives the request and validates: + - The existence and availability of the specified template + - The correctness and completeness of the provided parameters +3. The Fulfillment Service creates a new VirtualMachine custom resource (CR) in the appropriate namespace. +4. The O-SAC Operator detects the new VirtualMachine CR and triggers the reconciliation process. +5. The Operator, via AAP (Ansible Automation Platform), performs the following automation steps: + - Creates a dedicated namespace for the VM (if not already present) + - Provisions the required network resources (e.g., UDN L2 network) + - Creates the KubeVirt VirtualMachine resource according to the template and parameters + - Assigns a floating IP to the VM using ESI APIs + - ...other operations depending on the selected virtual machine template +6. The Operator monitors the status of the VM and updates the VirtualMachine CR status accordingly. +7. The tenant can query the status of the VM via the Fulfillment CLI or API, and access the VM using the assigned floating IP. + +The update process is the same as the creation workflow as it will be designed to be idempotent. + +#### Virtual machine deletion + +When a tenant requests the deletion of a VirtualMachine, the following workflow is executed: + +1. The tenant uses the Fulfillment CLI or API to request deletion of a VirtualMachine by specifying its identifier. +2. The Fulfillment Service receives the deletion request and validates: + - The existence of the specified VirtualMachine resource. + - That the tenant has permission to delete the resource. +3. The Fulfillment Service deletes the VirtualMachine custom resource (CR) from the appropriate namespace. +4. The O-SAC Operator detects the deletion of the VirtualMachine CR and triggers the cleanup process. +5. The Operator, via AAP (Ansible Automation Platform), performs the following automation steps: + - Deletes the KubeVirt VirtualMachine resource. + - Releases and deallocates any associated network resources (e.g., UDN L2 network, floating IPs via ESI APIs). + - ...other cleanup operations depending on the selected virtual machine template + - Deletes the dedicated namespace +6. The Operator updates the status of the deletion operation and ensures all resources are properly cleaned up. +7. The tenant can confirm the deletion via the Fulfillment CLI or API. + +This workflow ensures that all resources associated with the VirtualMachine are properly deprovisioned and that no orphaned resources remain. + +#### Virtual machine template management + +Virtual machine templates are managed by the provider using a GitOps workflow. The source of truth for templates resides in a version-controlled repository. A periodic job running in Ansible Automation Platform (AAP) is responsible for publishing the current set of templates to the Fulfillment Service. This ensures that any updates, additions, or removals of templates in the repository are automatically reflected in the Fulfillment Service, providing tenants with an up-to-date catalog of available VM templates. ### API Extensions -TBD - -### Implementation Details/Notes/Constraints - -#### VM creation minimal flow - -1. Tenant creates a VM Host, the automation: - 1. Creates a namespace - 2. Creates a UDN L2 network - 3. Creates a Kubevirt CR - -#### Tasks - -- Fulfilment-cli - * CRUD on VirtualMachineHost - * List VM templates -- Fulfillment-service: - * [Create VirtualMachineHost management API in fulfillment-service](https://github.com/innabox/issues/issues/201&sa=D&source=editors&ust=1757970010188331&usg=AOvVaw07B9NvURxHW07nF4X-1X41) - * [Publish VM templates endpoint in fulfillment-service](https://github.com/innabox/issues/issues/203&sa=D&source=editors&ust=1757970010188598&usg=AOvVaw35ilTLz7wwMg4_W6YdNeMB) - * [List VM templates](https://github.com/innabox/issues/issues/202&sa=D&source=editors&ust=1757970010188846&usg=AOvVaw2vg_K25W8xerGfSgYxrI4z) -- Cloudkit-operator: - * [Create VirtualMachineHost controller](https://github.com/innabox/issues/issues/204&sa=D&source=editors&ust=1757970010189157&usg=AOvVaw3brz_MVBsWOSWzjjv6Mv9P) -* AAP: - - [Create/Delete VM EDA and playbooks](https://github.com/innabox/issues/issues/207&sa=D&source=editors&ust=1757970010189421&usg=AOvVaw3xVLsU3dXhvFOKs1ZxGPzf - - [Create VM templates](https://github.com/innabox/issues/issues/205&sa=D&source=editors&ust=1757970010189609&usg=AOvVaw0buvglSiCKKLOSh-VBrxJH) - - [Publish VM templates to fulfillment-service](https://github.com/innabox/issues/issues/206&sa=D&source=editors&ust=1757970010189850&usg=AOvVaw1JCATg3Xsg0QA93zUyQUZf) - -Development: - -- Public API to manage networks + operator + AAP (?) -- Public API to manage VMs + operator + AAP -- Public API to publish template -- Create VM template in AAP + publication -- OS base image management + operator - -#### UDN layer 2 definition +#### VirtualMachine + +A tenant requests a virtual machine by requesting a VirtualMachine to the Fulfillment Service. Here is an example of request that creates a VM, using a template that let tenants to customize the amount of CPUs, memory and boot disk size: + +```json +{ + "object": { + "id": "myvm", + "spec": { + "state": "started", + "template": "ocp_virt_vm", + "template_parameters": { + "vm_cpu_cores": { + "value": 4 + }, + "vm_disk_size": { + "value": "30Gi" + }, + "vm_memory": { + "value": "4Gi" + } + } + } + } +} +``` +Once the virtual machine is created, tenants are able to review its current state: + +```json +{ + "@type": "type.googleapis.com/fulfillment.v1.VirtualMachine", + "id": "fecb9b9e-07ac-4d56-8b48-9d50aab71677", + "metadata": { + "creation_timestamp": "2025-09-17T08:14:17.569076Z", + "creators": [ + "guest" + ] + }, + "spec": { + "template": "ocp_virt_vm", + "template_parameters": { + "vm_cpu_cores": { + "value": 4 + }, + "vm_disk_size": { + "value": "30Gi" + }, + "vm_memory": { + "value": "4Gi" + } + } + }, + "status": { + "conditions": [ + { + "last_transition_time": "2025-09-19T17:32:24.054439350Z", + "message": "", + "status": "CONDITION_STATUS_FALSE", + "type": "VIRTUAL_MACHINE_CONDITION_TYPE_PROGRESSING" + }, + { + "last_transition_time": "2025-09-17T08:52:12.652582382Z", + "message": "", + "status": "CONDITION_STATUS_TRUE", + "type": "VIRTUAL_MACHINE_CONDITION_TYPE_READY" + }, + { + "status": "CONDITION_STATUS_FALSE", + "type": "VIRTUAL_MACHINE_CONDITION_TYPE_FAILED" + }, + { + "status": "CONDITION_STATUS_FALSE", + "type": "VIRTUAL_MACHINE_CONDITION_TYPE_DEGRADED" + } + ], + "state": "VIRTUAL_MACHINE_STATE_READY", + "internalIP": "10.0.0.1", + "externalIP": "193.1.2.3" + } +} ``` -apiVersion: v1 -kind: Namespace -metadata: - name: udn-dev - labels: - k8s.ovn.org/primary-user-defined-network: "" ---- -apiVersion: k8s.ovn.org/v1 -kind: UserDefinedNetwork -metadata: - name: udn-dev - namespace: udn-dev -spec: - layer2: - ipam: - lifecycle: Persistent - role: Primary - subnets: - - 10.200.0.0/16 - topology: Layer2 + +`internalIP` is the IP assigned directly to the virtual machine, and `externalIP` is the floating IP pointing to the virtual machine. + +#### VirtualMachineTemplate + +Virtual machine template are designed as Ansible roles, these roles must follow the follwing definition: + + + +There will be a peridioc job witch will publish the Ansible roles as virtual machine templates to the Fulfillment Service, so tenants will be able to reference them to create their virtual machines: + +```json +{ + "object": { + "id": "simple_vm", + "title": "Simple VM Template", + "description": "This template provisions a virtual machine. CPU, memory, and boot volume size is configurable", + "parameters": [ + { + "name": "vm_cpu_cores", + "default": { + "value": 2 + } + }, + { + "name": "vm_memory", + "default": { + "value": "2Gi" + } + }, + { + "name": "vm_disk_size", + "default": { + "value": "20Gi" + } + } + ] + } +} ``` -#### UDN localnet definition +### Implementation Details/Notes/Constraints -Requires NMState operator: -``` -apiVersion: nmstate.io/v1 -kind: NodeNetworkConfigurationPolicy -metadata: - name: mapping -spec: - nodeSelector: - node-role.kubernetes.io/worker: '' - desiredState: - ovn: - bridge-mappings: - - localnet: localnet1 - bridge: br-ex - state: present ---- -apiVersion: k8s.ovn.org/v1 -kind: ClusterUserDefinedNetwork -metadata: - name: cudn-localnet -spec: - namespaceSelector: - matchExpressions: - - key: kubernetes.io/metadata.name - operator: In - values: ["red", "blue"] - network: - topology: Localnet - localnet: - role: Secondary - physicalNetworkName: localnet1 - ipam: - mode: Disabled -``` +#### Virtual machines on HUB cluster -#### Live Migrations using UDN and KubeVirt - -- Kubevirt only allows/supports UDN L2 -- Localnet is not ideal due to the static nature of this approach and the VLAN is not necessarily the same on the destination node. (IE VxLAN endpoints) -- Complex hardware dependencies are hit and miss.  It all depends on the hardware, but there is the ability to migrate things like GPUs and NICs (SR-IOV) in specific circumstances. -- Here is an instance where we can use templates to avoid the complexity. -- Things to be aware of which could cause problems: - * Localnet networks which tie to specific h/w i/f., - * SR-IOV, - * DPDK, - * Node selecting or filtering logic, - * Possible DHCP issues. - -- Live migrations describe the type of migration and effected changes. - - * L2 UDN: Updates MAC learning tables, & ARP caches. - * L3 UDN: Updates node IP routing tables pointing to the new hosting node. - * L3 migrations should not be interpreted as allowing an IP address to change on the fly. - -#### UDN Topology Types - -UDN has topology types (L2/L3) with different block types highlighting the differences between one or the other. - -Layer2 (layer2: block) - -- Bridged networking - VMs appear on same broadcast domain -- DHCP/Static IP assignment - Needs IPAM (IP Address Management) for IP allocation -- East-west traffic flows directly between VMs without routing -- Flat network model - All VMs share the same subnet - -Layer3 (layer3: block): - -- Routed networking - Each node gets its own subnet slice -- Distributed subnets - hostSubnet defines how the CIDR is carved up per node -- Inter-node routing required for VM-to-VM communication -- Scalable addressing - Prevents IP exhaustion on large clusters - - Configuration Differences: - -- Layer2 needs IPAM for IP management - - ``` - layer2: - ipam: - lifecycle: Persistent # How long IPs are reserved - subnets: [10.200.0.0/16] # Flat subnet - ``` - -- Layer3 needs subnet distribution - - ``` - layer3: - subnets: - - cidr: 10.200.0.0/16 # Overall range - hostSubnet: 24 # Each node gets /24 slice - ``` - -#### Tasks - -- List options in kubevirt that prevents Live VM migration -- Try UDN + kubevirt - - Describe how Egress IP works - - Default Egress - - Why do we need an Egress IP - - Look at the manifests, will help to design API -- Describe Template and MVP - - VM Host / HostPool - - Disk - - Memory - - Storage - - Persistent vs Ephemeral - - SSH keys - - Cloud-init - - Network - - OS base image (Registry Service) - - Test: - - Create network - - Create VM in the network - - Communicate between Nodes (2 VMs) - - Floating IP / LoadBalancer -- Create a sequence diagram, showing service/operator/AAP responsibilities - - VM management - - Base image management - - Network management -- Get information about Metal LB and UDN’s floating IP is supposed to work in order to grab a public IP (and other related networking stuff) - -#### External resources - -- [Red Hat Sovereignty Technical Architecture](https://docs.google.com/presentation/d/1g23omSQ46NZ6qyDi-I4Q8mGNt-NQpKk0vOuaZf8Usqw/edit?slide%3Did.g344ca2ed691_0_2780%23slide%3Did.g344ca2ed691_0_2780&sa=D&source=editors&ust=1757970010209722&usg=AOvVaw0KkWOZaTJyX2nac50mzfW) -- [RHSovCloud-Network-Scenarios-0.1](https://docs.google.com/presentation/d/1_wYAbfoCcmuIOx-esCcOoSFJFR7tgkrypyyIHa7S4AY/edit?slide%3Did.g30e84f14e68_0_6149%23slide%3Did.g30e84f14e68_0_6149&sa=D&source=editors&ust=1757970010209955&usg=AOvVaw3oLbSpaYriqUYAxZ1TQlWX) -- [Red Hat Sovereignty Technical Architecture](https://docs.google.com/presentation/d/1g23omSQ46NZ6qyDi-I4Q8mGNt-NQpKk0vOuaZf8Usqw/edit?slide%3Did.g344ca2ed691_0_2780%23slide%3Did.g344ca2ed691_0_2780&sa=D&source=editors&ust=1757970010210162&usg=AOvVaw1NbKJtheHnk47tN4Ot5KVy) - -UDN:   - -- [User Defined Networks on OpenShift](https://docs.google.com/presentation/d/1Hx1Fzm1F9EkmqrmTjbMHPBAuIls-2-oK1IVrzFnW3L4/edit?slide%3Did.g31966f3f64c_0_855%23slide%3Did.g31966f3f64c_0_855&sa=D&source=editors&ust=1757970010210475&usg=AOvVaw3lrbnP41fTOVcfwnbXAuwW) -- -- [Chapter 2. Primary networks | Multiple networks | OpenShift Container Platform | 4.19 | Red Hat Documentation](https://docs.redhat.com/en/documentation/openshift_container_platform/4.19/html/multiple_networks/primary-networks%23about-user-defined-networks&sa=D&source=editors&ust=1757970010211384&usg=AOvVaw22g82g1WsGtCpIV_OOnpnt) -- [Chapter 3. Secondary networks | Multiple networks | OpenShift Container Platform | 4.19 | Red Hat Documentation](https://docs.redhat.com/en/documentation/openshift_container_platform/4.19/html/multiple_networks/secondary-networks%23configuration-localnet-switched-topology_configuring-additional-network-ovnk&sa=D&source=editors&ust=1757970010211903&usg=AOvVaw0q1VmdCCcREjz1O3xKPmt) -- [User defined networks in Red Hat OpenShift Virtualization](https://www.redhat.com/en/blog/user-defined-networks-red-hat-openshift-virtualization&sa=D&source=editors&ust=1757970010212242&usg=AOvVaw2KBAek4qp-jL8eF_I9gN5) +Virtual machines will be created on the HUB cluster that was selected by Fulfillment Service, it was discussed to create a dedicated cluster to handle VM workloads using Cluster-as-a-Service API, but since they are HostedCluster it won't increase the reliability of the solution, as their reliability are tied to the same HUB cluster. ### Risks and Mitigations @@ -377,71 +266,7 @@ TBD ## Open Questions [optional] -- The main point is to provide a public API that will select a management cluster to provision a VM (given requested resources CPU/Mem/GPU)? => not the priority -- What is the added value on top of kube virt? => multi-tenancy, multiple virt clusters, higher level networking primitives -- Since we plan to rely on the OCP stack (ACM/KubeVirt/Hypershift), is there something we want to make pluggable using Ansible? (access network configuration to the VM?) -- Do we need to provision infra on-demand to run kubevirt workload? -- What about networking isolation, how kubevirt works? What model to prioritize? -- Representation of resources a la AWS? `.` (e.g.: t3.xlarge (general purpose), c3.small (compute), m3.medium (memory)) -- Storage options? -- Should we focus on cluster provisioning or VM provisioning? -- Which networking model should we prioritize? The ability to connect on a physical network or the ability to create a virtual private network? -- How to make sure we co-collocate network and VM definitions together? -- From what I understand, there will be different ways to handle networking depending if we want VM workloads (UDN), BM workloads (Openstack), or BM+VM workloads (UDN localnet + Openstack). Would it make sense to create a templating system where the service provider would define pre-defined network setups, and have only one “Network” CR? - -- Our existing template concept should continue to be valuable for defining different kinds of VMs or even virtualized applications delivered as VMs. - -- That seem like an advanced use case to tackle later. It seems reasonable to start with an assumption that cloud providers will divide their compute hosts into logical virtualization clusters based largely on physical location and network topology. If they have 900 virt host, maybe they divide into 3 clusters of 300 nodes each. I wouldn't worry about "rebalancing" among different virt clusters at this point. - - Conversation from google doc comment: - - > * yes, makes sense - > - > * Are we thinking VMs will be served by the HCP hub worker nodes or the HCP spoke clusters? - > - > * A provider-run cluster will host VMs. For simplicity, we can start by making that the same cluster that the rest of the management tooling runs on. But there could be good reason to separate the virt-hosting cluster later. - > - > * Will networking become complex if we have a tenant cluster and a tenant VM that want to be on the same L2?  Technically doesn't the VM need to run on the tenant cluster for this? - > - > * No. There is no requirement that a VM needs to be located on the tenant cluster in order to access a particular L2 network. L2 networks are managed external to the clusters. - > - > * A UDN can be connected to a VLAN on the physical network. I assume that'll be a primary mechanism for establishing a "VPC" experience. A tenant's VMs will be connected to the UDN, whether those are individual VMs or openshift nodes in VMs. If the tenant also has a bare metal openshift cluster, then we'd depend on the physical network fabric to put those nodes onto a VLAN, which can be bridged to the UDN. - > - > * Basic UDN+virt info: https://www.redhat.com/en/blog/user-defined-networks-red-hat-openshift-virtualization - > - > * Lots of detail about UDNs: https://docs.redhat.com/en/documentation/openshift\_container\_platform/4.19/html/multiple\_networks/primary-networks#about-user-defined-networks - -- We need to start fitting this into a broader concept of a VPC. But generally, we should do a quick survey of what the options are for connecting a VM to other networks besides the pod network. - - > - We talked about the concept of network in the case of the baremetal fulfillment, so I was thinking that we could extend it to mkae it work for the VM case as well. - - > - Yes, though I think that comes with downsides IIRC, for example there won't be health checks on a VM if it's not on the pod network. Not a deal-breaker, but that's the sort of thing we need to understand. - - > - Also we could potentially simplify the network isolation story even if we're using real VLANs on the physical network. If we put the host OS in charge of connecting a VM to a particular VLAN, and give each host access to all the VLANs, that means we don't have to worry about about dynamically managing the network fabric like we do in the BM case. - - > - @mhrivnak@redhat.com Is the definition of VPC you're referring to the AWS definition? - - > - That's the inspiration, but we need to have our own definition. Private network, isolated from other tenants, controls around ingress and egress, perhaps continuity across management clusters, etc. Lots of things to consider. - - > - Sure.  Fundamentally VPC only allows layer 3 and they do some funky stuff that is hidden in layer 2.  For v1.0, I would say we stick to what is supported by OCP Virtualization. - -- I think I need more context to help me understand this question. - - > It's what we discussed in our last meeting.  - - > In case of UDN L2, VM and network definition need to be on the same HUB cluster. Knowing that network must be defined before the VM, what should condition the selection of the HUB cluster. The network or the VM? - - > Thinking about it, it's maybe not an interesting question, because a tenant may want to spawn VMs in that same network latter, and then face a issue getting the resources they want. - - > For localnet, we should not have this constraint, as the network is defined globally, and not local to a HUB cluster. - -- openstack? Is this just referencing that ESI uses some openstack parts? - -- I think we need some formality around defining networks, similar to a VPC construct found in many clouds. We should have some more focused design discussion on that topic. - -- Is there another option supported by openshift virt? - -- Could you add a bit of context for each of these explaining why it might cause a problem that we should be aware of? +- I think that without the concept of regions (mapped on HUB cluster?) in the Fulfillment Service, we are stuck with storage (as it is local to a HUB), and with private networking (for example, I want my web server with a floating IP to communicate internally with my DB). ## Test Plan From 10c605ab01cd8c07195eb1351651bd95aad0522e Mon Sep 17 00:00:00 2001 From: Adrien Gentil Date: Wed, 1 Oct 2025 09:21:21 +0200 Subject: [PATCH 3/7] expanded on workflows and VM statuses --- enhancements/vmaas/README.md | 198 +++++++++++++++++++++++++---------- 1 file changed, 142 insertions(+), 56 deletions(-) diff --git a/enhancements/vmaas/README.md b/enhancements/vmaas/README.md index 3c9e50bea..3aae816bf 100644 --- a/enhancements/vmaas/README.md +++ b/enhancements/vmaas/README.md @@ -18,7 +18,7 @@ superseded-by: This document proposes a service enabling tenants to easily create, manage, and operate virtual machines (VMs) within a self-service environment. The service will provide user-friendly APIs for provisioning, customizing, and controlling the lifecycle of VMs attaching storage, and accessing specialized hardware (GPUs). -This design is focused at delivering VM to tenants with an simple networking model, advanced features related to Virtual Data Center-as-a-Service (VDCaaS) will be part of another design. +This design aims to provide tenants with virtual machines using a straightforward networking model. More advanced capabilities, such as those related to Virtual Data Center-as-a-Service (VDCaaS), will be addressed in a separate proposal. ## Motivation @@ -32,7 +32,7 @@ Virtualization-as-a-Service (VMaaS) addresses the need for flexible, on-demand c - As a tenant, I want a VM that has access to specialized hardware (e.g.: GPU) - As a tenant, I want to manage the lifecycle of my VM (start/stop/terminate) - As a tenant, I want to connect on my VM through serial console -- As a tenant, I want to be able to expose network services to an external network +- As a tenant, I want to be able to expose network services to an external network ### Goals @@ -47,52 +47,61 @@ Virtualization-as-a-Service (VMaaS) addresses the need for flexible, on-demand c The following are explicitly out of scope for this proposal: - Implementing advanced VM orchestration features such as auto-scaling, and region placement +- Implementing Virtual Data Center-as-a-Service (VDCaaS) features such as virtual private networks - Automatically add physical resources to cope with the demand of VMs - Offering built-in backup, restore, or disaster recovery solutions for VMs or attached storage - Delivering a marketplace for third-party VM images or applications, we expect tenant to rely on en external image registry to distribute they OS base images +- Providing functionality for block storage management ## Proposal -The implementation of virtual machine fulfillment relies on two concepts: +The process of fulfilling virtual machine requests is based on two primary concepts: -* VirtualMachine: allows the management of a virtual machine by a tenant -* VirtualMachineTemplate: created by the provider, a template represent a pre-configuration for virtual machines that is made available to tenants. This pre-configuration is exposed to tenant though a template ID, and a set of parameters, parameters may be required or optional. +* **VirtualMachine**: Represents an individual virtual machine that a tenant can create and manage. +* **VirtualMachineTemplate**: Defined by the provider, this is a pre-configured blueprint for virtual machines. Each template is identified by a unique template ID and includes a set of parameters (some required, some optional) that tenants can specify when creating a VM. -At a high level, tenants will request a VirtualMachine by specifying: +To request a new VirtualMachine, tenants must provide: -* the ID of the template they want to use -* input parameters for the selected template -* the state of the virtual machine (started, stopped) +* The ID of the desired VirtualMachineTemplate +* Any required or optional parameters for that template +* The desired initial state of the VM (e.g., started or stopped) -We expect virtual machine fulfillment to follow the same workflow used for other fulfillment workflows; as such, we expect to update the -following existing O-SAC components: +The virtual machine fulfillment process will align with existing O-SAC fulfillment workflows. To support this, the following O-SAC components will be enhanced or updated: -* Fulfillment Service: Define the API for VirtualMachine and VirtualMachineTemplate -* Fulfillment CLI: Give the tenant access to the API -* O-SAC Operator: Manage and reconcile the Custom Resources for VirtualMachine -* O-SAC AAP: Use Ansible playbooks to perform the requested reconciliation operations of VirtualMachine by calling KubeVirt and ESI APIs (to assign floating IPs). +* **Fulfillment Service**: Defines and exposes the APIs for managing VirtualMachine and VirtualMachineTemplate resources. +* **Fulfillment CLI**: Provides tenants with command-line access to the Fulfillment Service APIs. +* **O-SAC Operator**: Monitors and reconciles VirtualMachine custom resources within the system. +* **O-SAC AAP (Ansible Automation Platform)**: Executes automation tasks (via Ansible playbooks) to reconcile VirtualMachine resources, including interactions with KubeVirt for VM lifecycle management and ESI APIs for assigning floating IPs. ### Workflow Description #### Virtual machine creation and update -1. The tenant uses the Fulfillment CLI to request the creation of a new VirtualMachine, specifying: - - The desired VirtualMachineTemplate ID - - Any required or optional parameters for the template (e.g., CPU, memory, disk size, network configuration) - - The initial state of the VM (started or stopped) -2. The Fulfillment Service receives the request and validates: - - The existence and availability of the specified template - - The correctness and completeness of the provided parameters -3. The Fulfillment Service creates a new VirtualMachine custom resource (CR) in the appropriate namespace. -4. The O-SAC Operator detects the new VirtualMachine CR and triggers the reconciliation process. -5. The Operator, via AAP (Ansible Automation Platform), performs the following automation steps: - - Creates a dedicated namespace for the VM (if not already present) - - Provisions the required network resources (e.g., UDN L2 network) - - Creates the KubeVirt VirtualMachine resource according to the template and parameters - - Assigns a floating IP to the VM using ESI APIs - - ...other operations depending on the selected virtual machine template -6. The Operator monitors the status of the VM and updates the VirtualMachine CR status accordingly. -7. The tenant can query the status of the VM via the Fulfillment CLI or API, and access the VM using the assigned floating IP. +1. The tenant initiates the creation of a new VirtualMachine using the Fulfillment CLI. The tenant must provide: + - The ID of the desired VirtualMachineTemplate + - All required and any optional parameters for the template (such as CPU, memory, disk size, network configuration) + - The desired initial state of the VM (e.g., started or stopped) + +2. The Fulfillment Service receives this request and performs validation to ensure: + - The specified template exists and is available + - All required parameters are provided and valid + +3. Upon successful validation, the Fulfillment Service creates a new VirtualMachine custom resource (CR) in the appropriate Hub and namespace. + +4. The O-SAC Operator detects the new VirtualMachine CR and begins the reconciliation process. + +5. The Operator, using Ansible Automation Platform (AAP), automates the following steps: + - Creates a dedicated namespace for the VM if one does not already exist + - Provisions necessary network resources, including: + - A UDN L2 network to provide network isolation + - A load balancer service with the VM as its backend + - Assignment of a floating IP to the load balancer service for external access + - Creates the KubeVirt VirtualMachine resource using the specified template and parameters + - Performs any additional operations required by the selected template + +6. The Operator continuously monitors the VM’s status and updates the VirtualMachine CR status to reflect the current state. + +7. The tenant can check the VM’s status at any time using the Fulfillment CLI or API, and can access the VM via the assigned floating IP. The update process is the same as the creation workflow as it will be designed to be idempotent. @@ -100,25 +109,34 @@ The update process is the same as the creation workflow as it will be designed t When a tenant requests the deletion of a VirtualMachine, the following workflow is executed: -1. The tenant uses the Fulfillment CLI or API to request deletion of a VirtualMachine by specifying its identifier. -2. The Fulfillment Service receives the deletion request and validates: - - The existence of the specified VirtualMachine resource. - - That the tenant has permission to delete the resource. -3. The Fulfillment Service deletes the VirtualMachine custom resource (CR) from the appropriate namespace. -4. The O-SAC Operator detects the deletion of the VirtualMachine CR and triggers the cleanup process. -5. The Operator, via AAP (Ansible Automation Platform), performs the following automation steps: +1. The tenant initiates the deletion of a VirtualMachine using the Fulfillment CLI or API by specifying its identifier. + +2. The Fulfillment Service receives the deletion request and performs validation to ensure: + - The specified VirtualMachine resource exists and is available. + - The tenant has permission to delete the resource. + +3. Upon successful validation, the Fulfillment Service deletes the VirtualMachine custom resource (CR) from the appropriate namespace. + +4. The O-SAC Operator detects the deletion of the VirtualMachine CR and begins the cleanup process. + +5. The Operator, using Ansible Automation Platform (AAP), automates the following steps: - Deletes the KubeVirt VirtualMachine resource. - - Releases and deallocates any associated network resources (e.g., UDN L2 network, floating IPs via ESI APIs). - - ...other cleanup operations depending on the selected virtual machine template - - Deletes the dedicated namespace + - Releases and deallocates any associated network resources, including: + - UDN L2 network + - Load balancer service + - Floating IPs via ESI APIs + - Performs any additional cleanup operations required by the selected virtual machine template. + - Deletes the dedicated namespace if it is no longer needed. + 6. The Operator updates the status of the deletion operation and ensures all resources are properly cleaned up. -7. The tenant can confirm the deletion via the Fulfillment CLI or API. + +7. The tenant can confirm the deletion and cleanup via the Fulfillment CLI or API. This workflow ensures that all resources associated with the VirtualMachine are properly deprovisioned and that no orphaned resources remain. #### Virtual machine template management -Virtual machine templates are managed by the provider using a GitOps workflow. The source of truth for templates resides in a version-controlled repository. A periodic job running in Ansible Automation Platform (AAP) is responsible for publishing the current set of templates to the Fulfillment Service. This ensures that any updates, additions, or removals of templates in the repository are automatically reflected in the Fulfillment Service, providing tenants with an up-to-date catalog of available VM templates. +Virtual machine templates are centrally managed by the provider using a GitOps approach. All templates are stored in a version-controlled repository, which acts as the single source of truth. At regular intervals, an automated job in Ansible Automation Platform (AAP) synchronizes the latest templates from this repository to the Fulfillment Service. As a result, any changes to the templates—such as updates, additions, or deletions—are automatically and consistently reflected in the Fulfillment Service. This process ensures that tenants always have access to the most current catalog of available VM templates. ### API Extensions @@ -149,7 +167,7 @@ A tenant requests a virtual machine by requesting a VirtualMachine to the Fulfil } ``` -Once the virtual machine is created, tenants are able to review its current state: +After creating a virtual machine, tenants can check its current status and details as follows: ```json { @@ -162,6 +180,7 @@ Once the virtual machine is created, tenants are able to review its current stat ] }, "spec": { + "state": "started", "template": "ocp_virt_vm", "template_parameters": { "vm_cpu_cores": { @@ -181,13 +200,49 @@ Once the virtual machine is created, tenants are able to review its current stat "last_transition_time": "2025-09-19T17:32:24.054439350Z", "message": "", "status": "CONDITION_STATUS_FALSE", - "type": "VIRTUAL_MACHINE_CONDITION_TYPE_PROGRESSING" + "type": "VIRTUAL_MACHINE_CONDITION_TYPE_PROVISIONNING" + }, + { + "last_transition_time": "2025-09-17T08:52:12.652582382Z", + "message": "", + "status": "CONDITION_STATUS_FALSE", + "type": "VIRTUAL_MACHINE_STATE_STARTING" }, { "last_transition_time": "2025-09-17T08:52:12.652582382Z", "message": "", "status": "CONDITION_STATUS_TRUE", - "type": "VIRTUAL_MACHINE_CONDITION_TYPE_READY" + "type": "VIRTUAL_MACHINE_STATE_RUNNING" + }, + { + "last_transition_time": "2025-09-17T08:52:12.652582382Z", + "message": "", + "status": "CONDITION_STATUS_FALSE", + "type": "VIRTUAL_MACHINE_STATE_STOPPING" + }, + { + "last_transition_time": "2025-09-17T08:52:12.652582382Z", + "message": "", + "status": "CONDITION_STATUS_FALSE", + "type": "VIRTUAL_MACHINE_STATE_STOPPED" + }, + { + "last_transition_time": "2025-09-17T08:52:12.652582382Z", + "message": "", + "status": "CONDITION_STATUS_FALSE", + "type": "VIRTUAL_MACHINE_STATE_TERMINATING" + }, + { + "last_transition_time": "2025-09-17T08:52:12.652582382Z", + "message": "", + "status": "CONDITION_STATUS_FALSE", + "type": "VIRTUAL_MACHINE_STATE_MIGRATING" + }, + { + "last_transition_time": "2025-09-17T08:52:12.652582382Z", + "message": "", + "status": "CONDITION_STATUS_FALSE", + "type": "VIRTUAL_MACHINE_STATE_PAUSED" }, { "status": "CONDITION_STATUS_FALSE", @@ -198,29 +253,56 @@ Once the virtual machine is created, tenants are able to review its current stat "type": "VIRTUAL_MACHINE_CONDITION_TYPE_DEGRADED" } ], - "state": "VIRTUAL_MACHINE_STATE_READY", + "state": "VIRTUAL_MACHINE_STATE_RUNNING", "internalIP": "10.0.0.1", "externalIP": "193.1.2.3" } } ``` -`internalIP` is the IP assigned directly to the virtual machine, and `externalIP` is the floating IP pointing to the virtual machine. +The status section provides two types of IP addresses for the virtual machine: + +- `internalIP`: The private IP address assigned to the VM on the internal network. + +- `externalIP`: The public (floating) IP address assigned to the VM. This address allows the VM to be accessed from outside the internal network, such as from the internet. + +The Virtual Machine can be in one of the following states, as reflected in the status section above: + +- **Provisioning** (`VIRTUAL_MACHINE_STATE_PROVISIONING`): The VM is being created. +- **Starting** (`VIRTUAL_MACHINE_STATE_STARTING`): The Pod for the Virtual Machine Instance (VMI) is being scheduled and started. +- **Running** (`VIRTUAL_MACHINE_STATE_RUNNING`): The VM is actively running inside its Pod. +- **Stopping** (`VIRTUAL_MACHINE_STATE_STOPPING`): The VM is in the process of shutting down. +- **Stopped** (`VIRTUAL_MACHINE_STATE_STOPPED`): The VM is not running. It exists as a VirtualMachine object, but there is no active VMI or Pod. +- **Terminating** (`VIRTUAL_MACHINE_STATE_TERMINATING`): The VM object is in the process of being deleted. +- **Migrating** (`VIRTUAL_MACHINE_STATE_MIGRATING`): The VM is in the process of being live-migrated to another node. +- **Paused** (`VIRTUAL_MACHINE_STATE_PAUSED`): The VM is in a suspended state. Its process is frozen, but its memory and resources are still allocated. + +These states are reported in the `type` field of the VM's `conditions` array in the status section. + #### VirtualMachineTemplate -Virtual machine template are designed as Ansible roles, these roles must follow the follwing definition: +Virtual machine templates are implemented as Ansible roles. Each role must include the following files: + +* `vm_template_role/meta/argument_specs.yaml`: Defines the [Ansible argument specification](https://docs.ansible.com/ansible/latest/dev_guide/developing_program_flow_modules.html#argument-spec). This file is required. +* `vm_template_role/meta/cloudkit.yaml`: Contains metadata for the template, including the title and description. This file is required. +Example of `cloudkit.yaml`: +```yaml +title: VM Template +description: > + This template provisions a virtual machine. +``` -There will be a peridioc job witch will publish the Ansible roles as virtual machine templates to the Fulfillment Service, so tenants will be able to reference them to create their virtual machines: +A periodic job will publish the Ansible roles as virtual machine templates to the Fulfillment Service, using the required files described above. The following is the API format used for this publication: ```json { "object": { - "id": "simple_vm", - "title": "Simple VM Template", - "description": "This template provisions a virtual machine. CPU, memory, and boot volume size is configurable", + "id": "vm_template_role", + "title": "VM Template", + "description": "This template provisions a virtual machine.", "parameters": [ { "name": "vm_cpu_cores", @@ -250,7 +332,11 @@ There will be a peridioc job witch will publish the Ansible roles as virtual mac #### Virtual machines on HUB cluster -Virtual machines will be created on the HUB cluster that was selected by Fulfillment Service, it was discussed to create a dedicated cluster to handle VM workloads using Cluster-as-a-Service API, but since they are HostedCluster it won't increase the reliability of the solution, as their reliability are tied to the same HUB cluster. +Virtual machines will be created on the HUB cluster chosen by the Fulfillment Service. Although there was discussion about creating a dedicated cluster for VM workloads using the Cluster-as-a-Service API, this approach would not improve reliability. This is because these dedicated clusters would still be implemented as HostedClusters, whose reliability ultimately depends on the same underlying HUB cluster. + +#### Netowrking + +Because O-SAC does not yet provide a VDCaaS (Virtual Data Center as a Service) layer, and the Fulfillment Service cannot guarantee that two virtual machines will be provisioned on the same HUB cluster, each virtual machine must be assigned a floating IP. This ensures that every VM is accessible regardless of where it is deployed. ### Risks and Mitigations @@ -266,7 +352,7 @@ TBD ## Open Questions [optional] -- I think that without the concept of regions (mapped on HUB cluster?) in the Fulfillment Service, we are stuck with storage (as it is local to a HUB), and with private networking (for example, I want my web server with a floating IP to communicate internally with my DB). +- I think that without the concept of regions (mapped on HUB clusters?) in the Fulfillment Service, we can't move much forward with storage (as it is local to a HUB), and with the ability to create VMs without floating IPs (for example, I want my web server with a floating IP to communicate internally with my DB). ## Test Plan From d44dcf62c65913954313f106f38037ba673a247f Mon Sep 17 00:00:00 2001 From: Adrien Gentil Date: Wed, 1 Oct 2025 11:19:34 +0200 Subject: [PATCH 4/7] wrap lines and expand drawback section --- enhancements/vmaas/README.md | 278 ++++++++++++++++++++++++++--------- 1 file changed, 205 insertions(+), 73 deletions(-) diff --git a/enhancements/vmaas/README.md b/enhancements/vmaas/README.md index 3aae816bf..343285218 100644 --- a/enhancements/vmaas/README.md +++ b/enhancements/vmaas/README.md @@ -16,29 +16,42 @@ superseded-by: ## Summary -This document proposes a service enabling tenants to easily create, manage, and operate virtual machines (VMs) within a self-service environment. The service will provide user-friendly APIs for provisioning, customizing, and controlling the lifecycle of VMs attaching storage, and accessing specialized hardware (GPUs). +This document proposes a service enabling tenants to easily create, manage, and +operate virtual machines (VMs) within a self-service environment. The service +will provide user-friendly APIs for provisioning, customizing, and controlling +the lifecycle of VMs attaching storage, and accessing specialized hardware +(GPUs). -This design aims to provide tenants with virtual machines using a straightforward networking model. More advanced capabilities, such as those related to Virtual Data Center-as-a-Service (VDCaaS), will be addressed in a separate proposal. +This design aims to provide tenants with virtual machines using a +straightforward networking model. More advanced capabilities, such as those +related to Virtual Data Center-as-a-Service (VDCaaS), will be addressed in a +separate proposal. ## Motivation -Virtualization-as-a-Service (VMaaS) addresses the need for flexible, on-demand compute resources within a multi-tenant environments. This is a need identified in the scope of O-SAC, and will benefit to the MOC. +Virtualization-as-a-Service (VMaaS) addresses the need for flexible, on-demand +compute resources within a multi-tenant environments. This is a need identified +in the scope of O-SAC, and will benefit to the MOC. ### User Stories -- As a provider, I want to define VM templates that my tenants will be able to use +- As a provider, I want to define VM templates that my tenants will be able to + use - As a tenant, I want to list a pre-defined VM templates - As a tenant, I want to create a VM based on a pre-defined template - As a tenant, I want a VM that has access to specialized hardware (e.g.: GPU) - As a tenant, I want to manage the lifecycle of my VM (start/stop/terminate) - As a tenant, I want to connect on my VM through serial console -- As a tenant, I want to be able to expose network services to an external network +- As a tenant, I want to be able to expose network services to an external + network ### Goals -- Provide a self-service API for tenants to create, manage, and operate virtual machines (VMs) with minimal operational overhead +- Provide a self-service API for tenants to create, manage, and operate virtual + machines (VMs) with minimal operational overhead - Support a catalog of pre-defined VM templates -- Offer access to specialized hardware (e.g., GPUs) for tenants that require it trhough the catalog of tamplates +- Offer access to specialized hardware (e.g., GPUs) for tenants that require it + through the catalog of templates - Use ESI to assign floating IPs to created VM so they can be accessed - Achieve high availability by supporting live migration of VMs @@ -46,19 +59,28 @@ Virtualization-as-a-Service (VMaaS) addresses the need for flexible, on-demand c The following are explicitly out of scope for this proposal: -- Implementing advanced VM orchestration features such as auto-scaling, and region placement -- Implementing Virtual Data Center-as-a-Service (VDCaaS) features such as virtual private networks +- Implementing advanced VM orchestration features such as auto-scaling, and + region placement +- Implementing Virtual Data Center-as-a-Service (VDCaaS) features such as + virtual private networks - Automatically add physical resources to cope with the demand of VMs -- Offering built-in backup, restore, or disaster recovery solutions for VMs or attached storage -- Delivering a marketplace for third-party VM images or applications, we expect tenant to rely on en external image registry to distribute they OS base images -- Providing functionality for block storage management +- Offering built-in backup, restore, or disaster recovery solutions for VMs or + attached storage +- Delivering a marketplace for third-party VM images or applications, we expect + tenant to rely on en external image registry to distribute they OS base images ## Proposal -The process of fulfilling virtual machine requests is based on two primary concepts: +The process of fulfilling virtual machine requests is based on two primary +concepts: -* **VirtualMachine**: Represents an individual virtual machine that a tenant can create and manage. -* **VirtualMachineTemplate**: Defined by the provider, this is a pre-configured blueprint for virtual machines. Each template is identified by a unique template ID and includes a set of parameters (some required, some optional) that tenants can specify when creating a VM. +* **VirtualMachine**: Represents an individual virtual machine that a tenant can + create and manage. Tenants can only see the virtual machines they created. +* **VirtualMachineTemplate**: Defined by the provider, this is a pre-configured + blueprint for virtual machines. Each template is identified by a unique + template ID and includes a set of parameters (some required, some optional) + that tenants can specify when creating a VM. Templates are available to all + tenants to use, they cannot edit them. To request a new VirtualMachine, tenants must provide: @@ -66,83 +88,138 @@ To request a new VirtualMachine, tenants must provide: * Any required or optional parameters for that template * The desired initial state of the VM (e.g., started or stopped) -The virtual machine fulfillment process will align with existing O-SAC fulfillment workflows. To support this, the following O-SAC components will be enhanced or updated: - -* **Fulfillment Service**: Defines and exposes the APIs for managing VirtualMachine and VirtualMachineTemplate resources. -* **Fulfillment CLI**: Provides tenants with command-line access to the Fulfillment Service APIs. -* **O-SAC Operator**: Monitors and reconciles VirtualMachine custom resources within the system. -* **O-SAC AAP (Ansible Automation Platform)**: Executes automation tasks (via Ansible playbooks) to reconcile VirtualMachine resources, including interactions with KubeVirt for VM lifecycle management and ESI APIs for assigning floating IPs. +The virtual machine fulfillment process will align with existing O-SAC +fulfillment workflows. To support this, the following O-SAC components will be +enhanced or updated: + +* **Fulfillment Service**: Defines and exposes the APIs for managing + VirtualMachine and VirtualMachineTemplate resources. +* **Fulfillment CLI**: Provides tenants with command-line access to the + Fulfillment Service APIs. +* **O-SAC Operator**: Monitors and reconciles VirtualMachine custom resources + within the system. +* **O-SAC AAP (Ansible Automation Platform)**: Executes automation tasks (via + Ansible playbooks) to reconcile VirtualMachine resources, including + interactions with KubeVirt for VM lifecycle management and ESI APIs for + assigning floating IPs. + +Under the hood, the virtual machines will be managed using OpenShift +Virtualization. Each VirtualMachine will be created in its own dedicated +namespace on the Hub cluster. The VM will be connected to a dedicated UDN L2 +network, which provides network isolation, mainly for 2 reasons: +- security: this means the VM cannot communicate with other workloads on the Hub + cluster using its internal IP address. +- Live migration: UDN L2 enables the virtual machine to be migrated live from + OpenShift nodes to others. Migration can happen when a node is degraded or in + maintenance. + +To enable external connectivity, a Kubernetes Service of type LoadBalancer will +be created in front of the VM. A floating (external) IP address will be assigned +to this LoadBalancer service, allowing the VM to be accessed from outside the +cluster. ### Workflow Description #### Virtual machine creation and update -1. The tenant initiates the creation of a new VirtualMachine using the Fulfillment CLI. The tenant must provide: +1. The tenant initiates the creation of a new VirtualMachine using the + Fulfillment CLI. The tenant must provide: - The ID of the desired VirtualMachineTemplate - - All required and any optional parameters for the template (such as CPU, memory, disk size, network configuration) + - All required and any optional parameters for the template (such as CPU, + memory, disk size, network configuration) - The desired initial state of the VM (e.g., started or stopped) -2. The Fulfillment Service receives this request and performs validation to ensure: +2. The Fulfillment Service receives this request and performs validation to + ensure: - The specified template exists and is available - All required parameters are provided and valid -3. Upon successful validation, the Fulfillment Service creates a new VirtualMachine custom resource (CR) in the appropriate Hub and namespace. +3. Upon successful validation, the Fulfillment Service creates a new + VirtualMachine custom resource (CR) in the appropriate Hub and namespace. -4. The O-SAC Operator detects the new VirtualMachine CR and begins the reconciliation process. +4. The O-SAC Operator detects the new VirtualMachine CR and begins the + reconciliation process. -5. The Operator, using Ansible Automation Platform (AAP), automates the following steps: +5. The Operator, using Ansible Automation Platform (AAP), automates the + following steps: - Creates a dedicated namespace for the VM if one does not already exist - Provisions necessary network resources, including: - A UDN L2 network to provide network isolation - A load balancer service with the VM as its backend - - Assignment of a floating IP to the load balancer service for external access - - Creates the KubeVirt VirtualMachine resource using the specified template and parameters + - Assignment of a floating IP to the load balancer service for external + access + - Creates the KubeVirt VirtualMachine resource using the specified template + and parameters - Performs any additional operations required by the selected template -6. The Operator continuously monitors the VM’s status and updates the VirtualMachine CR status to reflect the current state. +6. The Operator continuously monitors the VM’s status and updates the + VirtualMachine CR status to reflect the current state. -7. The tenant can check the VM’s status at any time using the Fulfillment CLI or API, and can access the VM via the assigned floating IP. +7. The tenant can check the VM’s status at any time using the Fulfillment CLI or + API, and can access the VM via the assigned floating IP. -The update process is the same as the creation workflow as it will be designed to be idempotent. +The update process is the same as the creation workflow as it will be designed +to be idempotent. #### Virtual machine deletion -When a tenant requests the deletion of a VirtualMachine, the following workflow is executed: +When a tenant requests the deletion of a VirtualMachine, the following workflow +is executed: -1. The tenant initiates the deletion of a VirtualMachine using the Fulfillment CLI or API by specifying its identifier. +1. The tenant initiates the deletion of a VirtualMachine using the Fulfillment + CLI or API by specifying its identifier. -2. The Fulfillment Service receives the deletion request and performs validation to ensure: +2. The Fulfillment Service receives the deletion request and performs validation + to ensure: - The specified VirtualMachine resource exists and is available. - The tenant has permission to delete the resource. -3. Upon successful validation, the Fulfillment Service deletes the VirtualMachine custom resource (CR) from the appropriate namespace. +3. Upon successful validation, the Fulfillment Service deletes the + VirtualMachine custom resource (CR) from the appropriate namespace. -4. The O-SAC Operator detects the deletion of the VirtualMachine CR and begins the cleanup process. +4. The O-SAC Operator detects the deletion of the VirtualMachine CR and begins + the cleanup process. -5. The Operator, using Ansible Automation Platform (AAP), automates the following steps: +5. The Operator, using Ansible Automation Platform (AAP), automates the + following steps: - Deletes the KubeVirt VirtualMachine resource. - - Releases and deallocates any associated network resources, including: + - Releases and deallocate any associated network resources, including: - UDN L2 network - Load balancer service - Floating IPs via ESI APIs - - Performs any additional cleanup operations required by the selected virtual machine template. + - Performs any additional cleanup operations required by the selected + virtual machine template. - Deletes the dedicated namespace if it is no longer needed. -6. The Operator updates the status of the deletion operation and ensures all resources are properly cleaned up. +6. The Operator updates the status of the deletion operation and ensures all + resources are properly cleaned up. -7. The tenant can confirm the deletion and cleanup via the Fulfillment CLI or API. +7. The tenant can confirm the deletion and cleanup via the Fulfillment CLI or + API. -This workflow ensures that all resources associated with the VirtualMachine are properly deprovisioned and that no orphaned resources remain. +This workflow ensures that all resources associated with the VirtualMachine are +properly deprovisioned and that no orphaned resources remain. #### Virtual machine template management -Virtual machine templates are centrally managed by the provider using a GitOps approach. All templates are stored in a version-controlled repository, which acts as the single source of truth. At regular intervals, an automated job in Ansible Automation Platform (AAP) synchronizes the latest templates from this repository to the Fulfillment Service. As a result, any changes to the templates—such as updates, additions, or deletions—are automatically and consistently reflected in the Fulfillment Service. This process ensures that tenants always have access to the most current catalog of available VM templates. +Virtual machine templates are centrally managed by the provider using a GitOps +approach. All templates are stored in a version-controlled repository, which +acts as the single source of truth. At regular intervals, an automated job in +Ansible Automation Platform (AAP) synchronizes the latest templates from this +repository to the Fulfillment Service. As a result, any changes to the +templates—such as updates, additions, or deletions—are automatically and +consistently reflected in the Fulfillment Service. This process ensures that +tenants always have access to the most current catalog of available VM +templates. ### API Extensions #### VirtualMachine -A tenant requests a virtual machine by requesting a VirtualMachine to the Fulfillment Service. Here is an example of request that creates a VM, using a template that let tenants to customize the amount of CPUs, memory and boot disk size: +A tenant requests a virtual machine by requesting a VirtualMachine to the +Fulfillment Service. Here is an example of request that creates a VM, using a +template that let tenants to customize the amount of CPUs, memory and boot disk +size: ```json { @@ -167,7 +244,8 @@ A tenant requests a virtual machine by requesting a VirtualMachine to the Fulfil } ``` -After creating a virtual machine, tenants can check its current status and details as follows: +After creating a virtual machine, tenants can check its current status and +details as follows: ```json { @@ -232,12 +310,6 @@ After creating a virtual machine, tenants can check its current status and detai "status": "CONDITION_STATUS_FALSE", "type": "VIRTUAL_MACHINE_STATE_TERMINATING" }, - { - "last_transition_time": "2025-09-17T08:52:12.652582382Z", - "message": "", - "status": "CONDITION_STATUS_FALSE", - "type": "VIRTUAL_MACHINE_STATE_MIGRATING" - }, { "last_transition_time": "2025-09-17T08:52:12.652582382Z", "message": "", @@ -262,30 +334,46 @@ After creating a virtual machine, tenants can check its current status and detai The status section provides two types of IP addresses for the virtual machine: -- `internalIP`: The private IP address assigned to the VM on the internal network. +- `internalIP`: The private IP address assigned to the VM on the internal + network. -- `externalIP`: The public (floating) IP address assigned to the VM. This address allows the VM to be accessed from outside the internal network, such as from the internet. +- `externalIP`: The public (floating) IP address assigned to the VM. This + address allows the VM to be accessed from outside the internal network, such + as from the internet. -The Virtual Machine can be in one of the following states, as reflected in the status section above: +The Virtual Machine can be in one of the following states, as reflected in the +status section above. These states are mapped from the underlying KubeVirt +VirtualMachine status conditions: -- **Provisioning** (`VIRTUAL_MACHINE_STATE_PROVISIONING`): The VM is being created. -- **Starting** (`VIRTUAL_MACHINE_STATE_STARTING`): The Pod for the Virtual Machine Instance (VMI) is being scheduled and started. -- **Running** (`VIRTUAL_MACHINE_STATE_RUNNING`): The VM is actively running inside its Pod. -- **Stopping** (`VIRTUAL_MACHINE_STATE_STOPPING`): The VM is in the process of shutting down. -- **Stopped** (`VIRTUAL_MACHINE_STATE_STOPPED`): The VM is not running. It exists as a VirtualMachine object, but there is no active VMI or Pod. -- **Terminating** (`VIRTUAL_MACHINE_STATE_TERMINATING`): The VM object is in the process of being deleted. -- **Migrating** (`VIRTUAL_MACHINE_STATE_MIGRATING`): The VM is in the process of being live-migrated to another node. -- **Paused** (`VIRTUAL_MACHINE_STATE_PAUSED`): The VM is in a suspended state. Its process is frozen, but its memory and resources are still allocated. +- **Provisioning** (`VIRTUAL_MACHINE_STATE_PROVISIONING`): The VM is being + created. +- **Starting** (`VIRTUAL_MACHINE_STATE_STARTING`): The Pod for the Virtual + Machine Instance (VMI) is being scheduled and started. +- **Running** (`VIRTUAL_MACHINE_STATE_RUNNING`): The VM is actively running + inside its Pod. +- **Stopping** (`VIRTUAL_MACHINE_STATE_STOPPING`): The VM is in the process of + shutting down. +- **Stopped** (`VIRTUAL_MACHINE_STATE_STOPPED`): The VM is not running. It + exists as a VirtualMachine object, but there is no active VMI or Pod. +- **Terminating** (`VIRTUAL_MACHINE_STATE_TERMINATING`): The VM object is in the + process of being deleted. +- **Paused** (`VIRTUAL_MACHINE_STATE_PAUSED`): The VM is in a suspended state. + Its process is frozen, but its memory and resources are still allocated. -These states are reported in the `type` field of the VM's `conditions` array in the status section. +These states are reported in the `type` field of the VM's `conditions` array in +the status section. #### VirtualMachineTemplate -Virtual machine templates are implemented as Ansible roles. Each role must include the following files: +Virtual machine templates are implemented as Ansible roles. Each role must +include the following files: -* `vm_template_role/meta/argument_specs.yaml`: Defines the [Ansible argument specification](https://docs.ansible.com/ansible/latest/dev_guide/developing_program_flow_modules.html#argument-spec). This file is required. -* `vm_template_role/meta/cloudkit.yaml`: Contains metadata for the template, including the title and description. This file is required. +* `vm_template_role/meta/argument_specs.yaml`: Defines the [Ansible argument + specification](https://docs.ansible.com/ansible/latest/dev_guide/developing_program_flow_modules.html#argument-spec). + This file is required. +* `vm_template_role/meta/cloudkit.yaml`: Contains metadata for the template, + including the title and description. This file is required. Example of `cloudkit.yaml`: @@ -295,7 +383,9 @@ description: > This template provisions a virtual machine. ``` -A periodic job will publish the Ansible roles as virtual machine templates to the Fulfillment Service, using the required files described above. The following is the API format used for this publication: +A periodic job will publish the Ansible roles as virtual machine templates to +the Fulfillment Service, using the required files described above. The following +is the API format used for this publication: ```json { @@ -332,11 +422,37 @@ A periodic job will publish the Ansible roles as virtual machine templates to th #### Virtual machines on HUB cluster -Virtual machines will be created on the HUB cluster chosen by the Fulfillment Service. Although there was discussion about creating a dedicated cluster for VM workloads using the Cluster-as-a-Service API, this approach would not improve reliability. This is because these dedicated clusters would still be implemented as HostedClusters, whose reliability ultimately depends on the same underlying HUB cluster. +Virtual machines will be created on the HUB cluster chosen by the Fulfillment +Service. Although there was discussion about creating a dedicated cluster for VM +workloads using the Cluster-as-a-Service API, this approach would not improve +reliability. This is because these dedicated clusters would still be implemented +as HostedClusters, whose reliability ultimately depends on the same underlying +HUB cluster. + +#### Networking + +Because O-SAC does not yet provide a VDCaaS (Virtual Data Center as a Service) +layer, and the Fulfillment Service cannot guarantee that two virtual machines +will be provisioned on the same HUB cluster, each virtual machine must be +assigned a floating IP. This ensures that every VM is accessible regardless of +where it is deployed. -#### Netowrking -Because O-SAC does not yet provide a VDCaaS (Virtual Data Center as a Service) layer, and the Fulfillment Service cannot guarantee that two virtual machines will be provisioned on the same HUB cluster, each virtual machine must be assigned a floating IP. This ensures that every VM is accessible regardless of where it is deployed. +#### Load balancer service type and MetalLB + +Even though a Kubernetes Service of type LoadBalancer requires explicit port +mappings to be defined, MetalLB will actually expose all ports on the assigned +external IP address. This means that, in practice, any port on the VM can be +accessed via the floating IP, regardless of the ports specified in the Service +manifest. + +#### UDN and network isolation + +Virtual machines that are provisioned on different UDN networks will still be +able to communicate with each other by using their assigned floating IPs. Since +each VM is given a floating IP for external access, network traffic between VMs +on separate UDN networks can be routed through these public endpoints, ensuring +connectivity even in the absence of a shared internal network. ### Risks and Mitigations @@ -344,7 +460,24 @@ TBD ### Drawbacks -TBD +The initial design of VMaaS has significant limitations due to the absence of +regions and zones in our API: + +- Tenants cannot create complex VM architectures that use both public and + private networks. Features such as security groups and network ACLs are also + missing. +- Tenants are unable to provision permanent block storage that can be attached + to and detached from different VMs as needed. + +We expect to revisit and improve this design once the VDCaaS (Virtual Data +Center as a Service) functionality is available. + +Another limitation is that currently, virtual machines can only be managed on +the local (HUB) cluster. In the future, there may be a need to manage virtual +machines on remote or dedicated clusters, especially if service providers deploy +clusters specifically for VM workloads. This would allow for enhanced security +and the use of specialized hardware. The development of such feature will be +part of another enhancement. ## Alternatives (Not Implemented) @@ -352,7 +485,6 @@ TBD ## Open Questions [optional] -- I think that without the concept of regions (mapped on HUB clusters?) in the Fulfillment Service, we can't move much forward with storage (as it is local to a HUB), and with the ability to create VMs without floating IPs (for example, I want my web server with a floating IP to communicate internally with my DB). ## Test Plan From 369783f36bf5fec1c1f1a225a2bc22e3bc1d1d46 Mon Sep 17 00:00:00 2001 From: Adrien Gentil Date: Thu, 9 Oct 2025 17:26:12 +0200 Subject: [PATCH 5/7] address comments --- enhancements/vmaas/README.md | 134 +++++++++++++++++------------------ 1 file changed, 64 insertions(+), 70 deletions(-) diff --git a/enhancements/vmaas/README.md b/enhancements/vmaas/README.md index 343285218..54daa63af 100644 --- a/enhancements/vmaas/README.md +++ b/enhancements/vmaas/README.md @@ -1,5 +1,5 @@ --- -title: Virtualization as a service +title: Virtual-Machine-as-a-service authors: - Adrien Gentil creation-date: 2025-09-15 @@ -11,7 +11,7 @@ replaces: superseded-by: --- -# Virtualization-as-a-Service +# Virtual-Machine-as-a-Service ## Summary @@ -29,15 +29,17 @@ separate proposal. ## Motivation -Virtualization-as-a-Service (VMaaS) addresses the need for flexible, on-demand -compute resources within a multi-tenant environments. This is a need identified -in the scope of O-SAC, and will benefit to the MOC. +Virtual-Machine-as-a-Service (VMaaS) addresses the need for flexible, on-demand +compute resources within a multi-tenant environment. Additionally, VMaaS enables +the sharing of specialized hardware resources, such as GPUs, across multiple +projects, maximizing hardware utilization and accessibility. This is a need +identified in the scope of O-SAC, and will benefit the MOC. ### User Stories - As a provider, I want to define VM templates that my tenants will be able to use -- As a tenant, I want to list a pre-defined VM templates +- As a tenant, I want to list pre-defined VM templates - As a tenant, I want to create a VM based on a pre-defined template - As a tenant, I want a VM that has access to specialized hardware (e.g.: GPU) - As a tenant, I want to manage the lifecycle of my VM (start/stop/terminate) @@ -67,12 +69,12 @@ The following are explicitly out of scope for this proposal: - Offering built-in backup, restore, or disaster recovery solutions for VMs or attached storage - Delivering a marketplace for third-party VM images or applications, we expect - tenant to rely on en external image registry to distribute they OS base images + tenant to rely on an external image registry to distribute OS base images ## Proposal The process of fulfilling virtual machine requests is based on two primary -concepts: +APIs: * **VirtualMachine**: Represents an individual virtual machine that a tenant can create and manage. Tenants can only see the virtual machines they created. @@ -103,21 +105,6 @@ enhanced or updated: interactions with KubeVirt for VM lifecycle management and ESI APIs for assigning floating IPs. -Under the hood, the virtual machines will be managed using OpenShift -Virtualization. Each VirtualMachine will be created in its own dedicated -namespace on the Hub cluster. The VM will be connected to a dedicated UDN L2 -network, which provides network isolation, mainly for 2 reasons: -- security: this means the VM cannot communicate with other workloads on the Hub - cluster using its internal IP address. -- Live migration: UDN L2 enables the virtual machine to be migrated live from - OpenShift nodes to others. Migration can happen when a node is degraded or in - maintenance. - -To enable external connectivity, a Kubernetes Service of type LoadBalancer will -be created in front of the VM. A floating (external) IP address will be assigned -to this LoadBalancer service, allowing the VM to be accessed from outside the -cluster. - ### Workflow Description #### Virtual machine creation and update @@ -142,8 +129,8 @@ cluster. 5. The Operator, using Ansible Automation Platform (AAP), automates the following steps: - - Creates a dedicated namespace for the VM if one does not already exist - - Provisions necessary network resources, including: + - Provisions necessary resources, including: + - Tenant's namespace - A UDN L2 network to provide network isolation - A load balancer service with the VM as its backend - Assignment of a floating IP to the load balancer service for external @@ -183,7 +170,8 @@ is executed: 5. The Operator, using Ansible Automation Platform (AAP), automates the following steps: - Deletes the KubeVirt VirtualMachine resource. - - Releases and deallocate any associated network resources, including: + - Releases and deallocate any associated resources, including: + - Tenant's namespace - UDN L2 network - Load balancer service - Floating IPs via ESI APIs @@ -419,69 +407,75 @@ is the API format used for this publication: ### Implementation Details/Notes/Constraints +This proposal is built upon OpenShift Virtualization, which leverages KubeVirt +to provide virtualization capabilities. By using OpenShift Virtualization, we +are able to offer tenants a self-service platform for creating, managing, and +operating virtual machines (VMs) with minimal operational overhead, directly +aligning with our goal of providing a user-friendly API for VM lifecycle +management. + +When a VM is created, it is connected to a dedicated UDN (User Defined Network) +that is assigned to the tenant. In other words, all VMs belonging to the same +tenant share a single UDN, while VMs from different tenants are placed on +separate UDNs. This design ensures strong network isolation between tenants. + +Even though all of a tenant's VMs share the same UDN (User Defined Network) in a +given HUB cluster, each VM still needs its own floating IP in order to be +accessible from other VMs. This is because the Fulfillment Service may provision +a tenant's VMs on different HUB clusters, as a consequence, internal network +connectivity between VMs cannot be guaranteed. The floating IP is also required +to access the VM from the outside. + +To allow services running on a VM to be accessible from outside the cluster, the +template may handle the port mapping through a parameter. Defining which ports +to be exposed as a template parameter provides flexibility: it allows us to +easily adjust how ports will be exposed in the future, especially when VDCaaS +(Virtual Data Center as a Service) will introduce new ways to manage external +access. -#### Virtual machines on HUB cluster - -Virtual machines will be created on the HUB cluster chosen by the Fulfillment -Service. Although there was discussion about creating a dedicated cluster for VM -workloads using the Cluster-as-a-Service API, this approach would not improve -reliability. This is because these dedicated clusters would still be implemented -as HostedClusters, whose reliability ultimately depends on the same underlying -HUB cluster. - -#### Networking - -Because O-SAC does not yet provide a VDCaaS (Virtual Data Center as a Service) -layer, and the Fulfillment Service cannot guarantee that two virtual machines -will be provisioned on the same HUB cluster, each virtual machine must be -assigned a floating IP. This ensures that every VM is accessible regardless of -where it is deployed. +### Risks and Mitigations +TBD -#### Load balancer service type and MetalLB +### Drawbacks -Even though a Kubernetes Service of type LoadBalancer requires explicit port -mappings to be defined, MetalLB will actually expose all ports on the assigned -external IP address. This means that, in practice, any port on the VM can be -accessed via the floating IP, regardless of the ports specified in the Service -manifest. +#### Virtual machines on HUB cluster -#### UDN and network isolation +Initially, all virtual machines will be created on the HUB cluster selected by +the Fulfillment Service. -Virtual machines that are provisioned on different UDN networks will still be -able to communicate with each other by using their assigned floating IPs. Since -each VM is given a floating IP for external access, network traffic between VMs -on separate UDN networks can be routed through these public endpoints, ensuring -connectivity even in the absence of a shared internal network. +In the future, we plan to support deploying virtual machines on remote, +dedicated clusters. This approach offers several advantages: -### Risks and Mitigations +- It provides stronger isolation between customer workloads and the cloud + provider’s management tools +- It allows the management (HUB) cluster to be upgraded, maintained, or migrated + independently from the clusters running customer VMs +- The clusters hosting VMs can be managed and monitored according to their own + uptime and operational requirements, even if they share hardware with the + management cluster (in case of Hosted Clusters) +- It enables the possibility of creating dedicated virtualization clusters for + individual tenants, which is important for tenants who require higher levels + of isolation -TBD +This feature will be part of a separate enhancement. -### Drawbacks +#### Networking -The initial design of VMaaS has significant limitations due to the absence of -regions and zones in our API: +While simple, the initial design of VMaaS networking has significant limitations: - Tenants cannot create complex VM architectures that use both public and private networks. Features such as security groups and network ACLs are also - missing. -- Tenants are unable to provision permanent block storage that can be attached - to and detached from different VMs as needed. + missing +- Each gets a floating IP assigned to it, this pool of IPs is limited +- OpenShift team doesn't recommend more than 100 UDNs per cluster, this solution + will then be limited to 100 tenants using VMs We expect to revisit and improve this design once the VDCaaS (Virtual Data Center as a Service) functionality is available. -Another limitation is that currently, virtual machines can only be managed on -the local (HUB) cluster. In the future, there may be a need to manage virtual -machines on remote or dedicated clusters, especially if service providers deploy -clusters specifically for VM workloads. This would allow for enhanced security -and the use of specialized hardware. The development of such feature will be -part of another enhancement. - ## Alternatives (Not Implemented) -TBD ## Open Questions [optional] From 184d82ef54bb8fb6182365a65378a126978456ef Mon Sep 17 00:00:00 2001 From: Adrien Gentil Date: Thu, 9 Oct 2025 17:51:47 +0200 Subject: [PATCH 6/7] VirtualMachine to ComputeInstance --- enhancements/vmaas/README.md | 92 ++++++++++++++++++------------------ 1 file changed, 46 insertions(+), 46 deletions(-) diff --git a/enhancements/vmaas/README.md b/enhancements/vmaas/README.md index 54daa63af..812a58cfb 100644 --- a/enhancements/vmaas/README.md +++ b/enhancements/vmaas/README.md @@ -73,20 +73,19 @@ The following are explicitly out of scope for this proposal: ## Proposal -The process of fulfilling virtual machine requests is based on two primary -APIs: +The process of fulfilling virtual machine requests is based on two primary APIs: -* **VirtualMachine**: Represents an individual virtual machine that a tenant can - create and manage. Tenants can only see the virtual machines they created. -* **VirtualMachineTemplate**: Defined by the provider, this is a pre-configured +* **ComputeInstance**: Represents an individual virtual machine that a tenant + can create and manage. Tenants can only see the virtual machines they created. +* **ComputeInstanceTemplate**: Defined by the provider, this is a pre-configured blueprint for virtual machines. Each template is identified by a unique template ID and includes a set of parameters (some required, some optional) that tenants can specify when creating a VM. Templates are available to all tenants to use, they cannot edit them. -To request a new VirtualMachine, tenants must provide: +To request a new ComputeInstance, tenants must provide: -* The ID of the desired VirtualMachineTemplate +* The ID of the desired ComputeInstanceTemplate * Any required or optional parameters for that template * The desired initial state of the VM (e.g., started or stopped) @@ -95,13 +94,13 @@ fulfillment workflows. To support this, the following O-SAC components will be enhanced or updated: * **Fulfillment Service**: Defines and exposes the APIs for managing - VirtualMachine and VirtualMachineTemplate resources. + ComputeInstance and ComputeInstanceTemplate resources. * **Fulfillment CLI**: Provides tenants with command-line access to the Fulfillment Service APIs. -* **O-SAC Operator**: Monitors and reconciles VirtualMachine custom resources +* **O-SAC Operator**: Monitors and reconciles ComputeInstance custom resources within the system. * **O-SAC AAP (Ansible Automation Platform)**: Executes automation tasks (via - Ansible playbooks) to reconcile VirtualMachine resources, including + Ansible playbooks) to reconcile ComputeInstance resources, including interactions with KubeVirt for VM lifecycle management and ESI APIs for assigning floating IPs. @@ -109,9 +108,9 @@ enhanced or updated: #### Virtual machine creation and update -1. The tenant initiates the creation of a new VirtualMachine using the +1. The tenant initiates the creation of a new ComputeInstance using the Fulfillment CLI. The tenant must provide: - - The ID of the desired VirtualMachineTemplate + - The ID of the desired ComputeInstanceTemplate - All required and any optional parameters for the template (such as CPU, memory, disk size, network configuration) - The desired initial state of the VM (e.g., started or stopped) @@ -122,9 +121,9 @@ enhanced or updated: - All required parameters are provided and valid 3. Upon successful validation, the Fulfillment Service creates a new - VirtualMachine custom resource (CR) in the appropriate Hub and namespace. + ComputeInstance custom resource (CR) in the appropriate Hub and namespace. -4. The O-SAC Operator detects the new VirtualMachine CR and begins the +4. The O-SAC Operator detects the new ComputeInstance CR and begins the reconciliation process. 5. The Operator, using Ansible Automation Platform (AAP), automates the @@ -140,7 +139,7 @@ enhanced or updated: - Performs any additional operations required by the selected template 6. The Operator continuously monitors the VM’s status and updates the - VirtualMachine CR status to reflect the current state. + ComputeInstance CR status to reflect the current state. 7. The tenant can check the VM’s status at any time using the Fulfillment CLI or API, and can access the VM via the assigned floating IP. @@ -150,21 +149,21 @@ to be idempotent. #### Virtual machine deletion -When a tenant requests the deletion of a VirtualMachine, the following workflow +When a tenant requests the deletion of a ComputeInstance, the following workflow is executed: -1. The tenant initiates the deletion of a VirtualMachine using the Fulfillment +1. The tenant initiates the deletion of a ComputeInstance using the Fulfillment CLI or API by specifying its identifier. 2. The Fulfillment Service receives the deletion request and performs validation to ensure: - - The specified VirtualMachine resource exists and is available. + - The specified ComputeInstance resource exists and is available. - The tenant has permission to delete the resource. 3. Upon successful validation, the Fulfillment Service deletes the - VirtualMachine custom resource (CR) from the appropriate namespace. + ComputeInstance custom resource (CR) from the appropriate namespace. -4. The O-SAC Operator detects the deletion of the VirtualMachine CR and begins +4. The O-SAC Operator detects the deletion of the ComputeInstance CR and begins the cleanup process. 5. The Operator, using Ansible Automation Platform (AAP), automates the @@ -185,7 +184,7 @@ is executed: 7. The tenant can confirm the deletion and cleanup via the Fulfillment CLI or API. -This workflow ensures that all resources associated with the VirtualMachine are +This workflow ensures that all resources associated with the ComputeInstance are properly deprovisioned and that no orphaned resources remain. #### Virtual machine template management @@ -202,9 +201,9 @@ templates. ### API Extensions -#### VirtualMachine +#### ComputeInstance -A tenant requests a virtual machine by requesting a VirtualMachine to the +A tenant requests a virtual machine by requesting a ComputeInstance to the Fulfillment Service. Here is an example of request that creates a VM, using a template that let tenants to customize the amount of CPUs, memory and boot disk size: @@ -237,7 +236,7 @@ details as follows: ```json { - "@type": "type.googleapis.com/fulfillment.v1.VirtualMachine", + "@type": "type.googleapis.com/fulfillment.v1.ComputeInstance", "id": "fecb9b9e-07ac-4d56-8b48-9d50aab71677", "metadata": { "creation_timestamp": "2025-09-17T08:14:17.569076Z", @@ -266,54 +265,54 @@ details as follows: "last_transition_time": "2025-09-19T17:32:24.054439350Z", "message": "", "status": "CONDITION_STATUS_FALSE", - "type": "VIRTUAL_MACHINE_CONDITION_TYPE_PROVISIONNING" + "type": "COMPUTE_INSTANCE_CONDITION_TYPE_PROVISIONNING" }, { "last_transition_time": "2025-09-17T08:52:12.652582382Z", "message": "", "status": "CONDITION_STATUS_FALSE", - "type": "VIRTUAL_MACHINE_STATE_STARTING" + "type": "COMPUTE_INSTANCE_STATE_STARTING" }, { "last_transition_time": "2025-09-17T08:52:12.652582382Z", "message": "", "status": "CONDITION_STATUS_TRUE", - "type": "VIRTUAL_MACHINE_STATE_RUNNING" + "type": "COMPUTE_INSTANCE_STATE_RUNNING" }, { "last_transition_time": "2025-09-17T08:52:12.652582382Z", "message": "", "status": "CONDITION_STATUS_FALSE", - "type": "VIRTUAL_MACHINE_STATE_STOPPING" + "type": "COMPUTE_INSTANCE_STATE_STOPPING" }, { "last_transition_time": "2025-09-17T08:52:12.652582382Z", "message": "", "status": "CONDITION_STATUS_FALSE", - "type": "VIRTUAL_MACHINE_STATE_STOPPED" + "type": "COMPUTE_INSTANCE_STATE_STOPPED" }, { "last_transition_time": "2025-09-17T08:52:12.652582382Z", "message": "", "status": "CONDITION_STATUS_FALSE", - "type": "VIRTUAL_MACHINE_STATE_TERMINATING" + "type": "COMPUTE_INSTANCE_STATE_TERMINATING" }, { "last_transition_time": "2025-09-17T08:52:12.652582382Z", "message": "", "status": "CONDITION_STATUS_FALSE", - "type": "VIRTUAL_MACHINE_STATE_PAUSED" + "type": "COMPUTE_INSTANCE_STATE_PAUSED" }, { "status": "CONDITION_STATUS_FALSE", - "type": "VIRTUAL_MACHINE_CONDITION_TYPE_FAILED" + "type": "COMPUTE_INSTANCE_CONDITION_TYPE_FAILED" }, { "status": "CONDITION_STATUS_FALSE", - "type": "VIRTUAL_MACHINE_CONDITION_TYPE_DEGRADED" + "type": "COMPUTE_INSTANCE_CONDITION_TYPE_DEGRADED" } ], - "state": "VIRTUAL_MACHINE_STATE_RUNNING", + "state": "COMPUTE_INSTANCE_STATE_RUNNING", "internalIP": "10.0.0.1", "externalIP": "193.1.2.3" } @@ -329,30 +328,30 @@ The status section provides two types of IP addresses for the virtual machine: address allows the VM to be accessed from outside the internal network, such as from the internet. -The Virtual Machine can be in one of the following states, as reflected in the +The virtual machine can be in one of the following states, as reflected in the status section above. These states are mapped from the underlying KubeVirt VirtualMachine status conditions: -- **Provisioning** (`VIRTUAL_MACHINE_STATE_PROVISIONING`): The VM is being +- **Provisioning** (`COMPUTE_INSTANCE_STATE_PROVISIONING`): The VM is being created. -- **Starting** (`VIRTUAL_MACHINE_STATE_STARTING`): The Pod for the Virtual +- **Starting** (`COMPUTE_INSTANCE_STATE_STARTING`): The Pod for the Virtual Machine Instance (VMI) is being scheduled and started. -- **Running** (`VIRTUAL_MACHINE_STATE_RUNNING`): The VM is actively running +- **Running** (`COMPUTE_INSTANCE_STATE_RUNNING`): The VM is actively running inside its Pod. -- **Stopping** (`VIRTUAL_MACHINE_STATE_STOPPING`): The VM is in the process of +- **Stopping** (`COMPUTE_INSTANCE_STATE_STOPPING`): The VM is in the process of shutting down. -- **Stopped** (`VIRTUAL_MACHINE_STATE_STOPPED`): The VM is not running. It - exists as a VirtualMachine object, but there is no active VMI or Pod. -- **Terminating** (`VIRTUAL_MACHINE_STATE_TERMINATING`): The VM object is in the - process of being deleted. -- **Paused** (`VIRTUAL_MACHINE_STATE_PAUSED`): The VM is in a suspended state. +- **Stopped** (`COMPUTE_INSTANCE_STATE_STOPPED`): The VM is not running. It + exists as a ComputeInstance object, but there is no active VMI or Pod. +- **Terminating** (`COMPUTE_INSTANCE_STATE_TERMINATING`): The VM object is in + the process of being deleted. +- **Paused** (`COMPUTE_INSTANCE_STATE_PAUSED`): The VM is in a suspended state. Its process is frozen, but its memory and resources are still allocated. These states are reported in the `type` field of the VM's `conditions` array in the status section. -#### VirtualMachineTemplate +#### ComputeInstanceTemplate Virtual machine templates are implemented as Ansible roles. Each role must include the following files: @@ -462,7 +461,8 @@ This feature will be part of a separate enhancement. #### Networking -While simple, the initial design of VMaaS networking has significant limitations: +While simple, the initial design of VMaaS networking has significant +limitations: - Tenants cannot create complex VM architectures that use both public and private networks. Features such as security groups and network ACLs are also From f6865333e55e5b44e737d71888076c8a1ca52466 Mon Sep 17 00:00:00 2001 From: Adrien Gentil Date: Mon, 13 Oct 2025 09:58:04 +0200 Subject: [PATCH 7/7] re-wording condition descriptions --- enhancements/vmaas/README.md | 48 +++++++++++++++++++----------------- 1 file changed, 25 insertions(+), 23 deletions(-) diff --git a/enhancements/vmaas/README.md b/enhancements/vmaas/README.md index 812a58cfb..15aa8b255 100644 --- a/enhancements/vmaas/README.md +++ b/enhancements/vmaas/README.md @@ -18,7 +18,7 @@ superseded-by: This document proposes a service enabling tenants to easily create, manage, and operate virtual machines (VMs) within a self-service environment. The service -will provide user-friendly APIs for provisioning, customizing, and controlling +will provide user-friendly APIs for provisioning, customizing, and controlling the lifecycle of VMs attaching storage, and accessing specialized hardware (GPUs). @@ -32,8 +32,8 @@ separate proposal. Virtual-Machine-as-a-Service (VMaaS) addresses the need for flexible, on-demand compute resources within a multi-tenant environment. Additionally, VMaaS enables the sharing of specialized hardware resources, such as GPUs, across multiple -projects, maximizing hardware utilization and accessibility. This is a need -identified in the scope of O-SAC, and will benefit the MOC. +projects, maximizing hardware utilization and accessibility. This addresses a +need identified in the scope of O-SAC, and will benefit the MOC. ### User Stories @@ -56,6 +56,9 @@ identified in the scope of O-SAC, and will benefit the MOC. through the catalog of templates - Use ESI to assign floating IPs to created VM so they can be accessed - Achieve high availability by supporting live migration of VMs +- Implement a quota system (to limit the number of VMs, storage, or resources + each tenant can use) is out of scope for this enhancement and will be + addressed in a separate proposal ### Non-Goals @@ -265,7 +268,7 @@ details as follows: "last_transition_time": "2025-09-19T17:32:24.054439350Z", "message": "", "status": "CONDITION_STATUS_FALSE", - "type": "COMPUTE_INSTANCE_CONDITION_TYPE_PROVISIONNING" + "type": "COMPUTE_INSTANCE_CONDITION_TYPE_PROVISIONING" }, { "last_transition_time": "2025-09-17T08:52:12.652582382Z", @@ -321,31 +324,30 @@ details as follows: The status section provides two types of IP addresses for the virtual machine: -- `internalIP`: The private IP address assigned to the VM on the internal - network. +- `internalIP`: The private IP address assigned to the virtual machine on the + internal network. -- `externalIP`: The public (floating) IP address assigned to the VM. This - address allows the VM to be accessed from outside the internal network, such - as from the internet. +- `externalIP`: The public (floating) IP address assigned to the virtual + machine. This address allows the virtual machine to be accessed from outside + the internal network, such as from the internet. The virtual machine can be in one of the following states, as reflected in the status section above. These states are mapped from the underlying KubeVirt VirtualMachine status conditions: -- **Provisioning** (`COMPUTE_INSTANCE_STATE_PROVISIONING`): The VM is being - created. -- **Starting** (`COMPUTE_INSTANCE_STATE_STARTING`): The Pod for the Virtual - Machine Instance (VMI) is being scheduled and started. -- **Running** (`COMPUTE_INSTANCE_STATE_RUNNING`): The VM is actively running - inside its Pod. -- **Stopping** (`COMPUTE_INSTANCE_STATE_STOPPING`): The VM is in the process of - shutting down. -- **Stopped** (`COMPUTE_INSTANCE_STATE_STOPPED`): The VM is not running. It - exists as a ComputeInstance object, but there is no active VMI or Pod. -- **Terminating** (`COMPUTE_INSTANCE_STATE_TERMINATING`): The VM object is in - the process of being deleted. -- **Paused** (`COMPUTE_INSTANCE_STATE_PAUSED`): The VM is in a suspended state. - Its process is frozen, but its memory and resources are still allocated. +- **Provisioning** (`COMPUTE_INSTANCE_STATE_PROVISIONING`): The virtual machine + is being created. +- **Starting** (`COMPUTE_INSTANCE_STATE_STARTING`): The virtual machine started. +- **Running** (`COMPUTE_INSTANCE_STATE_RUNNING`): The virtual machine is + actively running. +- **Stopping** (`COMPUTE_INSTANCE_STATE_STOPPING`): The virtual machine is in + the process of shutting down. +- **Stopped** (`COMPUTE_INSTANCE_STATE_STOPPED`): The virtual machine is + stopped. +- **Terminating** (`COMPUTE_INSTANCE_STATE_TERMINATING`): The virtual machine is + in the process of being deleted. +- **Paused** (`COMPUTE_INSTANCE_STATE_PAUSED`): The virtual machine is in a + suspended state. These states are reported in the `type` field of the VM's `conditions` array in the status section.