From 0f755fba2906ce290dae9d24526ff6e2e840a886 Mon Sep 17 00:00:00 2001 From: Austin Jamias Date: Tue, 23 Sep 2025 16:04:18 -0400 Subject: [PATCH 01/10] Add Baremetal Observability Enhancement --- .../baremetal-observability/README.md | 219 ++++++++++++++++++ 1 file changed, 219 insertions(+) create mode 100644 enhancements/baremetal-observability/README.md diff --git a/enhancements/baremetal-observability/README.md b/enhancements/baremetal-observability/README.md new file mode 100644 index 000000000..d9aa701f8 --- /dev/null +++ b/enhancements/baremetal-observability/README.md @@ -0,0 +1,219 @@ +--- +title: bare-metal-observability +authors: + - Austin Jamias +creation-date: 2025-09-24 +last-updated: 2025-09-24 +tracking-link: # link to the tracking ticket (for example: Github issue) that corresponds to this enhancement + - N/A +see-also: + - N/A +replaces: + - N/A +superseded-by: + - N/A +--- + +# Bare Metal Observability + +## Summary + +We want a way to gather bare metal metrics from tenant clusters and have them +available at some configurable endpoint for the cloud provider to manage. This +feature will be easily toggleable per tenant cluster. + +## Motivation + +Gathering baremetal metrics is an important aspect of observability in cloud +systems. It allows the cloud provider to make informed decisions on cost +estimates and abnormal hardware behavior. + +If the cloud provider decides to expose baremetal metrics to tenants, +it would allow them to also make informed decisions on cost estimates +and essential hardware metrics if they should use it for their projects. + +### User Stories + +* As a provider, I want to easily deploy and undeploy bare metal + observability using an automation application like Ansible + or an operator. +* As a provider, I want to be able to access a metric endpoint so I + can connect it to Prometheus compatible frontend applications and + stacks. +* As a provider, I want my metrics to come with labels about the + originating clusters and nodes so I can build my own RBAC system + to the metrics for the tenants. +* As a provider, I want to easily include another metric exporter if + I decide I want to use something other than Prometheus's IPMI and + SNMP exporters. + +### Goals + +This will be a success if this can be deployed as an optional feature +in the OSAC installation process. The implementation will rely on ESI +and OpenStack for retrieving information on where to get baremetal +metrics. This will be using Prometheus compatible applications +because the Prometheus is built-in to OpenShift. + +This will be a success if there is minimal to no communication needed +between the tenant and the provider regarding additional observability +management and configuration. + +### Non-Goals + +* We will not be looking into how to use metrics for billing purposes. +* We will not be looking into proxies/applications that would usually + go on top of the metric API endpoint. +* We will not be looking into collecting from in-OS exporters in bare + metal clusters. This may call for a later enhancement. +* We will not be looking into fine-grain access to metrics, but we + can implement metric labeling with cluster and node labels now to + help with implementing fine-grain access later. This will call for + a later enhancement. + +## Proposal + +The proposed implementation makes use of the [multi-target exporter pattern](https://prometheus.io/docs/guides/multi-target-exporter/#the-multi-target-exporter-pattern), +notably the [IPMI-exporter](https://github.com/prometheus-community/ipmi_exporter) +and the [SNMP-exporter](https://github.com/prometheus/snmp_exporter). +Multi-target exporters have properties perfect for our goals: +* the exporter does not have to run on the machine the metrics are taken from +* the exporter will get the target’s metrics via a network protocol +* the exporter can query multiple targets + +The following items describe the key points of the proposal: + +**Devices will be scraped remotely** + +The first proposed change comes from the idea that the provider +should not expect any tenant to run any metric exporter. Therefore, +the proposed implementation has the exporters to be ran on the hub +cluster and perform remote scrapes. + +**Discoverable** + +The nodes that host the exporters must be on the same network as the +devices that the exporters perform the scrape on. The way this might +happen might be through the fulfillment service. The nodes given to +the tenants should have either minimal or no direct read/write access +to the devices. + +**Distinguishable** + +The scraped metrics will need to be labeled with a cluster and node +identifier so that the provider can identify which metric belongs to +which cluster, node, and tenant, and the tenant can identify which node +each of their metric belongs to. + + +### Workflow Description + +**Deploying Baremetal Observability** +1. The cloud provider runs an ansible playbook with an argument that + indicates the want to deploy. +2. The resources needed for the baremetal metric collection system + gets deployed + +**Removing Baremetal Observability** +1. The cloud provider runs an ansible playbook with an argument that + indicates the want to remove. +2. The resources needed for the baremetal metric collection system + gets removed + +### API Extensions + +N/A + +### Implementation Details/Notes/Constraints + +For simplicity, this section will only talk about Prometheus's IPMI +exporter, but we also plan on deploying Prometheus's SNMP exporter, +and more exporters can easily be added aswell. + +**Exporter constraints**: + +Operational metrics are metrics that are scraped from exporters +running on the same host. Operational metrics from baremetal clusters +cannot be trusted because the tenant can theoretically supply fake +metrics to the cloud provider. + +**AAP**: + +The installation process will be described by a bunch of tasks in an +Ansible playbook. The playbook will need to be repeated any time a +new node gets added to the cloud provider's node pool. The task of +setting up an IPMI exporter to scrape metrics from all ESI nodes +looks like this: + +1. Connect the hub nodes to the BMC network +2. Get all IPMI IPs and ports to scrape from OpenStack +3. Create an IPMI exporter deployment on the hub cluster configured + with the username and password required to access the BMCs of each + node +4. Create an IPMI exporter service +5. Create an IPMI exporter ServiceMonitor connected to the IPMI + exporter service and configured to scrape the IPs and ports, and + discoverable by openshift-monitoring's Prometheus pods + +**Access to Metrics**: + +Access to metrics can be done through Prometheus or the Thanos +Querier (also the Observatorium API). Providers are able to +deploy their own Prometheus API compatible applications if they +choose to do so. + +### Risks and Mitigations + +Metrics will have to be proxied by some sort of fine-grain +access system to comply with privacy laws such as GDPR. + +### Drawbacks + +N/A + +## Alternatives (Not Implemented) + +An alternative for tenant OpenShift clusters is to have exporters +exist as a daemonset on tenant nodes and the metrics will be seen by +openshift-monitoring and then open-cluster-management-observability. +However, this means that the tenant could potentially supply any kind +of metric they want, and this would require two different +implementations for tenant OpenShift clusters and baremetal clusters. + +There are alternative metric collection projects out there, but +this will use Prometheus because it is part of OpenShift, and there +are a vast amount of supported Prometheus compatible exporters and +applications. + +## Open Questions [optional] + +N/A + +## Test Plan + +N/A + +## Graduation Criteria + +N/A + +### Removing a deprecated feature + +N/A + +## Upgrade / Downgrade Strategy + +N/A + +## Version Skew Strategy + +N/A + +## Support Procedures + +N/A + +## Infrastructure Needed [optional] + +N/A + From 1b300c2347cd76dbe0880d29a141ca3da182fab9 Mon Sep 17 00:00:00 2001 From: Austin Jamias Date: Wed, 24 Sep 2025 11:05:18 -0400 Subject: [PATCH 02/10] Fix GitHub Actions pre-commit Error --- .../baremetal-observability/README.md | 121 +++++++++--------- 1 file changed, 60 insertions(+), 61 deletions(-) diff --git a/enhancements/baremetal-observability/README.md b/enhancements/baremetal-observability/README.md index d9aa701f8..1561ac62b 100644 --- a/enhancements/baremetal-observability/README.md +++ b/enhancements/baremetal-observability/README.md @@ -18,63 +18,63 @@ superseded-by: ## Summary -We want a way to gather bare metal metrics from tenant clusters and have them -available at some configurable endpoint for the cloud provider to manage. This +We want a way to gather bare metal metrics from tenant clusters and have them +available at some configurable endpoint for the cloud provider to manage. This feature will be easily toggleable per tenant cluster. ## Motivation -Gathering baremetal metrics is an important aspect of observability in cloud -systems. It allows the cloud provider to make informed decisions on cost +Gathering baremetal metrics is an important aspect of observability in cloud +systems. It allows the cloud provider to make informed decisions on cost estimates and abnormal hardware behavior. -If the cloud provider decides to expose baremetal metrics to tenants, -it would allow them to also make informed decisions on cost estimates +If the cloud provider decides to expose baremetal metrics to tenants, +it would allow them to also make informed decisions on cost estimates and essential hardware metrics if they should use it for their projects. ### User Stories -* As a provider, I want to easily deploy and undeploy bare metal - observability using an automation application like Ansible +* As a provider, I want to easily deploy and undeploy bare metal + observability using an automation application like Ansible or an operator. -* As a provider, I want to be able to access a metric endpoint so I +* As a provider, I want to be able to access a metric endpoint so I can connect it to Prometheus compatible frontend applications and stacks. * As a provider, I want my metrics to come with labels about the - originating clusters and nodes so I can build my own RBAC system + originating clusters and nodes so I can build my own RBAC system to the metrics for the tenants. -* As a provider, I want to easily include another metric exporter if - I decide I want to use something other than Prometheus's IPMI and +* As a provider, I want to easily include another metric exporter if + I decide I want to use something other than Prometheus's IPMI and SNMP exporters. ### Goals -This will be a success if this can be deployed as an optional feature -in the OSAC installation process. The implementation will rely on ESI -and OpenStack for retrieving information on where to get baremetal -metrics. This will be using Prometheus compatible applications +This will be a success if this can be deployed as an optional feature +in the OSAC installation process. The implementation will rely on ESI +and OpenStack for retrieving information on where to get baremetal +metrics. This will be using Prometheus compatible applications because the Prometheus is built-in to OpenShift. -This will be a success if there is minimal to no communication needed -between the tenant and the provider regarding additional observability +This will be a success if there is minimal to no communication needed +between the tenant and the provider regarding additional observability management and configuration. ### Non-Goals * We will not be looking into how to use metrics for billing purposes. -* We will not be looking into proxies/applications that would usually +* We will not be looking into proxies/applications that would usually go on top of the metric API endpoint. -* We will not be looking into collecting from in-OS exporters in bare +* We will not be looking into collecting from in-OS exporters in bare metal clusters. This may call for a later enhancement. -* We will not be looking into fine-grain access to metrics, but we - can implement metric labeling with cluster and node labels now to - help with implementing fine-grain access later. This will call for +* We will not be looking into fine-grain access to metrics, but we + can implement metric labeling with cluster and node labels now to + help with implementing fine-grain access later. This will call for a later enhancement. ## Proposal The proposed implementation makes use of the [multi-target exporter pattern](https://prometheus.io/docs/guides/multi-target-exporter/#the-multi-target-exporter-pattern), -notably the [IPMI-exporter](https://github.com/prometheus-community/ipmi_exporter) +notably the [IPMI-exporter](https://github.com/prometheus-community/ipmi_exporter) and the [SNMP-exporter](https://github.com/prometheus/snmp_exporter). Multi-target exporters have properties perfect for our goals: * the exporter does not have to run on the machine the metrics are taken from @@ -85,39 +85,39 @@ The following items describe the key points of the proposal: **Devices will be scraped remotely** -The first proposed change comes from the idea that the provider -should not expect any tenant to run any metric exporter. Therefore, +The first proposed change comes from the idea that the provider +should not expect any tenant to run any metric exporter. Therefore, the proposed implementation has the exporters to be ran on the hub cluster and perform remote scrapes. **Discoverable** -The nodes that host the exporters must be on the same network as the -devices that the exporters perform the scrape on. The way this might -happen might be through the fulfillment service. The nodes given to -the tenants should have either minimal or no direct read/write access +The nodes that host the exporters must be on the same network as the +devices that the exporters perform the scrape on. The way this might +happen might be through the fulfillment service. The nodes given to +the tenants should have either minimal or no direct read/write access to the devices. **Distinguishable** The scraped metrics will need to be labeled with a cluster and node -identifier so that the provider can identify which metric belongs to -which cluster, node, and tenant, and the tenant can identify which node +identifier so that the provider can identify which metric belongs to +which cluster, node, and tenant, and the tenant can identify which node each of their metric belongs to. ### Workflow Description **Deploying Baremetal Observability** -1. The cloud provider runs an ansible playbook with an argument that +1. The cloud provider runs an ansible playbook with an argument that indicates the want to deploy. -2. The resources needed for the baremetal metric collection system +2. The resources needed for the baremetal metric collection system gets deployed **Removing Baremetal Observability** -1. The cloud provider runs an ansible playbook with an argument that +1. The cloud provider runs an ansible playbook with an argument that indicates the want to remove. -2. The resources needed for the baremetal metric collection system +2. The resources needed for the baremetal metric collection system gets removed ### API Extensions @@ -126,45 +126,45 @@ N/A ### Implementation Details/Notes/Constraints -For simplicity, this section will only talk about Prometheus's IPMI +For simplicity, this section will only talk about Prometheus's IPMI exporter, but we also plan on deploying Prometheus's SNMP exporter, and more exporters can easily be added aswell. **Exporter constraints**: -Operational metrics are metrics that are scraped from exporters -running on the same host. Operational metrics from baremetal clusters +Operational metrics are metrics that are scraped from exporters +running on the same host. Operational metrics from baremetal clusters cannot be trusted because the tenant can theoretically supply fake metrics to the cloud provider. -**AAP**: +**AAP**: -The installation process will be described by a bunch of tasks in an -Ansible playbook. The playbook will need to be repeated any time a -new node gets added to the cloud provider's node pool. The task of -setting up an IPMI exporter to scrape metrics from all ESI nodes +The installation process will be described by a bunch of tasks in an +Ansible playbook. The playbook will need to be repeated any time a +new node gets added to the cloud provider's node pool. The task of +setting up an IPMI exporter to scrape metrics from all ESI nodes looks like this: 1. Connect the hub nodes to the BMC network 2. Get all IPMI IPs and ports to scrape from OpenStack -3. Create an IPMI exporter deployment on the hub cluster configured - with the username and password required to access the BMCs of each +3. Create an IPMI exporter deployment on the hub cluster configured + with the username and password required to access the BMCs of each node 4. Create an IPMI exporter service -5. Create an IPMI exporter ServiceMonitor connected to the IPMI - exporter service and configured to scrape the IPs and ports, and +5. Create an IPMI exporter ServiceMonitor connected to the IPMI + exporter service and configured to scrape the IPs and ports, and discoverable by openshift-monitoring's Prometheus pods **Access to Metrics**: -Access to metrics can be done through Prometheus or the Thanos -Querier (also the Observatorium API). Providers are able to -deploy their own Prometheus API compatible applications if they +Access to metrics can be done through Prometheus or the Thanos +Querier (also the Observatorium API). Providers are able to +deploy their own Prometheus API compatible applications if they choose to do so. ### Risks and Mitigations -Metrics will have to be proxied by some sort of fine-grain +Metrics will have to be proxied by some sort of fine-grain access system to comply with privacy laws such as GDPR. ### Drawbacks @@ -173,16 +173,16 @@ N/A ## Alternatives (Not Implemented) -An alternative for tenant OpenShift clusters is to have exporters -exist as a daemonset on tenant nodes and the metrics will be seen by -openshift-monitoring and then open-cluster-management-observability. -However, this means that the tenant could potentially supply any kind -of metric they want, and this would require two different +An alternative for tenant OpenShift clusters is to have exporters +exist as a daemonset on tenant nodes and the metrics will be seen by +openshift-monitoring and then open-cluster-management-observability. +However, this means that the tenant could potentially supply any kind +of metric they want, and this would require two different implementations for tenant OpenShift clusters and baremetal clusters. -There are alternative metric collection projects out there, but -this will use Prometheus because it is part of OpenShift, and there -are a vast amount of supported Prometheus compatible exporters and +There are alternative metric collection projects out there, but +this will use Prometheus because it is part of OpenShift, and there +are a vast amount of supported Prometheus compatible exporters and applications. ## Open Questions [optional] @@ -216,4 +216,3 @@ N/A ## Infrastructure Needed [optional] N/A - From 414ac1f939bc3275986664d22fad900eb23d52ca Mon Sep 17 00:00:00 2001 From: Austin Jamias Date: Tue, 30 Sep 2025 15:40:07 -0400 Subject: [PATCH 03/10] 1st Feedback Updates --- .../baremetal-observability/README.md | 158 ++++++++++++------ 1 file changed, 104 insertions(+), 54 deletions(-) diff --git a/enhancements/baremetal-observability/README.md b/enhancements/baremetal-observability/README.md index 1561ac62b..bd1423459 100644 --- a/enhancements/baremetal-observability/README.md +++ b/enhancements/baremetal-observability/README.md @@ -3,7 +3,7 @@ title: bare-metal-observability authors: - Austin Jamias creation-date: 2025-09-24 -last-updated: 2025-09-24 +last-updated: 2025-09-30 tracking-link: # link to the tracking ticket (for example: Github issue) that corresponds to this enhancement - N/A see-also: @@ -20,7 +20,7 @@ superseded-by: We want a way to gather bare metal metrics from tenant clusters and have them available at some configurable endpoint for the cloud provider to manage. This -feature will be easily toggleable per tenant cluster. +feature will provide an interface compatible to many inventory services. ## Motivation @@ -35,29 +35,29 @@ and essential hardware metrics if they should use it for their projects. ### User Stories * As a provider, I want to easily deploy and undeploy bare metal - observability using an automation application like Ansible - or an operator. + observability using Ansible playbooks. * As a provider, I want to be able to access a metric endpoint so I can connect it to Prometheus compatible frontend applications and stacks. * As a provider, I want my metrics to come with labels about the originating clusters and nodes so I can build my own RBAC system to the metrics for the tenants. -* As a provider, I want to easily include another metric exporter if - I decide I want to use something other than Prometheus's IPMI and - SNMP exporters. +* As a provider, I want to easily include Ansible roles for other + metric exporters if I decide I want to use something other than + Prometheus's IPMI and SNMP exporters. +* As a provider, I want to easily swap different Ansible roles that + deploy using different inventory services. ### Goals This will be a success if this can be deployed as an optional feature -in the OSAC installation process. The implementation will rely on ESI -and OpenStack for retrieving information on where to get baremetal -metrics. This will be using Prometheus compatible applications -because the Prometheus is built-in to OpenShift. - -This will be a success if there is minimal to no communication needed -between the tenant and the provider regarding additional observability -management and configuration. +in the OSAC installation process. The implementation will rely on +Ansible roles providing an interface that can work with any inventory +service for retrieving information on where to get baremetal metrics. +The implementation will also rely on Prometheus compatible +applications because the Prometheus is already built in to OpenShift +(openshift-monitoring) and ACM observability +(open-cluster-management-observability). ### Non-Goals @@ -67,15 +67,15 @@ management and configuration. * We will not be looking into collecting from in-OS exporters in bare metal clusters. This may call for a later enhancement. * We will not be looking into fine-grain access to metrics, but we - can implement metric labeling with cluster and node labels now to - help with implementing fine-grain access later. This will call for - a later enhancement. + can implement metric labeling node labels now to help with + implementing fine-grain access later. This will call for a later + enhancement. ## Proposal The proposed implementation makes use of the [multi-target exporter pattern](https://prometheus.io/docs/guides/multi-target-exporter/#the-multi-target-exporter-pattern), notably the [IPMI-exporter](https://github.com/prometheus-community/ipmi_exporter) -and the [SNMP-exporter](https://github.com/prometheus/snmp_exporter). +and the [SNMP-exporter](https://github.com/prometheus/snmp_exporter) Multi-target exporters have properties perfect for our goals: * the exporter does not have to run on the machine the metrics are taken from * the exporter will get the target’s metrics via a network protocol @@ -87,38 +87,47 @@ The following items describe the key points of the proposal: The first proposed change comes from the idea that the provider should not expect any tenant to run any metric exporter. Therefore, -the proposed implementation has the exporters to be ran on the hub -cluster and perform remote scrapes. +the proposed implementation has the exporters to be ran on the ACM +hub cluster and perform remote scrapes. **Discoverable** -The nodes that host the exporters must be on the same network as the -devices that the exporters perform the scrape on. The way this might -happen might be through the fulfillment service. The nodes given to -the tenants should have either minimal or no direct read/write access -to the devices. +The nodes that host the exporters (the hub cluster) must be on the +same network as the BMCs (baseboard management controller). The nodes +given to the tenants should have either minimal or no direct +read/write access to the BMCs. **Distinguishable** -The scraped metrics will need to be labeled with a cluster and node -identifier so that the provider can identify which metric belongs to -which cluster, node, and tenant, and the tenant can identify which node -each of their metric belongs to. - +As of right now, the scraped metrics will need to be labeled with a +node identifier so that the provider can then identify which cluster +and tenant it belongs to. Once the enhancement proposal for the +baremetal fulfillment gets implemented, the implementation can switch +from using a node identifier to a cloudkit.Host and cloudkit.HostPool +resource identifier. ### Workflow Description **Deploying Baremetal Observability** -1. The cloud provider runs an ansible playbook with an argument that - indicates the want to deploy. -2. The resources needed for the baremetal metric collection system - gets deployed +1. The cloud provider navigates to either the Ansible web interface, + or prepares a POST request to send to the Ansible Automation + Platform API. +2. The cloud provider submits a request to run the + baremetal-observability playbook with extra-vars that indicate + the signal to create the system, which inventory service they are + using and which exporters to deploy. +3. The resources needed for the baremetal metric collection system + gets deployed. **Removing Baremetal Observability** -1. The cloud provider runs an ansible playbook with an argument that - indicates the want to remove. -2. The resources needed for the baremetal metric collection system - gets removed +1. The cloud provider navigates to either the Ansible web interface, + or prepares a POST request to send to the Ansible Automation + Platform API. +2. The cloud provider submits a request to run the + baremetal-observability playbook with extra-vars that indicate the + signal to destroy the system. +3. The Ansible job removes everything deployed by the + baremetal-observability template. ### API Extensions @@ -126,41 +135,82 @@ N/A ### Implementation Details/Notes/Constraints -For simplicity, this section will only talk about Prometheus's IPMI -exporter, but we also plan on deploying Prometheus's SNMP exporter, -and more exporters can easily be added aswell. +The directory structure focuses on keeping the inventory service and +types of exporters modular so that the cloud provider can pick and +choose for their needs. An example might look like this: + +```bash +└── roles + ├── baremetal_observability + │   ├── defaults + │   │   └── main.yaml + │   ├── meta + │   │   └── argument_specs.yaml + │   └── tasks + │   ├── create_baremetal_observability.yaml + │   ├── destroy_baremetal_observability.yaml + │   └── main.yaml + │ + │ # inventory services + ├── openstack + │   ├── defaults + │   │   └── main.yaml + │   ├── meta + │   │   └── argument_specs.yaml + │   └── tasks + │   └── main.yaml + │ + │ # metric exporters + ├── ipmi-exporter + │   └── tasks + │   └── main.yaml + ├── snmp-exporter + │ └── tasks + │ └── main.yaml +``` **Exporter constraints**: Operational metrics are metrics that are scraped from exporters running on the same host. Operational metrics from baremetal clusters cannot be trusted because the tenant can theoretically supply fake -metrics to the cloud provider. +metrics to the cloud provider. For now, we assume that all exporters +scrape from the BMC. **AAP**: The installation process will be described by a bunch of tasks in an Ansible playbook. The playbook will need to be repeated any time a -new node gets added to the cloud provider's node pool. The task of -setting up an IPMI exporter to scrape metrics from all ESI nodes +new node gets added to the cloud provider's node pool. To keep this +playbook flexible, the user will need to provide the following +parameters in the extra-vars flag when executing the playbook: + +* string bmo\_state: Either "present" or "absent". Set to "present" + to deploy baremetal-observability, set to "absent" to undeploy. +* string bmo\_inventory\_service: Set it to the name of the Ansible + role that will retrieve relevant information for the exporters +* string[] bmo\_exporters: Set it to a list of metric exporters you + wish to deploy + +The task of setting up exporters to scrape metrics from all nodes looks like this: 1. Connect the hub nodes to the BMC network -2. Get all IPMI IPs and ports to scrape from OpenStack -3. Create an IPMI exporter deployment on the hub cluster configured - with the username and password required to access the BMCs of each - node -4. Create an IPMI exporter service +2. Get all BMC IPs and ports to scrape from the inventory service +3. Create a k8s.Deployment for each exporter in bmo\_exporters on the + hub cluster configured with the username and password required to + access the BMCs of each node +4. Create a k8s.Service for each exporter 5. Create an IPMI exporter ServiceMonitor connected to the IPMI exporter service and configured to scrape the IPs and ports, and discoverable by openshift-monitoring's Prometheus pods **Access to Metrics**: -Access to metrics can be done through Prometheus or the Thanos -Querier (also the Observatorium API). Providers are able to -deploy their own Prometheus API compatible applications if they -choose to do so. +The cloud provider can access metrics through Prometheus or the +Thanos Querier (also the Observatorium API). Providers are able to +deploy their own Prometheus API compatible applications for tenants +to access metrics if they choose to do so. ### Risks and Mitigations From adcd086e630e85f559d9d27e732cf2f773fd64f0 Mon Sep 17 00:00:00 2001 From: Austin Jamias Date: Tue, 30 Sep 2025 16:05:44 -0400 Subject: [PATCH 04/10] Add Link to Implementation Repo --- enhancements/baremetal-observability/README.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/enhancements/baremetal-observability/README.md b/enhancements/baremetal-observability/README.md index bd1423459..ebf3a6e94 100644 --- a/enhancements/baremetal-observability/README.md +++ b/enhancements/baremetal-observability/README.md @@ -7,7 +7,7 @@ last-updated: 2025-09-30 tracking-link: # link to the tracking ticket (for example: Github issue) that corresponds to this enhancement - N/A see-also: - - N/A + - https://github.com/ajamias/baremetal-observability-aap replaces: - N/A superseded-by: From 2c3fdbfa51526c0c2e02b4dc913215b037895466 Mon Sep 17 00:00:00 2001 From: Austin Jamias <94348154+ajamias@users.noreply.github.com> Date: Tue, 14 Oct 2025 13:57:33 -0400 Subject: [PATCH 05/10] Update enhancements/baremetal-observability/README.md Co-authored-by: Lars Kellogg-Stedman --- enhancements/baremetal-observability/README.md | 3 +-- 1 file changed, 1 insertion(+), 2 deletions(-) diff --git a/enhancements/baremetal-observability/README.md b/enhancements/baremetal-observability/README.md index ebf3a6e94..0fd91d17c 100644 --- a/enhancements/baremetal-observability/README.md +++ b/enhancements/baremetal-observability/README.md @@ -94,8 +94,7 @@ hub cluster and perform remote scrapes. The nodes that host the exporters (the hub cluster) must be on the same network as the BMCs (baseboard management controller). The nodes -given to the tenants should have either minimal or no direct -read/write access to the BMCs. +given to the tenants should have no access to the BMCs. **Distinguishable** From 03f302693116b6a50237a31b8da454959d919db1 Mon Sep 17 00:00:00 2001 From: Austin Jamias <94348154+ajamias@users.noreply.github.com> Date: Tue, 14 Oct 2025 13:57:48 -0400 Subject: [PATCH 06/10] Update enhancements/baremetal-observability/README.md Co-authored-by: Lars Kellogg-Stedman --- enhancements/baremetal-observability/README.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/enhancements/baremetal-observability/README.md b/enhancements/baremetal-observability/README.md index 0fd91d17c..fff7bb3df 100644 --- a/enhancements/baremetal-observability/README.md +++ b/enhancements/baremetal-observability/README.md @@ -98,7 +98,7 @@ given to the tenants should have no access to the BMCs. **Distinguishable** -As of right now, the scraped metrics will need to be labeled with a +Initially, the scraped metrics will need to be labeled with a node identifier so that the provider can then identify which cluster and tenant it belongs to. Once the enhancement proposal for the baremetal fulfillment gets implemented, the implementation can switch From 8e39723eb5b7886677ce349c84e352f2d5b93db1 Mon Sep 17 00:00:00 2001 From: Austin Jamias <94348154+ajamias@users.noreply.github.com> Date: Tue, 14 Oct 2025 13:57:59 -0400 Subject: [PATCH 07/10] Update enhancements/baremetal-observability/README.md Co-authored-by: Lars Kellogg-Stedman --- enhancements/baremetal-observability/README.md | 7 +++---- 1 file changed, 3 insertions(+), 4 deletions(-) diff --git a/enhancements/baremetal-observability/README.md b/enhancements/baremetal-observability/README.md index fff7bb3df..86555d6db 100644 --- a/enhancements/baremetal-observability/README.md +++ b/enhancements/baremetal-observability/README.md @@ -111,10 +111,9 @@ resource identifier. 1. The cloud provider navigates to either the Ansible web interface, or prepares a POST request to send to the Ansible Automation Platform API. -2. The cloud provider submits a request to run the - baremetal-observability playbook with extra-vars that indicate - the signal to create the system, which inventory service they are - using and which exporters to deploy. +2. The cloud provider submits a request to run the baremetal-observability + playbook with extra-vars that instruct the playbook to create the system, + which inventory service they are using and which exporters to deploy. 3. The resources needed for the baremetal metric collection system gets deployed. From 052290f3aaeebb865a4546dd9599eaa60600099a Mon Sep 17 00:00:00 2001 From: Austin Jamias <94348154+ajamias@users.noreply.github.com> Date: Tue, 14 Oct 2025 13:58:17 -0400 Subject: [PATCH 08/10] Update enhancements/baremetal-observability/README.md Co-authored-by: Lars Kellogg-Stedman --- enhancements/baremetal-observability/README.md | 6 +++--- 1 file changed, 3 insertions(+), 3 deletions(-) diff --git a/enhancements/baremetal-observability/README.md b/enhancements/baremetal-observability/README.md index 86555d6db..3cc633159 100644 --- a/enhancements/baremetal-observability/README.md +++ b/enhancements/baremetal-observability/README.md @@ -183,11 +183,11 @@ new node gets added to the cloud provider's node pool. To keep this playbook flexible, the user will need to provide the following parameters in the extra-vars flag when executing the playbook: -* string bmo\_state: Either "present" or "absent". Set to "present" +* string `bmo_state`: Either "present" or "absent". Set to "present" to deploy baremetal-observability, set to "absent" to undeploy. -* string bmo\_inventory\_service: Set it to the name of the Ansible +* string `bmo_inventory_service`: Set it to the name of the Ansible role that will retrieve relevant information for the exporters -* string[] bmo\_exporters: Set it to a list of metric exporters you +* string[] `bmo_exporters`: Set it to a list of metric exporters you wish to deploy The task of setting up exporters to scrape metrics from all nodes From ea631c049cc0e9390950c13e4b5cac657e8004e7 Mon Sep 17 00:00:00 2001 From: Austin Jamias <94348154+ajamias@users.noreply.github.com> Date: Tue, 14 Oct 2025 13:58:28 -0400 Subject: [PATCH 09/10] Update enhancements/baremetal-observability/README.md Co-authored-by: Lars Kellogg-Stedman --- enhancements/baremetal-observability/README.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/enhancements/baremetal-observability/README.md b/enhancements/baremetal-observability/README.md index 3cc633159..e2f755000 100644 --- a/enhancements/baremetal-observability/README.md +++ b/enhancements/baremetal-observability/README.md @@ -193,7 +193,7 @@ parameters in the extra-vars flag when executing the playbook: The task of setting up exporters to scrape metrics from all nodes looks like this: -1. Connect the hub nodes to the BMC network +1. Grant the hub nodes access to the BMC network 2. Get all BMC IPs and ports to scrape from the inventory service 3. Create a k8s.Deployment for each exporter in bmo\_exporters on the hub cluster configured with the username and password required to From 8515ac1b6e96e329b078fd34cbbd429f8cc7bd2a Mon Sep 17 00:00:00 2001 From: Austin Jamias <94348154+ajamias@users.noreply.github.com> Date: Tue, 14 Oct 2025 13:58:43 -0400 Subject: [PATCH 10/10] Update enhancements/baremetal-observability/README.md Co-authored-by: Lars Kellogg-Stedman --- enhancements/baremetal-observability/README.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/enhancements/baremetal-observability/README.md b/enhancements/baremetal-observability/README.md index e2f755000..e1f64a220 100644 --- a/enhancements/baremetal-observability/README.md +++ b/enhancements/baremetal-observability/README.md @@ -195,7 +195,7 @@ looks like this: 1. Grant the hub nodes access to the BMC network 2. Get all BMC IPs and ports to scrape from the inventory service -3. Create a k8s.Deployment for each exporter in bmo\_exporters on the +3. Create a k8s.Deployment for each exporter in `bmo_exporters` on the hub cluster configured with the username and password required to access the BMCs of each node 4. Create a k8s.Service for each exporter