Skip to content
266 changes: 266 additions & 0 deletions enhancements/baremetal-observability/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,266 @@
---
title: bare-metal-observability
authors:
- Austin Jamias
creation-date: 2025-09-24
last-updated: 2025-09-30
tracking-link: # link to the tracking ticket (for example: Github issue) that corresponds to this enhancement
- N/A
see-also:
- https://github.com/ajamias/baremetal-observability-aap
replaces:
- N/A
superseded-by:
- N/A
---

# Bare Metal Observability

## Summary

We want a way to gather bare metal metrics from tenant clusters and have them
available at some configurable endpoint for the cloud provider to manage. This
feature will provide an interface compatible to many inventory services.

## Motivation

Gathering baremetal metrics is an important aspect of observability in cloud
systems. It allows the cloud provider to make informed decisions on cost
estimates and abnormal hardware behavior.

If the cloud provider decides to expose baremetal metrics to tenants,
it would allow them to also make informed decisions on cost estimates
and essential hardware metrics if they should use it for their projects.

### User Stories

* As a provider, I want to easily deploy and undeploy bare metal
observability using Ansible playbooks.
* As a provider, I want to be able to access a metric endpoint so I
can connect it to Prometheus compatible frontend applications and
stacks.
* As a provider, I want my metrics to come with labels about the
originating clusters and nodes so I can build my own RBAC system
to the metrics for the tenants.
* As a provider, I want to easily include Ansible roles for other
metric exporters if I decide I want to use something other than
Prometheus's IPMI and SNMP exporters.
* As a provider, I want to easily swap different Ansible roles that
deploy using different inventory services.

### Goals

This will be a success if this can be deployed as an optional feature
in the OSAC installation process. The implementation will rely on
Ansible roles providing an interface that can work with any inventory
service for retrieving information on where to get baremetal metrics.
The implementation will also rely on Prometheus compatible
applications because the Prometheus is already built in to OpenShift
(openshift-monitoring) and ACM observability
(open-cluster-management-observability).

### Non-Goals

* We will not be looking into how to use metrics for billing purposes.
* We will not be looking into proxies/applications that would usually
go on top of the metric API endpoint.
* We will not be looking into collecting from in-OS exporters in bare
metal clusters. This may call for a later enhancement.
* We will not be looking into fine-grain access to metrics, but we
can implement metric labeling node labels now to help with
implementing fine-grain access later. This will call for a later
enhancement.

## Proposal

The proposed implementation makes use of the [multi-target exporter pattern](https://prometheus.io/docs/guides/multi-target-exporter/#the-multi-target-exporter-pattern),
notably the [IPMI-exporter](https://github.com/prometheus-community/ipmi_exporter)
and the [SNMP-exporter](https://github.com/prometheus/snmp_exporter)
Multi-target exporters have properties perfect for our goals:
* the exporter does not have to run on the machine the metrics are taken from
* the exporter will get the target’s metrics via a network protocol
* the exporter can query multiple targets
Comment thread
ajamias marked this conversation as resolved.

The following items describe the key points of the proposal:

**Devices will be scraped remotely**

The first proposed change comes from the idea that the provider
should not expect any tenant to run any metric exporter. Therefore,
the proposed implementation has the exporters to be ran on the ACM
hub cluster and perform remote scrapes.

**Discoverable**

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This heading doesn't seem to fit the content of the section: you're describing how the devices will be discovered; you're describing some network requirements.


The nodes that host the exporters (the hub cluster) must be on the
same network as the BMCs (baseboard management controller). The nodes
Comment on lines +95 to +96

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There is no requirement that the hub cluster must be on the same network as the BMCs. The requirement is that the hub cluster (or where we are running the collectors) has access to the BMC network. This could be via direct attachment, but it could also be via routed access, vpn, etc.

given to the tenants should have no access to the BMCs.

**Distinguishable**

Initially, the scraped metrics will need to be labeled with a
node identifier so that the provider can then identify which cluster
and tenant it belongs to. Once the enhancement proposal for the
baremetal fulfillment gets implemented, the implementation can switch
from using a node identifier to a cloudkit.Host and cloudkit.HostPool
resource identifier.

### Workflow Description

**Deploying Baremetal Observability**
1. The cloud provider navigates to either the Ansible web interface,
or prepares a POST request to send to the Ansible Automation
Platform API.
2. The cloud provider submits a request to run the baremetal-observability
playbook with extra-vars that instruct the playbook to create the system,
which inventory service they are using and which exporters to deploy.
3. The resources needed for the baremetal metric collection system
gets deployed.

**Removing Baremetal Observability**
1. The cloud provider navigates to either the Ansible web interface,
or prepares a POST request to send to the Ansible Automation
Platform API.
2. The cloud provider submits a request to run the
baremetal-observability playbook with extra-vars that indicate the
signal to destroy the system.
3. The Ansible job removes everything deployed by the
baremetal-observability template.

### API Extensions

N/A

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@larsks @tzumainn @ajamias Are there any OpenStack webhooks that need to be created to kick off some Event Driven Ansible to update the IPMI exporter when nodes are added or deleted to the pool of nodes?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm not sure I understand - how would the IPMI exproter need to be updated?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I talked to Mainn asking if we would want event driven Ansible or have the user manually run Ansible jobs from a template to update the configs whenever a node would be added or removed from the available nodes. We agreed that the user should manually run the jobs because event driven Ansible would be hard to implement with different inventory services, and the templates could be integrated with Isaiah's templates. Is this right @tzumainn ?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think for testing purposes, it's fine to have the user manually run the Ansible jobs. In the longer-term, I think it makes sense to have these jobs have the option to be called automatically within the bare metal fulfillment process (and one thing that may have changed since we last spoke is that we're currently not planning on using templates during bare metal fulfillment). But I think there are policy questions there - should bare metal metrics always be collected? can it be turned off by the user? - that are worth exploring.

To summarize: I think for testing and proof-of-concept purposes, it's fine to run these playbooks manually for now; eventually we'll need to figure out when/how to automate that configuration.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Users shouldn't manually run Ansible to configure bare metal observability. The fulfillment-service should be able to report the nodes to observe, and there should be an AAP Job based on an AAP Template that can run. The Job would receive the list of tenant nodes and configures the observability on or off automatically. Otherwise there is no point for this proposal if it is a manual process anyway.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

To further expand upon my comment: what may be missing from this document is an explanation of when bare metal observability is active. Does O-SAC always run these playbooks? Is it configurable at the cloud provider level? Is it configurable at the tenant level? It may be worth bringing up these questions.

In terms of implementation steps, I think it makes perfect sense to detail a multi-stage approach. The first step is manual testing; the second is integrating these playbooks within existing bare metal fulfillment workflows.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Does O-SAC always run these playbooks? Is it configurable at the cloud provider level?

Ideally we would always configure the metrics endpoints, even if we're not actually consuming the data. This would make the metrics available to the providers own collection if they opt to go that route.

The difficulty here, of course, is that the scrape configurations will typically need to be customized for specific models and manufacturers, so it's not really something we can activate automatically.

I think the cloud provider gets to configure both (a) whether or not we are exposing bare metal metrics on prometheus-compatible endpoints, and (b) whether or not we are collecting those metrics into some sort of observability environment.

Is it configurable at the tenant level?

No. The provider owns the hardware.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ideally we would always configure the metrics endpoints, even if we're not actually consuming the data. This would make the metrics available to the providers own collection if they opt to go that route.

I think the cloud provider gets to configure both (a) whether or not we are exposing bare metal metrics on prometheus-compatible endpoints, and (b) whether or not we are collecting those metrics into some sort of observability environment.

I think that adding an OpenTelemetry Collector would fit in nicely. The OTEL collector would receive from enabled metric exporters (OLTP, Prometheus, etc) and push them to a selection of storage backends (ACM observability by default in our case), which allows the provider to configure the type of endpoint they want to use for their observability environment.

Does this sound like a better approach for the long term goal?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That sounds like a reasonable approach.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I like the OpenTelemetry Collector idea too!


### Implementation Details/Notes/Constraints

The directory structure focuses on keeping the inventory service and
types of exporters modular so that the cloud provider can pick and
choose for their needs. An example might look like this:

```bash
└── roles
├── baremetal_observability
│   ├── defaults
│   │   └── main.yaml
│   ├── meta
│   │   └── argument_specs.yaml
│   └── tasks
│   ├── create_baremetal_observability.yaml
│   ├── destroy_baremetal_observability.yaml
│   └── main.yaml
│
│ # inventory services
├── openstack
│   ├── defaults
│   │   └── main.yaml
│   ├── meta
│   │   └── argument_specs.yaml
│   └── tasks
│   └── main.yaml
│
│ # metric exporters
├── ipmi-exporter
│   └── tasks
│   └── main.yaml
├── snmp-exporter
│ └── tasks
│ └── main.yaml
```

**Exporter constraints**:

Operational metrics are metrics that are scraped from exporters
running on the same host. Operational metrics from baremetal clusters
cannot be trusted because the tenant can theoretically supply fake
metrics to the cloud provider. For now, we assume that all exporters
scrape from the BMC.

**AAP**:

The installation process will be described by a bunch of tasks in an
Ansible playbook. The playbook will need to be repeated any time a
new node gets added to the cloud provider's node pool. To keep this
playbook flexible, the user will need to provide the following
parameters in the extra-vars flag when executing the playbook:

* string `bmo_state`: Either "present" or "absent". Set to "present"
to deploy baremetal-observability, set to "absent" to undeploy.
* string `bmo_inventory_service`: Set it to the name of the Ansible
role that will retrieve relevant information for the exporters
* string[] `bmo_exporters`: Set it to a list of metric exporters you
wish to deploy

The task of setting up exporters to scrape metrics from all nodes
looks like this:

1. Grant the hub nodes access to the BMC network
2. Get all BMC IPs and ports to scrape from the inventory service
3. Create a k8s.Deployment for each exporter in `bmo_exporters` on the
hub cluster configured with the username and password required to
access the BMCs of each node
4. Create a k8s.Service for each exporter
5. Create an IPMI exporter ServiceMonitor connected to the IPMI
exporter service and configured to scrape the IPs and ports, and
discoverable by openshift-monitoring's Prometheus pods

**Access to Metrics**:

The cloud provider can access metrics through Prometheus or the
Thanos Querier (also the Observatorium API). Providers are able to
deploy their own Prometheus API compatible applications for tenants
to access metrics if they choose to do so.

### Risks and Mitigations

Metrics will have to be proxied by some sort of fine-grain
access system to comply with privacy laws such as GDPR.

### Drawbacks

N/A

## Alternatives (Not Implemented)

An alternative for tenant OpenShift clusters is to have exporters
exist as a daemonset on tenant nodes and the metrics will be seen by
openshift-monitoring and then open-cluster-management-observability.
However, this means that the tenant could potentially supply any kind
of metric they want, and this would require two different
implementations for tenant OpenShift clusters and baremetal clusters.

There are alternative metric collection projects out there, but
this will use Prometheus because it is part of OpenShift, and there
are a vast amount of supported Prometheus compatible exporters and
applications.

## Open Questions [optional]

N/A

## Test Plan

N/A

## Graduation Criteria

N/A

### Removing a deprecated feature

N/A

## Upgrade / Downgrade Strategy

N/A

## Version Skew Strategy

N/A

## Support Procedures

N/A

## Infrastructure Needed [optional]

N/A