Skip to content

Add Baremetal Observability Enhancement - #10

Closed
ajamias wants to merge 10 commits into
osac-project:mainfrom
ajamias:baremetal-observability
Closed

ajamias wants to merge 10 commits into
osac-project:mainfrom
ajamias:baremetal-observability

Conversation

@ajamias

@ajamias ajamias commented Sep 24, 2025

Copy link
Copy Markdown
Contributor

No description provided.

@computate computate left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just fix the spelling suggestions, and I added some additional feedback.

Comment thread enhancements/baremetal-observability/README.md Outdated
Comment thread enhancements/baremetal-observability/README.md Outdated
Comment thread enhancements/baremetal-observability/README.md Outdated
Comment thread enhancements/baremetal-observability/README.md Outdated

### API Extensions

N/A

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@larsks @tzumainn @ajamias Are there any OpenStack webhooks that need to be created to kick off some Event Driven Ansible to update the IPMI exporter when nodes are added or deleted to the pool of nodes?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm not sure I understand - how would the IPMI exproter need to be updated?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I talked to Mainn asking if we would want event driven Ansible or have the user manually run Ansible jobs from a template to update the configs whenever a node would be added or removed from the available nodes. We agreed that the user should manually run the jobs because event driven Ansible would be hard to implement with different inventory services, and the templates could be integrated with Isaiah's templates. Is this right @tzumainn ?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think for testing purposes, it's fine to have the user manually run the Ansible jobs. In the longer-term, I think it makes sense to have these jobs have the option to be called automatically within the bare metal fulfillment process (and one thing that may have changed since we last spoke is that we're currently not planning on using templates during bare metal fulfillment). But I think there are policy questions there - should bare metal metrics always be collected? can it be turned off by the user? - that are worth exploring.

To summarize: I think for testing and proof-of-concept purposes, it's fine to run these playbooks manually for now; eventually we'll need to figure out when/how to automate that configuration.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Users shouldn't manually run Ansible to configure bare metal observability. The fulfillment-service should be able to report the nodes to observe, and there should be an AAP Job based on an AAP Template that can run. The Job would receive the list of tenant nodes and configures the observability on or off automatically. Otherwise there is no point for this proposal if it is a manual process anyway.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

To further expand upon my comment: what may be missing from this document is an explanation of when bare metal observability is active. Does O-SAC always run these playbooks? Is it configurable at the cloud provider level? Is it configurable at the tenant level? It may be worth bringing up these questions.

In terms of implementation steps, I think it makes perfect sense to detail a multi-stage approach. The first step is manual testing; the second is integrating these playbooks within existing bare metal fulfillment workflows.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Does O-SAC always run these playbooks? Is it configurable at the cloud provider level?

Ideally we would always configure the metrics endpoints, even if we're not actually consuming the data. This would make the metrics available to the providers own collection if they opt to go that route.

The difficulty here, of course, is that the scrape configurations will typically need to be customized for specific models and manufacturers, so it's not really something we can activate automatically.

I think the cloud provider gets to configure both (a) whether or not we are exposing bare metal metrics on prometheus-compatible endpoints, and (b) whether or not we are collecting those metrics into some sort of observability environment.

Is it configurable at the tenant level?

No. The provider owns the hardware.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ideally we would always configure the metrics endpoints, even if we're not actually consuming the data. This would make the metrics available to the providers own collection if they opt to go that route.

I think the cloud provider gets to configure both (a) whether or not we are exposing bare metal metrics on prometheus-compatible endpoints, and (b) whether or not we are collecting those metrics into some sort of observability environment.

I think that adding an OpenTelemetry Collector would fit in nicely. The OTEL collector would receive from enabled metric exporters (OLTP, Prometheus, etc) and push them to a selection of storage backends (ACM observability by default in our case), which allows the provider to configure the type of endpoint they want to use for their observability environment.

Does this sound like a better approach for the long term goal?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That sounds like a reasonable approach.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I like the OpenTelemetry Collector idea too!

Comment thread enhancements/baremetal-observability/README.md Outdated
Comment thread enhancements/baremetal-observability/README.md Outdated
Comment on lines +160 to +163
Access to metrics can be done through Prometheus or the Thanos
Querier (also the Observatorium API). Providers are able to
deploy their own Prometheus API compatible applications if they
choose to do so.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The local-cluster Prometheus does not have tenant cluster metrics, so I think we might also mention that the bare metal metrics can be available to ACM Observability, as well as all the managed cluster metrics per cluster label.

The hypershift1 and hypershift2 clusters already have the the ipmi metrics in the allowlist. These are metrics are available to the cloud provider. The tenants will not be able to access bare metal metrics from their cluster unless they build another application to forward the metrics, or use the prom-keycloak-proxy.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You are right that the local-cluster (openshift-monitoring) Prometheus does not have tenant cluster metrics, but all ACM Observability metrics are stored in the local-cluster Prometheus instance. If you query one of the acm metrics, they will have a label with name prometheus and then the value is <namespace>/<pod_name>. It seems that all the ACM metrics are stored in openshift-monitoring/k8s (aka the local-cluster Prometheus).

This proposal also does not focus on how tenants will access data, that will be up to the cloud provider (or a future enhancement proposal).

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Comment thread enhancements/baremetal-observability/README.md Outdated
Comment thread enhancements/baremetal-observability/README.md
Comment thread enhancements/baremetal-observability/README.md Outdated
Comment thread enhancements/baremetal-observability/README.md Outdated

@tzumainn tzumainn left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Left a few comments!

Comment thread enhancements/baremetal-observability/README.md Outdated
Comment thread enhancements/baremetal-observability/README.md Outdated
@ajamias
ajamias force-pushed the baremetal-observability branch from ccc980f to 414ac1f Compare September 30, 2025 20:03
Comment on lines +192 to +193
* string[] bmo\_exporters: Set it to a list of metric exporters you
wish to deploy

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Keep in mind that for k8s operators, if it's first deployed with bmo_state: present, bmo_exporters: [ 'snmp' ], and then later deployed with bmo_state: present, bmo_exporters: [ 'ipmi' ], then snmp exporters would get uninstalled, and ipmi exporters would be installed instead.

Comment on lines +95 to +96
The nodes that host the exporters (the hub cluster) must be on the
same network as the BMCs (baseboard management controller). The nodes

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There is no requirement that the hub cluster must be on the same network as the BMCs. The requirement is that the hub cluster (or where we are running the collectors) has access to the BMC network. This could be via direct attachment, but it could also be via routed access, vpn, etc.

the proposed implementation has the exporters to be ran on the ACM
hub cluster and perform remote scrapes.

**Discoverable**

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This heading doesn't seem to fit the content of the section: you're describing how the devices will be discovered; you're describing some network requirements.

Comment thread enhancements/baremetal-observability/README.md Outdated
Comment thread enhancements/baremetal-observability/README.md Outdated
Comment thread enhancements/baremetal-observability/README.md Outdated

### API Extensions

N/A

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Does O-SAC always run these playbooks? Is it configurable at the cloud provider level?

Ideally we would always configure the metrics endpoints, even if we're not actually consuming the data. This would make the metrics available to the providers own collection if they opt to go that route.

The difficulty here, of course, is that the scrape configurations will typically need to be customized for specific models and manufacturers, so it's not really something we can activate automatically.

I think the cloud provider gets to configure both (a) whether or not we are exposing bare metal metrics on prometheus-compatible endpoints, and (b) whether or not we are collecting those metrics into some sort of observability environment.

Is it configurable at the tenant level?

No. The provider owns the hardware.

Comment thread enhancements/baremetal-observability/README.md Outdated
Comment thread enhancements/baremetal-observability/README.md Outdated
Comment thread enhancements/baremetal-observability/README.md Outdated
ajamias and others added 6 commits October 14, 2025 13:57
Co-authored-by: Lars Kellogg-Stedman <lars@oddbit.com>
Co-authored-by: Lars Kellogg-Stedman <lars@oddbit.com>
Co-authored-by: Lars Kellogg-Stedman <lars@oddbit.com>
Co-authored-by: Lars Kellogg-Stedman <lars@oddbit.com>
Co-authored-by: Lars Kellogg-Stedman <lars@oddbit.com>
Co-authored-by: Lars Kellogg-Stedman <lars@oddbit.com>
@computate

Copy link
Copy Markdown

Why was this closed?

ElayAharoni added a commit to ElayAharoni/enhancement-proposals that referenced this pull request Jun 29, 2026
- Resolve AdminNetworksPage topology view wording contradiction (issue osac-project#3)
- Specify IPv4/IPv6 CIDRs explicitly in FR-6 (issue osac-project#4)
- Pick side drawer pattern for subnet detail display in FR-12 (issue osac-project#5)
- Add Priority field to SecurityGroup rule specification in FR-18 (issue osac-project#6)
- Standardize PublicIP action terminology to 'Release' in FR-28 (issue osac-project#7)
- Document multi-NIC same-VN constraint rationale in FR-34 (issue osac-project#8)
- Align wizard empty-state flow with inline overlay pattern in FR-38 (issue osac-project#9)
- Define Retry action API contract in FR-42 (issue osac-project#10)
- Clarify Subnet endpoints are create/delete only in FR-45 (issue osac-project#11)
- Remove redundant NFR-9 (issue osac-project#12)
- Update Open Question 8.2 wording to match Non-Goals

Assisted-by: Claude Code <noreply@anthropic.com>
Signed-off-by: Elay Aharoni <elayaha@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants