Conversation
computate
left a comment
There was a problem hiding this comment.
Just fix the spelling suggestions, and I added some additional feedback.
|
|
||
| ### API Extensions | ||
|
|
||
| N/A |
There was a problem hiding this comment.
I'm not sure I understand - how would the IPMI exproter need to be updated?
There was a problem hiding this comment.
I talked to Mainn asking if we would want event driven Ansible or have the user manually run Ansible jobs from a template to update the configs whenever a node would be added or removed from the available nodes. We agreed that the user should manually run the jobs because event driven Ansible would be hard to implement with different inventory services, and the templates could be integrated with Isaiah's templates. Is this right @tzumainn ?
There was a problem hiding this comment.
I think for testing purposes, it's fine to have the user manually run the Ansible jobs. In the longer-term, I think it makes sense to have these jobs have the option to be called automatically within the bare metal fulfillment process (and one thing that may have changed since we last spoke is that we're currently not planning on using templates during bare metal fulfillment). But I think there are policy questions there - should bare metal metrics always be collected? can it be turned off by the user? - that are worth exploring.
To summarize: I think for testing and proof-of-concept purposes, it's fine to run these playbooks manually for now; eventually we'll need to figure out when/how to automate that configuration.
There was a problem hiding this comment.
Users shouldn't manually run Ansible to configure bare metal observability. The fulfillment-service should be able to report the nodes to observe, and there should be an AAP Job based on an AAP Template that can run. The Job would receive the list of tenant nodes and configures the observability on or off automatically. Otherwise there is no point for this proposal if it is a manual process anyway.
There was a problem hiding this comment.
To further expand upon my comment: what may be missing from this document is an explanation of when bare metal observability is active. Does O-SAC always run these playbooks? Is it configurable at the cloud provider level? Is it configurable at the tenant level? It may be worth bringing up these questions.
In terms of implementation steps, I think it makes perfect sense to detail a multi-stage approach. The first step is manual testing; the second is integrating these playbooks within existing bare metal fulfillment workflows.
There was a problem hiding this comment.
Does O-SAC always run these playbooks? Is it configurable at the cloud provider level?
Ideally we would always configure the metrics endpoints, even if we're not actually consuming the data. This would make the metrics available to the providers own collection if they opt to go that route.
The difficulty here, of course, is that the scrape configurations will typically need to be customized for specific models and manufacturers, so it's not really something we can activate automatically.
I think the cloud provider gets to configure both (a) whether or not we are exposing bare metal metrics on prometheus-compatible endpoints, and (b) whether or not we are collecting those metrics into some sort of observability environment.
Is it configurable at the tenant level?
No. The provider owns the hardware.
There was a problem hiding this comment.
Ideally we would always configure the metrics endpoints, even if we're not actually consuming the data. This would make the metrics available to the providers own collection if they opt to go that route.
I think the cloud provider gets to configure both (a) whether or not we are exposing bare metal metrics on prometheus-compatible endpoints, and (b) whether or not we are collecting those metrics into some sort of observability environment.
I think that adding an OpenTelemetry Collector would fit in nicely. The OTEL collector would receive from enabled metric exporters (OLTP, Prometheus, etc) and push them to a selection of storage backends (ACM observability by default in our case), which allows the provider to configure the type of endpoint they want to use for their observability environment.
Does this sound like a better approach for the long term goal?
There was a problem hiding this comment.
That sounds like a reasonable approach.
There was a problem hiding this comment.
I like the OpenTelemetry Collector idea too!
| Access to metrics can be done through Prometheus or the Thanos | ||
| Querier (also the Observatorium API). Providers are able to | ||
| deploy their own Prometheus API compatible applications if they | ||
| choose to do so. |
There was a problem hiding this comment.
The local-cluster Prometheus does not have tenant cluster metrics, so I think we might also mention that the bare metal metrics can be available to ACM Observability, as well as all the managed cluster metrics per cluster label.
The hypershift1 and hypershift2 clusters already have the the ipmi metrics in the allowlist. These are metrics are available to the cloud provider. The tenants will not be able to access bare metal metrics from their cluster unless they build another application to forward the metrics, or use the prom-keycloak-proxy.
There was a problem hiding this comment.
You are right that the local-cluster (openshift-monitoring) Prometheus does not have tenant cluster metrics, but all ACM Observability metrics are stored in the local-cluster Prometheus instance. If you query one of the acm metrics, they will have a label with name prometheus and then the value is <namespace>/<pod_name>. It seems that all the ACM metrics are stored in openshift-monitoring/k8s (aka the local-cluster Prometheus).
This proposal also does not focus on how tenants will access data, that will be up to the cloud provider (or a future enhancement proposal).
There was a problem hiding this comment.
- ACM hub local-cluster metrics are in both ACM and the local-cluster prometheus, but this is not true that "all ACM Observability metrics are stored in the local-cluster Prometheus instance". ACM managed cluster metrics are all stored in object storage instead, see: https://github.com/OCP-on-NERC/nerc-ocp-config/blob/main/acm-observability/base/observability.open-cluster-management.io/multiclusterobservabilities/observability/multiclusterobservability.yaml#L18-L20
- The hub cluster metrics allowlist defines custom metrics like
ipmi: https://github.com/OCP-on-NERC/nerc-ocp-config/blob/main/acm-observability/base/core/configmaps/observability-metrics-custom-allowlist/configmap.yaml#L43 - And the user workload monitoring configmap
observability-metrics-custom-allowlistin thesnmp-exporternamespace defines custom SNMP metrics in the MOC: https://github.com/OCP-on-NERC/nerc-ocp-config/blob/main/snmp-exporter/base/files/uwl_metrics_list.yaml
ccc980f to
414ac1f
Compare
| * string[] bmo\_exporters: Set it to a list of metric exporters you | ||
| wish to deploy |
There was a problem hiding this comment.
Keep in mind that for k8s operators, if it's first deployed with bmo_state: present, bmo_exporters: [ 'snmp' ], and then later deployed with bmo_state: present, bmo_exporters: [ 'ipmi' ], then snmp exporters would get uninstalled, and ipmi exporters would be installed instead.
| The nodes that host the exporters (the hub cluster) must be on the | ||
| same network as the BMCs (baseboard management controller). The nodes |
There was a problem hiding this comment.
There is no requirement that the hub cluster must be on the same network as the BMCs. The requirement is that the hub cluster (or where we are running the collectors) has access to the BMC network. This could be via direct attachment, but it could also be via routed access, vpn, etc.
| the proposed implementation has the exporters to be ran on the ACM | ||
| hub cluster and perform remote scrapes. | ||
|
|
||
| **Discoverable** |
There was a problem hiding this comment.
This heading doesn't seem to fit the content of the section: you're describing how the devices will be discovered; you're describing some network requirements.
|
|
||
| ### API Extensions | ||
|
|
||
| N/A |
There was a problem hiding this comment.
Does O-SAC always run these playbooks? Is it configurable at the cloud provider level?
Ideally we would always configure the metrics endpoints, even if we're not actually consuming the data. This would make the metrics available to the providers own collection if they opt to go that route.
The difficulty here, of course, is that the scrape configurations will typically need to be customized for specific models and manufacturers, so it's not really something we can activate automatically.
I think the cloud provider gets to configure both (a) whether or not we are exposing bare metal metrics on prometheus-compatible endpoints, and (b) whether or not we are collecting those metrics into some sort of observability environment.
Is it configurable at the tenant level?
No. The provider owns the hardware.
Co-authored-by: Lars Kellogg-Stedman <lars@oddbit.com>
Co-authored-by: Lars Kellogg-Stedman <lars@oddbit.com>
Co-authored-by: Lars Kellogg-Stedman <lars@oddbit.com>
Co-authored-by: Lars Kellogg-Stedman <lars@oddbit.com>
Co-authored-by: Lars Kellogg-Stedman <lars@oddbit.com>
Co-authored-by: Lars Kellogg-Stedman <lars@oddbit.com>
|
Why was this closed? |
- Resolve AdminNetworksPage topology view wording contradiction (issue osac-project#3) - Specify IPv4/IPv6 CIDRs explicitly in FR-6 (issue osac-project#4) - Pick side drawer pattern for subnet detail display in FR-12 (issue osac-project#5) - Add Priority field to SecurityGroup rule specification in FR-18 (issue osac-project#6) - Standardize PublicIP action terminology to 'Release' in FR-28 (issue osac-project#7) - Document multi-NIC same-VN constraint rationale in FR-34 (issue osac-project#8) - Align wizard empty-state flow with inline overlay pattern in FR-38 (issue osac-project#9) - Define Retry action API contract in FR-42 (issue osac-project#10) - Clarify Subnet endpoints are create/delete only in FR-45 (issue osac-project#11) - Remove redundant NFR-9 (issue osac-project#12) - Update Open Question 8.2 wording to match Non-Goals Assisted-by: Claude Code <noreply@anthropic.com> Signed-off-by: Elay Aharoni <elayaha@gmail.com>
No description provided.