diff --git a/CHANGELOG.md b/CHANGELOG.md index e5b5d3c73686..7028f4fcc711 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -8,6 +8,8 @@ - [ENHANCEMENT] Add Agent Operator Helm quickstart guide (@hjet) +- [ENHANCEMENT] Reorg Agent Operator quickstart guides (@hjet) + - [BUGFIX] Packaging: Use correct user/group env variables in RPM %post script (@simonc6372) - [BUGFIX] Validate logs config when using logs_instance with automatic logging processor (@mapno) diff --git a/docs/operator/custom-resource-quickstart.md b/docs/operator/custom-resource-quickstart.md new file mode 100644 index 000000000000..a0424468fa84 --- /dev/null +++ b/docs/operator/custom-resource-quickstart.md @@ -0,0 +1,394 @@ ++++ +title = "Custom Resource Quickstart" +weight = 120 ++++ +# Grafana Agent Operator Custom Resource Quickstart + +In this guide you'll learn how to deploy [Agent Operator]({{< relref "./_index.md" >}})'s custom resources into your Kubernetes cluster. + +You'll roll out the following custom resources (CRs): + +- A `GrafanaAgent` resource, which discovers one or more `MetricsInstance` and `LogsInstances` resources. +- A `MetricsInstance` resource that defines where to ship collected metrics. Under the hood, this rolls out a Grafana Agent StatefulSet that will scrape and ship metrics to a `remote_write` endpoint. +- A `ServiceMonitor` resource to collect cAdvisor and kubelet metrics. Under the hood, this configures the `MetricsInstance` / Agent StatefulSet. +- A `LogsInstance` resource that defines where to ship collected logs. Under the hood, this rolls out a Grafana Agent DaemonSet that will tail log files on your cluster nodes. +- A `PodLogs` resource to collect container logs from Kubernetes Pods. Under the hood, this configures the`LogsInstance` / Agent DaemonSet. + +To learn more about the custom resources Operator provides and their hierarchy, please consult [Operator architecture]({{< relref "./architecture.md" >}}). + +> **Note:** Agent Operator is currently in beta and its custom resources are subject to change as the project evolves. It currently supports the metrics and logs subsystems of Grafana Agent. Integrations and traces support is coming soon. + +By the end of this guide, you will be scraping and shipping cAdvisor and Kubelet metrics to a Prometheus-compatible metrics endpoint. You'll also be collecting and shipping your Pods' container logs to a Loki-compatible logs endpoint. + +## Prerequisites + +Before you begin, make sure that you have installed Agent Operator into your cluster. You can learn how to do this in: +- [Installing Grafana Agent Operator with Helm]({{< relref "./helm-getting-started.md" >}}) +- [Installing Grafana Agent Operator]({{< relref "./getting-started.md" >}}) + +## Step 1: Deploy GrafanaAgent resource + +In this step you'll roll out a `GrafanaAgent` resource. A `GrafanaAgent` resource discovers `MetricsInstance` and `LogsInstance` resources and defines the Grafana Agent image, Pod requests, limits, affinities, and tolerations. Pod attributes can only be defined at the GrafanaAgent level and are propagated to `MetricsInstance` and `LogsInstance` Pods. To learn more, please see the GrafanaAgent [Custom Resource Definition](https://github.com/grafana/agent/blob/main/production/operator/crds/monitoring.grafana.com_grafanaagents.yaml). + +> **Note:** Due to the variety of possible deployment architectures, the official Agent Operator Helm chart does not provide built-in templates for the custom resources described in this quickstart. These must be configured and deployed manually. However, you are encouraged to template and add the following manifests to your own in-house Helm charts and GitOps flows. + +Roll out the following manifests in your cluster: + +```yaml +apiVersion: monitoring.grafana.com/v1alpha1 +kind: GrafanaAgent +metadata: + name: grafana-agent + namespace: default + labels: + app: grafana-agent +spec: + image: grafana/agent:v0.20.0 + logLevel: info + serviceAccountName: grafana-agent + metrics: + instanceSelector: + matchLabels: + agent: grafana-agent-metrics + externalLabels: + cluster: cloud + + logs: + instanceSelector: + matchLabels: + agent: grafana-agent-logs + +--- + +apiVersion: v1 +kind: ServiceAccount +metadata: + name: grafana-agent + namespace: default + +--- + +apiVersion: rbac.authorization.k8s.io/v1 +kind: ClusterRole +metadata: + name: grafana-agent +rules: +- apiGroups: + - "" + resources: + - nodes + - nodes/proxy + - nodes/metrics + - services + - endpoints + - pods + verbs: + - get + - list + - watch +- apiGroups: + - networking.k8s.io + resources: + - ingresses + verbs: + - get + - list + - watch +- nonResourceURLs: + - /metrics + - /metrics/cadvisor + verbs: + - get + +--- + +apiVersion: rbac.authorization.k8s.io/v1 +kind: ClusterRoleBinding +metadata: + name: grafana-agent +roleRef: + apiGroup: rbac.authorization.k8s.io + kind: ClusterRole + name: grafana-agent +subjects: +- kind: ServiceAccount + name: grafana-agent + namespace: default +``` + +This creates a ServiceAccount, ClusterRole, and ClusterRoleBinding for the GrafanaAgent resource. It also creates a GrafanaAgent resource and specifies an Agent image version. Finally, the GrafanaAgent resource specifies `MetricsInstance` and `LogsInstance` selectors. These search for MetricsInstances and LogsInstances in the same namespace with labels matching `agent: grafana-agent-metrics` and `agent: grafana-agent-logs`, respectively. It also sets a `cluster: cloud` label for all metrics shipped your Prometheus-compatible endpoint. You should change this label to your desired cluster name. + +The full hierarchy of custom resources is as follows: + +- `GrafanaAgent` + - `MetricsInstance` + - `PodMonitor` + - `Probe` + - `ServiceMonitor` + - `LogsInstance` + - `PodLogs` + +Deploying a GrafanaAgent resource on its own will not spin up any Agent Pods. Agent Operator will create Agent Pods once MetricsInstance and LogsIntance resources have been created. In the next step, you'll roll out a `MetricsInstance` resource to scrape cAdvisor and Kubelet metrics and ship these to your Prometheus-compatible metrics endpoint. + +## Step 2: Deploy a MetricsInstance resource + +In this step you'll roll out a MetricsInstance resource. MetricsInstance resources define a `remote_write` sink for metrics and configure one or more selectors to watch for creation and updates to `*Monitor` objects. These objects allow you to define Agent scrape targets via K8s manifests: + +- [ServiceMonitors](https://github.com/prometheus-operator/prometheus-operator/blob/master/Documentation/api.md#servicemonitor) +- [PodMonitors](https://github.com/prometheus-operator/prometheus-operator/blob/master/Documentation/api.md#podmonitor) +- [Probes](https://github.com/prometheus-operator/prometheus-operator/blob/master/Documentation/api.md#probe) + +Roll out the following manifest into your cluster: + +```yaml +apiVersion: monitoring.grafana.com/v1alpha1 +kind: MetricsInstance +metadata: + name: primary + namespace: default + labels: + agent: grafana-agent-metrics +spec: + remoteWrite: + - url: your_remote_write_URL + basicAuth: + username: + name: primary-credentials-metrics + key: username + password: + name: primary-credentials-metrics + key: password + + # Supply an empty namespace selector to look in all namespaces. Remove + # this to only look in the same namespace as the MetricsInstance CR + serviceMonitorNamespaceSelector: {} + serviceMonitorSelector: + matchLabels: + instance: primary + + # Supply an empty namespace selector to look in all namespaces. Remove + # this to only look in the same namespace as the MetricsInstance CR. + podMonitorNamespaceSelector: {} + podMonitorSelector: + matchLabels: + instance: primary + + # Supply an empty namespace selector to look in all namespaces. Remove + # this to only look in the same namespace as the MetricsInstance CR. + probeNamespaceSelector: {} + probeSelector: + matchLabels: + instance: primary +``` + +Be sure to replace the `remote_write` URL and customize the namespace and label configuration as necessary. This will associate itself with the `agent: grafana-agent` GrafanaAgent resource deployed in the previous step, and watch for creation and updates to `*Monitors` monitors with the the `instance: primary` label. + +Once you've rolled out this manifest, create the `basicAuth` credentials using a Kubernetes Secret: + +```yaml +apiVersion: v1 +kind: Secret +metadata: + name: primary-credentials-metrics + namespace: default +stringData: + username: 'your_cloud_prometheus_username' + password: 'your_cloud_prometheus_API_key' +``` + +If you're using Grafana Cloud, you can find your hosted Prometheus endpoint username and password in the [Grafana Cloud Portal](https://grafana.com/profile/org ). You may wish to base64-encode these values yourself. In this case, please use `data` instead of `stringData`. + +Once you've rolled out the `MetricsInstance` and its Secret, you can confirm that the MetricsInstance Agent is up and running with `kubectl get pod`. Since we haven't defined any monitors yet, this Agent will not have any scrape targets defined. In the next step, we'll create scrape targets for the cAdvisor and kubelet endpoints exposed by the `kubelet` service in the cluster. + +## Step 3: Create ServiceMonitors for kubelet and cAdvisor endpoints + +In this step, you'll create ServiceMonitors for kubelet and cAdvisor metrics exposed by the `kubelet` Service. Every node in your cluster exposes kubelet and cadvisor metrics at `/metrics` and `/metrics/cadvisor` respectively. Agent Operator creates a `kubelet` service that exposes these Node endpoints so that they can be scraped using ServiceMonitors. + +To scrape these two endpoints, roll out the following two ServiceMonitors in your cluster: + +- Kubelet ServiceMonitor + +```yaml +apiVersion: monitoring.coreos.com/v1 +kind: ServiceMonitor +metadata: + labels: + instance: primary + name: kubelet-monitor + namespace: default +spec: + endpoints: + - bearerTokenFile: /var/run/secrets/kubernetes.io/serviceaccount/token + honorLabels: true + interval: 60s + metricRelabelings: + - action: keep + regex: kubelet_cgroup_manager_duration_seconds_count|go_goroutines|kubelet_pod_start_duration_seconds_count|kubelet_runtime_operations_total|kubelet_pleg_relist_duration_seconds_bucket|volume_manager_total_volumes|kubelet_volume_stats_capacity_bytes|container_cpu_usage_seconds_total|container_network_transmit_bytes_total|kubelet_runtime_operations_errors_total|container_network_receive_bytes_total|container_memory_swap|container_network_receive_packets_total|container_cpu_cfs_periods_total|container_cpu_cfs_throttled_periods_total|kubelet_running_pod_count|node_namespace_pod_container:container_cpu_usage_seconds_total:sum_rate|container_memory_working_set_bytes|storage_operation_errors_total|kubelet_pleg_relist_duration_seconds_count|kubelet_running_pods|rest_client_request_duration_seconds_bucket|process_resident_memory_bytes|storage_operation_duration_seconds_count|kubelet_running_containers|kubelet_runtime_operations_duration_seconds_bucket|kubelet_node_config_error|kubelet_cgroup_manager_duration_seconds_bucket|kubelet_running_container_count|kubelet_volume_stats_available_bytes|kubelet_volume_stats_inodes|container_memory_rss|kubelet_pod_worker_duration_seconds_count|kubelet_node_name|kubelet_pleg_relist_interval_seconds_bucket|container_network_receive_packets_dropped_total|kubelet_pod_worker_duration_seconds_bucket|container_start_time_seconds|container_network_transmit_packets_dropped_total|process_cpu_seconds_total|storage_operation_duration_seconds_bucket|container_memory_cache|container_network_transmit_packets_total|kubelet_volume_stats_inodes_used|up|rest_client_requests_total + sourceLabels: + - __name__ + - action: replace + targetLabel: job + replacement: integrations/kubernetes/kubelet + port: https-metrics + relabelings: + - sourceLabels: + - __metrics_path__ + targetLabel: metrics_path + scheme: https + tlsConfig: + insecureSkipVerify: true + namespaceSelector: + matchNames: + - default + selector: + matchLabels: + app.kubernetes.io/name: kubelet +``` + +- cAdvsior ServiceMonitor + +```yaml +apiVersion: monitoring.coreos.com/v1 +kind: ServiceMonitor +metadata: + labels: + instance: primary + name: cadvisor-monitor + namespace: default +spec: + endpoints: + - bearerTokenFile: /var/run/secrets/kubernetes.io/serviceaccount/token + honorLabels: true + honorTimestamps: false + interval: 60s + metricRelabelings: + - action: keep + regex: kubelet_cgroup_manager_duration_seconds_count|go_goroutines|kubelet_pod_start_duration_seconds_count|kubelet_runtime_operations_total|kubelet_pleg_relist_duration_seconds_bucket|volume_manager_total_volumes|kubelet_volume_stats_capacity_bytes|container_cpu_usage_seconds_total|container_network_transmit_bytes_total|kubelet_runtime_operations_errors_total|container_network_receive_bytes_total|container_memory_swap|container_network_receive_packets_total|container_cpu_cfs_periods_total|container_cpu_cfs_throttled_periods_total|kubelet_running_pod_count|node_namespace_pod_container:container_cpu_usage_seconds_total:sum_rate|container_memory_working_set_bytes|storage_operation_errors_total|kubelet_pleg_relist_duration_seconds_count|kubelet_running_pods|rest_client_request_duration_seconds_bucket|process_resident_memory_bytes|storage_operation_duration_seconds_count|kubelet_running_containers|kubelet_runtime_operations_duration_seconds_bucket|kubelet_node_config_error|kubelet_cgroup_manager_duration_seconds_bucket|kubelet_running_container_count|kubelet_volume_stats_available_bytes|kubelet_volume_stats_inodes|container_memory_rss|kubelet_pod_worker_duration_seconds_count|kubelet_node_name|kubelet_pleg_relist_interval_seconds_bucket|container_network_receive_packets_dropped_total|kubelet_pod_worker_duration_seconds_bucket|container_start_time_seconds|container_network_transmit_packets_dropped_total|process_cpu_seconds_total|storage_operation_duration_seconds_bucket|container_memory_cache|container_network_transmit_packets_total|kubelet_volume_stats_inodes_used|up|rest_client_requests_total + sourceLabels: + - __name__ + - action: replace + targetLabel: job + replacement: integrations/kubernetes/cadvisor + path: /metrics/cadvisor + port: https-metrics + relabelings: + - sourceLabels: + - __metrics_path__ + targetLabel: metrics_path + scheme: https + tlsConfig: + insecureSkipVerify: true + namespaceSelector: + matchNames: + - default + selector: + matchLabels: + app.kubernetes.io/name: kubelet +``` + +These two ServiceMonitors configure Agent to scrape all the Kubelet and cAdvisor endpoints in your Kubernetes cluster (one of each per Node). In addition, it defines a `job` label which you may change (it is preset here for compatibility with Grafana Cloud's Kubernetes integration), and allowlists a core set of Kubernetes metrics to reduce remote metrics usage. If you don't need this allowlist, you may omit it, however note that your metrics usage will increase significantly. + + When you're done, Agent should now be shipping Kubelet and cAdvisor metrics to your remote Prometheus endpoint. + +## Step 4: Deploy LogsInstance and PodLogs resources + +In this step, you'll deploy a LogsInstance resource to collect logs from your cluster nodes and ship these to your remote Loki endpoint. Under the hood, Agent Operator will deploy a DaemonSet of Agents in your cluster that will tail log files defined in PodLogs resources. + +Deploy the LogsInstance into your cluster: + +```yaml +apiVersion: monitoring.grafana.com/v1alpha1 +kind: LogsInstance +metadata: + name: primary + namespace: default + labels: + agent: grafana-agent-logs +spec: + clients: + - url: your_remote_logs_URL + basicAuth: + username: + name: primary-credentials-logs + key: username + password: + name: primary-credentials-logs + key: password + + # Supply an empty namespace selector to look in all namespaces. Remove + # this to only look in the same namespace as the LogsInstance CR + podLogsNamespaceSelector: {} + podLogsSelector: + matchLabels: + instance: primary +``` + +This LogsInstance will pick up PodLogs resources with the `instance: primary` label. Be sure to set the Loki URL to the correct push endpoint (for Grafana Cloud, this will be something like `logs-prod-us-central1.grafana.net/loki/api/v1/push`, however you should check the Cloud Portal to confirm). + +Also note that we are using the `agent: grafana-agent-logs` label here, which will associate this LogsInstance with the GrafanaAgent resource defined in Step 1. This means that it will inherit requests, limits, affinities and other properties defined in the GrafanaAgent custom resource. + +Create the Secret for the LogsInstance resource: + +```yaml +apiVersion: v1 +kind: Secret +metadata: + name: primary-credentials-logs + namespace: default +stringData: + username: 'your_username_here' + password: 'your_password_here' +``` + +If you're using Grafana Cloud, you can find your hosted Loki endpoint username and password in the [Grafana Cloud Portal](https://grafana.com/profile/org). You may wish to base64-encode these values yourself. In this case, please use `data` instead of `stringData`. + +Finally, we'll roll out a PodLogs resource to define our logging targets. Under the hood, Agent Operator will turn this into Agent config for the logs subsystem, and roll it out to the DaemonSet of logging agents. + +The following is a minimal working example which you should adapt to your production needs: + +```yaml +apiVersion: monitoring.grafana.com/v1alpha1 +kind: PodLogs +metadata: + labels: + instance: primary + name: kubernetes-pods + namespace: default +spec: + pipelineStages: + - docker: {} + namespaceSelector: + matchNames: + - default + selector: + matchLabels: {} +``` + +This tails container logs for all Pods in the `default` Namespace. You can restrict the set of Pods matched by using the `matchLabels` selector. You can also set additional `pipelineStages` and create `relabelings` to add or modify log line labels. To learn more about the PodLogs spec and available resource fields, please see the [PodLogs CRD](https://github.com/grafana/agent/blob/main/production/operator/crds/monitoring.grafana.com_podlogs.yaml). + +Under the hood, the above PodLogs resource will add the following labels to log lines: + +- `namespace` +- `service` +- `pod` +- `container` +- `job` + - Set to `PodLogs_namespace/PodLogs_name` +- `__path__` (the path to log files) + - Set to `/var/log/pods/*$1/*.log` where `$1` is `__meta_kubernetes_pod_uid/__meta_kubernetes_pod_container_name` + +To learn more about this config format and other available labels, please see the [Promtail Scraping](https://grafana.com/docs/loki/latest/clients/promtail/scraping/#promtail-scraping-service-discovery) reference documentation. Agent Operator will load this config into the LogsInstance agents automatically. + +At this point the DaemonSet of logging agents should be tailing your container logs, applying some default labels to the log lines, and shipping them to your remote Loki endpoint. + +## Conclusion + +At this point you've rolled out the following into your cluster: + +- A `GrafanaAgent` resource, which discovers one or more `MetricsInstance` and `LogsInstances` resources. +- A `MetricsInstance` resource that defines where to ship collected metrics. +- A `ServiceMonitor` resource to collect cAdvisor and kubelet metrics. +- A `LogsInstance` resource that defines where to ship collected logs. +- A `PodLogs` resource to collect container logs from Kubernetes Pods. + +You can verify that everything is working correctly by navigating to your Grafana instance and querying your Loki and Prometheus datasources. Operator support for Tempo and traces is coming soon. diff --git a/docs/operator/getting-started.md b/docs/operator/getting-started.md index 2daa50f1c8a3..9e795bc5e353 100644 --- a/docs/operator/getting-started.md +++ b/docs/operator/getting-started.md @@ -1,14 +1,24 @@ +++ -title = "Get started with Grafana Agent Operator" +title = "Installing Grafana Agent Operator" weight = 100 +++ -# Get started with Grafana Agent Operator +# Installing Grafana Agent Operator -An official Helm chart is planned to make it really easy to deploy the Grafana Agent -Operator on Kubernetes. For now, things must be done a little manually. +In this guide you'll learn how to deploy the [Grafana Agent Operator]({{< relref "./_index.md" >}}) into your Kubernetes cluster. This guide does *not* use Helm. To learn how to deploy Agent Operator using the [grafana-agent-operator Helm chart](https://github.com/grafana/helm-charts/tree/main/charts/agent-operator), please see [Installing Grafana Agent Operator with Helm]({{< relref "./helm-getting-started.md" >}}). -## Deploy CustomResourceDefinitions +> **Note:** Agent Operator is currently in beta and its custom resources are subject to change as the project evolves. It currently supports the metrics and logs subsystems of Grafana Agent. Integrations and traces support is coming soon. + +By the end of this guide, you'll have deloyed Agent Operator into your cluster. + +## Prerequisites + +Before you begin, make sure that you have the following available to you: + +- A Kubernetes cluster +- The `kubectl` command-line client installed and configured on your machine + +## Step 1: Deploy CustomResourceDefinitions Before you can write custom resources to describe a Grafana Agent deployment, you _must_ deploy the @@ -37,7 +47,7 @@ the documentation for each resource. For example, `kubectl explain GrafanaAgent` will describe the GrafanaAgent CRD, and `kubectl explain GrafanaAgent.spec` will give you information on its spec field. -## Install Agent Operator on Kubernetes +## Step 2: Install Agent Operator Use the following deployment to run the Operator, changing values as desired: @@ -127,7 +137,7 @@ subjects: namespace: default ``` -## Run Operator locally +### Run Operator locally Before running locally, _make sure your kubectl context is correct!_ Running locally uses your current kubectl context, and you probably don't want @@ -143,269 +153,6 @@ Afterwards, you can run the operator using `go run`: go run ./cmd/agent-operator ``` -## Deploy GrafanaAgent - -Now that the Operator is running, you can create a deployment of the -Grafana Agent. The first step is to create a GrafanaAgent resource. This -resource will discover a set of MetricsInstance resources. You can use -this example, which creates a GrafanaAgent and the appropriate ServiceAccount -for you: - -```yaml -apiVersion: monitoring.grafana.com/v1alpha1 -kind: GrafanaAgent -metadata: - name: grafana-agent - namespace: default - labels: - app: grafana-agent -spec: - image: grafana/agent:v0.20.0 - logLevel: info - serviceAccountName: grafana-agent - metrics: - instanceSelector: - matchLabels: - agent: grafana-agent - ---- - -apiVersion: v1 -kind: ServiceAccount -metadata: - name: grafana-agent - namespace: default - ---- - -apiVersion: rbac.authorization.k8s.io/v1 -kind: ClusterRole -metadata: - name: grafana-agent -rules: -- apiGroups: - - "" - resources: - - nodes - - nodes/proxy - - nodes/metrics - - services - - endpoints - - pods - verbs: - - get - - list - - watch -- apiGroups: - - networking.k8s.io - resources: - - ingresses - verbs: - - get - - list - - watch -- nonResourceURLs: - - /metrics - - /metrics/cadvisor - verbs: - - get - ---- - -apiVersion: rbac.authorization.k8s.io/v1 -kind: ClusterRoleBinding -metadata: - name: grafana-agent -roleRef: - apiGroup: rbac.authorization.k8s.io - kind: ClusterRole - name: grafana-agent -subjects: -- kind: ServiceAccount - name: grafana-agent - namespace: default -``` - -Note that this searches for MetricsInstances in the same namespace with the -label matching `agent: grafana-agent`. A MetricsInstance is a custom resource -that describes where to write collected metrics. Use this one as an example: - -```yaml -apiVersion: monitoring.grafana.com/v1alpha1 -kind: MetricsInstance -metadata: - name: primary - namespace: default - labels: - agent: grafana-agent -spec: - remoteWrite: - - url: https://prometheus-us-central1.grafana.net/api/prom/push - basicAuth: - username: - name: primary-credentials - key: username - password: - name: primary-credentials - key: password - - # Supply an empty namespace selector to look in all namespaces. Remove - # this to only look in the same namespace. - serviceMonitorNamespaceSelector: {} - serviceMonitorSelector: - matchLabels: - instance: primary - - # Supply an empty namespace selector to look in all namespaces. Remove - # this to only look in the same namespace. - podMonitorNamespaceSelector: {} - podMonitorSelector: - matchLabels: - instance: primary - - # Supply an empty namespace selector to look in all namespaces. Remove - # this to only look in the same namespace. - probeNamespaceSelector: {} - probeSelector: - matchLabels: - instance: primary -``` - -Replace the remoteWrite URL to match your vendor. If your vendor doesn't need -credentials, you may remove the `basicAuth` section. Otherwise, configure a -secret with the base64-encoded values of the username and password: - -```yaml -apiVersion: v1 -kind: Secret -metadata: - name: primary-credentials - namespace: default -data: - username: BASE64_ENCODED_USERNAME - password: BASE64_ENCODED_PASSWORD -``` +## Conclusion -The above configuration of MetricsInstance will discover all -[PodMonitors](https://github.com/prometheus-operator/prometheus-operator/blob/master/Documentation/api.md#podmonitor), -[Probes](https://github.com/prometheus-operator/prometheus-operator/blob/master/Documentation/api.md#probe), -and [ServiceMonitors](https://github.com/prometheus-operator/prometheus-operator/blob/master/Documentation/api.md#servicemonitor) -with a label matching `instance: primary`. Create resources as appropriate for -your environment. - -As an example, here is a ServiceMonitor that can collect metrics from `kube-dns`: - -```yaml -apiVersion: monitoring.coreos.com/v1 -kind: ServiceMonitor -metadata: - name: kube-dns - namespace: kube-system - labels: - instance: primary -spec: - selector: - matchLabels: - k8s-app: kube-dns - endpoints: - - port: metrics -``` - -## Monitor Kubelets - -The `--kubelet-service=default/kubelet` argument passed to the Grafana Agent -Operator container tells the Operator to manage a Service called `kubelet` in -the `default` namespace. The Service will have one endpoint created per Node in -the cluster, allowing a ServiceMonitor to scrape both Kubelet and cAdvisor -metrics. - -Use the following ServiceMonitor as a base to collect Kubelet and cAdvisor -metrics: - -```yaml -apiVersion: monitoring.coreos.com/v1 -kind: ServiceMonitor -metadata: - labels: - app.kubernetes.io/name: kubelet - instance: primary - name: kubelet - namespace: default -spec: - endpoints: - - bearerTokenFile: /var/run/secrets/kubernetes.io/serviceaccount/token - honorLabels: true - interval: 30s - metricRelabelings: - - action: drop - regex: kubelet_(pod_worker_latency_microseconds|pod_start_latency_microseconds|cgroup_manager_latency_microseconds|pod_worker_start_latency_microseconds|pleg_relist_latency_microseconds|pleg_relist_interval_microseconds|runtime_operations|runtime_operations_latency_microseconds|runtime_operations_errors|eviction_stats_age_microseconds|device_plugin_registration_count|device_plugin_alloc_latency_microseconds|network_plugin_operations_latency_microseconds) - sourceLabels: - - __name__ - - action: drop - regex: scheduler_(e2e_scheduling_latency_microseconds|scheduling_algorithm_predicate_evaluation|scheduling_algorithm_priority_evaluation|scheduling_algorithm_preemption_evaluation|scheduling_algorithm_latency_microseconds|binding_latency_microseconds|scheduling_latency_seconds) - sourceLabels: - - __name__ - - action: drop - regex: apiserver_(request_count|request_latencies|request_latencies_summary|dropped_requests|storage_data_key_generation_latencies_microseconds|storage_transformation_failures_total|storage_transformation_latencies_microseconds|proxy_tunnel_sync_latency_secs) - sourceLabels: - - __name__ - - action: drop - regex: kubelet_docker_(operations|operations_latency_microseconds|operations_errors|operations_timeout) - sourceLabels: - - __name__ - - action: drop - regex: reflector_(items_per_list|items_per_watch|list_duration_seconds|lists_total|short_watches_total|watch_duration_seconds|watches_total) - sourceLabels: - - __name__ - - action: drop - regex: etcd_(helper_cache_hit_count|helper_cache_miss_count|helper_cache_entry_count|object_counts|request_cache_get_latencies_summary|request_cache_add_latencies_summary|request_latencies_summary) - sourceLabels: - - __name__ - - action: drop - regex: transformation_(transformation_latencies_microseconds|failures_total) - sourceLabels: - - __name__ - - action: drop - regex: (admission_quota_controller_adds|admission_quota_controller_depth|admission_quota_controller_longest_running_processor_microseconds|admission_quota_controller_queue_latency|admission_quota_controller_unfinished_work_seconds|admission_quota_controller_work_duration|APIServiceOpenAPIAggregationControllerQueue1_adds|APIServiceOpenAPIAggregationControllerQueue1_depth|APIServiceOpenAPIAggregationControllerQueue1_longest_running_processor_microseconds|APIServiceOpenAPIAggregationControllerQueue1_queue_latency|APIServiceOpenAPIAggregationControllerQueue1_retries|APIServiceOpenAPIAggregationControllerQueue1_unfinished_work_seconds|APIServiceOpenAPIAggregationControllerQueue1_work_duration|APIServiceRegistrationController_adds|APIServiceRegistrationController_depth|APIServiceRegistrationController_longest_running_processor_microseconds|APIServiceRegistrationController_queue_latency|APIServiceRegistrationController_retries|APIServiceRegistrationController_unfinished_work_seconds|APIServiceRegistrationController_work_duration|autoregister_adds|autoregister_depth|autoregister_longest_running_processor_microseconds|autoregister_queue_latency|autoregister_retries|autoregister_unfinished_work_seconds|autoregister_work_duration|AvailableConditionController_adds|AvailableConditionController_depth|AvailableConditionController_longest_running_processor_microseconds|AvailableConditionController_queue_latency|AvailableConditionController_retries|AvailableConditionController_unfinished_work_seconds|AvailableConditionController_work_duration|crd_autoregistration_controller_adds|crd_autoregistration_controller_depth|crd_autoregistration_controller_longest_running_processor_microseconds|crd_autoregistration_controller_queue_latency|crd_autoregistration_controller_retries|crd_autoregistration_controller_unfinished_work_seconds|crd_autoregistration_controller_work_duration|crdEstablishing_adds|crdEstablishing_depth|crdEstablishing_longest_running_processor_microseconds|crdEstablishing_queue_latency|crdEstablishing_retries|crdEstablishing_unfinished_work_seconds|crdEstablishing_work_duration|crd_finalizer_adds|crd_finalizer_depth|crd_finalizer_longest_running_processor_microseconds|crd_finalizer_queue_latency|crd_finalizer_retries|crd_finalizer_unfinished_work_seconds|crd_finalizer_work_duration|crd_naming_condition_controller_adds|crd_naming_condition_controller_depth|crd_naming_condition_controller_longest_running_processor_microseconds|crd_naming_condition_controller_queue_latency|crd_naming_condition_controller_retries|crd_naming_condition_controller_unfinished_work_seconds|crd_naming_condition_controller_work_duration|crd_openapi_controller_adds|crd_openapi_controller_depth|crd_openapi_controller_longest_running_processor_microseconds|crd_openapi_controller_queue_latency|crd_openapi_controller_retries|crd_openapi_controller_unfinished_work_seconds|crd_openapi_controller_work_duration|DiscoveryController_adds|DiscoveryController_depth|DiscoveryController_longest_running_processor_microseconds|DiscoveryController_queue_latency|DiscoveryController_retries|DiscoveryController_unfinished_work_seconds|DiscoveryController_work_duration|kubeproxy_sync_proxy_rules_latency_microseconds|non_structural_schema_condition_controller_adds|non_structural_schema_condition_controller_depth|non_structural_schema_condition_controller_longest_running_processor_microseconds|non_structural_schema_condition_controller_queue_latency|non_structural_schema_condition_controller_retries|non_structural_schema_condition_controller_unfinished_work_seconds|non_structural_schema_condition_controller_work_duration|rest_client_request_latency_seconds|storage_operation_errors_total|storage_operation_status_count) - sourceLabels: - - __name__ - port: https-metrics - relabelings: - - sourceLabels: - - __metrics_path__ - targetLabel: metrics_path - scheme: https - tlsConfig: - insecureSkipVerify: true - - bearerTokenFile: /var/run/secrets/kubernetes.io/serviceaccount/token - honorLabels: true - honorTimestamps: false - interval: 30s - metricRelabelings: - - action: drop - regex: container_(network_tcp_usage_total|network_udp_usage_total|tasks_state|cpu_load_average_10s) - sourceLabels: - - __name__ - - action: drop - regex: (container_fs_.*|container_spec_.*|container_blkio_device_usage_total|container_file_descriptors|container_sockets|container_threads_max|container_threads|container_start_time_seconds|container_last_seen);; - sourceLabels: - - __name__ - - pod - - namespace - path: /metrics/cadvisor - port: https-metrics - relabelings: - - sourceLabels: - - __metrics_path__ - targetLabel: metrics_path - scheme: https - tlsConfig: - insecureSkipVerify: true - jobLabel: app.kubernetes.io/name - namespaceSelector: - matchNames: - - default - selector: - matchLabels: - app.kubernetes.io/name: kubelet -``` +With Agent Operator up and running, you can move on to setting up a `GrafanaAgent` custom resource. This will discover `MetricsInstance` and `LogsInstance` custom resources and endow them with Pod attributes (like requests and limits) defined in the `GrafanaAgent` spec. To learn how to do this, please see [Custom Resource Quickstart]({{< relref "./custom-resource-quickstart.md" >}}. diff --git a/docs/operator/helm-getting-started.md b/docs/operator/helm-getting-started.md index 18ae6d81fe71..379b7fa3101f 100644 --- a/docs/operator/helm-getting-started.md +++ b/docs/operator/helm-getting-started.md @@ -1,24 +1,14 @@ +++ -title = "Grafana Agent Operator Helm Quickstart" +title = "Installing Grafana Agent Operator with Helm" weight = 110 +++ -# Grafana Agent Operator Helm Quickstart +# Installing Grafana Agent Operator with Helm -In this guide you'll learn how to deploy the [Grafana Agent Operator]({{< relref "./_index.md" >}}) into your Kubernetes cluster using the [grafana-agent-operator Helm chart](https://github.com/grafana/helm-charts/tree/main/charts/agent-operator). - -You'll then deploy the following custom resources (CRs): - -- A `GrafanaAgent` resource, which discovers one or more `MetricsInstance` and `LogsInstances` resources. -- A `MetricsInstance` resource that defines where to ship collected metrics. Under the hood, this rolls out a Grafana Agent StatefulSet that will scrape and ship metrics to a `remote_write` endpoint. -- A `ServiceMonitor` resource to collect cAdvisor and kubelet metrics. Under the hood, this configures the `MetricsInstance` / Agent StatefulSet. -- A `LogsInstance` resource that defines where to ship collected logs. Under the hood, this rolls out a Grafana Agent DaemonSet that will tail log files on your cluster nodes. -- A `PodLogs` resource to collect container logs from Kubernetes Pods. Under the hood, this configures the`LogsInstance` / Agent DaemonSet. - -To learn more about the custom resources Operator provides and their hierarchy, please consult [Operator architecture]({{< relref "./architecture.md" >}}). +In this guide you'll learn how to deploy the [Grafana Agent Operator]({{< relref "./_index.md" >}}) into your Kubernetes cluster using the [grafana-agent-operator Helm chart](https://github.com/grafana/helm-charts/tree/main/charts/agent-operator). > **Note:** Agent Operator is currently in beta and its custom resources are subject to change as the project evolves. It currently supports the metrics and logs subsystems of Grafana Agent. Integrations and traces support is coming soon. -By the end of this guide, you'll have deloyed Agent Operator into your cluster and will be scraping and shipping cAdvisor and Kubelet metrics to a Prometheus-compatible metrics endpoint. You'll also be collecting and shipping your Pods' container logs to a Loki-compatible logs endpoint. +By the end of this guide, you'll have deloyed Agent Operator into your cluster. ## Prerequisites @@ -28,7 +18,7 @@ Before you begin, make sure that you have the following available to you: - The `kubectl` command-line client installed and configured on your machine - The `helm` command-line client installed and configured on your machine -## Step 1: Install Agent Operator Helm Chart +## Install Agent Operator Helm Chart In this step you'll install the [grafana-agent-operator Helm chart](https://github.com/grafana/helm-charts/tree/main/charts/agent-operator) into your Kubernetes cluster. This will install the latest version of Agent Operator and its [Custom Resource Definitions](https://github.com/grafana/agent/tree/main/production/operator/crds) (CRDs). By default the chart will configure the operator to maintain a Service that allows you scrape kubelets using a `ServiceMonitor`. @@ -70,371 +60,6 @@ kubectl get svc You should see an Agent Operator Pod in `RUNNING` state, and a `kubelet` Service. -With Agent Operator up and running, you can move on to setting up a `GrafanaAgent` custom resource. This will discover `MetricsInstance` and `LogsInstance` custom resources and endow them with Pod attributes (like requests and limits) defined in the `GrafanaAgent` spec. - -## Step 2: Deploy GrafanaAgent resource - -In this step you'll roll out a `GrafanaAgent` resource. A `GrafanaAgent` resource discovers `MetricsInstance` and `LogsInstance` resources and defines the Grafana Agent image, Pod requests, limits, affinities, and tolerations. Pod attributes can only be defined at the GrafanaAgent level and are propagated to `MetricsInstance` and `LogsInstance` Pods. To learn more, please see the GrafanaAgent [Custom Resource Definition](https://github.com/grafana/agent/blob/main/production/operator/crds/monitoring.grafana.com_grafanaagents.yaml). - -> **Note:** Due to the variety of possible deployment architectures, the official Agent Operator Helm chart does not provide built-in templates for the custom resources described in this quickstart. These must be configured and deployed manually. However, you are encouraged to template and add the following manifests to your own in-house Helm charts and GitOps flows. - -Roll out the following manifests in your cluster: - -```yaml -apiVersion: monitoring.grafana.com/v1alpha1 -kind: GrafanaAgent -metadata: - name: grafana-agent - namespace: default - labels: - app: grafana-agent -spec: - image: grafana/agent:v0.20.0 - logLevel: info - serviceAccountName: grafana-agent - metrics: - instanceSelector: - matchLabels: - agent: grafana-agent-metrics - externalLabels: - cluster: cloud - - logs: - instanceSelector: - matchLabels: - agent: grafana-agent-logs - ---- - -apiVersion: v1 -kind: ServiceAccount -metadata: - name: grafana-agent - namespace: default - ---- - -apiVersion: rbac.authorization.k8s.io/v1 -kind: ClusterRole -metadata: - name: grafana-agent -rules: -- apiGroups: - - "" - resources: - - nodes - - nodes/proxy - - nodes/metrics - - services - - endpoints - - pods - verbs: - - get - - list - - watch -- apiGroups: - - networking.k8s.io - resources: - - ingresses - verbs: - - get - - list - - watch -- nonResourceURLs: - - /metrics - - /metrics/cadvisor - verbs: - - get - ---- - -apiVersion: rbac.authorization.k8s.io/v1 -kind: ClusterRoleBinding -metadata: - name: grafana-agent -roleRef: - apiGroup: rbac.authorization.k8s.io - kind: ClusterRole - name: grafana-agent -subjects: -- kind: ServiceAccount - name: grafana-agent - namespace: default -``` - -This creates a ServiceAccount, ClusterRole, and ClusterRoleBinding for the GrafanaAgent resource. It also creates a GrafanaAgent resource and specifies an Agent image version. Finally, the GrafanaAgent resource specifies `MetricsInstance` and `LogsInstance` selectors. These search for MetricsInstances and LogsInstances in the same namespace with labels matching `agent: grafana-agent-metrics` and `agent: grafana-agent-logs`, respectively. It also sets a `cluster: cloud` label for all metrics shipped your Prometheus-compatible endpoint. You should change this label to your desired cluster name. - -The full hierarchy of custom resources is as follows: - -- `GrafanaAgent` - - `MetricsInstance` - - `PodMonitor` - - `Probe` - - `ServiceMonitor` - - `LogsInstance` - - `PodLogs` - -Deploying a GrafanaAgent resource on its own will not spin up any Agent Pods. Agent Operator will create Agent Pods once MetricsInstance and LogsIntance resources have been created. In the next step, you'll roll out a `MetricsInstance` resource to scrape cAdvisor and Kubelet metrics and ship these to your Prometheus-compatible metrics endpoint. - -## Step 3: Deploy a MetricsInstance resource - -In this step you'll roll out a MetricsInstance resource. MetricsInstance resources define a `remote_write` sink for metrics and configure one or more selectors to watch for creation and updates to `*Monitor` objects. These objects allow you to define Agent scrape targets via K8s manifests: - -- [ServiceMonitors](https://github.com/prometheus-operator/prometheus-operator/blob/master/Documentation/api.md#servicemonitor) -- [PodMonitors](https://github.com/prometheus-operator/prometheus-operator/blob/master/Documentation/api.md#podmonitor) -- [Probes](https://github.com/prometheus-operator/prometheus-operator/blob/master/Documentation/api.md#probe) - -Roll out the following manifest into your cluster: - -``` -apiVersion: monitoring.grafana.com/v1alpha1 -kind: MetricsInstance -metadata: - name: primary - namespace: default - labels: - agent: grafana-agent-metrics -spec: - remoteWrite: - - url: your_remote_write_URL - basicAuth: - username: - name: primary-credentials-metrics - key: username - password: - name: primary-credentials-metrics - key: password - - # Supply an empty namespace selector to look in all namespaces. Remove - # this to only look in the same namespace as the MetricsInstance CR - serviceMonitorNamespaceSelector: {} - serviceMonitorSelector: - matchLabels: - instance: primary - - # Supply an empty namespace selector to look in all namespaces. Remove - # this to only look in the same namespace as the MetricsInstance CR. - podMonitorNamespaceSelector: {} - podMonitorSelector: - matchLabels: - instance: primary - - # Supply an empty namespace selector to look in all namespaces. Remove - # this to only look in the same namespace as the MetricsInstance CR. - probeNamespaceSelector: {} - probeSelector: - matchLabels: - instance: primary -``` - -Be sure to replace the `remote_write` URL and customize the namespace and label configuration as necessary. This will associate itself with the `agent: grafana-agent` GrafanaAgent resource deployed in the previous step, and watch for creation and updates to `*Monitors` monitors with the the `instance: primary` label. - -Once you've rolled out this manifest, create the `basicAuth` credentials using a Kubernetes Secret: - -``` -apiVersion: v1 -kind: Secret -metadata: - name: primary-credentials-metrics - namespace: default -stringData: - username: 'your_cloud_prometheus_username' - password: 'your_cloud_prometheus_API_key' -``` - -If you're using Grafana Cloud, you can find your hosted Prometheus endpoint username and password in the [Grafana Cloud Portal](https://grafana.com/profile/org ). You may wish to base64-encode these values yourself. In this case, please use `data` instead of `stringData`. - -Once you've rolled out the `MetricsInstance` and its Secret, you can confirm that the MetricsInstance Agent is up and running with `kubectl get pod`. Since we haven't defined any monitors yet, this Agent will not have any scrape targets defined. In the next step, we'll create scrape targets for the cAdvisor and kubelet endpoints exposed by the `kubelet` service in the cluster. - -## Step 4: Create ServiceMonitors for kubelet and cAdvisor endpoints - -In this step, you'll create ServiceMonitors for kubelet and cAdvisor metrics exposed by the `kubelet` Service. Every node in your cluster exposes kubelet and cadvisor metrics at `/metrics` and `/metrics/cadvisor` respectively. Agent Operator creates a `kubelet` service that exposes these Node endpoints so that they can be scraped using ServiceMonitors. - -To scrape these two endpoints, roll out the following two ServiceMonitors in your cluster: - -- Kubelet ServiceMonitor - -```yaml -apiVersion: monitoring.coreos.com/v1 -kind: ServiceMonitor -metadata: - labels: - instance: primary - name: kubelet-monitor - namespace: default -spec: - endpoints: - - bearerTokenFile: /var/run/secrets/kubernetes.io/serviceaccount/token - honorLabels: true - interval: 60s - metricRelabelings: - - action: keep - regex: kubelet_cgroup_manager_duration_seconds_count|go_goroutines|kubelet_pod_start_duration_seconds_count|kubelet_runtime_operations_total|kubelet_pleg_relist_duration_seconds_bucket|volume_manager_total_volumes|kubelet_volume_stats_capacity_bytes|container_cpu_usage_seconds_total|container_network_transmit_bytes_total|kubelet_runtime_operations_errors_total|container_network_receive_bytes_total|container_memory_swap|container_network_receive_packets_total|container_cpu_cfs_periods_total|container_cpu_cfs_throttled_periods_total|kubelet_running_pod_count|node_namespace_pod_container:container_cpu_usage_seconds_total:sum_rate|container_memory_working_set_bytes|storage_operation_errors_total|kubelet_pleg_relist_duration_seconds_count|kubelet_running_pods|rest_client_request_duration_seconds_bucket|process_resident_memory_bytes|storage_operation_duration_seconds_count|kubelet_running_containers|kubelet_runtime_operations_duration_seconds_bucket|kubelet_node_config_error|kubelet_cgroup_manager_duration_seconds_bucket|kubelet_running_container_count|kubelet_volume_stats_available_bytes|kubelet_volume_stats_inodes|container_memory_rss|kubelet_pod_worker_duration_seconds_count|kubelet_node_name|kubelet_pleg_relist_interval_seconds_bucket|container_network_receive_packets_dropped_total|kubelet_pod_worker_duration_seconds_bucket|container_start_time_seconds|container_network_transmit_packets_dropped_total|process_cpu_seconds_total|storage_operation_duration_seconds_bucket|container_memory_cache|container_network_transmit_packets_total|kubelet_volume_stats_inodes_used|up|rest_client_requests_total - sourceLabels: - - __name__ - - action: replace - targetLabel: job - replacement: integrations/kubernetes/kubelet - port: https-metrics - relabelings: - - sourceLabels: - - __metrics_path__ - targetLabel: metrics_path - scheme: https - tlsConfig: - insecureSkipVerify: true - namespaceSelector: - matchNames: - - default - selector: - matchLabels: - app.kubernetes.io/name: kubelet -``` - -- cAdvsior ServiceMonitor - -```yaml -apiVersion: monitoring.coreos.com/v1 -kind: ServiceMonitor -metadata: - labels: - instance: primary - name: cadvisor-monitor - namespace: default -spec: - endpoints: - - bearerTokenFile: /var/run/secrets/kubernetes.io/serviceaccount/token - honorLabels: true - honorTimestamps: false - interval: 60s - metricRelabelings: - - action: keep - regex: kubelet_cgroup_manager_duration_seconds_count|go_goroutines|kubelet_pod_start_duration_seconds_count|kubelet_runtime_operations_total|kubelet_pleg_relist_duration_seconds_bucket|volume_manager_total_volumes|kubelet_volume_stats_capacity_bytes|container_cpu_usage_seconds_total|container_network_transmit_bytes_total|kubelet_runtime_operations_errors_total|container_network_receive_bytes_total|container_memory_swap|container_network_receive_packets_total|container_cpu_cfs_periods_total|container_cpu_cfs_throttled_periods_total|kubelet_running_pod_count|node_namespace_pod_container:container_cpu_usage_seconds_total:sum_rate|container_memory_working_set_bytes|storage_operation_errors_total|kubelet_pleg_relist_duration_seconds_count|kubelet_running_pods|rest_client_request_duration_seconds_bucket|process_resident_memory_bytes|storage_operation_duration_seconds_count|kubelet_running_containers|kubelet_runtime_operations_duration_seconds_bucket|kubelet_node_config_error|kubelet_cgroup_manager_duration_seconds_bucket|kubelet_running_container_count|kubelet_volume_stats_available_bytes|kubelet_volume_stats_inodes|container_memory_rss|kubelet_pod_worker_duration_seconds_count|kubelet_node_name|kubelet_pleg_relist_interval_seconds_bucket|container_network_receive_packets_dropped_total|kubelet_pod_worker_duration_seconds_bucket|container_start_time_seconds|container_network_transmit_packets_dropped_total|process_cpu_seconds_total|storage_operation_duration_seconds_bucket|container_memory_cache|container_network_transmit_packets_total|kubelet_volume_stats_inodes_used|up|rest_client_requests_total - sourceLabels: - - __name__ - - action: replace - targetLabel: job - replacement: integrations/kubernetes/cadvisor - path: /metrics/cadvisor - port: https-metrics - relabelings: - - sourceLabels: - - __metrics_path__ - targetLabel: metrics_path - scheme: https - tlsConfig: - insecureSkipVerify: true - namespaceSelector: - matchNames: - - default - selector: - matchLabels: - app.kubernetes.io/name: kubelet -``` - -These two ServiceMonitors configure Agent to scrape all the Kubelet and cAdvisor endpoints in your Kubernetes cluster (one of each per Node). In addition, it defines a `job` label which you may change (it is preset here for compatibility with Grafana Cloud's Kubernetes integration), and allowlists a core set of Kubernetes metrics to reduce remote metrics usage. If you don't need this allowlist, you may omit it, however note that your metrics usage will increase significantly. - - When you're done, Agent should now be shipping Kubelet and cAdvisor metrics to your remote Prometheus endpoint. - -## Step 5: Deploy LogsInstance and PodLogs resources - -In this step, you'll deploy a LogsInstance resource to collect logs from your cluster nodes and ship these to your remote Loki endpoint. Under the hood, Agent Operator will deploy a DaemonSet of Agents in your cluster that will tail log files defined in PodLogs resources. - -Deploy the LogsInstance into your cluster: - -```yaml -apiVersion: monitoring.grafana.com/v1alpha1 -kind: LogsInstance -metadata: - name: primary - namespace: default - labels: - agent: grafana-agent-logs -spec: - clients: - - url: your_remote_logs_URL - basicAuth: - username: - name: primary-credentials-logs - key: username - password: - name: primary-credentials-logs - key: password - - # Supply an empty namespace selector to look in all namespaces. Remove - # this to only look in the same namespace as the LogsInstance CR - podLogsNamespaceSelector: {} - podLogsSelector: - matchLabels: - instance: primary -``` - -This LogsInstance will pick up PodLogs resources with the `instance: primary` label. Be sure to set the Loki URL to the correct push endpoint (for Grafana Cloud, this will be something like `logs-prod-us-central1.grafana.net/loki/api/v1/push`, however you should check the Cloud Portal to confirm). - -Also note that we are using the `agent: grafana-agent-logs` label here, which will associate this LogsInstance with the GrafanaAgent resource defined in Step 2. This means that it will inherit requests, limits, affinities and other properties defined in the GrafanaAgent custom resource. - -Create the Secret for the LogsInstance resource: - -```yaml -apiVersion: v1 -kind: Secret -metadata: - name: primary-credentials-logs - namespace: default -stringData: - username: 'your_username_here' - password: 'your_password_here' -``` - -If you're using Grafana Cloud, you can find your hosted Loki endpoint username and password in the [Grafana Cloud Portal](https://grafana.com/profile/org). You may wish to base64-encode these values yourself. In this case, please use `data` instead of `stringData`. - -Finally, we'll roll out a PodLogs resource to define our logging targets. Under the hood, Agent Operator will turn this into Agent config for the logs subsystem, and roll it out to the DaemonSet of logging agents. - -The following is a minimal working example which you should adapt to your production needs: - -```yaml -apiVersion: monitoring.grafana.com/v1alpha1 -kind: PodLogs -metadata: - labels: - instance: primary - name: kubernetes-pods - namespace: default -spec: - pipelineStages: - - docker: {} - namespaceSelector: - matchNames: - - default - selector: - matchLabels: {} -``` - -This tails container logs for all Pods in the `default` Namespace. You can restrict the set of Pods matched by using the `matchLabels` selector. You can also set additional `pipelineStages` and create `relabelings` to add or modify log line labels. To learn more about the PodLogs spec and available resource fields, please see the [PodLogs CRD](https://github.com/grafana/agent/blob/main/production/operator/crds/monitoring.grafana.com_podlogs.yaml). - -Under the hood, the above PodLogs resource will add the following labels to log lines: - -- `namespace` -- `service` -- `pod` -- `container` -- `job` - - Set to `PodLogs_namespace/PodLogs_name` -- `__path__` (the path to log files) - - Set to `/var/log/pods/*$1/*.log` where `$1` is `__meta_kubernetes_pod_uid/__meta_kubernetes_pod_container_name` - -To learn more about this config format and other available labels, please see the [Promtail Scraping](https://grafana.com/docs/loki/latest/clients/promtail/scraping/#promtail-scraping-service-discovery) reference documentation. Agent Operator will load this config into the LogsInstance agents automatically. - -At this point the DaemonSet of logging agents should be tailing your container logs, applying some default labels to the log lines, and shipping them to your remote Loki endpoint. - ## Conclusion -At this point you've rolled out the following into your cluster: - -- A `GrafanaAgent` resource, which discovers one or more `MetricsInstance` and `LogsInstances` resources. -- A `MetricsInstance` resource that defines where to ship collected metrics. -- A `ServiceMonitor` resource to collect cAdvisor and kubelet metrics. -- A `LogsInstance` resource that defines where to ship collected logs. -- A `PodLogs` resource to collect container logs from Kubernetes Pods. - -You can verify that everything is working correctly by navigating to your Grafana instance and querying your Loki and Prometheus datasources. Operator support for Tempo and traces is coming soon. +With Agent Operator up and running, you can move on to setting up a `GrafanaAgent` custom resource. This will discover `MetricsInstance` and `LogsInstance` custom resources and endow them with Pod attributes (like requests and limits) defined in the `GrafanaAgent` spec. To learn how to do this, please see [Custom Resource Quickstart]({{< relref "./custom-resource-quickstart.md" >}}. diff --git a/production/README.md b/production/README.md index bcbd9ae3a808..346fe1209b8e 100644 --- a/production/README.md +++ b/production/README.md @@ -7,7 +7,7 @@ Here are some resources to help you run the Grafana Agent: - [Run the Agent locally](#running-the-agent-locally) - [Use the example Kubernetes configs](#use-the-example-kubernetes-configs) - [Grafana Cloud Kubernetes Quickstart Guides](#grafana-cloud-kubernetes-quickstart-guides) -- [Agent Operator Helm Quickstart](#agent-operator-helm-quickstart) +- [Agent Operator Helm Quickstart](#agent-operator-helm-quickstart-guide) - [Build the Agent from Source](#build-the-agent-from-source) - [Use our production Tanka configs](#use-our-production-tanka-configs) @@ -47,7 +47,7 @@ You can find them in the [Grafana Cloud documentation](https://grafana.com/docs/ ## Agent Operator Helm quickstart guide -This guide will show you how to deploy the [Grafana Agent Operator](../docs/operator/_index.md) into your Kubernetes cluster using the [grafana-agent-operator Helm chart](https://github.com/grafana/helm-charts/tree/main/charts/agent-operator). +This guide will show you how to deploy the [Grafana Agent Operator](https://grafana.com/docs/agent/latest/operator/) into your Kubernetes cluster using the [grafana-agent-operator Helm chart](https://github.com/grafana/helm-charts/tree/main/charts/agent-operator). You'll also deploy the following custom resources (CRs): - A `GrafanaAgent` resource, which discovers one or more `MetricsInstance` and `LogsInstances` resources. @@ -56,7 +56,7 @@ You'll also deploy the following custom resources (CRs): - A `LogsInstance` resource that defines where to ship collected logs. - A `PodLogs` resource to collect container logs from Kubernetes Pods. -You can find the guide [here](../docs/operator/helm-getting-started.md). +You can find the guide [here](https://grafana.com/docs/agent/latest/operator/helm-getting-started/). ## Build the Agent from source