-
Notifications
You must be signed in to change notification settings - Fork 208
docs: write observability setup guide (Prometheus + Grafana) #545
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Changes from all commits
File filter
Filter by extension
Conversations
Jump to
Diff view
Diff view
There are no files selected for viewing
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,286 @@ | ||
| # Observability Setup: Prometheus + Grafana | ||
|
|
||
| This guide walks through connecting sorokeep's built-in metrics endpoint to a Prometheus + Grafana observability stack. After completing it you'll have a working dashboard showing contract TTL health, alert activity, and extension costs — all updating automatically from the daemon's metrics. | ||
|
|
||
| ## Prerequisites | ||
|
|
||
| - A running sorokeep daemon (`sorokeep daemon`) | ||
| - [Prometheus](https://prometheus.io/download/) 2.x+ | ||
| - [Grafana](https://grafana.com/grafana/download/) 10.x+ | ||
| - Network connectivity between Prometheus and the sorokeep host | ||
|
|
||
| ## 1. Enable the Metrics Server | ||
|
|
||
| Start the daemon with the `--metrics-port` flag to expose a Prometheus-format `/metrics` endpoint: | ||
|
|
||
| ```bash | ||
| sorokeep daemon --network testnet --metrics-port 9464 | ||
| ``` | ||
|
|
||
| The metrics server also exposes: | ||
| - `/healthz` — liveness probe (always 200 while the process is running) | ||
| - `/readyz` — readiness probe (200 when DB and RPC are reachable, 503 otherwise) | ||
|
|
||
| You can verify the endpoint is working: | ||
|
|
||
| ```bash | ||
| curl http://localhost:9464/metrics | ||
| ``` | ||
|
|
||
| Sample output: | ||
|
|
||
| ``` | ||
|
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. 📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win Specify the fenced block language. This sample output triggers MD040; mark it as 🧰 Tools🪛 markdownlint-cli2 (0.23.1)[warning] 32-32: Fenced code blocks should have a language specified (MD040, fenced-code-language) 🤖 Prompt for AI AgentsSource: Linters/SAST tools |
||
| # HELP sorokeep_contracts_total Total number of watched contracts | ||
| # TYPE sorokeep_contracts_total gauge | ||
| sorokeep_contracts_total{network="testnet"} 3 | ||
|
|
||
| # HELP sorokeep_contract_entries_total Total number of contract entries across all contracts | ||
| # TYPE sorokeep_contract_entries_total gauge | ||
| sorokeep_contract_entries_total 10 | ||
|
|
||
| # HELP sorokeep_extensions_total Total number of TTL extensions performed | ||
| # TYPE sorokeep_extensions_total counter | ||
| sorokeep_extensions_total 42 | ||
|
|
||
| # HELP sorokeep_alerts_fired_total Total number of alerts fired | ||
| # TYPE sorokeep_alerts_fired_total counter | ||
| sorokeep_alerts_fired_total 7 | ||
| ``` | ||
|
|
||
| ### Metrics Port Reference | ||
|
|
||
| | Flag | Default | Description | | ||
| |------|---------|-------------| | ||
| | `--metrics-port` | Disabled | Port for the Prometheus metrics HTTP server | | ||
| | `--metrics-host` | `0.0.0.0` | Bind address for the metrics server | | ||
|
Comment on lines
+50
to
+55
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. 🔒 Security & Privacy | 🟠 Major | ⚡ Quick win Do not present The unauthenticated metrics endpoint exposes operational data and is reachable on every interface. Recommend binding to a private interface by default, or explicitly document firewall/reverse-proxy restrictions before exposing this port. 🤖 Prompt for AI Agents |
||
|
|
||
| ## 2. Configure Prometheus to Scrape sorokeep | ||
|
|
||
| Create or edit `prometheus.yml`: | ||
|
|
||
| ```yaml | ||
| scrape_configs: | ||
| - job_name: 'sorokeep' | ||
| scrape_interval: 15s | ||
| scrape_timeout: 10s | ||
| metrics_path: /metrics | ||
| static_configs: | ||
| - targets: | ||
| - 'localhost:9464' | ||
| labels: | ||
| service: sorokeep | ||
| network: testnet | ||
| ``` | ||
|
|
||
| Replace `localhost` with the sorokeep host address if Prometheus is running on a different machine. If sorokeep runs in a Docker container, use the container name or host's Docker bridge IP. | ||
|
Comment on lines
+67
to
+75
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. 🎯 Functional Correctness | 🟠 Major | ⚡ Quick win 🧩 Analysis chain🏁 Script executed: #!/bin/bash
set -euo pipefail
echo "== files =="
git ls-files | rg '(^|/)(prometheus|docker-compose|compose|observability)' || true
echo "== relevant docs slice =="
sed -n '1,260p' docs/observability.md 2>/dev/null | nl -ba | sed -n '1,260p'
echo "== search sorokeep/prometheus compose =="
rg -n "sorokeep|prometheus|9464|statsd|docker-compose|compose" -S . --glob '!node_modules' --glob '!target' | head -n 200Repository: AbdulmalikAlayande/sorokeep Length of output: 348 🏁 Script executed: #!/bin/bash
set -euxo pipefail
echo "== docs/observability.md relevant lines =="
awk '{printf "%6d\t%s\n", NR, $0}' docs/observability.md 2>/dev/null | sed -n '1,280p'
echo "== compose files =="
for f in docker-compose.yaml docker-compose.devnet.yaml; do
echo "--- $f ($(wc -l < "$f") lines) ---"
awk '{printf "%6d\t%s\n", NR, $0}' "$f" | sed -n '1,260p'
done
echo "== search for targets and ports =="
rg -n "targets:|localhost:9464|sorokeep:9464|9464|sorokeep|prometheus|statsd|expose:|ports:" -S . --glob '!node_modules' --glob '!target' | head -n 240Repository: AbdulmalikAlayande/sorokeep Length of output: 40020 Use the Compose service name for Prometheus scraping. Section 5 starts an 🤖 Prompt for AI Agents |
||
|
|
||
| Start Prometheus with the config: | ||
|
|
||
| ```bash | ||
| prometheus --config.file=prometheus.yml | ||
| ``` | ||
|
|
||
| Verify targets are up at `http://localhost:9090/targets` — the `sorokeep` job should show `UP`. | ||
|
|
||
| ### Available Metrics | ||
|
|
||
| | Metric | Type | Description | | ||
| |--------|------|-------------| | ||
| | `sorokeep_contracts_total` | Gauge | Watched contracts, labelled by network | | ||
| | `sorokeep_contract_entries_total` | Gauge | Total tracked ledger entries | | ||
| | `sorokeep_extensions_total` | Counter | TTL extensions performed | | ||
| | `sorokeep_extension_cost_xlm_total` | Counter | Total XLM spent on extensions | | ||
| | `sorokeep_alerts_fired_total` | Counter | Alerts fired across all contracts | | ||
| | `sorokeep_alerts_unresolved_total` | Gauge | Currently unresolved alerts | | ||
| | `sorokeep_channel_accounts_total` | Gauge | Configured channel accounts, labelled by network | | ||
|
|
||
| ## 3. Import the Grafana Dashboard | ||
|
|
||
| A pre-built Grafana dashboard for sorokeep is available at `resources/grafana/dashboard.json`. | ||
|
|
||
| ### Via the Grafana UI | ||
|
|
||
| 1. Open Grafana (`http://localhost:3000`, default login `admin`/`admin`) | ||
|
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. 🔒 Security & Privacy | 🟠 Major | ⚡ Quick win Remove the The guide documents the default Grafana password, sends it in a curl command, and hard-codes it in Compose. This encourages an immediately compromised Grafana instance. Require a secret/environment-provided password and use a token or environment variable for API imports. Suggested Compose change- - GF_SECURITY_ADMIN_PASSWORD=admin
+ - GF_SECURITY_ADMIN_PASSWORD=${GRAFANA_ADMIN_PASSWORD:?Set GRAFANA_ADMIN_PASSWORD}Also applies to: 111-120, 221-226 🤖 Prompt for AI Agents |
||
| 2. Navigate to **Dashboards → New → Import** | ||
| 3. Upload `resources/grafana/dashboard.json` or paste its contents | ||
| 4. Select the Prometheus data source that's scraping sorokeep | ||
| 5. Click **Import** | ||
|
|
||
| ### Via the API | ||
|
|
||
| ```bash | ||
| curl -X POST http://admin:admin@localhost:3000/api/dashboards/db \ | ||
| -H "Content-Type: application/json" \ | ||
| -d @- <<EOF | ||
| { | ||
| "dashboard": $(cat resources/grafana/dashboard.json), | ||
| "overwrite": true, | ||
| "message": "Imported sorokeep dashboard" | ||
| } | ||
| EOF | ||
| ``` | ||
|
|
||
| ### Dashboard Panels | ||
|
|
||
| | Panel | Description | | ||
| |-------|-------------| | ||
| | **Watched Contracts** | Total contracts per network (stat + sparkline) | | ||
| | **Entries Tracked** | Total tracked entries across all contracts | | ||
| | **Extensions (24h)** | Extension count and XLM cost over the last 24 hours | | ||
| | **Alerts Fired (24h)** | Alert count by severity over time | | ||
| | **Unresolved Alerts** | Current unresolved alert count | | ||
| | **Top Contracts by Cost** | Contracts with the highest extension costs | | ||
| | **TTL Distribution** | Remaining TTL distribution across entries | | ||
|
|
||
| ## 4. Alerting with Alertmanager | ||
|
|
||
| ### Example Alert Rules | ||
|
|
||
| Create `sorokeep_alerts.yml`: | ||
|
|
||
| ```yaml | ||
| groups: | ||
| - name: sorokeep | ||
| rules: | ||
| - alert: SorokeepDown | ||
| expr: up{job="sorokeep"} == 0 | ||
| for: 1m | ||
| annotations: | ||
| summary: "sorokeep instance {{ $labels.instance }} is down" | ||
| description: "Prometheus target {{ $labels.job }}/{{ $labels.instance }} has been unreachable for over 1 minute." | ||
|
|
||
| - alert: SorokeepHighUnresolvedAlerts | ||
| expr: sorokeep_alerts_unresolved_total > 0 | ||
| for: 5m | ||
| annotations: | ||
| summary: "Contract alerts are not being resolved" | ||
| description: "{{ $value }} alerts have been unresolved for more than 5 minutes." | ||
|
|
||
| - alert: SorokeepExtensionCostSpike | ||
| expr: rate(sorokeep_extension_cost_xlm_total[1h]) > 1 | ||
| for: 10m | ||
| annotations: | ||
| summary: "High extension cost rate detected" | ||
| description: "XLM extension cost rate is {{ $value }} XLM/hour — investigate unusual extension activity." | ||
|
Comment on lines
+159
to
+164
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. 🎯 Functional Correctness | 🟠 Major | ⚡ Quick win 🧩 Analysis chain🏁 Script executed: #!/bin/bash
set -euo pipefail
echo "== File snippet =="
sed -n '140,175p' docs/observability.md | cat -n
echo
echo "== Related counter names =="
rg -n "sorokeep_extension_cost|rate\\(|XLM|cost" docs/observability.md prometheus 2>/dev/null || trueRepository: AbdulmalikAlayande/sorokeep Length of output: 2479 🌐 Web query:
💡 Result: The rate function in Prometheus is a primary tool for calculating the per-second average rate of increase of a counter metric over a specified time range [1][2]. Key characteristics and usage: - Functionality: It takes a range vector (e.g., [5m]) and calculates the increase in the counter's value over that window, divided by the duration of the window to produce a per-second rate [1][3][2]. - Counter Resets: It automatically detects and compensates for counter resets (such as when a process restarts), ensuring the rate calculation remains accurate [1][4]. - Extrapolation: The function extrapolates to the ends of the time range, which helps account for missed scrapes or misalignment between scrape cycles and the requested time window [1][2]. - Best Practices: - Use only with counters (or native histograms where components behave like counters) [1][5]. - It is recommended for alerting and graphing slow-moving counters [1][4]. - When combining rate with an aggregation operator (like sum) or an _over_time function, always apply rate first to the raw counter before aggregating [1][5]. Comparison with similar functions: - rate(v range-vector): Calculates the per-second average rate over the entire specified time range, providing a smoothed trend [1][6]. - irate(v range-vector): Calculates an instantaneous rate based only on the last two data points in the range [6][7]. It is more responsive to sudden spikes but significantly noisier, making it less suitable for most alerting [6][7]. - increase(v range-vector): Returns the total absolute increase in a counter over the specified time range, rather than a per-second average [6][7]. Note that rate(v) is mathematically equivalent to increase(v) / range_duration [6][3]. Example query: To get the per-second average rate of HTTP requests over the last 5 minutes: rate(http_requests_total[5m]) [1][8] Citations:
Correct the extension-cost alert units.
🤖 Prompt for AI Agents |
||
| ``` | ||
|
|
||
| Include this file in your `prometheus.yml`: | ||
|
|
||
| ```yaml | ||
| rule_files: | ||
| - 'sorokeep_alerts.yml' | ||
| ``` | ||
|
|
||
| ### Wiring to Alertmanager | ||
|
|
||
| Configure Alertmanager to route sorokeep alerts to your preferred channels (email, Slack, PagerDuty, etc.): | ||
|
|
||
| ```yaml | ||
| # alertmanager.yml | ||
| route: | ||
| group_by: ['alertname'] | ||
| receiver: 'slack-team' | ||
| routes: | ||
| - match: | ||
| alertname: 'SorokeepDown' | ||
| receiver: 'pagerduty-ops' | ||
|
|
||
| receivers: | ||
| - name: 'slack-team' | ||
| slack_configs: | ||
| - api_url: 'https://hooks.slack.com/services/...' | ||
| channel: '#sorokeep-alerts' | ||
| ``` | ||
|
|
||
| ## 5. Docker Compose (Full Stack) | ||
|
|
||
| For a complete local observability stack including Prometheus, Grafana, and Alertmanager alongside sorokeep: | ||
|
|
||
| ```yaml | ||
| version: "3.8" | ||
|
|
||
| services: | ||
| sorokeep: | ||
| build: . | ||
| command: ["sorokeep", "daemon", "--network", "testnet", "--metrics-port", "9464"] | ||
| ports: | ||
| - "9464:9464" | ||
| volumes: | ||
| - sorokeep-data:/root/.sorokeep | ||
| restart: unless-stopped | ||
|
|
||
| prometheus: | ||
| image: prom/prometheus:latest | ||
| ports: | ||
| - "9090:9090" | ||
| volumes: | ||
| - ./prometheus.yml:/etc/prometheus/prometheus.yml | ||
| - ./sorokeep_alerts.yml:/etc/prometheus/sorokeep_alerts.yml | ||
| restart: unless-stopped | ||
|
|
||
| grafana: | ||
| image: grafana/grafana:latest | ||
| ports: | ||
| - "3000:3000" | ||
| environment: | ||
| - GF_SECURITY_ADMIN_PASSWORD=admin | ||
| volumes: | ||
| - grafana-data:/var/lib/grafana | ||
| - ./resources/grafana:/etc/grafana/provisioning/dashboards:ro | ||
|
Comment on lines
+221
to
+229
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. 🎯 Functional Correctness | 🟠 Major | ⚡ Quick win 🧩 Analysis chain🏁 Script executed: #!/bin/bash
set -euo pipefail
echo "== repo files matching grafana/dashboard/prometheus =="
git ls-files | rg -i '(^|/)(grafana|dashboard|prometheus).*|^docs/observability\.md$|^compose|docker-compose|docker-compose\.yml|docker-compose\.yaml|docker-compose.*\.ya?ml$' || true
echo
echo "== docs/observability.md around relevant section =="
if [ -f docs/observability.md ]; then
nl -ba docs/observability.md | sed -n '180,260p'
fi
echo
echo "== search for grafana provisioning references =="
rg -n "grafana|provisioning|promotion|provider|datasource|dashboard\.json|Prometheus|metrics" docs compose . -g '!**/node_modules/**' -g '!**/.git/**' 2>/dev/null | head -200 || trueRepository: AbdulmalikAlayande/sorokeep Length of output: 419 🏁 Script executed: #!/bin/bash
set -euo pipefail
echo "== docs/observability.md around relevant section =="
python3 - <<'PY'
from pathlib import Path
p=Path("docs/observability.md")
lines=p.read_text(errors="replace").splitlines()
for i,line in enumerate(lines[179:259], start=180):
print(f"{i:4}: {line}")
PY
echo
echo "== Compose grafana/config sections =="
python3 - <<'PY'
from pathlib import Path
for path in ["docker-compose.yaml","docker-compose.devnet.yaml"]:
p=Path(path)
if not p.exists():
continue
lines=p.read_text(errors="replace").splitlines()
seen=[]
for i,line in enumerate(lines,1):
if "grafana" in line.lower() or "dashboard" in line.lower() or "prometheus" in line.lower() or "provisioning" in line.lower() or "datasource" in line.lower() or "GF_" in line:
seen.append((i,line))
print(f"--- {path} ({len(lines)} lines) ---")
for i,line in seen[:80]:
print(f"{i}: {line}")
PY
echo
echo "== provision resource files =="
for f in $(git ls-files | rg -i 'resources/grafana|grafana.*ya?ml|datasource|dashboard\.json'; true); do
echo "--- $f ($(wc -l < "$f") lines) ---"
sed -n '1,160p' "$f"
doneRepository: AbdulmalikAlayande/sorokeep Length of output: 3057 🌐 Web query:
💡 Result: Grafana provides a system for automated configuration known as provisioning, which allows you to define data sources and dashboards using YAML files [1][2]. These files are processed by Grafana upon startup and can be continuously managed, supporting GitOps workflows [1][2]. Provisioning Directories The provisioning system typically looks for configuration files in directories defined by your Grafana installation. By default, these are structured as follows [2]: provisioning/ datasources/ dashboards/ Data Source Provisioning Data source configuration files are placed in the provisioning/datasources directory [1]. Each file contains a list of data sources under the datasources key [1][3]. If a data source with the same name already exists, Grafana will update it to match the provided configuration [1]. Dashboard Provisioning Dashboard configuration files are placed in the provisioning/dashboards directory [1][4]. Each file specifies a list of providers that instruct Grafana on where to load dashboard files from the local filesystem [1][4]. An example dashboard provider YAML configuration looks like this [1][4]: apiVersion: 1 providers: - name: 'My Dashboard Provider' orgId: 1 folder: '' type: file disableDeletion: false updateIntervalSeconds: 10 allowUiUpdates: false options: path: /var/lib/grafana/dashboards foldersFromFilesStructure: true Key Configuration Options: - name: A unique name for the provider [1][4]. - type: The provider type, which defaults to file [1][4]. - options.path: The filesystem path where your dashboard JSON files are located [1][4]. - foldersFromFilesStructure: When set to true, Grafana maps the filesystem directory structure to folders within the Grafana UI [1][4]. To use this, you must leave the folder and folderUid options unset [4]. Security and Paths For security, Grafana restricts the paths that can be used for provisioning [5]. You must configure allowed paths in your grafana.ini or custom.ini file using the permitted_provisioning_paths option under the [paths] section [5][6]. This option defines which directories Grafana is allowed to access for provisioning resources [5]. Citations:
Add Grafana dashboard and data-source provisioning configs. Mounting 🤖 Prompt for AI Agents |
||
| restart: unless-stopped | ||
|
|
||
| alertmanager: | ||
| image: prom/alertmanager:latest | ||
| ports: | ||
| - "9093:9093" | ||
| volumes: | ||
| - ./alertmanager.yml:/etc/alertmanager/alertmanager.yml | ||
| restart: unless-stopped | ||
|
|
||
| volumes: | ||
| sorokeep-data: | ||
| grafana-data: | ||
| ``` | ||
|
|
||
| ## 6. Troubleshooting | ||
|
|
||
| ### Scrape Target Down | ||
|
|
||
| **Symptom:** Prometheus shows `sorokeep` target as `DOWN`. | ||
|
|
||
| **Causes and fixes:** | ||
|
|
||
| | Cause | Check | Fix | | ||
| |-------|-------|-----| | ||
| | Metrics port not exposed | `curl http://localhost:9464/metrics` on the sorokeep host | Verify `--metrics-port` is set | | ||
| | Firewall blocking | `nc -zv <host> 9464` from the Prometheus host | Open port 9464 in the firewall / security group | | ||
| | Container networking | `docker ps` and check sorokeep container port mapping | Add `ports: - "9464:9464"` to the compose service | | ||
| | sorokeep not running | `systemctl status sorokeep` or check process list | Restart the daemon | | ||
| | Prometheus config error | `promtool check config prometheus.yml` | Fix any syntax errors | | ||
|
|
||
| ### No Metrics Data | ||
|
|
||
| **Symptom:** Target is `UP` but no metrics appear in Prometheus or Grafana panels are empty. | ||
|
|
||
| - Ensure the daemon has been running long enough to collect data (at least one monitor cycle) | ||
| - Check `http://localhost:9464/metrics` directly — if metrics show zero values, the daemon hasn't completed a cycle yet | ||
| - Verify the Prometheus `metrics_path` matches the endpoint (default `/metrics`) | ||
|
|
||
| ### High Scrape Latency | ||
|
|
||
| **Symptom:** Prometheus scrape durations are > 10s. | ||
|
|
||
| For large deployments with many contracts, increase the scrape interval in `prometheus.yml`: | ||
|
|
||
| ```yaml | ||
| scrape_configs: | ||
| - job_name: 'sorokeep' | ||
| scrape_interval: 30s | ||
| scrape_timeout: 15s | ||
| ``` | ||
|
|
||
| ### Grafana Dashboard Not Found | ||
|
|
||
| **Symptom:** The import screen says "Dashboard not found." | ||
|
|
||
| Ensure you're importing the correct file from `resources/grafana/dashboard.json`. If the file doesn't exist, generate it by running `sorokeep metrics --grafana-dashboard` (requires implementation of the dashboard export feature), or manually create a dashboard in the Grafana UI using the metric names listed above. | ||
|
Comment on lines
+282
to
+286
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win Do not recommend an unimplemented dashboard-export command. The troubleshooting step explicitly says 🤖 Prompt for AI Agents |
||
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win
🧩 Analysis chain
🏁 Script executed:
Repository: AbdulmalikAlayande/sorokeep
Length of output: 413
🏁 Script executed:
Repository: AbdulmalikAlayande/sorokeep
Length of output: 24419
Remove the
/metricsentry from the Roadmap.docs/observability.mddocuments the/metricsendpoint as a built-in feature enabled with--metrics-port, so keeping it under “Roadmap” inREADME.mdmislabels existing work as future work.🤖 Prompt for AI Agents