Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -677,7 +677,7 @@ Track overall progress on the [Sorokeep Roadmap board](https://github.com/Abdulm
> If you're a maintainer, follow that doc to create and link the board.

- Plugin interface for alert channels — so a new channel (Matrix, MS Teams, email) doesn't require touching core dispatch code or the DB schema
- Prometheus `/metrics` endpoint for teams with existing observability stacks
- Prometheus `/metrics` endpoint for teams with existing observability stacks — see [Observability Setup Guide](docs/observability.md)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
set -euo pipefail

echo "== locate files =="
git ls-files | rg '(^|/)README\.md$|docs/observability\.md$|observability' || true

echo "== README relevant lines =="
if [ -f README.md ]; then
  sed -n '660,690p' README.md | nl -ba -v660
fi

echo "== observability guide occurrences =="
if [ -f docs/observability.md ]; then
  rg -n "(/metrics|metrics-port|--metrics-port|built-in|Roadmap|future|observability)" docs/observability.md -C 2 || true
  wc -l docs/observability.md
fi

echo "== code occurrences for /metrics and metrics-port =="
rg -n "(/metrics|--metrics-port|metrics-port|metricsPort|metrics)" --glob '!node_modules' --glob '!dist' --glob '!build' . -C 2 || true

Repository: AbdulmalikAlayande/sorokeep

Length of output: 413


🏁 Script executed:

#!/bin/bash
set -u

echo "== README relevant lines =="
if [ -f README.md ]; then
  sed -n '660,690p' README.md | awk '{printf "%6d  %s\n", NR+659, $0}'
fi

echo "== observability guide occurrences =="
if [ -f docs/observability.md ]; then
  rg -n "(/metrics|metrics-port|--metrics-port|built-in|Roadmap|future|observability)" docs/observability.md -C 2 || true
  wc -l docs/observability.md
fi

echo "== code occurrences for /metrics and metrics-port =="
rg -n "(/metrics|--metrics-port|metrics-port|metricsPort|metrics)" --glob '!node_modules' --glob '!dist' --glob '!build' . -C 2 || true

Repository: AbdulmalikAlayande/sorokeep

Length of output: 24419


Remove the /metrics entry from the Roadmap.

docs/observability.md documents the /metrics endpoint as a built-in feature enabled with --metrics-port, so keeping it under “Roadmap” in README.md mislabels existing work as future work.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@README.md` at line 680, Remove the Prometheus `/metrics` bullet from the
Roadmap section in README.md, leaving the existing observability documentation
reference and other roadmap entries unchanged.

- Reusable GitHub Action wrapping `sorokeep check` for CI-integrated TTL checks
- Web dashboard for visual TTL monitoring
- Multi-contract batch operations
Expand Down
2 changes: 2 additions & 0 deletions docs/ARCHITECTURE.md
Original file line number Diff line number Diff line change
Expand Up @@ -84,4 +84,6 @@ Everything sorokeep knows lives in one SQLite file (`~/.sorokeep/sorokeep.db`, s
| Change to threshold/extension logic | `src/core/monitor.ts` or `src/core/extension.ts` — read the fault-isolation notes above before touching these |
| New RPC-derived data | `src/rpc/client.ts` first, then whatever core module consumes it |

For production deployments, see [Observability Setup](observability.md) — Prometheus scraping, Grafana dashboards, and Alertmanager rules.

If a change doesn't fit cleanly into one of these rows, that's usually a sign to open an issue and discuss the approach before writing code — see [CONTRIBUTING.md](../CONTRIBUTING.md#larger-contributions).
286 changes: 286 additions & 0 deletions docs/observability.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,286 @@
# Observability Setup: Prometheus + Grafana

This guide walks through connecting sorokeep's built-in metrics endpoint to a Prometheus + Grafana observability stack. After completing it you'll have a working dashboard showing contract TTL health, alert activity, and extension costs — all updating automatically from the daemon's metrics.

## Prerequisites

- A running sorokeep daemon (`sorokeep daemon`)
- [Prometheus](https://prometheus.io/download/) 2.x+
- [Grafana](https://grafana.com/grafana/download/) 10.x+
- Network connectivity between Prometheus and the sorokeep host

## 1. Enable the Metrics Server

Start the daemon with the `--metrics-port` flag to expose a Prometheus-format `/metrics` endpoint:

```bash
sorokeep daemon --network testnet --metrics-port 9464
```

The metrics server also exposes:
- `/healthz` — liveness probe (always 200 while the process is running)
- `/readyz` — readiness probe (200 when DB and RPC are reachable, 503 otherwise)

You can verify the endpoint is working:

```bash
curl http://localhost:9464/metrics
```

Sample output:

```

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win

Specify the fenced block language.

This sample output triggers MD040; mark it as text (or console) so Markdown lint passes.

🧰 Tools
🪛 markdownlint-cli2 (0.23.1)

[warning] 32-32: Fenced code blocks should have a language specified

(MD040, fenced-code-language)

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@docs/observability.md` at line 32, Update the fenced code block in the
observability documentation sample to specify a language identifier, using text
or console, so the Markdown lint rule passes.

Source: Linters/SAST tools

# HELP sorokeep_contracts_total Total number of watched contracts
# TYPE sorokeep_contracts_total gauge
sorokeep_contracts_total{network="testnet"} 3

# HELP sorokeep_contract_entries_total Total number of contract entries across all contracts
# TYPE sorokeep_contract_entries_total gauge
sorokeep_contract_entries_total 10

# HELP sorokeep_extensions_total Total number of TTL extensions performed
# TYPE sorokeep_extensions_total counter
sorokeep_extensions_total 42

# HELP sorokeep_alerts_fired_total Total number of alerts fired
# TYPE sorokeep_alerts_fired_total counter
sorokeep_alerts_fired_total 7
```

### Metrics Port Reference

| Flag | Default | Description |
|------|---------|-------------|
| `--metrics-port` | Disabled | Port for the Prometheus metrics HTTP server |
| `--metrics-host` | `0.0.0.0` | Bind address for the metrics server |
Comment on lines +50 to +55

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔒 Security & Privacy | 🟠 Major | ⚡ Quick win

Do not present 0.0.0.0 as an unqualified default.

The unauthenticated metrics endpoint exposes operational data and is reachable on every interface. Recommend binding to a private interface by default, or explicitly document firewall/reverse-proxy restrictions before exposing this port.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@docs/observability.md` around lines 50 - 55, Update the Metrics Port
Reference documentation for `--metrics-host` so it does not present `0.0.0.0` as
an unqualified default; recommend a private-interface binding or explicitly
state the firewall and reverse-proxy restrictions required before exposing the
unauthenticated metrics endpoint.


## 2. Configure Prometheus to Scrape sorokeep

Create or edit `prometheus.yml`:

```yaml
scrape_configs:
- job_name: 'sorokeep'
scrape_interval: 15s
scrape_timeout: 10s
metrics_path: /metrics
static_configs:
- targets:
- 'localhost:9464'
labels:
service: sorokeep
network: testnet
```

Replace `localhost` with the sorokeep host address if Prometheus is running on a different machine. If sorokeep runs in a Docker container, use the container name or host's Docker bridge IP.
Comment on lines +67 to +75

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟠 Major | ⚡ Quick win

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
set -euo pipefail

echo "== files =="
git ls-files | rg '(^|/)(prometheus|docker-compose|compose|observability)' || true

echo "== relevant docs slice =="
sed -n '1,260p' docs/observability.md 2>/dev/null | nl -ba | sed -n '1,260p'

echo "== search sorokeep/prometheus compose =="
rg -n "sorokeep|prometheus|9464|statsd|docker-compose|compose" -S . --glob '!node_modules' --glob '!target' | head -n 200

Repository: AbdulmalikAlayande/sorokeep

Length of output: 348


🏁 Script executed:

#!/bin/bash
set -euxo pipefail

echo "== docs/observability.md relevant lines =="
awk '{printf "%6d\t%s\n", NR, $0}' docs/observability.md 2>/dev/null | sed -n '1,280p'

echo "== compose files =="
for f in docker-compose.yaml docker-compose.devnet.yaml; do
  echo "--- $f ($(wc -l < "$f") lines) ---"
  awk '{printf "%6d\t%s\n", NR, $0}' "$f" | sed -n '1,260p'
done

echo "== search for targets and ports =="
rg -n "targets:|localhost:9464|sorokeep:9464|9464|sorokeep|prometheus|statsd|expose:|ports:" -S . --glob '!node_modules' --glob '!target' | head -n 240

Repository: AbdulmalikAlayande/sorokeep

Length of output: 40020


Use the Compose service name for Prometheus scraping.

Section 5 starts an observability Docker Compose stack with separate Prometheus and sorokeep services, so the localhost:9464 target is inside the Prometheus container. Update the Compose example Prometheus config to scrape the sorokeep service instead, e.g. targets: ['sorokeep:9464'].

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@docs/observability.md` around lines 67 - 75, Update the Prometheus scrape
target in the observability Docker Compose example from localhost:9464 to the
sorokeep Compose service address sorokeep:9464, while preserving the existing
labels and configuration structure.


Start Prometheus with the config:

```bash
prometheus --config.file=prometheus.yml
```

Verify targets are up at `http://localhost:9090/targets` — the `sorokeep` job should show `UP`.

### Available Metrics

| Metric | Type | Description |
|--------|------|-------------|
| `sorokeep_contracts_total` | Gauge | Watched contracts, labelled by network |
| `sorokeep_contract_entries_total` | Gauge | Total tracked ledger entries |
| `sorokeep_extensions_total` | Counter | TTL extensions performed |
| `sorokeep_extension_cost_xlm_total` | Counter | Total XLM spent on extensions |
| `sorokeep_alerts_fired_total` | Counter | Alerts fired across all contracts |
| `sorokeep_alerts_unresolved_total` | Gauge | Currently unresolved alerts |
| `sorokeep_channel_accounts_total` | Gauge | Configured channel accounts, labelled by network |

## 3. Import the Grafana Dashboard

A pre-built Grafana dashboard for sorokeep is available at `resources/grafana/dashboard.json`.

### Via the Grafana UI

1. Open Grafana (`http://localhost:3000`, default login `admin`/`admin`)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔒 Security & Privacy | 🟠 Major | ⚡ Quick win

Remove the admin/admin deployment path.

The guide documents the default Grafana password, sends it in a curl command, and hard-codes it in Compose. This encourages an immediately compromised Grafana instance. Require a secret/environment-provided password and use a token or environment variable for API imports.

Suggested Compose change
-      - GF_SECURITY_ADMIN_PASSWORD=admin
+      - GF_SECURITY_ADMIN_PASSWORD=${GRAFANA_ADMIN_PASSWORD:?Set GRAFANA_ADMIN_PASSWORD}

Also applies to: 111-120, 221-226

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@docs/observability.md` at line 103, Remove all admin/admin defaults from the
Grafana setup documentation and Compose configuration, requiring the Grafana
password to come from a secret or environment variable. Update the API import
curl commands to use a token or environment-provided credential, and revise the
referenced Grafana setup sections consistently.

2. Navigate to **Dashboards → New → Import**
3. Upload `resources/grafana/dashboard.json` or paste its contents
4. Select the Prometheus data source that's scraping sorokeep
5. Click **Import**

### Via the API

```bash
curl -X POST http://admin:admin@localhost:3000/api/dashboards/db \
-H "Content-Type: application/json" \
-d @- <<EOF
{
"dashboard": $(cat resources/grafana/dashboard.json),
"overwrite": true,
"message": "Imported sorokeep dashboard"
}
EOF
```

### Dashboard Panels

| Panel | Description |
|-------|-------------|
| **Watched Contracts** | Total contracts per network (stat + sparkline) |
| **Entries Tracked** | Total tracked entries across all contracts |
| **Extensions (24h)** | Extension count and XLM cost over the last 24 hours |
| **Alerts Fired (24h)** | Alert count by severity over time |
| **Unresolved Alerts** | Current unresolved alert count |
| **Top Contracts by Cost** | Contracts with the highest extension costs |
| **TTL Distribution** | Remaining TTL distribution across entries |

## 4. Alerting with Alertmanager

### Example Alert Rules

Create `sorokeep_alerts.yml`:

```yaml
groups:
- name: sorokeep
rules:
- alert: SorokeepDown
expr: up{job="sorokeep"} == 0
for: 1m
annotations:
summary: "sorokeep instance {{ $labels.instance }} is down"
description: "Prometheus target {{ $labels.job }}/{{ $labels.instance }} has been unreachable for over 1 minute."

- alert: SorokeepHighUnresolvedAlerts
expr: sorokeep_alerts_unresolved_total > 0
for: 5m
annotations:
summary: "Contract alerts are not being resolved"
description: "{{ $value }} alerts have been unresolved for more than 5 minutes."

- alert: SorokeepExtensionCostSpike
expr: rate(sorokeep_extension_cost_xlm_total[1h]) > 1
for: 10m
annotations:
summary: "High extension cost rate detected"
description: "XLM extension cost rate is {{ $value }} XLM/hour — investigate unusual extension activity."
Comment on lines +159 to +164

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟠 Major | ⚡ Quick win

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
set -euo pipefail

echo "== File snippet =="
sed -n '140,175p' docs/observability.md | cat -n

echo
echo "== Related counter names =="
rg -n "sorokeep_extension_cost|rate\\(|XLM|cost" docs/observability.md prometheus 2>/dev/null || true

Repository: AbdulmalikAlayande/sorokeep

Length of output: 2479


🌐 Web query:

Prometheus rate counter per second documentation

💡 Result:

The rate function in Prometheus is a primary tool for calculating the per-second average rate of increase of a counter metric over a specified time range [1][2]. Key characteristics and usage: - Functionality: It takes a range vector (e.g., [5m]) and calculates the increase in the counter's value over that window, divided by the duration of the window to produce a per-second rate [1][3][2]. - Counter Resets: It automatically detects and compensates for counter resets (such as when a process restarts), ensuring the rate calculation remains accurate [1][4]. - Extrapolation: The function extrapolates to the ends of the time range, which helps account for missed scrapes or misalignment between scrape cycles and the requested time window [1][2]. - Best Practices: - Use only with counters (or native histograms where components behave like counters) [1][5]. - It is recommended for alerting and graphing slow-moving counters [1][4]. - When combining rate with an aggregation operator (like sum) or an _over_time function, always apply rate first to the raw counter before aggregating [1][5]. Comparison with similar functions: - rate(v range-vector): Calculates the per-second average rate over the entire specified time range, providing a smoothed trend [1][6]. - irate(v range-vector): Calculates an instantaneous rate based only on the last two data points in the range [6][7]. It is more responsive to sudden spikes but significantly noisier, making it less suitable for most alerting [6][7]. - increase(v range-vector): Returns the total absolute increase in a counter over the specified time range, rather than a per-second average [6][7]. Note that rate(v) is mathematically equivalent to increase(v) / range_duration [6][3]. Example query: To get the per-second average rate of HTTP requests over the last 5 minutes: rate(http_requests_total[5m]) [1][8]

Citations:


Correct the extension-cost alert units.

rate(sorokeep_extension_cost_xlm_total[1h]) returns XLM/second, but the threshold is 1 XLM/second and the notification says XLM/hour. Either compare rate(...) * 3600 against an hourly threshold, or change the annotation to XLM/second so the alert and message match.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@docs/observability.md` around lines 159 - 164, Update the
SorokeepExtensionCostSpike alert so its Prometheus expression and notification
use consistent units: either multiply
rate(sorokeep_extension_cost_xlm_total[1h]) by 3600 and retain an hourly
threshold/message, or keep the expression unchanged and change the threshold and
description to XLM/second.

```

Include this file in your `prometheus.yml`:

```yaml
rule_files:
- 'sorokeep_alerts.yml'
```

### Wiring to Alertmanager

Configure Alertmanager to route sorokeep alerts to your preferred channels (email, Slack, PagerDuty, etc.):

```yaml
# alertmanager.yml
route:
group_by: ['alertname']
receiver: 'slack-team'
routes:
- match:
alertname: 'SorokeepDown'
receiver: 'pagerduty-ops'

receivers:
- name: 'slack-team'
slack_configs:
- api_url: 'https://hooks.slack.com/services/...'
channel: '#sorokeep-alerts'
```

## 5. Docker Compose (Full Stack)

For a complete local observability stack including Prometheus, Grafana, and Alertmanager alongside sorokeep:

```yaml
version: "3.8"

services:
sorokeep:
build: .
command: ["sorokeep", "daemon", "--network", "testnet", "--metrics-port", "9464"]
ports:
- "9464:9464"
volumes:
- sorokeep-data:/root/.sorokeep
restart: unless-stopped

prometheus:
image: prom/prometheus:latest
ports:
- "9090:9090"
volumes:
- ./prometheus.yml:/etc/prometheus/prometheus.yml
- ./sorokeep_alerts.yml:/etc/prometheus/sorokeep_alerts.yml
restart: unless-stopped

grafana:
image: grafana/grafana:latest
ports:
- "3000:3000"
environment:
- GF_SECURITY_ADMIN_PASSWORD=admin
volumes:
- grafana-data:/var/lib/grafana
- ./resources/grafana:/etc/grafana/provisioning/dashboards:ro
Comment on lines +221 to +229

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟠 Major | ⚡ Quick win

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
set -euo pipefail

echo "== repo files matching grafana/dashboard/prometheus =="
git ls-files | rg -i '(^|/)(grafana|dashboard|prometheus).*|^docs/observability\.md$|^compose|docker-compose|docker-compose\.yml|docker-compose\.yaml|docker-compose.*\.ya?ml$' || true

echo
echo "== docs/observability.md around relevant section =="
if [ -f docs/observability.md ]; then
  nl -ba docs/observability.md | sed -n '180,260p'
fi

echo
echo "== search for grafana provisioning references =="
rg -n "grafana|provisioning|promotion|provider|datasource|dashboard\.json|Prometheus|metrics" docs compose . -g '!**/node_modules/**' -g '!**/.git/**' 2>/dev/null | head -200 || true

Repository: AbdulmalikAlayande/sorokeep

Length of output: 419


🏁 Script executed:

#!/bin/bash
set -euo pipefail

echo "== docs/observability.md around relevant section =="
python3 - <<'PY'
from pathlib import Path
p=Path("docs/observability.md")
lines=p.read_text(errors="replace").splitlines()
for i,line in enumerate(lines[179:259], start=180):
    print(f"{i:4}: {line}")
PY

echo
echo "== Compose grafana/config sections =="
python3 - <<'PY'
from pathlib import Path
for path in ["docker-compose.yaml","docker-compose.devnet.yaml"]:
    p=Path(path)
    if not p.exists():
        continue
    lines=p.read_text(errors="replace").splitlines()
    seen=[]
    for i,line in enumerate(lines,1):
        if "grafana" in line.lower() or "dashboard" in line.lower() or "prometheus" in line.lower() or "provisioning" in line.lower() or "datasource" in line.lower() or "GF_" in line:
            seen.append((i,line))
    print(f"--- {path} ({len(lines)} lines) ---")
    for i,line in seen[:80]:
        print(f"{i}: {line}")
PY

echo
echo "== provision resource files =="
for f in $(git ls-files | rg -i 'resources/grafana|grafana.*ya?ml|datasource|dashboard\.json'; true); do
  echo "--- $f ($(wc -l < "$f") lines) ---"
  sed -n '1,160p' "$f"
done

Repository: AbdulmalikAlayande/sorokeep

Length of output: 3057


🌐 Web query:

Grafana provisioning documentation dashboard providers datasource provisioning dashboards directory YAML provider

💡 Result:

Grafana provides a system for automated configuration known as provisioning, which allows you to define data sources and dashboards using YAML files [1][2]. These files are processed by Grafana upon startup and can be continuously managed, supporting GitOps workflows [1][2]. Provisioning Directories The provisioning system typically looks for configuration files in directories defined by your Grafana installation. By default, these are structured as follows [2]: provisioning/ datasources/ dashboards/ Data Source Provisioning Data source configuration files are placed in the provisioning/datasources directory [1]. Each file contains a list of data sources under the datasources key [1][3]. If a data source with the same name already exists, Grafana will update it to match the provided configuration [1]. Dashboard Provisioning Dashboard configuration files are placed in the provisioning/dashboards directory [1][4]. Each file specifies a list of providers that instruct Grafana on where to load dashboard files from the local filesystem [1][4]. An example dashboard provider YAML configuration looks like this [1][4]: apiVersion: 1 providers: - name: 'My Dashboard Provider' orgId: 1 folder: '' type: file disableDeletion: false updateIntervalSeconds: 10 allowUiUpdates: false options: path: /var/lib/grafana/dashboards foldersFromFilesStructure: true Key Configuration Options: - name: A unique name for the provider [1][4]. - type: The provider type, which defaults to file [1][4]. - options.path: The filesystem path where your dashboard JSON files are located [1][4]. - foldersFromFilesStructure: When set to true, Grafana maps the filesystem directory structure to folders within the Grafana UI [1][4]. To use this, you must leave the folder and folderUid options unset [4]. Security and Paths For security, Grafana restricts the paths that can be used for provisioning [5]. You must configure allowed paths in your grafana.ini or custom.ini file using the permitted_provisioning_paths option under the [paths] section [5][6]. This option defines which directories Grafana is allowed to access for provisioning resources [5].

Citations:


Add Grafana dashboard and data-source provisioning configs.

Mounting resources/grafana into dashboards/ only makes dashboard files visible; Grafana still needs a provisioning/dashboards/<provider>.yaml pointing to the mounted path and a provisioning/datasources/prometheus.yaml datasource. Without these, the full-stack compose example starts Grafana but does not automatically import dashboard.json or make Prometheus available.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@docs/observability.md` around lines 221 - 229, Add Grafana provisioning
configuration alongside the compose example: create a dashboard provider YAML
under provisioning/dashboards that points to the mounted
/etc/grafana/provisioning/dashboards directory, and add a Prometheus datasource
YAML under provisioning/datasources with the compose service URL. Ensure the
existing grafana volume mount exposes these provisioning files so dashboard.json
is imported automatically and Prometheus is available.

restart: unless-stopped

alertmanager:
image: prom/alertmanager:latest
ports:
- "9093:9093"
volumes:
- ./alertmanager.yml:/etc/alertmanager/alertmanager.yml
restart: unless-stopped

volumes:
sorokeep-data:
grafana-data:
```

## 6. Troubleshooting

### Scrape Target Down

**Symptom:** Prometheus shows `sorokeep` target as `DOWN`.

**Causes and fixes:**

| Cause | Check | Fix |
|-------|-------|-----|
| Metrics port not exposed | `curl http://localhost:9464/metrics` on the sorokeep host | Verify `--metrics-port` is set |
| Firewall blocking | `nc -zv <host> 9464` from the Prometheus host | Open port 9464 in the firewall / security group |
| Container networking | `docker ps` and check sorokeep container port mapping | Add `ports: - "9464:9464"` to the compose service |
| sorokeep not running | `systemctl status sorokeep` or check process list | Restart the daemon |
| Prometheus config error | `promtool check config prometheus.yml` | Fix any syntax errors |

### No Metrics Data

**Symptom:** Target is `UP` but no metrics appear in Prometheus or Grafana panels are empty.

- Ensure the daemon has been running long enough to collect data (at least one monitor cycle)
- Check `http://localhost:9464/metrics` directly — if metrics show zero values, the daemon hasn't completed a cycle yet
- Verify the Prometheus `metrics_path` matches the endpoint (default `/metrics`)

### High Scrape Latency

**Symptom:** Prometheus scrape durations are > 10s.

For large deployments with many contracts, increase the scrape interval in `prometheus.yml`:

```yaml
scrape_configs:
- job_name: 'sorokeep'
scrape_interval: 30s
scrape_timeout: 15s
```

### Grafana Dashboard Not Found

**Symptom:** The import screen says "Dashboard not found."

Ensure you're importing the correct file from `resources/grafana/dashboard.json`. If the file doesn't exist, generate it by running `sorokeep metrics --grafana-dashboard` (requires implementation of the dashboard export feature), or manually create a dashboard in the Grafana UI using the metric names listed above.
Comment on lines +282 to +286

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Do not recommend an unimplemented dashboard-export command.

The troubleshooting step explicitly says sorokeep metrics --grafana-dashboard requires an unimplemented feature, so users following the recovery path will receive a command failure. Remove that command and point users to the checked-in dashboard file or manual import.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@docs/observability.md` around lines 282 - 286, Update the “Grafana Dashboard
Not Found” troubleshooting section to remove the unimplemented sorokeep metrics
--grafana-dashboard command. Direct users to import the checked-in
resources/grafana/dashboard.json file, or manually create/import a dashboard in
Grafana using the listed metric names.