Skip to content

feat(managed): spawn /dp/heartbeat worker at boot - #31

Merged
moonming merged 1 commit into
mainfrom
feat/dp-heartbeat
Apr 23, 2026
Merged

feat(managed): spawn /dp/heartbeat worker at boot#31
moonming merged 1 commit into
mainfrom
feat/dp-heartbeat

Conversation

@moonming

Copy link
Copy Markdown
Member

Summary

DP half of the liveness channel. Paired with the cp-api handler at api7/AISIX-Cloud#10. Stacks on #30 (/dp/register client); retargets to main when #30 merges.

After managed-mode bootstrap completes (either register or confirm-existing-bundle), main spawns one tokio task that POSTs /dp/heartbeat on a fixed interval for the life of the process.

Request

POST <heartbeat_url>
  Authorization: Bearer <dp_id>
  Content-Type: application/json

  { "dp_id": "...", "uptime_seconds": 1234, "version": "0.5.2" }

Auth is the DP id as Bearer (Phase 1). Phase 2 upgrades to mTLS client cert once cp-api terminates mTLS.

Worker lifecycle

  • Interval: value returned by the register response, clamped to [5s, 300s] as defence against a buggy CP reply (0 = burst loop, a week = useless).
  • Tick cadence: MissedTickBehavior::Delay — a slow beat doesn't burst catch-up beats afterwards.
  • Failure handling: individual failures log warn! and the ticker keeps running. A transient CP outage means "no dashboard update", not "DP stops trying".
  • Shutdown: tokio::select! on the shared watch::Receiver<bool>; graceful shutdown drains the in-flight request inside the 2s grace window.

Boot paths

Scenario Source of HeartbeatConfig
First boot (registration just ran) Registered { heartbeat_url, dp_id, heartbeat_interval } from the register response
Subsequent boot (cert bundle already on disk) dp_id read from managed.dp_id_file; URL = managed.cp_base_url + /dp/heartbeat; default interval 15s

If dp_id_file is unreadable/empty on the subsequent-boot path, the heartbeat worker is disabled with a warning — the DP should still proxy traffic, the Gateway page just won't see it as "live".

Files

  • crates/aisix-server/src/heartbeat.rs (new, ~220 lines) — config + spawn/run/send + tests
  • crates/aisix-server/src/main.rsmod heartbeat, pre-etcd heartbeat_cfg block, heartbeat_task awaited alongside watch_task at shutdown, load_heartbeat_config_from_disk helper

Tests (cargo test --workspace green, cargo clippy -D warnings clean)

Test Covers
send_posts_dp_id_and_bearer wiremock asserts Authorization header + body field shape matches cp-api's parser
send_propagates_non_success_body CP error body surfaces into the anyhow chain — operators see DP_NOT_FOUND etc. without decoding logs
run_stops_on_cancel spawn worker → observe first beat → flip cancel → task joins within 2s
sanitised_interval_clamps_extremes 10ms → 5s, 86400s → 300s

Explicitly out of scope

  • /dp/telemetry (next PR)
  • Local config snapshot so proxy serves from cache when etcd is unreachable (prd-09 §9.7.2)
  • mTLS upgrade for heartbeat auth (Phase 2, paired with the cp-api mTLS listener once that lands)

Copilot AI review requested due to automatic review settings April 23, 2026 09:36

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds the DP-side managed-mode liveness heartbeat worker so the control plane can track that a data plane instance is alive, complementing the existing managed-mode bootstrap/registration flow.

Changes:

  • Introduces a new heartbeat module implementing periodic POST /dp/heartbeat with interval clamping and wiremock-based tests.
  • Extends main managed-mode bootstrap to derive HeartbeatConfig from either registration response (first boot) or persisted dp_id + cp_base_url (subsequent boots).
  • Spawns the heartbeat task during startup and awaits it during shutdown.

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated 4 comments.

File Description
crates/aisix-server/src/main.rs Computes heartbeat config during managed bootstrap; spawns and joins heartbeat task; adds load_heartbeat_config_from_disk.
crates/aisix-server/src/heartbeat.rs New worker implementation (spawn/run/send) for periodic heartbeat POSTs, plus unit tests.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +270 to +271
if let Some(task) = heartbeat_task {
let _ = task.await;
Comment on lines +95 to +98
_ = cancel.changed() => {
if *cancel.borrow() {
tracing::info!("heartbeat shutting down");
return;
Comment on lines +90 to +104
match send(&client, &cfg, uptime).await {
Ok(()) => tracing::debug!("heartbeat ok"),
Err(e) => tracing::warn!(error = %e, "heartbeat failed"),
}
}
_ = cancel.changed() => {
if *cancel.borrow() {
tracing::info!("heartbeat shutting down");
return;
}
}
}
}
}

Ok(h) => Some(h),
Err(e) => {
tracing::warn!(error = %e,
"managed mode: heartbeat worker disabled (dp_id unreadable)");
@moonming
moonming changed the base branch from feat/dp-register to main April 23, 2026 12:31
DP half of the liveness channel. Paired with the cp-api handler at
api7/AISIX-Cloud#10. Stacks on #30 (/dp/register client).

## Behaviour

After the managed-mode bootstrap either registers or confirms an
existing bundle on disk, `main` spawns one tokio task that POSTs
`/dp/heartbeat` on a fixed interval. The task:

  - ticks at the interval returned by the register response (clamped
    to [5s, 300s] as defence against a buggy CP reply)
  - uses MissedTickBehavior::Delay so a slow tick doesn't burst
    catch-up beats afterwards
  - logs individual failures as warnings and keeps running; a
    transient CP outage means "no dashboard update" but not "DP
    stops trying"
  - cancels via the shared `watch::Receiver<bool>` so graceful
    shutdown drains the in-flight request

Request body: `{ dp_id, uptime_seconds, version }`.
Auth: `Authorization: Bearer <dp_id>` (Phase 1; Phase 2 upgrades to
mTLS once cp-api terminates mTLS).

## Boot paths (both covered)

1. **First boot**: register returns `Registered` with heartbeat_url +
   dp_id + interval; these flow straight into `HeartbeatConfig`.
2. **Subsequent boot**: bundle already on disk; `dp_id` is read from
   `managed.dp_id_file` and the URL is synthesised from
   `managed.cp_base_url` with a default 15s interval.
   Failure to read dp_id → worker disabled with a warning, not a
   hard boot failure (the DP should still proxy traffic).

## Files

- `crates/aisix-server/src/heartbeat.rs` (new)
  - `HeartbeatConfig` + `sanitised()` interval clamping
  - `spawn()` / `run()` / `send()` split so tests can drive each step
  - HTTP client built inside the worker (per-instance, not shared —
    the worker owns its lifetime)
- `crates/aisix-server/src/main.rs`
  - `mod heartbeat;`
  - New pre-etcd block constructing `heartbeat_cfg: Option<_>`
  - `heartbeat_task: Option<JoinHandle>` awaited at the end of run()
  - `load_heartbeat_config_from_disk()` helper for the "bundle exists
    from a prior boot" branch

## Tests (`cargo test --workspace` green, `cargo clippy -D warnings` clean)

- `heartbeat::tests::send_posts_dp_id_and_bearer` — wiremock asserts
  on the Authorization header + body fields the CP handler uses.
- `heartbeat::tests::send_propagates_non_success_body` — CP error
  body surfaces into the anyhow chain so operators see which error
  code fired without decoding logs.
- `heartbeat::tests::run_stops_on_cancel` — spawn the worker, observe
  a successful beat, flip cancel, assert the task returns inside a
  2s grace window.
- `heartbeat::tests::sanitised_interval_clamps_extremes` — 10ms → 5s
  and 86400s → 300s as sanity bounds.

## Explicitly out of scope (tracked)

- `/dp/telemetry` (next PR)
- Local config snapshot so proxy serves from cache when etcd is
  unreachable (prd-09 §9.7.2)
- mTLS upgrade for heartbeat auth (Phase 2; paired with the cp-api
  mTLS listener once that lands)
@moonming
moonming force-pushed the feat/dp-heartbeat branch from e519423 to ede0a69 Compare April 23, 2026 12:31
@moonming
moonming merged commit b597129 into main Apr 23, 2026
2 of 4 checks passed
moonming added a commit that referenced this pull request Apr 23, 2026
The same Docker image now serves both standalone and managed
(aisix.cloud tenant) deployments. Two pieces:

- config.managed.yaml — bootstrap template baked at
  /etc/aisix/config.managed.yaml. Has placeholder etcd endpoint
  (overwritten by /dp/register response), managed.enabled = true,
  and unbindable admin (defence-in-depth if managed mode somehow
  flipped off). All real per-DP secrets come from env vars.

- docker/entrypoint.sh — picks the config file via AISIX_CONFIG_PATH
  (default /etc/aisix/config.yaml). Standalone users mount their
  config at the default path; managed users point AISIX_CONFIG_PATH
  at the baked file and inject AISIX_MANAGED__REGISTRATION_TOKEN +
  AISIX_MANAGED__CP_BASE_URL.

Existing main.rs bootstrap (PR #30 + #31) already does the rest:
register-and-persist on first boot, reload bundle on subsequent
boots, spawn heartbeat worker.

Tests: parses_managed_block_with_register_fields locks the YAML
shape so any new required field on ManagedConfig fails CI loudly
instead of silently breaking the image.

Docs: docs/managed-mode.md walks operators through first boot,
restart semantics, env-var override matrix, and common errors.

This unblocks AISIX-Cloud E2E scenarios 2/3/4 — the test harness
can now `docker run` aisix with a deployment_token and have the DP
register itself without prebaked certs.
@jarvis9443
jarvis9443 deleted the feat/dp-heartbeat branch June 25, 2026 06:25
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants