feat(health): serve operator probes until accept fails - #111
cursor[bot] wants to merge 13 commits into
Conversation
RED: the domain health snapshot is not yet exposed as GET /live and GET /ready with fail-closed readiness and a safe as-built OpenAPI 3.2 contract.
Map RuntimeHealthSnapshot to GET /live and GET /ready with fail-closed readiness, RFC 9457 problem details for unsupported requests, and an as-built OpenAPI 3.2.0 document that lists only those operations.
Exercise HealthHttpResponse::content_type and GET /ready?capability= from the library test module so Linux line coverage can reach the remaining unexecuted instantiations.
* feat(health): bind a TCP listener for operator probes GET /live and GET /ready already exist as a request translator. Bind a blocking listener that serves one HTTP/1.1 request per accepted connection without adding public product routes, TLS, or SLO values. * docs(traceability): name Active PR #92 on the bound health listener * feat(health): answer probes from a PostgreSQL operational snapshot * test(health): require a composed PostgreSQL operational snapshot RED: runtime and relation probes exist, but callers still assemble liveness, backlog, integrity, and the postgres capability by hand. * feat(health): compose PostgreSQL probes into one operational snapshot Map runtime and relation probes plus caller-supplied backlog into a RuntimeHealthSnapshot. Probe failure becomes unknown/unready evidence and never returns a driver error. * feat(health): answer probes from a PostgreSQL operational snapshot Rebuild #95 on the current #92 listener plus the #93 snapshot composer. handle_postgres_health_http_request and accept_one_postgres_health_http observe the caller-owned store and reuse the existing probe translator. * docs(health): name #95 snapshot wiring on the current listener stack Restore the architecture mappings the rewrite dropped: TRACEABILITY module and capability rows, OPERABILITY fail-closed snapshot readiness, ADR-0014 as-built status, and the composed-snapshot changelog line. Co-authored-by: Seongho Bae <seonghobae@users.noreply.github.com> --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Seongho Bae <seonghobae@users.noreply.github.com> --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Seongho Bae <seonghobae@users.noreply.github.com>
The inherited #72 recovery fixture inserted a processing consumption row without claim_deadline_at. Migration 0019 requires that column for processing rows, and the deadline trigger is UPDATE-only, so exact-head CI failed closed. Seed a valid persisted claim and assert the deadline survives COPY restore. Co-authored-by: Seongho Bae <seonghobae@users.noreply.github.com>
…port RED: GET /live must not observe PostgreSQL, bare GET /ready must fail closed on a read-only store, incomplete or oversized requests must return 400 without echoing input, and the as-built OpenAPI /live operation must document HTTP 503. Co-authored-by: Seongho Bae <seonghobae@users.noreply.github.com>
Observe PostgreSQL only for GET /ready, after accept. Bare /ready requires postgres_operational_store. The listener applies a 2-second I/O timeout, rejects incomplete or oversized requests without echo, and adds Cache-Control: no-store plus Allow: GET on 405. Co-authored-by: Seongho Bae <seonghobae@users.noreply.github.com>
Operators can now run a blocking accept loop so a load balancer can keep asking GET /live and GET /ready. Interrupted accepts retry; any other accept, read, or write error stops the loop. TLS and keep-alive stay out of this slice. Co-authored-by: Seongho Bae <seonghobae@users.noreply.github.com>
Keep #91 as the predecessor translator head. Operators should land the serve-loop successor so a load balancer can keep probing. Co-authored-by: Seongho Bae <seonghobae@users.noreply.github.com>
There was a problem hiding this comment.
Do not merge #111 as-is. The serve loop stops on per-connection read/write errors, so one dropped load-balancer probe can take down GET /live and GET /ready for every later check.
Keep #91/#102 parked. Land a successor that keeps accepting after connection I/O errors and stops only when accept itself fails (Interrupted/ConnectionAborted/ConnectionReset retry; WouldBlock and other accept failures still stop so a closed or non-blocking listener does not spin).
Independent last-push approval is still required. Do not self-approve.
Sent by Cursor Automation: Fix Issues
| fn classify_serve_accept(result: io::Result<()>) -> ServeAcceptProgress { | ||
| match result { | ||
| Ok(()) => ServeAcceptProgress::Continue, | ||
| Err(error) if error.kind() == io::ErrorKind::Interrupted => ServeAcceptProgress::Continue, | ||
| Err(error) => ServeAcceptProgress::Stop(error), |
There was a problem hiding this comment.
classify_serve_accept treats every non-Interrupted accept_one_health_http_with error as fatal, including BrokenPipe, ConnectionReset, and write timeout after a successful accept.
A Kubernetes or load-balancer probe that sends GET /live and then resets before reading the body currently stops serve_health_http / serve_postgres_health_http. The next probe has no listener loop.
Split accept from connection I/O. Retry Interrupted, ConnectionAborted, and ConnectionReset on accept. Continue after any post-accept read/write error. Keep stopping on accept WouldBlock so the existing non-blocking shutdown contract still holds.
|
Closing as a proven superseded predecessor. Fresh compare shows #111 head |
Keep #91 as the predecessor translator head. Operators should land the serve-loop successor so a load balancer can keep probing. Co-authored-by: Seongho Bae <seonghobae@users.noreply.github.com>
Keep #91 as the predecessor translator head. Operators should land the serve-loop successor so a load balancer can keep probing. Co-authored-by: Seongho Bae <seonghobae@users.noreply.github.com>
Keep #91 as the predecessor translator head. Operators should land the serve-loop successor so a load balancer can keep probing. Co-authored-by: Seongho Bae <seonghobae@users.noreply.github.com>


Why
PR #91 at
79aa953exposes GET/liveand GET/readyand can accept one connection. Operators still cannot keep a probe server running. A load balancer that probes every few seconds needs a blocking accept loop.TDD
RED required
serve_health_httpto answer two successive probes and then return when accept fails. GREEN adds that loop plusserve_postgres_health_http. Interrupted accepts retry; any other accept, read, or write error stops the loop so a closed or non-blocking listener does not spin.Scope
serve_health_http/serve_health_http_withkeep accepting until accept fails.serve_postgres_health_httpreuses the same loop. GET/livestill has no store I/O. GET/readystill observes after accept. Bare/readystill requirespostgres_operational_store.Out of scope
79aa953so operators get the serve loopTest plan
cargo test --lib health --test health_http_listener_contract --test health_http_probe_contractcargo test --test documentation_architecture_contract --test traceability_active_pr_contractcargo clippy --all-targets -- -D warningscargo test --test postgres_health_http_contract(needsTEST_DATABASE_URL)Do not merge until exact-head checks and independent last-push approval are satisfied. Never self-approve.
Operator next action
Call
serve_health_httporserve_postgres_health_httpafterbind_health_http. Point liveness at GET/liveand readiness at GET/ready. Do not treat/liveHTTP 200 as write-readiness. Close or retarget #91 and draft #102 after this lands so the same probe stack does not race.