Skip to content

feat(health): serve operator probes until accept fails - #111

Closed
cursor[bot] wants to merge 13 commits into
mainfrom
cursor/bc-0ec5ca15-dfa1-48eb-b2db-416fb6ef72ef-6846
Closed

cursor[bot] wants to merge 13 commits into
mainfrom
cursor/bc-0ec5ca15-dfa1-48eb-b2db-416fb6ef72ef-6846

Conversation

@cursor

@cursor cursor Bot commented Aug 16, 2026

Copy link
Copy Markdown
Contributor

Why

PR #91 at 79aa953 exposes GET /live and GET /ready and can accept one connection. Operators still cannot keep a probe server running. A load balancer that probes every few seconds needs a blocking accept loop.

TDD

RED required serve_health_http to answer two successive probes and then return when accept fails. GREEN adds that loop plus serve_postgres_health_http. Interrupted accepts retry; any other accept, read, or write error stops the loop so a closed or non-blocking listener does not spin.

Scope

  • serve_health_http / serve_health_http_with keep accepting until accept fails.
  • serve_postgres_health_http reuses the same loop. GET /live still has no store I/O. GET /ready still observes after accept. Bare /ready still requires postgres_operational_store.
  • TRACEABILITY, OPERABILITY, ADR-0014, and CHANGELOG stay Active PR work, not protected-main truth.

Out of scope

Test plan

  • cargo test --lib health --test health_http_listener_contract --test health_http_probe_contract
  • cargo test --test documentation_architecture_contract --test traceability_active_pr_contract
  • cargo clippy --all-targets -- -D warnings
  • cargo test --test postgres_health_http_contract (needs TEST_DATABASE_URL)

Do not merge until exact-head checks and independent last-push approval are satisfied. Never self-approve.

Operator next action

Call serve_health_http or serve_postgres_health_http after bind_health_http. Point liveness at GET /live and readiness at GET /ready. Do not treat /live HTTP 200 as write-readiness. Close or retarget #91 and draft #102 after this lands so the same probe stack does not race.

Open in Web View Automation 

seonghobae and others added 12 commits August 16, 2026 19:51
RED: the domain health snapshot is not yet exposed as GET /live and
GET /ready with fail-closed readiness and a safe as-built OpenAPI 3.2
contract.
Map RuntimeHealthSnapshot to GET /live and GET /ready with fail-closed
readiness, RFC 9457 problem details for unsupported requests, and an
as-built OpenAPI 3.2.0 document that lists only those operations.
Exercise HealthHttpResponse::content_type and GET /ready?capability= from
the library test module so Linux line coverage can reach the remaining
unexecuted instantiations.
* feat(health): bind a TCP listener for operator probes

GET /live and GET /ready already exist as a request translator. Bind a
blocking listener that serves one HTTP/1.1 request per accepted
connection without adding public product routes, TLS, or SLO values.

* docs(traceability): name Active PR #92 on the bound health listener

* feat(health): answer probes from a PostgreSQL operational snapshot

* test(health): require a composed PostgreSQL operational snapshot

RED: runtime and relation probes exist, but callers still assemble
liveness, backlog, integrity, and the postgres capability by hand.

* feat(health): compose PostgreSQL probes into one operational snapshot

Map runtime and relation probes plus caller-supplied backlog into a
RuntimeHealthSnapshot. Probe failure becomes unknown/unready evidence
and never returns a driver error.

* feat(health): answer probes from a PostgreSQL operational snapshot

Rebuild #95 on the current #92 listener plus the #93 snapshot composer.
handle_postgres_health_http_request and accept_one_postgres_health_http
observe the caller-owned store and reuse the existing probe translator.

* docs(health): name #95 snapshot wiring on the current listener stack

Restore the architecture mappings the rewrite dropped: TRACEABILITY
module and capability rows, OPERABILITY fail-closed snapshot readiness,
ADR-0014 as-built status, and the composed-snapshot changelog line.

Co-authored-by: Seongho Bae <seonghobae@users.noreply.github.com>

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Seongho Bae <seonghobae@users.noreply.github.com>

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Seongho Bae <seonghobae@users.noreply.github.com>
Keep Active PR #91 as the remaining health HTTP stack after #76
landed on main. Drop the stale #76 active-work entry.

Co-authored-by: Seongho Bae <seonghobae@users.noreply.github.com>
The inherited #72 recovery fixture inserted a processing consumption
row without claim_deadline_at. Migration 0019 requires that column for
processing rows, and the deadline trigger is UPDATE-only, so exact-head
CI failed closed. Seed a valid persisted claim and assert the deadline
survives COPY restore.

Co-authored-by: Seongho Bae <seonghobae@users.noreply.github.com>
#92, #93, and #95 were absorbed into this branch. Keep TRACEABILITY and
ADR-0014 as-built status pointing at the remaining health HTTP PR so
closed stacked numbers are not treated as independent Active work.

Co-authored-by: Seongho Bae <seonghobae@users.noreply.github.com>
…port

RED: GET /live must not observe PostgreSQL, bare GET /ready must fail
closed on a read-only store, incomplete or oversized requests must
return 400 without echoing input, and the as-built OpenAPI /live
operation must document HTTP 503.

Co-authored-by: Seongho Bae <seonghobae@users.noreply.github.com>
Observe PostgreSQL only for GET /ready, after accept. Bare /ready
requires postgres_operational_store. The listener applies a 2-second
I/O timeout, rejects incomplete or oversized requests without echo,
and adds Cache-Control: no-store plus Allow: GET on 405.

Co-authored-by: Seongho Bae <seonghobae@users.noreply.github.com>
Operators can now run a blocking accept loop so a load balancer can keep
asking GET /live and GET /ready. Interrupted accepts retry; any other
accept, read, or write error stops the loop. TLS and keep-alive stay out
of this slice.

Co-authored-by: Seongho Bae <seonghobae@users.noreply.github.com>
@cursor
cursor Bot requested a review from seonghobae August 16, 2026 15:23
Keep #91 as the predecessor translator head. Operators should land the
serve-loop successor so a load balancer can keep probing.

Co-authored-by: Seongho Bae <seonghobae@users.noreply.github.com>

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do not merge #111 as-is. The serve loop stops on per-connection read/write errors, so one dropped load-balancer probe can take down GET /live and GET /ready for every later check.

Keep #91/#102 parked. Land a successor that keeps accepting after connection I/O errors and stops only when accept itself fails (Interrupted/ConnectionAborted/ConnectionReset retry; WouldBlock and other accept failures still stop so a closed or non-blocking listener does not spin).

Independent last-push approval is still required. Do not self-approve.

Open in Web View Automation 

Sent by Cursor Automation: Fix Issues

Comment thread src/health_http.rs
Comment on lines +208 to +212
fn classify_serve_accept(result: io::Result<()>) -> ServeAcceptProgress {
match result {
Ok(()) => ServeAcceptProgress::Continue,
Err(error) if error.kind() == io::ErrorKind::Interrupted => ServeAcceptProgress::Continue,
Err(error) => ServeAcceptProgress::Stop(error),

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

classify_serve_accept treats every non-Interrupted accept_one_health_http_with error as fatal, including BrokenPipe, ConnectionReset, and write timeout after a successful accept.

A Kubernetes or load-balancer probe that sends GET /live and then resets before reading the body currently stops serve_health_http / serve_postgres_health_http. The next probe has no listener loop.

Split accept from connection I/O. Retry Interrupted, ConnectionAborted, and ConnectionReset on accept. Continue after any post-accept read/write error. Keep stopping on accept WouldBlock so the existing non-blocking shutdown contract still holds.

Copy link
Copy Markdown
Contributor

Closing as a proven superseded predecessor. Fresh compare shows #111 head 9080ba17 is the merge base/ancestor of #132 head e7440491; #132 carries the continuous serve loop through dropped-connection/reset coverage and adds the deployable fail-closed health process entrypoint. Continue exact-head CI/review on #132; do not merge #111 separately.

@seonghobae seonghobae closed this Aug 16, 2026
cursor Bot pushed a commit that referenced this pull request Aug 17, 2026
Keep #91 as the predecessor translator head. Operators should land the
serve-loop successor so a load balancer can keep probing.

Co-authored-by: Seongho Bae <seonghobae@users.noreply.github.com>
cursor Bot pushed a commit that referenced this pull request Aug 18, 2026
Keep #91 as the predecessor translator head. Operators should land the
serve-loop successor so a load balancer can keep probing.

Co-authored-by: Seongho Bae <seonghobae@users.noreply.github.com>
cursor Bot pushed a commit that referenced this pull request Aug 18, 2026
Keep #91 as the predecessor translator head. Operators should land the
serve-loop successor so a load balancer can keep probing.

Co-authored-by: Seongho Bae <seonghobae@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants