Skip to content

[PROJ-1851] Preserve Jetstream progress under firehose load - #358

Closed
AndrewNordstrom wants to merge 3 commits into
mainfrom
dev/PROJ-1851-jetstream-backpressure
Closed

AndrewNordstrom wants to merge 3 commits into
mainfrom
dev/PROJ-1851-jetstream-backpressure

Conversation

@AndrewNordstrom

Copy link
Copy Markdown
Collaborator

Summary

  • replace Jetstream queue-overflow reconnect loops with bounded WebSocket pause/resume backpressure
  • expose cursor lag, queue depth, pause/resume, overload reconnect, and cumulative drop telemetry on operator health surfaces
  • degrade Jetstream health when event or cursor freshness reaches five minutes without coupling catch-up to readiness restarts
  • harden reconnect ownership so stale sockets, stale timers, and reconnect-time cursor read failures cannot create duplicate connections or skip safely completed progress

Incident evidence

  • production logged 1,828 Jetstream ingestion queue saturated events after 2026-07-12 07:57 EDT
  • the queue repeatedly reached 10,000 pending events with 20 active handlers
  • cursor lag was approximately 35,700 seconds and grew during a 30-second sample despite healthy process readiness
  • this change does not jump the cursor, discard the backlog, or alter production feed/governance contracts

Verification

  • focused Jetstream/health suite: 6 files, 55 tests passed
  • npm run verify: 153 files, 1,700 tests passed; root/CLI/SDK/legacy web/web-next builds passed
  • npm run docs:verify: passed
  • cd web-next && npm run lint: passed
  • cd web-next && npx tsc --noEmit: passed
  • git diff --check: passed
  • two completed local CodeRabbit passes produced 10 findings, all addressed; subsequent exact-diff retries stalled without a result, so hosted exact-head review is required

Rollout boundary

  • no environment, schema, systemd, production key, or feed URI change
  • production success requires zero new hard-limit drops/overload reconnects, increasing pause/resume counters under load, declining cursor lag, healthy readiness, and an unchanged non-empty public feed
  • rollback is the prior application SHA; no cursor jump or backlog deletion is authorized

Linear: PROJ-1851

@coderabbitai

coderabbitai Bot commented Jul 14, 2026 •

Copy link
Copy Markdown

Review Change Stack

Warning

Review limit reached

You’ve reached a temporary PR review limit under our Fair Usage Limits Policy.

Your recent review volume is higher than typical usage, so adaptive limits are currently applied.

Next review available in: 18 minutes

Your organization has reached its usage spending cap. Adjust your spending cap in the billing tab.

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Repository YAML (base), Organization UI (inherited)

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 2a7f7169-b70b-4952-9e70-c6c8d3ce5c46

📥 Commits

Reviewing files that changed from the base of the PR and between 8bd9aa8 and fc44c95.

📒 Files selected for processing (16)
  • docs/dev-journal.md
  • src/admin/routes/feed-health.ts
  • src/admin/routes/health.ts
  • src/admin/routes/vitals.ts
  • src/index.ts
  • src/ingestion/handlers/post-handler.ts
  • src/ingestion/jetstream-health.ts
  • src/ingestion/jetstream.ts
  • src/lib/health.ts
  • tests/feed-health-rescore.test.ts
  • tests/health-redaction.test.ts
  • tests/jetstream-backpressure.test.ts
  • tests/jetstream-lifecycle.test.ts
  • tests/jetstream-message-processing.test.ts
  • tests/post-handler-filtering.test.ts
  • tests/queue-saturation-metrics.test.ts

Walkthrough

Jetstream ingestion now applies bounded WebSocket backpressure, preserves safe cursor state across reconnects, reports runtime telemetry, and evaluates freshness using cursor lag and recent event activity. Admin health, feed-health, vitals, tests, and the development journal reflect these changes.

Changes

Jetstream ingestion reliability

Layer / File(s) Summary
Bounded ingestion and reconnect lifecycle
src/ingestion/jetstream.ts, src/ingestion/jetstream-health.ts, tests/jetstream-*.test.ts, tests/queue-saturation-metrics.test.ts
Inbound delivery pauses at the queue high-water mark and resumes after draining; reconnect timers, stale sockets, cursor fallback, and cumulative drop metrics are covered by tests.
Freshness health calculation and registration
src/lib/health.ts, src/index.ts, tests/health-redaction.test.ts
Jetstream health now includes runtime metrics and reports unhealthy status for disconnected sockets, stale events, or cursor lag at least five minutes.
Operator health and telemetry surfaces
src/admin/routes/feed-health.ts, src/admin/routes/health.ts, src/admin/routes/vitals.ts, tests/feed-health-rescore.test.ts, tests/health-redaction.test.ts, docs/dev-journal.md
Feed-health, health, and vitals expose cursor, queue, pause/resume, reconnect, and drop metrics; route assertions and the journal document the updated behavior.

Estimated code review effort: 4 (Complex) | ~45 minutes

Sequence Diagram(s)

sequenceDiagram
  participant JetstreamWebSocket
  participant JetstreamQueue
  participant processEvent
  participant HealthSurface
  JetstreamWebSocket->>JetstreamQueue: deliver events
  JetstreamQueue->>processEvent: process bounded work
  processEvent-->>JetstreamQueue: release completed slots
  JetstreamQueue->>JetstreamWebSocket: pause or resume delivery
  HealthSurface->>JetstreamQueue: read runtime state
  HealthSurface-->>HealthSurface: evaluate cursor and event freshness
Loading

Possibly related issues

  • PROJ-1851: Implements the issue’s backpressure, cursor-safety, telemetry, and freshness-health objectives.

Suggested labels: documentation, javascript

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 55.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title is concise and accurately summarizes the Jetstream backpressure/progress-preservation change.
Description check ✅ Passed The description clearly matches the Jetstream backpressure, telemetry, and health changes in this PR.
Linked Issues check ✅ Passed The implementation aligns with PROJ-1851: bounded pause/resume backpressure, cursor safety hardening, telemetry, and degraded health are all present.
Out of Scope Changes check ✅ Passed The changes stay focused on Jetstream ingestion, health surfaces, tests, and docs, with no obvious unrelated additions.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch dev/PROJ-1851-jetstream-backpressure
✨ Simplify code
  • Create PR with simplified code
  • Commit simplified code in branch dev/PROJ-1851-jetstream-backpressure

Warning

Review ran into problems

🔥 Problems

These MCP integrations need to be re-authenticated in the Integrations settings: Notion


Comment @coderabbitai help to get the list of available commands.

@AndrewNordstrom
AndrewNordstrom marked this pull request as ready for review July 14, 2026 00:48
@coderabbitai coderabbitai Bot added documentation Improvements or additions to documentation javascript Pull requests that update javascript code labels Jul 14, 2026
@AndrewNordstrom
AndrewNordstrom marked this pull request as draft July 14, 2026 01:03

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 9

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (2)
src/ingestion/jetstream.ts (2)

1263-1283: 🩺 Stability & Availability | 🟡 Minor | ⚡ Quick win

Reset all reconnect lifecycle state in the test helper.

reset() clears the new timer and counters but leaves isShuttingDown, useFallback, reconnect/failure counters, active sockets, and metrics intervals unchanged. Reconnect/fallback tests can therefore become order-dependent or leak work into later tests. Reset or explicitly close/clear every lifecycle resource, then test reset after a fallback and pending reconnect.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/ingestion/jetstream.ts` around lines 1263 - 1283, Update the test
helper’s reset() method to restore all reconnect lifecycle state, including
isShuttingDown, useFallback, reconnect/failure counters, active sockets, and
metrics intervals, in addition to the existing fields. Explicitly close or clear
any remaining lifecycle resources and reset related state so fallback and
pending-reconnect tests are isolated and do not leak work into later tests.

964-975: 🗄️ Data Integrity & Integration | 🟠 Major | 🏗️ Heavy lift

Reject stale Jetstream messages before processEvent. In src/ingestion/jetstream.ts:964, a buffered message from the previous socket can still reach processJetstreamMessageData after reconnect, so stale commit events still hit processEvent and its side effects; the generation checks only protect cursor/pin bookkeeping. Add a reconnect regression test that delivers an old-socket buffered message after the new connection is active.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/ingestion/jetstream.ts` around lines 964 - 975, Update the JetStream
message handling around the socket.on('message') callback and
processJetstreamMessageData to validate sessionGeneration and socket ownership
before invoking processEvent, dropping buffered messages from prior connections.
Preserve processing for messages from the active connection and keep existing
cursor/pin bookkeeping checks. Add a reconnect regression test that delivers a
buffered old-socket message after the new connection is active and verifies it
produces no processEvent side effects.

Source: Path instructions

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@src/ingestion/jetstream.ts`:
- Around line 245-257: Update stopJetstream() to mark shutdown, close/detach
inbound delivery, and wait until activeEventCount reaches zero before persisting
lastCursorUs; retain draining of queued waiters as appropriate. Guard
resumeInboundIfReady() (and any releaseSlot-triggered path) with the shutdown
state so teardown cannot reopen ingestion or accept new events. Add a regression
test using a slow handler that finishes after stopJetstream() begins and verify
the persisted cursor includes its completed work.
- Around line 37-41: Increase the reserved headroom in the threshold calculation
near PAUSE_QUEUE_THRESHOLD so pausing occurs early enough to account for frames
already buffered after ws.pause(). Ensure eventQueue remains below
MAX_PENDING_EVENTS and prevents handleQueueOverload() and dropped work during
sustained input; add a soak test for continuous traffic while paused that
asserts no dropped events or overload reconnects.

In `@src/lib/health.ts`:
- Around line 81-85: Bound the startup freshness logic in the health calculation
around lastEventAgeMs, eventIsFresh, and cursorIsFresh by tracking ingestion
start or connection time. When lastEventAt or cursorLagMs is absent, use that
start time as the fallback age so a connected instance with no event or cursor
becomes unhealthy at the five-minute JETSTREAM_FRESHNESS_LIMIT_MS boundary;
retain observed event and cursor ages once available, and add tests for the
exact boundary.

In `@tests/feed-health-rescore.test.ts`:
- Around line 253-269: Update the Fastify test around registerFeedHealthRoutes
and app.inject so the request and all assertions execute inside a try block,
with await app.close() in a finally block. Ensure the Fastify instance is closed
whether the injection or assertions succeed or fail.

In `@tests/health-redaction.test.ts`:
- Around line 80-113: Extend the calculateJetstreamHealth tests to cover the
just-below five-minute boundary at 299,999 ms for both event age and cursor lag.
Add assertions using new Date(nowMs - 299_999) and cursorLagMs: 299_999,
verifying status is healthy while preserving the existing exact-boundary
unhealthy cases.
- Around line 239-251: Extend the expected telemetry assertions in
tests/health-redaction.test.ts lines 239-251 to include cursor_us,
active_events, pause_count, resume_count, and overload_reconnect_count alongside
the existing JetStream fields. Also update tests/feed-health-rescore.test.ts
lines 192-194 to assert cursorUs, activeEvents, pauseCount, resumeCount,
overloadReconnectCount, and totalDroppedEvents, using the fixture values and
preserving the existing response contract checks.

In `@tests/jetstream-backpressure.test.ts`:
- Around line 65-153: Add isolated tests in the Jetstream backpressure suite for
exceptions from the flow-control socket’s pause and resume methods. Use throwing
spies, invoke the relevant backpressure transition, and assert overload
recovery/reconnect accounting for pause failures plus detachment, runtime pause
state, and close(1011, 'backpressure_resume_failed') behavior for resume
failures. Ensure each test resets and cleans up queue state and verifies
counters against the implementation’s actual behavior.

In `@tests/jetstream-lifecycle.test.ts`:
- Around line 33-40: Update the reconnect test to invoke MockWebSocket.emitClose
with code 1006 for the abnormal peer/network closure path instead of close. In
MockWebSocket.close, validate the supplied close code and reject invalid codes
such as 1006 before calling wsCloseMock or emitClose, matching real ws behavior;
preserve valid close handling.

In `@tests/queue-saturation-metrics.test.ts`:
- Around line 72-88: Add boundary-focused cases alongside the existing
cursor-lag test for __testJetstreamQueue: verify no cursor returns null cursorUs
and cursorLagMs, a cursor equal to the fixed current time reports zero lag, and
a future cursor follows the intended non-negative/error behavior without
exposing negative lag. Use fake timers with fixed system time and reset the
queue and timers for each case.

---

Outside diff comments:
In `@src/ingestion/jetstream.ts`:
- Around line 1263-1283: Update the test helper’s reset() method to restore all
reconnect lifecycle state, including isShuttingDown, useFallback,
reconnect/failure counters, active sockets, and metrics intervals, in addition
to the existing fields. Explicitly close or clear any remaining lifecycle
resources and reset related state so fallback and pending-reconnect tests are
isolated and do not leak work into later tests.
- Around line 964-975: Update the JetStream message handling around the
socket.on('message') callback and processJetstreamMessageData to validate
sessionGeneration and socket ownership before invoking processEvent, dropping
buffered messages from prior connections. Preserve processing for messages from
the active connection and keep existing cursor/pin bookkeeping checks. Add a
reconnect regression test that delivers a buffered old-socket message after the
new connection is active and verifies it produces no processEvent side effects.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository YAML (base), Organization UI (inherited)

Review profile: ASSERTIVE

Plan: Pro Plus

Run ID: 91dfbf08-5a16-4fda-a4a2-8a01a4b78603

📥 Commits

Reviewing files that changed from the base of the PR and between 64ae6f8 and 8bd9aa8.

📒 Files selected for processing (13)
  • docs/dev-journal.md
  • src/admin/routes/feed-health.ts
  • src/admin/routes/health.ts
  • src/admin/routes/vitals.ts
  • src/index.ts
  • src/ingestion/jetstream-health.ts
  • src/ingestion/jetstream.ts
  • src/lib/health.ts
  • tests/feed-health-rescore.test.ts
  • tests/health-redaction.test.ts
  • tests/jetstream-backpressure.test.ts
  • tests/jetstream-lifecycle.test.ts
  • tests/queue-saturation-metrics.test.ts

Comment thread src/ingestion/jetstream.ts
Comment thread src/ingestion/jetstream.ts
Comment thread src/lib/health.ts
Comment thread tests/feed-health-rescore.test.ts
Comment thread tests/health-redaction.test.ts
Comment thread tests/health-redaction.test.ts
Comment thread tests/jetstream-backpressure.test.ts
Comment thread tests/jetstream-lifecycle.test.ts
Comment thread tests/queue-saturation-metrics.test.ts
@AndrewNordstrom
AndrewNordstrom marked this pull request as ready for review July 14, 2026 01:30
@AndrewNordstrom
AndrewNordstrom marked this pull request as draft July 14, 2026 01:31
@AndrewNordstrom

Copy link
Copy Markdown
Collaborator Author

Queue control: returning this PR to draft while #360 remains the sole PI-review merge lane.

Current head 7cf05e021ed659cf268ecf9cfdff852880365a78 adds a lifecycle-hardening follow-up and appears to address the prior review findings, but it is still behind current main and has no completed CodeRabbit review on this exact head (the automatic attempt is rate-limited). Do not promote or merge until #360 is reviewed, deployed, and production-smoked; then rebase once and re-run the full incident evidence packet.

@AndrewNordstrom

Copy link
Copy Markdown
Collaborator Author

Exact-head preparation receipt for f248bc6:\n\n- local exact worktree: 153 files / 1,716 tests\n- npm run build, npm run verify, npm run docs:verify, web-next lint/typecheck/build: pass\n- GitHub backend/frontend/docs/report/quality/security/CodeQL gates: pass\n- production-shaped 5,000-event socket burst: 5,053.69 events/sec, 26 pause/resume cycles, zero drops/reconnects/cursor mismatch\n\nThe prior CodeRabbit findings are covered by the current code and regression tests. Automatic review skipped because this PR intentionally remains draft while #360 is the sole PI-review merge lane. After #360 deploys and smokes cleanly, this branch still needs one current-main integration, exact-head hosted review, and production recovery proof before merge.

@AndrewNordstrom
AndrewNordstrom force-pushed the dev/PROJ-1851-jetstream-backpressure branch from f248bc6 to fc44c95 Compare July 14, 2026 06:21
@AndrewNordstrom
AndrewNordstrom marked this pull request as ready for review July 14, 2026 06:23
@github-actions

Copy link
Copy Markdown

This PR has been inactive for 14 days. It will close in 7 days if there is no update.

@github-actions github-actions Bot added the stale label Aug 24, 2026
@github-actions github-actions Bot closed this Aug 31, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation javascript Pull requests that update javascript code stale

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant